fix(golang): bound allocations from UPX-packed binary headers (#5195)

* fix(golang): bound UPX decompression by what the input can justify

A UPX block's `sz_unc` was passed straight to `make([]byte, n)`, so four bytes of
attacker input spanned the full uint32 range. A ~300KB file could drive multi-GB
resident memory, which Go reports as a fatal runtime OOM that no `recover` on this
path can contain. Syft parses binaries out of arbitrary images, so that is reachable
from any scan.

Two bounds do the work, and UPX's own invariants make them exact:

- all blocks together reconstruct `p_filesize`, so a running remainder caps the sum.
  Bounding only per-block would leave the total at (block count x original size), so
  the remainder is the part that actually closes it
- `p_filesize` sizes the output buffer, so it is capped both absolutely and against
  the size of the file on disk. The absolute cap alone left a ~60 byte header able to
  claim 500MB; the ratio alone cannot work either, since LZMA encodes a run of N
  equal bytes in O(log N) and a `go:embed` of 120MB of zeros legitimately packs 209x

The output buffer is also allocated on the first block that survives validation
rather than up front, so a header claiming a large size with nothing decodable behind
it costs nothing.

Alongside that, several ways the block loop could be steered off its own buffer:

- `parseELFPTLoadOffsets` bounds-checked program headers with `phStart+phentsize >
  len(buf)`, which overflows for a `p_offset` near 2^64 and lets an out-of-range entry
  through into a slice index. Switched to the subtraction form, and a `phentsize`
  shorter than an ELF64 program header is rejected since the reads use fixed offsets
- block placement had the same overflow shape, and on failure it skipped the copy and
  fell through. `outputOffset` derives from the rejected offset, so a `p_offset` at the
  uint64 ceiling wrapped it to 0 and the next block landed on the reconstructed ELF
  header, still returning success. Out-of-range placement now ends the block chain,
  keeping the blocks already placed since those often carry `.go.buildinfo`
- the UPX 2-byte LZMA header carries `lc`/`lp` as nibbles and `pb` as three bits, so
  they can hold values LZMA does not permit. Folded into the props byte they wrap
  (`lc=15, lp=15, pb=7` gives 465, which truncates to 209) and the stream mis-decodes
  instead of failing
- blocks decode directly into a slice of the output buffer instead of a per-block
  buffer that is then copied, so one block no longer doubles peak memory
- the block count is capped: the size budget alone still allows millions of 12-byte
  blocks, each spinning up an LZMA reader

Verified against the `image-small-upx` fixture, a real `upx --best --lzma` binary:
four blocks, max `sz_unc` equal to `p_blocksize` exactly, blocks summing to 0.68 of
`p_filesize`, and a 1.96x expansion against the packed file, so every bound holds with
headroom on genuine output.

Note `p_blocksize` is deliberately not used as a ceiling on `sz_unc`. It adds nothing
the remainder does not already cover, and real output sits at exactly `p_blocksize`,
so it would run with zero headroom against a value the format does not promise.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* fix(golang): report UPX binaries we cannot unpack as unknowns

A packed Go binary we fail to unpack means its packages are silently missing from the
SBOM. `getBuildInfo` checked the decompression error for nil and discarded it, so a
rejected file surfaced only the original `buildinfo.Read` error, and that one is
deliberately silenced because it is usually just "not a Go binary".

`decompressUPX` now marks the cases worth reporting with `errUPXDecompress`: it got
past the header and the method dispatch and still could not unpack the file. Everything
else stays quiet, which matters more than it sounds. `upx` defaults to NRV2B unless
`--lzma` is passed and only LZMA is implemented here, so treating any decompression
failure as reportable would attach a golang-cataloger unknown to every packed non-Go
binary in an image. Making the reportable case opt in rather than the quiet case a
growing list also means a guard added later is silent by default.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* refactor(golang): drop the vendored xcoff parser

The only thing the golang cataloger read from it was `TargetMachine`, which is the same
two magic bytes `getGOARCHFromBin` had already matched on to dispatch there. So a full
XCOFF walk (string table, symbol table, every relocation table) ran to recover a value
that was already in hand.

`getGOARCHFromBin` reads those two bytes directly now. That is also strictly more
correct: `xcoff.NewFile` bailed with "no symbol table" when `symptr == 0`, so a stripped
XCOFF binary reported no arch at all.

Fixtures move to `testdata/xcoff/`, since they are no longer a package's own testdata.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* feat(internal/file): report a reader's real size behind one helper

Bounds get written as "does this claim exceed what the reader actually holds", and the
subtlety is that a reader's own answer cannot be trusted: an `io.SectionReader` reports
the nominal length it was built with, so one constructed with `1<<63-1` will happily
claim to hold 8EB. Confirming the last byte is readable is what separates a real size
from a nominal one.

Returns `(int64, bool)` rather than a bare size, since there are five ways to not know
and collapsing them to 0 makes every caller reinvent "0 means stop bounding".

The doc also names the wrapper trap: a type embedding `io.ReaderAt` as an interface
promotes only `ReadAt`, so it answers no size at all however large the reader beneath it
is, and a bound written against it silently does nothing.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* fix(golang): bound every section a reader can expand, not just the name table

`CheckSectionNameTable` bounds the one section `elf.NewFile` expands while parsing, which
is enough for a caller that only parses. It is not enough for one that goes on to read
sections: those are expanded lazily, so a 260KB ELF declaring a compressed `.symtab`
drove 1.3GB of allocation through `goversion`, which opens the file with `debug/elf`
inside its own package and cannot be routed through `elfutil.NewFile`.

`CheckAllSections` is that gate, and is deliberately named as the superset so the
relationship to the narrow one is structural rather than documented.

Also exports `ErrDeclaredSizeExceeded`. A refusal here costs the SBOM a package, so
callers report it as an unknown, and matching on an error string to decide that is not
something to build a reporting policy on.

A ruleguard rule now flags `buildinfo.Read` and `version.ReadExeFromReader` outside
`scan_binary.go`. Both open `debug/elf` internally and are only bounded there by the
wrappers that gate the reader first, which the existing rule on `elf.NewFile` could not
see through.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* feat(internal/spillbuf): sparse offset-addressed buffer that spills to disk

For output that arrives out of order, at offsets the input itself declares: a
decompressor placing extents, an archive rebuilding a file from chunks. A plain `[]byte`
cannot do that without either pre-sizing to a length the input claims or growing to the
furthest offset it names, and both hand a hostile input an allocation knob.

The load-bearing property is that it reports only what it actually stored. `Size` is the
contiguous run written from offset zero and a read past it is `io.EOF`, never the zeros
an unwritten region would hand back for free. Sparse storage serving holes as real
content is what lets a write offset stand in for output nothing produced, so anything
sizing an allocation against this reader (`saferio` in `debug/elf` does exactly that)
stays bounded by work actually done.

`FirstGap` answers where the next write fits, so callers filling holes do not need the
extent bookkeeping and cannot re-derive that rule wrongly, and it only reports gaps
lying entirely inside `[0, within)`. The extent type is unexported for the same reason.

The rest of the contract worth stating: reads and writes after `Close` fail with
`os.ErrClosed` rather than looking like an empty buffer, and a write that fails leaves
the buffer unchanged.

Memory is bounded by the limit rather than by how much is written: filling 1MB and
filling 64MB cost about the same. What lands on disk past that limit is the caller's
bound to set, not this package's. Growth is geometric and capped, because sizing the
tier to exactly what each write needs is quadratic once a caller streams through it in
small chunks.

The golang cataloger is the only consumer today, and nothing else in the tree has this
shape.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* fix(golang): rebuild UPX binaries through spillbuf and read the reconstruction everywhere

Two amplification paths were still reachable. A 128 byte input allocated 16MB by
trusting `sz_cpr` before weighing it against the bytes the file actually holds, and the
reconstruction was assembled in a `[]byte` sized from the header's declared original
size, so a 301KB input allocated 300MB.

The reconstruction is a `spillbuf.Buffer` now, which is what makes the second one go
away: blocks are placed at the offsets the file declares, nothing is pre-sized from a
claim, and the reader's length is the contiguous run actually rebuilt. That last part is
the invariant everything downstream leans on. A reader reporting the declared size would
serve the holes between placed blocks as zeros it never stored, and a 64 byte block
parked at `p_filesize-64` then made a 300KB input hand out a 64MB reader: measured, that
drove 269MB through `getBuildInfo` and returned nothing, silently.

Placement and storage are separate concerns. `upx.go` places blocks at the offsets the
file declares and asks the buffer how much was rebuilt and where the next gap is; it
knows nothing about temp files or memory limits.

The remaining bounds are ratios against the input wherever they can be, since an
absolute ceiling only ever drops a large legitimate binary. The one absolute left is on
`p_filesize`, because padding an input is nearly free, and the LZMA dictionary is paid
for in its own block's `sz_cpr`.

Interface change: `scanFile` takes a context and hands back the reconstruction for each
binary that turned out to be packed, and the caller owns it. Everything after the build
info read (crypto settings, arch, symbols, the version scan) reads it, so a packed
binary no longer reports its packages with no symbols at all.

Also narrows what gets reported. This cataloger runs on every executable in an image, so
reporting every `getBuildInfo` error attached an unknown to every corrupt ELF, odd PE and
non-Go Mach-O slice in it. Now only four gaps, the ones that cost the SBOM something
syft could have had: a packed file we could not decode (`errUPXDecompress`), a header
the size bounds refused (`errUPXSizeRefused`), a reconstruction that came up short
(`errUPXPartial`), and an ELF expansion `elfutil` declined
(`ErrDeclaredSizeExceeded`). An unimplemented UPX method stays quiet on purpose: `upx`
defaults to NRV2B and only LZMA is implemented here, so reporting it would attach an
unknown to most packed non-Go binaries in an image.

That filtering is a new pattern. No other cataloger picks which of its errors become
unknowns, and the reason this one does is that it runs on every executable in an image
rather than on files a glob already selected.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* chore: ignore local spec/ planning notes

`/specs` was already ignored; `/spec` is the same thing under the singular name and was
not, so working notes landed in commits.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* refactor(golang): name what the unpack hands back, not where it stores it

the scan's signatures said `*spillbuf.Buffer` all the way down, which told
every reader below the scan that it cares whether the reconstruction lives in
memory or on disk. it does not. two interfaces instead, each naming only what
its side actually uses:

- `unpackedContents` for the consumers: `ReaderAt`, `Closer`, and a `Size` that
  is the contiguous run rebuilt from zero
- `blockSink` for the decoder, which additionally needs `WriterAt` and
  `FirstGap` to place blocks

this costs one thing worth flagging. `Close` on a nil `*spillbuf.Buffer` was
nil-safe, so callers just deferred it; widened into an interface, a nil pointer
becomes a non-nil interface and the same call panics. `closeUnpacked` absorbs
that, and it is now the only way an owner releases contents. `readerFor`,
`seekerFor` and `closeUnpacked` are the three places that resolve the nil.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* docs(elfutil): say that CheckAllSections is the decompression-bomb gate

the doc led with "bounds every section syft can drive debug/elf into
decompressing", which describes the mechanism without ever naming what it is
defending against. a reader landing on the call site could not tell that this
is the decompression-bomb check.

state the attack instead: a compressed section header declares its own
decompressed size, debug/elf believes that number and allocates it on open, and
nothing forces it to match what the compressed bytes actually yield. the 260KB
ELF that drove 1.3GB of allocation now lives in the doc rather than only at the
one call site that happened to mention it.

comments only.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* docs(golang): move the overflow note onto the bounds test it describes

the "subtraction form" note sat above the phStart assignment with ten lines of
wrap analysis between it and the `if` it was talking about, so the subtraction
it names reads as missing. it is `hdrLen-phStart < phentsize`.

move it onto that test and name the form it is avoiding (`phStart+phentsize >
hdrLen`, which can overflow past the buffer end) so the referent is not several
paragraphs away.

comments only.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

* chore(golang): drop the internal/ README left over from xcoff

that README existed to record where the vendored xcoff package came from and
that Go keeps it internal, which is a provenance note worth having while
third-party code sat there. dropping xcoff took the reason with it, and
repurposing it to point at gotestdata just made it redundant: gotestdata has
its own README covering why it is under internal/, why it is not testdata/, and
what belongs in each.

internal/ now holds one self-documenting directory.

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>

---------

Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>
This commit is contained in:
Alex Goodman
2026-09-10 09:21:19 -04:00
committed by GitHub
parent 0c049806bd
commit 1d25dfe4da
25 changed files with 4755 additions and 1556 deletions
+1
View File
@@ -20,6 +20,7 @@ bin/
/.tool
/.task
/generate
/spec
/specs
mise.toml
.make/.make
+66
View File
@@ -0,0 +1,66 @@
package file
import "io"
// ReaderSize reports the number of bytes behind r. The bool is false when that cannot be determined, and
// callers must branch on it rather than treating the zero as a size.
//
// Binary parsers read sizes out of the files they are parsing and allocate against them. Those sizes are
// 32- and 64-bit fields under the control of whoever produced the file, so a truncated or hostile input
// can declare gigabytes that are not there. The read would fail afterwards either way; what matters is
// refusing before the allocation, and that needs a real byte count to weigh the claim against.
//
// The bool is the whole point of the signature. An earlier version returned 0 for every failure mode, and
// a caller that forgot to special-case it got a bound of "nothing is allowed" or, more often, silently
// fell back to an unbounded path. Making the caller name the failure keeps that from happening by
// omission.
//
// Both *bytes.Reader and *io.SectionReader answer directly. A file handle only has Seek, so fall back to
// a save-and-restore seek, which leaves the caller's cursor where it found it. That fallback is not
// atomic: it moves the cursor and puts it back, so a reader being read concurrently can observe the
// intermediate position.
//
// Neither answer is taken on faith, because a reader can report a length it cannot deliver: an
// *io.SectionReader answers with the length it was constructed with, and debug/elf and friends construct
// theirs with a nominal one, so io.NewSectionReader(r, 0, 1<<63-1) would otherwise report a bound that
// permits everything as if it had been measured. The last byte is read back to confirm the count, which is
// one ReadAt whatever the size, and the read is what the caller was going to weigh anyway.
//
// A wrapper that embeds io.ReaderAt as an interface promotes only ReadAt, so it answers no size here
// however sizable the reader underneath it is. A type meant to be bounded needs to forward Size() itself.
func ReaderSize(r io.ReaderAt) (int64, bool) {
size, ok := reportedSize(r)
if !ok {
return 0, false
}
// ReadAt may return io.EOF alongside a full read, so the count is what says the byte is there
var last [1]byte
if n, _ := r.ReadAt(last[:], size-1); n != len(last) {
return 0, false
}
return size, true
}
// reportedSize is what the reader says about itself, before ReaderSize checks whether it can back it up.
func reportedSize(r io.ReaderAt) (int64, bool) {
if sr, ok := r.(interface{ Size() int64 }); ok {
size := sr.Size()
return size, size > 0
}
s, ok := r.(io.Seeker)
if !ok {
return 0, false
}
cur, err := s.Seek(0, io.SeekCurrent)
if err != nil {
return 0, false
}
size, err := s.Seek(0, io.SeekEnd)
if err != nil {
return 0, false
}
if _, err := s.Seek(cur, io.SeekStart); err != nil {
return 0, false
}
return size, size > 0
}
+140
View File
@@ -0,0 +1,140 @@
package file
import (
"bytes"
"io"
"os"
"path/filepath"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func TestReaderSize(t *testing.T) {
tests := []struct {
name string
r io.ReaderAt
want int64
wantOK bool
}{
{
name: "bytes.Reader answers directly",
r: bytes.NewReader(make([]byte, 1234)),
want: 1234,
wantOK: true,
},
{
// an empty reader is not a size to bound against, so it reports not-ok along with everything
// else that cannot be measured
name: "empty bytes.Reader",
r: bytes.NewReader(nil),
want: 0,
},
{
name: "SectionReader over a real length",
r: io.NewSectionReader(bytes.NewReader(make([]byte, 500)), 100, 300),
want: 300,
wantOK: true,
},
{
name: "seek-only reader falls back to seeking",
r: seekOnly{bytes.NewReader(make([]byte, 4096))},
want: 4096,
wantOK: true,
},
{
name: "reader with neither Size nor Seek cannot be measured",
r: readAtOnly{bytes.NewReader(make([]byte, 4096))},
want: 0,
},
{
// the shape debug/elf, debug/pe and debug/macho build internally. Taking this at its word hands
// back a bound that permits everything, reported as if it had been measured.
name: "SectionReader over a nominal length reports what it can deliver",
r: io.NewSectionReader(bytes.NewReader(make([]byte, 500)), 0, 1<<63-1),
want: 0,
},
{
// a Size() a type simply got wrong is the same failure: the bytes are what decides
name: "a size the reader cannot back up is refused",
r: overstatedSize{ReaderAt: bytes.NewReader(make([]byte, 10)), size: 1 << 30},
want: 0,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
size, ok := ReaderSize(tt.r)
assert.Equal(t, tt.want, size)
assert.Equal(t, tt.wantOK, ok)
})
}
}
func TestReaderSize_FileHandle(t *testing.T) {
path := filepath.Join(t.TempDir(), "f.bin")
require.NoError(t, os.WriteFile(path, make([]byte, 7777), 0o600))
f, err := os.Open(path)
require.NoError(t, err)
defer f.Close()
// the cursor has to come back where it was found: the caller may be mid-read, and an *os.File is the
// one input here that carries a shared position
_, err = f.Seek(101, io.SeekStart)
require.NoError(t, err)
size, ok := ReaderSize(f)
assert.True(t, ok)
assert.Equal(t, int64(7777), size)
at, err := f.Seek(0, io.SeekCurrent)
require.NoError(t, err)
assert.Equal(t, int64(101), at, "measuring the file must not move the caller's cursor")
}
// TestReaderSize_WrapperMustForwardSize pins the shape that made every bound expressed against this
// helper go silently inert. A type that embeds io.ReaderAt as an interface promotes only ReadAt, so it
// answers 0 here however sizable the reader underneath is, and a caller that treats 0 as "fall back to
// unbounded" then stops bounding anything. Forwarding Size() is what fixes it, and this is the assertion
// that a new wrapper has to satisfy.
func TestReaderSize_WrapperMustForwardSize(t *testing.T) {
inner := bytes.NewReader(make([]byte, 2048))
size, ok := ReaderSize(embedsReaderAt{inner})
assert.False(t, ok, "a wrapper that only embeds the interface cannot be measured")
assert.Equal(t, int64(0), size)
size, ok = ReaderSize(forwardsSize{inner})
assert.True(t, ok, "a wrapper meant to be bounded has to forward Size itself")
assert.Equal(t, int64(2048), size)
}
type seekOnly struct{ inner *bytes.Reader }
func (s seekOnly) ReadAt(p []byte, off int64) (int, error) { return s.inner.ReadAt(p, off) }
func (s seekOnly) Seek(off int64, whence int) (int64, error) {
return s.inner.Seek(off, whence)
}
type readAtOnly struct{ inner *bytes.Reader }
func (r readAtOnly) ReadAt(p []byte, off int64) (int, error) { return r.inner.ReadAt(p, off) }
type embedsReaderAt struct{ io.ReaderAt }
type forwardsSize struct{ io.ReaderAt }
func (f forwardsSize) Size() int64 {
size, _ := ReaderSize(f.ReaderAt)
return size
}
// overstatedSize answers with a size larger than the bytes behind it, the way a wrapper that returns a
// declared length rather than a measured one would.
type overstatedSize struct {
io.ReaderAt
size int64
}
func (o overstatedSize) Size() int64 { return o.size }
+330
View File
@@ -0,0 +1,330 @@
// Package spillbuf provides a sparse, offset-addressed byte buffer that keeps a bounded amount in memory
// and spills the rest to a temp file.
//
// It exists for the shape where output arrives out of order, at offsets the input itself declares: a
// decompressor placing extents, an archive rebuilding a file from chunks. A plain []byte cannot do that
// without either pre-sizing to a length the input claims or growing to the furthest offset it names, and
// both hand a hostile input an allocation knob.
//
// The buffer reports only what it actually stored. Size is the contiguous run written from offset zero,
// and a read past it is io.EOF rather than the zeros an unwritten region would otherwise hand back. That
// distinction is the whole point: sparse storage serves holes as zeros for free, so a reader that reported
// a length covering them would let a write offset stand in for real output. Anything sizing an allocation
// against this reader (debug/elf's saferio does exactly that) would then allocate bytes nothing produced.
//
// Allocation is bounded by the memory limit rather than by how much is written: filling 1MB and filling
// 64MB cost about the same, because everything past the limit goes to the file. Only memory is bounded
// here: how much ends up on disk is whatever the caller writes, so bounding that stays the caller's job.
package spillbuf
import (
"errors"
"fmt"
"io"
"math"
"os"
"slices"
"github.com/anchore/syft/internal/tmpdir"
)
// DefaultMemLimit is how much a buffer keeps in memory before it spills, when the caller does not say.
// Small on purpose: the memory tier saves a temp file for the common small case, it is not a place to
// hold output. A limit large enough to matter is a limit large enough to be an amplification path.
const DefaultMemLimit int64 = 1 << 20 // 1MB
// ErrNoTempDir is returned when a buffer needs to spill and has nowhere to spill to.
var ErrNoTempDir = errors.New("no temp dir available to spill to")
// Buffer is a sparse, offset-addressed buffer. The zero value is not usable; call New.
//
// Storage is two tiers with a fixed boundary: bytes below memLimit live in mem, bytes at or past it live
// in file. Nothing ever moves between them, so an offset belongs to exactly one tier for the life of the
// buffer and the only case needing care is a range that straddles the boundary.
//
// Not safe for concurrent use. Every consumer writes from a single goroutine, and guarding it would cost
// each write an uncontended lock for no reader.
type Buffer struct {
td *tmpdir.TempDir
memLimit int64
// mem holds [0, memLimit). Grown geometrically to fit what has been written, never past the limit.
mem []byte
// file holds [memLimit, ...), created on the first write that reaches there, so a buffer that stays
// small never touches disk. remove deletes it on Close.
file *os.File
remove func()
// written is the set of ranges stored, sorted by start and merged, so runs of adjacent writes collapse
// to one entry rather than growing per call. Everything the buffer reports comes from it.
written []extent
closed bool
}
// extent is a half-open byte range that has been written.
//
// Unexported on purpose. What a caller needs to know about the contents is answered by Size and FirstGap;
// handing out the bookkeeping invites re-deriving those at the call site and getting them subtly wrong,
// which is the bug this package exists to prevent.
type extent struct {
start, end int64
}
// Option configures a Buffer.
type Option func(*Buffer)
// WithMemLimit sets how many bytes are held in memory before the buffer spills to disk. Zero sends every
// write to the temp file. Negative is treated as zero.
//
// Peak allocation is a small multiple of this, since growing the tier holds the old copy alongside the
// new one and the new one may be rounded up past what was asked for.
func WithMemLimit(n int64) Option {
return func(b *Buffer) {
b.memLimit = max(n, 0)
}
}
// New returns a buffer that spills into td.
//
// td may be nil only when every write is known to stay inside the memory limit; a write past it then
// fails with ErrNoTempDir rather than allocating. Nothing is created on disk until a write needs it.
func New(td *tmpdir.TempDir, opts ...Option) *Buffer {
b := &Buffer{td: td, memLimit: DefaultMemLimit}
for _, opt := range opts {
opt(b)
}
return b
}
// split divides a range of n bytes starting at off across the two tiers, reporting how many bytes fall in
// each. The file portion starts at off+inMem.
func (b *Buffer) split(off, n int64) (inMem, inFile int64) {
if off >= b.memLimit {
return 0, n
}
inMem = min(n, b.memLimit-off)
return inMem, n - inMem
}
// fileOffset translates a buffer offset at or past the limit into an offset within the spill file.
func (b *Buffer) fileOffset(off int64) int64 { return off - b.memLimit }
// WriteAt stores p at off. Writes may land anywhere, in any order, and may overlap. A write that fails
// leaves the buffer as it was.
func (b *Buffer) WriteAt(p []byte, off int64) (int, error) {
if b.closed {
return 0, fmt.Errorf("spillbuf: write to a closed buffer: %w", os.ErrClosed)
}
if off < 0 {
return 0, fmt.Errorf("spillbuf: negative offset %d", off)
}
if len(p) == 0 {
return 0, nil
}
// subtraction form so off+len(p) cannot wrap into a small in-range end
if off > math.MaxInt64-int64(len(p)) {
return 0, fmt.Errorf("spillbuf: write of %d bytes at %d overflows", len(p), off)
}
inMem, inFile := b.split(off, int64(len(p)))
// the fallible half goes first: once the memory tier is stamped the old bytes are gone, so a file
// write that fails afterward would report failure over a buffer it had already changed
if inFile > 0 {
if err := b.writeFile(p[inMem:], b.fileOffset(off+inMem)); err != nil {
return 0, err
}
}
if inMem > 0 {
b.writeMem(p[:inMem], off)
}
b.record(extent{start: off, end: off + int64(len(p))})
return len(p), nil
}
func (b *Buffer) writeMem(p []byte, off int64) {
if need := off + int64(len(p)); int64(len(b.mem)) < need {
// geometric, capped at the limit. Sizing to exactly what each write needs looks tidier and is
// quadratic: a caller streaming a block in 32KB chunks reallocates and copies the whole tier every
// chunk, so filling 1MB costs ~16MB of allocation.
grown := min(b.memLimit, max(need, 2*int64(len(b.mem))))
// sized to exactly the geometric target rather than grown through append: append rounds the new
// capacity up past what was asked for, and the tier is already capped, so the slack is pure waste
mem := make([]byte, grown)
copy(mem, b.mem)
b.mem = mem
}
copy(b.mem[off:], p)
}
func (b *Buffer) writeFile(p []byte, off int64) error {
if b.file == nil {
if b.td == nil {
return ErrNoTempDir
}
f, remove, err := b.td.NewFile("syft-spill-*.bin") //nolint:gocritic // removal happens in Close
if err != nil {
return fmt.Errorf("spillbuf: unable to create spill file: %w", err)
}
b.file, b.remove = f, remove
}
if _, err := b.file.WriteAt(p, off); err != nil {
return fmt.Errorf("spillbuf: unable to write to spill file: %w", err)
}
return nil
}
// ReadAt returns bytes this buffer actually stored. A read starting at or past Size is io.EOF, and one
// running past it is short, because everything beyond is a hole the buffer never held.
func (b *Buffer) ReadAt(p []byte, off int64) (int, error) {
if b.closed {
return 0, fmt.Errorf("spillbuf: read from a closed buffer: %w", os.ErrClosed)
}
if off < 0 {
return 0, fmt.Errorf("spillbuf: negative offset %d", off)
}
if len(p) == 0 {
return 0, nil
}
size := b.Size()
if off >= size {
return 0, io.EOF
}
// bounded by Size, so every byte in [off, off+want) was written and both tiers really hold theirs
want := min(int64(len(p)), size-off)
inMem, inFile := b.split(off, want)
// guarded rather than relying on a zero-length copy: when off is past the limit the slice expression
// itself is out of range, since mem is only as long as what was written into it
if inMem > 0 {
copy(p[:inMem], b.mem[off:off+inMem])
}
if inFile > 0 {
if err := readFullAt(b.file, p[inMem:want], b.fileOffset(off+inMem)); err != nil {
return int(inMem), fmt.Errorf("spillbuf: unable to read spill file: %w", err)
}
}
if want < int64(len(p)) {
return int(want), io.EOF
}
return int(want), nil
}
// readFullAt fills p from r. A short read is an error rather than a short return: the caller has already
// bounded the range by what was written, so the file failing to deliver means it is not what we left.
func readFullAt(r io.ReaderAt, p []byte, off int64) error {
for read := 0; read < len(p); {
n, err := r.ReadAt(p[read:], off+int64(read))
read += n
if err != nil {
return err
}
}
return nil
}
// Size reports the contiguous run written from offset zero, which is the length this buffer is willing to
// stand behind. Deliberately not the furthest offset written: see the package doc for why the two are not
// interchangeable.
//
// record keeps the set merged and sorted, so the run from zero is at most the first two entries and the
// loop stops there: the cost this adds to every ReadAt is effectively constant, not a scan of the set.
func (b *Buffer) Size() int64 {
var covered int64
for _, e := range b.written {
if e.start > covered {
break
}
covered = max(covered, e.end)
}
return covered
}
// FirstGap returns the offset of the earliest unwritten run of at least size bytes lying entirely within
// [0, within), and whether there is one. Callers that place output at offsets of their own choosing use it
// to fill what they left behind, without needing the buffer's bookkeeping.
//
// The arithmetic is in subtraction form throughout: within comes from caller input and the result is used
// as a write offset, so a wrap here would be a write nobody bounded.
func (b *Buffer) FirstGap(size, within int64) (int64, bool) {
if b.closed || size < 0 || within < 0 {
return 0, false
}
var at int64
for _, e := range b.written {
// once the cursor is at the bound there is nothing left to offer, and it only moves forward
if at >= within {
return 0, false
}
// the gap ahead of this extent ends at whichever comes first, the extent or the bound
end := min(e.start, within)
if end > at && end-at >= size {
return at, true
}
at = max(at, e.end)
}
if at <= within && within-at >= size {
return at, true
}
return 0, false
}
// record merges e into the written set, keeping it sorted by start with no overlapping or adjacent pairs.
func (b *Buffer) record(e extent) {
if e.end <= e.start {
return
}
// inserted in place rather than appended and re-sorted: the set is already ordered, so a run of
// ascending writes (the common case) stays linear
i, _ := slices.BinarySearchFunc(b.written, e, func(x, y extent) int {
return cmpInt64(x.start, y.start)
})
b.written = slices.Insert(b.written, i, e)
merged := b.written[:1]
for _, cur := range b.written[1:] {
last := &merged[len(merged)-1]
if cur.start <= last.end {
last.end = max(last.end, cur.end)
continue
}
merged = append(merged, cur)
}
b.written = merged
}
func cmpInt64(x, y int64) int {
switch {
case x < y:
return -1
case x > y:
return 1
}
return 0
}
// Close releases the spill file, if one was ever created, and drops the memory tier. Safe on a nil
// receiver and safe to call more than once. Reads and writes after it fail with os.ErrClosed.
func (b *Buffer) Close() error {
if b == nil {
return nil
}
b.mem = nil
b.written = nil
b.closed = true
if b.file == nil {
return nil
}
err := b.file.Close()
if b.remove != nil {
b.remove()
}
b.file, b.remove = nil, nil
return err
}
+872
View File
@@ -0,0 +1,872 @@
package spillbuf
import (
"bytes"
"errors"
"io"
"math"
"math/rand"
"os"
"path/filepath"
"runtime"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"github.com/anchore/syft/internal/tmpdir"
)
func newTestBuffer(t *testing.T, opts ...Option) *Buffer {
t.Helper()
b := New(tmpdir.FromPath(t.TempDir()), opts...)
t.Cleanup(func() { _ = b.Close() })
return b
}
func mustWrite(t *testing.T, b *Buffer, off int64, p []byte) {
t.Helper()
n, err := b.WriteAt(p, off)
require.NoError(t, err)
require.Equal(t, len(p), n, "WriteAt must report the full write")
}
// readAll returns everything the buffer is willing to hand back.
func readAll(t *testing.T, b *Buffer) []byte {
t.Helper()
out := make([]byte, b.Size())
if len(out) == 0 {
return nil
}
n, err := b.ReadAt(out, 0)
require.NoError(t, err, "a read of exactly Size bytes is not short")
require.Equal(t, len(out), n)
return out
}
func repeat(b byte, n int) []byte { return bytes.Repeat([]byte{b}, n) }
// --- the core property: report only what was stored -------------------------------------------------
// TestSizeIsTheContiguousPrefix is the property the package exists for. A sparse buffer hands its holes
// back as zeros it never stored, so reporting a length that covers them lets a write offset stand in for
// real output: whatever sizes an allocation against this reader then allocates bytes nothing produced.
func TestSizeIsTheContiguousPrefix(t *testing.T) {
tests := []struct {
name string
writes []extent // start is the offset, end-start the length
wantSize int64
}{
{
name: "nothing written",
wantSize: 0,
},
{
name: "one run from zero",
writes: []extent{{0, 100}},
wantSize: 100,
},
{
name: "a run that does not start at zero covers nothing",
writes: []extent{{10, 100}},
wantSize: 0,
},
{
name: "adjacent runs join",
writes: []extent{{0, 50}, {50, 100}},
wantSize: 100,
},
{
name: "a hole stops the prefix",
writes: []extent{{0, 50}, {60, 100}},
wantSize: 50,
},
{
name: "the hole is filled later",
writes: []extent{{0, 50}, {60, 100}, {50, 60}},
wantSize: 100,
},
{
name: "out of order still resolves",
writes: []extent{{60, 100}, {0, 50}, {50, 60}},
wantSize: 100,
},
{
name: "overlapping runs",
writes: []extent{{0, 60}, {40, 100}},
wantSize: 100,
},
{
name: "a run fully inside another",
writes: []extent{{0, 100}, {20, 40}},
wantSize: 100,
},
{
name: "far placement is not coverage",
writes: []extent{{0, 64}, {1 << 20, 1<<20 + 64}},
wantSize: 64,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
b := newTestBuffer(t)
for _, w := range tt.writes {
mustWrite(t, b, w.start, repeat('x', int(w.end-w.start)))
}
assert.Equal(t, tt.wantSize, b.Size(), "Size is the contiguous prefix")
})
}
}
// TestReadPastSizeIsEOFNotZeros is the other half of the same property. Sparse storage returns zeros for
// a hole for free; handing those out would make the hole indistinguishable from real content.
func TestReadPastSizeIsEOFNotZeros(t *testing.T) {
b := newTestBuffer(t)
mustWrite(t, b, 0, repeat('A', 64))
mustWrite(t, b, 4096, repeat('B', 64)) // past a hole
require.Equal(t, int64(64), b.Size())
t.Run("a read starting in the hole is EOF", func(t *testing.T) {
p := make([]byte, 16)
n, err := b.ReadAt(p, 100)
assert.Zero(t, n)
assert.ErrorIs(t, err, io.EOF)
})
t.Run("a read starting at the far extent is EOF, not the bytes there", func(t *testing.T) {
p := make([]byte, 64)
n, err := b.ReadAt(p, 4096)
assert.Zero(t, n)
assert.ErrorIs(t, err, io.EOF)
assert.Equal(t, make([]byte, 64), p, "nothing may be copied out")
})
t.Run("a read running off the end is short and says so", func(t *testing.T) {
p := make([]byte, 128)
n, err := b.ReadAt(p, 0)
assert.Equal(t, 64, n, "only the stored prefix comes back")
assert.ErrorIs(t, err, io.EOF)
assert.Equal(t, repeat('A', 64), p[:64])
assert.Equal(t, make([]byte, 64), p[64:], "the rest is untouched")
})
}
// TestSparsePlacementIsNotAnAllocationKnob pins the amplification directly: a tiny payload parked at a
// huge offset must not cost memory proportional to the offset.
func TestSparsePlacementIsNotAnAllocationKnob(t *testing.T) {
const far = 512 << 20 // 512MB out
b := newTestBuffer(t, WithMemLimit(1<<20))
mustWrite(t, b, 0, repeat('A', 64))
var before, after runtime.MemStats
runtime.GC()
runtime.ReadMemStats(&before)
mustWrite(t, b, far, repeat('B', 64))
runtime.ReadMemStats(&after)
allocated := after.TotalAlloc - before.TotalAlloc
t.Logf("writing 64 bytes at offset %d allocated %d bytes", far, allocated)
assert.Less(t, allocated, uint64(1<<20),
"a far placement must cost its own bytes, not its offset")
assert.Equal(t, int64(64), b.Size(), "and it must not count toward what the buffer will deliver")
}
// --- tiers ------------------------------------------------------------------------------------------
func TestMemoryTierNeverTouchesDisk(t *testing.T) {
dir := t.TempDir()
b := New(tmpdir.FromPath(dir), WithMemLimit(4096))
t.Cleanup(func() { _ = b.Close() })
mustWrite(t, b, 0, repeat('A', 4096)) // exactly the limit
assert.Equal(t, repeat('A', 4096), readAll(t, b))
assert.Nil(t, b.file, "a buffer that stays inside the limit must not create a file")
assert.Empty(t, dirEntries(t, dir), "and must leave nothing on disk")
}
func TestSpillCreatesTheFileOnlyWhenNeeded(t *testing.T) {
dir := t.TempDir()
b := New(tmpdir.FromPath(dir), WithMemLimit(4096))
t.Cleanup(func() { _ = b.Close() })
mustWrite(t, b, 0, repeat('A', 4096))
require.Empty(t, dirEntries(t, dir))
mustWrite(t, b, 4096, repeat('B', 1)) // one byte past the limit
assert.NotNil(t, b.file)
assert.Len(t, dirEntries(t, dir), 1, "exactly one spill file")
}
func TestWriteStraddlingTheLimit(t *testing.T) {
const limit = 1024
b := newTestBuffer(t, WithMemLimit(limit))
payload := make([]byte, 2048)
for i := range payload {
payload[i] = byte(i % 251)
}
// starts below the limit and ends well past it, so the write is split across both tiers
mustWrite(t, b, limit-512, payload)
require.Equal(t, int64(0), b.Size(), "nothing was written at zero yet")
mustWrite(t, b, 0, repeat('Z', limit-512))
got := readAll(t, b)
require.Len(t, got, limit-512+2048)
assert.Equal(t, repeat('Z', limit-512), got[:limit-512], "the memory-only part")
assert.Equal(t, payload, got[limit-512:], "the straddling part reads back whole")
}
func TestMemLimitZeroSendsEverythingToDisk(t *testing.T) {
dir := t.TempDir()
b := New(tmpdir.FromPath(dir), WithMemLimit(0))
t.Cleanup(func() { _ = b.Close() })
mustWrite(t, b, 0, repeat('A', 16))
assert.Equal(t, repeat('A', 16), readAll(t, b))
assert.Len(t, dirEntries(t, dir), 1, "with no memory tier the first byte spills")
assert.Nil(t, b.mem)
}
func TestNegativeMemLimitIsTreatedAsZero(t *testing.T) {
b := newTestBuffer(t, WithMemLimit(-1))
mustWrite(t, b, 0, repeat('A', 8))
assert.Equal(t, repeat('A', 8), readAll(t, b))
}
func TestDefaultMemLimitApplies(t *testing.T) {
b := New(nil)
t.Cleanup(func() { _ = b.Close() })
assert.Equal(t, DefaultMemLimit, b.memLimit)
}
// --- no temp dir ------------------------------------------------------------------------------------
func TestNoTempDir(t *testing.T) {
t.Run("a write inside the memory limit needs no temp dir", func(t *testing.T) {
b := New(nil, WithMemLimit(4096))
t.Cleanup(func() { _ = b.Close() })
mustWrite(t, b, 0, repeat('A', 4096))
assert.Equal(t, repeat('A', 4096), readAll(t, b))
})
t.Run("a write past it fails rather than allocating", func(t *testing.T) {
b := New(nil, WithMemLimit(4096))
t.Cleanup(func() { _ = b.Close() })
_, err := b.WriteAt(repeat('B', 1), 4096)
require.Error(t, err)
assert.ErrorIs(t, err, ErrNoTempDir)
})
}
// --- input validation -------------------------------------------------------------------------------
func TestWriteAtRejectsBadInput(t *testing.T) {
b := newTestBuffer(t)
t.Run("negative offset", func(t *testing.T) {
_, err := b.WriteAt([]byte("x"), -1)
assert.Error(t, err)
})
t.Run("an offset plus length that would overflow", func(t *testing.T) {
_, err := b.WriteAt(repeat('x', 16), math.MaxInt64-8)
require.Error(t, err)
assert.Contains(t, err.Error(), "overflow",
"the check must be a subtraction, not off+len wrapping into an in-range end")
})
t.Run("an empty write is a no-op, not an extent", func(t *testing.T) {
n, err := b.WriteAt(nil, 500)
require.NoError(t, err)
assert.Zero(t, n)
assert.Empty(t, b.written, "a zero-length write records nothing")
})
}
func TestReadAtRejectsBadInput(t *testing.T) {
b := newTestBuffer(t)
mustWrite(t, b, 0, repeat('A', 16))
t.Run("negative offset", func(t *testing.T) {
_, err := b.ReadAt(make([]byte, 4), -1)
assert.Error(t, err)
})
t.Run("empty read", func(t *testing.T) {
n, err := b.ReadAt(nil, 0)
assert.NoError(t, err)
assert.Zero(t, n)
})
t.Run("read from an empty buffer", func(t *testing.T) {
empty := newTestBuffer(t)
n, err := empty.ReadAt(make([]byte, 4), 0)
assert.Zero(t, n)
assert.ErrorIs(t, err, io.EOF)
})
}
// TestReadAtHonorsTheReaderAtContract pins what io.ReaderAt requires: a short read must come with a
// non-nil error, and a full read must not invent one. Parsers built on io.ReaderAt (debug/elf among them)
// rely on this to tell "the file ends here" from "try again".
func TestReadAtHonorsTheReaderAtContract(t *testing.T) {
b := newTestBuffer(t, WithMemLimit(64))
mustWrite(t, b, 0, repeat('A', 200)) // spans both tiers
for _, size := range []int{1, 63, 64, 65, 199, 200} {
p := make([]byte, size)
n, err := b.ReadAt(p, 0)
assert.Equal(t, size, n, "size=%d", size)
assert.NoError(t, err, "a read fully inside the buffer must not report EOF (size=%d)", size)
}
p := make([]byte, 201)
n, err := b.ReadAt(p, 0)
assert.Equal(t, 200, n)
assert.ErrorIs(t, err, io.EOF, "one byte past the end is short and must say so")
}
func TestReadAtEveryOffsetAndLength(t *testing.T) {
const limit = 32
b := newTestBuffer(t, WithMemLimit(limit))
want := make([]byte, 100)
for i := range want {
want[i] = byte(i)
}
mustWrite(t, b, 0, want)
// every offset x length pair, so a tier-boundary off-by-one cannot hide
for off := 0; off <= len(want); off++ {
for l := 0; l <= len(want)-off+2; l++ {
p := make([]byte, l)
n, err := b.ReadAt(p, int64(off))
expected := len(want) - off
if l < expected {
expected = l
}
assert.Equal(t, expected, n, "off=%d len=%d", off, l)
assert.Equal(t, want[off:off+expected], p[:expected], "off=%d len=%d", off, l)
switch {
case l == 0:
assert.NoError(t, err, "off=%d len=0", off)
case off >= len(want):
assert.ErrorIs(t, err, io.EOF, "off=%d len=%d", off, l)
case l > expected:
assert.ErrorIs(t, err, io.EOF, "off=%d len=%d", off, l)
default:
assert.NoError(t, err, "off=%d len=%d", off, l)
}
}
}
}
// --- overwrite semantics ----------------------------------------------------------------------------
func TestOverwrite(t *testing.T) {
tests := []struct {
name string
memLimit int64
}{
{"in memory", 1 << 20},
{"on disk", 0},
{"across the boundary", 8},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
b := newTestBuffer(t, WithMemLimit(tt.memLimit))
mustWrite(t, b, 0, repeat('A', 16))
mustWrite(t, b, 4, repeat('B', 8))
want := append(append(repeat('A', 4), repeat('B', 8)...), repeat('A', 4)...)
assert.Equal(t, want, readAll(t, b), "the later write wins over the range it covers")
assert.Equal(t, int64(16), b.Size(), "an overwrite does not extend the buffer")
})
}
}
// --- extents ----------------------------------------------------------------------------------------
func TestWrittenExtentsAreSortedAndMerged(t *testing.T) {
tests := []struct {
name string
writes []extent
want []extent
}{
{
name: "disjoint stay separate, in order",
writes: []extent{{100, 150}, {0, 50}},
want: []extent{{0, 50}, {100, 150}},
},
{
name: "adjacent merge",
writes: []extent{{0, 50}, {50, 100}},
want: []extent{{0, 100}},
},
{
name: "overlapping merge",
writes: []extent{{0, 60}, {40, 100}},
want: []extent{{0, 100}},
},
{
name: "a write bridging two runs collapses all three",
writes: []extent{{0, 20}, {80, 100}, {20, 80}},
want: []extent{{0, 100}},
},
{
name: "a contained write changes nothing",
writes: []extent{{0, 100}, {20, 40}},
want: []extent{{0, 100}},
},
{
name: "identical writes collapse",
writes: []extent{{0, 10}, {0, 10}, {0, 10}},
want: []extent{{0, 10}},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
b := newTestBuffer(t)
for _, w := range tt.writes {
mustWrite(t, b, w.start, repeat('x', int(w.end-w.start)))
}
assert.Equal(t, tt.want, b.written)
})
}
}
func TestWrittenExtentsStayCompactUnderManyAdjacentWrites(t *testing.T) {
b := newTestBuffer(t)
for i := int64(0); i < 1000; i++ {
mustWrite(t, b, i*8, repeat('A', 8))
}
assert.Len(t, b.written, 1, "a run of adjacent writes must collapse to one extent, not 1000")
assert.Equal(t, int64(8000), b.Size())
}
// --- lifecycle --------------------------------------------------------------------------------------
func TestCloseRemovesTheSpillFile(t *testing.T) {
dir := t.TempDir()
b := New(tmpdir.FromPath(dir), WithMemLimit(0))
mustWrite(t, b, 0, repeat('A', 32))
require.Len(t, dirEntries(t, dir), 1)
require.NoError(t, b.Close())
assert.Empty(t, dirEntries(t, dir), "Close must take the spill file with it")
}
func TestCloseIsIdempotent(t *testing.T) {
b := New(tmpdir.FromPath(t.TempDir()), WithMemLimit(0))
mustWrite(t, b, 0, repeat('A', 32))
require.NoError(t, b.Close())
assert.NoError(t, b.Close(), "a second Close is a no-op, not a double-close")
assert.NoError(t, b.Close())
}
func TestCloseOnNilReceiver(t *testing.T) {
var b *Buffer
assert.NotPanics(t, func() {
assert.NoError(t, b.Close())
})
}
func TestCloseWithoutASpillFile(t *testing.T) {
b := New(nil, WithMemLimit(1<<20))
mustWrite(t, b, 0, repeat('A', 32))
assert.NoError(t, b.Close(), "a memory-only buffer closes cleanly with no file to remove")
}
func TestWriteAfterCloseFails(t *testing.T) {
b := New(tmpdir.FromPath(t.TempDir()))
mustWrite(t, b, 0, repeat('A', 8))
require.NoError(t, b.Close())
_, err := b.WriteAt(repeat('B', 8), 0)
require.Error(t, err, "a closed buffer must refuse writes rather than resurrect its file")
assert.ErrorIs(t, err, os.ErrClosed)
}
// TestReadAfterCloseFails pins that a closed buffer reads as closed rather than as empty. Close drops the
// extents, so without the check every read would come back io.EOF and a caller could not tell a buffer it
// still owns from one somebody already released.
func TestReadAfterCloseFails(t *testing.T) {
b := New(tmpdir.FromPath(t.TempDir()))
mustWrite(t, b, 0, repeat('A', 8))
require.NoError(t, b.Close())
n, err := b.ReadAt(make([]byte, 8), 0)
assert.Zero(t, n)
assert.ErrorIs(t, err, os.ErrClosed)
at, ok := b.FirstGap(8, 100)
assert.False(t, ok, "and a closed buffer offers nowhere to write")
assert.Zero(t, at)
}
// TestWriteFailureLeavesTheBufferUnchanged pins the ordering inside WriteAt: the fallible half runs first,
// so a spill file that cannot be created does not take the memory tier down with it.
func TestWriteFailureLeavesTheBufferUnchanged(t *testing.T) {
// FromPath takes the directory as given, so a path that does not exist makes file creation fail
b := New(tmpdir.FromPath(filepath.Join(t.TempDir(), "no-such-dir")), WithMemLimit(16))
t.Cleanup(func() { _ = b.Close() })
mustWrite(t, b, 0, repeat('A', 8))
_, err := b.WriteAt(repeat('B', 32), 0) // straddles the limit, so it needs the spill file
require.Error(t, err)
assert.Equal(t, int64(8), b.Size(), "a failed write records nothing")
assert.Nil(t, b.file, "and leaves no half-made spill file behind")
p := make([]byte, 8)
n, err := b.ReadAt(p, 0)
require.NoError(t, err)
require.Equal(t, 8, n)
assert.Equal(t, repeat('A', 8), p, "the memory tier still holds what it held")
mustWrite(t, b, 8, repeat('C', 8)) // inside the limit, so it still needs no file
assert.Equal(t, append(repeat('A', 8), repeat('C', 8)...), readAll(t, b))
}
// --- interface compliance ---------------------------------------------------------------------------
func TestBufferSatisfiesTheStandardInterfaces(t *testing.T) {
b := newTestBuffer(t)
var (
_ io.WriterAt = b
_ io.ReaderAt = b
_ io.Closer = b
)
assert.Implements(t, (*io.ReaderAt)(nil), b)
assert.Implements(t, (*io.WriterAt)(nil), b)
}
// TestWorksThroughIOOffsetWriter pins the composition every consumer uses: stream a decoded block into
// the buffer at a chosen offset without the writer knowing where it lands.
func TestWorksThroughIOOffsetWriter(t *testing.T) {
b := newTestBuffer(t, WithMemLimit(16))
w := io.NewOffsetWriter(b, 8)
n, err := io.Copy(w, bytes.NewReader(repeat('B', 32)))
require.NoError(t, err)
require.Equal(t, int64(32), n)
mustWrite(t, b, 0, repeat('A', 8))
got := readAll(t, b)
assert.Equal(t, append(repeat('A', 8), repeat('B', 32)...), got)
}
// TestWorksWithIOSectionReader pins the other direction: parsers wrap an io.ReaderAt in a SectionReader,
// and it must see a consistent length.
func TestWorksWithIOSectionReader(t *testing.T) {
b := newTestBuffer(t)
mustWrite(t, b, 0, repeat('A', 100))
sr := io.NewSectionReader(b, 0, b.Size())
got, err := io.ReadAll(sr)
require.NoError(t, err)
assert.Equal(t, repeat('A', 100), got)
}
// --- model-based randomized testing -----------------------------------------------------------------
// TestAgainstAReferenceModel drives random writes through the buffer and a dumb reference at the same
// time, then compares every readable byte. This is what catches tier-boundary and merge bugs that
// hand-written cases miss.
func TestAgainstAReferenceModel(t *testing.T) {
const universe = 8192
for _, limit := range []int64{0, 1, 64, 1000, 4096, universe * 2} {
t.Run("memLimit="+itoa(limit), func(t *testing.T) {
for seed := int64(0); seed < 40; seed++ {
rng := rand.New(rand.NewSource(seed)) //nolint:gosec // deterministic test input, not crypto
b := newTestBuffer(t, WithMemLimit(limit))
model := make([]byte, universe)
written := make([]bool, universe)
for op := 0; op < 40; op++ {
off := rng.Int63n(universe)
l := rng.Int63n(256) + 1
if off+l > universe {
l = universe - off
}
payload := make([]byte, l)
for i := range payload {
payload[i] = byte(rng.Intn(256))
}
mustWrite(t, b, off, payload)
copy(model[off:], payload)
for i := off; i < off+l; i++ {
written[i] = true
}
}
// the reference contiguous prefix
var want int64
for want < universe && written[want] {
want++
}
require.Equal(t, want, b.Size(), "seed=%d limit=%d", seed, limit)
if want == 0 {
continue
}
assert.Equal(t, model[:want], readAll(t, b), "seed=%d limit=%d", seed, limit)
}
})
}
}
// TestRandomReadsMatchTheModel checks partial reads at arbitrary offsets, which the whole-buffer
// comparison above cannot reach.
func TestRandomReadsMatchTheModel(t *testing.T) {
rng := rand.New(rand.NewSource(7)) //nolint:gosec // deterministic test input, not crypto
b := newTestBuffer(t, WithMemLimit(97))
model := make([]byte, 1000)
for i := range model {
model[i] = byte(rng.Intn(256))
}
// written in random-sized chunks, in order, so the whole range is covered
for off := 0; off < len(model); {
l := rng.Intn(50) + 1
if off+l > len(model) {
l = len(model) - off
}
mustWrite(t, b, int64(off), model[off:off+l])
off += l
}
require.Equal(t, int64(len(model)), b.Size())
for i := 0; i < 500; i++ {
off := rng.Intn(len(model))
l := rng.Intn(len(model)-off) + 1
p := make([]byte, l)
n, err := b.ReadAt(p, int64(off))
require.NoError(t, err)
require.Equal(t, l, n)
require.Equal(t, model[off:off+l], p, "off=%d len=%d", off, l)
}
}
// --- helpers ----------------------------------------------------------------------------------------
func dirEntries(t *testing.T, dir string) []os.DirEntry {
t.Helper()
entries, err := os.ReadDir(dir)
if errors.Is(err, os.ErrNotExist) {
return nil
}
require.NoError(t, err)
// tmpdir.NewFile creates the spill file inside the root it is given
var files []os.DirEntry
for _, e := range entries {
if e.IsDir() {
sub, err := os.ReadDir(filepath.Join(dir, e.Name()))
require.NoError(t, err)
for _, s := range sub {
files = append(files, s)
}
continue
}
files = append(files, e)
}
return files
}
func itoa(n int64) string {
if n == 0 {
return "0"
}
var digits []byte
for n > 0 {
digits = append([]byte{byte('0' + n%10)}, digits...)
n /= 10
}
return string(digits)
}
// TestMemoryTierGrowsAmortized is the regression test for quadratic growth. Sizing the tier to exactly
// what each write needs reallocates and copies the whole thing per call, so a caller streaming a block in
// small chunks (which is what io.Copy through an OffsetWriter does) pays O(n^2): filling a 1MB tier in
// 32KB chunks cost ~16MB of allocation and showed up as a 19MB spike decompressing a 16MB payload.
func TestMemoryTierGrowsAmortized(t *testing.T) {
const (
limit = 1 << 20
chunk = 32 << 10
)
b := newTestBuffer(t, WithMemLimit(limit))
var before, after runtime.MemStats
runtime.GC()
runtime.ReadMemStats(&before)
for off := int64(0); off < limit; off += chunk {
mustWrite(t, b, off, repeat('A', chunk))
}
runtime.ReadMemStats(&after)
allocated := after.TotalAlloc - before.TotalAlloc
t.Logf("filling a %d byte tier in %d byte chunks allocated %d bytes", limit, chunk, allocated)
assert.Less(t, allocated, uint64(4*limit),
"growth has to be amortized, not a fresh copy of the whole tier per write")
assert.Equal(t, int64(limit), b.Size())
}
// TestMemoryTierNeverExceedsTheLimit pins the cap that geometric growth could otherwise overshoot.
func TestMemoryTierNeverExceedsTheLimit(t *testing.T) {
const limit = 1000
b := newTestBuffer(t, WithMemLimit(limit))
for off := int64(0); off < 4*limit; off += 100 {
mustWrite(t, b, off, repeat('A', 100))
assert.LessOrEqual(t, int64(len(b.mem)), int64(limit),
"the memory tier may never grow past its own limit (off=%d)", off)
}
assert.Equal(t, int64(4*limit), b.Size())
}
// TestFirstGap covers the placement question a caller writing at offsets of its own choosing has to ask.
// It moved here from the UPX cataloger, which used to walk an exported extent list to answer it itself.
func TestFirstGap(t *testing.T) {
tests := []struct {
name string
writes []extent
size int64
within int64
wantAt int64
wantOK bool
}{
{
name: "an empty buffer starts at zero",
size: 16,
within: 100,
wantAt: 0, wantOK: true,
},
{
name: "the gap between two runs comes first",
writes: []extent{{0, 1000}, {1016, 2000}},
size: 16,
within: 3000,
wantAt: 1000, wantOK: true,
},
{
name: "with the gap filled the tail is what is left",
writes: []extent{{0, 2000}},
size: 1000,
within: 3000,
wantAt: 2000, wantOK: true,
},
{
name: "a gap too small is skipped for a later one",
writes: []extent{{0, 100}, {108, 200}},
size: 50,
within: 1000,
wantAt: 200, wantOK: true,
},
{
name: "nowhere left to fit is reported rather than squeezed in",
writes: []extent{{0, 1000}, {1016, 2000}},
size: 1001,
within: 3000,
wantOK: false,
},
{
name: "a run past the declared total must not wrap the remaining-space subtraction",
writes: []extent{{0, 200}},
size: 8,
within: 100,
wantOK: false,
},
{
name: "exactly filling the tail is allowed",
writes: []extent{{0, 900}},
size: 100,
within: 1000,
wantAt: 900, wantOK: true,
},
{
name: "a gap between two runs is still cut off by the bound",
writes: []extent{{0, 10}, {1000, 1010}},
size: 45,
within: 50,
wantOK: false,
},
{
name: "a gap between two runs that the bound still leaves room in",
writes: []extent{{0, 10}, {1000, 1010}},
size: 30,
within: 50,
wantAt: 10, wantOK: true,
},
{
name: "a negative size is refused rather than treated as zero",
size: -1,
within: 100,
wantOK: false,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
b := newTestBuffer(t)
for _, w := range tt.writes {
mustWrite(t, b, w.start, repeat('x', int(w.end-w.start)))
}
at, ok := b.FirstGap(tt.size, tt.within)
assert.Equal(t, tt.wantOK, ok)
if tt.wantOK {
assert.Equal(t, tt.wantAt, at)
}
})
}
}
// TestFirstGapNeverOverlapsWhatWasWritten is the property behind the table: whatever offset comes back has
// to be usable, so writing size bytes there must not land on top of anything already stored.
func TestFirstGapNeverOverlapsWhatWasWritten(t *testing.T) {
rng := rand.New(rand.NewSource(11)) //nolint:gosec // deterministic test input, not crypto
const universe = 4096
for seed := 0; seed < 200; seed++ {
b := newTestBuffer(t)
occupied := make([]bool, universe)
for op := 0; op < 6; op++ {
// deliberately reaches past universe too, so the bound has to cut a between-extents gap short
off := rng.Int63n(2 * universe)
l := rng.Int63n(200) + 1
mustWrite(t, b, off, repeat('x', int(l)))
for i := off; i < off+l && i < universe; i++ {
occupied[i] = true
}
}
size := rng.Int63n(100) + 1
at, ok := b.FirstGap(size, universe)
if !ok {
continue
}
require.LessOrEqual(t, at+size, int64(universe), "the result must fit inside the declared total")
for i := at; i < at+size; i++ {
require.False(t, occupied[i], "FirstGap returned %d for %d bytes, but %d is taken", at, size, i)
}
}
}
+51 -4
View File
@@ -35,6 +35,7 @@ package elfutil
import (
"debug/elf"
"encoding/binary"
"errors"
"fmt"
"io"
"math"
@@ -55,6 +56,15 @@ import (
// sit just under it. That is deliberate, since the sections syft actually reads are few.
const maxDeclaredSectionSize uint64 = 128 * intFile.MB
// ErrDeclaredSizeExceeded marks a file this package refused to parse because a section declared more
// decompressed bytes than maxDeclaredSectionSize allows.
//
// Exported so a caller can tell "syft declined to expand this" apart from "this file is broken". The
// distinction matters where the two are reported differently: the refusal is a real gap in the SBOM and
// worth surfacing, while a corrupt or non-Go executable is neither, and the golang cataloger sees far
// more of the latter than the former.
var ErrDeclaredSizeExceeded = errors.New("declared decompressed size over the limit")
// legacyZlibHeaderSize is the size of the .zdebug header: the "ZLIB" magic plus a big-endian size.
const legacyZlibHeaderSize = 12
@@ -102,6 +112,39 @@ func NewFile(r io.ReaderAt) (*elf.File, error) {
return f, nil
}
// CheckAllSections rejects ELF decompression bombs before debug/elf can expand one, without keeping the
// parse.
//
// The attack it stops: a compressed section's header declares its own decompressed size, and debug/elf
// believes that number, allocating it up front when the section is opened. Nothing forces the declared
// size to match what the compressed bytes actually yield, so a small file can name an enormous one. In
// the case this was written for, a 260KB ELF declaring a compressed .symtab drove 1.3GB of allocation.
// This walks the section headers and refuses any reachable section declaring more than
// maxDeclaredSectionSize, returning ErrDeclaredSizeExceeded.
//
// It is a strict superset of CheckSectionNameTable, which is what "All" in the name marks: reach for that
// narrower one only where a full elf.NewFile parse on every file is not worth paying for. Like it, this
// is for a caller whose debug/elf call is made inside another package, but it is the whole check rather
// than the eager half: goversion reads .symtab and the string table it links, and those are expanded
// lazily, so the name table bound alone leaves them unbounded.
//
// Anything that is not an ELF this package understands passes through untouched, since a caller reaching
// for this parses other containers too (goversion takes PE and Mach-O). A file elf.NewFile cannot parse
// passes through as well: describing a malformed file is the caller's job, not this gate's.
func CheckAllSections(r io.ReaderAt) error {
if _, _, ok := identify(r); !ok {
return nil
}
if err := CheckSectionNameTable(r); err != nil {
return err
}
f, err := elf.NewFile(r)
if err != nil {
return nil //nolint:nilerr // the caller's own parse reports the format error
}
return checkReachableSections(f)
}
// checkReachableSections bounds every section syft can drive debug/elf into decompressing. This runs
// after the parse because (*Section).Open expands lazily, so nothing has been allocated yet.
func checkReachableSections(f *elf.File) error {
@@ -109,8 +152,8 @@ func checkReachableSections(f *elf.File) error {
s := f.Sections[i]
declared, claimed := declaredSectionSize(s)
if claimed && declared > maxDeclaredSectionSize {
return fmt.Errorf("elf section %q declares %d decompressed bytes, over the %d byte limit",
s.Name, declared, maxDeclaredSectionSize)
return fmt.Errorf("%w: elf section %q declares %d decompressed bytes, over the %d byte limit",
ErrDeclaredSizeExceeded, s.Name, declared, maxDeclaredSectionSize)
}
}
return nil
@@ -233,6 +276,10 @@ func zdebugDeclaredSize(s *elf.Section) (uint64, bool) {
// debug/elf call is made for them inside another package, debug/buildinfo being the one syft reaches:
// gate the reader on this and the eager section-name table read is bounded, which is everything that
// package expands today, since it reaches .go.buildinfo through the program headers instead.
//
// This is deliberately the eager half only, so it is the wrong gate for a caller whose parser goes on to
// read sections of its own. Use CheckAllSections for those: a package that reads symbols expands sections
// this never looks at, and the difference is the whole reason both exist.
func CheckSectionNameTable(r io.ReaderAt) error {
class, order, ok := identify(r)
if !ok {
@@ -252,8 +299,8 @@ func CheckSectionNameTable(r io.ReaderAt) error {
return nil //nolint:nilerr // truncated section; elf.NewFile will say so
}
if declared > maxDeclaredSectionSize {
return fmt.Errorf("elf section name table declares %d decompressed bytes, over the %d byte limit",
declared, maxDeclaredSectionSize)
return fmt.Errorf("%w: elf section name table declares %d decompressed bytes, over the %d byte limit",
ErrDeclaredSizeExceeded, declared, maxDeclaredSectionSize)
}
return nil
}
+84
View File
@@ -341,6 +341,9 @@ func TestNewFile_RejectsOversizedReachableSections(t *testing.T) {
// the fixtures are hand-assembled ELF bytes, so a bad one would satisfy require.Error on
// its own. Naming the section proves the rejection is the one we set up, and naming the
// limit proves it came from this package rather than from debug/elf.
// the sentinel is what the callers key their reporting on, so it is asserted here rather
// than only from the packages downstream that consume it
assert.ErrorIs(t, err, ErrDeclaredSizeExceeded)
assert.Contains(t, err.Error(), tt.wantErr)
assert.Contains(t, err.Error(), fmt.Sprint(maxDeclaredSectionSize))
})
@@ -610,6 +613,87 @@ func TestNewFile_ExtendedShnumOverflowStillChecks(t *testing.T) {
require.Error(t, err, "a bogus section count must not disable the check")
}
// TestCheckAllSections covers the standalone gate, used by callers whose debug/elf parse happens inside
// another package (goversion, here) rather than through NewFile.
func TestCheckAllSections(t *testing.T) {
tests := []struct {
name string
data func(t *testing.T) []byte
// wantErr is empty when CheckAllSections must return nil.
wantErr string
// nameTableMisses marks a fixture CheckSectionNameTable alone does not reject, since the oversized
// section here is one debug/elf only expands lazily, after the parse CheckAllSections goes on to do.
nameTableMisses bool
}{
{
name: "non-ELF reader",
data: func(t *testing.T) []byte { return []byte("not an elf") },
},
{
name: "valid ident but a body elf.NewFile cannot parse",
data: func(t *testing.T) []byte {
full := buildELF(t, elf.ELFCLASS64, binary.LittleEndian, nil, buildOpts{})
truncated := full[:binary.Size(elf.Header64{})+4]
_, err := elf.NewFile(bytes.NewReader(truncated))
require.Error(t, err, "fixture is supposed to be rejected by debug/elf")
return truncated
},
},
{
name: "clean small ELF",
data: func(t *testing.T) []byte {
data := buildELF(t, elf.ELFCLASS64, binary.LittleEndian,
[]section{{name: ".text", typ: elf.SHT_PROGBITS, flags: elf.SHF_ALLOC}}, buildOpts{})
_, err := elf.NewFile(bytes.NewReader(data))
require.NoError(t, err, "fixture is supposed to be a file debug/elf accepts")
return data
},
},
{
name: "oversized section name table",
data: func(t *testing.T) []byte {
return buildELF(t, elf.ELFCLASS64, binary.LittleEndian, nil,
buildOpts{nameTable: &section{compressed: true, declaredSize: overLimit}})
},
wantErr: "section name table",
},
{
name: "oversized .symtab, expanded lazily rather than during the parse",
data: func(t *testing.T) []byte {
data := buildELF(t, elf.ELFCLASS64, binary.LittleEndian, []section{
{name: ".symtab", typ: elf.SHT_SYMTAB, link: 2, compressed: true, declaredSize: overLimit},
{name: ".strtab", typ: elf.SHT_STRTAB},
}, buildOpts{})
_, err := elf.NewFile(bytes.NewReader(data))
require.NoError(t, err, "fixture is supposed to be a file debug/elf accepts")
return data
},
wantErr: `".symtab"`,
nameTableMisses: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
data := tt.data(t)
if tt.nameTableMisses {
require.NoError(t, CheckSectionNameTable(bytes.NewReader(data)),
"the name table check alone doesn't reach a section debug/elf only expands lazily")
}
err := CheckAllSections(bytes.NewReader(data))
if tt.wantErr == "" {
require.NoError(t, err)
return
}
require.Error(t, err)
assert.ErrorIs(t, err, ErrDeclaredSizeExceeded)
assert.Contains(t, err.Error(), tt.wantErr)
})
}
}
func TestNewFile_PassesThroughNonELF(t *testing.T) {
tests := []struct {
name string
+60 -3
View File
@@ -4,7 +4,11 @@ import (
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"github.com/anchore/syft/syft/artifact"
"github.com/anchore/syft/syft/cataloging"
"github.com/anchore/syft/syft/pkg"
"github.com/anchore/syft/syft/pkg/cataloger/internal/pkgtest"
)
@@ -15,6 +19,11 @@ func Test_PackageCataloger_Binary(t *testing.T) {
fixture string
expectedPkgs []string
expectedRels []string
// wantErr is set where the fixture is expected to catalog cleanly. Without it the tester discards
// the cataloger's error, so an unknown newly attached to every file in the image cannot fail this
// table: packages and relationships still match. The packed fixture is the one that needs it, since
// the UPX reporting policy decides per file whether a gap is worth an unknown.
wantErr require.ErrorAssertionFunc
}{
{
name: "simple module with dependencies",
@@ -50,6 +59,7 @@ func Test_PackageCataloger_Binary(t *testing.T) {
{
name: "upx compressed binary",
fixture: "image-small-upx",
wantErr: require.NoError,
expectedPkgs: []string{
"anchore.io/not/real @ v1.0.0 (/run-me)",
"github.com/andybalholm/brotli @ v1.1.1 (/run-me)",
@@ -115,16 +125,63 @@ func Test_PackageCataloger_Binary(t *testing.T) {
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
pkgtest.NewCatalogTester().
tester := pkgtest.NewCatalogTester().
WithImageResolver(t, test.fixture).
ExpectsPackageStrings(test.expectedPkgs).
ExpectsRelationshipStrings(test.expectedRels).
TestCataloger(t, NewGoModuleBinaryCataloger(DefaultCatalogerConfig()))
ExpectsRelationshipStrings(test.expectedRels)
if test.wantErr != nil {
tester = tester.WithErrorAssertion(test.wantErr)
}
tester.TestCataloger(t, NewGoModuleBinaryCataloger(DefaultCatalogerConfig()))
})
}
}
// Test_PackageCataloger_Binary_SymbolsFromPackedBinary covers what the packed path used to miss: the
// pclntab is compressed along with everything else, so a UPX binary reported its packages with no symbols
// at all until the unpacked contents were threaded through the rest of the scan instead of only the
// build info read.
func Test_PackageCataloger_Binary_SymbolsFromPackedBinary(t *testing.T) {
cfg := DefaultCatalogerConfig()
cfg.CaptureSymbols = cataloging.SymbolScopeAll
symbolCounts := func(t *testing.T, pkgs []pkg.Package, _ []artifact.Relationship) map[string]int {
t.Helper()
counts := make(map[string]int)
for _, p := range pkgs {
meta, ok := p.Metadata.(pkg.GolangBinaryBuildinfoEntry)
require.True(t, ok, "unexpected metadata on %s", p.Name)
for _, names := range meta.Symbols {
counts[p.Name] += len(names)
}
}
return counts
}
var packed, unpacked map[string]int
pkgtest.NewCatalogTester().
WithImageResolver(t, "image-small").
ExpectsAssertion(func(t *testing.T, pkgs []pkg.Package, rels []artifact.Relationship) {
unpacked = symbolCounts(t, pkgs, rels)
}).
TestCataloger(t, NewGoModuleBinaryCataloger(cfg))
pkgtest.NewCatalogTester().
WithImageResolver(t, "image-small-upx").
// a real `upx --best --lzma` binary must unpack without leaving a gap behind. The reconstruction is
// truncated to the contiguous extent rebuilt, and a chain that stops short of p_filesize is now a
// reported partial, so this is the assertion that catches that firing on well-formed output.
WithErrorAssertion(require.NoError).
ExpectsAssertion(func(t *testing.T, pkgs []pkg.Package, rels []artifact.Relationship) {
packed = symbolCounts(t, pkgs, rels)
}).
TestCataloger(t, NewGoModuleBinaryCataloger(cfg))
require.NotEmpty(t, unpacked, "the unpacked fixture is the control: it must carry symbols")
assert.Equal(t, unpacked, packed, "the packed binary must report the same symbols as the one it was packed from")
}
func Test_Mod_Cataloger_Globs(t *testing.T) {
tests := []struct {
name string
@@ -1,4 +0,0 @@
xcoff
-----
The code in this package comes from: https://github.com/golang/go/tree/master/src/internal/xcoff -- it was copied over to add support for [xcoff](https://en.wikipedia.org/wiki/XCOFF) binaries. Golang keeps this package as internal, forbidding its external use.
@@ -1,688 +0,0 @@
// Copyright 2018 The Go Authors. All rights reserved.
// Use of this source code is governed by a BSD-style
// license that can be found in the LICENSE file.
// Package xcoff implements access to XCOFF (Extended Common Object File Format) files.
//nolint:all
package xcoff
import (
"encoding/binary"
"fmt"
"io"
"os"
"strings"
)
// SectionHeader holds information about an XCOFF section header.
type SectionHeader struct {
Name string
VirtualAddress uint64
Size uint64
Type uint32
Relptr uint64
Nreloc uint32
}
type Section struct {
SectionHeader
Relocs []Reloc
io.ReaderAt
sr *io.SectionReader
}
// AuxiliaryCSect holds information about an XCOFF symbol in an AUX_CSECT entry.
type AuxiliaryCSect struct {
Length int64
StorageMappingClass int
SymbolType int
}
// AuxiliaryFcn holds information about an XCOFF symbol in an AUX_FCN entry.
type AuxiliaryFcn struct {
Size int64
}
type Symbol struct {
Name string
Value uint64
SectionNumber int
StorageClass int
AuxFcn AuxiliaryFcn
AuxCSect AuxiliaryCSect
}
type Reloc struct {
VirtualAddress uint64
Symbol *Symbol
Signed bool
InstructionFixed bool
Length uint8
Type uint8
}
// ImportedSymbol holds information about an imported XCOFF symbol.
type ImportedSymbol struct {
Name string
Library string
}
// FileHeader holds information about an XCOFF file header.
type FileHeader struct {
TargetMachine uint16
}
// A File represents an open XCOFF file.
type File struct {
FileHeader
Sections []*Section
Symbols []*Symbol
StringTable []byte
LibraryPaths []string
closer io.Closer
}
// Open opens the named file using os.Open and prepares it for use as an XCOFF binary.
func Open(name string) (*File, error) {
f, err := os.Open(name)
if err != nil {
return nil, err
}
ff, err := NewFile(f)
if err != nil {
f.Close()
return nil, err
}
ff.closer = f
return ff, nil
}
// Close closes the File.
// If the File was created using NewFile directly instead of Open,
// Close has no effect.
func (f *File) Close() error {
var err error
if f.closer != nil {
err = f.closer.Close()
f.closer = nil
}
return err
}
// Section returns the first section with the given name, or nil if no such
// section exists.
// Xcoff have section's name limited to 8 bytes. Some sections like .gosymtab
// can be trunked but this method will still find them.
func (f *File) Section(name string) *Section {
for _, s := range f.Sections {
if s.Name == name || (len(name) > 8 && s.Name == name[:8]) {
return s
}
}
return nil
}
// SectionByType returns the first section in f with the
// given type, or nil if there is no such section.
func (f *File) SectionByType(typ uint32) *Section {
for _, s := range f.Sections {
if s.Type == typ {
return s
}
}
return nil
}
// cstring converts ASCII byte sequence b to string.
// It stops once it finds 0 or reaches end of b.
func cstring(b []byte) string {
var i int
for i = 0; i < len(b) && b[i] != 0; i++ {
}
return string(b[:i])
}
// getString extracts a string from an XCOFF string table.
func getString(st []byte, offset uint32) (string, bool) {
if offset < 4 || int(offset) >= len(st) {
return "", false
}
return cstring(st[offset:]), true
}
// NewFile creates a new File for accessing an XCOFF binary in an underlying reader.
func NewFile(r io.ReaderAt) (*File, error) {
sr := io.NewSectionReader(r, 0, 1<<63-1)
// Read XCOFF target machine
var magic uint16
if err := binary.Read(sr, binary.BigEndian, &magic); err != nil {
return nil, err
}
if magic != U802TOCMAGIC && magic != U64_TOCMAGIC {
return nil, fmt.Errorf("unrecognised XCOFF magic: 0x%x", magic)
}
f := new(File)
f.TargetMachine = magic
// Read XCOFF file header
if _, err := sr.Seek(0, io.SeekStart); err != nil {
return nil, err
}
var nscns uint16
var symptr uint64
var nsyms int32
var opthdr uint16
var hdrsz int
switch f.TargetMachine {
case U802TOCMAGIC:
fhdr := new(FileHeader32)
if err := binary.Read(sr, binary.BigEndian, fhdr); err != nil {
return nil, err
}
nscns = fhdr.Fnscns
symptr = uint64(fhdr.Fsymptr)
nsyms = fhdr.Fnsyms
opthdr = fhdr.Fopthdr
hdrsz = FILHSZ_32
case U64_TOCMAGIC:
fhdr := new(FileHeader64)
if err := binary.Read(sr, binary.BigEndian, fhdr); err != nil {
return nil, err
}
nscns = fhdr.Fnscns
symptr = fhdr.Fsymptr
nsyms = fhdr.Fnsyms
opthdr = fhdr.Fopthdr
hdrsz = FILHSZ_64
}
if symptr == 0 || nsyms <= 0 {
return nil, fmt.Errorf("no symbol table")
}
// Read string table (located right after symbol table).
offset := symptr + uint64(nsyms)*SYMESZ
if _, err := sr.Seek(int64(offset), io.SeekStart); err != nil {
return nil, err
}
// The first 4 bytes contain the length (in bytes).
var l uint32
if err := binary.Read(sr, binary.BigEndian, &l); err != nil {
return nil, err
}
if l > 4 {
if _, err := sr.Seek(int64(offset), io.SeekStart); err != nil {
return nil, err
}
f.StringTable = make([]byte, l)
if _, err := io.ReadFull(sr, f.StringTable); err != nil {
return nil, err
}
}
// Read section headers
if _, err := sr.Seek(int64(hdrsz)+int64(opthdr), io.SeekStart); err != nil {
return nil, err
}
f.Sections = make([]*Section, nscns)
for i := 0; i < int(nscns); i++ {
var scnptr uint64
s := new(Section)
switch f.TargetMachine {
case U802TOCMAGIC:
shdr := new(SectionHeader32)
if err := binary.Read(sr, binary.BigEndian, shdr); err != nil {
return nil, err
}
s.Name = cstring(shdr.Sname[:])
s.VirtualAddress = uint64(shdr.Svaddr)
s.Size = uint64(shdr.Ssize)
scnptr = uint64(shdr.Sscnptr)
s.Type = shdr.Sflags
s.Relptr = uint64(shdr.Srelptr)
s.Nreloc = uint32(shdr.Snreloc)
case U64_TOCMAGIC:
shdr := new(SectionHeader64)
if err := binary.Read(sr, binary.BigEndian, shdr); err != nil {
return nil, err
}
s.Name = cstring(shdr.Sname[:])
s.VirtualAddress = shdr.Svaddr
s.Size = shdr.Ssize
scnptr = shdr.Sscnptr
s.Type = shdr.Sflags
s.Relptr = shdr.Srelptr
s.Nreloc = shdr.Snreloc
}
r2 := r
if scnptr == 0 { // .bss must have all 0s
r2 = zeroReaderAt{}
}
s.sr = io.NewSectionReader(r2, int64(scnptr), int64(s.Size))
s.ReaderAt = s.sr
f.Sections[i] = s
}
// Symbol map needed by relocation
var idxToSym = make(map[int]*Symbol)
// Read symbol table
if _, err := sr.Seek(int64(symptr), io.SeekStart); err != nil {
return nil, err
}
f.Symbols = make([]*Symbol, 0)
for i := 0; i < int(nsyms); i++ {
var numaux int
var ok, needAuxFcn bool
sym := new(Symbol)
switch f.TargetMachine {
case U802TOCMAGIC:
se := new(SymEnt32)
if err := binary.Read(sr, binary.BigEndian, se); err != nil {
return nil, err
}
numaux = int(se.Nnumaux)
sym.SectionNumber = int(se.Nscnum)
sym.StorageClass = int(se.Nsclass)
sym.Value = uint64(se.Nvalue)
needAuxFcn = se.Ntype&SYM_TYPE_FUNC != 0 && numaux > 1
zeroes := binary.BigEndian.Uint32(se.Nname[:4])
if zeroes != 0 {
sym.Name = cstring(se.Nname[:])
} else {
offset := binary.BigEndian.Uint32(se.Nname[4:])
sym.Name, ok = getString(f.StringTable, offset)
if !ok {
goto skip
}
}
case U64_TOCMAGIC:
se := new(SymEnt64)
if err := binary.Read(sr, binary.BigEndian, se); err != nil {
return nil, err
}
numaux = int(se.Nnumaux)
sym.SectionNumber = int(se.Nscnum)
sym.StorageClass = int(se.Nsclass)
sym.Value = se.Nvalue
needAuxFcn = se.Ntype&SYM_TYPE_FUNC != 0 && numaux > 1
sym.Name, ok = getString(f.StringTable, se.Noffset)
if !ok {
goto skip
}
}
if sym.StorageClass != C_EXT && sym.StorageClass != C_WEAKEXT && sym.StorageClass != C_HIDEXT {
goto skip
}
// Must have at least one csect auxiliary entry.
if numaux < 1 || i+numaux >= int(nsyms) {
goto skip
}
if sym.SectionNumber > int(nscns) {
goto skip
}
if sym.SectionNumber == 0 {
sym.Value = 0
} else {
sym.Value -= f.Sections[sym.SectionNumber-1].VirtualAddress
}
idxToSym[i] = sym
// If this symbol is a function, it must retrieve its size from
// its AUX_FCN entry.
// It can happen that a function symbol doesn't have any AUX_FCN.
// In this case, needAuxFcn is false and their size will be set to 0.
if needAuxFcn {
switch f.TargetMachine {
case U802TOCMAGIC:
aux := new(AuxFcn32)
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
return nil, err
}
sym.AuxFcn.Size = int64(aux.Xfsize)
case U64_TOCMAGIC:
aux := new(AuxFcn64)
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
return nil, err
}
sym.AuxFcn.Size = int64(aux.Xfsize)
}
}
// Read csect auxiliary entry (by convention, it is the last).
if !needAuxFcn {
if _, err := sr.Seek(int64(numaux-1)*SYMESZ, io.SeekCurrent); err != nil {
return nil, err
}
}
i += numaux
numaux = 0
switch f.TargetMachine {
case U802TOCMAGIC:
aux := new(AuxCSect32)
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
return nil, err
}
sym.AuxCSect.SymbolType = int(aux.Xsmtyp & 0x7)
sym.AuxCSect.StorageMappingClass = int(aux.Xsmclas)
sym.AuxCSect.Length = int64(aux.Xscnlen)
case U64_TOCMAGIC:
aux := new(AuxCSect64)
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
return nil, err
}
sym.AuxCSect.SymbolType = int(aux.Xsmtyp & 0x7)
sym.AuxCSect.StorageMappingClass = int(aux.Xsmclas)
sym.AuxCSect.Length = int64(aux.Xscnlenhi)<<32 | int64(aux.Xscnlenlo)
}
f.Symbols = append(f.Symbols, sym)
skip:
i += numaux // Skip auxiliary entries
if _, err := sr.Seek(int64(numaux)*SYMESZ, io.SeekCurrent); err != nil {
return nil, err
}
}
// Read relocations
// Only for .data or .text section
for _, sect := range f.Sections {
if sect.Type != STYP_TEXT && sect.Type != STYP_DATA {
continue
}
sect.Relocs = make([]Reloc, sect.Nreloc)
if sect.Relptr == 0 {
continue
}
if _, err := sr.Seek(int64(sect.Relptr), io.SeekStart); err != nil {
return nil, err
}
for i := uint32(0); i < sect.Nreloc; i++ {
switch f.TargetMachine {
case U802TOCMAGIC:
rel := new(Reloc32)
if err := binary.Read(sr, binary.BigEndian, rel); err != nil {
return nil, err
}
sect.Relocs[i].VirtualAddress = uint64(rel.Rvaddr)
sect.Relocs[i].Symbol = idxToSym[int(rel.Rsymndx)]
sect.Relocs[i].Type = rel.Rtype
sect.Relocs[i].Length = rel.Rsize&0x3F + 1
if rel.Rsize&0x80 != 0 {
sect.Relocs[i].Signed = true
}
if rel.Rsize&0x40 != 0 {
sect.Relocs[i].InstructionFixed = true
}
case U64_TOCMAGIC:
rel := new(Reloc64)
if err := binary.Read(sr, binary.BigEndian, rel); err != nil {
return nil, err
}
sect.Relocs[i].VirtualAddress = rel.Rvaddr
sect.Relocs[i].Symbol = idxToSym[int(rel.Rsymndx)]
sect.Relocs[i].Type = rel.Rtype
sect.Relocs[i].Length = rel.Rsize&0x3F + 1
if rel.Rsize&0x80 != 0 {
sect.Relocs[i].Signed = true
}
if rel.Rsize&0x40 != 0 {
sect.Relocs[i].InstructionFixed = true
}
}
}
}
return f, nil
}
// zeroReaderAt is ReaderAt that reads 0s.
type zeroReaderAt struct{}
// ReadAt writes len(p) 0s into p.
func (w zeroReaderAt) ReadAt(p []byte, off int64) (n int, err error) {
for i := range p {
p[i] = 0
}
return len(p), nil
}
// Data reads and returns the contents of the XCOFF section s.
func (s *Section) Data() ([]byte, error) {
dat := make([]byte, s.sr.Size())
n, err := s.sr.ReadAt(dat, 0)
if n == len(dat) {
err = nil
}
return dat[:n], err
}
// CSect reads and returns the contents of a csect.
// func (f *File) CSect(name string) []byte {
// for _, sym := range f.Symbols {
// if sym.Name == name && sym.AuxCSect.SymbolType == XTY_SD {
// if i := sym.SectionNumber - 1; 0 <= i && i < len(f.Sections) {
// s := f.Sections[i]
// if sym.Value+uint64(sym.AuxCSect.Length) <= s.Size {
// dat := make([]byte, sym.AuxCSect.Length)
// _, err := s.sr.ReadAt(dat, int64(sym.Value))
// if err != nil {
// return nil
// }
// return dat
// }
// }
// break
// }
// }
// return nil
// }
// func (f *File) DWARF() (*dwarf.Data, error) {
// // There are many other DWARF sections, but these
// // are the ones the debug/dwarf package uses.
// // Don't bother loading others.
// var subtypes = [...]uint32{SSUBTYP_DWABREV, SSUBTYP_DWINFO, SSUBTYP_DWLINE, SSUBTYP_DWRNGES, SSUBTYP_DWSTR}
// var dat [len(subtypes)][]byte
// for i, subtype := range subtypes {
// s := f.SectionByType(STYP_DWARF | subtype)
// if s != nil {
// b, err := s.Data()
// if err != nil && uint64(len(b)) < s.Size {
// return nil, err
// }
// dat[i] = b
// }
// }
// abbrev, info, line, ranges, str := dat[0], dat[1], dat[2], dat[3], dat[4]
// return dwarf.New(abbrev, nil, nil, info, line, nil, ranges, str)
// }
// readImportIDs returns the import file IDs stored inside the .loader section.
// Library name pattern is either path/base/member or base/member
func (f *File) readImportIDs(s *Section) ([]string, error) {
// Read loader header
if _, err := s.sr.Seek(0, io.SeekStart); err != nil {
return nil, err
}
var istlen uint32
var nimpid int32
var impoff uint64
switch f.TargetMachine {
case U802TOCMAGIC:
lhdr := new(LoaderHeader32)
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
return nil, err
}
istlen = lhdr.Listlen
nimpid = lhdr.Lnimpid
impoff = uint64(lhdr.Limpoff)
case U64_TOCMAGIC:
lhdr := new(LoaderHeader64)
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
return nil, err
}
istlen = lhdr.Listlen
nimpid = lhdr.Lnimpid
impoff = lhdr.Limpoff
}
// Read loader import file ID table
if _, err := s.sr.Seek(int64(impoff), io.SeekStart); err != nil {
return nil, err
}
table := make([]byte, istlen)
if _, err := io.ReadFull(s.sr, table); err != nil {
return nil, err
}
offset := 0
// First import file ID is the default LIBPATH value
libpath := cstring(table[offset:])
f.LibraryPaths = strings.Split(libpath, ":")
offset += len(libpath) + 3 // 3 null bytes
all := make([]string, 0)
for i := 1; i < int(nimpid); i++ {
impidpath := cstring(table[offset:])
offset += len(impidpath) + 1
impidbase := cstring(table[offset:])
offset += len(impidbase) + 1
impidmem := cstring(table[offset:])
offset += len(impidmem) + 1
var path string
if len(impidpath) > 0 {
path = impidpath + "/" + impidbase + "/" + impidmem
} else {
path = impidbase + "/" + impidmem
}
all = append(all, path)
}
return all, nil
}
// ImportedSymbols returns the names of all symbols
// referred to by the binary f that are expected to be
// satisfied by other libraries at dynamic load time.
// It does not return weak symbols.
func (f *File) ImportedSymbols() ([]ImportedSymbol, error) {
s := f.SectionByType(STYP_LOADER)
if s == nil {
return nil, nil
}
// Read loader header
if _, err := s.sr.Seek(0, io.SeekStart); err != nil {
return nil, err
}
var stlen uint32
var stoff uint64
var nsyms int32
var symoff uint64
switch f.TargetMachine {
case U802TOCMAGIC:
lhdr := new(LoaderHeader32)
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
return nil, err
}
stlen = lhdr.Lstlen
stoff = uint64(lhdr.Lstoff)
nsyms = lhdr.Lnsyms
symoff = LDHDRSZ_32
case U64_TOCMAGIC:
lhdr := new(LoaderHeader64)
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
return nil, err
}
stlen = lhdr.Lstlen
stoff = lhdr.Lstoff
nsyms = lhdr.Lnsyms
symoff = lhdr.Lsymoff
}
// Read loader section string table
if _, err := s.sr.Seek(int64(stoff), io.SeekStart); err != nil {
return nil, err
}
st := make([]byte, stlen)
if _, err := io.ReadFull(s.sr, st); err != nil {
return nil, err
}
// Read imported libraries
libs, err := f.readImportIDs(s)
if err != nil {
return nil, err
}
// Read loader symbol table
if _, err := s.sr.Seek(int64(symoff), io.SeekStart); err != nil {
return nil, err
}
all := make([]ImportedSymbol, 0)
for i := 0; i < int(nsyms); i++ {
var name string
var ifile int32
var ok bool
switch f.TargetMachine {
case U802TOCMAGIC:
ldsym := new(LoaderSymbol32)
if err := binary.Read(s.sr, binary.BigEndian, ldsym); err != nil {
return nil, err
}
if ldsym.Lsmtype&0x40 == 0 {
continue // Imported symbols only
}
zeroes := binary.BigEndian.Uint32(ldsym.Lname[:4])
if zeroes != 0 {
name = cstring(ldsym.Lname[:])
} else {
offset := binary.BigEndian.Uint32(ldsym.Lname[4:])
name, ok = getString(st, offset)
if !ok {
continue
}
}
ifile = ldsym.Lifile
case U64_TOCMAGIC:
ldsym := new(LoaderSymbol64)
if err := binary.Read(s.sr, binary.BigEndian, ldsym); err != nil {
return nil, err
}
if ldsym.Lsmtype&0x40 == 0 {
continue // Imported symbols only
}
name, ok = getString(st, ldsym.Loffset)
if !ok {
continue
}
ifile = ldsym.Lifile
}
var sym ImportedSymbol
sym.Name = name
if ifile >= 1 && int(ifile) <= len(libs) {
sym.Library = libs[ifile-1]
}
all = append(all, sym)
}
return all, nil
}
// ImportedLibraries returns the names of all libraries
// referred to by the binary f that are expected to be
// linked with the binary at dynamic link time.
func (f *File) ImportedLibraries() ([]string, error) {
s := f.SectionByType(STYP_LOADER)
if s == nil {
return nil, nil
}
all, err := f.readImportIDs(s)
return all, err
}
@@ -1,102 +0,0 @@
// Copyright 2018 The Go Authors. All rights reserved.
// Use of this source code is governed by a BSD-style
// license that can be found in the LICENSE file.
package xcoff
import (
"reflect"
"testing"
)
type fileTest struct {
file string
hdr FileHeader
sections []*SectionHeader
needed []string
}
var fileTests = []fileTest{
{
"testdata/gcc-ppc32-aix-dwarf2-exec",
FileHeader{U802TOCMAGIC},
[]*SectionHeader{
{".text", 0x10000290, 0x00000bbd, STYP_TEXT, 0x7ae6, 0x36},
{".data", 0x20000e4d, 0x00000437, STYP_DATA, 0x7d02, 0x2b},
{".bss", 0x20001284, 0x0000021c, STYP_BSS, 0, 0},
{".loader", 0x00000000, 0x000004b3, STYP_LOADER, 0, 0},
{".dwline", 0x00000000, 0x000000df, STYP_DWARF | SSUBTYP_DWLINE, 0x7eb0, 0x7},
{".dwinfo", 0x00000000, 0x00000314, STYP_DWARF | SSUBTYP_DWINFO, 0x7ef6, 0xa},
{".dwabrev", 0x00000000, 0x000000d6, STYP_DWARF | SSUBTYP_DWABREV, 0, 0},
{".dwarnge", 0x00000000, 0x00000020, STYP_DWARF | SSUBTYP_DWARNGE, 0x7f5a, 0x2},
{".dwloc", 0x00000000, 0x00000074, STYP_DWARF | SSUBTYP_DWLOC, 0, 0},
{".debug", 0x00000000, 0x00005e4f, STYP_DEBUG, 0, 0},
},
[]string{"libc.a/shr.o"},
},
{
"testdata/gcc-ppc64-aix-dwarf2-exec",
FileHeader{U64_TOCMAGIC},
[]*SectionHeader{
{".text", 0x10000480, 0x00000afd, STYP_TEXT, 0x8322, 0x34},
{".data", 0x20000f7d, 0x000002f3, STYP_DATA, 0x85fa, 0x25},
{".bss", 0x20001270, 0x00000428, STYP_BSS, 0, 0},
{".loader", 0x00000000, 0x00000535, STYP_LOADER, 0, 0},
{".dwline", 0x00000000, 0x000000b4, STYP_DWARF | SSUBTYP_DWLINE, 0x8800, 0x4},
{".dwinfo", 0x00000000, 0x0000036a, STYP_DWARF | SSUBTYP_DWINFO, 0x8838, 0x7},
{".dwabrev", 0x00000000, 0x000000b5, STYP_DWARF | SSUBTYP_DWABREV, 0, 0},
{".dwarnge", 0x00000000, 0x00000040, STYP_DWARF | SSUBTYP_DWARNGE, 0x889a, 0x2},
{".dwloc", 0x00000000, 0x00000062, STYP_DWARF | SSUBTYP_DWLOC, 0, 0},
{".debug", 0x00000000, 0x00006605, STYP_DEBUG, 0, 0},
},
[]string{"libc.a/shr_64.o"},
},
}
func TestOpen(t *testing.T) {
for i := range fileTests {
tt := &fileTests[i]
f, err := Open(tt.file)
if err != nil {
t.Error(err)
continue
}
if !reflect.DeepEqual(f.FileHeader, tt.hdr) {
t.Errorf("open %s:\n\thave %#v\n\twant %#v\n", tt.file, f.FileHeader, tt.hdr)
continue
}
for i, sh := range f.Sections {
if i >= len(tt.sections) {
break
}
have := &sh.SectionHeader
want := tt.sections[i]
if !reflect.DeepEqual(have, want) {
t.Errorf("open %s, section %d:\n\thave %#v\n\twant %#v\n", tt.file, i, have, want)
}
}
tn := len(tt.sections)
fn := len(f.Sections)
if tn != fn {
t.Errorf("open %s: len(Sections) = %d, want %d", tt.file, fn, tn)
}
tl := tt.needed
fl, err := f.ImportedLibraries()
if err != nil {
t.Error(err)
}
if !reflect.DeepEqual(tl, fl) {
t.Errorf("open %s: loader import = %v, want %v", tt.file, tl, fl)
}
}
}
func TestOpenFailure(t *testing.T) {
filename := "file.go" // not an XCOFF object file
_, err := Open(filename) // don't crash
if err == nil {
t.Errorf("open %s: succeeded unexpectedly", filename)
}
}
@@ -1,373 +0,0 @@
// The code in this package comes from:
// https://github.com/golang/go/tree/master/src/internal/xcoff
// it was copied over to add support for xcoff binaries.
// Golang keeps this package as internal, forbidding its external use.
// Copyright 2018 The Go Authors. All rights reserved.
// Use of this source code is governed by a BSD-style
// license that can be found in the LICENSE file.
//nolint:all
package xcoff
// File Header.
type FileHeader32 struct {
Fmagic uint16 // Target machine
Fnscns uint16 // Number of sections
Ftimedat int32 // Time and date of file creation
Fsymptr uint32 // Byte offset to symbol table start
Fnsyms int32 // Number of entries in symbol table
Fopthdr uint16 // Number of bytes in optional header
Fflags uint16 // Flags
}
type FileHeader64 struct {
Fmagic uint16 // Target machine
Fnscns uint16 // Number of sections
Ftimedat int32 // Time and date of file creation
Fsymptr uint64 // Byte offset to symbol table start
Fopthdr uint16 // Number of bytes in optional header
Fflags uint16 // Flags
Fnsyms int32 // Number of entries in symbol table
}
const (
FILHSZ_32 = 20
FILHSZ_64 = 24
)
const (
U802TOCMAGIC = 0737 // AIX 32-bit XCOFF
U64_TOCMAGIC = 0767 // AIX 64-bit XCOFF
)
// Flags that describe the type of the object file.
const (
F_RELFLG = 0x0001
F_EXEC = 0x0002
F_LNNO = 0x0004
F_FDPR_PROF = 0x0010
F_FDPR_OPTI = 0x0020
F_DSA = 0x0040
F_VARPG = 0x0100
F_DYNLOAD = 0x1000
F_SHROBJ = 0x2000
F_LOADONLY = 0x4000
)
// Section Header.
type SectionHeader32 struct {
Sname [8]byte // Section name
Spaddr uint32 // Physical address
Svaddr uint32 // Virtual address
Ssize uint32 // Section size
Sscnptr uint32 // Offset in file to raw data for section
Srelptr uint32 // Offset in file to relocation entries for section
Slnnoptr uint32 // Offset in file to line number entries for section
Snreloc uint16 // Number of relocation entries
Snlnno uint16 // Number of line number entries
Sflags uint32 // Flags to define the section type
}
type SectionHeader64 struct {
Sname [8]byte // Section name
Spaddr uint64 // Physical address
Svaddr uint64 // Virtual address
Ssize uint64 // Section size
Sscnptr uint64 // Offset in file to raw data for section
Srelptr uint64 // Offset in file to relocation entries for section
Slnnoptr uint64 // Offset in file to line number entries for section
Snreloc uint32 // Number of relocation entries
Snlnno uint32 // Number of line number entries
Sflags uint32 // Flags to define the section type
Spad uint32 // Needs to be 72 bytes long
}
// Flags defining the section type.
const (
STYP_DWARF = 0x0010
STYP_TEXT = 0x0020
STYP_DATA = 0x0040
STYP_BSS = 0x0080
STYP_EXCEPT = 0x0100
STYP_INFO = 0x0200
STYP_TDATA = 0x0400
STYP_TBSS = 0x0800
STYP_LOADER = 0x1000
STYP_DEBUG = 0x2000
STYP_TYPCHK = 0x4000
STYP_OVRFLO = 0x8000
)
const (
SSUBTYP_DWINFO = 0x10000 // DWARF info section
SSUBTYP_DWLINE = 0x20000 // DWARF line-number section
SSUBTYP_DWPBNMS = 0x30000 // DWARF public names section
SSUBTYP_DWPBTYP = 0x40000 // DWARF public types section
SSUBTYP_DWARNGE = 0x50000 // DWARF aranges section
SSUBTYP_DWABREV = 0x60000 // DWARF abbreviation section
SSUBTYP_DWSTR = 0x70000 // DWARF strings section
SSUBTYP_DWRNGES = 0x80000 // DWARF ranges section
SSUBTYP_DWLOC = 0x90000 // DWARF location lists section
SSUBTYP_DWFRAME = 0xA0000 // DWARF frames section
SSUBTYP_DWMAC = 0xB0000 // DWARF macros section
)
// Symbol Table Entry.
type SymEnt32 struct {
Nname [8]byte // Symbol name
Nvalue uint32 // Symbol value
Nscnum int16 // Section number of symbol
Ntype uint16 // Basic and derived type specification
Nsclass int8 // Storage class of symbol
Nnumaux int8 // Number of auxiliary entries
}
type SymEnt64 struct {
Nvalue uint64 // Symbol value
Noffset uint32 // Offset of the name in string table or .debug section
Nscnum int16 // Section number of symbol
Ntype uint16 // Basic and derived type specification
Nsclass int8 // Storage class of symbol
Nnumaux int8 // Number of auxiliary entries
}
const SYMESZ = 18
const (
// Nscnum
N_DEBUG = -2
N_ABS = -1
N_UNDEF = 0
//Ntype
SYM_V_INTERNAL = 0x1000
SYM_V_HIDDEN = 0x2000
SYM_V_PROTECTED = 0x3000
SYM_V_EXPORTED = 0x4000
SYM_TYPE_FUNC = 0x0020 // is function
)
// Storage Class.
const (
C_NULL = 0 // Symbol table entry marked for deletion
C_EXT = 2 // External symbol
C_STAT = 3 // Static symbol
C_BLOCK = 100 // Beginning or end of inner block
C_FCN = 101 // Beginning or end of function
C_FILE = 103 // Source file name and compiler information
C_HIDEXT = 107 // Unnamed external symbol
C_BINCL = 108 // Beginning of include file
C_EINCL = 109 // End of include file
C_WEAKEXT = 111 // Weak external symbol
C_DWARF = 112 // DWARF symbol
C_GSYM = 128 // Global variable
C_LSYM = 129 // Automatic variable allocated on stack
C_PSYM = 130 // Argument to subroutine allocated on stack
C_RSYM = 131 // Register variable
C_RPSYM = 132 // Argument to function or procedure stored in register
C_STSYM = 133 // Statically allocated symbol
C_BCOMM = 135 // Beginning of common block
C_ECOML = 136 // Local member of common block
C_ECOMM = 137 // End of common block
C_DECL = 140 // Declaration of object
C_ENTRY = 141 // Alternate entry
C_FUN = 142 // Function or procedure
C_BSTAT = 143 // Beginning of static block
C_ESTAT = 144 // End of static block
C_GTLS = 145 // Global thread-local variable
C_STTLS = 146 // Static thread-local variable
)
// File Auxiliary Entry
type AuxFile64 struct {
Xfname [8]byte // Name or offset inside string table
Xftype uint8 // Source file string type
Xauxtype uint8 // Type of auxiliary entry
}
// Function Auxiliary Entry
type AuxFcn32 struct {
Xexptr uint32 // File offset to exception table entry
Xfsize uint32 // Size of function in bytes
Xlnnoptr uint32 // File pointer to line number
Xendndx uint32 // Symbol table index of next entry
Xpad uint16 // Unused
}
type AuxFcn64 struct {
Xlnnoptr uint64 // File pointer to line number
Xfsize uint32 // Size of function in bytes
Xendndx uint32 // Symbol table index of next entry
Xpad uint8 // Unused
Xauxtype uint8 // Type of auxiliary entry
}
type AuxSect64 struct {
Xscnlen uint64 // section length
Xnreloc uint64 // Num RLDs
pad uint8
Xauxtype uint8 // Type of auxiliary entry
}
// csect Auxiliary Entry.
type AuxCSect32 struct {
Xscnlen int32 // Length or symbol table index
Xparmhash uint32 // Offset of parameter type-check string
Xsnhash uint16 // .typchk section number
Xsmtyp uint8 // Symbol alignment and type
Xsmclas uint8 // Storage-mapping class
Xstab uint32 // Reserved
Xsnstab uint16 // Reserved
}
type AuxCSect64 struct {
Xscnlenlo uint32 // Lower 4 bytes of length or symbol table index
Xparmhash uint32 // Offset of parameter type-check string
Xsnhash uint16 // .typchk section number
Xsmtyp uint8 // Symbol alignment and type
Xsmclas uint8 // Storage-mapping class
Xscnlenhi int32 // Upper 4 bytes of length or symbol table index
Xpad uint8 // Unused
Xauxtype uint8 // Type of auxiliary entry
}
// Auxiliary type
// const (
// _AUX_EXCEPT = 255
// _AUX_FCN = 254
// _AUX_SYM = 253
// _AUX_FILE = 252
// _AUX_CSECT = 251
// _AUX_SECT = 250
// )
// Symbol type field.
const (
XTY_ER = 0 // External reference
XTY_SD = 1 // Section definition
XTY_LD = 2 // Label definition
XTY_CM = 3 // Common csect definition
)
// Defines for File auxiliary definitions: x_ftype field of x_file
const (
XFT_FN = 0 // Source File Name
XFT_CT = 1 // Compile Time Stamp
XFT_CV = 2 // Compiler Version Number
XFT_CD = 128 // Compiler Defined Information
)
// Storage-mapping class.
const (
XMC_PR = 0 // Program code
XMC_RO = 1 // Read-only constant
XMC_DB = 2 // Debug dictionary table
XMC_TC = 3 // TOC entry
XMC_UA = 4 // Unclassified
XMC_RW = 5 // Read/Write data
XMC_GL = 6 // Global linkage
XMC_XO = 7 // Extended operation
XMC_SV = 8 // 32-bit supervisor call descriptor
XMC_BS = 9 // BSS class
XMC_DS = 10 // Function descriptor
XMC_UC = 11 // Unnamed FORTRAN common
XMC_TC0 = 15 // TOC anchor
XMC_TD = 16 // Scalar data entry in the TOC
XMC_SV64 = 17 // 64-bit supervisor call descriptor
XMC_SV3264 = 18 // Supervisor call descriptor for both 32-bit and 64-bit
XMC_TL = 20 // Read/Write thread-local data
XMC_UL = 21 // Read/Write thread-local data (.tbss)
XMC_TE = 22 // TOC entry
)
// Loader Header.
type LoaderHeader32 struct {
Lversion int32 // Loader section version number
Lnsyms int32 // Number of symbol table entries
Lnreloc int32 // Number of relocation table entries
Listlen uint32 // Length of import file ID string table
Lnimpid int32 // Number of import file IDs
Limpoff uint32 // Offset to start of import file IDs
Lstlen uint32 // Length of string table
Lstoff uint32 // Offset to start of string table
}
type LoaderHeader64 struct {
Lversion int32 // Loader section version number
Lnsyms int32 // Number of symbol table entries
Lnreloc int32 // Number of relocation table entries
Listlen uint32 // Length of import file ID string table
Lnimpid int32 // Number of import file IDs
Lstlen uint32 // Length of string table
Limpoff uint64 // Offset to start of import file IDs
Lstoff uint64 // Offset to start of string table
Lsymoff uint64 // Offset to start of symbol table
Lrldoff uint64 // Offset to start of relocation entries
}
const (
LDHDRSZ_32 = 32
LDHDRSZ_64 = 56
)
// Loader Symbol.
type LoaderSymbol32 struct {
Lname [8]byte // Symbol name or byte offset into string table
Lvalue uint32 // Address field
Lscnum int16 // Section number containing symbol
Lsmtype int8 // Symbol type, export, import flags
Lsmclas int8 // Symbol storage class
Lifile int32 // Import file ID; ordinal of import file IDs
Lparm uint32 // Parameter type-check field
}
type LoaderSymbol64 struct {
Lvalue uint64 // Address field
Loffset uint32 // Byte offset into string table of symbol name
Lscnum int16 // Section number containing symbol
Lsmtype int8 // Symbol type, export, import flags
Lsmclas int8 // Symbol storage class
Lifile int32 // Import file ID; ordinal of import file IDs
Lparm uint32 // Parameter type-check field
}
type Reloc32 struct {
Rvaddr uint32 // (virtual) address of reference
Rsymndx uint32 // Index into symbol table
Rsize uint8 // Sign and reloc bit len
Rtype uint8 // Toc relocation type
}
type Reloc64 struct {
Rvaddr uint64 // (virtual) address of reference
Rsymndx uint32 // Index into symbol table
Rsize uint8 // Sign and reloc bit len
Rtype uint8 // Toc relocation type
}
const (
R_POS = 0x00 // A(sym) Positive Relocation
R_NEG = 0x01 // -A(sym) Negative Relocation
R_REL = 0x02 // A(sym-*) Relative to self
R_TOC = 0x03 // A(sym-TOC) Relative to TOC
R_TRL = 0x12 // A(sym-TOC) TOC Relative indirect load.
R_TRLA = 0x13 // A(sym-TOC) TOC Rel load address. modifiable inst
R_GL = 0x05 // A(external TOC of sym) Global Linkage
R_TCL = 0x06 // A(local TOC of sym) Local object TOC address
R_RL = 0x0C // A(sym) Pos indirect load. modifiable instruction
R_RLA = 0x0D // A(sym) Pos Load Address. modifiable instruction
R_REF = 0x0F // AL0(sym) Non relocating ref. No garbage collect
R_BA = 0x08 // A(sym) Branch absolute. Cannot modify instruction
R_RBA = 0x18 // A(sym) Branch absolute. modifiable instruction
R_BR = 0x0A // A(sym-*) Branch rel to self. non modifiable
R_RBR = 0x1A // A(sym-*) Branch rel to self. modifiable instr
R_TLS = 0x20 // General-dynamic reference to TLS symbol
R_TLS_IE = 0x21 // Initial-exec reference to TLS symbol
R_TLS_LD = 0x22 // Local-dynamic reference to TLS symbol
R_TLS_LE = 0x23 // Local-exec reference to TLS symbol
R_TLSM = 0x24 // Module reference to TLS symbol
R_TLSML = 0x25 // Module reference to local (own) module
R_TOCU = 0x30 // Relative to TOC - high order bits
R_TOCL = 0x31 // Relative to TOC - low order bits
)
+19 -11
View File
@@ -5,12 +5,14 @@ import (
"context"
"debug/macho"
"debug/pe"
"encoding/binary"
"errors"
"fmt"
"io"
"regexp"
"runtime/debug"
"slices"
"strconv"
"strings"
"sync"
"time"
@@ -25,7 +27,6 @@ import (
"github.com/anchore/syft/syft/internal/unionreader"
"github.com/anchore/syft/syft/pkg"
"github.com/anchore/syft/syft/pkg/cataloger/generic"
"github.com/anchore/syft/syft/pkg/cataloger/golang/internal/xcoff"
)
const goArch = "GOARCH"
@@ -120,12 +121,19 @@ func (c *goBinaryCataloger) parseGoBinary(ctx context.Context, resolver file.Res
}
defer internal.CloseAndLogError(reader.ReadCloser, reader.RealPath)
mods, errs := scanFile(reader.Location, unionReader, c.symbolSelector.enabled())
mods, errs := scanFile(ctx, reader.Location, unionReader, c.symbolSelector.enabled())
// scanFile hands back the reconstruction for each binary that turned out to be packed. The readers
// below want it too, so ownership ends here and not inside the scan.
defer func() {
for _, mod := range mods {
closeUnpacked(mod.unpacked)
}
}()
var rels []artifact.Relationship
for _, mod := range mods {
var depPkgs []pkg.Package
mainPkg, depPkgs := c.buildGoPkgInfo(ctx, resolver, reader.Location, mod, mod.arch, unionReader)
mainPkg, depPkgs := c.buildGoPkgInfo(ctx, resolver, reader.Location, mod, mod.arch, seekerFor(mod.unpacked, unionReader))
if mainPkg != nil {
rels = createModuleRelationships(*mainPkg, depPkgs)
pkgs = append(pkgs, *mainPkg)
@@ -171,7 +179,7 @@ func moduleEqual(lhs, rhs *debug.Module) bool {
var emptyModule debug.Module
var moduleFromPartialPackageBuild = debug.Module{Path: "command-line-arguments"}
func (c *goBinaryCataloger) buildGoPkgInfo(ctx context.Context, resolver file.Resolver, location file.Location, mod *extendedBuildInfo, arch string, reader io.ReadSeekCloser) (*pkg.Package, []pkg.Package) {
func (c *goBinaryCataloger) buildGoPkgInfo(ctx context.Context, resolver file.Resolver, location file.Location, mod *extendedBuildInfo, arch string, reader io.ReadSeeker) (*pkg.Package, []pkg.Package) {
if mod == nil {
return nil, nil
}
@@ -240,7 +248,7 @@ func missingMainModule(mod *extendedBuildInfo) bool {
return mod.Main == moduleFromPartialPackageBuild
}
func (c *goBinaryCataloger) makeGoMainPackage(ctx context.Context, resolver file.Resolver, mod *extendedBuildInfo, arch string, location file.Location, reader io.ReadSeekCloser, symbols map[string][]string) pkg.Package {
func (c *goBinaryCataloger) makeGoMainPackage(ctx context.Context, resolver file.Resolver, mod *extendedBuildInfo, arch string, location file.Location, reader io.ReadSeeker, symbols map[string][]string) pkg.Package {
gbs := getBuildSettings(mod.Settings)
lics := c.licenseResolver.getLicenses(ctx, resolver, mod.Main.Path, mod.Main.Version)
gover, experiments := getExperimentsFromVersion(mod.GoVersion)
@@ -283,7 +291,7 @@ func (c *goBinaryCataloger) makeGoMainPackage(ctx context.Context, resolver file
// the only thing that seems to work is to just look for version strings following both \x00 and \x00.L for now
var semverPattern = regexp.MustCompile(`(\x00|\x{FFFD})(.L)?(?P<version>v?(\d+\.\d+\.\d+[-\w]*[+\w]*))\x00`)
func (c *goBinaryCataloger) findMainModuleVersion(metadata *pkg.GolangBinaryBuildinfoEntry, gbs pkg.KeyValues, reader io.ReadSeekCloser) string {
func (c *goBinaryCataloger) findMainModuleVersion(metadata *pkg.GolangBinaryBuildinfoEntry, gbs pkg.KeyValues, reader io.ReadSeeker) string {
vcsVersion, hasVersion := gbs.Get("vcs.revision")
timestamp, hasTimestamp := gbs.Get("vcs.time")
@@ -416,11 +424,11 @@ func getGOARCHFromBin(r io.ReaderAt) (string, error) {
}
arch = f.Cpu.String()
case bytes.HasPrefix(ident, []byte{0x01, 0xDF}) || bytes.HasPrefix(ident, []byte{0x01, 0xF7}):
f, err := xcoff.NewFile(r)
if err != nil {
return "", fmt.Errorf("unrecognized file format: %w", err)
}
arch = fmt.Sprintf("%d", f.TargetMachine)
// XCOFF's target machine *is* the magic we just matched on (0737 / 0767), so there is nothing to
// parse: a full XCOFF walk would read the string table, every symbol and every relocation only to
// hand back these two bytes. Reading them directly also keeps stripped binaries working, which a
// walk does not since it needs a symbol table to get that far.
arch = strconv.Itoa(int(binary.BigEndian.Uint16(ident[:2])))
default:
return "", errUnrecognizedFormat
}
@@ -108,24 +108,27 @@ func Test_getGOARCHFromBin(t *testing.T) {
},
{
name: "xcoff-32bit",
filepath: "internal/xcoff/testdata/gcc-ppc32-aix-dwarf2-exec",
filepath: "testdata/xcoff/gcc-ppc32-aix-dwarf2-exec",
expected: strconv.Itoa(0x1DF),
},
{
name: "xcoff-64bit",
filepath: "internal/xcoff/testdata/gcc-ppc64-aix-dwarf2-exec",
filepath: "testdata/xcoff/gcc-ppc64-aix-dwarf2-exec",
expected: strconv.Itoa(0x1F7),
},
}
for _, tt := range tests {
f, err := os.Open(tt.filepath)
require.NoError(t, err)
arch, err := getGOARCHFromBin(f)
require.NoError(t, err, "test name: %s", tt.name)
assert.Equal(t, tt.expected, arch)
}
t.Run(tt.name, func(t *testing.T) {
f, err := os.Open(tt.filepath)
require.NoError(t, err)
t.Cleanup(func() { _ = f.Close() })
arch, err := getGOARCHFromBin(f)
require.NoError(t, err)
assert.Equal(t, tt.expected, arch)
})
}
}
func TestBuildGoPkgInfo(t *testing.T) {
+241 -60
View File
@@ -1,7 +1,9 @@
package golang
import (
"context"
"debug/buildinfo"
"errors"
"fmt"
"io"
"runtime/debug"
@@ -21,10 +23,32 @@ type extendedBuildInfo struct {
cryptoSettings []string
arch string
symbols []binarySymbol
// unpacked is the reconstruction when this binary was UPX-packed, and nil when it was not. It is
// carried on the build info because the readers downstream of the scan want it too, since the packed
// bytes hold no readable version strings either. The caller of scanFile owns it and must Close it.
unpacked unpackedContents
}
// unpackedContents is the reconstruction of a UPX-packed binary, as everything downstream of the unpack
// uses it: bytes at offsets, the length it is willing to stand behind, and the temp file underneath to
// release when the scan is done with it. internal/spillbuf is what implements it, and naming that type in
// these signatures would say the readers care where the bytes live, which they do not.
//
// A nil value means the binary was not packed, and unpackUPX only ever returns a literal nil for that: a
// nil *spillbuf.Buffer widened into this interface is a non-nil interface holding a nil pointer, so it
// reads as present everywhere it is checked and then panics on first use. readerFor, seekerFor and
// closeUnpacked are the only places that check.
type unpackedContents interface {
io.ReaderAt
io.Closer
// Size is the contiguous run rebuilt from offset zero, which is all the reconstruction stands behind.
Size() int64
}
// scanFile scans file to try to report the Go and module versions.
func scanFile(location file.Location, reader unionreader.UnionReader, captureSymbols bool) ([]*extendedBuildInfo, error) {
func scanFile(ctx context.Context, location file.Location, reader unionreader.UnionReader, captureSymbols bool) ([]*extendedBuildInfo, error) {
// NOTE: multiple readers are returned to cover universal binaries, which are files
// with more than one binary
readers, errs := unionreader.GetReaders(reader)
@@ -35,53 +59,134 @@ func scanFile(location file.Location, reader unionreader.UnionReader, captureSym
var builds []*extendedBuildInfo
for _, r := range readers {
bi, err := getBuildInfo(r, location)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang buildinfo")
continue
build, err := scanReader(ctx, location, r, captureSymbols)
errs = unknown.Join(errs, err)
if build != nil {
builds = append(builds, build)
}
// it's possible the reader just isn't a go binary, in which case just skip it
if bi == nil {
continue
}
v, err := getCryptoInformation(r)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang version info")
// don't skip this build info.
// we can still catalog packages, even if we can't get the crypto information
errs = unknown.Appendf(errs, location, "unable to read golang version info: %w", err)
}
v = append(v, getNativeFIPSSettings(bi.Settings)...)
arch := getGOARCH(bi.Settings)
if arch == "" {
arch, err = getGOARCHFromBin(r)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang arch info")
// don't skip this build info.
// we can still catalog packages, even if we can't get the arch information
errs = unknown.Appendf(errs, location, "unable to read golang arch info: %w", err)
}
}
var symbols []binarySymbol
if captureSymbols {
symbols, err = getSymbols(r)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang symbol info")
// don't skip this build info.
// we can still catalog packages, even if we can't get the symbol information
errs = unknown.Appendf(errs, location, "unable to read golang symbol info: %w", err)
}
}
builds = append(builds, &extendedBuildInfo{BuildInfo: bi, cryptoSettings: v, arch: arch, symbols: symbols})
}
return builds, errs
}
// scanReader reports the build info, crypto settings, arch and symbols of a single binary. Everything is
// read from the same reader, which is the unpacked contents when the binary turns out to be UPX-packed:
// the packed bytes carry no readable pclntab or version strings either, so nothing downstream of the
// build info should be looking at them.
func scanReader(ctx context.Context, location file.Location, r io.ReaderAt, captureSymbols bool) (*extendedBuildInfo, error) {
var errs error
// unpacked is nil when there was nothing to unpack; readerFor turns that into the reader every parser
// below should use, so none of them has to ask whether this binary was packed.
// err can arrive alongside a usable bi: a partial reconstruction still carries build info, and the
// bytes it lost are a gap worth reporting rather than a reason to drop the binary.
unpacked, bi, err := readContentsAndBuildInfo(ctx, r)
// ownership of unpacked transfers to the returned extendedBuildInfo on success. Until then it is held
// here: the parsers below panic on malformed input (which is why getBuildInfo has a recover), and an
// unwind past this point would leave a reconstruction of up to maxUPXOriginalSize on disk for the rest
// of the scan: the temp root is swept, but not until the run ends.
ownContents := true
defer func() {
if ownContents {
closeUnpacked(unpacked)
}
}()
if err != nil {
if isCancelled(err) {
return nil, err
}
log.WithFields("file", location.RealPath, "error", err).Trace("unable to fully read golang buildinfo")
if reportableGap(err) {
// the build info is either missing or it is not, and the same err covers both: a partial
// reconstruction still hands back everything .go.buildinfo carried, and what it lost is the
// non-loadable tail the readers below want. Saying "unable to read golang buildinfo" there
// describes a failure that did not happen.
if bi != nil {
errs = unknown.Appendf(errs, location, "golang binary read incompletely: %w", err)
} else {
errs = unknown.Appendf(errs, location, "unable to read golang buildinfo: %w", err)
}
}
if bi == nil {
return nil, errs
}
}
// it's possible the reader just isn't a go binary, in which case just skip it
if bi == nil {
return nil, errs
}
// resolved once, here: every parser below reads the same bytes, and the choice of which reader that is
// belongs at the top rather than repeated at each call site
contents := readerFor(unpacked, r)
// everything below reads those contents through a parser that expands ELF sections, and each bounds its own
// reads: getSymbols opens the file with elfutil.NewFile, getCryptoInformation gates on
// elfutil.CheckAllSections. getBuildInfo's CheckSectionNameTable is not what makes these safe, since it
// bounds only the section elf.NewFile expands as it parses.
v, err := getCryptoInformation(contents)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang version info")
// don't skip this build info.
// we can still catalog packages, even if we can't get the crypto information
errs = unknown.Appendf(errs, location, "unable to read golang version info: %w", err)
}
v = append(v, getNativeFIPSSettings(bi.Settings)...)
arch := getGOARCH(bi.Settings)
if arch == "" {
arch, err = getGOARCHFromBin(contents)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang arch info")
// don't skip this build info.
// we can still catalog packages, even if we can't get the arch information
errs = unknown.Appendf(errs, location, "unable to read golang arch info: %w", err)
}
}
var symbols []binarySymbol
if captureSymbols {
symbols, err = getSymbols(contents)
if err != nil {
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang symbol info")
// don't skip this build info.
// we can still catalog packages, even if we can't get the symbol information
errs = unknown.Appendf(errs, location, "unable to read golang symbol info: %w", err)
}
}
ownContents = false
return &extendedBuildInfo{BuildInfo: bi, cryptoSettings: v, arch: arch, symbols: symbols, unpacked: unpacked}, errs
}
// reportableGap reports whether err is a gap this cataloger chose to leave, rather than a file it was
// never going to catalog. A packed file we failed to decode, a reconstruction that came up short, and a
// file we declined to expand (from either the ELF section bound or the UPX header bounds) all cost the
// SBOM something real. This cataloger runs on every executable in an image, so reporting anything else
// would attach an unknown to every corrupt, truncated or non-Go binary in it, which is the noise the
// quiet-by-default policy in upx.go exists to avoid.
//
// One real gap is deliberately left off this list: a UPX method we have not implemented. It is the same
// cost to the SBOM as a decode failure, but it fires on most packed binaries rather than on rare ones.
// See errUPXDecompress in upx.go for the reasoning.
func reportableGap(err error) bool {
return errors.Is(err, errUPXDecompress) ||
errors.Is(err, errUPXSizeRefused) ||
errors.Is(err, errUPXPartial) ||
errors.Is(err, elfutil.ErrDeclaredSizeExceeded)
}
func getCryptoInformation(reader io.ReaderAt) ([]string, error) {
// goversion opens the file with debug/elf itself and reads .symtab plus the string table it links.
// Those are expanded lazily, so the section-name table bound getBuildInfo already applied does not
// reach them. CheckAllSections is the decompression-bomb gate: a compressed section declares its own
// decompressed size and debug/elf allocates that much on open, so a 260KB ELF declaring a compressed
// .symtab drove 1.3GB of allocation before this was here.
if err := elfutil.CheckAllSections(reader); err != nil {
return nil, err
}
v, err := version.ReadExeFromReader(reader)
if err != nil {
return nil, err
@@ -123,6 +228,99 @@ func getNativeFIPSSettings(settings []debug.BuildSetting) []string {
return cryptoSettings
}
// readContentsAndBuildInfo returns the contents this binary should be read from, along with its build
// info. A UPX-packed binary is unpacked first and read from the reconstruction.
//
// If the reconstruction carries no build info, the bytes as they were found are tried before giving up.
// The "UPX!" magic is located by an unanchored scan over the first 8KB and every field behind it is
// attacker-controlled, so an ordinary Go binary carrying those four bytes can be made to produce a
// plausible header and a reconstruction of nothing. Without the retry that binary reports no packages at
// all, which turns a false positive here into a way to hide a dependency list.
//
// A gap in the unpack is returned even when build info came through, so it is reported rather than papered
// over: the readers after this one (crypto settings, arch, symbols) read the same contents, and the bytes
// a partial reconstruction lost are exactly the non-loadable tail those depend on. That only holds when
// the contents are a reconstruction: a gap read off the bytes as they were found is not a gap at all, and
// both places below that hand those bytes back drop it.
func readContentsAndBuildInfo(ctx context.Context, r io.ReaderAt) (unpackedContents, *debug.BuildInfo, error) {
unpacked, unpackErr := unpackUPX(ctx, r)
if isCancelled(unpackErr) {
return unpacked, nil, unpackErr
}
bi, err := getBuildInfo(readerFor(unpacked, r))
if err == nil && bi != nil {
if unpacked == nil {
// nothing was rebuilt, so nothing downstream is short: every reader below this one reads the
// same bytes getBuildInfo just read in full. unpackUPX returns the input as it was found
// alongside a refusal (the size bounds) or a failure (block 1 not decoding, or no temp dir),
// and readable build info in those bytes is itself the evidence the file was not really
// packed. A genuinely packed binary carries no readable .go.buildinfo in its packed bytes, so
// it never reaches here and its gap is still reported below.
return unpacked, bi, nil
}
return unpacked, bi, unpackErr
}
if unpacked != nil {
if foundBI, foundErr := getBuildInfo(r); foundErr == nil && foundBI != nil {
log.WithFields("error", err, "unpackError", unpackErr).
Trace("UPX reconstruction carried no build info, reading the binary as it was found")
closeUnpacked(unpacked)
// unpackErr is deliberately dropped. Every reader after this one reads the bytes as they were
// found, not the reconstruction, so nothing downstream is short: the header this binary
// carried was not describing real UPX output in the first place. Returning the gap here would
// put an unknown on a binary the cataloger read completely.
return nil, foundBI, nil
}
}
// a packed file we could not unpack is the more specific gap, so it wins over whatever buildinfo made
// of the bytes it was handed
if unpackErr != nil {
return unpacked, nil, unpackErr
}
return unpacked, bi, err
}
// isCancelled reports whether err means the scan was called off, which is a reason to stop rather than a
// gap in the SBOM.
func isCancelled(err error) bool {
return errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded)
}
// readerFor returns where a scanned binary should be read from: the reconstruction when it was packed,
// the input as it was found otherwise.
//
// The nil check lives here, in one place, on purpose: every reader below the scan goes through here, so no
// parser has to ask whether this binary was packed.
func readerFor(unpacked unpackedContents, found io.ReaderAt) io.ReaderAt {
if unpacked == nil {
return found
}
return unpacked
}
// seekerFor is readerFor for the readers after the scan, which seek. found belongs to the caller and is
// reused for every later module, so this deliberately hands back no Closer.
func seekerFor(unpacked unpackedContents, found io.ReadSeeker) io.ReadSeeker {
if unpacked == nil {
return found
}
// the SectionReader snapshots Size() here, which is correct: the reconstruction is complete by the time
// anything seeks it. It also gives each caller its own cursor over the same contents.
return io.NewSectionReader(unpacked, 0, unpacked.Size())
}
// closeUnpacked releases a reconstruction, if the binary turned out to be packed at all. Nothing to unpack
// is the common case rather than an edge one, so every owner of an unpackedContents closes through here.
func closeUnpacked(unpacked unpackedContents) {
if unpacked == nil {
return
}
_ = unpacked.Close()
}
// readBuildInfo bounds the reader before handing it to debug/buildinfo, which opens ELF files with
// debug/elf itself rather than through elfutil. elf.NewFile expands the section-name string table as it
// parses, so an unbounded read here is reachable no matter how little of the file buildinfo goes on to
@@ -134,7 +332,7 @@ func readBuildInfo(r io.ReaderAt) (*debug.BuildInfo, error) {
return buildinfo.Read(r)
}
func getBuildInfo(r io.ReaderAt, location file.Location) (bi *debug.BuildInfo, err error) {
func getBuildInfo(r io.ReaderAt) (bi *debug.BuildInfo, err error) {
defer func() {
if r := recover(); r != nil {
// this can happen in cases where a malformed binary is passed in can be initially parsed, but not
@@ -144,24 +342,7 @@ func getBuildInfo(r io.ReaderAt, location file.Location) (bi *debug.BuildInfo, e
}
}()
// try to read buildinfo from the binary directly
bi, err = readBuildInfo(r)
if err == nil {
return bi, nil
}
// if direct read fails and this looks like a UPX-compressed binary,
// try to decompress and read the buildinfo from the decompressed data
if isUPXCompressed(r) {
log.WithFields("path", location.RealPath).Trace("detected UPX-compressed Go binary, attempting decompression to read the build info")
decompressed, decompErr := decompressUPX(r)
if decompErr == nil {
bi, err = readBuildInfo(decompressed)
if err == nil {
return bi, nil
}
}
}
// note: the stdlib does not export the error we need to check for
if err != nil {
@@ -6,13 +6,15 @@ import (
"debug/buildinfo"
"debug/elf"
"encoding/binary"
"os"
"runtime"
"testing"
"github.com/kastenhq/goversion/version"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"github.com/anchore/syft/syft/file"
"github.com/anchore/syft/syft/internal/elfutil"
)
// Test_getBuildInfo_compressedSectionBomb covers the reason readBuildInfo exists: debug/buildinfo opens
@@ -33,14 +35,122 @@ func Test_getBuildInfo_compressedSectionBomb(t *testing.T) {
assert.Greater(t, unguarded, uint64(declared), "fixture did not actually deliver the declared bytes")
guarded := measureAlloc(t, func() {
_, err := getBuildInfo(bytes.NewReader(bomb), file.NewLocation("bomb"))
_, err := getBuildInfo(bytes.NewReader(bomb))
require.Error(t, err)
assert.Contains(t, err.Error(), "over the")
assert.ErrorIs(t, err, elfutil.ErrDeclaredSizeExceeded)
})
assert.Less(t, guarded, uint64(32<<20), "getBuildInfo allocated far more than the input warrants")
t.Logf("unguarded allocated %d bytes, guarded allocated %d bytes", unguarded, guarded)
}
// Test_getCryptoInformation_compressedSymtabBomb covers the second unbounded door into debug/elf:
// goversion opens the file itself and reads .symtab plus the string table it links, and those are
// expanded lazily, so the section-name table bound getBuildInfo applies never reaches them.
func Test_getCryptoInformation_compressedSymtabBomb(t *testing.T) {
const declared = 256 << 20 // comfortably over elfutil's bound, small enough to allocate in a test
bomb := elfWithCompressedSymtab(t, declared)
t.Logf("%d byte fixture declares a %d byte symbol table", len(bomb), declared)
// the unguarded path is the thing being defended against: prove the fixture really is a bomb
unguarded := measureAlloc(t, func() {
_, err := version.ReadExeFromReader(bytes.NewReader(bomb))
t.Logf("goversion err: %v", err)
})
assert.Greater(t, unguarded, uint64(declared), "fixture did not actually deliver the declared bytes")
guarded := measureAlloc(t, func() {
_, err := getCryptoInformation(bytes.NewReader(bomb))
require.ErrorIs(t, err, elfutil.ErrDeclaredSizeExceeded)
})
assert.Less(t, guarded, uint64(32<<20), "getCryptoInformation allocated far more than the input warrants")
t.Logf("unguarded allocated %d bytes, guarded allocated %d bytes", unguarded, guarded)
}
// Test_getCryptoInformation_passesThroughNonELF keeps the gate from becoming a container filter: goversion
// reads PE and Mach-O too, and CheckAllSections has to leave those to it.
func Test_getCryptoInformation_passesThroughNonELF(t *testing.T) {
runMakeTarget(t, "archs")
for _, name := range []string{"hello-win-amd64", "hello-mach-o-arm64"} {
t.Run(name, func(t *testing.T) {
f, err := os.Open("testdata/archs/binaries/" + name)
require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, f.Close()) })
_, err = getCryptoInformation(f)
require.NoError(t, err)
})
}
}
// elfWithCompressedSymtab builds a minimal ELF64 whose .symtab is SHF_COMPRESSED, declaring `declared`
// decompressed bytes and genuinely delivering them. The name table is left uncompressed so the file gets
// past the bound getBuildInfo already applies, which is the point: this is the section that bound misses.
func elfWithCompressedSymtab(t *testing.T, declared uint64) []byte {
t.Helper()
payload := make([]byte, declared)
var compressed bytes.Buffer
zw := zlib.NewWriter(&compressed)
_, err := zw.Write(payload)
require.NoError(t, err)
require.NoError(t, zw.Close())
names := []byte("\x00.shstrtab\x00.symtab\x00.strtab\x00")
nameOf := func(s string) uint32 {
i := bytes.Index(names, []byte("\x00"+s+"\x00"))
require.GreaterOrEqual(t, i, 0)
return uint32(i + 1)
}
ehsize := uint64(binary.Size(elf.Header64{}))
shentsize := uint64(binary.Size(elf.Section64{}))
chdrsize := uint64(binary.Size(elf.Chdr64{}))
shoff := ehsize
namesOff := shoff + 4*shentsize
symOff := namesOff + uint64(len(names))
symSize := chdrsize + uint64(compressed.Len())
strOff := symOff + symSize
var ident [16]byte
copy(ident[:], elf.ELFMAG)
ident[elf.EI_CLASS] = byte(elf.ELFCLASS64)
ident[elf.EI_DATA] = byte(elf.ELFDATA2LSB)
ident[elf.EI_VERSION] = byte(elf.EV_CURRENT)
buf := &bytes.Buffer{}
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Header64{
Ident: ident, Type: uint16(elf.ET_REL), Machine: uint16(elf.EM_X86_64),
Version: uint32(elf.EV_CURRENT), Shoff: shoff, Ehsize: uint16(ehsize),
Shentsize: uint16(shentsize), Shnum: 4, Shstrndx: 1,
}))
// the null section
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{}))
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{
Name: nameOf(".shstrtab"), Type: uint32(elf.SHT_STRTAB), Off: namesOff,
Size: uint64(len(names)), Addralign: 1,
}))
// the bomb: sh_size covers the whole zlib stream, ch_size is what debug/elf expands to
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{
Name: nameOf(".symtab"), Type: uint32(elf.SHT_SYMTAB), Flags: uint64(elf.SHF_COMPRESSED),
Off: symOff, Size: symSize, Link: 3, Entsize: 24, Addralign: 1,
}))
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{
Name: nameOf(".strtab"), Type: uint32(elf.SHT_STRTAB), Off: strOff, Size: 1, Addralign: 1,
}))
buf.Write(names)
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Chdr64{
Type: uint32(elf.COMPRESS_ZLIB), Size: declared, Addralign: 1,
}))
buf.Write(compressed.Bytes())
buf.Write([]byte{0})
return buf.Bytes()
}
// measureAlloc reports the bytes allocated while fn ran. TotalAlloc is process-wide, so a test using this
// must not call t.Parallel: another test's allocations would land in the measurement.
func measureAlloc(t *testing.T, fn func()) uint64 {
t.Helper()
var before, after runtime.MemStats
@@ -51,14 +161,33 @@ func measureAlloc(t *testing.T, fn func()) uint64 {
return after.TotalAlloc - before.TotalAlloc
}
// elfWithCompressedNameTable builds a minimal ELF64 whose only real section is a SHF_COMPRESSED
// .shstrtab declaring `declared` decompressed bytes and genuinely delivering them.
// elfWithCompressedNameTable builds a minimal ELF64 whose only real section is a SHF_COMPRESSED .shstrtab
// declaring `declared` decompressed bytes and genuinely delivering them. Only the test that runs the
// unguarded path needs delivery; use elfDeclaringNameTable everywhere else, since building this one costs
// `declared` bytes of allocation plus a zlib compress of them.
func elfWithCompressedNameTable(t *testing.T, declared uint64) []byte {
t.Helper()
return elfNameTableFixture(t, declared, true)
}
// elfDeclaringNameTable declares `declared` bytes without delivering them. CheckSectionNameTable reads the
// declared size out of the compression header and refuses before decompressing anything, so a test of the
// bound itself never needs the stream to be real.
func elfDeclaringNameTable(t *testing.T, declared uint64) []byte {
t.Helper()
return elfNameTableFixture(t, declared, false)
}
func elfNameTableFixture(t *testing.T, declared uint64, deliver bool) []byte {
t.Helper()
// the decompressed name table only has to start with the section names; the rest is padding that
// exists purely to make the declared size real
payload := make([]byte, declared)
size := declared
if !deliver {
size = 64
}
payload := make([]byte, size)
copy(payload, "\x00.shstrtab\x00")
var compressed bytes.Buffer
@@ -8,8 +8,6 @@ import (
"github.com/kastenhq/goversion/version"
"github.com/stretchr/testify/assert"
"github.com/anchore/syft/syft/file"
)
func Test_getBuildInfo(t *testing.T) {
@@ -33,7 +31,7 @@ func Test_getBuildInfo(t *testing.T) {
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
gotBi, err := getBuildInfo(tt.args.r, file.Location{})
gotBi, err := getBuildInfo(tt.args.r)
if !tt.wantErr(t, err, fmt.Sprintf("getBuildInfo(%v)", tt.args.r)) {
return
}
@@ -0,0 +1,96 @@
package golang
import (
"bytes"
"context"
"os"
"path/filepath"
"runtime/debug"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"github.com/anchore/syft/internal/spillbuf"
"github.com/anchore/syft/internal/tmpdir"
"github.com/anchore/syft/syft/file"
"github.com/anchore/syft/syft/internal/fileresolver"
"github.com/anchore/syft/syft/internal/unionreader"
)
// TestScanFile_UnpackedBinaryReadsTheFileItWasGiven is the regression test for the shape of bug that
// modelling "not packed" as nil invites. The reader was once an io.ReadSeekCloser holding a nil pointer,
// which is a non-nil interface, so every ordinary Go binary took the packed branch: it seeked a nil file
// for its version and nil-dereferenced on close. The panic was recovered by the task executor, so the
// only symptom was the go-binary cataloger reporting nothing at all.
//
// The nil is now the signal rather than a hazard: unpacked is nil when there was nothing to unpack, and
// readerFor, seekerFor and closeUnpacked are the only places that resolve it.
//
// Docker-gated like everything else that needs a real Go binary; the point is that a plain binary reads
// from the file it was handed, which needs a real one.
func TestScanFile_UnpackedBinaryReadsTheFileItWasGiven(t *testing.T) {
runMakeTarget(t, "archs")
f, err := os.Open(filepath.Join("testdata", "archs", "binaries", "hello-linux-arm"))
require.NoError(t, err)
t.Cleanup(func() { _ = f.Close() })
ur, err := unionreader.GetUnionReader(f)
require.NoError(t, err)
builds, _ := scanFile(context.Background(), file.NewLocation("hello-linux-arm"), ur, false)
require.NotEmpty(t, builds, "a plain Go binary must still produce build info")
for _, b := range builds {
assert.Nil(t, b.unpacked, "an unpacked binary must not have left a reconstruction behind")
assert.Same(t, ur, seekerFor(b.unpacked, ur), "the readers after the scan must get the file itself")
assert.NotPanics(t, func() { closeUnpacked(b.unpacked); closeUnpacked(b.unpacked) },
"closing nothing has to be safe and idempotent, since a defer over every build calls it")
}
}
// TestParseGoBinary_PlainBinaryStillYieldsPackages is the same regression one layer out, at the entry
// point the cataloger actually calls. The typed-nil panic fired from a defer here and was swallowed by
// the task executor, so the failure mode is an empty result rather than an error.
func TestParseGoBinary_PlainBinaryStillYieldsPackages(t *testing.T) {
runMakeTarget(t, "archs")
f, err := os.Open(filepath.Join("testdata", "archs", "binaries", "hello-linux-arm"))
require.NoError(t, err)
t.Cleanup(func() { _ = f.Close() })
c := newGoBinaryCataloger(DefaultCatalogerConfig())
pkgs, _, err := c.parseGoBinary(context.Background(), fileresolver.Empty{}, nil,
file.NewLocationReadCloser(file.NewLocation("hello-linux-arm"), f))
require.NoError(t, err)
assert.NotEmpty(t, pkgs, "an ordinary Go binary must still produce packages")
}
// TestMakeGoMainPackage_VersionComesFromTheUnpackedContents pins the reader threading: for a packed
// binary the version scan has to read the reconstruction, since the packed bytes carry no readable
// version string. Nothing else covers this, because the from-contents scan is off by default and the
// Docker fixture gets its version from ldflags before the scan is reached.
func TestMakeGoMainPackage_VersionComesFromTheUnpackedContents(t *testing.T) {
// the version pattern wants the string NUL-delimited
unpacked := spillbuf.New(tmpdir.FromPath(t.TempDir()))
t.Cleanup(func() { _ = unpacked.Close() })
_, err := unpacked.WriteAt([]byte("\x00v9.9.9\x00"), 0)
require.NoError(t, err)
// what the packed file on disk would have said, which must not be what comes out
packed := &nopReadSeekCloser{bytes.NewReader([]byte("\x00v1.1.1\x00"))}
c := &goBinaryCataloger{
licenseResolver: newGoLicenseResolver("", CatalogerConfig{}),
mainModuleVersion: MainModuleVersionConfig{FromContents: true},
}
mod := &extendedBuildInfo{
BuildInfo: &debug.BuildInfo{Main: debug.Module{Path: "github.com/anchore/syft", Version: devel}},
unpacked: unpacked,
}
got := c.makeGoMainPackage(context.Background(), fileresolver.Empty{}, mod, "amd64",
file.NewLocation("packed"), seekerFor(mod.unpacked, packed), nil)
assert.Equal(t, "v9.9.9", got.Version, "the version must come from the reconstruction, not the packed bytes")
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,759 @@
package golang
import (
"bytes"
"context"
"debug/elf"
"encoding/binary"
"fmt"
"io"
"os"
"path/filepath"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
intFile "github.com/anchore/syft/internal/file"
"github.com/anchore/syft/internal/tmpdir"
"github.com/anchore/syft/internal/unknown"
"github.com/anchore/syft/syft/file"
"github.com/anchore/syft/syft/internal/elfutil"
"github.com/anchore/syft/syft/internal/unionreader"
)
// the reconstruction's storage, its Close semantics and the contiguous-prefix rule it reports as Size all
// belong to internal/spillbuf now, and are tested there. What is left in this file is what the golang
// cataloger does with them.
// sparseFixture is a UPX file whose first block carries an ELF header declaring a second PT_LOAD segment
// at farOffset, so a tiny third block is placed there. declaredShstrtab, when non-zero, makes that ELF
// header also declare a section name table of that size sitting inside the gap, which is what turns the
// reader's length into an allocation for anything that parses it.
func sparseFixture(t *testing.T, declared uint32, farOffset, declaredShstrtab uint64, pad int) []byte {
t.Helper()
phdrs := make([]byte, 112)
binary.LittleEndian.PutUint32(phdrs[0:4], 1) // phdr[0] PT_LOAD
binary.LittleEndian.PutUint64(phdrs[8:16], 0) // p_offset 0
binary.LittleEndian.PutUint32(phdrs[56:60], 1)
binary.LittleEndian.PutUint64(phdrs[64:72], farOffset) // phdr[1] p_offset, out at the declared end
hdr := make([]byte, 64)
copy(hdr, []byte{0x7f, 'E', 'L', 'F'})
hdr[4], hdr[5], hdr[6] = 2, 1, 1 // ELFCLASS64, little endian, EV_CURRENT
binary.LittleEndian.PutUint16(hdr[0x10:0x12], uint16(elf.ET_EXEC))
binary.LittleEndian.PutUint16(hdr[0x12:0x14], uint16(elf.EM_X86_64))
binary.LittleEndian.PutUint32(hdr[0x14:0x18], 1)
binary.LittleEndian.PutUint64(hdr[0x20:0x28], 64) // e_phoff
binary.LittleEndian.PutUint16(hdr[0x36:0x38], 56) // e_phentsize
binary.LittleEndian.PutUint16(hdr[0x38:0x3a], 2) // e_phnum
block1 := append(hdr, phdrs...)
if declaredShstrtab > 0 {
sh := make([]byte, 128) // two section headers; index 1 is the name table
binary.LittleEndian.PutUint32(sh[64+4:64+8], uint32(elf.SHT_STRTAB))
binary.LittleEndian.PutUint64(sh[64+24:64+32], farOffset/2) // sh_offset, inside the gap
binary.LittleEndian.PutUint64(sh[64+32:64+40], declaredShstrtab)
binary.LittleEndian.PutUint64(block1[0x28:0x30], 64+112) // e_shoff
binary.LittleEndian.PutUint16(block1[0x3a:0x3c], 64) // e_shentsize
binary.LittleEndian.PutUint16(block1[0x3c:0x3e], 2) // e_shnum
binary.LittleEndian.PutUint16(block1[0x3e:0x40], 1) // e_shstrndx
block1 = append(block1, sh...)
}
data := buildUPXHeader(declared, 4096)
data = append(data, blockFor(t, block1)...)
data = append(data, blockFor(t, bytes.Repeat([]byte("B"), 32))...) // sequential
data = append(data, blockFor(t, bytes.Repeat([]byte("C"), 64))...) // -> ptLoadOffsets[1]
data = append(data, make([]byte, 12)...) // end marker
return padTo(data, pad)
}
// TestDecompressUPX_ReaderLengthIsWhatWasRebuilt covers the amplification that survives moving the
// reconstruction from the heap to a temp file.
//
// A block's destination comes from a PT_LOAD p_offset in the first block, bounded only by p_filesize, so a
// 64 byte block parked at p_filesize-64 must not make the reconstruction report the full declared size
// while holding only a few hundred real bytes. A sparse file hands its holes back as zeros it never stored,
// so every parser downstream that sizes against "what this reader will deliver" -- which is the whole
// argument for writing to disk rather than the heap -- would otherwise allocate against a number the input
// never paid for.
func TestDecompressUPX_ReaderLengthIsWhatWasRebuilt(t *testing.T) {
const declared = 64 << 20
const input = 300_000
data := sparseFixture(t, declared, declared-64, 0, input)
out, err := unpack(t, data)
// the hole is exactly what makes this partial: the blocks past it are given up, and that is reported
// rather than logged, since a caller reading this reconstruction is missing bytes the header declared
require.ErrorIs(t, err, errUPXPartial)
require.NotNil(t, out, "a short reconstruction is still the contents to read from")
// the ELF header block plus the two small blocks, and nothing for the hole
assert.Less(t, out.Size(), int64(4096),
"the reader must be sized by the bytes actually rebuilt, not by where a block was parked")
_, err = out.ReadAt(make([]byte, 512), int64(declared)/2)
assert.ErrorIs(t, err, io.EOF, "the hole must not read back as free zeros")
}
func TestDecompressUPX_SparsePlacementIsNotAnAllocationKnob(t *testing.T) {
// the same fixture, with the reconstructed ELF also declaring a 60MB section name table sitting in the
// hole. This is the end of the exploit chain: debug/elf grows its read as the reads succeed, and
// against a hole they all succeed.
const declared = 64 << 20
const input = 300_000
data := sparseFixture(t, declared, declared-64, 60<<20, input)
out, err := unpack(t, data)
require.ErrorIs(t, err, errUPXPartial)
require.NotNil(t, out)
allocated := measureAlloc(t, func() {
_, _ = getBuildInfo(out)
})
// debug/elf's own first chunk is ~10MB regardless of what a file declares, so the bound to assert is
// that the declared 60MB is not reachable, not that this is free
assert.Less(t, allocated, uint64(16*intFile.MB),
"a 300KB input must not drive tens of MB of allocation through the reconstruction")
}
func TestReadPTLoadOffsets_ShortReadIsNotTrusted(t *testing.T) {
// on a short read the tail of the buffer is zeros the file never provided, and a p_offset of zero
// parsed out of it would place a later block over the ELF header.
phdrs := make([]byte, 112)
binary.LittleEndian.PutUint32(phdrs[0:4], 1)
binary.LittleEndian.PutUint64(phdrs[8:16], 0x1000)
binary.LittleEndian.PutUint32(phdrs[56:60], 1)
binary.LittleEndian.PutUint64(phdrs[64:72], 0x2000)
full := buildELF64(64, 56, 2, phdrs)
assert.Equal(t, []uint64{0x1000, 0x2000}, readPTLoadOffsets(bytes.NewReader(full), 0, uint32(len(full))),
"the whole header is readable, so both segments come back")
// the same header with the second program header cut off. The block claims it is there; the file is
// not that long.
truncated := full[:64+56+8]
got := readPTLoadOffsets(bytes.NewReader(truncated), 0, uint32(len(full)))
assert.Equal(t, []uint64{0x1000}, got,
"only the segment actually present may be trusted; a zero read out of the gap is not an offset")
}
func TestDecompressUPX_StoredBlockIsPlaced(t *testing.T) {
// method 0 is an extent UPX could not compress and wrote verbatim, and real output carries a few of
// them: the alignment padding between PT_LOAD segments is only a handful of bytes. Nothing covered
// this path outside the Docker-backed fixture.
payload := bytes.Repeat([]byte("S"), 48)
stored := make([]byte, 12)
binary.LittleEndian.PutUint32(stored[0:4], uint32(len(payload))) // sz_unc
binary.LittleEndian.PutUint32(stored[4:8], uint32(len(payload))) // sz_cpr, equal for a stored block
stored[8] = upxMethodStored
data := buildUPXHeader(4096, 4096)
data = append(data, stored...)
data = append(data, payload...)
data = append(data, make([]byte, 12)...)
data = padTo(data, 512)
out, err := unpack(t, data)
// the fixture declares 4096 and delivers one 48 byte extent, so the chain is short by construction and
// the partial goes with it. Real output rebuilds p_filesize exactly; what is under test is placement.
require.ErrorIs(t, err, errUPXPartial)
assert.Equal(t, payload, readAll(t, out), "a stored block is copied through as-is")
}
func TestDecompressUPX_StoredBlockWithMismatchedSizesIsQuiet(t *testing.T) {
// a copy is only meaningful when the two sizes agree. A mismatch is how a stray b_info-shaped run of
// bytes looks, so it ends the chain rather than failing the file.
stored := make([]byte, 12)
binary.LittleEndian.PutUint32(stored[0:4], 64) // sz_unc
binary.LittleEndian.PutUint32(stored[4:8], 32) // sz_cpr, disagrees
stored[8] = upxMethodStored
data := padTo(append(buildUPXHeader(4096, 4096), stored...), 512)
_, err := unpack(t, data)
require.Error(t, err)
// a stored block whose sizes disagree is not a block we can read, which on the first block is the
// unsupported-method refusal. Either way it must stay clear of errUPXDecompress.
assert.ErrorIs(t, err, errUnsupportedUPXMethod, "a copy whose sizes disagree is not a copy")
assert.NotErrorIs(t, err, errUPXDecompress, "not a packed binary, so it stays quiet")
}
// TestScanReader_ReportingIsNarrow pins which gaps this cataloger claims. It runs against every
// executable in an image, so reporting each file debug/buildinfo cannot parse would attach an unknown to
// most files in a typical one. Only a bound syft itself chose to enforce, and a packed file it could not
// unpack, are real gaps.
func TestScanReader_ReportingIsNarrow(t *testing.T) {
tests := []struct {
name string
data []byte
report bool
}{
{
name: "a corrupt ELF is not this cataloger's gap",
data: append([]byte{0x7f, 'E', 'L', 'F', 2, 1, 1}, bytes.Repeat([]byte{0xAB}, 512)...),
},
{
name: "neither is a file that is not an executable at all",
data: bytes.Repeat([]byte("not a binary"), 64),
},
{
name: "nor a truncated PE",
data: append([]byte("MZ\x90\x00"), bytes.Repeat([]byte{0}, 512)...),
},
{
name: "a section name table syft declined to expand is",
data: elfDeclaringNameTable(t, 256<<20),
// this is elfutil.ErrDeclaredSizeExceeded: syft chose not to expand it, so the packages behind
// it really are missing from the SBOM because of a decision made here
report: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
loc := file.NewLocation("subject")
_, err := scanReader(context.Background(), loc, bytes.NewReader(tt.data), false)
if !tt.report {
assert.NoError(t, err, "this cataloger sees every executable in an image; it must stay quiet here")
return
}
require.Error(t, err)
var coordErr *unknown.CoordinateError
require.ErrorAs(t, err, &coordErr, "a reported gap has to carry the location it belongs to")
assert.Equal(t, loc.Coordinates, coordErr.Coordinates)
assert.ErrorIs(t, coordErr.Reason, elfutil.ErrDeclaredSizeExceeded,
"the reported gap has to carry the sentinel the split keys on, not just a message")
})
}
}
func TestScanReader_RefusedExpansionIsReportedAsSuch(t *testing.T) {
// the sentinel is what lets the narrow reporting above tell "syft declined to expand this" apart from
// "this file is broken", so the two must not collapse into one string match
_, err := readBuildInfo(bytes.NewReader(elfDeclaringNameTable(t, 256<<20)))
require.Error(t, err)
assert.ErrorIs(t, err, elfutil.ErrDeclaredSizeExceeded)
}
// TestReadContentsAndBuildInfo_FakeUPXHeaderCannotHideAGoBinary covers the evasion vector the unpacking
// opened up. The "UPX!" magic is found by an unanchored substring scan over the first 8KB and every field
// behind it is attacker-controlled, so an ordinary Go binary can be made to produce a plausible header and
// a reconstruction of nothing. Reading from that reconstruction unconditionally means the binary reports
// no packages at all, which turns a false positive into a way to hide a dependency list.
// goELF64Fixture returns a real, unpacked Go binary that clears the ELF64 little-endian container gate in
// parseUPXInfo. hello-linux-arm is ELFCLASS32 and is rejected before the UPX magic scan runs, so a test
// that splices a header into it exercises nothing.
func goELF64Fixture(t *testing.T) []byte {
t.Helper()
runMakeTarget(t, "archs")
original, err := os.ReadFile(filepath.Join("testdata", "archs", "binaries", "hello-linux-ppc64le"))
require.NoError(t, err)
require.Equal(t, byte(2), original[4], "the container gate only accepts ELF64")
require.Equal(t, byte(1), original[5], "the container gate only accepts little-endian")
return original
}
// spliceFakeUPXHeader writes a plausible UPX header plus one decodable block into padding inside the
// magic scan window, leaving the ELF itself intact. The offset is found rather than hardcoded: it has to
// sit past the section header table (writing over that corrupts the binary and the test then proves
// nothing) and inside upxMagicScanWindow (past it the magic is never seen).
func spliceFakeUPXHeader(t *testing.T, original []byte) []byte {
t.Helper()
return spliceUPXChain(t, original, 64<<10, blockFor(t, []byte("junk")))
}
// spliceUPXChain is spliceFakeUPXHeader with a caller-supplied p_filesize and b_info chain, so a test can
// choose whether the fake header lands on a clean short reconstruction, one that gives something up, or a
// refusal by the size bounds with nothing unpacked at all.
func spliceUPXChain(t *testing.T, original []byte, originalSize uint32, chain []byte) []byte {
t.Helper()
header := buildUPXHeader(originalSize, 4096)[64:] // drop the stub, the real ELF header is already here
block := chain
need := len(header) + len(block)
shoff := binary.LittleEndian.Uint64(original[0x28:0x30])
shentsize := binary.LittleEndian.Uint16(original[0x3a:0x3c])
shnum := binary.LittleEndian.Uint16(original[0x3c:0x3e])
after := int(shoff) + int(shentsize)*int(shnum)
at := -1
for i := after; i+need <= upxMagicScanWindow && i+need <= len(original); i++ {
if bytes.Equal(original[i:i+need], make([]byte, need)) {
at = i
break
}
}
require.GreaterOrEqual(t, at, 0, "no padding in the scan window big enough to splice a header into")
spiked := append([]byte(nil), original...)
copy(spiked[at:], header)
copy(spiked[at+len(header):], block)
// the splice must not have broken the binary, or the fallback below would be covering for a corrupt
// ELF rather than for a fake header. A claim the size bounds refuse is a deliberate caller choice, so
// only assert the header parses when it was meant to.
if originalSize <= maxUPXOriginalSize {
require.NotNil(t, mustParseUPX(t, spiked), "the spliced header has to actually parse as UPX")
}
return spiked
}
// mustParseUPX reports the UPX header the scan finds in data, proving the splice is what the parser sees.
func mustParseUPX(t *testing.T, data []byte) *upxInfo {
t.Helper()
r := bytes.NewReader(data)
info, err := parseUPXInfo(r, int64(len(data)))
require.NoError(t, err)
return info
}
func TestScanFile_FakeUPXHeaderStillYieldsPackages(t *testing.T) {
// the same evasion one layer out, at the entry point the cataloger calls. The header has to actually be
// spliced: run this against the untouched binary and it passes with the fallback deleted.
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
original := goELF64Fixture(t)
spiked := spliceFakeUPXHeader(t, original)
ur, err := unionreader.GetUnionReader(io.NopCloser(bytes.NewReader(spiked)))
require.NoError(t, err)
builds, _ := scanFile(ctx, file.NewLocation("hello-linux-ppc64le"), ur, false)
require.NotEmpty(t, builds, "a Go binary must not be hidden by four bytes of fake magic")
for _, b := range builds {
assert.Nil(t, b.unpacked,
"the reconstruction carried nothing, so the scan has to fall back to the bytes as they were found")
closeUnpacked(b.unpacked)
}
}
// TestScanReader_CancellationStopsTheScan pins the direction a swallowed ctx.Err() got wrong. Cancellation
// is not a gap in the SBOM, so it is not reported as an unknown, but it does mean stop: returning "nothing
// to unpack, no error" sent the scan on to parse the packed bytes and build packages out of them after the
// caller had already called it off.
func TestScanReader_CancellationStopsTheScan(t *testing.T) {
// enough blocks that the loop checks ctx at least once
blocks := make([][]byte, 64)
for i := range blocks {
blocks[i] = bytes.Repeat([]byte("A"), 1024)
}
fixture := buildUPXFile(t, 64*1024, 4096, blocks, nil)
ctx, cancel := context.WithCancel(tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir())))
cancel()
build, err := scanReader(ctx, file.NewLocation("/cancelled"), bytes.NewReader(fixture), false)
assert.Nil(t, build)
require.Error(t, err, "a cancelled scan has to stop rather than fall through to the packed bytes")
assert.ErrorIs(t, err, context.Canceled)
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
assert.Empty(t, coordErrs, "cancellation is not a gap in the SBOM")
}
// TestScanReader_SizeRefusalIsReported covers the reporting split for a header that really is UPX and
// claims more than the bounds allow. That is a file we declined to expand, which is the same category as
// elfutil.ErrDeclaredSizeExceeded and gets the same unknown. The ratio clears a measured 209x worst case at
// 256, so a legitimate binary crossing it must not vanish from the SBOM silently.
func TestScanReader_SizeRefusalIsReported(t *testing.T) {
// a 4KB input claiming 512MB: pays no ratio and clears no ceiling
data := padTo(buildUPXHeader(maxUPXOriginalSize, 0x1000), 4096)
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
build, err := scanReader(ctx, file.NewLocation("/refused"), bytes.NewReader(data), false)
assert.Nil(t, build)
require.Error(t, err)
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
require.NotEmpty(t, coordErrs,
"a file we declined to expand is a gap we chose to leave, so it is reported")
assert.ErrorIs(t, coordErrs[0].Reason, errUPXSizeRefused)
}
// corruptLZMABlock builds a b_info whose stream carries valid LZMA parameters and nothing else that
// decodes. This is what the loader stub behind the last real block looks like to readChainBlock: the
// method byte is one we implement and the sizes are in range, so the chain accepts it and the decoder is
// the thing that says no.
func corruptLZMABlock(szUnc uint32) []byte {
stream := make([]byte, 64)
stream[0] = 0x00 // pb = 0
stream[1] = 0x00 // lc = 0, lp = 0
for i := 2; i < len(stream); i++ {
stream[i] = byte(i * 7) // not a range-coded anything
}
b := make([]byte, 12)
binary.LittleEndian.PutUint32(b[0:4], szUnc)
binary.LittleEndian.PutUint32(b[4:8], uint32(len(stream)))
b[8] = 14 // b_method = LZMA
return append(b, stream...)
}
// TestDecompressUPX_UndecodableLaterBlockKeepsWhatCameBefore covers a chain that gives up a good
// reconstruction over garbage past the end of it.
//
// readChainBlock accepts any b_info-shaped run of bytes whose method it knows and whose sizes fit the
// file, so the loader stub behind the last real block parses as a block roughly one time in 256 on the
// method byte alone. A decode failure on that block must not discard every block already placed and turn a
// fully recoverable binary into zero packages: the two sibling conditions on the same block (an overrun,
// and a placement that does not fit) already keep their partial output, and this one must too.
func TestDecompressUPX_UndecodableLaterBlockKeepsWhatCameBefore(t *testing.T) {
head := bytes.Repeat([]byte("H"), 256)
data := buildUPXHeader(4096, 4096)
data = append(data, blockFor(t, head)...)
data = append(data, corruptLZMABlock(64)...)
data = padTo(data, 1024)
out, err := unpack(t, data)
require.ErrorIs(t, err, errUPXPartial, "the blocks already placed are kept, and what was lost is reported")
require.NotErrorIs(t, err, errUPXDecompress, "a good reconstruction is not discarded over trailing garbage")
require.NotNil(t, out)
assert.Equal(t, head, readAll(t, out), "everything decoded before the bad block stays placed")
}
// TestDecompressUPX_UndecodableFirstBlockStillFailsTheFile is the other half: with nothing placed there is
// no partial output to keep, so the file really is one we could not unpack.
func TestDecompressUPX_UndecodableFirstBlockStillFailsTheFile(t *testing.T) {
data := padTo(append(buildUPXHeader(4096, 4096), corruptLZMABlock(64)...), 1024)
_, err := unpack(t, data)
require.ErrorIs(t, err, errUPXDecompress, "a packed file with nothing readable in it is a gap in the SBOM")
assert.NotErrorIs(t, err, errUPXPartial)
}
// TestScanReader_PartialReconstructionIsReported is the regression test for a partial unpack vanishing
// silently. The reconstruction is truncated to the contiguous prefix, which is the right bound, but the
// bytes past the gap are gone: on a real binary that is the non-loadable tail holding the section headers,
// so the symbols and often the build info go with it. Reported at Trace, the binary contributed no
// packages and no unknown, which is the one outcome the reporting split in upx.go exists to prevent.
func TestScanReader_PartialReconstructionIsReported(t *testing.T) {
const declared = 64 << 20
data := sparseFixture(t, declared, declared-64, 0, 300_000)
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
build, err := scanReader(ctx, file.NewLocation("/partial"), bytes.NewReader(data), false)
assert.Nil(t, build)
require.Error(t, err)
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
require.NotEmpty(t, coordErrs, "a reconstruction that lost bytes is a gap this cataloger chose to leave")
assert.ErrorIs(t, coordErrs[0].Reason, errUPXPartial)
}
// TestReadContentsAndBuildInfo_FallbackToFoundBytesReportsNoGap pins the retry direction. A gap is
// reported because the readers after it (crypto settings, arch, symbols) would otherwise read a short
// reconstruction with nothing saying so, which is what
// TestReadContentsAndBuildInfo_GapSurvivesBuildInfoFromTheReconstruction covers. That reasoning runs out
// here: the retry hands back the bytes as they were found, every reader below reads those, and the header
// that produced the short reconstruction was not describing real UPX output to begin with. Reporting it
// would put "UPX reconstruction is incomplete" on a binary that was never packed and was read in full.
func TestReadContentsAndBuildInfo_FallbackToFoundBytesReportsNoGap(t *testing.T) {
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
// a chain that decodes one block and then hits a block it cannot read: the reconstruction is real but
// short, and the build info comes from the fallback to the bytes as they were found
chain := append(blockFor(t, []byte("junk")), corruptLZMABlock(64)...)
spiked := spliceUPXChain(t, goELF64Fixture(t), 64<<10, chain)
unpacked, bi, err := readContentsAndBuildInfo(ctx, bytes.NewReader(spiked))
t.Cleanup(func() { closeUnpacked(unpacked) })
require.NotNil(t, bi, "the fallback still finds the build info")
require.Nil(t, unpacked, "the reconstruction was released; the caller reads the bytes as they were found")
assert.NoError(t, err, "nothing downstream reads the short reconstruction, so there is no gap to report")
}
// TestUnpackUPX_CancelledContextComesBackAsCtxErr covers the cancellation arm in unpackUPX directly: the
// scan-level test above it passes even with this arm deleted, since readContentsAndBuildInfo does not
// re-check ctx.Err() on its own.
func TestUnpackUPX_CancelledContextComesBackAsCtxErr(t *testing.T) {
payload := bytes.Repeat([]byte("A"), 2048)
blocks := make([][]byte, 64)
for i := range blocks {
blocks[i] = payload
}
data := padTo(buildUPXFile(t, 4096*64, 2048, blocks, nil), 4096)
ctx, cancel := context.WithCancel(tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir())))
cancel()
out, err := unpackUPX(ctx, bytes.NewReader(data))
require.ErrorIs(t, err, context.Canceled)
assert.NotErrorIs(t, err, errUPXDecompress, "a cancelled scan is not a gap in the SBOM")
assert.NotErrorIs(t, err, errUPXPartial, "nor is it a short reconstruction")
assert.Nil(t, out, "a cancelled unpack hands back nothing; the caller reads the input as it was found")
}
// TestParseUPXInfo_HeaderRunsPastTheScanWindow covers the guard on the l_info+p_info read: the magic can
// sit close enough to the end of what was actually read that the 20 bytes behind it are not there.
func TestParseUPXInfo_HeaderRunsPastTheScanWindow(t *testing.T) {
data := append(packedELFStub(), upxMagic...) // magic at 64, nothing behind it
r := bytes.NewReader(data)
_, err := parseUPXInfo(r, int64(len(data)))
require.Error(t, err)
assert.ErrorIs(t, err, errUPXImplausibleHeader)
assert.NotErrorIs(t, err, errUPXSizeRefused, "a header we could not even read is not a refusal to expand")
}
// TestReadContentsAndBuildInfo_GapSurvivesBuildInfoFromTheReconstruction is the same property one path
// over: here the reconstruction itself carries the build info, so the early return must not throw the
// unpack error away. The chain rebuilds a whole Go binary and then hits a block it cannot read, which is
// exactly the shape of a real packed binary whose tail extents did not all come back.
func TestReadContentsAndBuildInfo_GapSurvivesBuildInfoFromTheReconstruction(t *testing.T) {
payload := goELF64Fixture(t)
data := buildUPXHeader(uint32(len(payload)+64), 4096)
data = append(data, blockFor(t, payload)...)
data = append(data, corruptLZMABlock(64)...)
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
contents, bi, err := readContentsAndBuildInfo(ctx, bytes.NewReader(data))
require.NotNil(t, contents)
t.Cleanup(func() { closeUnpacked(contents) })
require.NotNil(t, contents, "the build info has to come from the reconstruction here")
require.NotNil(t, bi, "the rebuilt binary is a whole Go binary")
assert.ErrorIs(t, err, errUPXPartial,
"the block the chain gave up is still a gap, even though the build info came through")
}
// TestScanReader_PartialReconstructionStillYieldsItsPackages is the other direction of the reporting
// split, and the one that keeps the new reporting from becoming a regression of its own: a gap is
// something to report alongside the packages, not instead of them. A binary whose chain gave up its tail
// still has its module list, and dropping it would trade a silent false negative for a loud one.
func TestScanReader_PartialReconstructionStillYieldsItsPackages(t *testing.T) {
payload := goELF64Fixture(t)
data := buildUPXHeader(uint32(len(payload)+64), 4096)
data = append(data, blockFor(t, payload)...)
data = append(data, corruptLZMABlock(64)...)
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
build, err := scanReader(ctx, file.NewLocation("/partial-but-usable"), bytes.NewReader(data), false)
require.NotNil(t, build, "a reported gap must not cost the packages that did come through")
t.Cleanup(func() { closeUnpacked(build.unpacked) })
assert.NotNil(t, build.BuildInfo)
require.Error(t, err)
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
require.NotEmpty(t, coordErrs, "and the gap is still reported")
assert.ErrorIs(t, coordErrs[0].Reason, errUPXPartial)
}
// sizeProbeSpy counts the size probes made against it. intFile.ReaderSize answers from Size() when the
// reader has one, which is what this intercepts.
type sizeProbeSpy struct {
*bytes.Reader
probes int
}
func (s *sizeProbeSpy) Size() int64 {
s.probes++
return s.Reader.Size()
}
// TestUnpackUPX_ContainerGateComesBeforeTheSizeProbe pins the ordering that keeps this cheap. The
// cataloger runs over every file in an image and almost none are packed ELF, so nothing above the six
// byte ident check should cost more than that: ReaderSize seeks to the end and reads the last byte back,
// which over a squashfs or tar-backed reader is a real seek and decompress.
func TestUnpackUPX_ContainerGateComesBeforeTheSizeProbe(t *testing.T) {
tests := []struct {
name string
data []byte
probed bool
}{
{name: "not an ELF at all", data: bytes.Repeat([]byte("just some bytes"), 64)},
{name: "a 32-bit ELF", data: func() []byte { b := padTo(packedELFStub(), 256); b[4] = 1; return b }()},
{name: "a big-endian ELF64", data: func() []byte { b := padTo(packedELFStub(), 256); b[5] = 2; return b }()},
{name: "an ELF64 we would try to unpack", data: padTo(packedELFStub(), 256), probed: true},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
spy := &sizeProbeSpy{Reader: bytes.NewReader(tt.data)}
out, err := unpackUPX(context.Background(), spy)
require.NoError(t, err)
t.Cleanup(func() { closeUnpacked(out) })
if tt.probed {
assert.Positive(t, spy.probes, "a candidate container does get measured")
return
}
assert.Zero(t, spy.probes, "a file we will never unpack must not be measured first")
})
}
}
// TestScanReader_ChainEndingShortIsReported covers the fourth shape a short reconstruction comes in: a
// chain that simply ran out of b_info structures before rebuilding p_filesize must be reported, the same
// as one that gave up mid-walk, one that hit the block cap, and one with blocks stranded past a hole. Every
// quiet exit in readChainBlock produces that shape (an unreadable b_info, the end marker, a zero sz_cpr,
// compressed data past the end of the input, an unimplemented method past the first block), and the
// reconstruction is then truncated to the covered prefix, which can drop the section headers and leave the
// binary contributing neither packages nor an unknown.
func TestScanReader_ChainEndingShortIsReported(t *testing.T) {
// one 2048 byte block against a declared 4096, then the end marker: nothing failed, the chain just ran
// out. stopped is nil, covered equals the furthest placement, and covered is half of total.
payload := bytes.Repeat([]byte("A"), 2048)
data := buildUPXHeader(4096, 2048)
data = append(data, blockFor(t, payload)...)
data = append(data, make([]byte, 12)...) // end marker
data = padTo(data, 512)
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
build, err := scanReader(ctx, file.NewLocation("/short-chain"), bytes.NewReader(data), false)
assert.Nil(t, build)
require.Error(t, err)
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
require.NotEmpty(t, coordErrs, "a chain that ended before rebuilding p_filesize is a reported gap")
assert.ErrorIs(t, coordErrs[0].Reason, errUPXPartial)
}
// TestScanReader_SizeRefusalOnAnUnpackedBinaryIsNotReported is the other direction of
// TestScanReader_SizeRefusalIsReported. The magic is found by an unanchored scan over the first 8KB, so an
// ordinary Go binary can carry a "UPX!" that parses into a header the size bounds then refuse. Nothing is
// unpacked in that case: unpackUPX hands back the input exactly as it was found, every reader below reads
// those bytes in full, and the build info proves they are complete. Reporting the refusal there put
// "golang binary read incompletely" on a binary the cataloger read completely, with every package present.
func TestScanReader_SizeRefusalOnAnUnpackedBinaryIsNotReported(t *testing.T) {
original := goELF64Fixture(t)
// a header whose p_filesize no input this size could ever pay for, so parseUPXInfo refuses it before
// decompressUPX is reached and there is no reconstruction at any point
spiked := spliceUPXChain(t, original, 0xFFFFFFFF, blockFor(t, []byte("junk")))
_, perr := parseUPXInfo(bytes.NewReader(spiked), int64(len(spiked)))
require.ErrorIs(t, perr, errUPXSizeRefused, "the splice has to be refused by the size bounds")
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
build, err := scanReader(ctx, file.NewLocation("/spiked"), bytes.NewReader(spiked), false)
if build != nil {
t.Cleanup(func() { closeUnpacked(build.unpacked) })
}
require.NotNil(t, build, "the binary is a complete Go binary and must still be cataloged")
assert.Nil(t, build.unpacked, "nothing was unpacked, so these are the bytes as found")
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
assert.Empty(t, coordErrs, "a file read in full must not be reported as read incompletely")
}
// TestDecompressUPX_DenseChainReportsNoGap is the guard for the covered >= total early return in
// finalExtent, and it is the fixture every other one in this package is not: dense. A chain that rebuilds
// exactly p_filesize must come back with no partial reason attached.
//
// This matters because reporting a chain that ended short (the arm below that early return) is only safe if
// well-formed output reaches full coverage. Every other crafted fixture here declares more than it delivers
// and asserts errUPXPartial, and the one test that sees real `upx --best --lzma` output needs Docker, so
// without this the early return has no local guard and the new arm's false-positive risk is untested.
//
// Layout below tiles 320 bytes through all three placement rules: block 1 at zero, block 2 sequentially
// behind it, block 3 at a PT_LOAD offset from the headers in block 1, and block 4 into the hole those left.
func TestDecompressUPX_DenseChainReportsNoGap(t *testing.T) {
const (
headerBlock = 64 + 112 // ELF64 header plus two program headers
block2Size = 24 // sequential, so [176, 200)
ptLoad1 = 256 // block 3 lands here, leaving [200, 256) behind
block3Size = 64 // [256, 320)
holeSize = ptLoad1 - (headerBlock + block2Size)
total = ptLoad1 + block3Size
)
phdrs := make([]byte, 112)
binary.LittleEndian.PutUint32(phdrs[0:4], 1) // phdr[0] PT_LOAD
binary.LittleEndian.PutUint64(phdrs[8:16], 0) // p_offset 0, the extent block 2 covers
binary.LittleEndian.PutUint32(phdrs[56:60], 1)
binary.LittleEndian.PutUint64(phdrs[64:72], ptLoad1) // phdr[1] p_offset, where block 3 goes
block1 := buildELF64(64, 56, 2, phdrs)
require.Len(t, block1, headerBlock, "the first block is the original ELF headers")
data := buildUPXHeader(total, 4096)
data = append(data, blockFor(t, block1)...) // placed at 0
data = append(data, blockFor(t, bytes.Repeat([]byte("S"), block2Size))...) // sequential
data = append(data, blockFor(t, bytes.Repeat([]byte("P"), block3Size))...) // -> ptLoadOffsets[1]
data = append(data, blockFor(t, bytes.Repeat([]byte("F"), holeSize))...) // -> firstHole
data = append(data, make([]byte, 12)...) // end marker
data = padTo(data, 1024)
out, err := unpack(t, data)
require.NotNil(t, out)
require.NoError(t, err,
"a chain that rebuilt every byte p_filesize declared has given nothing up, so there is no gap to report")
got := readAll(t, out)
assert.Len(t, got, total, "the reconstruction is exactly what was declared")
assert.Equal(t, bytes.Repeat([]byte("F"), holeSize), got[headerBlock+block2Size:ptLoad1],
"the hole between the sequential extent and the PT_LOAD one is filled by the block after them")
assert.Equal(t, bytes.Repeat([]byte("P"), block3Size), got[ptLoad1:],
"block 3 lands at the PT_LOAD offset the headers declared")
}
// TestFinalExtent pins which reason a short reconstruction comes back with, not merely that it is short.
// Every arm wraps errUPXPartial and nothing else read the messages, so three of the four were
// interchangeable: deleting the block-cap arm and the stranded-blocks arm left the whole package green
// because both fell through to the arm below them. finalExtent is a pure function, so this costs nothing.
func TestFinalExtent(t *testing.T) {
tests := []struct {
name string
covered uint64
blockNum int
total uint64
stopped error
wantMsg string // empty means no gap is reported
}{
{
name: "a chain that rebuilt everything reports nothing, whatever it tripped over next",
covered: 100,
blockNum: 2,
total: 100,
stopped: fmt.Errorf("%w: ignored once coverage is complete", errUPXPartial),
},
{
name: "a mid-walk stop is more specific than anything the extents show",
covered: 100,
blockNum: 2,
total: 200,
stopped: fmt.Errorf("%w: block 2 did not decode", errUPXPartial),
wantMsg: "block 2 did not decode",
},
{
name: "hitting the block cap is named as the cap",
covered: 100,
blockNum: maxUPXBlocks,
total: 200,
wantMsg: "block cap",
},
{
name: "a chain that simply ended is named as such",
covered: 100,
blockNum: 3,
total: 200,
wantMsg: "chain ended after 100 of the declared 200",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got := finalExtent(tt.covered, tt.blockNum, tt.total, tt.stopped)
if tt.wantMsg == "" {
assert.NoError(t, got)
return
}
require.Error(t, got)
assert.ErrorIs(t, got, errUPXPartial)
assert.Contains(t, got.Error(), tt.wantMsg,
"the reason has to say which shape this was, or the arms are interchangeable")
})
}
}
File diff suppressed because it is too large Load Diff
+43 -65
View File
@@ -8,57 +8,64 @@ import (
"github.com/stretchr/testify/require"
)
func TestIsUPXCompressed(t *testing.T) {
// every case here is one the magic scan must reject, asserted on parseUPXInfo's errNotUPX result.
//
// Every negative case that is meant to exercise the magic scan has to carry a real ELF64 little-endian
// ident, or the container gate rejects it first and the subtest passes without the scan running at all
// (the gate and a missing magic both return errNotUPX, so nothing in the assertion can tell them apart).
// TestParseUPXInfo_OnlyELF64LittleEndian owns the gate; these own the scan.
func TestParseUPXInfo_MagicDetection(t *testing.T) {
tests := []struct {
name string
data []byte
expected bool
name string
data []byte
foundMagic bool // the magic was located (the header may still be rejected as implausible)
}{
{
name: "contains UPX magic at start",
data: append([]byte("UPX!"), make([]byte, 100)...),
expected: true,
name: "contains UPX magic at start",
data: append(append(packedELFStub(), []byte("UPX!")...), make([]byte, 100)...),
foundMagic: true,
},
{
name: "contains UPX magic with offset",
data: append(append(make([]byte, 500), []byte("UPX!")...), make([]byte, 100)...),
expected: true,
name: "contains UPX magic with offset",
data: append(append(append(packedELFStub(), make([]byte, 500)...), []byte("UPX!")...), make([]byte, 100)...),
foundMagic: true,
},
{
name: "no UPX magic",
data: []byte("\x7FELF" + string(make([]byte, 100))),
expected: false,
// UPX packs PE and Mach-O with this same l_info/p_info layout, and nothing here can place
// their blocks, so the container has to be checked before the magic is trusted
name: "UPX magic but not an ELF container",
data: append([]byte("MZ\x90\x00UPX!"), make([]byte, 100)...),
},
{
name: "empty data",
data: []byte{},
expected: false,
name: "no UPX magic",
data: append(packedELFStub(), make([]byte, 100)...),
},
{
name: "partial UPX magic",
data: []byte("UPX"),
expected: false,
// three of the four magic bytes must not match
name: "partial UPX magic",
data: append(append(packedELFStub(), []byte("UPX")...), make([]byte, 100)...),
},
{
// never reaches the scan: the six byte ident read comes up short, so the container gate
// refuses it first. Kept because an empty reader is a shape the cataloger really is handed.
name: "empty data",
data: []byte{},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
reader := bytes.NewReader(tt.data)
result := isUPXCompressed(reader)
assert.Equal(t, tt.expected, result)
_, err := parseUPXInfo(bytes.NewReader(tt.data), int64(len(tt.data)))
require.Error(t, err, "none of these fixtures is a usable UPX header")
if tt.foundMagic {
assert.NotErrorIs(t, err, errNotUPX, "the magic was found, so the header was parsed and rejected")
} else {
assert.ErrorIs(t, err, errNotUPX)
}
})
}
}
func TestParseUPXInfo_NotUPX(t *testing.T) {
data := []byte("\x7FELF" + string(make([]byte, 100)))
reader := bytes.NewReader(data)
_, err := parseUPXInfo(reader)
require.Error(t, err)
assert.ErrorIs(t, err, errNotUPX)
}
func TestParseUPXInfo_ValidHeader(t *testing.T) {
// construct a minimal valid UPX header matching actual format
// l_info: checksum (4) + magic (4) + lsize (2) + version (1) + format (1)
@@ -84,45 +91,16 @@ func TestParseUPXInfo_ValidHeader(t *testing.T) {
14, 0, 0, 0, // method=LZMA, filter info
}
data := append(append(lInfo, pInfo...), bInfo...)
data = append(data, make([]byte, 100)...) // padding
// padded so the declared 1MB stays within maxUPXExpansion of the fixture's own size; a real UPX file
// carries the compressed data this header describes, and the bound is measured against that.
data := append(append(append(packedELFStub(), lInfo...), pInfo...), bInfo...)
data = append(data, make([]byte, 0x100000/maxUPXExpansion)...)
reader := bytes.NewReader(data)
info, err := parseUPXInfo(reader)
info, err := parseUPXInfo(reader, sizeOf(t, reader))
require.NoError(t, err)
assert.Equal(t, uint8(14), info.version)
assert.Equal(t, uint8(22), info.format)
assert.Equal(t, uint32(0x100000), info.originalSize)
}
func TestDecompressUPX_UnsupportedMethod(t *testing.T) {
// construct a header with an unsupported compression method
lInfo := []byte{
0, 0, 0, 0, // l_checksum
'U', 'P', 'X', '!',
0, 0, // l_lsize
14, 22, // version, format
}
pInfo := []byte{
0, 0, 0, 0, // p_progid
0x00, 0x01, 0x00, 0x00, // p_filesize = 256 bytes (small for test)
0, 0, 0x10, 0, // p_blocksize
}
bInfo := []byte{
0x00, 0x01, 0x00, 0x00, // sz_unc = 256
0x80, 0x00, 0x00, 0x00, // sz_cpr = 128
99, 0, 0, 0, // unsupported method
}
data := append(append(lInfo, pInfo...), bInfo...)
data = append(data, make([]byte, 1000)...)
reader := bytes.NewReader(data)
_, err := decompressUPX(reader)
require.Error(t, err)
assert.ErrorIs(t, err, errUnsupportedUPXMethod)
}
+14
View File
@@ -103,3 +103,17 @@ func noDirectELFOpen(m dsl.Matcher) {
Where(!m.File().PkgPath.Matches(`/syft/internal/elfutil$`)).
Report("do not open ELF files with debug/elf directly; use elfutil.NewFile, which bounds declared decompressed section sizes")
}
// nolint:unused
func noUnboundedELFParsers(m dsl.Matcher) {
// buildinfo.Read and goversion's ReadExeFromReader each open debug/elf internally and expand sections
// sized by the file itself, the same unbounded allocation noDirectELFOpen guards against above. They
// are only safe in this repo because scan_binary.go gates the reader before calling them: readBuildInfo
// runs elfutil.CheckSectionNameTable first, getCryptoInformation runs elfutil.CheckAllSections first.
m.Match(
`buildinfo.Read($_)`,
`version.ReadExeFromReader($_)`,
).
Where(!(m.File().Name.Matches(`scan_binary\.go$`) && m.File().PkgPath.Matches(`/syft/pkg/cataloger/golang$`))).
Report("do not call buildinfo.Read/goversion.ReadExeFromReader directly; route through the bounded wrappers in the golang cataloger (readBuildInfo, getCryptoInformation), or gate the reader with elfutil.CheckSectionNameTable/elfutil.CheckAllSections first")
}