mirror of
https://github.com/anchore/syft.git
synced 2026-10-11 21:57:22 +02:00
fix(golang): bound allocations from UPX-packed binary headers (#5195)
* fix(golang): bound UPX decompression by what the input can justify A UPX block's `sz_unc` was passed straight to `make([]byte, n)`, so four bytes of attacker input spanned the full uint32 range. A ~300KB file could drive multi-GB resident memory, which Go reports as a fatal runtime OOM that no `recover` on this path can contain. Syft parses binaries out of arbitrary images, so that is reachable from any scan. Two bounds do the work, and UPX's own invariants make them exact: - all blocks together reconstruct `p_filesize`, so a running remainder caps the sum. Bounding only per-block would leave the total at (block count x original size), so the remainder is the part that actually closes it - `p_filesize` sizes the output buffer, so it is capped both absolutely and against the size of the file on disk. The absolute cap alone left a ~60 byte header able to claim 500MB; the ratio alone cannot work either, since LZMA encodes a run of N equal bytes in O(log N) and a `go:embed` of 120MB of zeros legitimately packs 209x The output buffer is also allocated on the first block that survives validation rather than up front, so a header claiming a large size with nothing decodable behind it costs nothing. Alongside that, several ways the block loop could be steered off its own buffer: - `parseELFPTLoadOffsets` bounds-checked program headers with `phStart+phentsize > len(buf)`, which overflows for a `p_offset` near 2^64 and lets an out-of-range entry through into a slice index. Switched to the subtraction form, and a `phentsize` shorter than an ELF64 program header is rejected since the reads use fixed offsets - block placement had the same overflow shape, and on failure it skipped the copy and fell through. `outputOffset` derives from the rejected offset, so a `p_offset` at the uint64 ceiling wrapped it to 0 and the next block landed on the reconstructed ELF header, still returning success. Out-of-range placement now ends the block chain, keeping the blocks already placed since those often carry `.go.buildinfo` - the UPX 2-byte LZMA header carries `lc`/`lp` as nibbles and `pb` as three bits, so they can hold values LZMA does not permit. Folded into the props byte they wrap (`lc=15, lp=15, pb=7` gives 465, which truncates to 209) and the stream mis-decodes instead of failing - blocks decode directly into a slice of the output buffer instead of a per-block buffer that is then copied, so one block no longer doubles peak memory - the block count is capped: the size budget alone still allows millions of 12-byte blocks, each spinning up an LZMA reader Verified against the `image-small-upx` fixture, a real `upx --best --lzma` binary: four blocks, max `sz_unc` equal to `p_blocksize` exactly, blocks summing to 0.68 of `p_filesize`, and a 1.96x expansion against the packed file, so every bound holds with headroom on genuine output. Note `p_blocksize` is deliberately not used as a ceiling on `sz_unc`. It adds nothing the remainder does not already cover, and real output sits at exactly `p_blocksize`, so it would run with zero headroom against a value the format does not promise. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * fix(golang): report UPX binaries we cannot unpack as unknowns A packed Go binary we fail to unpack means its packages are silently missing from the SBOM. `getBuildInfo` checked the decompression error for nil and discarded it, so a rejected file surfaced only the original `buildinfo.Read` error, and that one is deliberately silenced because it is usually just "not a Go binary". `decompressUPX` now marks the cases worth reporting with `errUPXDecompress`: it got past the header and the method dispatch and still could not unpack the file. Everything else stays quiet, which matters more than it sounds. `upx` defaults to NRV2B unless `--lzma` is passed and only LZMA is implemented here, so treating any decompression failure as reportable would attach a golang-cataloger unknown to every packed non-Go binary in an image. Making the reportable case opt in rather than the quiet case a growing list also means a guard added later is silent by default. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * refactor(golang): drop the vendored xcoff parser The only thing the golang cataloger read from it was `TargetMachine`, which is the same two magic bytes `getGOARCHFromBin` had already matched on to dispatch there. So a full XCOFF walk (string table, symbol table, every relocation table) ran to recover a value that was already in hand. `getGOARCHFromBin` reads those two bytes directly now. That is also strictly more correct: `xcoff.NewFile` bailed with "no symbol table" when `symptr == 0`, so a stripped XCOFF binary reported no arch at all. Fixtures move to `testdata/xcoff/`, since they are no longer a package's own testdata. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * feat(internal/file): report a reader's real size behind one helper Bounds get written as "does this claim exceed what the reader actually holds", and the subtlety is that a reader's own answer cannot be trusted: an `io.SectionReader` reports the nominal length it was built with, so one constructed with `1<<63-1` will happily claim to hold 8EB. Confirming the last byte is readable is what separates a real size from a nominal one. Returns `(int64, bool)` rather than a bare size, since there are five ways to not know and collapsing them to 0 makes every caller reinvent "0 means stop bounding". The doc also names the wrapper trap: a type embedding `io.ReaderAt` as an interface promotes only `ReadAt`, so it answers no size at all however large the reader beneath it is, and a bound written against it silently does nothing. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * fix(golang): bound every section a reader can expand, not just the name table `CheckSectionNameTable` bounds the one section `elf.NewFile` expands while parsing, which is enough for a caller that only parses. It is not enough for one that goes on to read sections: those are expanded lazily, so a 260KB ELF declaring a compressed `.symtab` drove 1.3GB of allocation through `goversion`, which opens the file with `debug/elf` inside its own package and cannot be routed through `elfutil.NewFile`. `CheckAllSections` is that gate, and is deliberately named as the superset so the relationship to the narrow one is structural rather than documented. Also exports `ErrDeclaredSizeExceeded`. A refusal here costs the SBOM a package, so callers report it as an unknown, and matching on an error string to decide that is not something to build a reporting policy on. A ruleguard rule now flags `buildinfo.Read` and `version.ReadExeFromReader` outside `scan_binary.go`. Both open `debug/elf` internally and are only bounded there by the wrappers that gate the reader first, which the existing rule on `elf.NewFile` could not see through. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * feat(internal/spillbuf): sparse offset-addressed buffer that spills to disk For output that arrives out of order, at offsets the input itself declares: a decompressor placing extents, an archive rebuilding a file from chunks. A plain `[]byte` cannot do that without either pre-sizing to a length the input claims or growing to the furthest offset it names, and both hand a hostile input an allocation knob. The load-bearing property is that it reports only what it actually stored. `Size` is the contiguous run written from offset zero and a read past it is `io.EOF`, never the zeros an unwritten region would hand back for free. Sparse storage serving holes as real content is what lets a write offset stand in for output nothing produced, so anything sizing an allocation against this reader (`saferio` in `debug/elf` does exactly that) stays bounded by work actually done. `FirstGap` answers where the next write fits, so callers filling holes do not need the extent bookkeeping and cannot re-derive that rule wrongly, and it only reports gaps lying entirely inside `[0, within)`. The extent type is unexported for the same reason. The rest of the contract worth stating: reads and writes after `Close` fail with `os.ErrClosed` rather than looking like an empty buffer, and a write that fails leaves the buffer unchanged. Memory is bounded by the limit rather than by how much is written: filling 1MB and filling 64MB cost about the same. What lands on disk past that limit is the caller's bound to set, not this package's. Growth is geometric and capped, because sizing the tier to exactly what each write needs is quadratic once a caller streams through it in small chunks. The golang cataloger is the only consumer today, and nothing else in the tree has this shape. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * fix(golang): rebuild UPX binaries through spillbuf and read the reconstruction everywhere Two amplification paths were still reachable. A 128 byte input allocated 16MB by trusting `sz_cpr` before weighing it against the bytes the file actually holds, and the reconstruction was assembled in a `[]byte` sized from the header's declared original size, so a 301KB input allocated 300MB. The reconstruction is a `spillbuf.Buffer` now, which is what makes the second one go away: blocks are placed at the offsets the file declares, nothing is pre-sized from a claim, and the reader's length is the contiguous run actually rebuilt. That last part is the invariant everything downstream leans on. A reader reporting the declared size would serve the holes between placed blocks as zeros it never stored, and a 64 byte block parked at `p_filesize-64` then made a 300KB input hand out a 64MB reader: measured, that drove 269MB through `getBuildInfo` and returned nothing, silently. Placement and storage are separate concerns. `upx.go` places blocks at the offsets the file declares and asks the buffer how much was rebuilt and where the next gap is; it knows nothing about temp files or memory limits. The remaining bounds are ratios against the input wherever they can be, since an absolute ceiling only ever drops a large legitimate binary. The one absolute left is on `p_filesize`, because padding an input is nearly free, and the LZMA dictionary is paid for in its own block's `sz_cpr`. Interface change: `scanFile` takes a context and hands back the reconstruction for each binary that turned out to be packed, and the caller owns it. Everything after the build info read (crypto settings, arch, symbols, the version scan) reads it, so a packed binary no longer reports its packages with no symbols at all. Also narrows what gets reported. This cataloger runs on every executable in an image, so reporting every `getBuildInfo` error attached an unknown to every corrupt ELF, odd PE and non-Go Mach-O slice in it. Now only four gaps, the ones that cost the SBOM something syft could have had: a packed file we could not decode (`errUPXDecompress`), a header the size bounds refused (`errUPXSizeRefused`), a reconstruction that came up short (`errUPXPartial`), and an ELF expansion `elfutil` declined (`ErrDeclaredSizeExceeded`). An unimplemented UPX method stays quiet on purpose: `upx` defaults to NRV2B and only LZMA is implemented here, so reporting it would attach an unknown to most packed non-Go binaries in an image. That filtering is a new pattern. No other cataloger picks which of its errors become unknowns, and the reason this one does is that it runs on every executable in an image rather than on files a glob already selected. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * chore: ignore local spec/ planning notes `/specs` was already ignored; `/spec` is the same thing under the singular name and was not, so working notes landed in commits. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * refactor(golang): name what the unpack hands back, not where it stores it the scan's signatures said `*spillbuf.Buffer` all the way down, which told every reader below the scan that it cares whether the reconstruction lives in memory or on disk. it does not. two interfaces instead, each naming only what its side actually uses: - `unpackedContents` for the consumers: `ReaderAt`, `Closer`, and a `Size` that is the contiguous run rebuilt from zero - `blockSink` for the decoder, which additionally needs `WriterAt` and `FirstGap` to place blocks this costs one thing worth flagging. `Close` on a nil `*spillbuf.Buffer` was nil-safe, so callers just deferred it; widened into an interface, a nil pointer becomes a non-nil interface and the same call panics. `closeUnpacked` absorbs that, and it is now the only way an owner releases contents. `readerFor`, `seekerFor` and `closeUnpacked` are the three places that resolve the nil. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * docs(elfutil): say that CheckAllSections is the decompression-bomb gate the doc led with "bounds every section syft can drive debug/elf into decompressing", which describes the mechanism without ever naming what it is defending against. a reader landing on the call site could not tell that this is the decompression-bomb check. state the attack instead: a compressed section header declares its own decompressed size, debug/elf believes that number and allocates it on open, and nothing forces it to match what the compressed bytes actually yield. the 260KB ELF that drove 1.3GB of allocation now lives in the doc rather than only at the one call site that happened to mention it. comments only. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * docs(golang): move the overflow note onto the bounds test it describes the "subtraction form" note sat above the phStart assignment with ten lines of wrap analysis between it and the `if` it was talking about, so the subtraction it names reads as missing. it is `hdrLen-phStart < phentsize`. move it onto that test and name the form it is avoiding (`phStart+phentsize > hdrLen`, which can overflow past the buffer end) so the referent is not several paragraphs away. comments only. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> * chore(golang): drop the internal/ README left over from xcoff that README existed to record where the vendored xcoff package came from and that Go keeps it internal, which is a provenance note worth having while third-party code sat there. dropping xcoff took the reason with it, and repurposing it to point at gotestdata just made it redundant: gotestdata has its own README covering why it is under internal/, why it is not testdata/, and what belongs in each. internal/ now holds one self-documenting directory. Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com> --------- Signed-off-by: Alex Goodman <wagoodman@users.noreply.github.com>
This commit is contained in:
@@ -20,6 +20,7 @@ bin/
|
||||
/.tool
|
||||
/.task
|
||||
/generate
|
||||
/spec
|
||||
/specs
|
||||
mise.toml
|
||||
.make/.make
|
||||
|
||||
@@ -0,0 +1,66 @@
|
||||
package file
|
||||
|
||||
import "io"
|
||||
|
||||
// ReaderSize reports the number of bytes behind r. The bool is false when that cannot be determined, and
|
||||
// callers must branch on it rather than treating the zero as a size.
|
||||
//
|
||||
// Binary parsers read sizes out of the files they are parsing and allocate against them. Those sizes are
|
||||
// 32- and 64-bit fields under the control of whoever produced the file, so a truncated or hostile input
|
||||
// can declare gigabytes that are not there. The read would fail afterwards either way; what matters is
|
||||
// refusing before the allocation, and that needs a real byte count to weigh the claim against.
|
||||
//
|
||||
// The bool is the whole point of the signature. An earlier version returned 0 for every failure mode, and
|
||||
// a caller that forgot to special-case it got a bound of "nothing is allowed" or, more often, silently
|
||||
// fell back to an unbounded path. Making the caller name the failure keeps that from happening by
|
||||
// omission.
|
||||
//
|
||||
// Both *bytes.Reader and *io.SectionReader answer directly. A file handle only has Seek, so fall back to
|
||||
// a save-and-restore seek, which leaves the caller's cursor where it found it. That fallback is not
|
||||
// atomic: it moves the cursor and puts it back, so a reader being read concurrently can observe the
|
||||
// intermediate position.
|
||||
//
|
||||
// Neither answer is taken on faith, because a reader can report a length it cannot deliver: an
|
||||
// *io.SectionReader answers with the length it was constructed with, and debug/elf and friends construct
|
||||
// theirs with a nominal one, so io.NewSectionReader(r, 0, 1<<63-1) would otherwise report a bound that
|
||||
// permits everything as if it had been measured. The last byte is read back to confirm the count, which is
|
||||
// one ReadAt whatever the size, and the read is what the caller was going to weigh anyway.
|
||||
//
|
||||
// A wrapper that embeds io.ReaderAt as an interface promotes only ReadAt, so it answers no size here
|
||||
// however sizable the reader underneath it is. A type meant to be bounded needs to forward Size() itself.
|
||||
func ReaderSize(r io.ReaderAt) (int64, bool) {
|
||||
size, ok := reportedSize(r)
|
||||
if !ok {
|
||||
return 0, false
|
||||
}
|
||||
// ReadAt may return io.EOF alongside a full read, so the count is what says the byte is there
|
||||
var last [1]byte
|
||||
if n, _ := r.ReadAt(last[:], size-1); n != len(last) {
|
||||
return 0, false
|
||||
}
|
||||
return size, true
|
||||
}
|
||||
|
||||
// reportedSize is what the reader says about itself, before ReaderSize checks whether it can back it up.
|
||||
func reportedSize(r io.ReaderAt) (int64, bool) {
|
||||
if sr, ok := r.(interface{ Size() int64 }); ok {
|
||||
size := sr.Size()
|
||||
return size, size > 0
|
||||
}
|
||||
s, ok := r.(io.Seeker)
|
||||
if !ok {
|
||||
return 0, false
|
||||
}
|
||||
cur, err := s.Seek(0, io.SeekCurrent)
|
||||
if err != nil {
|
||||
return 0, false
|
||||
}
|
||||
size, err := s.Seek(0, io.SeekEnd)
|
||||
if err != nil {
|
||||
return 0, false
|
||||
}
|
||||
if _, err := s.Seek(cur, io.SeekStart); err != nil {
|
||||
return 0, false
|
||||
}
|
||||
return size, size > 0
|
||||
}
|
||||
@@ -0,0 +1,140 @@
|
||||
package file
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
func TestReaderSize(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
r io.ReaderAt
|
||||
want int64
|
||||
wantOK bool
|
||||
}{
|
||||
{
|
||||
name: "bytes.Reader answers directly",
|
||||
r: bytes.NewReader(make([]byte, 1234)),
|
||||
want: 1234,
|
||||
wantOK: true,
|
||||
},
|
||||
{
|
||||
// an empty reader is not a size to bound against, so it reports not-ok along with everything
|
||||
// else that cannot be measured
|
||||
name: "empty bytes.Reader",
|
||||
r: bytes.NewReader(nil),
|
||||
want: 0,
|
||||
},
|
||||
{
|
||||
name: "SectionReader over a real length",
|
||||
r: io.NewSectionReader(bytes.NewReader(make([]byte, 500)), 100, 300),
|
||||
want: 300,
|
||||
wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "seek-only reader falls back to seeking",
|
||||
r: seekOnly{bytes.NewReader(make([]byte, 4096))},
|
||||
want: 4096,
|
||||
wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "reader with neither Size nor Seek cannot be measured",
|
||||
r: readAtOnly{bytes.NewReader(make([]byte, 4096))},
|
||||
want: 0,
|
||||
},
|
||||
{
|
||||
// the shape debug/elf, debug/pe and debug/macho build internally. Taking this at its word hands
|
||||
// back a bound that permits everything, reported as if it had been measured.
|
||||
name: "SectionReader over a nominal length reports what it can deliver",
|
||||
r: io.NewSectionReader(bytes.NewReader(make([]byte, 500)), 0, 1<<63-1),
|
||||
want: 0,
|
||||
},
|
||||
{
|
||||
// a Size() a type simply got wrong is the same failure: the bytes are what decides
|
||||
name: "a size the reader cannot back up is refused",
|
||||
r: overstatedSize{ReaderAt: bytes.NewReader(make([]byte, 10)), size: 1 << 30},
|
||||
want: 0,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
size, ok := ReaderSize(tt.r)
|
||||
assert.Equal(t, tt.want, size)
|
||||
assert.Equal(t, tt.wantOK, ok)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestReaderSize_FileHandle(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "f.bin")
|
||||
require.NoError(t, os.WriteFile(path, make([]byte, 7777), 0o600))
|
||||
f, err := os.Open(path)
|
||||
require.NoError(t, err)
|
||||
defer f.Close()
|
||||
|
||||
// the cursor has to come back where it was found: the caller may be mid-read, and an *os.File is the
|
||||
// one input here that carries a shared position
|
||||
_, err = f.Seek(101, io.SeekStart)
|
||||
require.NoError(t, err)
|
||||
|
||||
size, ok := ReaderSize(f)
|
||||
assert.True(t, ok)
|
||||
assert.Equal(t, int64(7777), size)
|
||||
|
||||
at, err := f.Seek(0, io.SeekCurrent)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, int64(101), at, "measuring the file must not move the caller's cursor")
|
||||
}
|
||||
|
||||
// TestReaderSize_WrapperMustForwardSize pins the shape that made every bound expressed against this
|
||||
// helper go silently inert. A type that embeds io.ReaderAt as an interface promotes only ReadAt, so it
|
||||
// answers 0 here however sizable the reader underneath is, and a caller that treats 0 as "fall back to
|
||||
// unbounded" then stops bounding anything. Forwarding Size() is what fixes it, and this is the assertion
|
||||
// that a new wrapper has to satisfy.
|
||||
func TestReaderSize_WrapperMustForwardSize(t *testing.T) {
|
||||
inner := bytes.NewReader(make([]byte, 2048))
|
||||
|
||||
size, ok := ReaderSize(embedsReaderAt{inner})
|
||||
assert.False(t, ok, "a wrapper that only embeds the interface cannot be measured")
|
||||
assert.Equal(t, int64(0), size)
|
||||
|
||||
size, ok = ReaderSize(forwardsSize{inner})
|
||||
assert.True(t, ok, "a wrapper meant to be bounded has to forward Size itself")
|
||||
assert.Equal(t, int64(2048), size)
|
||||
}
|
||||
|
||||
type seekOnly struct{ inner *bytes.Reader }
|
||||
|
||||
func (s seekOnly) ReadAt(p []byte, off int64) (int, error) { return s.inner.ReadAt(p, off) }
|
||||
func (s seekOnly) Seek(off int64, whence int) (int64, error) {
|
||||
return s.inner.Seek(off, whence)
|
||||
}
|
||||
|
||||
type readAtOnly struct{ inner *bytes.Reader }
|
||||
|
||||
func (r readAtOnly) ReadAt(p []byte, off int64) (int, error) { return r.inner.ReadAt(p, off) }
|
||||
|
||||
type embedsReaderAt struct{ io.ReaderAt }
|
||||
|
||||
type forwardsSize struct{ io.ReaderAt }
|
||||
|
||||
func (f forwardsSize) Size() int64 {
|
||||
size, _ := ReaderSize(f.ReaderAt)
|
||||
return size
|
||||
}
|
||||
|
||||
// overstatedSize answers with a size larger than the bytes behind it, the way a wrapper that returns a
|
||||
// declared length rather than a measured one would.
|
||||
type overstatedSize struct {
|
||||
io.ReaderAt
|
||||
size int64
|
||||
}
|
||||
|
||||
func (o overstatedSize) Size() int64 { return o.size }
|
||||
@@ -0,0 +1,330 @@
|
||||
// Package spillbuf provides a sparse, offset-addressed byte buffer that keeps a bounded amount in memory
|
||||
// and spills the rest to a temp file.
|
||||
//
|
||||
// It exists for the shape where output arrives out of order, at offsets the input itself declares: a
|
||||
// decompressor placing extents, an archive rebuilding a file from chunks. A plain []byte cannot do that
|
||||
// without either pre-sizing to a length the input claims or growing to the furthest offset it names, and
|
||||
// both hand a hostile input an allocation knob.
|
||||
//
|
||||
// The buffer reports only what it actually stored. Size is the contiguous run written from offset zero,
|
||||
// and a read past it is io.EOF rather than the zeros an unwritten region would otherwise hand back. That
|
||||
// distinction is the whole point: sparse storage serves holes as zeros for free, so a reader that reported
|
||||
// a length covering them would let a write offset stand in for real output. Anything sizing an allocation
|
||||
// against this reader (debug/elf's saferio does exactly that) would then allocate bytes nothing produced.
|
||||
//
|
||||
// Allocation is bounded by the memory limit rather than by how much is written: filling 1MB and filling
|
||||
// 64MB cost about the same, because everything past the limit goes to the file. Only memory is bounded
|
||||
// here: how much ends up on disk is whatever the caller writes, so bounding that stays the caller's job.
|
||||
package spillbuf
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"math"
|
||||
"os"
|
||||
"slices"
|
||||
|
||||
"github.com/anchore/syft/internal/tmpdir"
|
||||
)
|
||||
|
||||
// DefaultMemLimit is how much a buffer keeps in memory before it spills, when the caller does not say.
|
||||
// Small on purpose: the memory tier saves a temp file for the common small case, it is not a place to
|
||||
// hold output. A limit large enough to matter is a limit large enough to be an amplification path.
|
||||
const DefaultMemLimit int64 = 1 << 20 // 1MB
|
||||
|
||||
// ErrNoTempDir is returned when a buffer needs to spill and has nowhere to spill to.
|
||||
var ErrNoTempDir = errors.New("no temp dir available to spill to")
|
||||
|
||||
// Buffer is a sparse, offset-addressed buffer. The zero value is not usable; call New.
|
||||
//
|
||||
// Storage is two tiers with a fixed boundary: bytes below memLimit live in mem, bytes at or past it live
|
||||
// in file. Nothing ever moves between them, so an offset belongs to exactly one tier for the life of the
|
||||
// buffer and the only case needing care is a range that straddles the boundary.
|
||||
//
|
||||
// Not safe for concurrent use. Every consumer writes from a single goroutine, and guarding it would cost
|
||||
// each write an uncontended lock for no reader.
|
||||
type Buffer struct {
|
||||
td *tmpdir.TempDir
|
||||
memLimit int64
|
||||
|
||||
// mem holds [0, memLimit). Grown geometrically to fit what has been written, never past the limit.
|
||||
mem []byte
|
||||
|
||||
// file holds [memLimit, ...), created on the first write that reaches there, so a buffer that stays
|
||||
// small never touches disk. remove deletes it on Close.
|
||||
file *os.File
|
||||
remove func()
|
||||
|
||||
// written is the set of ranges stored, sorted by start and merged, so runs of adjacent writes collapse
|
||||
// to one entry rather than growing per call. Everything the buffer reports comes from it.
|
||||
written []extent
|
||||
|
||||
closed bool
|
||||
}
|
||||
|
||||
// extent is a half-open byte range that has been written.
|
||||
//
|
||||
// Unexported on purpose. What a caller needs to know about the contents is answered by Size and FirstGap;
|
||||
// handing out the bookkeeping invites re-deriving those at the call site and getting them subtly wrong,
|
||||
// which is the bug this package exists to prevent.
|
||||
type extent struct {
|
||||
start, end int64
|
||||
}
|
||||
|
||||
// Option configures a Buffer.
|
||||
type Option func(*Buffer)
|
||||
|
||||
// WithMemLimit sets how many bytes are held in memory before the buffer spills to disk. Zero sends every
|
||||
// write to the temp file. Negative is treated as zero.
|
||||
//
|
||||
// Peak allocation is a small multiple of this, since growing the tier holds the old copy alongside the
|
||||
// new one and the new one may be rounded up past what was asked for.
|
||||
func WithMemLimit(n int64) Option {
|
||||
return func(b *Buffer) {
|
||||
b.memLimit = max(n, 0)
|
||||
}
|
||||
}
|
||||
|
||||
// New returns a buffer that spills into td.
|
||||
//
|
||||
// td may be nil only when every write is known to stay inside the memory limit; a write past it then
|
||||
// fails with ErrNoTempDir rather than allocating. Nothing is created on disk until a write needs it.
|
||||
func New(td *tmpdir.TempDir, opts ...Option) *Buffer {
|
||||
b := &Buffer{td: td, memLimit: DefaultMemLimit}
|
||||
for _, opt := range opts {
|
||||
opt(b)
|
||||
}
|
||||
return b
|
||||
}
|
||||
|
||||
// split divides a range of n bytes starting at off across the two tiers, reporting how many bytes fall in
|
||||
// each. The file portion starts at off+inMem.
|
||||
func (b *Buffer) split(off, n int64) (inMem, inFile int64) {
|
||||
if off >= b.memLimit {
|
||||
return 0, n
|
||||
}
|
||||
inMem = min(n, b.memLimit-off)
|
||||
return inMem, n - inMem
|
||||
}
|
||||
|
||||
// fileOffset translates a buffer offset at or past the limit into an offset within the spill file.
|
||||
func (b *Buffer) fileOffset(off int64) int64 { return off - b.memLimit }
|
||||
|
||||
// WriteAt stores p at off. Writes may land anywhere, in any order, and may overlap. A write that fails
|
||||
// leaves the buffer as it was.
|
||||
func (b *Buffer) WriteAt(p []byte, off int64) (int, error) {
|
||||
if b.closed {
|
||||
return 0, fmt.Errorf("spillbuf: write to a closed buffer: %w", os.ErrClosed)
|
||||
}
|
||||
if off < 0 {
|
||||
return 0, fmt.Errorf("spillbuf: negative offset %d", off)
|
||||
}
|
||||
if len(p) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
// subtraction form so off+len(p) cannot wrap into a small in-range end
|
||||
if off > math.MaxInt64-int64(len(p)) {
|
||||
return 0, fmt.Errorf("spillbuf: write of %d bytes at %d overflows", len(p), off)
|
||||
}
|
||||
|
||||
inMem, inFile := b.split(off, int64(len(p)))
|
||||
// the fallible half goes first: once the memory tier is stamped the old bytes are gone, so a file
|
||||
// write that fails afterward would report failure over a buffer it had already changed
|
||||
if inFile > 0 {
|
||||
if err := b.writeFile(p[inMem:], b.fileOffset(off+inMem)); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
if inMem > 0 {
|
||||
b.writeMem(p[:inMem], off)
|
||||
}
|
||||
|
||||
b.record(extent{start: off, end: off + int64(len(p))})
|
||||
return len(p), nil
|
||||
}
|
||||
|
||||
func (b *Buffer) writeMem(p []byte, off int64) {
|
||||
if need := off + int64(len(p)); int64(len(b.mem)) < need {
|
||||
// geometric, capped at the limit. Sizing to exactly what each write needs looks tidier and is
|
||||
// quadratic: a caller streaming a block in 32KB chunks reallocates and copies the whole tier every
|
||||
// chunk, so filling 1MB costs ~16MB of allocation.
|
||||
grown := min(b.memLimit, max(need, 2*int64(len(b.mem))))
|
||||
// sized to exactly the geometric target rather than grown through append: append rounds the new
|
||||
// capacity up past what was asked for, and the tier is already capped, so the slack is pure waste
|
||||
mem := make([]byte, grown)
|
||||
copy(mem, b.mem)
|
||||
b.mem = mem
|
||||
}
|
||||
copy(b.mem[off:], p)
|
||||
}
|
||||
|
||||
func (b *Buffer) writeFile(p []byte, off int64) error {
|
||||
if b.file == nil {
|
||||
if b.td == nil {
|
||||
return ErrNoTempDir
|
||||
}
|
||||
f, remove, err := b.td.NewFile("syft-spill-*.bin") //nolint:gocritic // removal happens in Close
|
||||
if err != nil {
|
||||
return fmt.Errorf("spillbuf: unable to create spill file: %w", err)
|
||||
}
|
||||
b.file, b.remove = f, remove
|
||||
}
|
||||
if _, err := b.file.WriteAt(p, off); err != nil {
|
||||
return fmt.Errorf("spillbuf: unable to write to spill file: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// ReadAt returns bytes this buffer actually stored. A read starting at or past Size is io.EOF, and one
|
||||
// running past it is short, because everything beyond is a hole the buffer never held.
|
||||
func (b *Buffer) ReadAt(p []byte, off int64) (int, error) {
|
||||
if b.closed {
|
||||
return 0, fmt.Errorf("spillbuf: read from a closed buffer: %w", os.ErrClosed)
|
||||
}
|
||||
if off < 0 {
|
||||
return 0, fmt.Errorf("spillbuf: negative offset %d", off)
|
||||
}
|
||||
if len(p) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
size := b.Size()
|
||||
if off >= size {
|
||||
return 0, io.EOF
|
||||
}
|
||||
|
||||
// bounded by Size, so every byte in [off, off+want) was written and both tiers really hold theirs
|
||||
want := min(int64(len(p)), size-off)
|
||||
inMem, inFile := b.split(off, want)
|
||||
|
||||
// guarded rather than relying on a zero-length copy: when off is past the limit the slice expression
|
||||
// itself is out of range, since mem is only as long as what was written into it
|
||||
if inMem > 0 {
|
||||
copy(p[:inMem], b.mem[off:off+inMem])
|
||||
}
|
||||
if inFile > 0 {
|
||||
if err := readFullAt(b.file, p[inMem:want], b.fileOffset(off+inMem)); err != nil {
|
||||
return int(inMem), fmt.Errorf("spillbuf: unable to read spill file: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
if want < int64(len(p)) {
|
||||
return int(want), io.EOF
|
||||
}
|
||||
return int(want), nil
|
||||
}
|
||||
|
||||
// readFullAt fills p from r. A short read is an error rather than a short return: the caller has already
|
||||
// bounded the range by what was written, so the file failing to deliver means it is not what we left.
|
||||
func readFullAt(r io.ReaderAt, p []byte, off int64) error {
|
||||
for read := 0; read < len(p); {
|
||||
n, err := r.ReadAt(p[read:], off+int64(read))
|
||||
read += n
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Size reports the contiguous run written from offset zero, which is the length this buffer is willing to
|
||||
// stand behind. Deliberately not the furthest offset written: see the package doc for why the two are not
|
||||
// interchangeable.
|
||||
//
|
||||
// record keeps the set merged and sorted, so the run from zero is at most the first two entries and the
|
||||
// loop stops there: the cost this adds to every ReadAt is effectively constant, not a scan of the set.
|
||||
func (b *Buffer) Size() int64 {
|
||||
var covered int64
|
||||
for _, e := range b.written {
|
||||
if e.start > covered {
|
||||
break
|
||||
}
|
||||
covered = max(covered, e.end)
|
||||
}
|
||||
return covered
|
||||
}
|
||||
|
||||
// FirstGap returns the offset of the earliest unwritten run of at least size bytes lying entirely within
|
||||
// [0, within), and whether there is one. Callers that place output at offsets of their own choosing use it
|
||||
// to fill what they left behind, without needing the buffer's bookkeeping.
|
||||
//
|
||||
// The arithmetic is in subtraction form throughout: within comes from caller input and the result is used
|
||||
// as a write offset, so a wrap here would be a write nobody bounded.
|
||||
func (b *Buffer) FirstGap(size, within int64) (int64, bool) {
|
||||
if b.closed || size < 0 || within < 0 {
|
||||
return 0, false
|
||||
}
|
||||
|
||||
var at int64
|
||||
for _, e := range b.written {
|
||||
// once the cursor is at the bound there is nothing left to offer, and it only moves forward
|
||||
if at >= within {
|
||||
return 0, false
|
||||
}
|
||||
// the gap ahead of this extent ends at whichever comes first, the extent or the bound
|
||||
end := min(e.start, within)
|
||||
if end > at && end-at >= size {
|
||||
return at, true
|
||||
}
|
||||
at = max(at, e.end)
|
||||
}
|
||||
if at <= within && within-at >= size {
|
||||
return at, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// record merges e into the written set, keeping it sorted by start with no overlapping or adjacent pairs.
|
||||
func (b *Buffer) record(e extent) {
|
||||
if e.end <= e.start {
|
||||
return
|
||||
}
|
||||
|
||||
// inserted in place rather than appended and re-sorted: the set is already ordered, so a run of
|
||||
// ascending writes (the common case) stays linear
|
||||
i, _ := slices.BinarySearchFunc(b.written, e, func(x, y extent) int {
|
||||
return cmpInt64(x.start, y.start)
|
||||
})
|
||||
b.written = slices.Insert(b.written, i, e)
|
||||
|
||||
merged := b.written[:1]
|
||||
for _, cur := range b.written[1:] {
|
||||
last := &merged[len(merged)-1]
|
||||
if cur.start <= last.end {
|
||||
last.end = max(last.end, cur.end)
|
||||
continue
|
||||
}
|
||||
merged = append(merged, cur)
|
||||
}
|
||||
b.written = merged
|
||||
}
|
||||
|
||||
func cmpInt64(x, y int64) int {
|
||||
switch {
|
||||
case x < y:
|
||||
return -1
|
||||
case x > y:
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// Close releases the spill file, if one was ever created, and drops the memory tier. Safe on a nil
|
||||
// receiver and safe to call more than once. Reads and writes after it fail with os.ErrClosed.
|
||||
func (b *Buffer) Close() error {
|
||||
if b == nil {
|
||||
return nil
|
||||
}
|
||||
b.mem = nil
|
||||
b.written = nil
|
||||
b.closed = true
|
||||
if b.file == nil {
|
||||
return nil
|
||||
}
|
||||
err := b.file.Close()
|
||||
if b.remove != nil {
|
||||
b.remove()
|
||||
}
|
||||
b.file, b.remove = nil, nil
|
||||
return err
|
||||
}
|
||||
@@ -0,0 +1,872 @@
|
||||
package spillbuf
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"errors"
|
||||
"io"
|
||||
"math"
|
||||
"math/rand"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
"github.com/anchore/syft/internal/tmpdir"
|
||||
)
|
||||
|
||||
func newTestBuffer(t *testing.T, opts ...Option) *Buffer {
|
||||
t.Helper()
|
||||
b := New(tmpdir.FromPath(t.TempDir()), opts...)
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
return b
|
||||
}
|
||||
|
||||
func mustWrite(t *testing.T, b *Buffer, off int64, p []byte) {
|
||||
t.Helper()
|
||||
n, err := b.WriteAt(p, off)
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, len(p), n, "WriteAt must report the full write")
|
||||
}
|
||||
|
||||
// readAll returns everything the buffer is willing to hand back.
|
||||
func readAll(t *testing.T, b *Buffer) []byte {
|
||||
t.Helper()
|
||||
out := make([]byte, b.Size())
|
||||
if len(out) == 0 {
|
||||
return nil
|
||||
}
|
||||
n, err := b.ReadAt(out, 0)
|
||||
require.NoError(t, err, "a read of exactly Size bytes is not short")
|
||||
require.Equal(t, len(out), n)
|
||||
return out
|
||||
}
|
||||
|
||||
func repeat(b byte, n int) []byte { return bytes.Repeat([]byte{b}, n) }
|
||||
|
||||
// --- the core property: report only what was stored -------------------------------------------------
|
||||
|
||||
// TestSizeIsTheContiguousPrefix is the property the package exists for. A sparse buffer hands its holes
|
||||
// back as zeros it never stored, so reporting a length that covers them lets a write offset stand in for
|
||||
// real output: whatever sizes an allocation against this reader then allocates bytes nothing produced.
|
||||
func TestSizeIsTheContiguousPrefix(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
writes []extent // start is the offset, end-start the length
|
||||
wantSize int64
|
||||
}{
|
||||
{
|
||||
name: "nothing written",
|
||||
wantSize: 0,
|
||||
},
|
||||
{
|
||||
name: "one run from zero",
|
||||
writes: []extent{{0, 100}},
|
||||
wantSize: 100,
|
||||
},
|
||||
{
|
||||
name: "a run that does not start at zero covers nothing",
|
||||
writes: []extent{{10, 100}},
|
||||
wantSize: 0,
|
||||
},
|
||||
{
|
||||
name: "adjacent runs join",
|
||||
writes: []extent{{0, 50}, {50, 100}},
|
||||
wantSize: 100,
|
||||
},
|
||||
{
|
||||
name: "a hole stops the prefix",
|
||||
writes: []extent{{0, 50}, {60, 100}},
|
||||
wantSize: 50,
|
||||
},
|
||||
{
|
||||
name: "the hole is filled later",
|
||||
writes: []extent{{0, 50}, {60, 100}, {50, 60}},
|
||||
wantSize: 100,
|
||||
},
|
||||
{
|
||||
name: "out of order still resolves",
|
||||
writes: []extent{{60, 100}, {0, 50}, {50, 60}},
|
||||
wantSize: 100,
|
||||
},
|
||||
{
|
||||
name: "overlapping runs",
|
||||
writes: []extent{{0, 60}, {40, 100}},
|
||||
wantSize: 100,
|
||||
},
|
||||
{
|
||||
name: "a run fully inside another",
|
||||
writes: []extent{{0, 100}, {20, 40}},
|
||||
wantSize: 100,
|
||||
},
|
||||
{
|
||||
name: "far placement is not coverage",
|
||||
writes: []extent{{0, 64}, {1 << 20, 1<<20 + 64}},
|
||||
wantSize: 64,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
for _, w := range tt.writes {
|
||||
mustWrite(t, b, w.start, repeat('x', int(w.end-w.start)))
|
||||
}
|
||||
assert.Equal(t, tt.wantSize, b.Size(), "Size is the contiguous prefix")
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestReadPastSizeIsEOFNotZeros is the other half of the same property. Sparse storage returns zeros for
|
||||
// a hole for free; handing those out would make the hole indistinguishable from real content.
|
||||
func TestReadPastSizeIsEOFNotZeros(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
mustWrite(t, b, 0, repeat('A', 64))
|
||||
mustWrite(t, b, 4096, repeat('B', 64)) // past a hole
|
||||
|
||||
require.Equal(t, int64(64), b.Size())
|
||||
|
||||
t.Run("a read starting in the hole is EOF", func(t *testing.T) {
|
||||
p := make([]byte, 16)
|
||||
n, err := b.ReadAt(p, 100)
|
||||
assert.Zero(t, n)
|
||||
assert.ErrorIs(t, err, io.EOF)
|
||||
})
|
||||
|
||||
t.Run("a read starting at the far extent is EOF, not the bytes there", func(t *testing.T) {
|
||||
p := make([]byte, 64)
|
||||
n, err := b.ReadAt(p, 4096)
|
||||
assert.Zero(t, n)
|
||||
assert.ErrorIs(t, err, io.EOF)
|
||||
assert.Equal(t, make([]byte, 64), p, "nothing may be copied out")
|
||||
})
|
||||
|
||||
t.Run("a read running off the end is short and says so", func(t *testing.T) {
|
||||
p := make([]byte, 128)
|
||||
n, err := b.ReadAt(p, 0)
|
||||
assert.Equal(t, 64, n, "only the stored prefix comes back")
|
||||
assert.ErrorIs(t, err, io.EOF)
|
||||
assert.Equal(t, repeat('A', 64), p[:64])
|
||||
assert.Equal(t, make([]byte, 64), p[64:], "the rest is untouched")
|
||||
})
|
||||
}
|
||||
|
||||
// TestSparsePlacementIsNotAnAllocationKnob pins the amplification directly: a tiny payload parked at a
|
||||
// huge offset must not cost memory proportional to the offset.
|
||||
func TestSparsePlacementIsNotAnAllocationKnob(t *testing.T) {
|
||||
const far = 512 << 20 // 512MB out
|
||||
|
||||
b := newTestBuffer(t, WithMemLimit(1<<20))
|
||||
mustWrite(t, b, 0, repeat('A', 64))
|
||||
|
||||
var before, after runtime.MemStats
|
||||
runtime.GC()
|
||||
runtime.ReadMemStats(&before)
|
||||
mustWrite(t, b, far, repeat('B', 64))
|
||||
runtime.ReadMemStats(&after)
|
||||
|
||||
allocated := after.TotalAlloc - before.TotalAlloc
|
||||
t.Logf("writing 64 bytes at offset %d allocated %d bytes", far, allocated)
|
||||
assert.Less(t, allocated, uint64(1<<20),
|
||||
"a far placement must cost its own bytes, not its offset")
|
||||
assert.Equal(t, int64(64), b.Size(), "and it must not count toward what the buffer will deliver")
|
||||
}
|
||||
|
||||
// --- tiers ------------------------------------------------------------------------------------------
|
||||
|
||||
func TestMemoryTierNeverTouchesDisk(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
b := New(tmpdir.FromPath(dir), WithMemLimit(4096))
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
|
||||
mustWrite(t, b, 0, repeat('A', 4096)) // exactly the limit
|
||||
|
||||
assert.Equal(t, repeat('A', 4096), readAll(t, b))
|
||||
assert.Nil(t, b.file, "a buffer that stays inside the limit must not create a file")
|
||||
assert.Empty(t, dirEntries(t, dir), "and must leave nothing on disk")
|
||||
}
|
||||
|
||||
func TestSpillCreatesTheFileOnlyWhenNeeded(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
b := New(tmpdir.FromPath(dir), WithMemLimit(4096))
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
|
||||
mustWrite(t, b, 0, repeat('A', 4096))
|
||||
require.Empty(t, dirEntries(t, dir))
|
||||
|
||||
mustWrite(t, b, 4096, repeat('B', 1)) // one byte past the limit
|
||||
assert.NotNil(t, b.file)
|
||||
assert.Len(t, dirEntries(t, dir), 1, "exactly one spill file")
|
||||
}
|
||||
|
||||
func TestWriteStraddlingTheLimit(t *testing.T) {
|
||||
const limit = 1024
|
||||
b := newTestBuffer(t, WithMemLimit(limit))
|
||||
|
||||
payload := make([]byte, 2048)
|
||||
for i := range payload {
|
||||
payload[i] = byte(i % 251)
|
||||
}
|
||||
// starts below the limit and ends well past it, so the write is split across both tiers
|
||||
mustWrite(t, b, limit-512, payload)
|
||||
|
||||
require.Equal(t, int64(0), b.Size(), "nothing was written at zero yet")
|
||||
mustWrite(t, b, 0, repeat('Z', limit-512))
|
||||
|
||||
got := readAll(t, b)
|
||||
require.Len(t, got, limit-512+2048)
|
||||
assert.Equal(t, repeat('Z', limit-512), got[:limit-512], "the memory-only part")
|
||||
assert.Equal(t, payload, got[limit-512:], "the straddling part reads back whole")
|
||||
}
|
||||
|
||||
func TestMemLimitZeroSendsEverythingToDisk(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
b := New(tmpdir.FromPath(dir), WithMemLimit(0))
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
|
||||
mustWrite(t, b, 0, repeat('A', 16))
|
||||
assert.Equal(t, repeat('A', 16), readAll(t, b))
|
||||
assert.Len(t, dirEntries(t, dir), 1, "with no memory tier the first byte spills")
|
||||
assert.Nil(t, b.mem)
|
||||
}
|
||||
|
||||
func TestNegativeMemLimitIsTreatedAsZero(t *testing.T) {
|
||||
b := newTestBuffer(t, WithMemLimit(-1))
|
||||
mustWrite(t, b, 0, repeat('A', 8))
|
||||
assert.Equal(t, repeat('A', 8), readAll(t, b))
|
||||
}
|
||||
|
||||
func TestDefaultMemLimitApplies(t *testing.T) {
|
||||
b := New(nil)
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
assert.Equal(t, DefaultMemLimit, b.memLimit)
|
||||
}
|
||||
|
||||
// --- no temp dir ------------------------------------------------------------------------------------
|
||||
|
||||
func TestNoTempDir(t *testing.T) {
|
||||
t.Run("a write inside the memory limit needs no temp dir", func(t *testing.T) {
|
||||
b := New(nil, WithMemLimit(4096))
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
|
||||
mustWrite(t, b, 0, repeat('A', 4096))
|
||||
assert.Equal(t, repeat('A', 4096), readAll(t, b))
|
||||
})
|
||||
|
||||
t.Run("a write past it fails rather than allocating", func(t *testing.T) {
|
||||
b := New(nil, WithMemLimit(4096))
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
|
||||
_, err := b.WriteAt(repeat('B', 1), 4096)
|
||||
require.Error(t, err)
|
||||
assert.ErrorIs(t, err, ErrNoTempDir)
|
||||
})
|
||||
}
|
||||
|
||||
// --- input validation -------------------------------------------------------------------------------
|
||||
|
||||
func TestWriteAtRejectsBadInput(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
|
||||
t.Run("negative offset", func(t *testing.T) {
|
||||
_, err := b.WriteAt([]byte("x"), -1)
|
||||
assert.Error(t, err)
|
||||
})
|
||||
|
||||
t.Run("an offset plus length that would overflow", func(t *testing.T) {
|
||||
_, err := b.WriteAt(repeat('x', 16), math.MaxInt64-8)
|
||||
require.Error(t, err)
|
||||
assert.Contains(t, err.Error(), "overflow",
|
||||
"the check must be a subtraction, not off+len wrapping into an in-range end")
|
||||
})
|
||||
|
||||
t.Run("an empty write is a no-op, not an extent", func(t *testing.T) {
|
||||
n, err := b.WriteAt(nil, 500)
|
||||
require.NoError(t, err)
|
||||
assert.Zero(t, n)
|
||||
assert.Empty(t, b.written, "a zero-length write records nothing")
|
||||
})
|
||||
}
|
||||
|
||||
func TestReadAtRejectsBadInput(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
mustWrite(t, b, 0, repeat('A', 16))
|
||||
|
||||
t.Run("negative offset", func(t *testing.T) {
|
||||
_, err := b.ReadAt(make([]byte, 4), -1)
|
||||
assert.Error(t, err)
|
||||
})
|
||||
|
||||
t.Run("empty read", func(t *testing.T) {
|
||||
n, err := b.ReadAt(nil, 0)
|
||||
assert.NoError(t, err)
|
||||
assert.Zero(t, n)
|
||||
})
|
||||
|
||||
t.Run("read from an empty buffer", func(t *testing.T) {
|
||||
empty := newTestBuffer(t)
|
||||
n, err := empty.ReadAt(make([]byte, 4), 0)
|
||||
assert.Zero(t, n)
|
||||
assert.ErrorIs(t, err, io.EOF)
|
||||
})
|
||||
}
|
||||
|
||||
// TestReadAtHonorsTheReaderAtContract pins what io.ReaderAt requires: a short read must come with a
|
||||
// non-nil error, and a full read must not invent one. Parsers built on io.ReaderAt (debug/elf among them)
|
||||
// rely on this to tell "the file ends here" from "try again".
|
||||
func TestReadAtHonorsTheReaderAtContract(t *testing.T) {
|
||||
b := newTestBuffer(t, WithMemLimit(64))
|
||||
mustWrite(t, b, 0, repeat('A', 200)) // spans both tiers
|
||||
|
||||
for _, size := range []int{1, 63, 64, 65, 199, 200} {
|
||||
p := make([]byte, size)
|
||||
n, err := b.ReadAt(p, 0)
|
||||
assert.Equal(t, size, n, "size=%d", size)
|
||||
assert.NoError(t, err, "a read fully inside the buffer must not report EOF (size=%d)", size)
|
||||
}
|
||||
|
||||
p := make([]byte, 201)
|
||||
n, err := b.ReadAt(p, 0)
|
||||
assert.Equal(t, 200, n)
|
||||
assert.ErrorIs(t, err, io.EOF, "one byte past the end is short and must say so")
|
||||
}
|
||||
|
||||
func TestReadAtEveryOffsetAndLength(t *testing.T) {
|
||||
const limit = 32
|
||||
b := newTestBuffer(t, WithMemLimit(limit))
|
||||
|
||||
want := make([]byte, 100)
|
||||
for i := range want {
|
||||
want[i] = byte(i)
|
||||
}
|
||||
mustWrite(t, b, 0, want)
|
||||
|
||||
// every offset x length pair, so a tier-boundary off-by-one cannot hide
|
||||
for off := 0; off <= len(want); off++ {
|
||||
for l := 0; l <= len(want)-off+2; l++ {
|
||||
p := make([]byte, l)
|
||||
n, err := b.ReadAt(p, int64(off))
|
||||
|
||||
expected := len(want) - off
|
||||
if l < expected {
|
||||
expected = l
|
||||
}
|
||||
assert.Equal(t, expected, n, "off=%d len=%d", off, l)
|
||||
assert.Equal(t, want[off:off+expected], p[:expected], "off=%d len=%d", off, l)
|
||||
|
||||
switch {
|
||||
case l == 0:
|
||||
assert.NoError(t, err, "off=%d len=0", off)
|
||||
case off >= len(want):
|
||||
assert.ErrorIs(t, err, io.EOF, "off=%d len=%d", off, l)
|
||||
case l > expected:
|
||||
assert.ErrorIs(t, err, io.EOF, "off=%d len=%d", off, l)
|
||||
default:
|
||||
assert.NoError(t, err, "off=%d len=%d", off, l)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- overwrite semantics ----------------------------------------------------------------------------
|
||||
|
||||
func TestOverwrite(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
memLimit int64
|
||||
}{
|
||||
{"in memory", 1 << 20},
|
||||
{"on disk", 0},
|
||||
{"across the boundary", 8},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
b := newTestBuffer(t, WithMemLimit(tt.memLimit))
|
||||
mustWrite(t, b, 0, repeat('A', 16))
|
||||
mustWrite(t, b, 4, repeat('B', 8))
|
||||
|
||||
want := append(append(repeat('A', 4), repeat('B', 8)...), repeat('A', 4)...)
|
||||
assert.Equal(t, want, readAll(t, b), "the later write wins over the range it covers")
|
||||
assert.Equal(t, int64(16), b.Size(), "an overwrite does not extend the buffer")
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// --- extents ----------------------------------------------------------------------------------------
|
||||
|
||||
func TestWrittenExtentsAreSortedAndMerged(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
writes []extent
|
||||
want []extent
|
||||
}{
|
||||
{
|
||||
name: "disjoint stay separate, in order",
|
||||
writes: []extent{{100, 150}, {0, 50}},
|
||||
want: []extent{{0, 50}, {100, 150}},
|
||||
},
|
||||
{
|
||||
name: "adjacent merge",
|
||||
writes: []extent{{0, 50}, {50, 100}},
|
||||
want: []extent{{0, 100}},
|
||||
},
|
||||
{
|
||||
name: "overlapping merge",
|
||||
writes: []extent{{0, 60}, {40, 100}},
|
||||
want: []extent{{0, 100}},
|
||||
},
|
||||
{
|
||||
name: "a write bridging two runs collapses all three",
|
||||
writes: []extent{{0, 20}, {80, 100}, {20, 80}},
|
||||
want: []extent{{0, 100}},
|
||||
},
|
||||
{
|
||||
name: "a contained write changes nothing",
|
||||
writes: []extent{{0, 100}, {20, 40}},
|
||||
want: []extent{{0, 100}},
|
||||
},
|
||||
{
|
||||
name: "identical writes collapse",
|
||||
writes: []extent{{0, 10}, {0, 10}, {0, 10}},
|
||||
want: []extent{{0, 10}},
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
for _, w := range tt.writes {
|
||||
mustWrite(t, b, w.start, repeat('x', int(w.end-w.start)))
|
||||
}
|
||||
assert.Equal(t, tt.want, b.written)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestWrittenExtentsStayCompactUnderManyAdjacentWrites(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
for i := int64(0); i < 1000; i++ {
|
||||
mustWrite(t, b, i*8, repeat('A', 8))
|
||||
}
|
||||
assert.Len(t, b.written, 1, "a run of adjacent writes must collapse to one extent, not 1000")
|
||||
assert.Equal(t, int64(8000), b.Size())
|
||||
}
|
||||
|
||||
// --- lifecycle --------------------------------------------------------------------------------------
|
||||
|
||||
func TestCloseRemovesTheSpillFile(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
b := New(tmpdir.FromPath(dir), WithMemLimit(0))
|
||||
mustWrite(t, b, 0, repeat('A', 32))
|
||||
require.Len(t, dirEntries(t, dir), 1)
|
||||
|
||||
require.NoError(t, b.Close())
|
||||
assert.Empty(t, dirEntries(t, dir), "Close must take the spill file with it")
|
||||
}
|
||||
|
||||
func TestCloseIsIdempotent(t *testing.T) {
|
||||
b := New(tmpdir.FromPath(t.TempDir()), WithMemLimit(0))
|
||||
mustWrite(t, b, 0, repeat('A', 32))
|
||||
|
||||
require.NoError(t, b.Close())
|
||||
assert.NoError(t, b.Close(), "a second Close is a no-op, not a double-close")
|
||||
assert.NoError(t, b.Close())
|
||||
}
|
||||
|
||||
func TestCloseOnNilReceiver(t *testing.T) {
|
||||
var b *Buffer
|
||||
assert.NotPanics(t, func() {
|
||||
assert.NoError(t, b.Close())
|
||||
})
|
||||
}
|
||||
|
||||
func TestCloseWithoutASpillFile(t *testing.T) {
|
||||
b := New(nil, WithMemLimit(1<<20))
|
||||
mustWrite(t, b, 0, repeat('A', 32))
|
||||
assert.NoError(t, b.Close(), "a memory-only buffer closes cleanly with no file to remove")
|
||||
}
|
||||
|
||||
func TestWriteAfterCloseFails(t *testing.T) {
|
||||
b := New(tmpdir.FromPath(t.TempDir()))
|
||||
mustWrite(t, b, 0, repeat('A', 8))
|
||||
require.NoError(t, b.Close())
|
||||
|
||||
_, err := b.WriteAt(repeat('B', 8), 0)
|
||||
require.Error(t, err, "a closed buffer must refuse writes rather than resurrect its file")
|
||||
assert.ErrorIs(t, err, os.ErrClosed)
|
||||
}
|
||||
|
||||
// TestReadAfterCloseFails pins that a closed buffer reads as closed rather than as empty. Close drops the
|
||||
// extents, so without the check every read would come back io.EOF and a caller could not tell a buffer it
|
||||
// still owns from one somebody already released.
|
||||
func TestReadAfterCloseFails(t *testing.T) {
|
||||
b := New(tmpdir.FromPath(t.TempDir()))
|
||||
mustWrite(t, b, 0, repeat('A', 8))
|
||||
require.NoError(t, b.Close())
|
||||
|
||||
n, err := b.ReadAt(make([]byte, 8), 0)
|
||||
assert.Zero(t, n)
|
||||
assert.ErrorIs(t, err, os.ErrClosed)
|
||||
|
||||
at, ok := b.FirstGap(8, 100)
|
||||
assert.False(t, ok, "and a closed buffer offers nowhere to write")
|
||||
assert.Zero(t, at)
|
||||
}
|
||||
|
||||
// TestWriteFailureLeavesTheBufferUnchanged pins the ordering inside WriteAt: the fallible half runs first,
|
||||
// so a spill file that cannot be created does not take the memory tier down with it.
|
||||
func TestWriteFailureLeavesTheBufferUnchanged(t *testing.T) {
|
||||
// FromPath takes the directory as given, so a path that does not exist makes file creation fail
|
||||
b := New(tmpdir.FromPath(filepath.Join(t.TempDir(), "no-such-dir")), WithMemLimit(16))
|
||||
t.Cleanup(func() { _ = b.Close() })
|
||||
|
||||
mustWrite(t, b, 0, repeat('A', 8))
|
||||
|
||||
_, err := b.WriteAt(repeat('B', 32), 0) // straddles the limit, so it needs the spill file
|
||||
require.Error(t, err)
|
||||
|
||||
assert.Equal(t, int64(8), b.Size(), "a failed write records nothing")
|
||||
assert.Nil(t, b.file, "and leaves no half-made spill file behind")
|
||||
|
||||
p := make([]byte, 8)
|
||||
n, err := b.ReadAt(p, 0)
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, 8, n)
|
||||
assert.Equal(t, repeat('A', 8), p, "the memory tier still holds what it held")
|
||||
|
||||
mustWrite(t, b, 8, repeat('C', 8)) // inside the limit, so it still needs no file
|
||||
assert.Equal(t, append(repeat('A', 8), repeat('C', 8)...), readAll(t, b))
|
||||
}
|
||||
|
||||
// --- interface compliance ---------------------------------------------------------------------------
|
||||
|
||||
func TestBufferSatisfiesTheStandardInterfaces(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
var (
|
||||
_ io.WriterAt = b
|
||||
_ io.ReaderAt = b
|
||||
_ io.Closer = b
|
||||
)
|
||||
assert.Implements(t, (*io.ReaderAt)(nil), b)
|
||||
assert.Implements(t, (*io.WriterAt)(nil), b)
|
||||
}
|
||||
|
||||
// TestWorksThroughIOOffsetWriter pins the composition every consumer uses: stream a decoded block into
|
||||
// the buffer at a chosen offset without the writer knowing where it lands.
|
||||
func TestWorksThroughIOOffsetWriter(t *testing.T) {
|
||||
b := newTestBuffer(t, WithMemLimit(16))
|
||||
|
||||
w := io.NewOffsetWriter(b, 8)
|
||||
n, err := io.Copy(w, bytes.NewReader(repeat('B', 32)))
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, int64(32), n)
|
||||
|
||||
mustWrite(t, b, 0, repeat('A', 8))
|
||||
|
||||
got := readAll(t, b)
|
||||
assert.Equal(t, append(repeat('A', 8), repeat('B', 32)...), got)
|
||||
}
|
||||
|
||||
// TestWorksWithIOSectionReader pins the other direction: parsers wrap an io.ReaderAt in a SectionReader,
|
||||
// and it must see a consistent length.
|
||||
func TestWorksWithIOSectionReader(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
mustWrite(t, b, 0, repeat('A', 100))
|
||||
|
||||
sr := io.NewSectionReader(b, 0, b.Size())
|
||||
got, err := io.ReadAll(sr)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, repeat('A', 100), got)
|
||||
}
|
||||
|
||||
// --- model-based randomized testing -----------------------------------------------------------------
|
||||
|
||||
// TestAgainstAReferenceModel drives random writes through the buffer and a dumb reference at the same
|
||||
// time, then compares every readable byte. This is what catches tier-boundary and merge bugs that
|
||||
// hand-written cases miss.
|
||||
func TestAgainstAReferenceModel(t *testing.T) {
|
||||
const universe = 8192
|
||||
|
||||
for _, limit := range []int64{0, 1, 64, 1000, 4096, universe * 2} {
|
||||
t.Run("memLimit="+itoa(limit), func(t *testing.T) {
|
||||
for seed := int64(0); seed < 40; seed++ {
|
||||
rng := rand.New(rand.NewSource(seed)) //nolint:gosec // deterministic test input, not crypto
|
||||
b := newTestBuffer(t, WithMemLimit(limit))
|
||||
|
||||
model := make([]byte, universe)
|
||||
written := make([]bool, universe)
|
||||
|
||||
for op := 0; op < 40; op++ {
|
||||
off := rng.Int63n(universe)
|
||||
l := rng.Int63n(256) + 1
|
||||
if off+l > universe {
|
||||
l = universe - off
|
||||
}
|
||||
payload := make([]byte, l)
|
||||
for i := range payload {
|
||||
payload[i] = byte(rng.Intn(256))
|
||||
}
|
||||
|
||||
mustWrite(t, b, off, payload)
|
||||
copy(model[off:], payload)
|
||||
for i := off; i < off+l; i++ {
|
||||
written[i] = true
|
||||
}
|
||||
}
|
||||
|
||||
// the reference contiguous prefix
|
||||
var want int64
|
||||
for want < universe && written[want] {
|
||||
want++
|
||||
}
|
||||
|
||||
require.Equal(t, want, b.Size(), "seed=%d limit=%d", seed, limit)
|
||||
if want == 0 {
|
||||
continue
|
||||
}
|
||||
assert.Equal(t, model[:want], readAll(t, b), "seed=%d limit=%d", seed, limit)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestRandomReadsMatchTheModel checks partial reads at arbitrary offsets, which the whole-buffer
|
||||
// comparison above cannot reach.
|
||||
func TestRandomReadsMatchTheModel(t *testing.T) {
|
||||
rng := rand.New(rand.NewSource(7)) //nolint:gosec // deterministic test input, not crypto
|
||||
b := newTestBuffer(t, WithMemLimit(97))
|
||||
|
||||
model := make([]byte, 1000)
|
||||
for i := range model {
|
||||
model[i] = byte(rng.Intn(256))
|
||||
}
|
||||
// written in random-sized chunks, in order, so the whole range is covered
|
||||
for off := 0; off < len(model); {
|
||||
l := rng.Intn(50) + 1
|
||||
if off+l > len(model) {
|
||||
l = len(model) - off
|
||||
}
|
||||
mustWrite(t, b, int64(off), model[off:off+l])
|
||||
off += l
|
||||
}
|
||||
require.Equal(t, int64(len(model)), b.Size())
|
||||
|
||||
for i := 0; i < 500; i++ {
|
||||
off := rng.Intn(len(model))
|
||||
l := rng.Intn(len(model)-off) + 1
|
||||
p := make([]byte, l)
|
||||
n, err := b.ReadAt(p, int64(off))
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, l, n)
|
||||
require.Equal(t, model[off:off+l], p, "off=%d len=%d", off, l)
|
||||
}
|
||||
}
|
||||
|
||||
// --- helpers ----------------------------------------------------------------------------------------
|
||||
|
||||
func dirEntries(t *testing.T, dir string) []os.DirEntry {
|
||||
t.Helper()
|
||||
entries, err := os.ReadDir(dir)
|
||||
if errors.Is(err, os.ErrNotExist) {
|
||||
return nil
|
||||
}
|
||||
require.NoError(t, err)
|
||||
|
||||
// tmpdir.NewFile creates the spill file inside the root it is given
|
||||
var files []os.DirEntry
|
||||
for _, e := range entries {
|
||||
if e.IsDir() {
|
||||
sub, err := os.ReadDir(filepath.Join(dir, e.Name()))
|
||||
require.NoError(t, err)
|
||||
for _, s := range sub {
|
||||
files = append(files, s)
|
||||
}
|
||||
continue
|
||||
}
|
||||
files = append(files, e)
|
||||
}
|
||||
return files
|
||||
}
|
||||
|
||||
func itoa(n int64) string {
|
||||
if n == 0 {
|
||||
return "0"
|
||||
}
|
||||
var digits []byte
|
||||
for n > 0 {
|
||||
digits = append([]byte{byte('0' + n%10)}, digits...)
|
||||
n /= 10
|
||||
}
|
||||
return string(digits)
|
||||
}
|
||||
|
||||
// TestMemoryTierGrowsAmortized is the regression test for quadratic growth. Sizing the tier to exactly
|
||||
// what each write needs reallocates and copies the whole thing per call, so a caller streaming a block in
|
||||
// small chunks (which is what io.Copy through an OffsetWriter does) pays O(n^2): filling a 1MB tier in
|
||||
// 32KB chunks cost ~16MB of allocation and showed up as a 19MB spike decompressing a 16MB payload.
|
||||
func TestMemoryTierGrowsAmortized(t *testing.T) {
|
||||
const (
|
||||
limit = 1 << 20
|
||||
chunk = 32 << 10
|
||||
)
|
||||
b := newTestBuffer(t, WithMemLimit(limit))
|
||||
|
||||
var before, after runtime.MemStats
|
||||
runtime.GC()
|
||||
runtime.ReadMemStats(&before)
|
||||
for off := int64(0); off < limit; off += chunk {
|
||||
mustWrite(t, b, off, repeat('A', chunk))
|
||||
}
|
||||
runtime.ReadMemStats(&after)
|
||||
|
||||
allocated := after.TotalAlloc - before.TotalAlloc
|
||||
t.Logf("filling a %d byte tier in %d byte chunks allocated %d bytes", limit, chunk, allocated)
|
||||
assert.Less(t, allocated, uint64(4*limit),
|
||||
"growth has to be amortized, not a fresh copy of the whole tier per write")
|
||||
assert.Equal(t, int64(limit), b.Size())
|
||||
}
|
||||
|
||||
// TestMemoryTierNeverExceedsTheLimit pins the cap that geometric growth could otherwise overshoot.
|
||||
func TestMemoryTierNeverExceedsTheLimit(t *testing.T) {
|
||||
const limit = 1000
|
||||
b := newTestBuffer(t, WithMemLimit(limit))
|
||||
|
||||
for off := int64(0); off < 4*limit; off += 100 {
|
||||
mustWrite(t, b, off, repeat('A', 100))
|
||||
assert.LessOrEqual(t, int64(len(b.mem)), int64(limit),
|
||||
"the memory tier may never grow past its own limit (off=%d)", off)
|
||||
}
|
||||
assert.Equal(t, int64(4*limit), b.Size())
|
||||
}
|
||||
|
||||
// TestFirstGap covers the placement question a caller writing at offsets of its own choosing has to ask.
|
||||
// It moved here from the UPX cataloger, which used to walk an exported extent list to answer it itself.
|
||||
func TestFirstGap(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
writes []extent
|
||||
size int64
|
||||
within int64
|
||||
wantAt int64
|
||||
wantOK bool
|
||||
}{
|
||||
{
|
||||
name: "an empty buffer starts at zero",
|
||||
size: 16,
|
||||
within: 100,
|
||||
wantAt: 0, wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "the gap between two runs comes first",
|
||||
writes: []extent{{0, 1000}, {1016, 2000}},
|
||||
size: 16,
|
||||
within: 3000,
|
||||
wantAt: 1000, wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "with the gap filled the tail is what is left",
|
||||
writes: []extent{{0, 2000}},
|
||||
size: 1000,
|
||||
within: 3000,
|
||||
wantAt: 2000, wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "a gap too small is skipped for a later one",
|
||||
writes: []extent{{0, 100}, {108, 200}},
|
||||
size: 50,
|
||||
within: 1000,
|
||||
wantAt: 200, wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "nowhere left to fit is reported rather than squeezed in",
|
||||
writes: []extent{{0, 1000}, {1016, 2000}},
|
||||
size: 1001,
|
||||
within: 3000,
|
||||
wantOK: false,
|
||||
},
|
||||
{
|
||||
name: "a run past the declared total must not wrap the remaining-space subtraction",
|
||||
writes: []extent{{0, 200}},
|
||||
size: 8,
|
||||
within: 100,
|
||||
wantOK: false,
|
||||
},
|
||||
{
|
||||
name: "exactly filling the tail is allowed",
|
||||
writes: []extent{{0, 900}},
|
||||
size: 100,
|
||||
within: 1000,
|
||||
wantAt: 900, wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "a gap between two runs is still cut off by the bound",
|
||||
writes: []extent{{0, 10}, {1000, 1010}},
|
||||
size: 45,
|
||||
within: 50,
|
||||
wantOK: false,
|
||||
},
|
||||
{
|
||||
name: "a gap between two runs that the bound still leaves room in",
|
||||
writes: []extent{{0, 10}, {1000, 1010}},
|
||||
size: 30,
|
||||
within: 50,
|
||||
wantAt: 10, wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "a negative size is refused rather than treated as zero",
|
||||
size: -1,
|
||||
within: 100,
|
||||
wantOK: false,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
b := newTestBuffer(t)
|
||||
for _, w := range tt.writes {
|
||||
mustWrite(t, b, w.start, repeat('x', int(w.end-w.start)))
|
||||
}
|
||||
|
||||
at, ok := b.FirstGap(tt.size, tt.within)
|
||||
assert.Equal(t, tt.wantOK, ok)
|
||||
if tt.wantOK {
|
||||
assert.Equal(t, tt.wantAt, at)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestFirstGapNeverOverlapsWhatWasWritten is the property behind the table: whatever offset comes back has
|
||||
// to be usable, so writing size bytes there must not land on top of anything already stored.
|
||||
func TestFirstGapNeverOverlapsWhatWasWritten(t *testing.T) {
|
||||
rng := rand.New(rand.NewSource(11)) //nolint:gosec // deterministic test input, not crypto
|
||||
const universe = 4096
|
||||
|
||||
for seed := 0; seed < 200; seed++ {
|
||||
b := newTestBuffer(t)
|
||||
occupied := make([]bool, universe)
|
||||
|
||||
for op := 0; op < 6; op++ {
|
||||
// deliberately reaches past universe too, so the bound has to cut a between-extents gap short
|
||||
off := rng.Int63n(2 * universe)
|
||||
l := rng.Int63n(200) + 1
|
||||
mustWrite(t, b, off, repeat('x', int(l)))
|
||||
for i := off; i < off+l && i < universe; i++ {
|
||||
occupied[i] = true
|
||||
}
|
||||
}
|
||||
|
||||
size := rng.Int63n(100) + 1
|
||||
at, ok := b.FirstGap(size, universe)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
require.LessOrEqual(t, at+size, int64(universe), "the result must fit inside the declared total")
|
||||
for i := at; i < at+size; i++ {
|
||||
require.False(t, occupied[i], "FirstGap returned %d for %d bytes, but %d is taken", at, size, i)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -35,6 +35,7 @@ package elfutil
|
||||
import (
|
||||
"debug/elf"
|
||||
"encoding/binary"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"math"
|
||||
@@ -55,6 +56,15 @@ import (
|
||||
// sit just under it. That is deliberate, since the sections syft actually reads are few.
|
||||
const maxDeclaredSectionSize uint64 = 128 * intFile.MB
|
||||
|
||||
// ErrDeclaredSizeExceeded marks a file this package refused to parse because a section declared more
|
||||
// decompressed bytes than maxDeclaredSectionSize allows.
|
||||
//
|
||||
// Exported so a caller can tell "syft declined to expand this" apart from "this file is broken". The
|
||||
// distinction matters where the two are reported differently: the refusal is a real gap in the SBOM and
|
||||
// worth surfacing, while a corrupt or non-Go executable is neither, and the golang cataloger sees far
|
||||
// more of the latter than the former.
|
||||
var ErrDeclaredSizeExceeded = errors.New("declared decompressed size over the limit")
|
||||
|
||||
// legacyZlibHeaderSize is the size of the .zdebug header: the "ZLIB" magic plus a big-endian size.
|
||||
const legacyZlibHeaderSize = 12
|
||||
|
||||
@@ -102,6 +112,39 @@ func NewFile(r io.ReaderAt) (*elf.File, error) {
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// CheckAllSections rejects ELF decompression bombs before debug/elf can expand one, without keeping the
|
||||
// parse.
|
||||
//
|
||||
// The attack it stops: a compressed section's header declares its own decompressed size, and debug/elf
|
||||
// believes that number, allocating it up front when the section is opened. Nothing forces the declared
|
||||
// size to match what the compressed bytes actually yield, so a small file can name an enormous one. In
|
||||
// the case this was written for, a 260KB ELF declaring a compressed .symtab drove 1.3GB of allocation.
|
||||
// This walks the section headers and refuses any reachable section declaring more than
|
||||
// maxDeclaredSectionSize, returning ErrDeclaredSizeExceeded.
|
||||
//
|
||||
// It is a strict superset of CheckSectionNameTable, which is what "All" in the name marks: reach for that
|
||||
// narrower one only where a full elf.NewFile parse on every file is not worth paying for. Like it, this
|
||||
// is for a caller whose debug/elf call is made inside another package, but it is the whole check rather
|
||||
// than the eager half: goversion reads .symtab and the string table it links, and those are expanded
|
||||
// lazily, so the name table bound alone leaves them unbounded.
|
||||
//
|
||||
// Anything that is not an ELF this package understands passes through untouched, since a caller reaching
|
||||
// for this parses other containers too (goversion takes PE and Mach-O). A file elf.NewFile cannot parse
|
||||
// passes through as well: describing a malformed file is the caller's job, not this gate's.
|
||||
func CheckAllSections(r io.ReaderAt) error {
|
||||
if _, _, ok := identify(r); !ok {
|
||||
return nil
|
||||
}
|
||||
if err := CheckSectionNameTable(r); err != nil {
|
||||
return err
|
||||
}
|
||||
f, err := elf.NewFile(r)
|
||||
if err != nil {
|
||||
return nil //nolint:nilerr // the caller's own parse reports the format error
|
||||
}
|
||||
return checkReachableSections(f)
|
||||
}
|
||||
|
||||
// checkReachableSections bounds every section syft can drive debug/elf into decompressing. This runs
|
||||
// after the parse because (*Section).Open expands lazily, so nothing has been allocated yet.
|
||||
func checkReachableSections(f *elf.File) error {
|
||||
@@ -109,8 +152,8 @@ func checkReachableSections(f *elf.File) error {
|
||||
s := f.Sections[i]
|
||||
declared, claimed := declaredSectionSize(s)
|
||||
if claimed && declared > maxDeclaredSectionSize {
|
||||
return fmt.Errorf("elf section %q declares %d decompressed bytes, over the %d byte limit",
|
||||
s.Name, declared, maxDeclaredSectionSize)
|
||||
return fmt.Errorf("%w: elf section %q declares %d decompressed bytes, over the %d byte limit",
|
||||
ErrDeclaredSizeExceeded, s.Name, declared, maxDeclaredSectionSize)
|
||||
}
|
||||
}
|
||||
return nil
|
||||
@@ -233,6 +276,10 @@ func zdebugDeclaredSize(s *elf.Section) (uint64, bool) {
|
||||
// debug/elf call is made for them inside another package, debug/buildinfo being the one syft reaches:
|
||||
// gate the reader on this and the eager section-name table read is bounded, which is everything that
|
||||
// package expands today, since it reaches .go.buildinfo through the program headers instead.
|
||||
//
|
||||
// This is deliberately the eager half only, so it is the wrong gate for a caller whose parser goes on to
|
||||
// read sections of its own. Use CheckAllSections for those: a package that reads symbols expands sections
|
||||
// this never looks at, and the difference is the whole reason both exist.
|
||||
func CheckSectionNameTable(r io.ReaderAt) error {
|
||||
class, order, ok := identify(r)
|
||||
if !ok {
|
||||
@@ -252,8 +299,8 @@ func CheckSectionNameTable(r io.ReaderAt) error {
|
||||
return nil //nolint:nilerr // truncated section; elf.NewFile will say so
|
||||
}
|
||||
if declared > maxDeclaredSectionSize {
|
||||
return fmt.Errorf("elf section name table declares %d decompressed bytes, over the %d byte limit",
|
||||
declared, maxDeclaredSectionSize)
|
||||
return fmt.Errorf("%w: elf section name table declares %d decompressed bytes, over the %d byte limit",
|
||||
ErrDeclaredSizeExceeded, declared, maxDeclaredSectionSize)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -341,6 +341,9 @@ func TestNewFile_RejectsOversizedReachableSections(t *testing.T) {
|
||||
// the fixtures are hand-assembled ELF bytes, so a bad one would satisfy require.Error on
|
||||
// its own. Naming the section proves the rejection is the one we set up, and naming the
|
||||
// limit proves it came from this package rather than from debug/elf.
|
||||
// the sentinel is what the callers key their reporting on, so it is asserted here rather
|
||||
// than only from the packages downstream that consume it
|
||||
assert.ErrorIs(t, err, ErrDeclaredSizeExceeded)
|
||||
assert.Contains(t, err.Error(), tt.wantErr)
|
||||
assert.Contains(t, err.Error(), fmt.Sprint(maxDeclaredSectionSize))
|
||||
})
|
||||
@@ -610,6 +613,87 @@ func TestNewFile_ExtendedShnumOverflowStillChecks(t *testing.T) {
|
||||
require.Error(t, err, "a bogus section count must not disable the check")
|
||||
}
|
||||
|
||||
// TestCheckAllSections covers the standalone gate, used by callers whose debug/elf parse happens inside
|
||||
// another package (goversion, here) rather than through NewFile.
|
||||
func TestCheckAllSections(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
data func(t *testing.T) []byte
|
||||
// wantErr is empty when CheckAllSections must return nil.
|
||||
wantErr string
|
||||
// nameTableMisses marks a fixture CheckSectionNameTable alone does not reject, since the oversized
|
||||
// section here is one debug/elf only expands lazily, after the parse CheckAllSections goes on to do.
|
||||
nameTableMisses bool
|
||||
}{
|
||||
{
|
||||
name: "non-ELF reader",
|
||||
data: func(t *testing.T) []byte { return []byte("not an elf") },
|
||||
},
|
||||
{
|
||||
name: "valid ident but a body elf.NewFile cannot parse",
|
||||
data: func(t *testing.T) []byte {
|
||||
full := buildELF(t, elf.ELFCLASS64, binary.LittleEndian, nil, buildOpts{})
|
||||
truncated := full[:binary.Size(elf.Header64{})+4]
|
||||
_, err := elf.NewFile(bytes.NewReader(truncated))
|
||||
require.Error(t, err, "fixture is supposed to be rejected by debug/elf")
|
||||
return truncated
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "clean small ELF",
|
||||
data: func(t *testing.T) []byte {
|
||||
data := buildELF(t, elf.ELFCLASS64, binary.LittleEndian,
|
||||
[]section{{name: ".text", typ: elf.SHT_PROGBITS, flags: elf.SHF_ALLOC}}, buildOpts{})
|
||||
_, err := elf.NewFile(bytes.NewReader(data))
|
||||
require.NoError(t, err, "fixture is supposed to be a file debug/elf accepts")
|
||||
return data
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "oversized section name table",
|
||||
data: func(t *testing.T) []byte {
|
||||
return buildELF(t, elf.ELFCLASS64, binary.LittleEndian, nil,
|
||||
buildOpts{nameTable: §ion{compressed: true, declaredSize: overLimit}})
|
||||
},
|
||||
wantErr: "section name table",
|
||||
},
|
||||
{
|
||||
name: "oversized .symtab, expanded lazily rather than during the parse",
|
||||
data: func(t *testing.T) []byte {
|
||||
data := buildELF(t, elf.ELFCLASS64, binary.LittleEndian, []section{
|
||||
{name: ".symtab", typ: elf.SHT_SYMTAB, link: 2, compressed: true, declaredSize: overLimit},
|
||||
{name: ".strtab", typ: elf.SHT_STRTAB},
|
||||
}, buildOpts{})
|
||||
_, err := elf.NewFile(bytes.NewReader(data))
|
||||
require.NoError(t, err, "fixture is supposed to be a file debug/elf accepts")
|
||||
return data
|
||||
},
|
||||
wantErr: `".symtab"`,
|
||||
nameTableMisses: true,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
data := tt.data(t)
|
||||
|
||||
if tt.nameTableMisses {
|
||||
require.NoError(t, CheckSectionNameTable(bytes.NewReader(data)),
|
||||
"the name table check alone doesn't reach a section debug/elf only expands lazily")
|
||||
}
|
||||
|
||||
err := CheckAllSections(bytes.NewReader(data))
|
||||
if tt.wantErr == "" {
|
||||
require.NoError(t, err)
|
||||
return
|
||||
}
|
||||
require.Error(t, err)
|
||||
assert.ErrorIs(t, err, ErrDeclaredSizeExceeded)
|
||||
assert.Contains(t, err.Error(), tt.wantErr)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestNewFile_PassesThroughNonELF(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
|
||||
@@ -4,7 +4,11 @@ import (
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
"github.com/anchore/syft/syft/artifact"
|
||||
"github.com/anchore/syft/syft/cataloging"
|
||||
"github.com/anchore/syft/syft/pkg"
|
||||
"github.com/anchore/syft/syft/pkg/cataloger/internal/pkgtest"
|
||||
)
|
||||
|
||||
@@ -15,6 +19,11 @@ func Test_PackageCataloger_Binary(t *testing.T) {
|
||||
fixture string
|
||||
expectedPkgs []string
|
||||
expectedRels []string
|
||||
// wantErr is set where the fixture is expected to catalog cleanly. Without it the tester discards
|
||||
// the cataloger's error, so an unknown newly attached to every file in the image cannot fail this
|
||||
// table: packages and relationships still match. The packed fixture is the one that needs it, since
|
||||
// the UPX reporting policy decides per file whether a gap is worth an unknown.
|
||||
wantErr require.ErrorAssertionFunc
|
||||
}{
|
||||
{
|
||||
name: "simple module with dependencies",
|
||||
@@ -50,6 +59,7 @@ func Test_PackageCataloger_Binary(t *testing.T) {
|
||||
{
|
||||
name: "upx compressed binary",
|
||||
fixture: "image-small-upx",
|
||||
wantErr: require.NoError,
|
||||
expectedPkgs: []string{
|
||||
"anchore.io/not/real @ v1.0.0 (/run-me)",
|
||||
"github.com/andybalholm/brotli @ v1.1.1 (/run-me)",
|
||||
@@ -115,16 +125,63 @@ func Test_PackageCataloger_Binary(t *testing.T) {
|
||||
|
||||
for _, test := range tests {
|
||||
t.Run(test.name, func(t *testing.T) {
|
||||
pkgtest.NewCatalogTester().
|
||||
tester := pkgtest.NewCatalogTester().
|
||||
WithImageResolver(t, test.fixture).
|
||||
ExpectsPackageStrings(test.expectedPkgs).
|
||||
ExpectsRelationshipStrings(test.expectedRels).
|
||||
TestCataloger(t, NewGoModuleBinaryCataloger(DefaultCatalogerConfig()))
|
||||
ExpectsRelationshipStrings(test.expectedRels)
|
||||
if test.wantErr != nil {
|
||||
tester = tester.WithErrorAssertion(test.wantErr)
|
||||
}
|
||||
tester.TestCataloger(t, NewGoModuleBinaryCataloger(DefaultCatalogerConfig()))
|
||||
})
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
// Test_PackageCataloger_Binary_SymbolsFromPackedBinary covers what the packed path used to miss: the
|
||||
// pclntab is compressed along with everything else, so a UPX binary reported its packages with no symbols
|
||||
// at all until the unpacked contents were threaded through the rest of the scan instead of only the
|
||||
// build info read.
|
||||
func Test_PackageCataloger_Binary_SymbolsFromPackedBinary(t *testing.T) {
|
||||
cfg := DefaultCatalogerConfig()
|
||||
cfg.CaptureSymbols = cataloging.SymbolScopeAll
|
||||
|
||||
symbolCounts := func(t *testing.T, pkgs []pkg.Package, _ []artifact.Relationship) map[string]int {
|
||||
t.Helper()
|
||||
counts := make(map[string]int)
|
||||
for _, p := range pkgs {
|
||||
meta, ok := p.Metadata.(pkg.GolangBinaryBuildinfoEntry)
|
||||
require.True(t, ok, "unexpected metadata on %s", p.Name)
|
||||
for _, names := range meta.Symbols {
|
||||
counts[p.Name] += len(names)
|
||||
}
|
||||
}
|
||||
return counts
|
||||
}
|
||||
|
||||
var packed, unpacked map[string]int
|
||||
pkgtest.NewCatalogTester().
|
||||
WithImageResolver(t, "image-small").
|
||||
ExpectsAssertion(func(t *testing.T, pkgs []pkg.Package, rels []artifact.Relationship) {
|
||||
unpacked = symbolCounts(t, pkgs, rels)
|
||||
}).
|
||||
TestCataloger(t, NewGoModuleBinaryCataloger(cfg))
|
||||
|
||||
pkgtest.NewCatalogTester().
|
||||
WithImageResolver(t, "image-small-upx").
|
||||
// a real `upx --best --lzma` binary must unpack without leaving a gap behind. The reconstruction is
|
||||
// truncated to the contiguous extent rebuilt, and a chain that stops short of p_filesize is now a
|
||||
// reported partial, so this is the assertion that catches that firing on well-formed output.
|
||||
WithErrorAssertion(require.NoError).
|
||||
ExpectsAssertion(func(t *testing.T, pkgs []pkg.Package, rels []artifact.Relationship) {
|
||||
packed = symbolCounts(t, pkgs, rels)
|
||||
}).
|
||||
TestCataloger(t, NewGoModuleBinaryCataloger(cfg))
|
||||
|
||||
require.NotEmpty(t, unpacked, "the unpacked fixture is the control: it must carry symbols")
|
||||
assert.Equal(t, unpacked, packed, "the packed binary must report the same symbols as the one it was packed from")
|
||||
}
|
||||
|
||||
func Test_Mod_Cataloger_Globs(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
|
||||
@@ -1,4 +0,0 @@
|
||||
xcoff
|
||||
-----
|
||||
|
||||
The code in this package comes from: https://github.com/golang/go/tree/master/src/internal/xcoff -- it was copied over to add support for [xcoff](https://en.wikipedia.org/wiki/XCOFF) binaries. Golang keeps this package as internal, forbidding its external use.
|
||||
@@ -1,688 +0,0 @@
|
||||
// Copyright 2018 The Go Authors. All rights reserved.
|
||||
// Use of this source code is governed by a BSD-style
|
||||
// license that can be found in the LICENSE file.
|
||||
|
||||
// Package xcoff implements access to XCOFF (Extended Common Object File Format) files.
|
||||
|
||||
//nolint:all
|
||||
package xcoff
|
||||
|
||||
import (
|
||||
"encoding/binary"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// SectionHeader holds information about an XCOFF section header.
|
||||
type SectionHeader struct {
|
||||
Name string
|
||||
VirtualAddress uint64
|
||||
Size uint64
|
||||
Type uint32
|
||||
Relptr uint64
|
||||
Nreloc uint32
|
||||
}
|
||||
|
||||
type Section struct {
|
||||
SectionHeader
|
||||
Relocs []Reloc
|
||||
io.ReaderAt
|
||||
sr *io.SectionReader
|
||||
}
|
||||
|
||||
// AuxiliaryCSect holds information about an XCOFF symbol in an AUX_CSECT entry.
|
||||
type AuxiliaryCSect struct {
|
||||
Length int64
|
||||
StorageMappingClass int
|
||||
SymbolType int
|
||||
}
|
||||
|
||||
// AuxiliaryFcn holds information about an XCOFF symbol in an AUX_FCN entry.
|
||||
type AuxiliaryFcn struct {
|
||||
Size int64
|
||||
}
|
||||
|
||||
type Symbol struct {
|
||||
Name string
|
||||
Value uint64
|
||||
SectionNumber int
|
||||
StorageClass int
|
||||
AuxFcn AuxiliaryFcn
|
||||
AuxCSect AuxiliaryCSect
|
||||
}
|
||||
|
||||
type Reloc struct {
|
||||
VirtualAddress uint64
|
||||
Symbol *Symbol
|
||||
Signed bool
|
||||
InstructionFixed bool
|
||||
Length uint8
|
||||
Type uint8
|
||||
}
|
||||
|
||||
// ImportedSymbol holds information about an imported XCOFF symbol.
|
||||
type ImportedSymbol struct {
|
||||
Name string
|
||||
Library string
|
||||
}
|
||||
|
||||
// FileHeader holds information about an XCOFF file header.
|
||||
type FileHeader struct {
|
||||
TargetMachine uint16
|
||||
}
|
||||
|
||||
// A File represents an open XCOFF file.
|
||||
type File struct {
|
||||
FileHeader
|
||||
Sections []*Section
|
||||
Symbols []*Symbol
|
||||
StringTable []byte
|
||||
LibraryPaths []string
|
||||
|
||||
closer io.Closer
|
||||
}
|
||||
|
||||
// Open opens the named file using os.Open and prepares it for use as an XCOFF binary.
|
||||
func Open(name string) (*File, error) {
|
||||
f, err := os.Open(name)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
ff, err := NewFile(f)
|
||||
if err != nil {
|
||||
f.Close()
|
||||
return nil, err
|
||||
}
|
||||
ff.closer = f
|
||||
return ff, nil
|
||||
}
|
||||
|
||||
// Close closes the File.
|
||||
// If the File was created using NewFile directly instead of Open,
|
||||
// Close has no effect.
|
||||
func (f *File) Close() error {
|
||||
var err error
|
||||
if f.closer != nil {
|
||||
err = f.closer.Close()
|
||||
f.closer = nil
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
// Section returns the first section with the given name, or nil if no such
|
||||
// section exists.
|
||||
// Xcoff have section's name limited to 8 bytes. Some sections like .gosymtab
|
||||
// can be trunked but this method will still find them.
|
||||
func (f *File) Section(name string) *Section {
|
||||
for _, s := range f.Sections {
|
||||
if s.Name == name || (len(name) > 8 && s.Name == name[:8]) {
|
||||
return s
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// SectionByType returns the first section in f with the
|
||||
// given type, or nil if there is no such section.
|
||||
func (f *File) SectionByType(typ uint32) *Section {
|
||||
for _, s := range f.Sections {
|
||||
if s.Type == typ {
|
||||
return s
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// cstring converts ASCII byte sequence b to string.
|
||||
// It stops once it finds 0 or reaches end of b.
|
||||
func cstring(b []byte) string {
|
||||
var i int
|
||||
for i = 0; i < len(b) && b[i] != 0; i++ {
|
||||
}
|
||||
return string(b[:i])
|
||||
}
|
||||
|
||||
// getString extracts a string from an XCOFF string table.
|
||||
func getString(st []byte, offset uint32) (string, bool) {
|
||||
if offset < 4 || int(offset) >= len(st) {
|
||||
return "", false
|
||||
}
|
||||
return cstring(st[offset:]), true
|
||||
}
|
||||
|
||||
// NewFile creates a new File for accessing an XCOFF binary in an underlying reader.
|
||||
func NewFile(r io.ReaderAt) (*File, error) {
|
||||
sr := io.NewSectionReader(r, 0, 1<<63-1)
|
||||
// Read XCOFF target machine
|
||||
var magic uint16
|
||||
if err := binary.Read(sr, binary.BigEndian, &magic); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if magic != U802TOCMAGIC && magic != U64_TOCMAGIC {
|
||||
return nil, fmt.Errorf("unrecognised XCOFF magic: 0x%x", magic)
|
||||
}
|
||||
|
||||
f := new(File)
|
||||
f.TargetMachine = magic
|
||||
|
||||
// Read XCOFF file header
|
||||
if _, err := sr.Seek(0, io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var nscns uint16
|
||||
var symptr uint64
|
||||
var nsyms int32
|
||||
var opthdr uint16
|
||||
var hdrsz int
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
fhdr := new(FileHeader32)
|
||||
if err := binary.Read(sr, binary.BigEndian, fhdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
nscns = fhdr.Fnscns
|
||||
symptr = uint64(fhdr.Fsymptr)
|
||||
nsyms = fhdr.Fnsyms
|
||||
opthdr = fhdr.Fopthdr
|
||||
hdrsz = FILHSZ_32
|
||||
case U64_TOCMAGIC:
|
||||
fhdr := new(FileHeader64)
|
||||
if err := binary.Read(sr, binary.BigEndian, fhdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
nscns = fhdr.Fnscns
|
||||
symptr = fhdr.Fsymptr
|
||||
nsyms = fhdr.Fnsyms
|
||||
opthdr = fhdr.Fopthdr
|
||||
hdrsz = FILHSZ_64
|
||||
}
|
||||
|
||||
if symptr == 0 || nsyms <= 0 {
|
||||
return nil, fmt.Errorf("no symbol table")
|
||||
}
|
||||
|
||||
// Read string table (located right after symbol table).
|
||||
offset := symptr + uint64(nsyms)*SYMESZ
|
||||
if _, err := sr.Seek(int64(offset), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// The first 4 bytes contain the length (in bytes).
|
||||
var l uint32
|
||||
if err := binary.Read(sr, binary.BigEndian, &l); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if l > 4 {
|
||||
if _, err := sr.Seek(int64(offset), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
f.StringTable = make([]byte, l)
|
||||
if _, err := io.ReadFull(sr, f.StringTable); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
|
||||
// Read section headers
|
||||
if _, err := sr.Seek(int64(hdrsz)+int64(opthdr), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
f.Sections = make([]*Section, nscns)
|
||||
for i := 0; i < int(nscns); i++ {
|
||||
var scnptr uint64
|
||||
s := new(Section)
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
shdr := new(SectionHeader32)
|
||||
if err := binary.Read(sr, binary.BigEndian, shdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
s.Name = cstring(shdr.Sname[:])
|
||||
s.VirtualAddress = uint64(shdr.Svaddr)
|
||||
s.Size = uint64(shdr.Ssize)
|
||||
scnptr = uint64(shdr.Sscnptr)
|
||||
s.Type = shdr.Sflags
|
||||
s.Relptr = uint64(shdr.Srelptr)
|
||||
s.Nreloc = uint32(shdr.Snreloc)
|
||||
case U64_TOCMAGIC:
|
||||
shdr := new(SectionHeader64)
|
||||
if err := binary.Read(sr, binary.BigEndian, shdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
s.Name = cstring(shdr.Sname[:])
|
||||
s.VirtualAddress = shdr.Svaddr
|
||||
s.Size = shdr.Ssize
|
||||
scnptr = shdr.Sscnptr
|
||||
s.Type = shdr.Sflags
|
||||
s.Relptr = shdr.Srelptr
|
||||
s.Nreloc = shdr.Snreloc
|
||||
}
|
||||
r2 := r
|
||||
if scnptr == 0 { // .bss must have all 0s
|
||||
r2 = zeroReaderAt{}
|
||||
}
|
||||
s.sr = io.NewSectionReader(r2, int64(scnptr), int64(s.Size))
|
||||
s.ReaderAt = s.sr
|
||||
f.Sections[i] = s
|
||||
}
|
||||
|
||||
// Symbol map needed by relocation
|
||||
var idxToSym = make(map[int]*Symbol)
|
||||
|
||||
// Read symbol table
|
||||
if _, err := sr.Seek(int64(symptr), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
f.Symbols = make([]*Symbol, 0)
|
||||
for i := 0; i < int(nsyms); i++ {
|
||||
var numaux int
|
||||
var ok, needAuxFcn bool
|
||||
sym := new(Symbol)
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
se := new(SymEnt32)
|
||||
if err := binary.Read(sr, binary.BigEndian, se); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
numaux = int(se.Nnumaux)
|
||||
sym.SectionNumber = int(se.Nscnum)
|
||||
sym.StorageClass = int(se.Nsclass)
|
||||
sym.Value = uint64(se.Nvalue)
|
||||
needAuxFcn = se.Ntype&SYM_TYPE_FUNC != 0 && numaux > 1
|
||||
zeroes := binary.BigEndian.Uint32(se.Nname[:4])
|
||||
if zeroes != 0 {
|
||||
sym.Name = cstring(se.Nname[:])
|
||||
} else {
|
||||
offset := binary.BigEndian.Uint32(se.Nname[4:])
|
||||
sym.Name, ok = getString(f.StringTable, offset)
|
||||
if !ok {
|
||||
goto skip
|
||||
}
|
||||
}
|
||||
case U64_TOCMAGIC:
|
||||
se := new(SymEnt64)
|
||||
if err := binary.Read(sr, binary.BigEndian, se); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
numaux = int(se.Nnumaux)
|
||||
sym.SectionNumber = int(se.Nscnum)
|
||||
sym.StorageClass = int(se.Nsclass)
|
||||
sym.Value = se.Nvalue
|
||||
needAuxFcn = se.Ntype&SYM_TYPE_FUNC != 0 && numaux > 1
|
||||
sym.Name, ok = getString(f.StringTable, se.Noffset)
|
||||
if !ok {
|
||||
goto skip
|
||||
}
|
||||
}
|
||||
if sym.StorageClass != C_EXT && sym.StorageClass != C_WEAKEXT && sym.StorageClass != C_HIDEXT {
|
||||
goto skip
|
||||
}
|
||||
// Must have at least one csect auxiliary entry.
|
||||
if numaux < 1 || i+numaux >= int(nsyms) {
|
||||
goto skip
|
||||
}
|
||||
|
||||
if sym.SectionNumber > int(nscns) {
|
||||
goto skip
|
||||
}
|
||||
if sym.SectionNumber == 0 {
|
||||
sym.Value = 0
|
||||
} else {
|
||||
sym.Value -= f.Sections[sym.SectionNumber-1].VirtualAddress
|
||||
}
|
||||
|
||||
idxToSym[i] = sym
|
||||
|
||||
// If this symbol is a function, it must retrieve its size from
|
||||
// its AUX_FCN entry.
|
||||
// It can happen that a function symbol doesn't have any AUX_FCN.
|
||||
// In this case, needAuxFcn is false and their size will be set to 0.
|
||||
if needAuxFcn {
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
aux := new(AuxFcn32)
|
||||
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sym.AuxFcn.Size = int64(aux.Xfsize)
|
||||
case U64_TOCMAGIC:
|
||||
aux := new(AuxFcn64)
|
||||
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sym.AuxFcn.Size = int64(aux.Xfsize)
|
||||
}
|
||||
}
|
||||
|
||||
// Read csect auxiliary entry (by convention, it is the last).
|
||||
if !needAuxFcn {
|
||||
if _, err := sr.Seek(int64(numaux-1)*SYMESZ, io.SeekCurrent); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
i += numaux
|
||||
numaux = 0
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
aux := new(AuxCSect32)
|
||||
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sym.AuxCSect.SymbolType = int(aux.Xsmtyp & 0x7)
|
||||
sym.AuxCSect.StorageMappingClass = int(aux.Xsmclas)
|
||||
sym.AuxCSect.Length = int64(aux.Xscnlen)
|
||||
case U64_TOCMAGIC:
|
||||
aux := new(AuxCSect64)
|
||||
if err := binary.Read(sr, binary.BigEndian, aux); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sym.AuxCSect.SymbolType = int(aux.Xsmtyp & 0x7)
|
||||
sym.AuxCSect.StorageMappingClass = int(aux.Xsmclas)
|
||||
sym.AuxCSect.Length = int64(aux.Xscnlenhi)<<32 | int64(aux.Xscnlenlo)
|
||||
}
|
||||
f.Symbols = append(f.Symbols, sym)
|
||||
skip:
|
||||
i += numaux // Skip auxiliary entries
|
||||
if _, err := sr.Seek(int64(numaux)*SYMESZ, io.SeekCurrent); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
|
||||
// Read relocations
|
||||
// Only for .data or .text section
|
||||
for _, sect := range f.Sections {
|
||||
if sect.Type != STYP_TEXT && sect.Type != STYP_DATA {
|
||||
continue
|
||||
}
|
||||
sect.Relocs = make([]Reloc, sect.Nreloc)
|
||||
if sect.Relptr == 0 {
|
||||
continue
|
||||
}
|
||||
if _, err := sr.Seek(int64(sect.Relptr), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for i := uint32(0); i < sect.Nreloc; i++ {
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
rel := new(Reloc32)
|
||||
if err := binary.Read(sr, binary.BigEndian, rel); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sect.Relocs[i].VirtualAddress = uint64(rel.Rvaddr)
|
||||
sect.Relocs[i].Symbol = idxToSym[int(rel.Rsymndx)]
|
||||
sect.Relocs[i].Type = rel.Rtype
|
||||
sect.Relocs[i].Length = rel.Rsize&0x3F + 1
|
||||
|
||||
if rel.Rsize&0x80 != 0 {
|
||||
sect.Relocs[i].Signed = true
|
||||
}
|
||||
if rel.Rsize&0x40 != 0 {
|
||||
sect.Relocs[i].InstructionFixed = true
|
||||
}
|
||||
|
||||
case U64_TOCMAGIC:
|
||||
rel := new(Reloc64)
|
||||
if err := binary.Read(sr, binary.BigEndian, rel); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sect.Relocs[i].VirtualAddress = rel.Rvaddr
|
||||
sect.Relocs[i].Symbol = idxToSym[int(rel.Rsymndx)]
|
||||
sect.Relocs[i].Type = rel.Rtype
|
||||
sect.Relocs[i].Length = rel.Rsize&0x3F + 1
|
||||
if rel.Rsize&0x80 != 0 {
|
||||
sect.Relocs[i].Signed = true
|
||||
}
|
||||
if rel.Rsize&0x40 != 0 {
|
||||
sect.Relocs[i].InstructionFixed = true
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// zeroReaderAt is ReaderAt that reads 0s.
|
||||
type zeroReaderAt struct{}
|
||||
|
||||
// ReadAt writes len(p) 0s into p.
|
||||
func (w zeroReaderAt) ReadAt(p []byte, off int64) (n int, err error) {
|
||||
for i := range p {
|
||||
p[i] = 0
|
||||
}
|
||||
return len(p), nil
|
||||
}
|
||||
|
||||
// Data reads and returns the contents of the XCOFF section s.
|
||||
func (s *Section) Data() ([]byte, error) {
|
||||
dat := make([]byte, s.sr.Size())
|
||||
n, err := s.sr.ReadAt(dat, 0)
|
||||
if n == len(dat) {
|
||||
err = nil
|
||||
}
|
||||
return dat[:n], err
|
||||
}
|
||||
|
||||
// CSect reads and returns the contents of a csect.
|
||||
// func (f *File) CSect(name string) []byte {
|
||||
// for _, sym := range f.Symbols {
|
||||
// if sym.Name == name && sym.AuxCSect.SymbolType == XTY_SD {
|
||||
// if i := sym.SectionNumber - 1; 0 <= i && i < len(f.Sections) {
|
||||
// s := f.Sections[i]
|
||||
// if sym.Value+uint64(sym.AuxCSect.Length) <= s.Size {
|
||||
// dat := make([]byte, sym.AuxCSect.Length)
|
||||
// _, err := s.sr.ReadAt(dat, int64(sym.Value))
|
||||
// if err != nil {
|
||||
// return nil
|
||||
// }
|
||||
// return dat
|
||||
// }
|
||||
// }
|
||||
// break
|
||||
// }
|
||||
// }
|
||||
// return nil
|
||||
// }
|
||||
|
||||
// func (f *File) DWARF() (*dwarf.Data, error) {
|
||||
// // There are many other DWARF sections, but these
|
||||
// // are the ones the debug/dwarf package uses.
|
||||
// // Don't bother loading others.
|
||||
// var subtypes = [...]uint32{SSUBTYP_DWABREV, SSUBTYP_DWINFO, SSUBTYP_DWLINE, SSUBTYP_DWRNGES, SSUBTYP_DWSTR}
|
||||
// var dat [len(subtypes)][]byte
|
||||
// for i, subtype := range subtypes {
|
||||
// s := f.SectionByType(STYP_DWARF | subtype)
|
||||
// if s != nil {
|
||||
// b, err := s.Data()
|
||||
// if err != nil && uint64(len(b)) < s.Size {
|
||||
// return nil, err
|
||||
// }
|
||||
// dat[i] = b
|
||||
// }
|
||||
// }
|
||||
|
||||
// abbrev, info, line, ranges, str := dat[0], dat[1], dat[2], dat[3], dat[4]
|
||||
// return dwarf.New(abbrev, nil, nil, info, line, nil, ranges, str)
|
||||
// }
|
||||
|
||||
// readImportIDs returns the import file IDs stored inside the .loader section.
|
||||
// Library name pattern is either path/base/member or base/member
|
||||
func (f *File) readImportIDs(s *Section) ([]string, error) {
|
||||
// Read loader header
|
||||
if _, err := s.sr.Seek(0, io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var istlen uint32
|
||||
var nimpid int32
|
||||
var impoff uint64
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
lhdr := new(LoaderHeader32)
|
||||
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
istlen = lhdr.Listlen
|
||||
nimpid = lhdr.Lnimpid
|
||||
impoff = uint64(lhdr.Limpoff)
|
||||
case U64_TOCMAGIC:
|
||||
lhdr := new(LoaderHeader64)
|
||||
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
istlen = lhdr.Listlen
|
||||
nimpid = lhdr.Lnimpid
|
||||
impoff = lhdr.Limpoff
|
||||
}
|
||||
|
||||
// Read loader import file ID table
|
||||
if _, err := s.sr.Seek(int64(impoff), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
table := make([]byte, istlen)
|
||||
if _, err := io.ReadFull(s.sr, table); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
offset := 0
|
||||
// First import file ID is the default LIBPATH value
|
||||
libpath := cstring(table[offset:])
|
||||
f.LibraryPaths = strings.Split(libpath, ":")
|
||||
offset += len(libpath) + 3 // 3 null bytes
|
||||
all := make([]string, 0)
|
||||
for i := 1; i < int(nimpid); i++ {
|
||||
impidpath := cstring(table[offset:])
|
||||
offset += len(impidpath) + 1
|
||||
impidbase := cstring(table[offset:])
|
||||
offset += len(impidbase) + 1
|
||||
impidmem := cstring(table[offset:])
|
||||
offset += len(impidmem) + 1
|
||||
var path string
|
||||
if len(impidpath) > 0 {
|
||||
path = impidpath + "/" + impidbase + "/" + impidmem
|
||||
} else {
|
||||
path = impidbase + "/" + impidmem
|
||||
}
|
||||
all = append(all, path)
|
||||
}
|
||||
|
||||
return all, nil
|
||||
}
|
||||
|
||||
// ImportedSymbols returns the names of all symbols
|
||||
// referred to by the binary f that are expected to be
|
||||
// satisfied by other libraries at dynamic load time.
|
||||
// It does not return weak symbols.
|
||||
func (f *File) ImportedSymbols() ([]ImportedSymbol, error) {
|
||||
s := f.SectionByType(STYP_LOADER)
|
||||
if s == nil {
|
||||
return nil, nil
|
||||
}
|
||||
// Read loader header
|
||||
if _, err := s.sr.Seek(0, io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var stlen uint32
|
||||
var stoff uint64
|
||||
var nsyms int32
|
||||
var symoff uint64
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
lhdr := new(LoaderHeader32)
|
||||
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
stlen = lhdr.Lstlen
|
||||
stoff = uint64(lhdr.Lstoff)
|
||||
nsyms = lhdr.Lnsyms
|
||||
symoff = LDHDRSZ_32
|
||||
case U64_TOCMAGIC:
|
||||
lhdr := new(LoaderHeader64)
|
||||
if err := binary.Read(s.sr, binary.BigEndian, lhdr); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
stlen = lhdr.Lstlen
|
||||
stoff = lhdr.Lstoff
|
||||
nsyms = lhdr.Lnsyms
|
||||
symoff = lhdr.Lsymoff
|
||||
}
|
||||
|
||||
// Read loader section string table
|
||||
if _, err := s.sr.Seek(int64(stoff), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
st := make([]byte, stlen)
|
||||
if _, err := io.ReadFull(s.sr, st); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// Read imported libraries
|
||||
libs, err := f.readImportIDs(s)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// Read loader symbol table
|
||||
if _, err := s.sr.Seek(int64(symoff), io.SeekStart); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
all := make([]ImportedSymbol, 0)
|
||||
for i := 0; i < int(nsyms); i++ {
|
||||
var name string
|
||||
var ifile int32
|
||||
var ok bool
|
||||
switch f.TargetMachine {
|
||||
case U802TOCMAGIC:
|
||||
ldsym := new(LoaderSymbol32)
|
||||
if err := binary.Read(s.sr, binary.BigEndian, ldsym); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if ldsym.Lsmtype&0x40 == 0 {
|
||||
continue // Imported symbols only
|
||||
}
|
||||
zeroes := binary.BigEndian.Uint32(ldsym.Lname[:4])
|
||||
if zeroes != 0 {
|
||||
name = cstring(ldsym.Lname[:])
|
||||
} else {
|
||||
offset := binary.BigEndian.Uint32(ldsym.Lname[4:])
|
||||
name, ok = getString(st, offset)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
}
|
||||
ifile = ldsym.Lifile
|
||||
case U64_TOCMAGIC:
|
||||
ldsym := new(LoaderSymbol64)
|
||||
if err := binary.Read(s.sr, binary.BigEndian, ldsym); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if ldsym.Lsmtype&0x40 == 0 {
|
||||
continue // Imported symbols only
|
||||
}
|
||||
name, ok = getString(st, ldsym.Loffset)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
ifile = ldsym.Lifile
|
||||
}
|
||||
var sym ImportedSymbol
|
||||
sym.Name = name
|
||||
if ifile >= 1 && int(ifile) <= len(libs) {
|
||||
sym.Library = libs[ifile-1]
|
||||
}
|
||||
all = append(all, sym)
|
||||
}
|
||||
|
||||
return all, nil
|
||||
}
|
||||
|
||||
// ImportedLibraries returns the names of all libraries
|
||||
// referred to by the binary f that are expected to be
|
||||
// linked with the binary at dynamic link time.
|
||||
func (f *File) ImportedLibraries() ([]string, error) {
|
||||
s := f.SectionByType(STYP_LOADER)
|
||||
if s == nil {
|
||||
return nil, nil
|
||||
}
|
||||
all, err := f.readImportIDs(s)
|
||||
return all, err
|
||||
}
|
||||
@@ -1,102 +0,0 @@
|
||||
// Copyright 2018 The Go Authors. All rights reserved.
|
||||
// Use of this source code is governed by a BSD-style
|
||||
// license that can be found in the LICENSE file.
|
||||
|
||||
package xcoff
|
||||
|
||||
import (
|
||||
"reflect"
|
||||
"testing"
|
||||
)
|
||||
|
||||
type fileTest struct {
|
||||
file string
|
||||
hdr FileHeader
|
||||
sections []*SectionHeader
|
||||
needed []string
|
||||
}
|
||||
|
||||
var fileTests = []fileTest{
|
||||
{
|
||||
"testdata/gcc-ppc32-aix-dwarf2-exec",
|
||||
FileHeader{U802TOCMAGIC},
|
||||
[]*SectionHeader{
|
||||
{".text", 0x10000290, 0x00000bbd, STYP_TEXT, 0x7ae6, 0x36},
|
||||
{".data", 0x20000e4d, 0x00000437, STYP_DATA, 0x7d02, 0x2b},
|
||||
{".bss", 0x20001284, 0x0000021c, STYP_BSS, 0, 0},
|
||||
{".loader", 0x00000000, 0x000004b3, STYP_LOADER, 0, 0},
|
||||
{".dwline", 0x00000000, 0x000000df, STYP_DWARF | SSUBTYP_DWLINE, 0x7eb0, 0x7},
|
||||
{".dwinfo", 0x00000000, 0x00000314, STYP_DWARF | SSUBTYP_DWINFO, 0x7ef6, 0xa},
|
||||
{".dwabrev", 0x00000000, 0x000000d6, STYP_DWARF | SSUBTYP_DWABREV, 0, 0},
|
||||
{".dwarnge", 0x00000000, 0x00000020, STYP_DWARF | SSUBTYP_DWARNGE, 0x7f5a, 0x2},
|
||||
{".dwloc", 0x00000000, 0x00000074, STYP_DWARF | SSUBTYP_DWLOC, 0, 0},
|
||||
{".debug", 0x00000000, 0x00005e4f, STYP_DEBUG, 0, 0},
|
||||
},
|
||||
[]string{"libc.a/shr.o"},
|
||||
},
|
||||
{
|
||||
"testdata/gcc-ppc64-aix-dwarf2-exec",
|
||||
FileHeader{U64_TOCMAGIC},
|
||||
[]*SectionHeader{
|
||||
{".text", 0x10000480, 0x00000afd, STYP_TEXT, 0x8322, 0x34},
|
||||
{".data", 0x20000f7d, 0x000002f3, STYP_DATA, 0x85fa, 0x25},
|
||||
{".bss", 0x20001270, 0x00000428, STYP_BSS, 0, 0},
|
||||
{".loader", 0x00000000, 0x00000535, STYP_LOADER, 0, 0},
|
||||
{".dwline", 0x00000000, 0x000000b4, STYP_DWARF | SSUBTYP_DWLINE, 0x8800, 0x4},
|
||||
{".dwinfo", 0x00000000, 0x0000036a, STYP_DWARF | SSUBTYP_DWINFO, 0x8838, 0x7},
|
||||
{".dwabrev", 0x00000000, 0x000000b5, STYP_DWARF | SSUBTYP_DWABREV, 0, 0},
|
||||
{".dwarnge", 0x00000000, 0x00000040, STYP_DWARF | SSUBTYP_DWARNGE, 0x889a, 0x2},
|
||||
{".dwloc", 0x00000000, 0x00000062, STYP_DWARF | SSUBTYP_DWLOC, 0, 0},
|
||||
{".debug", 0x00000000, 0x00006605, STYP_DEBUG, 0, 0},
|
||||
},
|
||||
[]string{"libc.a/shr_64.o"},
|
||||
},
|
||||
}
|
||||
|
||||
func TestOpen(t *testing.T) {
|
||||
for i := range fileTests {
|
||||
tt := &fileTests[i]
|
||||
|
||||
f, err := Open(tt.file)
|
||||
if err != nil {
|
||||
t.Error(err)
|
||||
continue
|
||||
}
|
||||
if !reflect.DeepEqual(f.FileHeader, tt.hdr) {
|
||||
t.Errorf("open %s:\n\thave %#v\n\twant %#v\n", tt.file, f.FileHeader, tt.hdr)
|
||||
continue
|
||||
}
|
||||
|
||||
for i, sh := range f.Sections {
|
||||
if i >= len(tt.sections) {
|
||||
break
|
||||
}
|
||||
have := &sh.SectionHeader
|
||||
want := tt.sections[i]
|
||||
if !reflect.DeepEqual(have, want) {
|
||||
t.Errorf("open %s, section %d:\n\thave %#v\n\twant %#v\n", tt.file, i, have, want)
|
||||
}
|
||||
}
|
||||
tn := len(tt.sections)
|
||||
fn := len(f.Sections)
|
||||
if tn != fn {
|
||||
t.Errorf("open %s: len(Sections) = %d, want %d", tt.file, fn, tn)
|
||||
}
|
||||
tl := tt.needed
|
||||
fl, err := f.ImportedLibraries()
|
||||
if err != nil {
|
||||
t.Error(err)
|
||||
}
|
||||
if !reflect.DeepEqual(tl, fl) {
|
||||
t.Errorf("open %s: loader import = %v, want %v", tt.file, tl, fl)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestOpenFailure(t *testing.T) {
|
||||
filename := "file.go" // not an XCOFF object file
|
||||
_, err := Open(filename) // don't crash
|
||||
if err == nil {
|
||||
t.Errorf("open %s: succeeded unexpectedly", filename)
|
||||
}
|
||||
}
|
||||
@@ -1,373 +0,0 @@
|
||||
// The code in this package comes from:
|
||||
// https://github.com/golang/go/tree/master/src/internal/xcoff
|
||||
// it was copied over to add support for xcoff binaries.
|
||||
// Golang keeps this package as internal, forbidding its external use.
|
||||
|
||||
// Copyright 2018 The Go Authors. All rights reserved.
|
||||
// Use of this source code is governed by a BSD-style
|
||||
// license that can be found in the LICENSE file.
|
||||
|
||||
//nolint:all
|
||||
package xcoff
|
||||
|
||||
// File Header.
|
||||
type FileHeader32 struct {
|
||||
Fmagic uint16 // Target machine
|
||||
Fnscns uint16 // Number of sections
|
||||
Ftimedat int32 // Time and date of file creation
|
||||
Fsymptr uint32 // Byte offset to symbol table start
|
||||
Fnsyms int32 // Number of entries in symbol table
|
||||
Fopthdr uint16 // Number of bytes in optional header
|
||||
Fflags uint16 // Flags
|
||||
}
|
||||
|
||||
type FileHeader64 struct {
|
||||
Fmagic uint16 // Target machine
|
||||
Fnscns uint16 // Number of sections
|
||||
Ftimedat int32 // Time and date of file creation
|
||||
Fsymptr uint64 // Byte offset to symbol table start
|
||||
Fopthdr uint16 // Number of bytes in optional header
|
||||
Fflags uint16 // Flags
|
||||
Fnsyms int32 // Number of entries in symbol table
|
||||
}
|
||||
|
||||
const (
|
||||
FILHSZ_32 = 20
|
||||
FILHSZ_64 = 24
|
||||
)
|
||||
const (
|
||||
U802TOCMAGIC = 0737 // AIX 32-bit XCOFF
|
||||
U64_TOCMAGIC = 0767 // AIX 64-bit XCOFF
|
||||
)
|
||||
|
||||
// Flags that describe the type of the object file.
|
||||
const (
|
||||
F_RELFLG = 0x0001
|
||||
F_EXEC = 0x0002
|
||||
F_LNNO = 0x0004
|
||||
F_FDPR_PROF = 0x0010
|
||||
F_FDPR_OPTI = 0x0020
|
||||
F_DSA = 0x0040
|
||||
F_VARPG = 0x0100
|
||||
F_DYNLOAD = 0x1000
|
||||
F_SHROBJ = 0x2000
|
||||
F_LOADONLY = 0x4000
|
||||
)
|
||||
|
||||
// Section Header.
|
||||
type SectionHeader32 struct {
|
||||
Sname [8]byte // Section name
|
||||
Spaddr uint32 // Physical address
|
||||
Svaddr uint32 // Virtual address
|
||||
Ssize uint32 // Section size
|
||||
Sscnptr uint32 // Offset in file to raw data for section
|
||||
Srelptr uint32 // Offset in file to relocation entries for section
|
||||
Slnnoptr uint32 // Offset in file to line number entries for section
|
||||
Snreloc uint16 // Number of relocation entries
|
||||
Snlnno uint16 // Number of line number entries
|
||||
Sflags uint32 // Flags to define the section type
|
||||
}
|
||||
|
||||
type SectionHeader64 struct {
|
||||
Sname [8]byte // Section name
|
||||
Spaddr uint64 // Physical address
|
||||
Svaddr uint64 // Virtual address
|
||||
Ssize uint64 // Section size
|
||||
Sscnptr uint64 // Offset in file to raw data for section
|
||||
Srelptr uint64 // Offset in file to relocation entries for section
|
||||
Slnnoptr uint64 // Offset in file to line number entries for section
|
||||
Snreloc uint32 // Number of relocation entries
|
||||
Snlnno uint32 // Number of line number entries
|
||||
Sflags uint32 // Flags to define the section type
|
||||
Spad uint32 // Needs to be 72 bytes long
|
||||
}
|
||||
|
||||
// Flags defining the section type.
|
||||
const (
|
||||
STYP_DWARF = 0x0010
|
||||
STYP_TEXT = 0x0020
|
||||
STYP_DATA = 0x0040
|
||||
STYP_BSS = 0x0080
|
||||
STYP_EXCEPT = 0x0100
|
||||
STYP_INFO = 0x0200
|
||||
STYP_TDATA = 0x0400
|
||||
STYP_TBSS = 0x0800
|
||||
STYP_LOADER = 0x1000
|
||||
STYP_DEBUG = 0x2000
|
||||
STYP_TYPCHK = 0x4000
|
||||
STYP_OVRFLO = 0x8000
|
||||
)
|
||||
const (
|
||||
SSUBTYP_DWINFO = 0x10000 // DWARF info section
|
||||
SSUBTYP_DWLINE = 0x20000 // DWARF line-number section
|
||||
SSUBTYP_DWPBNMS = 0x30000 // DWARF public names section
|
||||
SSUBTYP_DWPBTYP = 0x40000 // DWARF public types section
|
||||
SSUBTYP_DWARNGE = 0x50000 // DWARF aranges section
|
||||
SSUBTYP_DWABREV = 0x60000 // DWARF abbreviation section
|
||||
SSUBTYP_DWSTR = 0x70000 // DWARF strings section
|
||||
SSUBTYP_DWRNGES = 0x80000 // DWARF ranges section
|
||||
SSUBTYP_DWLOC = 0x90000 // DWARF location lists section
|
||||
SSUBTYP_DWFRAME = 0xA0000 // DWARF frames section
|
||||
SSUBTYP_DWMAC = 0xB0000 // DWARF macros section
|
||||
)
|
||||
|
||||
// Symbol Table Entry.
|
||||
type SymEnt32 struct {
|
||||
Nname [8]byte // Symbol name
|
||||
Nvalue uint32 // Symbol value
|
||||
Nscnum int16 // Section number of symbol
|
||||
Ntype uint16 // Basic and derived type specification
|
||||
Nsclass int8 // Storage class of symbol
|
||||
Nnumaux int8 // Number of auxiliary entries
|
||||
}
|
||||
|
||||
type SymEnt64 struct {
|
||||
Nvalue uint64 // Symbol value
|
||||
Noffset uint32 // Offset of the name in string table or .debug section
|
||||
Nscnum int16 // Section number of symbol
|
||||
Ntype uint16 // Basic and derived type specification
|
||||
Nsclass int8 // Storage class of symbol
|
||||
Nnumaux int8 // Number of auxiliary entries
|
||||
}
|
||||
|
||||
const SYMESZ = 18
|
||||
|
||||
const (
|
||||
// Nscnum
|
||||
N_DEBUG = -2
|
||||
N_ABS = -1
|
||||
N_UNDEF = 0
|
||||
|
||||
//Ntype
|
||||
SYM_V_INTERNAL = 0x1000
|
||||
SYM_V_HIDDEN = 0x2000
|
||||
SYM_V_PROTECTED = 0x3000
|
||||
SYM_V_EXPORTED = 0x4000
|
||||
SYM_TYPE_FUNC = 0x0020 // is function
|
||||
)
|
||||
|
||||
// Storage Class.
|
||||
const (
|
||||
C_NULL = 0 // Symbol table entry marked for deletion
|
||||
C_EXT = 2 // External symbol
|
||||
C_STAT = 3 // Static symbol
|
||||
C_BLOCK = 100 // Beginning or end of inner block
|
||||
C_FCN = 101 // Beginning or end of function
|
||||
C_FILE = 103 // Source file name and compiler information
|
||||
C_HIDEXT = 107 // Unnamed external symbol
|
||||
C_BINCL = 108 // Beginning of include file
|
||||
C_EINCL = 109 // End of include file
|
||||
C_WEAKEXT = 111 // Weak external symbol
|
||||
C_DWARF = 112 // DWARF symbol
|
||||
C_GSYM = 128 // Global variable
|
||||
C_LSYM = 129 // Automatic variable allocated on stack
|
||||
C_PSYM = 130 // Argument to subroutine allocated on stack
|
||||
C_RSYM = 131 // Register variable
|
||||
C_RPSYM = 132 // Argument to function or procedure stored in register
|
||||
C_STSYM = 133 // Statically allocated symbol
|
||||
C_BCOMM = 135 // Beginning of common block
|
||||
C_ECOML = 136 // Local member of common block
|
||||
C_ECOMM = 137 // End of common block
|
||||
C_DECL = 140 // Declaration of object
|
||||
C_ENTRY = 141 // Alternate entry
|
||||
C_FUN = 142 // Function or procedure
|
||||
C_BSTAT = 143 // Beginning of static block
|
||||
C_ESTAT = 144 // End of static block
|
||||
C_GTLS = 145 // Global thread-local variable
|
||||
C_STTLS = 146 // Static thread-local variable
|
||||
)
|
||||
|
||||
// File Auxiliary Entry
|
||||
type AuxFile64 struct {
|
||||
Xfname [8]byte // Name or offset inside string table
|
||||
Xftype uint8 // Source file string type
|
||||
Xauxtype uint8 // Type of auxiliary entry
|
||||
}
|
||||
|
||||
// Function Auxiliary Entry
|
||||
type AuxFcn32 struct {
|
||||
Xexptr uint32 // File offset to exception table entry
|
||||
Xfsize uint32 // Size of function in bytes
|
||||
Xlnnoptr uint32 // File pointer to line number
|
||||
Xendndx uint32 // Symbol table index of next entry
|
||||
Xpad uint16 // Unused
|
||||
}
|
||||
type AuxFcn64 struct {
|
||||
Xlnnoptr uint64 // File pointer to line number
|
||||
Xfsize uint32 // Size of function in bytes
|
||||
Xendndx uint32 // Symbol table index of next entry
|
||||
Xpad uint8 // Unused
|
||||
Xauxtype uint8 // Type of auxiliary entry
|
||||
}
|
||||
|
||||
type AuxSect64 struct {
|
||||
Xscnlen uint64 // section length
|
||||
Xnreloc uint64 // Num RLDs
|
||||
pad uint8
|
||||
Xauxtype uint8 // Type of auxiliary entry
|
||||
}
|
||||
|
||||
// csect Auxiliary Entry.
|
||||
type AuxCSect32 struct {
|
||||
Xscnlen int32 // Length or symbol table index
|
||||
Xparmhash uint32 // Offset of parameter type-check string
|
||||
Xsnhash uint16 // .typchk section number
|
||||
Xsmtyp uint8 // Symbol alignment and type
|
||||
Xsmclas uint8 // Storage-mapping class
|
||||
Xstab uint32 // Reserved
|
||||
Xsnstab uint16 // Reserved
|
||||
}
|
||||
|
||||
type AuxCSect64 struct {
|
||||
Xscnlenlo uint32 // Lower 4 bytes of length or symbol table index
|
||||
Xparmhash uint32 // Offset of parameter type-check string
|
||||
Xsnhash uint16 // .typchk section number
|
||||
Xsmtyp uint8 // Symbol alignment and type
|
||||
Xsmclas uint8 // Storage-mapping class
|
||||
Xscnlenhi int32 // Upper 4 bytes of length or symbol table index
|
||||
Xpad uint8 // Unused
|
||||
Xauxtype uint8 // Type of auxiliary entry
|
||||
}
|
||||
|
||||
// Auxiliary type
|
||||
// const (
|
||||
// _AUX_EXCEPT = 255
|
||||
// _AUX_FCN = 254
|
||||
// _AUX_SYM = 253
|
||||
// _AUX_FILE = 252
|
||||
// _AUX_CSECT = 251
|
||||
// _AUX_SECT = 250
|
||||
// )
|
||||
|
||||
// Symbol type field.
|
||||
const (
|
||||
XTY_ER = 0 // External reference
|
||||
XTY_SD = 1 // Section definition
|
||||
XTY_LD = 2 // Label definition
|
||||
XTY_CM = 3 // Common csect definition
|
||||
)
|
||||
|
||||
// Defines for File auxiliary definitions: x_ftype field of x_file
|
||||
const (
|
||||
XFT_FN = 0 // Source File Name
|
||||
XFT_CT = 1 // Compile Time Stamp
|
||||
XFT_CV = 2 // Compiler Version Number
|
||||
XFT_CD = 128 // Compiler Defined Information
|
||||
)
|
||||
|
||||
// Storage-mapping class.
|
||||
const (
|
||||
XMC_PR = 0 // Program code
|
||||
XMC_RO = 1 // Read-only constant
|
||||
XMC_DB = 2 // Debug dictionary table
|
||||
XMC_TC = 3 // TOC entry
|
||||
XMC_UA = 4 // Unclassified
|
||||
XMC_RW = 5 // Read/Write data
|
||||
XMC_GL = 6 // Global linkage
|
||||
XMC_XO = 7 // Extended operation
|
||||
XMC_SV = 8 // 32-bit supervisor call descriptor
|
||||
XMC_BS = 9 // BSS class
|
||||
XMC_DS = 10 // Function descriptor
|
||||
XMC_UC = 11 // Unnamed FORTRAN common
|
||||
XMC_TC0 = 15 // TOC anchor
|
||||
XMC_TD = 16 // Scalar data entry in the TOC
|
||||
XMC_SV64 = 17 // 64-bit supervisor call descriptor
|
||||
XMC_SV3264 = 18 // Supervisor call descriptor for both 32-bit and 64-bit
|
||||
XMC_TL = 20 // Read/Write thread-local data
|
||||
XMC_UL = 21 // Read/Write thread-local data (.tbss)
|
||||
XMC_TE = 22 // TOC entry
|
||||
)
|
||||
|
||||
// Loader Header.
|
||||
type LoaderHeader32 struct {
|
||||
Lversion int32 // Loader section version number
|
||||
Lnsyms int32 // Number of symbol table entries
|
||||
Lnreloc int32 // Number of relocation table entries
|
||||
Listlen uint32 // Length of import file ID string table
|
||||
Lnimpid int32 // Number of import file IDs
|
||||
Limpoff uint32 // Offset to start of import file IDs
|
||||
Lstlen uint32 // Length of string table
|
||||
Lstoff uint32 // Offset to start of string table
|
||||
}
|
||||
|
||||
type LoaderHeader64 struct {
|
||||
Lversion int32 // Loader section version number
|
||||
Lnsyms int32 // Number of symbol table entries
|
||||
Lnreloc int32 // Number of relocation table entries
|
||||
Listlen uint32 // Length of import file ID string table
|
||||
Lnimpid int32 // Number of import file IDs
|
||||
Lstlen uint32 // Length of string table
|
||||
Limpoff uint64 // Offset to start of import file IDs
|
||||
Lstoff uint64 // Offset to start of string table
|
||||
Lsymoff uint64 // Offset to start of symbol table
|
||||
Lrldoff uint64 // Offset to start of relocation entries
|
||||
}
|
||||
|
||||
const (
|
||||
LDHDRSZ_32 = 32
|
||||
LDHDRSZ_64 = 56
|
||||
)
|
||||
|
||||
// Loader Symbol.
|
||||
type LoaderSymbol32 struct {
|
||||
Lname [8]byte // Symbol name or byte offset into string table
|
||||
Lvalue uint32 // Address field
|
||||
Lscnum int16 // Section number containing symbol
|
||||
Lsmtype int8 // Symbol type, export, import flags
|
||||
Lsmclas int8 // Symbol storage class
|
||||
Lifile int32 // Import file ID; ordinal of import file IDs
|
||||
Lparm uint32 // Parameter type-check field
|
||||
}
|
||||
|
||||
type LoaderSymbol64 struct {
|
||||
Lvalue uint64 // Address field
|
||||
Loffset uint32 // Byte offset into string table of symbol name
|
||||
Lscnum int16 // Section number containing symbol
|
||||
Lsmtype int8 // Symbol type, export, import flags
|
||||
Lsmclas int8 // Symbol storage class
|
||||
Lifile int32 // Import file ID; ordinal of import file IDs
|
||||
Lparm uint32 // Parameter type-check field
|
||||
}
|
||||
|
||||
type Reloc32 struct {
|
||||
Rvaddr uint32 // (virtual) address of reference
|
||||
Rsymndx uint32 // Index into symbol table
|
||||
Rsize uint8 // Sign and reloc bit len
|
||||
Rtype uint8 // Toc relocation type
|
||||
}
|
||||
|
||||
type Reloc64 struct {
|
||||
Rvaddr uint64 // (virtual) address of reference
|
||||
Rsymndx uint32 // Index into symbol table
|
||||
Rsize uint8 // Sign and reloc bit len
|
||||
Rtype uint8 // Toc relocation type
|
||||
}
|
||||
|
||||
const (
|
||||
R_POS = 0x00 // A(sym) Positive Relocation
|
||||
R_NEG = 0x01 // -A(sym) Negative Relocation
|
||||
R_REL = 0x02 // A(sym-*) Relative to self
|
||||
R_TOC = 0x03 // A(sym-TOC) Relative to TOC
|
||||
R_TRL = 0x12 // A(sym-TOC) TOC Relative indirect load.
|
||||
|
||||
R_TRLA = 0x13 // A(sym-TOC) TOC Rel load address. modifiable inst
|
||||
R_GL = 0x05 // A(external TOC of sym) Global Linkage
|
||||
R_TCL = 0x06 // A(local TOC of sym) Local object TOC address
|
||||
R_RL = 0x0C // A(sym) Pos indirect load. modifiable instruction
|
||||
R_RLA = 0x0D // A(sym) Pos Load Address. modifiable instruction
|
||||
R_REF = 0x0F // AL0(sym) Non relocating ref. No garbage collect
|
||||
R_BA = 0x08 // A(sym) Branch absolute. Cannot modify instruction
|
||||
R_RBA = 0x18 // A(sym) Branch absolute. modifiable instruction
|
||||
R_BR = 0x0A // A(sym-*) Branch rel to self. non modifiable
|
||||
R_RBR = 0x1A // A(sym-*) Branch rel to self. modifiable instr
|
||||
|
||||
R_TLS = 0x20 // General-dynamic reference to TLS symbol
|
||||
R_TLS_IE = 0x21 // Initial-exec reference to TLS symbol
|
||||
R_TLS_LD = 0x22 // Local-dynamic reference to TLS symbol
|
||||
R_TLS_LE = 0x23 // Local-exec reference to TLS symbol
|
||||
R_TLSM = 0x24 // Module reference to TLS symbol
|
||||
R_TLSML = 0x25 // Module reference to local (own) module
|
||||
|
||||
R_TOCU = 0x30 // Relative to TOC - high order bits
|
||||
R_TOCL = 0x31 // Relative to TOC - low order bits
|
||||
)
|
||||
@@ -5,12 +5,14 @@ import (
|
||||
"context"
|
||||
"debug/macho"
|
||||
"debug/pe"
|
||||
"encoding/binary"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"regexp"
|
||||
"runtime/debug"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
@@ -25,7 +27,6 @@ import (
|
||||
"github.com/anchore/syft/syft/internal/unionreader"
|
||||
"github.com/anchore/syft/syft/pkg"
|
||||
"github.com/anchore/syft/syft/pkg/cataloger/generic"
|
||||
"github.com/anchore/syft/syft/pkg/cataloger/golang/internal/xcoff"
|
||||
)
|
||||
|
||||
const goArch = "GOARCH"
|
||||
@@ -120,12 +121,19 @@ func (c *goBinaryCataloger) parseGoBinary(ctx context.Context, resolver file.Res
|
||||
}
|
||||
defer internal.CloseAndLogError(reader.ReadCloser, reader.RealPath)
|
||||
|
||||
mods, errs := scanFile(reader.Location, unionReader, c.symbolSelector.enabled())
|
||||
mods, errs := scanFile(ctx, reader.Location, unionReader, c.symbolSelector.enabled())
|
||||
// scanFile hands back the reconstruction for each binary that turned out to be packed. The readers
|
||||
// below want it too, so ownership ends here and not inside the scan.
|
||||
defer func() {
|
||||
for _, mod := range mods {
|
||||
closeUnpacked(mod.unpacked)
|
||||
}
|
||||
}()
|
||||
|
||||
var rels []artifact.Relationship
|
||||
for _, mod := range mods {
|
||||
var depPkgs []pkg.Package
|
||||
mainPkg, depPkgs := c.buildGoPkgInfo(ctx, resolver, reader.Location, mod, mod.arch, unionReader)
|
||||
mainPkg, depPkgs := c.buildGoPkgInfo(ctx, resolver, reader.Location, mod, mod.arch, seekerFor(mod.unpacked, unionReader))
|
||||
if mainPkg != nil {
|
||||
rels = createModuleRelationships(*mainPkg, depPkgs)
|
||||
pkgs = append(pkgs, *mainPkg)
|
||||
@@ -171,7 +179,7 @@ func moduleEqual(lhs, rhs *debug.Module) bool {
|
||||
var emptyModule debug.Module
|
||||
var moduleFromPartialPackageBuild = debug.Module{Path: "command-line-arguments"}
|
||||
|
||||
func (c *goBinaryCataloger) buildGoPkgInfo(ctx context.Context, resolver file.Resolver, location file.Location, mod *extendedBuildInfo, arch string, reader io.ReadSeekCloser) (*pkg.Package, []pkg.Package) {
|
||||
func (c *goBinaryCataloger) buildGoPkgInfo(ctx context.Context, resolver file.Resolver, location file.Location, mod *extendedBuildInfo, arch string, reader io.ReadSeeker) (*pkg.Package, []pkg.Package) {
|
||||
if mod == nil {
|
||||
return nil, nil
|
||||
}
|
||||
@@ -240,7 +248,7 @@ func missingMainModule(mod *extendedBuildInfo) bool {
|
||||
return mod.Main == moduleFromPartialPackageBuild
|
||||
}
|
||||
|
||||
func (c *goBinaryCataloger) makeGoMainPackage(ctx context.Context, resolver file.Resolver, mod *extendedBuildInfo, arch string, location file.Location, reader io.ReadSeekCloser, symbols map[string][]string) pkg.Package {
|
||||
func (c *goBinaryCataloger) makeGoMainPackage(ctx context.Context, resolver file.Resolver, mod *extendedBuildInfo, arch string, location file.Location, reader io.ReadSeeker, symbols map[string][]string) pkg.Package {
|
||||
gbs := getBuildSettings(mod.Settings)
|
||||
lics := c.licenseResolver.getLicenses(ctx, resolver, mod.Main.Path, mod.Main.Version)
|
||||
gover, experiments := getExperimentsFromVersion(mod.GoVersion)
|
||||
@@ -283,7 +291,7 @@ func (c *goBinaryCataloger) makeGoMainPackage(ctx context.Context, resolver file
|
||||
// the only thing that seems to work is to just look for version strings following both \x00 and \x00.L for now
|
||||
var semverPattern = regexp.MustCompile(`(\x00|\x{FFFD})(.L)?(?P<version>v?(\d+\.\d+\.\d+[-\w]*[+\w]*))\x00`)
|
||||
|
||||
func (c *goBinaryCataloger) findMainModuleVersion(metadata *pkg.GolangBinaryBuildinfoEntry, gbs pkg.KeyValues, reader io.ReadSeekCloser) string {
|
||||
func (c *goBinaryCataloger) findMainModuleVersion(metadata *pkg.GolangBinaryBuildinfoEntry, gbs pkg.KeyValues, reader io.ReadSeeker) string {
|
||||
vcsVersion, hasVersion := gbs.Get("vcs.revision")
|
||||
timestamp, hasTimestamp := gbs.Get("vcs.time")
|
||||
|
||||
@@ -416,11 +424,11 @@ func getGOARCHFromBin(r io.ReaderAt) (string, error) {
|
||||
}
|
||||
arch = f.Cpu.String()
|
||||
case bytes.HasPrefix(ident, []byte{0x01, 0xDF}) || bytes.HasPrefix(ident, []byte{0x01, 0xF7}):
|
||||
f, err := xcoff.NewFile(r)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("unrecognized file format: %w", err)
|
||||
}
|
||||
arch = fmt.Sprintf("%d", f.TargetMachine)
|
||||
// XCOFF's target machine *is* the magic we just matched on (0737 / 0767), so there is nothing to
|
||||
// parse: a full XCOFF walk would read the string table, every symbol and every relocation only to
|
||||
// hand back these two bytes. Reading them directly also keeps stripped binaries working, which a
|
||||
// walk does not since it needs a symbol table to get that far.
|
||||
arch = strconv.Itoa(int(binary.BigEndian.Uint16(ident[:2])))
|
||||
default:
|
||||
return "", errUnrecognizedFormat
|
||||
}
|
||||
|
||||
@@ -108,24 +108,27 @@ func Test_getGOARCHFromBin(t *testing.T) {
|
||||
},
|
||||
{
|
||||
name: "xcoff-32bit",
|
||||
filepath: "internal/xcoff/testdata/gcc-ppc32-aix-dwarf2-exec",
|
||||
filepath: "testdata/xcoff/gcc-ppc32-aix-dwarf2-exec",
|
||||
expected: strconv.Itoa(0x1DF),
|
||||
},
|
||||
{
|
||||
name: "xcoff-64bit",
|
||||
filepath: "internal/xcoff/testdata/gcc-ppc64-aix-dwarf2-exec",
|
||||
filepath: "testdata/xcoff/gcc-ppc64-aix-dwarf2-exec",
|
||||
expected: strconv.Itoa(0x1F7),
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
f, err := os.Open(tt.filepath)
|
||||
require.NoError(t, err)
|
||||
arch, err := getGOARCHFromBin(f)
|
||||
require.NoError(t, err, "test name: %s", tt.name)
|
||||
assert.Equal(t, tt.expected, arch)
|
||||
}
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
f, err := os.Open(tt.filepath)
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = f.Close() })
|
||||
|
||||
arch, err := getGOARCHFromBin(f)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, tt.expected, arch)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildGoPkgInfo(t *testing.T) {
|
||||
|
||||
@@ -1,7 +1,9 @@
|
||||
package golang
|
||||
|
||||
import (
|
||||
"context"
|
||||
"debug/buildinfo"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"runtime/debug"
|
||||
@@ -21,10 +23,32 @@ type extendedBuildInfo struct {
|
||||
cryptoSettings []string
|
||||
arch string
|
||||
symbols []binarySymbol
|
||||
|
||||
// unpacked is the reconstruction when this binary was UPX-packed, and nil when it was not. It is
|
||||
// carried on the build info because the readers downstream of the scan want it too, since the packed
|
||||
// bytes hold no readable version strings either. The caller of scanFile owns it and must Close it.
|
||||
unpacked unpackedContents
|
||||
}
|
||||
|
||||
// unpackedContents is the reconstruction of a UPX-packed binary, as everything downstream of the unpack
|
||||
// uses it: bytes at offsets, the length it is willing to stand behind, and the temp file underneath to
|
||||
// release when the scan is done with it. internal/spillbuf is what implements it, and naming that type in
|
||||
// these signatures would say the readers care where the bytes live, which they do not.
|
||||
//
|
||||
// A nil value means the binary was not packed, and unpackUPX only ever returns a literal nil for that: a
|
||||
// nil *spillbuf.Buffer widened into this interface is a non-nil interface holding a nil pointer, so it
|
||||
// reads as present everywhere it is checked and then panics on first use. readerFor, seekerFor and
|
||||
// closeUnpacked are the only places that check.
|
||||
type unpackedContents interface {
|
||||
io.ReaderAt
|
||||
io.Closer
|
||||
|
||||
// Size is the contiguous run rebuilt from offset zero, which is all the reconstruction stands behind.
|
||||
Size() int64
|
||||
}
|
||||
|
||||
// scanFile scans file to try to report the Go and module versions.
|
||||
func scanFile(location file.Location, reader unionreader.UnionReader, captureSymbols bool) ([]*extendedBuildInfo, error) {
|
||||
func scanFile(ctx context.Context, location file.Location, reader unionreader.UnionReader, captureSymbols bool) ([]*extendedBuildInfo, error) {
|
||||
// NOTE: multiple readers are returned to cover universal binaries, which are files
|
||||
// with more than one binary
|
||||
readers, errs := unionreader.GetReaders(reader)
|
||||
@@ -35,53 +59,134 @@ func scanFile(location file.Location, reader unionreader.UnionReader, captureSym
|
||||
|
||||
var builds []*extendedBuildInfo
|
||||
for _, r := range readers {
|
||||
bi, err := getBuildInfo(r, location)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang buildinfo")
|
||||
|
||||
continue
|
||||
build, err := scanReader(ctx, location, r, captureSymbols)
|
||||
errs = unknown.Join(errs, err)
|
||||
if build != nil {
|
||||
builds = append(builds, build)
|
||||
}
|
||||
// it's possible the reader just isn't a go binary, in which case just skip it
|
||||
if bi == nil {
|
||||
continue
|
||||
}
|
||||
|
||||
v, err := getCryptoInformation(r)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang version info")
|
||||
// don't skip this build info.
|
||||
// we can still catalog packages, even if we can't get the crypto information
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang version info: %w", err)
|
||||
}
|
||||
v = append(v, getNativeFIPSSettings(bi.Settings)...)
|
||||
arch := getGOARCH(bi.Settings)
|
||||
if arch == "" {
|
||||
arch, err = getGOARCHFromBin(r)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang arch info")
|
||||
// don't skip this build info.
|
||||
// we can still catalog packages, even if we can't get the arch information
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang arch info: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
var symbols []binarySymbol
|
||||
if captureSymbols {
|
||||
symbols, err = getSymbols(r)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang symbol info")
|
||||
// don't skip this build info.
|
||||
// we can still catalog packages, even if we can't get the symbol information
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang symbol info: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
builds = append(builds, &extendedBuildInfo{BuildInfo: bi, cryptoSettings: v, arch: arch, symbols: symbols})
|
||||
}
|
||||
return builds, errs
|
||||
}
|
||||
|
||||
// scanReader reports the build info, crypto settings, arch and symbols of a single binary. Everything is
|
||||
// read from the same reader, which is the unpacked contents when the binary turns out to be UPX-packed:
|
||||
// the packed bytes carry no readable pclntab or version strings either, so nothing downstream of the
|
||||
// build info should be looking at them.
|
||||
func scanReader(ctx context.Context, location file.Location, r io.ReaderAt, captureSymbols bool) (*extendedBuildInfo, error) {
|
||||
var errs error
|
||||
|
||||
// unpacked is nil when there was nothing to unpack; readerFor turns that into the reader every parser
|
||||
// below should use, so none of them has to ask whether this binary was packed.
|
||||
// err can arrive alongside a usable bi: a partial reconstruction still carries build info, and the
|
||||
// bytes it lost are a gap worth reporting rather than a reason to drop the binary.
|
||||
unpacked, bi, err := readContentsAndBuildInfo(ctx, r)
|
||||
|
||||
// ownership of unpacked transfers to the returned extendedBuildInfo on success. Until then it is held
|
||||
// here: the parsers below panic on malformed input (which is why getBuildInfo has a recover), and an
|
||||
// unwind past this point would leave a reconstruction of up to maxUPXOriginalSize on disk for the rest
|
||||
// of the scan: the temp root is swept, but not until the run ends.
|
||||
ownContents := true
|
||||
defer func() {
|
||||
if ownContents {
|
||||
closeUnpacked(unpacked)
|
||||
}
|
||||
}()
|
||||
|
||||
if err != nil {
|
||||
if isCancelled(err) {
|
||||
return nil, err
|
||||
}
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to fully read golang buildinfo")
|
||||
if reportableGap(err) {
|
||||
// the build info is either missing or it is not, and the same err covers both: a partial
|
||||
// reconstruction still hands back everything .go.buildinfo carried, and what it lost is the
|
||||
// non-loadable tail the readers below want. Saying "unable to read golang buildinfo" there
|
||||
// describes a failure that did not happen.
|
||||
if bi != nil {
|
||||
errs = unknown.Appendf(errs, location, "golang binary read incompletely: %w", err)
|
||||
} else {
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang buildinfo: %w", err)
|
||||
}
|
||||
}
|
||||
if bi == nil {
|
||||
return nil, errs
|
||||
}
|
||||
}
|
||||
|
||||
// it's possible the reader just isn't a go binary, in which case just skip it
|
||||
if bi == nil {
|
||||
return nil, errs
|
||||
}
|
||||
|
||||
// resolved once, here: every parser below reads the same bytes, and the choice of which reader that is
|
||||
// belongs at the top rather than repeated at each call site
|
||||
contents := readerFor(unpacked, r)
|
||||
|
||||
// everything below reads those contents through a parser that expands ELF sections, and each bounds its own
|
||||
// reads: getSymbols opens the file with elfutil.NewFile, getCryptoInformation gates on
|
||||
// elfutil.CheckAllSections. getBuildInfo's CheckSectionNameTable is not what makes these safe, since it
|
||||
// bounds only the section elf.NewFile expands as it parses.
|
||||
v, err := getCryptoInformation(contents)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang version info")
|
||||
// don't skip this build info.
|
||||
// we can still catalog packages, even if we can't get the crypto information
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang version info: %w", err)
|
||||
}
|
||||
v = append(v, getNativeFIPSSettings(bi.Settings)...)
|
||||
arch := getGOARCH(bi.Settings)
|
||||
if arch == "" {
|
||||
arch, err = getGOARCHFromBin(contents)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang arch info")
|
||||
// don't skip this build info.
|
||||
// we can still catalog packages, even if we can't get the arch information
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang arch info: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
var symbols []binarySymbol
|
||||
if captureSymbols {
|
||||
symbols, err = getSymbols(contents)
|
||||
if err != nil {
|
||||
log.WithFields("file", location.RealPath, "error", err).Trace("unable to read golang symbol info")
|
||||
// don't skip this build info.
|
||||
// we can still catalog packages, even if we can't get the symbol information
|
||||
errs = unknown.Appendf(errs, location, "unable to read golang symbol info: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
ownContents = false
|
||||
return &extendedBuildInfo{BuildInfo: bi, cryptoSettings: v, arch: arch, symbols: symbols, unpacked: unpacked}, errs
|
||||
}
|
||||
|
||||
// reportableGap reports whether err is a gap this cataloger chose to leave, rather than a file it was
|
||||
// never going to catalog. A packed file we failed to decode, a reconstruction that came up short, and a
|
||||
// file we declined to expand (from either the ELF section bound or the UPX header bounds) all cost the
|
||||
// SBOM something real. This cataloger runs on every executable in an image, so reporting anything else
|
||||
// would attach an unknown to every corrupt, truncated or non-Go binary in it, which is the noise the
|
||||
// quiet-by-default policy in upx.go exists to avoid.
|
||||
//
|
||||
// One real gap is deliberately left off this list: a UPX method we have not implemented. It is the same
|
||||
// cost to the SBOM as a decode failure, but it fires on most packed binaries rather than on rare ones.
|
||||
// See errUPXDecompress in upx.go for the reasoning.
|
||||
func reportableGap(err error) bool {
|
||||
return errors.Is(err, errUPXDecompress) ||
|
||||
errors.Is(err, errUPXSizeRefused) ||
|
||||
errors.Is(err, errUPXPartial) ||
|
||||
errors.Is(err, elfutil.ErrDeclaredSizeExceeded)
|
||||
}
|
||||
|
||||
func getCryptoInformation(reader io.ReaderAt) ([]string, error) {
|
||||
// goversion opens the file with debug/elf itself and reads .symtab plus the string table it links.
|
||||
// Those are expanded lazily, so the section-name table bound getBuildInfo already applied does not
|
||||
// reach them. CheckAllSections is the decompression-bomb gate: a compressed section declares its own
|
||||
// decompressed size and debug/elf allocates that much on open, so a 260KB ELF declaring a compressed
|
||||
// .symtab drove 1.3GB of allocation before this was here.
|
||||
if err := elfutil.CheckAllSections(reader); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
v, err := version.ReadExeFromReader(reader)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -123,6 +228,99 @@ func getNativeFIPSSettings(settings []debug.BuildSetting) []string {
|
||||
return cryptoSettings
|
||||
}
|
||||
|
||||
// readContentsAndBuildInfo returns the contents this binary should be read from, along with its build
|
||||
// info. A UPX-packed binary is unpacked first and read from the reconstruction.
|
||||
//
|
||||
// If the reconstruction carries no build info, the bytes as they were found are tried before giving up.
|
||||
// The "UPX!" magic is located by an unanchored scan over the first 8KB and every field behind it is
|
||||
// attacker-controlled, so an ordinary Go binary carrying those four bytes can be made to produce a
|
||||
// plausible header and a reconstruction of nothing. Without the retry that binary reports no packages at
|
||||
// all, which turns a false positive here into a way to hide a dependency list.
|
||||
//
|
||||
// A gap in the unpack is returned even when build info came through, so it is reported rather than papered
|
||||
// over: the readers after this one (crypto settings, arch, symbols) read the same contents, and the bytes
|
||||
// a partial reconstruction lost are exactly the non-loadable tail those depend on. That only holds when
|
||||
// the contents are a reconstruction: a gap read off the bytes as they were found is not a gap at all, and
|
||||
// both places below that hand those bytes back drop it.
|
||||
func readContentsAndBuildInfo(ctx context.Context, r io.ReaderAt) (unpackedContents, *debug.BuildInfo, error) {
|
||||
unpacked, unpackErr := unpackUPX(ctx, r)
|
||||
if isCancelled(unpackErr) {
|
||||
return unpacked, nil, unpackErr
|
||||
}
|
||||
|
||||
bi, err := getBuildInfo(readerFor(unpacked, r))
|
||||
if err == nil && bi != nil {
|
||||
if unpacked == nil {
|
||||
// nothing was rebuilt, so nothing downstream is short: every reader below this one reads the
|
||||
// same bytes getBuildInfo just read in full. unpackUPX returns the input as it was found
|
||||
// alongside a refusal (the size bounds) or a failure (block 1 not decoding, or no temp dir),
|
||||
// and readable build info in those bytes is itself the evidence the file was not really
|
||||
// packed. A genuinely packed binary carries no readable .go.buildinfo in its packed bytes, so
|
||||
// it never reaches here and its gap is still reported below.
|
||||
return unpacked, bi, nil
|
||||
}
|
||||
return unpacked, bi, unpackErr
|
||||
}
|
||||
|
||||
if unpacked != nil {
|
||||
if foundBI, foundErr := getBuildInfo(r); foundErr == nil && foundBI != nil {
|
||||
log.WithFields("error", err, "unpackError", unpackErr).
|
||||
Trace("UPX reconstruction carried no build info, reading the binary as it was found")
|
||||
closeUnpacked(unpacked)
|
||||
// unpackErr is deliberately dropped. Every reader after this one reads the bytes as they were
|
||||
// found, not the reconstruction, so nothing downstream is short: the header this binary
|
||||
// carried was not describing real UPX output in the first place. Returning the gap here would
|
||||
// put an unknown on a binary the cataloger read completely.
|
||||
return nil, foundBI, nil
|
||||
}
|
||||
}
|
||||
|
||||
// a packed file we could not unpack is the more specific gap, so it wins over whatever buildinfo made
|
||||
// of the bytes it was handed
|
||||
if unpackErr != nil {
|
||||
return unpacked, nil, unpackErr
|
||||
}
|
||||
return unpacked, bi, err
|
||||
}
|
||||
|
||||
// isCancelled reports whether err means the scan was called off, which is a reason to stop rather than a
|
||||
// gap in the SBOM.
|
||||
func isCancelled(err error) bool {
|
||||
return errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded)
|
||||
}
|
||||
|
||||
// readerFor returns where a scanned binary should be read from: the reconstruction when it was packed,
|
||||
// the input as it was found otherwise.
|
||||
//
|
||||
// The nil check lives here, in one place, on purpose: every reader below the scan goes through here, so no
|
||||
// parser has to ask whether this binary was packed.
|
||||
func readerFor(unpacked unpackedContents, found io.ReaderAt) io.ReaderAt {
|
||||
if unpacked == nil {
|
||||
return found
|
||||
}
|
||||
return unpacked
|
||||
}
|
||||
|
||||
// seekerFor is readerFor for the readers after the scan, which seek. found belongs to the caller and is
|
||||
// reused for every later module, so this deliberately hands back no Closer.
|
||||
func seekerFor(unpacked unpackedContents, found io.ReadSeeker) io.ReadSeeker {
|
||||
if unpacked == nil {
|
||||
return found
|
||||
}
|
||||
// the SectionReader snapshots Size() here, which is correct: the reconstruction is complete by the time
|
||||
// anything seeks it. It also gives each caller its own cursor over the same contents.
|
||||
return io.NewSectionReader(unpacked, 0, unpacked.Size())
|
||||
}
|
||||
|
||||
// closeUnpacked releases a reconstruction, if the binary turned out to be packed at all. Nothing to unpack
|
||||
// is the common case rather than an edge one, so every owner of an unpackedContents closes through here.
|
||||
func closeUnpacked(unpacked unpackedContents) {
|
||||
if unpacked == nil {
|
||||
return
|
||||
}
|
||||
_ = unpacked.Close()
|
||||
}
|
||||
|
||||
// readBuildInfo bounds the reader before handing it to debug/buildinfo, which opens ELF files with
|
||||
// debug/elf itself rather than through elfutil. elf.NewFile expands the section-name string table as it
|
||||
// parses, so an unbounded read here is reachable no matter how little of the file buildinfo goes on to
|
||||
@@ -134,7 +332,7 @@ func readBuildInfo(r io.ReaderAt) (*debug.BuildInfo, error) {
|
||||
return buildinfo.Read(r)
|
||||
}
|
||||
|
||||
func getBuildInfo(r io.ReaderAt, location file.Location) (bi *debug.BuildInfo, err error) {
|
||||
func getBuildInfo(r io.ReaderAt) (bi *debug.BuildInfo, err error) {
|
||||
defer func() {
|
||||
if r := recover(); r != nil {
|
||||
// this can happen in cases where a malformed binary is passed in can be initially parsed, but not
|
||||
@@ -144,24 +342,7 @@ func getBuildInfo(r io.ReaderAt, location file.Location) (bi *debug.BuildInfo, e
|
||||
}
|
||||
}()
|
||||
|
||||
// try to read buildinfo from the binary directly
|
||||
bi, err = readBuildInfo(r)
|
||||
if err == nil {
|
||||
return bi, nil
|
||||
}
|
||||
|
||||
// if direct read fails and this looks like a UPX-compressed binary,
|
||||
// try to decompress and read the buildinfo from the decompressed data
|
||||
if isUPXCompressed(r) {
|
||||
log.WithFields("path", location.RealPath).Trace("detected UPX-compressed Go binary, attempting decompression to read the build info")
|
||||
decompressed, decompErr := decompressUPX(r)
|
||||
if decompErr == nil {
|
||||
bi, err = readBuildInfo(decompressed)
|
||||
if err == nil {
|
||||
return bi, nil
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// note: the stdlib does not export the error we need to check for
|
||||
if err != nil {
|
||||
|
||||
@@ -6,13 +6,15 @@ import (
|
||||
"debug/buildinfo"
|
||||
"debug/elf"
|
||||
"encoding/binary"
|
||||
"os"
|
||||
"runtime"
|
||||
"testing"
|
||||
|
||||
"github.com/kastenhq/goversion/version"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
"github.com/anchore/syft/syft/file"
|
||||
"github.com/anchore/syft/syft/internal/elfutil"
|
||||
)
|
||||
|
||||
// Test_getBuildInfo_compressedSectionBomb covers the reason readBuildInfo exists: debug/buildinfo opens
|
||||
@@ -33,14 +35,122 @@ func Test_getBuildInfo_compressedSectionBomb(t *testing.T) {
|
||||
assert.Greater(t, unguarded, uint64(declared), "fixture did not actually deliver the declared bytes")
|
||||
|
||||
guarded := measureAlloc(t, func() {
|
||||
_, err := getBuildInfo(bytes.NewReader(bomb), file.NewLocation("bomb"))
|
||||
_, err := getBuildInfo(bytes.NewReader(bomb))
|
||||
require.Error(t, err)
|
||||
assert.Contains(t, err.Error(), "over the")
|
||||
assert.ErrorIs(t, err, elfutil.ErrDeclaredSizeExceeded)
|
||||
})
|
||||
assert.Less(t, guarded, uint64(32<<20), "getBuildInfo allocated far more than the input warrants")
|
||||
t.Logf("unguarded allocated %d bytes, guarded allocated %d bytes", unguarded, guarded)
|
||||
}
|
||||
|
||||
// Test_getCryptoInformation_compressedSymtabBomb covers the second unbounded door into debug/elf:
|
||||
// goversion opens the file itself and reads .symtab plus the string table it links, and those are
|
||||
// expanded lazily, so the section-name table bound getBuildInfo applies never reaches them.
|
||||
func Test_getCryptoInformation_compressedSymtabBomb(t *testing.T) {
|
||||
const declared = 256 << 20 // comfortably over elfutil's bound, small enough to allocate in a test
|
||||
|
||||
bomb := elfWithCompressedSymtab(t, declared)
|
||||
t.Logf("%d byte fixture declares a %d byte symbol table", len(bomb), declared)
|
||||
|
||||
// the unguarded path is the thing being defended against: prove the fixture really is a bomb
|
||||
unguarded := measureAlloc(t, func() {
|
||||
_, err := version.ReadExeFromReader(bytes.NewReader(bomb))
|
||||
t.Logf("goversion err: %v", err)
|
||||
})
|
||||
assert.Greater(t, unguarded, uint64(declared), "fixture did not actually deliver the declared bytes")
|
||||
|
||||
guarded := measureAlloc(t, func() {
|
||||
_, err := getCryptoInformation(bytes.NewReader(bomb))
|
||||
require.ErrorIs(t, err, elfutil.ErrDeclaredSizeExceeded)
|
||||
})
|
||||
assert.Less(t, guarded, uint64(32<<20), "getCryptoInformation allocated far more than the input warrants")
|
||||
t.Logf("unguarded allocated %d bytes, guarded allocated %d bytes", unguarded, guarded)
|
||||
}
|
||||
|
||||
// Test_getCryptoInformation_passesThroughNonELF keeps the gate from becoming a container filter: goversion
|
||||
// reads PE and Mach-O too, and CheckAllSections has to leave those to it.
|
||||
func Test_getCryptoInformation_passesThroughNonELF(t *testing.T) {
|
||||
runMakeTarget(t, "archs")
|
||||
|
||||
for _, name := range []string{"hello-win-amd64", "hello-mach-o-arm64"} {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
f, err := os.Open("testdata/archs/binaries/" + name)
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { require.NoError(t, f.Close()) })
|
||||
|
||||
_, err = getCryptoInformation(f)
|
||||
require.NoError(t, err)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// elfWithCompressedSymtab builds a minimal ELF64 whose .symtab is SHF_COMPRESSED, declaring `declared`
|
||||
// decompressed bytes and genuinely delivering them. The name table is left uncompressed so the file gets
|
||||
// past the bound getBuildInfo already applies, which is the point: this is the section that bound misses.
|
||||
func elfWithCompressedSymtab(t *testing.T, declared uint64) []byte {
|
||||
t.Helper()
|
||||
|
||||
payload := make([]byte, declared)
|
||||
var compressed bytes.Buffer
|
||||
zw := zlib.NewWriter(&compressed)
|
||||
_, err := zw.Write(payload)
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, zw.Close())
|
||||
|
||||
names := []byte("\x00.shstrtab\x00.symtab\x00.strtab\x00")
|
||||
nameOf := func(s string) uint32 {
|
||||
i := bytes.Index(names, []byte("\x00"+s+"\x00"))
|
||||
require.GreaterOrEqual(t, i, 0)
|
||||
return uint32(i + 1)
|
||||
}
|
||||
|
||||
ehsize := uint64(binary.Size(elf.Header64{}))
|
||||
shentsize := uint64(binary.Size(elf.Section64{}))
|
||||
chdrsize := uint64(binary.Size(elf.Chdr64{}))
|
||||
shoff := ehsize
|
||||
namesOff := shoff + 4*shentsize
|
||||
symOff := namesOff + uint64(len(names))
|
||||
symSize := chdrsize + uint64(compressed.Len())
|
||||
strOff := symOff + symSize
|
||||
|
||||
var ident [16]byte
|
||||
copy(ident[:], elf.ELFMAG)
|
||||
ident[elf.EI_CLASS] = byte(elf.ELFCLASS64)
|
||||
ident[elf.EI_DATA] = byte(elf.ELFDATA2LSB)
|
||||
ident[elf.EI_VERSION] = byte(elf.EV_CURRENT)
|
||||
|
||||
buf := &bytes.Buffer{}
|
||||
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Header64{
|
||||
Ident: ident, Type: uint16(elf.ET_REL), Machine: uint16(elf.EM_X86_64),
|
||||
Version: uint32(elf.EV_CURRENT), Shoff: shoff, Ehsize: uint16(ehsize),
|
||||
Shentsize: uint16(shentsize), Shnum: 4, Shstrndx: 1,
|
||||
}))
|
||||
// the null section
|
||||
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{}))
|
||||
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{
|
||||
Name: nameOf(".shstrtab"), Type: uint32(elf.SHT_STRTAB), Off: namesOff,
|
||||
Size: uint64(len(names)), Addralign: 1,
|
||||
}))
|
||||
// the bomb: sh_size covers the whole zlib stream, ch_size is what debug/elf expands to
|
||||
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{
|
||||
Name: nameOf(".symtab"), Type: uint32(elf.SHT_SYMTAB), Flags: uint64(elf.SHF_COMPRESSED),
|
||||
Off: symOff, Size: symSize, Link: 3, Entsize: 24, Addralign: 1,
|
||||
}))
|
||||
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Section64{
|
||||
Name: nameOf(".strtab"), Type: uint32(elf.SHT_STRTAB), Off: strOff, Size: 1, Addralign: 1,
|
||||
}))
|
||||
buf.Write(names)
|
||||
require.NoError(t, binary.Write(buf, binary.LittleEndian, elf.Chdr64{
|
||||
Type: uint32(elf.COMPRESS_ZLIB), Size: declared, Addralign: 1,
|
||||
}))
|
||||
buf.Write(compressed.Bytes())
|
||||
buf.Write([]byte{0})
|
||||
|
||||
return buf.Bytes()
|
||||
}
|
||||
|
||||
// measureAlloc reports the bytes allocated while fn ran. TotalAlloc is process-wide, so a test using this
|
||||
// must not call t.Parallel: another test's allocations would land in the measurement.
|
||||
func measureAlloc(t *testing.T, fn func()) uint64 {
|
||||
t.Helper()
|
||||
var before, after runtime.MemStats
|
||||
@@ -51,14 +161,33 @@ func measureAlloc(t *testing.T, fn func()) uint64 {
|
||||
return after.TotalAlloc - before.TotalAlloc
|
||||
}
|
||||
|
||||
// elfWithCompressedNameTable builds a minimal ELF64 whose only real section is a SHF_COMPRESSED
|
||||
// .shstrtab declaring `declared` decompressed bytes and genuinely delivering them.
|
||||
// elfWithCompressedNameTable builds a minimal ELF64 whose only real section is a SHF_COMPRESSED .shstrtab
|
||||
// declaring `declared` decompressed bytes and genuinely delivering them. Only the test that runs the
|
||||
// unguarded path needs delivery; use elfDeclaringNameTable everywhere else, since building this one costs
|
||||
// `declared` bytes of allocation plus a zlib compress of them.
|
||||
func elfWithCompressedNameTable(t *testing.T, declared uint64) []byte {
|
||||
t.Helper()
|
||||
return elfNameTableFixture(t, declared, true)
|
||||
}
|
||||
|
||||
// elfDeclaringNameTable declares `declared` bytes without delivering them. CheckSectionNameTable reads the
|
||||
// declared size out of the compression header and refuses before decompressing anything, so a test of the
|
||||
// bound itself never needs the stream to be real.
|
||||
func elfDeclaringNameTable(t *testing.T, declared uint64) []byte {
|
||||
t.Helper()
|
||||
return elfNameTableFixture(t, declared, false)
|
||||
}
|
||||
|
||||
func elfNameTableFixture(t *testing.T, declared uint64, deliver bool) []byte {
|
||||
t.Helper()
|
||||
|
||||
// the decompressed name table only has to start with the section names; the rest is padding that
|
||||
// exists purely to make the declared size real
|
||||
payload := make([]byte, declared)
|
||||
size := declared
|
||||
if !deliver {
|
||||
size = 64
|
||||
}
|
||||
payload := make([]byte, size)
|
||||
copy(payload, "\x00.shstrtab\x00")
|
||||
|
||||
var compressed bytes.Buffer
|
||||
|
||||
@@ -8,8 +8,6 @@ import (
|
||||
|
||||
"github.com/kastenhq/goversion/version"
|
||||
"github.com/stretchr/testify/assert"
|
||||
|
||||
"github.com/anchore/syft/syft/file"
|
||||
)
|
||||
|
||||
func Test_getBuildInfo(t *testing.T) {
|
||||
@@ -33,7 +31,7 @@ func Test_getBuildInfo(t *testing.T) {
|
||||
}
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
gotBi, err := getBuildInfo(tt.args.r, file.Location{})
|
||||
gotBi, err := getBuildInfo(tt.args.r)
|
||||
if !tt.wantErr(t, err, fmt.Sprintf("getBuildInfo(%v)", tt.args.r)) {
|
||||
return
|
||||
}
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
package golang
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"runtime/debug"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
"github.com/anchore/syft/internal/spillbuf"
|
||||
"github.com/anchore/syft/internal/tmpdir"
|
||||
"github.com/anchore/syft/syft/file"
|
||||
"github.com/anchore/syft/syft/internal/fileresolver"
|
||||
"github.com/anchore/syft/syft/internal/unionreader"
|
||||
)
|
||||
|
||||
// TestScanFile_UnpackedBinaryReadsTheFileItWasGiven is the regression test for the shape of bug that
|
||||
// modelling "not packed" as nil invites. The reader was once an io.ReadSeekCloser holding a nil pointer,
|
||||
// which is a non-nil interface, so every ordinary Go binary took the packed branch: it seeked a nil file
|
||||
// for its version and nil-dereferenced on close. The panic was recovered by the task executor, so the
|
||||
// only symptom was the go-binary cataloger reporting nothing at all.
|
||||
//
|
||||
// The nil is now the signal rather than a hazard: unpacked is nil when there was nothing to unpack, and
|
||||
// readerFor, seekerFor and closeUnpacked are the only places that resolve it.
|
||||
//
|
||||
// Docker-gated like everything else that needs a real Go binary; the point is that a plain binary reads
|
||||
// from the file it was handed, which needs a real one.
|
||||
func TestScanFile_UnpackedBinaryReadsTheFileItWasGiven(t *testing.T) {
|
||||
runMakeTarget(t, "archs")
|
||||
f, err := os.Open(filepath.Join("testdata", "archs", "binaries", "hello-linux-arm"))
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = f.Close() })
|
||||
|
||||
ur, err := unionreader.GetUnionReader(f)
|
||||
require.NoError(t, err)
|
||||
|
||||
builds, _ := scanFile(context.Background(), file.NewLocation("hello-linux-arm"), ur, false)
|
||||
require.NotEmpty(t, builds, "a plain Go binary must still produce build info")
|
||||
|
||||
for _, b := range builds {
|
||||
assert.Nil(t, b.unpacked, "an unpacked binary must not have left a reconstruction behind")
|
||||
assert.Same(t, ur, seekerFor(b.unpacked, ur), "the readers after the scan must get the file itself")
|
||||
assert.NotPanics(t, func() { closeUnpacked(b.unpacked); closeUnpacked(b.unpacked) },
|
||||
"closing nothing has to be safe and idempotent, since a defer over every build calls it")
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseGoBinary_PlainBinaryStillYieldsPackages is the same regression one layer out, at the entry
|
||||
// point the cataloger actually calls. The typed-nil panic fired from a defer here and was swallowed by
|
||||
// the task executor, so the failure mode is an empty result rather than an error.
|
||||
func TestParseGoBinary_PlainBinaryStillYieldsPackages(t *testing.T) {
|
||||
runMakeTarget(t, "archs")
|
||||
f, err := os.Open(filepath.Join("testdata", "archs", "binaries", "hello-linux-arm"))
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = f.Close() })
|
||||
|
||||
c := newGoBinaryCataloger(DefaultCatalogerConfig())
|
||||
pkgs, _, err := c.parseGoBinary(context.Background(), fileresolver.Empty{}, nil,
|
||||
file.NewLocationReadCloser(file.NewLocation("hello-linux-arm"), f))
|
||||
|
||||
require.NoError(t, err)
|
||||
assert.NotEmpty(t, pkgs, "an ordinary Go binary must still produce packages")
|
||||
}
|
||||
|
||||
// TestMakeGoMainPackage_VersionComesFromTheUnpackedContents pins the reader threading: for a packed
|
||||
// binary the version scan has to read the reconstruction, since the packed bytes carry no readable
|
||||
// version string. Nothing else covers this, because the from-contents scan is off by default and the
|
||||
// Docker fixture gets its version from ldflags before the scan is reached.
|
||||
func TestMakeGoMainPackage_VersionComesFromTheUnpackedContents(t *testing.T) {
|
||||
// the version pattern wants the string NUL-delimited
|
||||
unpacked := spillbuf.New(tmpdir.FromPath(t.TempDir()))
|
||||
t.Cleanup(func() { _ = unpacked.Close() })
|
||||
_, err := unpacked.WriteAt([]byte("\x00v9.9.9\x00"), 0)
|
||||
require.NoError(t, err)
|
||||
|
||||
// what the packed file on disk would have said, which must not be what comes out
|
||||
packed := &nopReadSeekCloser{bytes.NewReader([]byte("\x00v1.1.1\x00"))}
|
||||
|
||||
c := &goBinaryCataloger{
|
||||
licenseResolver: newGoLicenseResolver("", CatalogerConfig{}),
|
||||
mainModuleVersion: MainModuleVersionConfig{FromContents: true},
|
||||
}
|
||||
mod := &extendedBuildInfo{
|
||||
BuildInfo: &debug.BuildInfo{Main: debug.Module{Path: "github.com/anchore/syft", Version: devel}},
|
||||
unpacked: unpacked,
|
||||
}
|
||||
|
||||
got := c.makeGoMainPackage(context.Background(), fileresolver.Empty{}, mod, "amd64",
|
||||
file.NewLocation("packed"), seekerFor(mod.unpacked, packed), nil)
|
||||
|
||||
assert.Equal(t, "v9.9.9", got.Version, "the version must come from the reconstruction, not the packed bytes")
|
||||
}
|
||||
+829
-229
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,759 @@
|
||||
package golang
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"debug/elf"
|
||||
"encoding/binary"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
|
||||
intFile "github.com/anchore/syft/internal/file"
|
||||
"github.com/anchore/syft/internal/tmpdir"
|
||||
"github.com/anchore/syft/internal/unknown"
|
||||
"github.com/anchore/syft/syft/file"
|
||||
"github.com/anchore/syft/syft/internal/elfutil"
|
||||
"github.com/anchore/syft/syft/internal/unionreader"
|
||||
)
|
||||
|
||||
// the reconstruction's storage, its Close semantics and the contiguous-prefix rule it reports as Size all
|
||||
// belong to internal/spillbuf now, and are tested there. What is left in this file is what the golang
|
||||
// cataloger does with them.
|
||||
|
||||
// sparseFixture is a UPX file whose first block carries an ELF header declaring a second PT_LOAD segment
|
||||
// at farOffset, so a tiny third block is placed there. declaredShstrtab, when non-zero, makes that ELF
|
||||
// header also declare a section name table of that size sitting inside the gap, which is what turns the
|
||||
// reader's length into an allocation for anything that parses it.
|
||||
func sparseFixture(t *testing.T, declared uint32, farOffset, declaredShstrtab uint64, pad int) []byte {
|
||||
t.Helper()
|
||||
|
||||
phdrs := make([]byte, 112)
|
||||
binary.LittleEndian.PutUint32(phdrs[0:4], 1) // phdr[0] PT_LOAD
|
||||
binary.LittleEndian.PutUint64(phdrs[8:16], 0) // p_offset 0
|
||||
binary.LittleEndian.PutUint32(phdrs[56:60], 1)
|
||||
binary.LittleEndian.PutUint64(phdrs[64:72], farOffset) // phdr[1] p_offset, out at the declared end
|
||||
|
||||
hdr := make([]byte, 64)
|
||||
copy(hdr, []byte{0x7f, 'E', 'L', 'F'})
|
||||
hdr[4], hdr[5], hdr[6] = 2, 1, 1 // ELFCLASS64, little endian, EV_CURRENT
|
||||
binary.LittleEndian.PutUint16(hdr[0x10:0x12], uint16(elf.ET_EXEC))
|
||||
binary.LittleEndian.PutUint16(hdr[0x12:0x14], uint16(elf.EM_X86_64))
|
||||
binary.LittleEndian.PutUint32(hdr[0x14:0x18], 1)
|
||||
binary.LittleEndian.PutUint64(hdr[0x20:0x28], 64) // e_phoff
|
||||
binary.LittleEndian.PutUint16(hdr[0x36:0x38], 56) // e_phentsize
|
||||
binary.LittleEndian.PutUint16(hdr[0x38:0x3a], 2) // e_phnum
|
||||
|
||||
block1 := append(hdr, phdrs...)
|
||||
if declaredShstrtab > 0 {
|
||||
sh := make([]byte, 128) // two section headers; index 1 is the name table
|
||||
binary.LittleEndian.PutUint32(sh[64+4:64+8], uint32(elf.SHT_STRTAB))
|
||||
binary.LittleEndian.PutUint64(sh[64+24:64+32], farOffset/2) // sh_offset, inside the gap
|
||||
binary.LittleEndian.PutUint64(sh[64+32:64+40], declaredShstrtab)
|
||||
binary.LittleEndian.PutUint64(block1[0x28:0x30], 64+112) // e_shoff
|
||||
binary.LittleEndian.PutUint16(block1[0x3a:0x3c], 64) // e_shentsize
|
||||
binary.LittleEndian.PutUint16(block1[0x3c:0x3e], 2) // e_shnum
|
||||
binary.LittleEndian.PutUint16(block1[0x3e:0x40], 1) // e_shstrndx
|
||||
block1 = append(block1, sh...)
|
||||
}
|
||||
|
||||
data := buildUPXHeader(declared, 4096)
|
||||
data = append(data, blockFor(t, block1)...)
|
||||
data = append(data, blockFor(t, bytes.Repeat([]byte("B"), 32))...) // sequential
|
||||
data = append(data, blockFor(t, bytes.Repeat([]byte("C"), 64))...) // -> ptLoadOffsets[1]
|
||||
data = append(data, make([]byte, 12)...) // end marker
|
||||
return padTo(data, pad)
|
||||
}
|
||||
|
||||
// TestDecompressUPX_ReaderLengthIsWhatWasRebuilt covers the amplification that survives moving the
|
||||
// reconstruction from the heap to a temp file.
|
||||
//
|
||||
// A block's destination comes from a PT_LOAD p_offset in the first block, bounded only by p_filesize, so a
|
||||
// 64 byte block parked at p_filesize-64 must not make the reconstruction report the full declared size
|
||||
// while holding only a few hundred real bytes. A sparse file hands its holes back as zeros it never stored,
|
||||
// so every parser downstream that sizes against "what this reader will deliver" -- which is the whole
|
||||
// argument for writing to disk rather than the heap -- would otherwise allocate against a number the input
|
||||
// never paid for.
|
||||
func TestDecompressUPX_ReaderLengthIsWhatWasRebuilt(t *testing.T) {
|
||||
const declared = 64 << 20
|
||||
const input = 300_000
|
||||
|
||||
data := sparseFixture(t, declared, declared-64, 0, input)
|
||||
|
||||
out, err := unpack(t, data)
|
||||
// the hole is exactly what makes this partial: the blocks past it are given up, and that is reported
|
||||
// rather than logged, since a caller reading this reconstruction is missing bytes the header declared
|
||||
require.ErrorIs(t, err, errUPXPartial)
|
||||
require.NotNil(t, out, "a short reconstruction is still the contents to read from")
|
||||
|
||||
// the ELF header block plus the two small blocks, and nothing for the hole
|
||||
assert.Less(t, out.Size(), int64(4096),
|
||||
"the reader must be sized by the bytes actually rebuilt, not by where a block was parked")
|
||||
|
||||
_, err = out.ReadAt(make([]byte, 512), int64(declared)/2)
|
||||
assert.ErrorIs(t, err, io.EOF, "the hole must not read back as free zeros")
|
||||
}
|
||||
|
||||
func TestDecompressUPX_SparsePlacementIsNotAnAllocationKnob(t *testing.T) {
|
||||
// the same fixture, with the reconstructed ELF also declaring a 60MB section name table sitting in the
|
||||
// hole. This is the end of the exploit chain: debug/elf grows its read as the reads succeed, and
|
||||
// against a hole they all succeed.
|
||||
const declared = 64 << 20
|
||||
const input = 300_000
|
||||
|
||||
data := sparseFixture(t, declared, declared-64, 60<<20, input)
|
||||
|
||||
out, err := unpack(t, data)
|
||||
require.ErrorIs(t, err, errUPXPartial)
|
||||
require.NotNil(t, out)
|
||||
|
||||
allocated := measureAlloc(t, func() {
|
||||
_, _ = getBuildInfo(out)
|
||||
})
|
||||
|
||||
// debug/elf's own first chunk is ~10MB regardless of what a file declares, so the bound to assert is
|
||||
// that the declared 60MB is not reachable, not that this is free
|
||||
assert.Less(t, allocated, uint64(16*intFile.MB),
|
||||
"a 300KB input must not drive tens of MB of allocation through the reconstruction")
|
||||
}
|
||||
|
||||
func TestReadPTLoadOffsets_ShortReadIsNotTrusted(t *testing.T) {
|
||||
// on a short read the tail of the buffer is zeros the file never provided, and a p_offset of zero
|
||||
// parsed out of it would place a later block over the ELF header.
|
||||
phdrs := make([]byte, 112)
|
||||
binary.LittleEndian.PutUint32(phdrs[0:4], 1)
|
||||
binary.LittleEndian.PutUint64(phdrs[8:16], 0x1000)
|
||||
binary.LittleEndian.PutUint32(phdrs[56:60], 1)
|
||||
binary.LittleEndian.PutUint64(phdrs[64:72], 0x2000)
|
||||
full := buildELF64(64, 56, 2, phdrs)
|
||||
|
||||
assert.Equal(t, []uint64{0x1000, 0x2000}, readPTLoadOffsets(bytes.NewReader(full), 0, uint32(len(full))),
|
||||
"the whole header is readable, so both segments come back")
|
||||
|
||||
// the same header with the second program header cut off. The block claims it is there; the file is
|
||||
// not that long.
|
||||
truncated := full[:64+56+8]
|
||||
got := readPTLoadOffsets(bytes.NewReader(truncated), 0, uint32(len(full)))
|
||||
assert.Equal(t, []uint64{0x1000}, got,
|
||||
"only the segment actually present may be trusted; a zero read out of the gap is not an offset")
|
||||
}
|
||||
|
||||
func TestDecompressUPX_StoredBlockIsPlaced(t *testing.T) {
|
||||
// method 0 is an extent UPX could not compress and wrote verbatim, and real output carries a few of
|
||||
// them: the alignment padding between PT_LOAD segments is only a handful of bytes. Nothing covered
|
||||
// this path outside the Docker-backed fixture.
|
||||
payload := bytes.Repeat([]byte("S"), 48)
|
||||
stored := make([]byte, 12)
|
||||
binary.LittleEndian.PutUint32(stored[0:4], uint32(len(payload))) // sz_unc
|
||||
binary.LittleEndian.PutUint32(stored[4:8], uint32(len(payload))) // sz_cpr, equal for a stored block
|
||||
stored[8] = upxMethodStored
|
||||
|
||||
data := buildUPXHeader(4096, 4096)
|
||||
data = append(data, stored...)
|
||||
data = append(data, payload...)
|
||||
data = append(data, make([]byte, 12)...)
|
||||
data = padTo(data, 512)
|
||||
|
||||
out, err := unpack(t, data)
|
||||
// the fixture declares 4096 and delivers one 48 byte extent, so the chain is short by construction and
|
||||
// the partial goes with it. Real output rebuilds p_filesize exactly; what is under test is placement.
|
||||
require.ErrorIs(t, err, errUPXPartial)
|
||||
assert.Equal(t, payload, readAll(t, out), "a stored block is copied through as-is")
|
||||
}
|
||||
|
||||
func TestDecompressUPX_StoredBlockWithMismatchedSizesIsQuiet(t *testing.T) {
|
||||
// a copy is only meaningful when the two sizes agree. A mismatch is how a stray b_info-shaped run of
|
||||
// bytes looks, so it ends the chain rather than failing the file.
|
||||
stored := make([]byte, 12)
|
||||
binary.LittleEndian.PutUint32(stored[0:4], 64) // sz_unc
|
||||
binary.LittleEndian.PutUint32(stored[4:8], 32) // sz_cpr, disagrees
|
||||
stored[8] = upxMethodStored
|
||||
|
||||
data := padTo(append(buildUPXHeader(4096, 4096), stored...), 512)
|
||||
|
||||
_, err := unpack(t, data)
|
||||
require.Error(t, err)
|
||||
// a stored block whose sizes disagree is not a block we can read, which on the first block is the
|
||||
// unsupported-method refusal. Either way it must stay clear of errUPXDecompress.
|
||||
assert.ErrorIs(t, err, errUnsupportedUPXMethod, "a copy whose sizes disagree is not a copy")
|
||||
assert.NotErrorIs(t, err, errUPXDecompress, "not a packed binary, so it stays quiet")
|
||||
}
|
||||
|
||||
// TestScanReader_ReportingIsNarrow pins which gaps this cataloger claims. It runs against every
|
||||
// executable in an image, so reporting each file debug/buildinfo cannot parse would attach an unknown to
|
||||
// most files in a typical one. Only a bound syft itself chose to enforce, and a packed file it could not
|
||||
// unpack, are real gaps.
|
||||
func TestScanReader_ReportingIsNarrow(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
data []byte
|
||||
report bool
|
||||
}{
|
||||
{
|
||||
name: "a corrupt ELF is not this cataloger's gap",
|
||||
data: append([]byte{0x7f, 'E', 'L', 'F', 2, 1, 1}, bytes.Repeat([]byte{0xAB}, 512)...),
|
||||
},
|
||||
{
|
||||
name: "neither is a file that is not an executable at all",
|
||||
data: bytes.Repeat([]byte("not a binary"), 64),
|
||||
},
|
||||
{
|
||||
name: "nor a truncated PE",
|
||||
data: append([]byte("MZ\x90\x00"), bytes.Repeat([]byte{0}, 512)...),
|
||||
},
|
||||
{
|
||||
name: "a section name table syft declined to expand is",
|
||||
data: elfDeclaringNameTable(t, 256<<20),
|
||||
// this is elfutil.ErrDeclaredSizeExceeded: syft chose not to expand it, so the packages behind
|
||||
// it really are missing from the SBOM because of a decision made here
|
||||
report: true,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
loc := file.NewLocation("subject")
|
||||
_, err := scanReader(context.Background(), loc, bytes.NewReader(tt.data), false)
|
||||
|
||||
if !tt.report {
|
||||
assert.NoError(t, err, "this cataloger sees every executable in an image; it must stay quiet here")
|
||||
return
|
||||
}
|
||||
require.Error(t, err)
|
||||
var coordErr *unknown.CoordinateError
|
||||
require.ErrorAs(t, err, &coordErr, "a reported gap has to carry the location it belongs to")
|
||||
assert.Equal(t, loc.Coordinates, coordErr.Coordinates)
|
||||
assert.ErrorIs(t, coordErr.Reason, elfutil.ErrDeclaredSizeExceeded,
|
||||
"the reported gap has to carry the sentinel the split keys on, not just a message")
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestScanReader_RefusedExpansionIsReportedAsSuch(t *testing.T) {
|
||||
// the sentinel is what lets the narrow reporting above tell "syft declined to expand this" apart from
|
||||
// "this file is broken", so the two must not collapse into one string match
|
||||
_, err := readBuildInfo(bytes.NewReader(elfDeclaringNameTable(t, 256<<20)))
|
||||
require.Error(t, err)
|
||||
assert.ErrorIs(t, err, elfutil.ErrDeclaredSizeExceeded)
|
||||
}
|
||||
|
||||
// TestReadContentsAndBuildInfo_FakeUPXHeaderCannotHideAGoBinary covers the evasion vector the unpacking
|
||||
// opened up. The "UPX!" magic is found by an unanchored substring scan over the first 8KB and every field
|
||||
// behind it is attacker-controlled, so an ordinary Go binary can be made to produce a plausible header and
|
||||
// a reconstruction of nothing. Reading from that reconstruction unconditionally means the binary reports
|
||||
// no packages at all, which turns a false positive into a way to hide a dependency list.
|
||||
// goELF64Fixture returns a real, unpacked Go binary that clears the ELF64 little-endian container gate in
|
||||
// parseUPXInfo. hello-linux-arm is ELFCLASS32 and is rejected before the UPX magic scan runs, so a test
|
||||
// that splices a header into it exercises nothing.
|
||||
func goELF64Fixture(t *testing.T) []byte {
|
||||
t.Helper()
|
||||
runMakeTarget(t, "archs")
|
||||
original, err := os.ReadFile(filepath.Join("testdata", "archs", "binaries", "hello-linux-ppc64le"))
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, byte(2), original[4], "the container gate only accepts ELF64")
|
||||
require.Equal(t, byte(1), original[5], "the container gate only accepts little-endian")
|
||||
return original
|
||||
}
|
||||
|
||||
// spliceFakeUPXHeader writes a plausible UPX header plus one decodable block into padding inside the
|
||||
// magic scan window, leaving the ELF itself intact. The offset is found rather than hardcoded: it has to
|
||||
// sit past the section header table (writing over that corrupts the binary and the test then proves
|
||||
// nothing) and inside upxMagicScanWindow (past it the magic is never seen).
|
||||
func spliceFakeUPXHeader(t *testing.T, original []byte) []byte {
|
||||
t.Helper()
|
||||
return spliceUPXChain(t, original, 64<<10, blockFor(t, []byte("junk")))
|
||||
}
|
||||
|
||||
// spliceUPXChain is spliceFakeUPXHeader with a caller-supplied p_filesize and b_info chain, so a test can
|
||||
// choose whether the fake header lands on a clean short reconstruction, one that gives something up, or a
|
||||
// refusal by the size bounds with nothing unpacked at all.
|
||||
func spliceUPXChain(t *testing.T, original []byte, originalSize uint32, chain []byte) []byte {
|
||||
t.Helper()
|
||||
|
||||
header := buildUPXHeader(originalSize, 4096)[64:] // drop the stub, the real ELF header is already here
|
||||
block := chain
|
||||
need := len(header) + len(block)
|
||||
|
||||
shoff := binary.LittleEndian.Uint64(original[0x28:0x30])
|
||||
shentsize := binary.LittleEndian.Uint16(original[0x3a:0x3c])
|
||||
shnum := binary.LittleEndian.Uint16(original[0x3c:0x3e])
|
||||
after := int(shoff) + int(shentsize)*int(shnum)
|
||||
|
||||
at := -1
|
||||
for i := after; i+need <= upxMagicScanWindow && i+need <= len(original); i++ {
|
||||
if bytes.Equal(original[i:i+need], make([]byte, need)) {
|
||||
at = i
|
||||
break
|
||||
}
|
||||
}
|
||||
require.GreaterOrEqual(t, at, 0, "no padding in the scan window big enough to splice a header into")
|
||||
|
||||
spiked := append([]byte(nil), original...)
|
||||
copy(spiked[at:], header)
|
||||
copy(spiked[at+len(header):], block)
|
||||
|
||||
// the splice must not have broken the binary, or the fallback below would be covering for a corrupt
|
||||
// ELF rather than for a fake header. A claim the size bounds refuse is a deliberate caller choice, so
|
||||
// only assert the header parses when it was meant to.
|
||||
if originalSize <= maxUPXOriginalSize {
|
||||
require.NotNil(t, mustParseUPX(t, spiked), "the spliced header has to actually parse as UPX")
|
||||
}
|
||||
return spiked
|
||||
}
|
||||
|
||||
// mustParseUPX reports the UPX header the scan finds in data, proving the splice is what the parser sees.
|
||||
func mustParseUPX(t *testing.T, data []byte) *upxInfo {
|
||||
t.Helper()
|
||||
r := bytes.NewReader(data)
|
||||
info, err := parseUPXInfo(r, int64(len(data)))
|
||||
require.NoError(t, err)
|
||||
return info
|
||||
}
|
||||
|
||||
func TestScanFile_FakeUPXHeaderStillYieldsPackages(t *testing.T) {
|
||||
// the same evasion one layer out, at the entry point the cataloger calls. The header has to actually be
|
||||
// spliced: run this against the untouched binary and it passes with the fallback deleted.
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
original := goELF64Fixture(t)
|
||||
|
||||
spiked := spliceFakeUPXHeader(t, original)
|
||||
|
||||
ur, err := unionreader.GetUnionReader(io.NopCloser(bytes.NewReader(spiked)))
|
||||
require.NoError(t, err)
|
||||
|
||||
builds, _ := scanFile(ctx, file.NewLocation("hello-linux-ppc64le"), ur, false)
|
||||
require.NotEmpty(t, builds, "a Go binary must not be hidden by four bytes of fake magic")
|
||||
for _, b := range builds {
|
||||
assert.Nil(t, b.unpacked,
|
||||
"the reconstruction carried nothing, so the scan has to fall back to the bytes as they were found")
|
||||
closeUnpacked(b.unpacked)
|
||||
}
|
||||
}
|
||||
|
||||
// TestScanReader_CancellationStopsTheScan pins the direction a swallowed ctx.Err() got wrong. Cancellation
|
||||
// is not a gap in the SBOM, so it is not reported as an unknown, but it does mean stop: returning "nothing
|
||||
// to unpack, no error" sent the scan on to parse the packed bytes and build packages out of them after the
|
||||
// caller had already called it off.
|
||||
func TestScanReader_CancellationStopsTheScan(t *testing.T) {
|
||||
// enough blocks that the loop checks ctx at least once
|
||||
blocks := make([][]byte, 64)
|
||||
for i := range blocks {
|
||||
blocks[i] = bytes.Repeat([]byte("A"), 1024)
|
||||
}
|
||||
fixture := buildUPXFile(t, 64*1024, 4096, blocks, nil)
|
||||
|
||||
ctx, cancel := context.WithCancel(tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir())))
|
||||
cancel()
|
||||
|
||||
build, err := scanReader(ctx, file.NewLocation("/cancelled"), bytes.NewReader(fixture), false)
|
||||
assert.Nil(t, build)
|
||||
require.Error(t, err, "a cancelled scan has to stop rather than fall through to the packed bytes")
|
||||
assert.ErrorIs(t, err, context.Canceled)
|
||||
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
|
||||
assert.Empty(t, coordErrs, "cancellation is not a gap in the SBOM")
|
||||
}
|
||||
|
||||
// TestScanReader_SizeRefusalIsReported covers the reporting split for a header that really is UPX and
|
||||
// claims more than the bounds allow. That is a file we declined to expand, which is the same category as
|
||||
// elfutil.ErrDeclaredSizeExceeded and gets the same unknown. The ratio clears a measured 209x worst case at
|
||||
// 256, so a legitimate binary crossing it must not vanish from the SBOM silently.
|
||||
func TestScanReader_SizeRefusalIsReported(t *testing.T) {
|
||||
// a 4KB input claiming 512MB: pays no ratio and clears no ceiling
|
||||
data := padTo(buildUPXHeader(maxUPXOriginalSize, 0x1000), 4096)
|
||||
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
build, err := scanReader(ctx, file.NewLocation("/refused"), bytes.NewReader(data), false)
|
||||
|
||||
assert.Nil(t, build)
|
||||
require.Error(t, err)
|
||||
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
|
||||
require.NotEmpty(t, coordErrs,
|
||||
"a file we declined to expand is a gap we chose to leave, so it is reported")
|
||||
assert.ErrorIs(t, coordErrs[0].Reason, errUPXSizeRefused)
|
||||
}
|
||||
|
||||
// corruptLZMABlock builds a b_info whose stream carries valid LZMA parameters and nothing else that
|
||||
// decodes. This is what the loader stub behind the last real block looks like to readChainBlock: the
|
||||
// method byte is one we implement and the sizes are in range, so the chain accepts it and the decoder is
|
||||
// the thing that says no.
|
||||
func corruptLZMABlock(szUnc uint32) []byte {
|
||||
stream := make([]byte, 64)
|
||||
stream[0] = 0x00 // pb = 0
|
||||
stream[1] = 0x00 // lc = 0, lp = 0
|
||||
for i := 2; i < len(stream); i++ {
|
||||
stream[i] = byte(i * 7) // not a range-coded anything
|
||||
}
|
||||
b := make([]byte, 12)
|
||||
binary.LittleEndian.PutUint32(b[0:4], szUnc)
|
||||
binary.LittleEndian.PutUint32(b[4:8], uint32(len(stream)))
|
||||
b[8] = 14 // b_method = LZMA
|
||||
return append(b, stream...)
|
||||
}
|
||||
|
||||
// TestDecompressUPX_UndecodableLaterBlockKeepsWhatCameBefore covers a chain that gives up a good
|
||||
// reconstruction over garbage past the end of it.
|
||||
//
|
||||
// readChainBlock accepts any b_info-shaped run of bytes whose method it knows and whose sizes fit the
|
||||
// file, so the loader stub behind the last real block parses as a block roughly one time in 256 on the
|
||||
// method byte alone. A decode failure on that block must not discard every block already placed and turn a
|
||||
// fully recoverable binary into zero packages: the two sibling conditions on the same block (an overrun,
|
||||
// and a placement that does not fit) already keep their partial output, and this one must too.
|
||||
func TestDecompressUPX_UndecodableLaterBlockKeepsWhatCameBefore(t *testing.T) {
|
||||
head := bytes.Repeat([]byte("H"), 256)
|
||||
|
||||
data := buildUPXHeader(4096, 4096)
|
||||
data = append(data, blockFor(t, head)...)
|
||||
data = append(data, corruptLZMABlock(64)...)
|
||||
data = padTo(data, 1024)
|
||||
|
||||
out, err := unpack(t, data)
|
||||
require.ErrorIs(t, err, errUPXPartial, "the blocks already placed are kept, and what was lost is reported")
|
||||
require.NotErrorIs(t, err, errUPXDecompress, "a good reconstruction is not discarded over trailing garbage")
|
||||
require.NotNil(t, out)
|
||||
|
||||
assert.Equal(t, head, readAll(t, out), "everything decoded before the bad block stays placed")
|
||||
}
|
||||
|
||||
// TestDecompressUPX_UndecodableFirstBlockStillFailsTheFile is the other half: with nothing placed there is
|
||||
// no partial output to keep, so the file really is one we could not unpack.
|
||||
func TestDecompressUPX_UndecodableFirstBlockStillFailsTheFile(t *testing.T) {
|
||||
data := padTo(append(buildUPXHeader(4096, 4096), corruptLZMABlock(64)...), 1024)
|
||||
|
||||
_, err := unpack(t, data)
|
||||
require.ErrorIs(t, err, errUPXDecompress, "a packed file with nothing readable in it is a gap in the SBOM")
|
||||
assert.NotErrorIs(t, err, errUPXPartial)
|
||||
}
|
||||
|
||||
// TestScanReader_PartialReconstructionIsReported is the regression test for a partial unpack vanishing
|
||||
// silently. The reconstruction is truncated to the contiguous prefix, which is the right bound, but the
|
||||
// bytes past the gap are gone: on a real binary that is the non-loadable tail holding the section headers,
|
||||
// so the symbols and often the build info go with it. Reported at Trace, the binary contributed no
|
||||
// packages and no unknown, which is the one outcome the reporting split in upx.go exists to prevent.
|
||||
func TestScanReader_PartialReconstructionIsReported(t *testing.T) {
|
||||
const declared = 64 << 20
|
||||
data := sparseFixture(t, declared, declared-64, 0, 300_000)
|
||||
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
build, err := scanReader(ctx, file.NewLocation("/partial"), bytes.NewReader(data), false)
|
||||
|
||||
assert.Nil(t, build)
|
||||
require.Error(t, err)
|
||||
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
|
||||
require.NotEmpty(t, coordErrs, "a reconstruction that lost bytes is a gap this cataloger chose to leave")
|
||||
assert.ErrorIs(t, coordErrs[0].Reason, errUPXPartial)
|
||||
}
|
||||
|
||||
// TestReadContentsAndBuildInfo_FallbackToFoundBytesReportsNoGap pins the retry direction. A gap is
|
||||
// reported because the readers after it (crypto settings, arch, symbols) would otherwise read a short
|
||||
// reconstruction with nothing saying so, which is what
|
||||
// TestReadContentsAndBuildInfo_GapSurvivesBuildInfoFromTheReconstruction covers. That reasoning runs out
|
||||
// here: the retry hands back the bytes as they were found, every reader below reads those, and the header
|
||||
// that produced the short reconstruction was not describing real UPX output to begin with. Reporting it
|
||||
// would put "UPX reconstruction is incomplete" on a binary that was never packed and was read in full.
|
||||
func TestReadContentsAndBuildInfo_FallbackToFoundBytesReportsNoGap(t *testing.T) {
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
|
||||
// a chain that decodes one block and then hits a block it cannot read: the reconstruction is real but
|
||||
// short, and the build info comes from the fallback to the bytes as they were found
|
||||
chain := append(blockFor(t, []byte("junk")), corruptLZMABlock(64)...)
|
||||
spiked := spliceUPXChain(t, goELF64Fixture(t), 64<<10, chain)
|
||||
|
||||
unpacked, bi, err := readContentsAndBuildInfo(ctx, bytes.NewReader(spiked))
|
||||
t.Cleanup(func() { closeUnpacked(unpacked) })
|
||||
|
||||
require.NotNil(t, bi, "the fallback still finds the build info")
|
||||
require.Nil(t, unpacked, "the reconstruction was released; the caller reads the bytes as they were found")
|
||||
assert.NoError(t, err, "nothing downstream reads the short reconstruction, so there is no gap to report")
|
||||
}
|
||||
|
||||
// TestUnpackUPX_CancelledContextComesBackAsCtxErr covers the cancellation arm in unpackUPX directly: the
|
||||
// scan-level test above it passes even with this arm deleted, since readContentsAndBuildInfo does not
|
||||
// re-check ctx.Err() on its own.
|
||||
func TestUnpackUPX_CancelledContextComesBackAsCtxErr(t *testing.T) {
|
||||
payload := bytes.Repeat([]byte("A"), 2048)
|
||||
blocks := make([][]byte, 64)
|
||||
for i := range blocks {
|
||||
blocks[i] = payload
|
||||
}
|
||||
data := padTo(buildUPXFile(t, 4096*64, 2048, blocks, nil), 4096)
|
||||
|
||||
ctx, cancel := context.WithCancel(tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir())))
|
||||
cancel()
|
||||
|
||||
out, err := unpackUPX(ctx, bytes.NewReader(data))
|
||||
require.ErrorIs(t, err, context.Canceled)
|
||||
assert.NotErrorIs(t, err, errUPXDecompress, "a cancelled scan is not a gap in the SBOM")
|
||||
assert.NotErrorIs(t, err, errUPXPartial, "nor is it a short reconstruction")
|
||||
assert.Nil(t, out, "a cancelled unpack hands back nothing; the caller reads the input as it was found")
|
||||
}
|
||||
|
||||
// TestParseUPXInfo_HeaderRunsPastTheScanWindow covers the guard on the l_info+p_info read: the magic can
|
||||
// sit close enough to the end of what was actually read that the 20 bytes behind it are not there.
|
||||
func TestParseUPXInfo_HeaderRunsPastTheScanWindow(t *testing.T) {
|
||||
data := append(packedELFStub(), upxMagic...) // magic at 64, nothing behind it
|
||||
|
||||
r := bytes.NewReader(data)
|
||||
_, err := parseUPXInfo(r, int64(len(data)))
|
||||
require.Error(t, err)
|
||||
assert.ErrorIs(t, err, errUPXImplausibleHeader)
|
||||
assert.NotErrorIs(t, err, errUPXSizeRefused, "a header we could not even read is not a refusal to expand")
|
||||
}
|
||||
|
||||
// TestReadContentsAndBuildInfo_GapSurvivesBuildInfoFromTheReconstruction is the same property one path
|
||||
// over: here the reconstruction itself carries the build info, so the early return must not throw the
|
||||
// unpack error away. The chain rebuilds a whole Go binary and then hits a block it cannot read, which is
|
||||
// exactly the shape of a real packed binary whose tail extents did not all come back.
|
||||
func TestReadContentsAndBuildInfo_GapSurvivesBuildInfoFromTheReconstruction(t *testing.T) {
|
||||
payload := goELF64Fixture(t)
|
||||
|
||||
data := buildUPXHeader(uint32(len(payload)+64), 4096)
|
||||
data = append(data, blockFor(t, payload)...)
|
||||
data = append(data, corruptLZMABlock(64)...)
|
||||
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
contents, bi, err := readContentsAndBuildInfo(ctx, bytes.NewReader(data))
|
||||
require.NotNil(t, contents)
|
||||
t.Cleanup(func() { closeUnpacked(contents) })
|
||||
|
||||
require.NotNil(t, contents, "the build info has to come from the reconstruction here")
|
||||
require.NotNil(t, bi, "the rebuilt binary is a whole Go binary")
|
||||
assert.ErrorIs(t, err, errUPXPartial,
|
||||
"the block the chain gave up is still a gap, even though the build info came through")
|
||||
}
|
||||
|
||||
// TestScanReader_PartialReconstructionStillYieldsItsPackages is the other direction of the reporting
|
||||
// split, and the one that keeps the new reporting from becoming a regression of its own: a gap is
|
||||
// something to report alongside the packages, not instead of them. A binary whose chain gave up its tail
|
||||
// still has its module list, and dropping it would trade a silent false negative for a loud one.
|
||||
func TestScanReader_PartialReconstructionStillYieldsItsPackages(t *testing.T) {
|
||||
payload := goELF64Fixture(t)
|
||||
|
||||
data := buildUPXHeader(uint32(len(payload)+64), 4096)
|
||||
data = append(data, blockFor(t, payload)...)
|
||||
data = append(data, corruptLZMABlock(64)...)
|
||||
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
build, err := scanReader(ctx, file.NewLocation("/partial-but-usable"), bytes.NewReader(data), false)
|
||||
|
||||
require.NotNil(t, build, "a reported gap must not cost the packages that did come through")
|
||||
t.Cleanup(func() { closeUnpacked(build.unpacked) })
|
||||
assert.NotNil(t, build.BuildInfo)
|
||||
|
||||
require.Error(t, err)
|
||||
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
|
||||
require.NotEmpty(t, coordErrs, "and the gap is still reported")
|
||||
assert.ErrorIs(t, coordErrs[0].Reason, errUPXPartial)
|
||||
}
|
||||
|
||||
// sizeProbeSpy counts the size probes made against it. intFile.ReaderSize answers from Size() when the
|
||||
// reader has one, which is what this intercepts.
|
||||
type sizeProbeSpy struct {
|
||||
*bytes.Reader
|
||||
probes int
|
||||
}
|
||||
|
||||
func (s *sizeProbeSpy) Size() int64 {
|
||||
s.probes++
|
||||
return s.Reader.Size()
|
||||
}
|
||||
|
||||
// TestUnpackUPX_ContainerGateComesBeforeTheSizeProbe pins the ordering that keeps this cheap. The
|
||||
// cataloger runs over every file in an image and almost none are packed ELF, so nothing above the six
|
||||
// byte ident check should cost more than that: ReaderSize seeks to the end and reads the last byte back,
|
||||
// which over a squashfs or tar-backed reader is a real seek and decompress.
|
||||
func TestUnpackUPX_ContainerGateComesBeforeTheSizeProbe(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
data []byte
|
||||
probed bool
|
||||
}{
|
||||
{name: "not an ELF at all", data: bytes.Repeat([]byte("just some bytes"), 64)},
|
||||
{name: "a 32-bit ELF", data: func() []byte { b := padTo(packedELFStub(), 256); b[4] = 1; return b }()},
|
||||
{name: "a big-endian ELF64", data: func() []byte { b := padTo(packedELFStub(), 256); b[5] = 2; return b }()},
|
||||
{name: "an ELF64 we would try to unpack", data: padTo(packedELFStub(), 256), probed: true},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
spy := &sizeProbeSpy{Reader: bytes.NewReader(tt.data)}
|
||||
out, err := unpackUPX(context.Background(), spy)
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { closeUnpacked(out) })
|
||||
|
||||
if tt.probed {
|
||||
assert.Positive(t, spy.probes, "a candidate container does get measured")
|
||||
return
|
||||
}
|
||||
assert.Zero(t, spy.probes, "a file we will never unpack must not be measured first")
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestScanReader_ChainEndingShortIsReported covers the fourth shape a short reconstruction comes in: a
|
||||
// chain that simply ran out of b_info structures before rebuilding p_filesize must be reported, the same
|
||||
// as one that gave up mid-walk, one that hit the block cap, and one with blocks stranded past a hole. Every
|
||||
// quiet exit in readChainBlock produces that shape (an unreadable b_info, the end marker, a zero sz_cpr,
|
||||
// compressed data past the end of the input, an unimplemented method past the first block), and the
|
||||
// reconstruction is then truncated to the covered prefix, which can drop the section headers and leave the
|
||||
// binary contributing neither packages nor an unknown.
|
||||
func TestScanReader_ChainEndingShortIsReported(t *testing.T) {
|
||||
// one 2048 byte block against a declared 4096, then the end marker: nothing failed, the chain just ran
|
||||
// out. stopped is nil, covered equals the furthest placement, and covered is half of total.
|
||||
payload := bytes.Repeat([]byte("A"), 2048)
|
||||
data := buildUPXHeader(4096, 2048)
|
||||
data = append(data, blockFor(t, payload)...)
|
||||
data = append(data, make([]byte, 12)...) // end marker
|
||||
data = padTo(data, 512)
|
||||
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
build, err := scanReader(ctx, file.NewLocation("/short-chain"), bytes.NewReader(data), false)
|
||||
|
||||
assert.Nil(t, build)
|
||||
require.Error(t, err)
|
||||
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
|
||||
require.NotEmpty(t, coordErrs, "a chain that ended before rebuilding p_filesize is a reported gap")
|
||||
assert.ErrorIs(t, coordErrs[0].Reason, errUPXPartial)
|
||||
}
|
||||
|
||||
// TestScanReader_SizeRefusalOnAnUnpackedBinaryIsNotReported is the other direction of
|
||||
// TestScanReader_SizeRefusalIsReported. The magic is found by an unanchored scan over the first 8KB, so an
|
||||
// ordinary Go binary can carry a "UPX!" that parses into a header the size bounds then refuse. Nothing is
|
||||
// unpacked in that case: unpackUPX hands back the input exactly as it was found, every reader below reads
|
||||
// those bytes in full, and the build info proves they are complete. Reporting the refusal there put
|
||||
// "golang binary read incompletely" on a binary the cataloger read completely, with every package present.
|
||||
func TestScanReader_SizeRefusalOnAnUnpackedBinaryIsNotReported(t *testing.T) {
|
||||
original := goELF64Fixture(t)
|
||||
// a header whose p_filesize no input this size could ever pay for, so parseUPXInfo refuses it before
|
||||
// decompressUPX is reached and there is no reconstruction at any point
|
||||
spiked := spliceUPXChain(t, original, 0xFFFFFFFF, blockFor(t, []byte("junk")))
|
||||
|
||||
_, perr := parseUPXInfo(bytes.NewReader(spiked), int64(len(spiked)))
|
||||
require.ErrorIs(t, perr, errUPXSizeRefused, "the splice has to be refused by the size bounds")
|
||||
|
||||
ctx := tmpdir.WithValue(context.Background(), tmpdir.FromPath(t.TempDir()))
|
||||
build, err := scanReader(ctx, file.NewLocation("/spiked"), bytes.NewReader(spiked), false)
|
||||
if build != nil {
|
||||
t.Cleanup(func() { closeUnpacked(build.unpacked) })
|
||||
}
|
||||
|
||||
require.NotNil(t, build, "the binary is a complete Go binary and must still be cataloged")
|
||||
assert.Nil(t, build.unpacked, "nothing was unpacked, so these are the bytes as found")
|
||||
coordErrs, _ := unknown.ExtractCoordinateErrors(err)
|
||||
assert.Empty(t, coordErrs, "a file read in full must not be reported as read incompletely")
|
||||
}
|
||||
|
||||
// TestDecompressUPX_DenseChainReportsNoGap is the guard for the covered >= total early return in
|
||||
// finalExtent, and it is the fixture every other one in this package is not: dense. A chain that rebuilds
|
||||
// exactly p_filesize must come back with no partial reason attached.
|
||||
//
|
||||
// This matters because reporting a chain that ended short (the arm below that early return) is only safe if
|
||||
// well-formed output reaches full coverage. Every other crafted fixture here declares more than it delivers
|
||||
// and asserts errUPXPartial, and the one test that sees real `upx --best --lzma` output needs Docker, so
|
||||
// without this the early return has no local guard and the new arm's false-positive risk is untested.
|
||||
//
|
||||
// Layout below tiles 320 bytes through all three placement rules: block 1 at zero, block 2 sequentially
|
||||
// behind it, block 3 at a PT_LOAD offset from the headers in block 1, and block 4 into the hole those left.
|
||||
func TestDecompressUPX_DenseChainReportsNoGap(t *testing.T) {
|
||||
const (
|
||||
headerBlock = 64 + 112 // ELF64 header plus two program headers
|
||||
block2Size = 24 // sequential, so [176, 200)
|
||||
ptLoad1 = 256 // block 3 lands here, leaving [200, 256) behind
|
||||
block3Size = 64 // [256, 320)
|
||||
holeSize = ptLoad1 - (headerBlock + block2Size)
|
||||
total = ptLoad1 + block3Size
|
||||
)
|
||||
|
||||
phdrs := make([]byte, 112)
|
||||
binary.LittleEndian.PutUint32(phdrs[0:4], 1) // phdr[0] PT_LOAD
|
||||
binary.LittleEndian.PutUint64(phdrs[8:16], 0) // p_offset 0, the extent block 2 covers
|
||||
binary.LittleEndian.PutUint32(phdrs[56:60], 1)
|
||||
binary.LittleEndian.PutUint64(phdrs[64:72], ptLoad1) // phdr[1] p_offset, where block 3 goes
|
||||
|
||||
block1 := buildELF64(64, 56, 2, phdrs)
|
||||
require.Len(t, block1, headerBlock, "the first block is the original ELF headers")
|
||||
|
||||
data := buildUPXHeader(total, 4096)
|
||||
data = append(data, blockFor(t, block1)...) // placed at 0
|
||||
data = append(data, blockFor(t, bytes.Repeat([]byte("S"), block2Size))...) // sequential
|
||||
data = append(data, blockFor(t, bytes.Repeat([]byte("P"), block3Size))...) // -> ptLoadOffsets[1]
|
||||
data = append(data, blockFor(t, bytes.Repeat([]byte("F"), holeSize))...) // -> firstHole
|
||||
data = append(data, make([]byte, 12)...) // end marker
|
||||
data = padTo(data, 1024)
|
||||
|
||||
out, err := unpack(t, data)
|
||||
require.NotNil(t, out)
|
||||
require.NoError(t, err,
|
||||
"a chain that rebuilt every byte p_filesize declared has given nothing up, so there is no gap to report")
|
||||
|
||||
got := readAll(t, out)
|
||||
assert.Len(t, got, total, "the reconstruction is exactly what was declared")
|
||||
assert.Equal(t, bytes.Repeat([]byte("F"), holeSize), got[headerBlock+block2Size:ptLoad1],
|
||||
"the hole between the sequential extent and the PT_LOAD one is filled by the block after them")
|
||||
assert.Equal(t, bytes.Repeat([]byte("P"), block3Size), got[ptLoad1:],
|
||||
"block 3 lands at the PT_LOAD offset the headers declared")
|
||||
}
|
||||
|
||||
// TestFinalExtent pins which reason a short reconstruction comes back with, not merely that it is short.
|
||||
// Every arm wraps errUPXPartial and nothing else read the messages, so three of the four were
|
||||
// interchangeable: deleting the block-cap arm and the stranded-blocks arm left the whole package green
|
||||
// because both fell through to the arm below them. finalExtent is a pure function, so this costs nothing.
|
||||
func TestFinalExtent(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
covered uint64
|
||||
blockNum int
|
||||
total uint64
|
||||
stopped error
|
||||
wantMsg string // empty means no gap is reported
|
||||
}{
|
||||
{
|
||||
name: "a chain that rebuilt everything reports nothing, whatever it tripped over next",
|
||||
covered: 100,
|
||||
blockNum: 2,
|
||||
total: 100,
|
||||
stopped: fmt.Errorf("%w: ignored once coverage is complete", errUPXPartial),
|
||||
},
|
||||
{
|
||||
name: "a mid-walk stop is more specific than anything the extents show",
|
||||
covered: 100,
|
||||
blockNum: 2,
|
||||
total: 200,
|
||||
stopped: fmt.Errorf("%w: block 2 did not decode", errUPXPartial),
|
||||
wantMsg: "block 2 did not decode",
|
||||
},
|
||||
{
|
||||
name: "hitting the block cap is named as the cap",
|
||||
covered: 100,
|
||||
blockNum: maxUPXBlocks,
|
||||
total: 200,
|
||||
wantMsg: "block cap",
|
||||
},
|
||||
{
|
||||
name: "a chain that simply ended is named as such",
|
||||
covered: 100,
|
||||
blockNum: 3,
|
||||
total: 200,
|
||||
wantMsg: "chain ended after 100 of the declared 200",
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
got := finalExtent(tt.covered, tt.blockNum, tt.total, tt.stopped)
|
||||
|
||||
if tt.wantMsg == "" {
|
||||
assert.NoError(t, got)
|
||||
return
|
||||
}
|
||||
require.Error(t, got)
|
||||
assert.ErrorIs(t, got, errUPXPartial)
|
||||
assert.Contains(t, got.Error(), tt.wantMsg,
|
||||
"the reason has to say which shape this was, or the arms are interchangeable")
|
||||
})
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -8,57 +8,64 @@ import (
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
func TestIsUPXCompressed(t *testing.T) {
|
||||
// every case here is one the magic scan must reject, asserted on parseUPXInfo's errNotUPX result.
|
||||
//
|
||||
// Every negative case that is meant to exercise the magic scan has to carry a real ELF64 little-endian
|
||||
// ident, or the container gate rejects it first and the subtest passes without the scan running at all
|
||||
// (the gate and a missing magic both return errNotUPX, so nothing in the assertion can tell them apart).
|
||||
// TestParseUPXInfo_OnlyELF64LittleEndian owns the gate; these own the scan.
|
||||
func TestParseUPXInfo_MagicDetection(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
data []byte
|
||||
expected bool
|
||||
name string
|
||||
data []byte
|
||||
foundMagic bool // the magic was located (the header may still be rejected as implausible)
|
||||
}{
|
||||
{
|
||||
name: "contains UPX magic at start",
|
||||
data: append([]byte("UPX!"), make([]byte, 100)...),
|
||||
expected: true,
|
||||
name: "contains UPX magic at start",
|
||||
data: append(append(packedELFStub(), []byte("UPX!")...), make([]byte, 100)...),
|
||||
foundMagic: true,
|
||||
},
|
||||
{
|
||||
name: "contains UPX magic with offset",
|
||||
data: append(append(make([]byte, 500), []byte("UPX!")...), make([]byte, 100)...),
|
||||
expected: true,
|
||||
name: "contains UPX magic with offset",
|
||||
data: append(append(append(packedELFStub(), make([]byte, 500)...), []byte("UPX!")...), make([]byte, 100)...),
|
||||
foundMagic: true,
|
||||
},
|
||||
{
|
||||
name: "no UPX magic",
|
||||
data: []byte("\x7FELF" + string(make([]byte, 100))),
|
||||
expected: false,
|
||||
// UPX packs PE and Mach-O with this same l_info/p_info layout, and nothing here can place
|
||||
// their blocks, so the container has to be checked before the magic is trusted
|
||||
name: "UPX magic but not an ELF container",
|
||||
data: append([]byte("MZ\x90\x00UPX!"), make([]byte, 100)...),
|
||||
},
|
||||
{
|
||||
name: "empty data",
|
||||
data: []byte{},
|
||||
expected: false,
|
||||
name: "no UPX magic",
|
||||
data: append(packedELFStub(), make([]byte, 100)...),
|
||||
},
|
||||
{
|
||||
name: "partial UPX magic",
|
||||
data: []byte("UPX"),
|
||||
expected: false,
|
||||
// three of the four magic bytes must not match
|
||||
name: "partial UPX magic",
|
||||
data: append(append(packedELFStub(), []byte("UPX")...), make([]byte, 100)...),
|
||||
},
|
||||
{
|
||||
// never reaches the scan: the six byte ident read comes up short, so the container gate
|
||||
// refuses it first. Kept because an empty reader is a shape the cataloger really is handed.
|
||||
name: "empty data",
|
||||
data: []byte{},
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
reader := bytes.NewReader(tt.data)
|
||||
result := isUPXCompressed(reader)
|
||||
assert.Equal(t, tt.expected, result)
|
||||
_, err := parseUPXInfo(bytes.NewReader(tt.data), int64(len(tt.data)))
|
||||
require.Error(t, err, "none of these fixtures is a usable UPX header")
|
||||
if tt.foundMagic {
|
||||
assert.NotErrorIs(t, err, errNotUPX, "the magic was found, so the header was parsed and rejected")
|
||||
} else {
|
||||
assert.ErrorIs(t, err, errNotUPX)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseUPXInfo_NotUPX(t *testing.T) {
|
||||
data := []byte("\x7FELF" + string(make([]byte, 100)))
|
||||
reader := bytes.NewReader(data)
|
||||
|
||||
_, err := parseUPXInfo(reader)
|
||||
require.Error(t, err)
|
||||
assert.ErrorIs(t, err, errNotUPX)
|
||||
}
|
||||
|
||||
func TestParseUPXInfo_ValidHeader(t *testing.T) {
|
||||
// construct a minimal valid UPX header matching actual format
|
||||
// l_info: checksum (4) + magic (4) + lsize (2) + version (1) + format (1)
|
||||
@@ -84,45 +91,16 @@ func TestParseUPXInfo_ValidHeader(t *testing.T) {
|
||||
14, 0, 0, 0, // method=LZMA, filter info
|
||||
}
|
||||
|
||||
data := append(append(lInfo, pInfo...), bInfo...)
|
||||
data = append(data, make([]byte, 100)...) // padding
|
||||
// padded so the declared 1MB stays within maxUPXExpansion of the fixture's own size; a real UPX file
|
||||
// carries the compressed data this header describes, and the bound is measured against that.
|
||||
data := append(append(append(packedELFStub(), lInfo...), pInfo...), bInfo...)
|
||||
data = append(data, make([]byte, 0x100000/maxUPXExpansion)...)
|
||||
|
||||
reader := bytes.NewReader(data)
|
||||
info, err := parseUPXInfo(reader)
|
||||
info, err := parseUPXInfo(reader, sizeOf(t, reader))
|
||||
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, uint8(14), info.version)
|
||||
assert.Equal(t, uint8(22), info.format)
|
||||
assert.Equal(t, uint32(0x100000), info.originalSize)
|
||||
}
|
||||
|
||||
func TestDecompressUPX_UnsupportedMethod(t *testing.T) {
|
||||
// construct a header with an unsupported compression method
|
||||
lInfo := []byte{
|
||||
0, 0, 0, 0, // l_checksum
|
||||
'U', 'P', 'X', '!',
|
||||
0, 0, // l_lsize
|
||||
14, 22, // version, format
|
||||
}
|
||||
|
||||
pInfo := []byte{
|
||||
0, 0, 0, 0, // p_progid
|
||||
0x00, 0x01, 0x00, 0x00, // p_filesize = 256 bytes (small for test)
|
||||
0, 0, 0x10, 0, // p_blocksize
|
||||
}
|
||||
|
||||
bInfo := []byte{
|
||||
0x00, 0x01, 0x00, 0x00, // sz_unc = 256
|
||||
0x80, 0x00, 0x00, 0x00, // sz_cpr = 128
|
||||
99, 0, 0, 0, // unsupported method
|
||||
}
|
||||
|
||||
data := append(append(lInfo, pInfo...), bInfo...)
|
||||
data = append(data, make([]byte, 1000)...)
|
||||
|
||||
reader := bytes.NewReader(data)
|
||||
_, err := decompressUPX(reader)
|
||||
|
||||
require.Error(t, err)
|
||||
assert.ErrorIs(t, err, errUnsupportedUPXMethod)
|
||||
}
|
||||
|
||||
@@ -103,3 +103,17 @@ func noDirectELFOpen(m dsl.Matcher) {
|
||||
Where(!m.File().PkgPath.Matches(`/syft/internal/elfutil$`)).
|
||||
Report("do not open ELF files with debug/elf directly; use elfutil.NewFile, which bounds declared decompressed section sizes")
|
||||
}
|
||||
|
||||
// nolint:unused
|
||||
func noUnboundedELFParsers(m dsl.Matcher) {
|
||||
// buildinfo.Read and goversion's ReadExeFromReader each open debug/elf internally and expand sections
|
||||
// sized by the file itself, the same unbounded allocation noDirectELFOpen guards against above. They
|
||||
// are only safe in this repo because scan_binary.go gates the reader before calling them: readBuildInfo
|
||||
// runs elfutil.CheckSectionNameTable first, getCryptoInformation runs elfutil.CheckAllSections first.
|
||||
m.Match(
|
||||
`buildinfo.Read($_)`,
|
||||
`version.ReadExeFromReader($_)`,
|
||||
).
|
||||
Where(!(m.File().Name.Matches(`scan_binary\.go$`) && m.File().PkgPath.Matches(`/syft/pkg/cataloger/golang$`))).
|
||||
Report("do not call buildinfo.Read/goversion.ReadExeFromReader directly; route through the bounded wrappers in the golang cataloger (readBuildInfo, getCryptoInformation), or gate the reader with elfutil.CheckSectionNameTable/elfutil.CheckAllSections first")
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user