svcadm(8)을 검색하려면 섹션에서 8 을 선택하고, 맨 페이지 이름에 svcadm을 입력하고 검색을 누른다.
bpf(4)
The Berkeley Packet Filter provides a raw interface to data link
layers in a protocol independent fashion. All packets on the
network, even those destined for other hosts, are accessible
through this mechanism. The packet filter appears as a character
special device, After opening the device, the file descriptor
must be bound to a specific network interface with the ioctl. A
given interface can be shared by multiple listeners, and the fil‐
ter underlying each descriptor will see an identical packet
stream. A separate device file is required for each minor de‐
vice. If a file is in use, the open will fail and will be set to
Associated with each open instance of a file is a user-settable
packet filter. Whenever a packet is received by an interface,
all file descriptors listening on that interface apply their fil‐
ter. Each descriptor that accepts the packet receives its own
copy. The packet filter will support any link level protocol
that has fixed length headers. Currently, only Ethernet, and
drivers have been modified to interact with Since packet data is
in network byte order, applications should use the macros to ex‐
tract multi-byte values. A packet can be sent out on the network
by writing to a file descriptor. The writes are unbuffered,
meaning only one packet can be processed per write. Currently,
only writes to Ethernets and links are supported. devices de‐
liver packet data to the application via memory buffers provided
by the application. The buffer mode is set using the ioctl, and
read using the ioctl. By default, devices operate in the mode,
in which packet data is copied explicitly from kernel to user
memory using the system call. The user process will declare a
fixed buffer size that will be used both for sizing internal
buffers and for all operations on the file. This size is queried
using the ioctl, and is set using the ioctl. Note that an indi‐
vidual packet larger than the buffer size is necessarily trun‐
cated. devices may also operate in the mode, in which packet
data is written directly into two user memory buffers by the ker‐
nel, avoiding both system call and copying overhead. Buffers are
of fixed (and equal) size, page-aligned, and an even multiple of
the page size. The maximum zero-copy buffer size is returned by
the ioctl. Note that an individual packet larger than the buffer
size is necessarily truncated. The user process registers two
memory buffers using the ioctl, which accepts a pointer as an ar‐
gument: struct bpf_zbuf { void *bz_bufa; void
*bz_bufb; size_t bz_buflen; }; is a pointer to the user‐
space address of the first buffer that will be filled, and is a
pointer to the second buffer. will then cycle between the two
buffers as they fill and are acknowledged. Each buffer begins
with a fixed-length header to hold synchronization and data
length information for the buffer: struct bpf_zbuf_header {
volatile u_int bzh_kernel_gen; /* Kernel generation
number. */ volatile u_int bzh_kernel_len; /* Length of
data in the buffer. */ volatile u_int bzh_user_gen; /*
User generation number. */ /* ...padding for future
use... */ }; The header structure of each buffer, including all
padding, should be zeroed before it is configured using Remaining
space in the buffer will be used by the kernel to store packet
data, laid out in the same format as with buffered read mode.
The kernel and the user process follow a simple acknowledgement
protocol via the buffer header to synchronize access to the
buffer: when the header generation numbers, and hold the same
value, the kernel owns the buffer, and when they differ, user‐
space owns the buffer. While the kernel owns the buffer, the
contents are unstable and may change asynchronously; while the
user process owns the buffer, its contents are stable and will
not be changed until the buffer has been acknowledged. Initial‐
izing the buffer headers to all 0's before registering the buffer
has the effect of assigning initial ownership of both buffers to
the kernel. The kernel signals that a buffer has been assigned
to userspace by modifying and userspace acknowledges the buffer
and returns it to the kernel by setting the value of to the value
of In order to avoid caching and memory re-ordering effects, the
user process must use atomic operations and memory barriers when
checking for and acknowledging buffers: #include <ma‐
chine/atomic.h>
/*
* Return ownership of a buffer to the kernel for reuse.
*/ static void buffer_acknowledge(struct bpf_zbuf_header *bzh) {
atomic_store_rel_int(&bzh->bzh_user_gen, bzh->bzh_ker‐
nel_gen); }
/*
* Check whether a buffer has been assigned to userspace by the
kernel.
* Return true if userspace owns the buffer, and false otherwise.
*/ static int buffer_check(struct bpf_zbuf_header *bzh) {
return (bzh->bzh_user_gen !=
atomic_load_acq_int(&bzh->bzh_kernel_gen)); } The user process
may force the assignment of the next buffer, if any data is pend‐
ing, to userspace using the ioctl. This allows the user process
to retrieve data in a partially filled buffer before the buffer
is full, such as following a timeout; the process must recheck
for buffer ownership using the header generation numbers, as the
buffer will not be assigned to userspace if no data was present.
As in the buffered read mode, and may be used to sleep awaiting
the availability of a completed buffer. They will return a read‐
able file descriptor when ownership of the next buffer is as‐
signed to user space. In the current implementation, the kernel
may assign zero, one, or both buffers to the user process; how‐
ever, an earlier implementation maintained the invariant that at
most one buffer could be assigned to the user process at a time.
In order to both ensure progress and high performance, user
processes should acknowledge a completely processed buffer as
quickly as possible, returning it for reuse, and not block wait‐
ing on a second buffer while holding another buffer. The command
codes below are defined in All commands require these includes:
#include <sys/types.h> #include <sys/time.h>
#include <sys/ioctl.h> #include <net/bpf.h> Addi‐
tionally, and require and In addition to the following commands
may be applied to any open file. The (third) argument to should
be a pointer to the type indicated. Returns the required buffer
length for reads on files. Sets the buffer length for reads on
files. The buffer must be set before the file is attached to an
interface with If the requested buffer size cannot be accommo‐
dated, the closest allowable size will be set and returned in the
argument. A read call will result in if it is passed a buffer
that is not this size. Returns the type of the data link layer
underlying the attached interface. is returned if no interface
has been specified. The device types, prefixed with are defined
in Forces the interface into promiscuous mode. All packets, not
just those destined for the local host, are processed. Since
more than one file can be listening on a given interface, a lis‐
tener that opened its interface non-promiscuously may receive
packets promiscuously. This problem can be remedied with an ap‐
propriate filter. Flushes the buffer of incoming packets, and
resets the statistics that are returned by BIOCGSTATS. Returns
the name of the hardware interface that the file is listening on.
The name is returned in the ifr_name field of the structure. All
other fields are undefined. Sets the hardware interface asso‐
ciate with the file. This command must be performed before any
packets can be read. The device is indicated by name using the
field of the structure. Additionally, performs the actions of
Set or get the read timeout parameter. The argument specifies
the length of time to wait before timing out on a read request.
This parameter is initialized to zero by indicating no timeout.
Returns the following structure of packet statistics: struct
bpf_stat { u_int bs_recv; /* number of packets re‐
ceived */ u_int bs_drop; /* number of packets dropped
*/ }; The fields are: the number of packets received by the de‐
scriptor since opened or reset (including any buffered since the
last read call); and the number of packets which were accepted by
the filter but dropped by the kernel because of buffer overflows
(i.e., the application's reads are not keeping up with the packet
traffic). Enable or disable based on the truth value of the ar‐
gument. When immediate mode is enabled, reads return immediately
upon packet reception. Otherwise, a read will block until either
the kernel buffer becomes full or a timeout occurs. This is use‐
ful for programs like which must respond to messages in real
time. The default for a new file is off. Sets the read filter
program used by the kernel to discard uninteresting packets. An
array of instructions and its length is passed in using the fol‐
lowing structure: struct bpf_program { int bf_len;
struct bpf_insn *bf_insns; }; The filter program is
pointed to by the field while its length in units of is given by
the field. See section for an explanation of the filter lan‐
guage. The only difference between and is performs the actions
of while does not. Sets the write filter program used by the
kernel to control what type of packets can be written to the in‐
terface. See the command for more information on the filter pro‐
gram. Returns the major and minor version numbers of the filter
language currently recognized by the kernel. Before installing a
filter, applications must check that the current version is com‐
patible with the running kernel. Version numbers are compatible
if the major numbers match and the application minor is less than
or equal to the kernel minor. The kernel version number is re‐
turned in the following structure: struct bpf_version {
u_short bv_major;
u_short bv_minor; }; The current version numbers are
given by and from An incompatible filter may result in undefined
behavior (most likely, an error returned by or haphazard packet
matching). Set or get the status of the flag. Set to zero if
the link level source address should be filled in automatically
by the interface output routine. Set to one if the link level
source address will be written, as provided, to the wire. This
flag is initialized to zero by default. These commands are obso‐
lete but left for compatibility. Use and instead. Set or get
the flag determining whether locally generated packets on the in‐
terface should be returned by BPF. Set to zero to see only in‐
coming packets on the interface. Set to one to see packets orig‐
inating locally and remotely on the interface. This flag is ini‐
tialized to one by default. Set or get the setting determining
whether incoming, outgoing, or all packets on the interface
should be returned by BPF. Set to to see only incoming packets
on the interface. Set to to see packets originating locally and
remotely on the interface. Set to to see only outgoing packets
on the interface. This setting is initialized to by default.
Set or get format and resolution of the time stamps returned by
BPF. Set to or to get time stamps in 64-bit format. Set to or
to get time stamps in 64-bit format. Set to or to get time
stamps in 64-bit format. Set to to ignore time stamp. All
64-bit time stamp formats are wrapped in The and are analogs of
corresponding formats without _FAST suffix but do not perform a
full time counter query, so their accuracy is one timer tick.
The and store the time elapsed since kernel boot. This setting
is initialized to by default. Set packet feedback mode. This
allows injected packets to be fed back as input to the interface
when output via the interface is successful. When direction is
set, injected outgoing packet is not returned by BPF to avoid du‐
plication. This flag is initialized to zero by default. Set the
locked flag on the descriptor. This prevents the execution of
ioctl commands which could change the underlying operating para‐
meters of the device. Get or set the current buffering mode;
possible values are buffered read mode, and zero-copy buffer
mode. Set the current zero-copy buffer locations; buffer loca‐
tions may be set only once zero-copy buffer mode has been se‐
lected, and prior to attaching to an interface. Buffers must be
of identical size, page-aligned, and an integer multiple of pages
in size. The three fields and must be filled out. If buffers
have already been set for this device, the ioctl will fail. Get
the largest individual zero-copy buffer size allowed. As two
buffers are used in zero-copy buffer mode, the limit (in prac‐
tice) is twice the returned size. As zero-copy buffers consume
kernel address space, conservative selection of buffer size is
suggested, especially when there are multiple descriptors in use
on 32-bit systems. Force ownership of the next buffer to be as‐
signed to userspace, if any data present in the buffer. If no
data is present, the buffer will remain owned by the kernel.
This allows consumers of zero-copy buffering to implement time‐
outs and retrieve partially filled buffers. In order to handle
the case where no data is present in the buffer and therefore
ownership is not assigned, the user process must check against
One of the following structures is prepended to each packet re‐
turned by or via a zero-copy buffer: struct bpf_xhdr {
struct bpf_ts bh_tstamp; /* time stamp */
uint32_t bh_caplen; /* length of captured por‐
tion */ uint32_t bh_datalen; /* original length
of packet */ u_short bh_hdrlen; /* length of
bpf header (this struct
plus alignment padding) */ };
struct bpf_hdr { struct timeval bh_tstamp; /* time
stamp */ uint32_t bh_caplen; /* length of cap‐
tured portion */ uint32_t bh_datalen; /* origi‐
nal length of packet */ u_short bh_hdrlen; /*
length of bpf header (this struct
plus alignment padding)
*/ }; The fields, whose values are stored in host order, and are:
The time at which the packet was processed by the packet filter.
The length of the captured portion of the packet. This is the
minimum of the truncation amount specified by the filter and the
length of the packet. The length of the packet off the wire.
This value is independent of the truncation amount specified by
the filter. The length of the header, which may not be equal to
or The field exists to account for padding between the header and
the link level protocol. The purpose here is to guarantee proper
alignment of the packet data structures, which is required on
alignment sensitive architectures and improves performance on
many other architectures. The packet filter ensures that the and
the network layer header will be word aligned. Currently, is
used when the time stamp is set to or for backward compatibility
reasons. Otherwise, is used. However, may be deprecated in the
near future. Suitable precautions must be taken when accessing
the link layer protocol fields on alignment restricted machines.
(This is not a problem on an Ethernet, since the type field is a
short falling on an even offset, and the addresses are probably
accessed in a bytewise fashion). Additionally, individual pack‐
ets are padded so that each starts on a word boundary. This re‐
quires that an application has some knowledge of how to get from
packet to packet. The macro is defined in to facilitate this
process. It rounds up its argument to the nearest word aligned
value (where a word is bytes wide). For example, if points to
the start of a packet, this expression will advance it to the
next packet: For the alignment mechanisms to work properly, the
buffer passed to must itself be word aligned. The function will
always return an aligned buffer. A filter program is an array of
instructions, with all branches forwardly directed, terminated by
a instruction. Each instruction performs some action on the
pseudo-machine state, which consists of an accumulator, index
register, scratch memory store, and implicit program counter.
The following structure defines the instruction format: struct
bpf_insn { u_short code; u_char jt;
u_char jf; u_long k; }; The field is used in
different ways by different instructions, and the and fields are
used as offsets by the branch instructions. The opcodes are en‐
coded in a semi-hierarchical fashion. There are eight classes of
instructions: and Various other mode and operator bits are or'd
into the class to give the actual instructions. The classes and
modes are defined in Below are the semantics for each defined in‐
struction. We use the convention that A is the accumulator, X is
the index register, P[] packet data, and M[] scratch memory
store. P[i:n] gives the data at byte offset in the packet, in‐
terpreted as a word (n=4), unsigned halfword (n=2), or unsigned
byte (n=1). M[i] gives the i'th word in the scratch memory
store, which is only addressed in word units. The memory store
is indexed from 0 to - 1. and are the corresponding fields in
the instruction definition. refers to the length of the packet.
These instructions copy a value into the accumulator. The type
of the source operand is specified by an and can be a constant
packet data at a fixed offset packet data at a variable offset
the packet length or a word in the scratch memory store For and
the data size must be specified as a word halfword or byte The
semantics of all the recognized instructions follow.
BPF_LD+BPF_W+BPF_ABS A <- P[k:4] BPF_LD+BPF_H+BPF_ABS A <-
P[k:2] BPF_LD+BPF_B+BPF_ABS A <- P[k:1]
BPF_LD+BPF_W+BPF_IND A <- P[X+k:4] BPF_LD+BPF_H+BPF_IND A
<- P[X+k:2] BPF_LD+BPF_B+BPF_IND A <- P[X+k:1]
BPF_LD+BPF_W+BPF_LEN A <- len BPF_LD+BPF_IMM A <- k
BPF_LD+BPF_MEM A <- M[k] These instructions load a value
into the index register. Note that the addressing modes are more
restrictive than those of the accumulator loads, but they include
a hack for efficiently loading the IP header length.
BPF_LDX+BPF_W+BPF_IMM X <- k BPF_LDX+BPF_W+BPF_MEM X <- M[k]
BPF_LDX+BPF_W+BPF_LEN X <- len BPF_LDX+BPF_B+BPF_MSH X <-
4*(P[k:1]&0xf) This instruction stores the accumulator into the
scratch memory. We do not need an addressing mode since there is
only one possibility for the destination.
BPF_ST M[k] <- A This instruction stores the in‐
dex register in the scratch memory store.
BPF_STX M[k] <- X The alu instructions perform
operations between the accumulator and index register or con‐
stant, and store the result back in the accumulator. For binary
operations, a source mode is required or
BPF_ALU+BPF_ADD+BPF_K A <- A + k BPF_ALU+BPF_SUB+BPF_K A <- A
- k BPF_ALU+BPF_MUL+BPF_K A <- A * k BPF_ALU+BPF_DIV+BPF_K A
<- A / k BPF_ALU+BPF_MOD+BPF_K A <- A % k
BPF_ALU+BPF_AND+BPF_K A <- A & k BPF_ALU+BPF_OR+BPF_K A <- A
| k BPF_ALU+BPF_XOR+BPF_K A <- A ^ k BPF_ALU+BPF_LSH+BPF_K A
<- A << k BPF_ALU+BPF_RSH+BPF_K A <- A >> k
BPF_ALU+BPF_ADD+BPF_X A <- A + X BPF_ALU+BPF_SUB+BPF_X A <- A
- X BPF_ALU+BPF_MUL+BPF_X A <- A * X BPF_ALU+BPF_DIV+BPF_X A
<- A / X BPF_ALU+BPF_MOD+BPF_X A <- A % X
BPF_ALU+BPF_AND+BPF_X A <- A & X BPF_ALU+BPF_OR+BPF_X A <- A
| X BPF_ALU+BPF_XOR+BPF_X A <- A ^ X BPF_ALU+BPF_LSH+BPF_X A
<- A << X BPF_ALU+BPF_RSH+BPF_X A <- A >> X
BPF_ALU+BPF_NEG A <- -A The jump instructions alter flow
of control. Conditional jumps compare the accumulator against a
constant or the index register If the result is true (or non-
zero), the true branch is taken, otherwise the false branch is
taken. Jump offsets are encoded in 8 bits so the longest jump is
256 instructions. However, the jump always opcode uses the 32
bit field as the offset, allowing arbitrarily distant destina‐
tions. All conditionals use unsigned comparison conventions.
BPF_JMP+BPF_JA pc += k BPF_JMP+BPF_JGT+BPF_K pc += (A
> k) ? jt : jf BPF_JMP+BPF_JGE+BPF_K pc += (A >= k) ? jt : jf
BPF_JMP+BPF_JEQ+BPF_K pc += (A == k) ? jt : jf
BPF_JMP+BPF_JSET+BPF_K pc += (A & k) ? jt : jf
BPF_JMP+BPF_JGT+BPF_X pc += (A > X) ? jt : jf
BPF_JMP+BPF_JGE+BPF_X pc += (A >= X) ? jt : jf
BPF_JMP+BPF_JEQ+BPF_X pc += (A == X) ? jt : jf
BPF_JMP+BPF_JSET+BPF_X pc += (A & X) ? jt : jf The return in‐
structions terminate the filter program and specify the amount of
packet to accept (i.e., they return the truncation amount). A
return value of zero indicates that the packet should be ignored.
The return value is either a constant or the accumulator
BPF_RET+BPF_A accept A bytes
BPF_RET+BPF_K accept k bytes The miscellaneous category
was created for anything that does not fit into the above
classes, and for any new instructions that might need to be
added. Currently, these are the register transfer instructions
that copy the index register to the accumulator or vice versa.
BPF_MISC+BPF_TAX X <- A BPF_MISC+BPF_TXA A <- X The
interface provides the following macros to facilitate array ini‐
tializers: and A set of variables controls the behaviour of the
subsystem Various programs use BPF to send (but not receive) raw
packets (cdpd, lldpd, dhcpd, dhcp relays, etc. are good examples
of such programs). They do not need incoming packets to be send
to them. Turning this option on makes new BPF users to be at‐
tached to write-only interface list until program explicitly
specifies read filter via This removes any performance degrada‐
tion for high-speed interfaces. Binary interface for retrieving
general statistics. Permits zero-copy to be used with net BPF
readers. Use with caution. Maximum number of instructions that
BPF program can contain. Use option to determine approximate
number of instruction for any filter. Maximum buffer size to al‐
locate for packets buffer. Default buffer size to allocate for
packets buffer. The following filter is taken from the Reverse
ARP Daemon. It accepts only Reverse ARP requests. struct
bpf_insn insns[] = { BPF_STMT(BPF_LD+BPF_H+BPF_ABS, 12),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, ETHERTYPE_REVARP, 0, 3),
BPF_STMT(BPF_LD+BPF_H+BPF_ABS, 20),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, REVARP_REQUEST, 0, 1),
BPF_STMT(BPF_RET+BPF_K, sizeof(struct ether_arp) +
sizeof(struct ether_header)),
BPF_STMT(BPF_RET+BPF_K, 0), }; This filter accepts only
IP packets between host 128.3.112.15 and 128.3.112.35. struct
bpf_insn insns[] = { BPF_STMT(BPF_LD+BPF_H+BPF_ABS, 12),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, ETHERTYPE_IP, 0, 8),
BPF_STMT(BPF_LD+BPF_W+BPF_ABS, 26),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, 0x8003700f, 0, 2),
BPF_STMT(BPF_LD+BPF_W+BPF_ABS, 30),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, 0x80037023, 3, 4),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, 0x80037023, 0, 3),
BPF_STMT(BPF_LD+BPF_W+BPF_ABS, 30),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, 0x8003700f, 0, 1),
BPF_STMT(BPF_RET+BPF_K, (u_int)-1),
BPF_STMT(BPF_RET+BPF_K, 0), }; Finally, this filter re‐
turns only TCP finger packets. We must parse the IP header to
reach the TCP header. The instruction checks that the IP frag‐
ment offset is 0 so we are sure that we have a TCP header.
struct bpf_insn insns[] = {
BPF_STMT(BPF_LD+BPF_H+BPF_ABS, 12),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, ETHERTYPE_IP, 0, 10),
BPF_STMT(BPF_LD+BPF_B+BPF_ABS, 23),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, IPPROTO_TCP, 0, 8),
BPF_STMT(BPF_LD+BPF_H+BPF_ABS, 20),
BPF_JUMP(BPF_JMP+BPF_JSET+BPF_K, 0x1fff, 6, 0),
BPF_STMT(BPF_LDX+BPF_B+BPF_MSH, 14),
BPF_STMT(BPF_LD+BPF_H+BPF_IND, 14),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, 79, 2, 0),
BPF_STMT(BPF_LD+BPF_H+BPF_IND, 16),
BPF_JUMP(BPF_JMP+BPF_JEQ+BPF_K, 79, 0, 1),
BPF_STMT(BPF_RET+BPF_K, (u_int)-1),
BPF_STMT(BPF_RET+BPF_K, 0), }; The Enet packet filter was
created in 1980 by Mike Accetta and Rick Rashid at Carnegie-Mel‐
lon University. Jeffrey Mogul, at Stanford, ported the code to
and continued its development from 1983 on. Since then, it has
evolved into the Ultrix Packet Filter at a module under and of
Lawrence Berkeley Laboratory, implemented BPF in Summer 1990.
Much of the design is due to Support for zero-copy buffers was
added by under contract to Seccuris Inc. The read buffer must be
of a fixed size (returned by the ioctl). A file that does not
request promiscuous mode may receive promiscuously received pack‐
ets as a side effect of another file requesting this mode on the
same hardware interface. This could be fixed in the kernel with
additional processing overhead. However, we favor the model
where all files must assume that the interface is promiscuous,
and if so desired, must utilize a filter to reject foreign pack‐
ets. Data link protocols with variable length headers are not
currently supported. The and settings have been observed to work
incorrectly on some interface types, including those with hard‐
ware loopback rather than software loopback, and point-to-point
interfaces. They appear to function correctly on a broad range
of Ethernet-style interfaces.