AvroSharp and Apache.Avro

Apache.Avro is the Apache Software Foundation's C# library for Avro. AvroSharp is an independent implementation of the same specification, written from the specification for current .NET. The two read and write the same bytes: AvroSharp's interop tests check it against Apache.Avro 1.12.2 in both directions, and against files that Apache Avro Java wrote.

This page covers what AvroSharp does differently, where Apache.Avro is the better fit, and how to move code over.

On this page:

At a glance

AvroSharp Apache.Avro 1.12.2
Speed faster on every benchmark: records 1.95–6.18×, container reads up to 22×, schema parsing up to 2.07× (benchmarks) the baseline
Allocations no more than Apache.Avro on any benchmark; a container file written with under 6 KB (except with xz) 5.7 MB for the same container file
Code generation a source generator: types are generated while the project builds (guide); or the avrosharp tool (CLI) the avrogen tool
Serializers generated, with no reflection; Native AOT and trimming compatible, checked by a Native AOT build in CI specific reader and writer driven by the schema and reflection at run time
Codecs null, deflate, snappy, zstandard, bzip2, xz: two packages, fully managed null and deflate, and a satellite package per codec
Schema registries Confluent, Confluent GUID, Apicurio and AWS Glue framing, with no client dependency; serializers for Confluent.Kafka, KafkaFlow, Azure Schema Registry and AWS Glue in add-on packages, new in 1.0.0 (integrations) not included; Confluent's and Microsoft's Avro serializers are built on it
Specification follows it where Apache.Avro 1.12.2 deviates (details) five deviations pinned by AvroSharp's tests
Hostile input bounded memory and nesting for malformed or hostile data and schemas; fuzzed nightly —
Targets .NET 8, 9 and 10, .NET Standard 2.0 and 2.1 (so .NET Framework) .NET Standard 2.0 and 2.1
License MIT Apache 2.0
Maturity 1.0: stable, following semantic versioning the long-standing reference implementation

Speed and allocations

A release requires every AvroSharp benchmark to be faster than its Apache.Avro counterpart and to allocate no more (the performance gate). On the last full run written up (i7-12800H, .NET 10):

  • Generated records read 3.95× and write 6.18× faster; the generic model reads 1.95× and writes 3.93× faster.
  • Container files read 2.9–7.4× faster with the fast codecs, and up to 22× with xz; writes are 3.5–10× faster.
  • Reading an older schema version is 2.72× faster with generated code.
  • Writes allocate close to nothing: under 6 KB for a whole container file, against 5.7 MB (xz takes 777 KB).

Where the gain comes from:

  • Generated serializers call the writer and reader directly, in schema order, with no schema walk, boxing or virtual calls per value.
  • Spans and buffer writers: AvroWriter and AvroReader work over Span<byte>, IBufferWriter<byte> and ReadOnlySequence<byte>, with pooled buffers and no per-value streams.
  • A value type for generic data: AvroValue holds any Avro value without boxing, and arrays of primitives are stored as the primitives.
  • Fast primitives: varints decoded and encoded without a loop per byte where the CPU allows, and bulk array reads, each kept only if it beats Apache.Avro on every tested CPU, with one exception for single-value writes on Zen+ CPUs (rules).

The benchmarks page has the full table and how to run it yourself.

Generated code instead of reflection

AvroSharp's source generator turns .avsc files into C# types as the project builds: there is no separate generation step to run and no generated code to check in, and the types update as the schemas change. Each record gets serializers that write and read its fields directly, so they work with Native AOT and trimming; CI publishes a Native AOT app and fails on any trim or AOT warning.

Generated types also get what hand-written code usually adds around Avro:

  • ToAvroBytes()/FromAvroBytes(), and writing into caller memory without allocating (TryWriteAvroBytes);
  • reading into an existing instance, reusing its collections (ReadFrom);
  • schema defaults applied by the constructor;
  • DateOnly, TimeOnly, DateTimeOffset, Guid and decimal for logical types, with decimals written exactly or rejected, never rounded;
  • IAvroSerializable<T> on .NET 8 and later, so files, streams and messages take the type with no delegates.

The generator also works the other way round. A partial class you wrote, marked [AvroSerializable], gets its schema and the same serializers from its members. Apache.Avro's reflect API does that job at run time: ReflectWriter<T> and ReflectReader<T> read and write existing classes, against a schema you write, by reflection. With [AvroSerializable] the schema comes from the class, and the code is generated at compile time. On the order record of the benchmarks, that's 4.7× faster reads and 7.2× faster writes than Apache.Avro (the run).

Following the specification

AvroSharp follows the Avro 1.12 specification, and matches Apache Avro Java where the specification leaves a choice. Where Apache.Avro 1.12.2 (C#) differs from the specification, AvroSharp's tests pin Apache's behavior, so a change in a later Apache release is noticed:

Case Specification and AvroSharp Apache.Avro 1.12.2
uuid on fixed(16) a valid logical type (1.12) rejects the schema
fixed of size 0 valid rejects it
a fixed with a logical type, in canonical form the fixed definition stays the definition is dropped, leaving only the name
the empty namespace "" inside a namespace the null namespace inherits the enclosing namespace
a null-namespace type referenced from inside a namespace found, as in Java an undefined name

Also checked against Java:

  • Canonical form and fingerprints pass all 34 of Apache's test vectors.
  • ToJson() writes a schema as Java's Schema.toString() does, byte for byte (attribute order, and numbers as Java prints them), so registering AvroSharp's text finds the version a Java client registered.
  • Schema resolution picks union branches by full name, then by unqualified name, then by promotion, as Java does, and applies reader aliases before names.
  • Container files from Java, in every codec, are read, and Java reads the ones AvroSharp writes.
  • A container file's schema is read as Java reads it, without validating names or defaults: a field named user-id, or a "default": null on a string field, is common in files that Java and older tools wrote, and the data is readable. Java's other leniencies there are accepted too: a "doc" or an enum's "default" of null, and a namespace that is not a string. Schemas given to AvroSchema.Parse are still validated.
  • Logical types don't take part in schema resolution, as in Java. A decimal of another scale or precision is read with the reader's, so 1.23 written with scale 2 reads as 0.123 with scale 3. The specification says such decimals don't match, but Java reads them, so AvroSharp does too, and files Java reads stay readable. AvroSchemaCompatibility warns about it (DecimalChanged), and fails it with WarningsAsErrors.

Hostile input

Avro data often comes from outside the process: files, messages, and schemas embedded in them. AvroSharp bounds what malformed or hostile input can make it do, and reports it as AvroDataException or AvroSchemaException:

  • Memory: every length is checked against the remaining input before anything is allocated; container blocks, schemas in file headers, stream objects and zero-size values have limits.
  • Nesting: records, arrays and maps are limited in depth, and the thread's stack is checked, so a hostile schema cannot overflow the stack.
  • Codecs: a corrupt block ends in InvalidDataException, never in a library's internal exception.
  • Fuzzing: libFuzzer targets cover the readers, and run every night.

The limits are options (GenericDatumReaderOptions, AvroFileReaderOptions, AvroStreamOptions) with defaults that fit ordinary data.

More of the Avro ecosystem in one library

  • Every codec in the specification: null and deflate in AvroSharp, and snappy, zstandard, bzip2 and xz in AvroSharp.Codecs, on fully managed libraries with no native binaries.
  • Schema registries: Confluent (4-byte ID and GUID), Apicurio and AWS Glue wire framing (AvroRegistryMessage), with an ID resolver you supply (IAvroSchemaIdResolver), and no registry client dependency (guide). Serializers for Confluent.Kafka, KafkaFlow, Azure Schema Registry and AWS Glue, without Apache.Avro, are add-on packages, new in 1.0.0 (integrations).
  • Single-object encoding (AvroMessage) with a schema store (AvroSchemaStore) selecting the writer schema by fingerprint (guide).
  • Streams of objects without a container, for sockets and pipes (guide).
  • Asynchronous container files, with no synchronous I/O, and pipelined reading that decompresses the next block while the current one is decoded.
  • JSON encoding of the generic model, checked against Apache.Avro's JSON encoder and decoder.
  • Canonical forms and fingerprints (CRC-64-AVRO, MD5, SHA-256: SchemaFingerprint), also from the command line (CLI).

When Apache.Avro is the better fit

  • You serialize classes you can't change. Apache.Avro's reflect API works with any class, at run time. AvroSharp's [AvroSerializable] needs a partial class it can add to, and doesn't support init-only members or primary constructors yet.
  • You want the Apache Software Foundation's implementation, maintained alongside the other Avro languages.

Moving from Apache.Avro

Migrating from Apache.Avro maps the APIs one by one, with code for both libraries. The two libraries interoperate on the wire, so services can move one at a time. Within one codebase:

  1. Generate types that work with both. Reference AvroSharp.Generators and set <AvroSharpApacheCompatible>true</AvroSharpApacheCompatible>. The generated classes replace avrogen's: code written for them compiles unchanged (the same property names, _SCHEMA and Schema), and Apache's SpecificDatumWriter<T>/SpecificDatumReader<T> still accept them. See the compatibility mode.
  2. Move call sites one at a time to AvroSharp's serializers: order.ToAvroBytes(), AvroFileWriter.Create<Order>(stream), AvroFileReader.Open<Order>(stream), AvroMessage, AvroRegistryMessage.
  3. Drop the compatibility mode and the Apache.Avro reference when nothing uses Apache's API any more. The types then use .NET types for logical types (DateOnly, decimal) and, by default, PascalCase properties; set AvroSharpPropertyNames to avro to keep the field names.

If you check in generated code, the avrosharp tool takes the place of avrogen, with --apache-compatible for step 1.


Apache Avro, Avro and Apache are trademarks of The Apache Software Foundation. AvroSharp is an independent project and is not endorsed by or affiliated with the Apache Software Foundation.