Releases: quickwit-oss/tantivy
Tantivy v0.22
What's Changed
Tantivy 0.22 will be able to read indices created with Tantivy 0.21.
Bugfixes
- Fix null byte handling in JSON paths (null bytes in json keys caused panic during indexing) #2345(@PSeitz)
- Fix bug that can cause
get_docids_for_value_range
to panic. #2295(@fulmicoton) - Avoid 1 document indices by increase min memory to 15MB for indexing #2176(@PSeitz)
- Fix merge panic for JSON fields #2284(@PSeitz)
- Fix bug occuring when merging JSON object indexed with positions. #2253(@fulmicoton)
- Fix empty DateHistogram gap bug #2183(@PSeitz)
- Fix range query end check (fields with less than 1 value per doc are affected) #2226(@PSeitz)
- Handle exclusive out of bounds ranges on fastfield range queries #2174(@PSeitz)
Breaking API Changes
- rename ReloadPolicy onCommit to onCommitWithDelay #2235(@giovannicuccu)
- Move exports from the root into modules #2220(@PSeitz)
- Accept field name instead of
Field
in FilterCollector #2196(@PSeitz) - remove deprecated IntOptions and DateTime #2353(@PSeitz)
Features/Improvements
-
Tantivy documents as a trait: Index data directly without converting to tantivy types first #2071(@ChillFish8)
-
encode some part of posting list as -1 instead of direct values (smaller inverted indices) #2185(@trinity-1686a)
-
Aggregation
- Support to deserialize f64 from string #2311(@PSeitz)
- Add a top_hits aggregator #2198(@ditsuke)
- Support bool type in term aggregation #2318(@PSeitz)
- Support ip adresses in term aggregation #2319(@PSeitz)
- Support date type in term aggregation #2172(@PSeitz)
- Support escaped dot when addressing field #2250(@PSeitz)
-
Add ExistsQuery to check documents that have a value #2160(@imotov)
-
Memory/Performance
- Faster TopN: replace BinaryHeap with TopNComputer #2186(@PSeitz)
- reduce number of allocations during indexing #2257(@PSeitz)
- Less Memory while indexing: docid deltas while indexing #2249(@PSeitz)
- Faster indexing: use term hashmap in fastfield #2243(@PSeitz)
- term hashmap remove copy in is_empty, unused unordered_id #2229(@PSeitz)
- add method to fetch block of first values in columnar #2330(@PSeitz)
- Faster aggregations: add fast path for full columns in fetch_block #2328(@PSeitz)
- Faster sstable loading: use fst for sstable index #2268(@trinity-1686a)
-
QueryParser
- allow newline where we allow space in query parser #2302(@trinity-1686a)
- allow some mixing of occur and bool in strict query parser #2323(@trinity-1686a)
- handle * inside term in lenient query parser #2228(@trinity-1686a)
- add support for exists query syntax in query parser #2170(@trinity-1686a)
-
report if a term matched when warming up posting list #2309(@trinity-1686a)
-
Support json fields in FuzzyTermQuery #2173(@PingXia-at)
-
Read list of fields encoded in term dictionary for JSON fields #2184(@PSeitz)
-
Forward regex parser errors #2288(@adamreichold)
-
Make FacetCounts defaultable and cloneable. #2322(@adamreichold)
New Contributors
- @imotov made their first contribution in #2160
- @PingXia-at made their first contribution in #2173
- @giovannicuccu made their first contribution in #2235
- @BlackHoleFox made their first contribution in #2265
- @ditsuke made their first contribution in #2282
- @MochiXu made their first contribution in #2312
Full Changelog: 0.21...v0.22.0
Tantivy v0.21
Bugfixes
- Fix track fast field memory consumption, which led to higher memory consumption than the budget allowed during indexing #2148#2147(@PSeitz)
- Fix a regression from 0.20 where sort index by date wasn't working anymore #2124(@PSeitz)
- Fix getting the root facet on the
FacetCollector
. #2086(@adamreichold) - Align numerical type priority order of columnar and query. #2088(@fmassot)
Breaking Changes
- Remove support for Brotli and Snappy compression #2123(@adamreichold)
Features/Improvements
- Implement lenient query parser #2129(@trinity-1686a)
- order_by_u64_field and order_by_fast_field allow sorting in ascending and descending order #2111(@naveenann)
- Allow dynamic filters in text analyzer builder #2110(@fulmicoton @fmassot)
- Aggregation
- Add missing parameter for term aggregation #2149#2103(@PSeitz)
- Add missing parameter for percentiles #2157(@PSeitz)
- Add missing parameter for stats,min,max,count,sum,avg #2151(@PSeitz)
- Improve aggregation deserialization error message #2150(@PSeitz)
- Add validation for type Bytes to term_agg #2077(@PSeitz)
- Alternative mixed field collection #2135(@PSeitz)
- Add missing query_terms impl for TermSetQuery. #2120(@adamreichold)
- Minor improvements to OwnedBytes #2134(@adamreichold)
- Remove allocations in split compound words #2080(@PSeitz)
- Ngram tokenizer now returns an error with invalid arguments #2102(@fmassot)
- Make TextAnalyzerBuilder public #2097(@adamreichold)
- Return an error when tokenizer is not found while indexing #2093(@naveenann)
- Delayed column opening during merge #2132(@PSeitz)
New Contributors
- @naveenann made their first contribution in #2093
- @cjrh made their first contribution in #2144
- @GodTamIt made their first contribution in #2145
- @ethever made their first contribution in #2165
Full Changelog: 0.20.2...0.21
0.20.1
Tantivy v0.20
What's Changed
Bugfixes
- Fix phrase queries with slop (slop supports now transpositions, algorithm that carries slop so far for num terms > 2) #2031#2020(@PSeitz)
- Handle error for exists on MMapDirectory #1988 (@PSeitz)
- Aggregation
Features/Improvements
- Add PhrasePrefixQuery #1842 (@trinity-1686a)
- Add
coerce
option for text and numbers types (convert the value instead of returning an error during indexing) #1904 (@PSeitz) - Add regex tokenizer #1759(@mkleen)
- Move tokenizer API to seperate crate. Having a seperate crate with a stable API will allow us to use tokenizers with different tantivy versions. #1767 (@PSeitz)
- Columnar crate: New fast field handling (@fulmicoton @PSeitz) #1806#1809
- Support for fast fields with optional values. Previously tantivy supported only single-valued and multi-value fast fields. The encoding of optional fast fields is now very compact.
- Fast field Support for JSON (schemaless fast fields). Support multiple types on the same column. #1876 (@fulmicoton)
- Unified access for fast fields over different cardinalities.
- Unified storage for typed and untyped fields.
- Move fastfield codecs into columnar. #1782 (@fulmicoton)
- Sparse dense index for optional values #1716 (@PSeitz)
- Switch to nanosecond precision in DateTime fastfield #2016 (@PSeitz)
- Aggregation
- Add
date_histogram
aggregation (onlyfixed_interval
for now) #1900 (@PSeitz) - Add
percentiles
aggregations #1984 (@PSeitz) - [breaking] Drop JSON support on intermediate agg result (we use postcard as format in
quickwit
to send intermediate results) #1992 (@PSeitz) - Set memory limit in bytes for aggregations after which they abort (Previously there was only the bucket limit) #1942#1957(@PSeitz)
- Add support for u64,i64,f64 fields in term aggregation #1883 (@PSeitz)
- Add count, min, max, and sum aggregations #1794 (@guilload)
- Switch to Aggregation without serde_untagged => better deserialization errors. #2003 (@PSeitz)
- Switch to ms in histogram for date type (ES compatibility) #2045 (@PSeitz)
- Reduce term aggregation memory consumption #2013 (@PSeitz)
- Reduce agg memory consumption: Replace generic aggregation collector (which has a high memory requirement per instance) in aggregation tree with optimized versions behind a trait.
- Split term collection count and sub_agg (Faster term agg with less memory consumption for cases without sub-aggs) #1921 (@PSeitz)
- Schemaless aggregations: In combination with stacker tantivy supports now schemaless aggregations via the JSON type.
- Perf: Fetch blocks of vals in aggregation for all cardinality #1950 (@PSeitz)
- Add
Searcher
with disabled scoring viaEnableScoring::Disabled
#1780 (@shikhar)- Enable tokenizer on json fields #2053 (@PSeitz)
- Enforcing "NOT" and "-" queries consistency in UserInputAst #1609 (@denis Bazhenov)
- Faster indexing
- Faster search
- Make BM25 scoring more flexible #1855 (@alexcole)
- Switch fs2 to fs4 as it is now unmaintained and does not support illumos #1944 (@Toasterson)
- Made BooleanWeight and BoostWeight public #1991 (@fulmicoton)
- Make index compatible with virtual drives on Windows #1843 (@yukun Guo)
- Auto downgrade index record option, instead of vint error #1857 (@PSeitz)
- Enable range query on fast field for u64 compatible types #1762 (@PSeitz) [#1876]
- sstable
- Isolating sstable and stacker in independant crates. #1718 (@fulmicoton)
- New sstable format #1943#1953 (@trinity-1686a)
- Use DeltaReader directly to implement Dictionnary::ord_to_term #1928 (@trinity-1686a)
- Use DeltaReader directly to implement Dictionnary::term_ord #1925 (@trinity-1686a)
- Add seperate tokenizer manager for fast fields #2019 (@PSeitz)
- Make construction of LevenshteinAutomatonBuilder for FuzzyTermQuery instances lazy. #1756 (@adamreichold)
- Added support for madvise when opening an mmaped Index #2036 (@fulmicoton)
- Rename
DatePrecision
toDateTimePrecision
#2051 (@guilload) - Query Parser
- Quotation mark can now be used for phrase queries. #2050 (@fulmicoton)
- PhrasePrefixQuery is supported in the query parser via:
field:"phrase ter"*
#2044 (@adamreichold)
- Docs
New Contributors
- @mhlakhani made their first contribution in #1733
- @pinkforest made their first contribution in #1746
- @DawChihLiou made their first contribution in #1737
- @mkleen made their first contribution in #1759
- @lonre made their first contribution in #1803
- @gyk made their first contribution in #1843
- @alexcole made their first contribution in #1855
- @Toasterson made their first contribution in #1944
- @vsop-479 made their first contribution in #1970
- @Tony-X made their first contribution in #1985
- @RTEnzyme made their first contribution in #1999
- @tottoto made their first contribution in #2018
- @nyurik made their first contribution in #2038
- @bazhenov made their first contribution in #1609
- @lavrd made their first contribution in #1422
- @tnxbutno made their first contribu...
Tantivy v0.19.2
Fixes an issue in the skip list deserialization, which deserialized the byte start offset incorrectly as u32.
get_doc
will fail for any docs that live in a block with start offset larger than u32::MAX (~4GB).
Causes index corruption, if a segment with a doc store file larger 4GB is merged. (@PSeitz)
Tantivy v0.19.1
Hotfix on handling user input for get_docid_for_value_range (@PSeitz)
Tantivy v0.19
What's Changed
Bugfixes
- Fix missing fieldnorms for u64, i64, f64, bool, bytes and date #1620 (@PSeitz)
- Fix interpolation overflow in linear interpolation fastfield codec #1480 (@PSeitz @fulmicoton)
Features/Improvements
- Add support for
IN
in queryparser , e.g.field: IN [val1 val2 val3]
#1683 (@trinity-1686a) - Skip score calculation, when no scoring is required #1646 (@PSeitz)
- Limit fast fields to u32 (
get_val(u32)
) #1644 (@PSeitz) - The
DateTime
type has been updated to hold timestamps with microseconds precision.
DateOptions
andDatePrecision
have been added to configure Date fields. The precision is used to hint on fast values compression. Otherwise, seconds precision is used everywhere else (i.e terms, indexing) #1396 (@evanxg852000) - Add IP address field type #1553 (@PSeitz)
- Add boolean field type #1382 (@boraarslan)
- Remove Searcher pool and make
Searcher
cloneable. (@PSeitz) - Validate settings on create #1570 (@PSeitz)
- Detect and apply gcd on fastfield codecs #1418 (@PSeitz)
- Doc store
- Make
tantivy::TantivyError
cloneable #1402 (@PSeitz) - Add support for phrase slop in query language #1393 (@saroh)
- Aggregation
- Faster indexing
New Contributors
- @ryanrussell made their first contribution in #1380
- @boraarslan made their first contribution in #1382
- @pier-oliviert made their first contribution in #1415
- @kianmeng made their first contribution in #1445
- @akr4 made their first contribution in #1473
- @waywardmonkeys made their first contribution in #1524
- @nigel-andrews made their first contribution in #1608
- @theduke made their first contribution in #1624
Full Changelog: 0.18...0.19
Tantivy v0.18.1
- Hotfix on positions. #1629 (@fmassot, @fulmicoton, @PSeitz)
Tantivy v0.18
Tantivy v0.17
- LogMergePolicy now triggers merges if the ratio of deleted documents reaches a threshold (@shikhar @fulmicoton) #115
- Adds a searcher Warmer API (@shikhar @fulmicoton)
- Change to non-strict schema. Ignore fields in data which are not defined in schema. Previously this returned an error. #1211
- Facets are necessarily indexed. Existing index with indexed facets should work out of the box. Index without facets that are marked with index: false should be broken (but they were already broken in a sense). (@fulmicoton) #1195 .
- Bugfix that could in theory impact durability in theory on some filesystems #1224
- Schema now offers not indexing fieldnorms (@lpouget) #922
- Reduce the number of fsync calls #1225
- Fix opening bytes index with dynamic codec (@PSeitz) #1278
- Added an aggregation collector compatible with Elasticsearch (@PSeitz)
- Added a JSON schema type @fulmicoton #1251
- Added support for slop in phrase queries @halvorboe #1068