Community post
Improving Hive’s semantic search performance

HiveSense is a HAF app that creates semantic embeddings for Hive posts. These embeddings will be used by the new HiveSense API calls to find posts that are similar to each other and to find posts that match a user’s search request.
A while back, one of our devs posted on the early work we did developing a semantic search API server for hive posts called HiveSense.
Today’s post describes the follow-on work in the past 2.5 months as we prepare for an official release of HiveSense as part of the standard HAF API stack, so this post assumes you’ve previously read the original post about HiveSense.
This post is even more technical than the previous one, so the primary audience is devs interested in practical considerations associated with using semantic search in their apps or who would like to contribute to HiveSense development in the future.
Performance Optimizations
- Use Ollama’s batch API to generate multiple embeddings in a single call
- Add support for, and default to, using 16 bit precision floating point vectors. This reduced the 768-dimension embeddings table size from 100GB to 50GB.
- Use PySBD for sentence detection instead of spacy (smaller)
- Dramatically reduced size of docker images (many changes here, but biggest win was removal of pgai from default config)
- Allow HiveSense to process data concurrently while hivemind is still in massive sync (originally Hivesense could only be synced after hivemind was in livesync, and hivemind takes about 2.5 days to sync on a very fast system). Since HiveSense syncs faster than hivemind, and can now sync in parallel with hivemind except for the time to create the HNSW index, HiveSense now only adds 1.5 hours to overall time to sync up a HAF API node from scratch.
- Redesigned worker thread implementation
Recall Optimizations (Recall here basically is referring to the “quality” of the search results)
- Chunk long posts on sentence boundaries to capture semantic meaning better in the chunks
- Chunk all of a post instead of limiting a post to a max of 3 chunks. This creates more embeddings, but allows for better matching of long posts, especially if they switch topics.
- Target chunk size based on number of tokens rather than raw character count
- Prepend the title to the post so the title will be used when calculating the embedding of the post’s first chunk
- Change the minimum word count for posts to a minimum token count, make the token count configurable, and default the minimum to 75 tokens
- Add many options to ease switching to better embedding models in the future and to examine tradeoffs between models
- Use query prefix to improve embeddings generated for search queries
- Don’t discard non-ASCII characters (these were being removed for embedding calculations)
- Improve filtering of HTML
- Increased m and ef_construction to improve recall quality, especially for “small queries” that sometimes got poor search results.
Miscellaneous changes
- Track the number of tokens in each post and order the embeddings generated for each post
- Allow filtering out shorter posts at search time (previously this could only be done at index time)
- Change API results to make paging deterministic (in progress)
Redesign of embeddings tables and indexing methodology
Originally we just used a single table with 768-dimension vectors and built a HNSW index on this table. Both the table and the index were originally 100GB each in size (e.g. 200GB total storage required by HiveSense). Our first optimization was to use 16-bit precision numbers instead of 32-bit to cut storage requirements in half (100GB in total size, which seemed like a reasonable amount of storage).
But another problem we found was that it was very time consuming to create an HNSW index this large. On systems without a LOT of memory, it would take quite a few days. On systems with 128GB of RAM installed, this time could be cut down to around 8.5 hours (the current code that computes this index really favors having sufficient memory for the index creation), but this seemed a steep requirement for most API servers (we have internal servers that have this much memory, but the cloud servers we rent only have 64GB).
The solution we arrived at was to create secondary smaller embeddings table and create an HNSW index on that smaller table, then use the larger embeddings table for the final similarity computation.
HiveSense uses principal component analysis (PCA) to generate a second embeddings table with much smaller 128-dimension vectors (table size 9GB) and a much smaller HNSW index on this table (the new HNSW index is only 16GB). This new index only takes 1.5 hours to build and only requires 28GB of memory (can be build in 4.5 hours on systems with less memory).
Storage-wise, with all the optimizations, we reduced total storage usage from 100+100=200GB down to 50+9+16=75GB.
This approach also dramatically speeds up API query time as we’re searching a much smaller index, but we don’t have full statistics for this yet (our guess is somewhere between 3x and 10x faster).
Of course, we did have one concern about this approach: we needed to ensure it didn’t negatively effect recall results. To ensure this, we compared search results for various queries between a full brute force search of the embeddings and a search using the new index to ensure the results didn’t significantly change.
New Sync Mode for HiveSense
A normal CPU is sufficient to generate embeddings for short text phrases like those used for search queries, but generating semantic embeddings for posts is too computationally intensive, so a GPU is required to generate them at a reasonable speed.
We didn’t want to force API node operators to have a GPU, so HiveSense can be configured to operate in two different modes: independent mode and sync mode.
In the independent mode, HiveSense expects to have access to one or more Ollama servers with GPUs providing computation power.
In sync mode, the embeddings for posts are fetched from another HiveSense server, so the local HiveSense server only needs to compute embeddings for user search queries (which can computed with an Ollama server powered just by a reasonable CPU).
As we don’t expect most current API nodes to have access to a GPU (our primary API node, api.hive.blog doesn’t), we expect most API node operators will configure HiveSense to operate in sync mode, sparing their server from repeating the expensive computations required for computing post embeddings.
What’s next for HiveSense?
We need to change the API to stabilize paging of search results based on our new approach: we will return 1000 results, with the first 20 results including permlink + summary results for the post, and the remaining results just providing permlinks. Client side apps will need to fetch further post summaries in case the user pages beyond the first page.
We need to update our app testing API server, api.syncad.com, with the new stack so that Hive apps can add support for HiveSense and perform “real-world” testing.
Finally, we need to officially release HiveSense along the other updated HAF apps. Currently I expect that to happen near the end of this quarter (sometime in September).
Replies (31)
that is a real important update. Thanks aöpt for workong on it!
Great news. Hive search can really use these improvements for content discoverability.
不明觉厉👍
$WINE
very technical information
What would we do without you? It's good that you exist! Well done!
This like a kind of encouraging for hive communities
This is a very important update. Thank you very much for the hard work on this.
I can't wait for September,
Thank you for taking this initiative
This highly creative
Excellent, thorough explanation that provides a glimpse into how to navigate successfully.
From what you Just said Hive is done 😂😂😂😂😂😂🎯
😂😂😂😂😂
Blocktrades can you please stop downvoting my original content with your alt accounts thanks 🙏🏾
😂😂😂😂😂😂😂😂😂😂😂😂😂
Don't be jealous So sad @themarkymark aka @theycallmemarky Aka @marky @gogreenbuddy Aka @usainvote @buildawhale Aka @punkteam Aka @ipromote
@letusbuyhive Aka @sagarkothari88
please explain why you keep downvoting my original content and comments 🤔
Please explain to everyone why 🤔
@meno @steevc @crimsonclad @azircon let's not forget blocktrades 😂 please tell your racist friend to stop downvoting my original content
How can we we allow a mentally ill person to have so much control on Hive 😂
It's unbelievable you downvoted this Goodbye Auntie R.I.P 🙏🏾 You fucked up big time it's clear blocktrades is in control 😂
If nothing can be done about your downvote abuse then Hive is a dead project 👎🏾
It's crazy that the person doing the downvoting is also farming the shit out of Hive 😂😂😂😂
If anyone wants to speak to me privately send me a message on Instagram @kgakakillerg
You can never stop the truth with lies 😂😂😂😂😂😂
It's crazy we have a man who pretends to be a millionaire running around Hive downvoting people for fun 🤔 and nothing has been done 🤔
It's crazy that someone is trying to bully me on Hive 😂😂😂😂😂😂😂😂
It's unbelievable that no one has came to help as people are afraid to say anything or they get downvoted too 😂😂😂😂😂😂
There's no real freedom on Hive if one person can cause so much harm
I feel sorry for you 🙏🏾
One thing you need to remember you can't take anything with you 😂😂😂😂
Go and enjoy your life have a bit of fun let your hair down go out and meet some real people in the flesh 😉
Your actions show that you must own Hive Blockchain 🤔 why do your actions go unchallenged 🤔
You are now stalking me 😂😂😂😂
Are you jealous about the DHF 😂😂😂😂
Aren't you farming enough rewards 😂😂😂😂
I heard you were into very young girls 😂😂😂😂😂😂
Is that why you are always online 🤔
Stalking people
No one is scared of you 😂😂😂😂😂😂😂
You are bad for Hive
Power down and go away get a life
https://hive.blog/hive-135178/@crimsonclad/re-kgakakillerg-sxllhv
https://hive.blog/hive-148441/@hivewatchers/svftu9
https://hive.blog/hive-148441/@hivewatchers/svdjjz
https://hive.blog/hive-176853/@steevc/re-kgakakillerg-syyy4x
https://hive.blog/dev/@howo/re-kgakakillerg-szhax7
https://hive.blog/hive/@steevc/follow-friday-respect
https://hive.blog/hive-127022/@shmoogleosukami/re-kgakakillerg-t0hcxc
It's unbelievable that they downvoted this Goodbye Auntie R.I.P 🙏🏾
It's clear you need help 🙏🏾
Hive is being held down by downvoting whales
Do you get a buzz out of downvoting people 🤔
You should really try to get outside more spend some time in the real world 🌎🌍
Why are you pushing people away to blurt and steemit
You are all definitely going to hell 😂😂😂😂😂😂😂😂
Why do you want to make enemies all over the world 🤔
Blocktrades stop making a fool of yourself 😂😂😂😂😂😂
Steevc please explain to your friend they have been exposed 😂😂😂😂😂
Downvotes are weak like you 😂😂😂😂😂😂😂
Why don't you go and spend your millions of dollar's you have 😂😂😂😂😂😂😂😂😂😂😂😂
Blocktrades please explain why you keep downvoting my original content with your alt accounts 🤔
You are so sad it's unreal 😂😂😂😂😂
You must know that you can't hide on Hive 😂😂😂😂😂😂😂
If you want everyone to leave Hive keep doing what you are doing 😂😂😂😂😂😂😂😂😂😂
Just remembered who started this
I'll be here to turn the lights off 😂😂😂😂😂😂
You are still stalking me 😂😂😂😂😂 it proves you have no life outside of Hive 😂😂😂😂😂😂😂👎🏾👎🏾👎🏾👎🏾👎🏾👎🏾👎🏾😂😂😂😂😂
Blocktrades please can you stop downvoting my original content with your alt accounts 🤔
Also is bullying people ok on Hive 🤔
You are only making Hive look like a big scam
Why do you keep stalking me I'm not gay sorry I can't help you 🙏🏾
Just tell me what the issue is 🤔
You are that stupid you set up an account called letusbuyhive to downvote people and support your farming Hive friend's 😂😂😂😂😂 and @buildawhale
Please get some help 🙏🏾
It's clear who has mental health issues that's why you should really stop pointing fingers at others 😂😂😂😂😂😂😂😂😂
It's clear blocktrades is behind this 😂😂😂😂😂
Can you please explain why you keep downvoting my original content I don't want to hear it's because disagreement of rewards I don't make any 😂😂😂😂😂
Please move on with your life blocktrades and leave me alone thank you 🙏🏾
Please tell everyone why you keep downvoting my original content 🤔
Still stalking me 😂😂😂😂😂😂
https://hive.blog/hive-127466/@steevc/re-blocktrades-t0kint
https://hive.blog/hive-127466/@blocktrades/t0lq41
The way you are stalking me it's clear you have no Life outside of Hive 😂😂😂😂😂😂😂
Keep downvoting people away 😂😂😂😂😂
Some people never learn 😂😂😂😂😂😂😂🤣🤣🤣🤣🤣
Is life that bad 😂😂😂😂😂
You can't win this 😂😂😂😂😂😂😂😂😂
Why is blocktrades.us no longer live?
Details here: https://hive.blog/blocktrades/@blocktrades/blocktrades-ending-its-cryptocurrency-trading-service-as-of-june-30th-2023-today
@blocktrades excellent information
Why do you keep downvoting my original content and comments
What's the problem 🤔
https://hive.blog/hive-148441/@kgakakillerg/t0yxxn
You farm enough Hive with your alt accounts 😂😂
No one is scared of you 😂 you are destroying the thing you say you care about
I know steevc is your friend and he has something to do with this
What you must understand no one is scared of you @blocktrades Aka @themarkymark @buildawhale @usainvote
Please explain what the problem is instead of downvoting my original content
Do you realize how stupid it looks I can understand you have no kid's so this is your life
Why not get a life Instead of trying to treat Hive like it's your kid or something that you can control fully
You are making have centralised by yourself 😂
Can we speak to each other like adults please stop acting like a child you are a old man please act your age 🙏🏾
Stop defending something that you know is wrong
Do you not ever ask yourself why Hive hasn't took off well just look at what you are doing and those close to you
It doesn't look good does it 🤔
💯 Original Content downvoted for no reason Please read and view
@steevc please tell your friends to stop downvoting my original content please 🙏🏾
@themarkymark @buildawhale @letsusbuyhive please stop downvoting my original content
@crimsonclad please do your job 🙏🏾
https://hive.blog/hive-135178/@crimsonclad/re-kgakakillerg-sxllhv
https://hive.blog/hive-148441/@hivewatchers/svftu9
https://hive.blog/hive-148441/@hivewatchers/svdjjz
https://hive.blog/hive-176853/@steevc/re-kgakakillerg-syyy4x
https://hive.blog/dev/@howo/re-kgakakillerg-szhax7
https://hive.blog/hive/@steevc/follow-friday-respect
https://hive.blog/hive-127022/@shmoogleosukami/re-kgakakillerg-t0hcxc
It's unbelievable that they downvoted this Goodbye Auntie R.I.P 🙏🏾
Comments being downvoted by blocktrades https://hive.blog/hive-170744/@kgakakillerg/t0ns3b
https://hive.blog/hive-127466/@steevc/re-blocktrades-t0kint
https://hive.blog/hive-127466/@blocktrades/t0lq41
https://hive.blog/hive/@ureka.stats/the-untrending-report-hive-downvote-analysis-2025-06-29-20250629143829
https://hive.blog/hive-127466/@kgakakillerg/t0m1vn
https://hive.blog/hive-108278/@kgakakillerg/t0rfo8
https://hive.blog/hive-127466/@kgakakillerg/t0vcl7
Congratulations @blocktrades! You have completed the following achievement on the Hive blockchain And have been rewarded with New badge(s)
Your next target is to reach 210000 upvotes.
You can view your badges on your board and compare yourself to others in the Ranking
If you no longer want to receive notifications, reply to this comment with the word
STOPIndeed the hivesense project is a welcome development. I can't wait to see more and more good developments on Hive.
Very interesting updates. It’s really cool to see all the changes and improvements.
Great news! Indeed, its a nonstop innovation for more efficient discovery!
This is not ok RE: Morning Run: Child's Play
Please stop downvoting my original content and comments with your alt accounts themarkymark and buildawhale it's clear you are behind them accounts otherwise you would have been downvoting them already for farming Hive with the buildawhale comment farm why do you delegate Hive power to @buildawhale and @usainvote 🤔 and so much Hive going to your other account @alpha 🤔
Please leave me alone and stop exposing yourself and others 🙏🏾
Do you not understand how stupid it looks @blocktrades https://hive.blog/hive/@ureka.stats/the-untrending-report-hive-downvote-analysis-2025-09-07-20250907013710
Please stop trying to bully me with downvotes it won't work just leave me alone 🙏🏾
No one is scared of downvotes @blocktrades please leave me alone 🙏🏾
RE: Improving Hive’s semantic search performance
Please leave me alone https://hive.blog/hive/@ureka.stats/the-untrending-report-hive-downvote-analysis-16-09-2025-20250916181314
What you are doing is classed as cyber bullying
Very important information.
Interesting information, it will surely be very useful for the ecosystem.
Great update! I really appreciate how you keep improving Hive's performance and making things run smoother for everyone. The changes in the embeddings tables and indexing may be technical, but they show how much dedication goes into keeping Hive fast and reliable. Thanks for all your hard work!
https://peakd.com/hive-124838/@meno/re-ackza-t43f2y
“Stock exchanges India, Hong Kong, and Australia Feel the Impact of the US–China Trade War”
Hey — we noticed you're holding VKBT, and wanted to invite you personally. We just opened a bridge so it can cross to our chain and actually move. There's finally somewhere for it to go: https://peakd.com/@angelicalist/hivers-guide-to-the-melek-prana-kula-ecosystem-20260830 Glad you held. 🙏