Speech order method and device based on LLM and fusion recall strategy and medium

CN122842579APending Publication Date: 2026-09-29BEIJING FENYANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611185485.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这种方法难以处理口语化、简略化、带有修饰词(如“上次订的那个饮料”“卖得最好的牛奶”)的查询,也无法进行语义层面的相关性排序

Benefits of technology

[0018]本发明的技术效果在于:本发明针对现有技术中的智能下单存在的语音识别与业务理解脱节、缺乏面向行业的深度语义召回能力、未能融入动态业务策略进行智能排序与决策,导致语音下单的准确性、便捷性和业务贴合度不足的技术缺陷,提出了基于大语言模型(LLM)与融合多种召回策略的AI语音下单方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842579A_ABST
    Figure CN122842579A_ABST
Patent Text Reader

Abstract

This invention proposes a voice ordering method, device, and medium based on LLM and a fusion recall strategy, relating to the field of artificial intelligence technology. The method includes: receiving a voice ordering command input by a user through a terminal and sending the command to a cloud server; the cloud server transcribes and recognizes the voice ordering command to obtain the order text, inputs the order text into an LLM deployed on the cloud server, and the LLM performs thought chain reasoning and structuring on the order text to generate a structured query object; using the structured query object, M-way recall requests are initiated in parallel, and the results are fused to obtain a Top-N preliminary candidate set; based on the preliminary candidate set, multi-objective optimization is performed to obtain a Top-K candidate set, which is then sent to the terminal; and a product order is generated based on the user's selection of the Top-K candidate set. This invention solves the technical problem of fuzzy commands failing to match precisely.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a voice ordering method, apparatus, and medium based on LLM and a fusion recall strategy. Background Technology

[0002] Existing technologies that utilize voice ordering primarily include: General-purpose speech recognition (ASR) and natural language understanding (NLP) services can convert speech into text and perform basic intent recognition and entity extraction. However, these are general-domain models and are not well adapted to the product databases, professional terminology, and expression habits specific to the FMCG industry, thus failing to directly guarantee accuracy in order placement scenarios.

[0003] Keyword-based product search: After a text query, products are retrieved through fuzzy matching or inverted indexes in the database. This method struggles to handle colloquial, abbreviated queries, queries with modifiers (such as "the drink I ordered last time" or "the best-selling milk"), and it also cannot perform semantic relevance ranking.

[0004] Simple voice ordering plugins: Some ERP or mobile sales tools have integrated voice input boxes, but they only use voice as a substitute for text input. Subsequent processes still require manual completion of product selection, verification, and confirmation, resulting in low levels of intelligence.

[0005] It is evident that the main shortcomings of existing technologies are: the disconnect between speech recognition and business understanding, the lack of deep semantic recall capabilities tailored to specific industries, and the failure to integrate dynamic business strategies for intelligent sorting and decision-making, resulting in insufficient accuracy, convenience, and business fit in voice ordering. Summary of the Invention

[0006] In view of one or more technical defects in the prior art, the present invention proposes the following technical solution.

[0007] A voice ordering method based on LLM and fusion recall strategy, the method includes: The receiving step involves receiving a voice order instruction input by the user through the terminal and sending the voice order instruction to the cloud server. In the reasoning step, the cloud server transcribes and recognizes the voice order command to obtain the order text, inputs the order text into the LLM deployed on the cloud server, and the LLM performs thought chain reasoning and structured processing on the order text to generate a structured query object. The fusion step involves using structured query objects to initiate M-way recall requests in parallel, and then fusing the results of the M-way recall requests to obtain a preliminary Top-N candidate set, where M≥2 and N≥2. The reordering step involves performing multi-objective optimization based on the initial Top-N candidate set to obtain a Top-K candidate set, where K≥2, and then sending the Top-K candidate set to the terminal. The generation step involves displaying the Top-K candidate set on the terminal and generating a product order based on the user's selection of the Top-K candidate set.

[0008] Furthermore, the method also includes a fine-tuning step, which records the user's click, modification, confirmation, and rejection operations on the Top-K candidate set displayed on the terminal, forming "voice order instruction - final ordered product" pairs as positive samples, adding them to the training dataset, and periodically fine-tuning the LLM.

[0009] Furthermore, the M-path recall request includes: vector semantic recall request: encoding the query description into a high-dimensional embedding vector, performing an approximate nearest neighbor search in the vector database, and recalling semantically similar products; inverted index precise recall request: performing precise matching on specific brand names and SKU codes to accurately recall corresponding products; and graph association recall: querying the current store's "frequently purchased list" and "related best-selling products" based on a knowledge graph to recall related products.

[0010] Furthermore, the operation of fusing the results of the M-way recall requests is as follows: calculate the fusion score of the structured query object, which is calculated as Score_final(item)=α×Norm(Vector_Score)+β×Norm(BM25_Score)+γ×Norm(Graph_Weight); Wherein, Score_final(item) represents the fusion score of the structured query object, Norm(.) represents the normalization function, Vector_Score represents the vector semantic recall request result, BM25_Score represents the inverted index precise recall request result, Graph_Weight represents the graph association recall result, and α, β, and γ represent the corresponding fusion weight coefficients, which are dynamically adjusted according to the recall type, which is either fuzzy or precise.

[0011] Furthermore, the re-ranking step involves: calculating the score of each product in the initial Top-N candidate set according to a multi-objective optimization strategy, and re-ranking the products based on their scores to obtain the Top-K candidate set. The score is calculated as follows: Score = w1 × Semantic_Sim + w2 × User_Preference + w3 × Business_Boost, where Semantic_Sim represents the "semantic matching score" of the candidate products calculated in real time, User_Preference represents the "user's historical repurchase rate", and Business_Boost is calculated based on the real-time inventory status and gross profit contribution. Business_Boost = Base_Score + I(is_promotion)w_promo + I(high_stock)w_stock, where I(high_stock)w_stock represents the high inventory status score, I(is_promotion)w_promo represents the gross profit contribution score, and Base_Score represents the base score.

[0012] This invention also proposes a voice ordering device based on LLM and a fusion recall strategy, the device comprising: The receiving unit receives the voice order instruction input by the user through the terminal and sends the voice order instruction to the cloud server; The reasoning unit, the cloud server, transcribes and recognizes the voice order command to obtain the order text, inputs the order text into the LLM deployed on the cloud server, and the LLM performs thought chain reasoning and structured processing on the order text to generate a structured query object; The fusion unit initiates M-way recall requests in parallel using structured query objects, and performs fusion processing on the results of the M-way recall requests to obtain a preliminary Top-N candidate set, where M≥2 and N≥2. The reordering unit performs multi-objective optimization based on the initial Top-N candidate set to obtain a Top-K candidate set, where K≥2, and sends the Top-K candidate set to the terminal. The generation unit displays the Top-K candidate set on the terminal and generates a product order based on the user's selection of the Top-K candidate set.

[0013] Furthermore, the device also includes a fine-tuning unit that records the user's click, modification, confirmation, and rejection operations on the Top-K candidate set displayed on the terminal, forming "voice order command - final ordered product" pairs as positive samples, adding them to the training dataset, and periodically fine-tuning the LLM.

[0014] Furthermore, the M-path recall request includes: vector semantic recall request: encoding the query description into a high-dimensional embedding vector, performing an approximate nearest neighbor search in the vector database, and recalling semantically similar products; inverted index precise recall request: performing precise matching on specific brand names and SKU codes to accurately recall corresponding products; and graph association recall: querying the current store's "frequently purchased list" and "related best-selling products" based on a knowledge graph to recall related products.

[0015] Furthermore, the operation of fusing the results of the M-way recall requests is as follows: calculate the fusion score of the structured query object, which is calculated as Score_final(item)=α×Norm(Vector_Score)+β×Norm(BM25_Score)+γ×Norm(Graph_Weight); Wherein, Score_final(item) represents the fusion score of the structured query object, Norm(.) represents the normalization function, Vector_Score represents the vector semantic recall request result, BM25_Score represents the inverted index precise recall request result, Graph_Weight represents the graph association recall result, and α, β, and γ represent the corresponding fusion weight coefficients, which are dynamically adjusted according to the recall type, which is either fuzzy or precise.

[0016] Furthermore, the re-ranking step involves: calculating the score of each product in the initial Top-N candidate set according to a multi-objective optimization strategy, and re-ranking the products based on their scores to obtain the Top-K candidate set. The score is calculated as follows: Score = w1 × Semantic_Sim + w2 × User_Preference + w3 × Business_Boost, where Semantic_Sim represents the "semantic matching score" of the candidate products calculated in real time, User_Preference represents the "user's historical repurchase rate", and Business_Boost is calculated based on the real-time inventory status and gross profit contribution. Business_Boost = Base_Score + I(is_promotion)w_promo + I(high_stock)w_stock, where I(high_stock)w_stock represents the high inventory status score, I(is_promotion)w_promo represents the gross profit contribution score, and Base_Score represents the base score.

[0017] The present invention also proposes a computer-readable storage medium storing computer program code, which, when executed by a computer, performs any of the methods described above.

[0018] The technical advantages of this invention are as follows: This invention addresses the technical shortcomings of existing intelligent ordering technologies, such as the disconnect between speech recognition and business understanding, the lack of deep semantic recall capabilities for specific industries, and the failure to integrate dynamic business strategies for intelligent sorting and decision-making, resulting in insufficient accuracy, convenience, and business fit in voice ordering. This invention proposes an AI voice ordering method based on a Large Language Model (LLM) and integrating multiple recall strategies.

[0019] This method employs a distributed architecture that collaborates across "end-edge-cloud," deeply integrating the semantic understanding capabilities of generative artificial intelligence with the efficient retrieval capabilities of search engines. It constructs a closed-loop processing system that transforms unstructured voice commands into structured business orders. The cloud server transcribes and recognizes the voice order command to obtain the order text. This text is then input into an LLM deployed on the cloud server. The LLM performs thought chain reasoning and structuring on the order text to generate a structured query object. The structured query object is used to initiate M parallel recall requests, and the results of these requests are fused to obtain a Top-N preliminary candidate set. Multi-objective optimization based on the Top-N preliminary candidate set yields a Top-K candidate set. The Top-K candidate set is displayed on the terminal, and a product order is generated based on the user's selection from the Top-K candidate set. This method solves the mapping problem from "non-standard commands" to "precise SKUs," addressing the limitations of existing keyword searches that cannot handle vague commands such as "like last time" or "the cheaper one." This invention achieves a seamless connection between human natural language and machine database language by leveraging the contextual reasoning capabilities of LLM and the fuzzy matching capabilities of vector retrieval. This is something that traditional regular expression matching or keyword retrieval techniques cannot achieve. Attached Figure Description

[0020] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.

[0021] Figure 1 This is a flowchart of a voice ordering method based on LLM and fusion recall strategy according to an embodiment of the present invention; Figure 2 This is a structural diagram of a voice ordering device based on an LLM and fusion recall strategy according to an embodiment of the present invention. Detailed Implementation

[0022] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] Figure 1 This invention illustrates a voice ordering method based on LLM and a fusion recall strategy, the method comprising: In step S101, the system receives a voice order command input by the user through the terminal and sends the command to the cloud server. In this invention, the terminal can initiate audio acquisition based on user touch operations. The terminal uses spectral subtraction to remove background steady-state noise and utilizes automatic gain control (AGC) to normalize the volume level. The audio data is framed and encoded in Opus format, and pushed to the cloud server gateway in real time via a WebSocket long connection. The gateway is responsible for authentication and flow control.

[0025] In reasoning step S102, the cloud server transcribes and recognizes the voice order command to obtain the order text, and inputs the order text into the LLM deployed on the cloud server. The LLM performs chain-of-thought reasoning and structuring on the order text to generate a structured query object. The cloud ASR service receives the audio stream and uses an end-to-end acoustic model based on the Conformer architecture for transcription. The system loads a hot word map customized for the FMCG industry (including unique brand abbreviations, specification terms, and promotional jargon), and uses a weighted finite-state converter (WFST) for decoding, significantly improving the accuracy of professional terminology recognition. The transcribed text is then input into the LLM. Context injection: The system automatically extracts the current user's profile features (historical preferences, region), current conversation history, and temporal and spatial states, constructs a System Prompt, and injects it into the LLM. Reasoning generation: The LLM uses chain-of-thought technology to parse complex logical instructions. For example, "Same as last time, but change A to B" is parsed as: query historical orders → copy order items → delete product A → add product B.

[0026] Output object: Generates a standardized structured query object, which includes a list of product entities, attribute modifiers (such as "full case" and "promotional items"), and filtering conditions.

[0027] The cloud server, loaded with a large-scale language model with 7B / 13B parameters, fine-tuned using a full corpus of FMCG (Fast Moving Consumer Goods) industry language, serves as the system's "brain," a core template capable of deeply understanding language meaning and making analytical decisions. It primarily includes Intent Recognition, Slot Filling, Inverse Text Normalization, and Prompt Engineering. Its output does not generate a single keyword, but rather a structured intermediate representation (IR) containing explicit entities (product names), implicit constraints (price ranges), and contextual referents (related stores).

[0028] In the fusion step S103, M-way recall requests are initiated in parallel using structured query objects, and the results of the M-way recall requests are fused to obtain a preliminary Top-N candidate set, where M≥2 and N≥2. In the reordering step S104, a Top-K candidate set is obtained by performing multi-objective optimization based on the Top-N preliminary candidate set, where K≥2, and the Top-K candidate set is sent to the terminal. In step S105, the Top-K candidate set is displayed on the terminal, and a product order is generated based on the user's selection of the Top-K candidate set.

[0029] This invention addresses the technical shortcomings of existing intelligent ordering technologies, such as the disconnect between speech recognition and business understanding, the lack of deep semantic recall capabilities tailored to specific industries, and the failure to integrate dynamic business strategies for intelligent sorting and decision-making, resulting in insufficient accuracy, convenience, and business relevance in voice ordering. To address these shortcomings, this invention proposes an AI voice ordering method based on a Large Language Model (LLM) and incorporating multiple recall strategies. This method employs a distributed architecture that collaborates across "end-edge-cloud," deeply integrating the semantic understanding capabilities of generative artificial intelligence with the efficient retrieval capabilities of search engines. It constructs a closed-loop processing system that transforms unstructured voice commands into structured business orders. The cloud server transcribes and recognizes the voice order command to obtain the order text. This text is then input into an LLM deployed on the cloud server. The LLM performs thought chain reasoning and structuring on the order text to generate a structured query object. The structured query object is used to initiate M parallel recall requests, and the results of these requests are fused to obtain a Top-N preliminary candidate set. Multi-objective optimization based on the Top-N preliminary candidate set yields a Top-K candidate set. The Top-K candidate set is displayed on the terminal, and a product order is generated based on the user's selection from the Top-K candidate set. This method solves the mapping problem from "non-standard commands" to "precise SKUs," addressing the limitations of existing keyword searches that cannot handle vague commands such as "like last time" or "the cheaper one." This invention achieves a seamless connection between human natural language and machine database language by leveraging the contextual reasoning capabilities of LLM and the fuzzy matching capabilities of vector retrieval. This is something that traditional regular expression matching or keyword retrieval techniques cannot achieve, and it is one of the key inventive concepts of this invention.

[0030] In one embodiment, the method further includes a fine-tuning step S106, which records the user's click, modification, confirmation, and rejection operations on the Top-K candidate set displayed on the terminal, forming "voice order instruction - final ordered product" pairs as positive samples, which are added to the training dataset. The LLM is then fine-tuned periodically based on this training dataset. In this invention, after adding the "voice instruction - final ordered product" pairs as positive samples to the training dataset, the RLHF (Reinforcement Learning Based on Human Feedback) process is periodically triggered to fine-tune the LLM and re-ranking model. This allows the system to become increasingly accurate in understanding a specific user's accent, expression habits, and business preferences as the number of uses increases. By incorporating a user behavior feedback-based RLHF mechanism, each use by the salesperson trains the system. Over time, the system's understanding of specific regional accents and specific store preferences will automatically converge to the optimal state, which is one of the key inventive concepts of this invention.

[0031] In one embodiment, the M-path recall request includes: a vector semantic recall request: encoding the query description into a high-dimensional embedding vector, performing an approximate nearest neighbor search in the vector database, and recalling semantically similar products; an inverted index precise recall request: performing precise matching on explicit brand names and SKU codes to accurately recall corresponding products; and a knowledge graph association recall request: querying the current store's "frequently purchased list" and "related best-selling products" based on a knowledge graph to recall related products.

[0032] This invention employs a composite storage and retrieval layer comprised of a vector database (Milvus), a high-performance inverted index (Elasticsearch), and a graph database (Neo4j). This layer is used for preprocessing product vectorization, and performs parallel three-way recall: semantic vector fuzzy matching, keyword precise matching and expansion, and knowledge graph-based recommendation. This resolves the contradiction between generalized spoken language comprehension and accurate SKU search, and supports initial reranking to ensure the accuracy of subsequent candidate set generation. This invention constructs a dynamic knowledge update mechanism based on RAG (Retrieval Augmentation Generation), solving the problem of long retraining cycles and inability to identify newly launched products in traditional models. This invention uses the product database as an external knowledge base (RAG architecture). New product listings only require updating the vector index, eliminating the need for model retraining. The LLM (Limited Language Model) can then identify new products through retrieval, exhibiting extremely high timeliness—another key inventive concept of this invention.

[0033] In one embodiment, the operation of fusing the results of M-way recall requests involves calculating the fusion score of the structured query object, using the following method: Score_final(item) = α × Norm(Vector_Score) + β × Norm(BM25_Score) + γ × Norm(Graph_Weight); where Score_final(item) represents the fusion score of the structured query object, Norm(.) represents the normalization function, Vector_Score represents the vector semantic recall request result, BM25_Score represents the inverted index exact recall request result, Graph_Weight represents the graph association recall result, and α, β, and γ represent the fusion weight coefficients of the three recall strategies, respectively. α, β, and γ are constrained to be non-negative real numbers and satisfy α + β + γ = 1. α, β, and / or γ are dynamically adjusted according to the recall type, which can be fuzzy or exact. For example, when the query contains a specific barcode, the β weight is automatically increased.

[0034] In this invention, three recall requests are initiated in parallel based on the Query Object. Of course, other recall requests can be set according to actual needs. Path A (Vector Semantic Recall): Encode the query description into a high-dimensional embedding vector, perform an approximate nearest neighbor (ANN) search in the vector database, and recall semantically similar products (solving ambiguous descriptions such as "that red drink").

[0035] Path B (Inverted Index Precise Recall): Precisely matches specific brand names and SKU codes to ensure absolute accuracy of core products.

[0036] Path C (Graph Association Recall): Based on the knowledge graph, query the current store's "frequently purchased list" and "related best-selling products" (e.g., recommend ham sausage when buying instant noodles).

[0037] This invention uses the RRF (Reciprocal Rank Fusion) algorithm to merge and deduplicate the results from the three paths to form a preliminary Top-N candidate set. The weighted fusion algorithm calculation formula can balance the accuracy and coverage of recall, which is another important inventive concept of this invention.

[0038] In one embodiment, the re-ranking step S104 is performed as follows: Calculate the score of each product in the Top-N preliminary candidate set according to the multi-objective optimization strategy, and re-rank the products according to the scores to obtain the Top-K candidate set. The score is calculated as follows: Score = w1 × Semantic_Sim + w2 × User_Preference + w3 × Business_Boost, where Semantic_Sim represents the "semantic matching score" of the candidate products calculated in real time, User_Preference represents the "user's historical repurchase rate", and Business_Boost is calculated based on the real-time inventory status and gross profit contribution. Business_Boost = Base_Score + I(is_promotion)w_promo + I(high_stock)w_stock, where I(high_stock)w_stock represents the high inventory status score, I(is_promotion)w_promo represents the gross profit contribution score, and Base_Score represents the base score.

[0039] In this invention, dynamic business weighting logic is introduced during the reordering phase: Business_Boost = Base_Score + I(is_promotion)w_promo + I(high_stock)w_stock. By adjusting the parameters w_promo and w_stock, operators can flexibly adjust the system's recommendation bias (e.g., end-of-month sales push mode, new product promotion mode) without modifying the code. Introducing the Business_Boost factor during the sorting stage makes the ordering system not only an efficiency tool but also an automated marketing engine (e.g., automatically prioritizing promotional items and overstocked inventory). This is another important inventive concept of this invention.

[0040] This invention constructs a domain-adaptive large language model parser: fine-tuning the large language model for the FMCG industry to enable it to accurately understand industry terminology, resolve fuzzy references, and structure colloquial instructions. It transforms non-standard voice instructions into structured query objects rich in business semantics, laying a high-precision foundation for subsequent retrieval and ranking.

[0041] This invention presents a multi-stage recall framework integrating semantics, rules, and collaborative signals: it innovatively combines vectorized semantic search, a dynamic business rule base, and collaborative filtering algorithms to form a hierarchical recall network. While ensuring recall speed, it significantly improves recall coverage and relevance when facing diverse and colloquial queries.

[0042] This invention constructs a business strategy-driven dynamic re-ranking mechanism: it designs a re-ranking model that integrates real-time business characteristics (promotions, gross profit, inventory) with semantic matching, and adds a flexibly configurable business strategy engine. This ensures that the ranking results not only match user intent but also proactively align with the company's real-time sales strategy, maximizing business value.

[0043] This invention establishes an end-to-end closed loop from voice to order generation: it constructs a fully automated system from voice input, intelligent parsing, candidate recall, business sorting to final order generation and confirmation, truly realizing "ordering with your voice", greatly reducing manual intervention and improving operational efficiency and experience.

[0044] Figure 2 A voice ordering device based on LLM and a fusion recall strategy is shown. The device includes: The receiving unit 201 receives voice order commands input by the user through the terminal and sends the voice order commands to the cloud server. In this invention, the terminal can initiate audio acquisition based on user touch operations. The terminal uses spectral subtraction to remove background steady-state noise and utilizes automatic gain control (AGC) to normalize the volume level. The audio data is framed and encoded in Opus format, and pushed to the cloud server gateway in real time via a WebSocket long connection. The gateway is responsible for authentication and flow control.

[0045] In the inference unit 202, the cloud server transcribes and recognizes the voice order command to obtain the order text, which is then input into the LLM deployed on the cloud server. The LLM performs chain-of-thought reasoning and structuring on the order text to generate a structured query object. The cloud-based ASR service receives the audio stream and uses an end-to-end acoustic model based on the Conformer architecture for transcription. The system loads a heatmap customized for the FMCG industry (containing unique brand abbreviations, specification terms, and promotional jargon), and uses a weighted finite-state converter (WFST) for decoding, significantly improving the accuracy of professional terminology recognition. The transcribed text is input into the LLM. Context injection: The system automatically extracts the current user's profile features (historical preferences, region), current conversation history, and temporal and spatial states, constructs a System Prompt, and injects it into the LLM. Inference generation: The LLM uses chain-of-thought technology to parse complex logical commands. For example, "Same as last time, but change A to B" is parsed as: query historical orders → copy order items → delete product A → add product B.

[0046] Output object: Generates a standardized structured query object, which includes a list of product entities, attribute modifiers (such as "full case" and "promotional items"), and filtering conditions.

[0047] The cloud server, loaded with a large-scale 7B / 13B language model fine-tuned from the entire FMCG industry corpus, serves as the system's "brain," a core template capable of deeply understanding language meaning and making analytical decisions. It primarily includes Intent Recognition, Slot Filling, Inverse Text Normalization, and Prompt Engineering. Its output is not a single keyword, but a structured intermediate representation (IR) containing explicit entities (product names), implicit constraints (price ranges), and contextual referents (related stores).

[0048] Fusion unit 203 initiates M-way recall requests in parallel using structured query objects, and performs fusion processing on the results of the M-way recall requests to obtain a preliminary Top-N candidate set, where M≥2 and N≥2; The reordering unit 204 performs multi-objective optimization based on the initial Top-N candidate set to obtain a Top-K candidate set, where K≥2, and sends the Top-K candidate set to the terminal. The generation unit 205 displays the Top-K candidate set in the terminal and generates a product order based on the user's selection of the Top-K candidate set.

[0049] In one embodiment, the device further includes a fine-tuning unit 206, which records the user's click, modification, confirmation, and rejection operations on the Top-K candidate set displayed on the terminal, forming "voice order command - final ordered item" pairs as positive samples, adding them to the training dataset, and periodically fine-tuning the LLM based on the training dataset. One embodiment of the present invention proposes a computer storage medium storing a computer program. When the computer program on the computer storage medium is executed by a processor, the above-described method is implemented. The computer storage medium can be a hard disk, DVD, CD, flash memory, or other storage device.

[0050] For ease of description, the above-described apparatus is divided into various functional units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.

[0051] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the apparatus described in various embodiments or some parts of the embodiments of this application.

[0052] Finally, it should be noted that the above embodiments are for illustration only and not for limiting the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention without departing from the spirit and scope of the present invention. Any modifications or partial substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A voice ordering method based on LLM and a fusion recall strategy, characterized in that, The method includes: The receiving step involves receiving a voice order instruction input by the user through the terminal and sending the voice order instruction to the cloud server. In the reasoning step, the cloud server transcribes and recognizes the voice order command to obtain the order text, inputs the order text into the LLM deployed on the cloud server, and the LLM performs thought chain reasoning and structured processing on the order text to generate a structured query object. The fusion step involves using structured query objects to initiate M-way recall requests in parallel, and then fusing the results of the M-way recall requests to obtain a preliminary Top-N candidate set, where M≥2 and N≥2. The reordering step involves performing multi-objective optimization based on the initial Top-N candidate set to obtain a Top-K candidate set, where K≥2, and then sending the Top-K candidate set to the terminal. The generation step involves displaying the Top-K candidate set on the terminal and generating a product order based on the user's selection of the Top-K candidate set.

2. The method according to claim 1, characterized in that, The method also includes a fine-tuning step, which records the user's click, modification, confirmation, and rejection operations on the Top-K candidate set displayed on the terminal, forming a voice order instruction-final order product pair as positive samples, adding them to the training dataset, and periodically fine-tuning the LLM.

3. The method according to claim 2, characterized in that, The M-path recall request includes: Vector semantic recall request: Encode the query description into a high-dimensional embedding vector, perform an approximate nearest neighbor search in the vector database, and recall semantically similar products; Inverted index precise recall request: precisely match specific brand names and SKU codes to accurately recall corresponding products; Knowledge graph-based recall: Based on the knowledge graph, query the current store's frequently purchased list and related best-selling products, and recall related products.

4. The method according to claim 3, characterized in that, The operation to fuse the results of M-way recall requests is as follows: calculate the fusion score of the structured query object, which is calculated as Score_final(item)=α×Norm(Vector_Score)+β×Norm(BM25_Score)+γ×Norm(Graph_Weight); Wherein, Score_final(item) represents the fusion score of the structured query object, Norm(.) represents the normalization function, Vector_Score represents the vector semantic recall request result, BM25_Score represents the inverted index precise recall request result, Graph_Weight represents the graph association recall result, and α, β, and γ represent the corresponding fusion weight coefficients, which are dynamically adjusted according to the recall type, which is either fuzzy or precise.

5. The method according to claim 4, characterized in that, The re-ranking step is as follows: Calculate the score of each product in the initial Top-N candidate set according to the multi-objective optimization strategy, and re-rank them according to the scores to obtain the Top-K candidate set. The score is calculated as follows: Score = w1 × Semantic_Sim + w2 × User_Preference + w3 × Business_Boost, where Semantic_Sim represents the semantic matching score of the candidate product calculated in real time, User_Preference represents the user's historical repurchase rate, and Business_Boost is calculated based on the real-time inventory status and gross profit contribution. Business_Boost = Base_Score + I(is_promotion)w_promo + I(high_stock)w_stock, where I(high_stock)w_stock represents the high inventory status score, I(is_promotion)w_promo represents the gross profit contribution score, and Base_Score represents the base score.

6. A voice ordering device based on LLM and a fusion recall strategy, characterized in that, The device includes: The receiving unit receives the voice order instruction input by the user through the terminal and sends the voice order instruction to the cloud server; The reasoning unit, the cloud server, transcribes and recognizes the voice order command to obtain the order text, inputs the order text into the LLM deployed on the cloud server, and the LLM performs thought chain reasoning and structured processing on the order text to generate a structured query object; The fusion unit initiates M-way recall requests in parallel using structured query objects, and performs fusion processing on the results of the M-way recall requests to obtain a preliminary Top-N candidate set, where M≥2 and N≥2. The reordering unit performs multi-objective optimization based on the initial Top-N candidate set to obtain a Top-K candidate set, where K≥2, and sends the Top-K candidate set to the terminal. The generation unit displays the Top-K candidate set on the terminal and generates a product order based on the user's selection of the Top-K candidate set.

7. The apparatus according to claim 6, characterized in that, The device also includes a fine-tuning unit that records the user's clicks, modifications, confirmations, and rejections on the Top-K candidate set displayed on the terminal, forming voice order instructions and final ordered product pairs as positive samples, adding them to the training dataset, and periodically fine-tuning the LLM.

8. The apparatus according to claim 7, characterized in that, The M-path recall requests include: Vector semantic recall request: Encoding the query description into a high-dimensional embedding vector, performing an approximate nearest neighbor search in the vector database, and recalling semantically similar products; Inverted index precise recall request: Performing precise matching on specific brand names and SKU codes to accurately recall corresponding products; Graph association recall: Retrieving the current store's frequently purchased list and related best-selling products based on a knowledge graph, and recalling related products.

9. The apparatus according to claim 8, characterized in that, The operation to fuse the results of M-way recall requests is as follows: calculate the fusion score of the structured query object, which is calculated as Score_final(item)=α×Norm(Vector_Score)+β×Norm(BM25_Score)+γ×Norm(Graph_Weight); Wherein, Score_final(item) represents the fusion score of the structured query object, Norm(.) represents the normalization function, Vector_Score represents the vector semantic recall request result, BM25_Score represents the inverted index precise recall request result, Graph_Weight represents the graph association recall result, and α, β, and γ represent the corresponding fusion weight coefficients, which are dynamically adjusted according to the recall type, which is either fuzzy or precise.

10. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1-5.