An intelligent shopping guide speech recognition method and system based on transfer learning

By adopting the adaptive learning rate adjustment mechanism and BERT model in the intelligent shopping guide system, the dynamic adaptation problems of speech recognition and natural language understanding in the supermarket environment are solved, and rapid learning and stable optimization of new and old knowledge are achieved, improving the accuracy and efficiency of query comprehension.

CN120071936BActive Publication Date: 2025-07-11INSPUR SMART SUPPLY CHAIN TECH (SHANDONG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510525914.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-11
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the prior art, the speech recognition and natural language understanding model of the intelligent shopping guide system in the supermarket environment is difficult to effectively respond to massive, dynamically changing product information and diversified customer queries, resulting in insufficient understanding ability, and the existing learning rate adjustment strategy cannot meet the needs of rapid adaptation and stable optimization of new and old knowledge.

Method used

The intelligent shopping guide method based on transfer learning is adopted, combined with the adaptive learning rate adjustment mechanism, and the comprehensive adaptive indicators of entity novelty and model prediction confidence are calculated, the learning step size is dynamically adjusted, and the fine-tuning process of pre-trained language models is optimized, especially the end-to-end training is used using the BERT-base-chinese model, and the BIO labeling system and supermarket-related data are used for labeling to ensure the efficient adaptation of the model in the supermarket environment.

Benefits of technology

It significantly improves the accuracy and robustness of the model in the dynamic supermarket environment, can quickly adapt to new products and complex queries, and provides smarter and more efficient shopping guide services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071936B_ABST
    Figure CN120071936B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing, and particularly to an intelligent shopping guide speech recognition method and system based on transfer learning. During the speech recognition and optimization stage of the intelligent shopping guide, a pre-trained language model is used for transfer learning fine-tuning, and a comprehensive adaptation index is obtained by dynamically combining the entity novelty score and the model prediction confidence of the training batch. Further, the learning step size is adjusted in real time in combination with the ratio of the model loss, so that the model can more intelligently adapt to new products, promotion information, and queries of different difficulties, optimizing the overall shopping guide experience. The present invention applies transfer learning to fine-tune the pre-trained language model for the supermarket shopping guide NLU task, and adopts a multi-stage adaptive learning rate adjustment strategy that integrates entity novelty, prediction confidence, and loss dynamics during fine-tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to an intelligent shopping guide speech recognition method and system based on transfer learning. Background Art

[0002] With the development of modern retail industry, large comprehensive supermarkets (hypermarkets) have become the main places for urban residents' daily shopping. Such hypermarkets usually cover a vast area and have tens of thousands of commodity categories, including food, fresh produce, daily necessities, household appliances, etc., and commodity information (such as prices, promotions, inventory, specific placement locations) is in a continuous dynamic change. When shopping in such a complex environment, customers often face the dilemma of information asymmetry: it is difficult to quickly find the exact location of the desired commodity, it is not easy to obtain the latest price or participate in preferential activities in real time, and they lack understanding of the dazzling array of commodities, especially newly launched products. In order to improve the shopping experience of customers and the operation efficiency of shopping malls, the introduction of intelligent shopping guide services has become an industry trend. The intelligent shopping guide system based on voice interaction, through the interaction interfaces deployed on service desks, self-service terminals or mobile robots, allows customers to ask questions in the most natural way, and theoretically can provide more convenient and efficient services than traditional touch-screen queries or limited manual shopping guides.

[0003] However, to achieve truly intelligent, accurate and fluent voice shopping guides, the core challenge lies in the system's deep understanding ability of customers' natural language queries. Voice queries in a hypermarket environment have significant characteristics: First, they involve a large amount of continuously updated domain-specific vocabulary such as commodity names, brand names, specifications, promotion terms, etc., and general speech recognition and natural language processing models often have serious out-of-vocabulary (OOV) problems and semantic understanding biases; Second, customers' questioning methods are flexible and diverse, with colloquial expressions, ellipses, references, and even vague descriptions being common; Third, the query intentions are complex and may involve multiple aspects such as price, location, attributes, evaluations, recommendations, etc. Therefore, the system not only needs to accurately convert speech into text (ASR), but more importantly, it needs to precisely understand the intention behind the text and extract all key information entities (NLU). Directly training a complex NLU model from scratch for each hypermarket to cope with the above challenges requires a large amount of high-quality domain-annotated data, which is costly and unrealistic. Therefore, using large language models pre-trained on a large amount of general text (such as BERT, etc.) and fine-tuning them with a relatively small amount of domain-annotated data through transfer learning to quickly adapt to the specific field of hypermarket shopping guides has become the most promising technical path currently.

[0004] Currently, the fine-tuning technology based on transfer learning has been widely applied to various NLP tasks. In the field of intelligent shopping guides, there have also been attempts to use pre-trained models for NLU optimization. However, during the fine-tuning process, the setting of the learning rate is a crucial and tricky part. In the existing technologies, the commonly adopted methods include: (1) using a fixed and very small learning rate for global fine-tuning; (2) adopting a preset learning rate decay strategy, such as linear decay, exponential decay, or decay by steps. The main drawbacks of these methods are that their adjustment methods fail to fully consider the dynamics and heterogeneity of the training data in the supermarket shopping guide scenario.

[0005] Specifically, when new product names or rare promotional terms appear in the training data, the model needs to quickly learn this new knowledge with a relatively large step size; while when dealing with very common and simple queries (such as "How much is an apple"), it needs to be fine-tuned stably with a smaller step size to avoid disrupting the learned patterns. A fixed learning rate cannot meet these two requirements simultaneously, which may lead to slow learning of new knowledge or over-adjustment and oscillation on simple data. The preset decay strategy has nothing to do with the content of the training data and cannot respond to the specific challenges encountered during the training process in real time. Although there are some general adaptive learning rate algorithms (such as Adam, RMSProp, etc.), they mainly adjust the learning rate according to historical gradient information and do not explicitly combine the specific characteristics of the domain task. Therefore, when the existing learning rate adjustment strategies are applied to the fine-tuning of the supermarket shopping guide NLU model, it is often difficult to achieve the optimal convergence speed, stability, and final model performance, which limits the further improvement of the understanding ability of the intelligent shopping guide system. Summary of the Invention

[0006] In view of the problem that the above fixed learning rate cannot meet the two requirements simultaneously, in the first aspect, the present invention proposes an intelligent shopping guide speech recognition method based on transfer learning, including: obtaining the inquiry speech of the customer for automatic speech recognition to obtain the initial recognition text; performing natural language understanding on the initial recognition text based on a pre-trained language model to identify the query intention of the customer and the entities in the query text; constructing a query instruction according to the query intention and entities and retrieving in the supermarket commodity information database; generating a reply content based on the retrieved relevant information and outputting it through voice or / and interface; using manually labeled data to adjust the parameters of the pre-trained language model through backpropagation and parameter update; the parameter update includes an adaptive learning rate , there is:

[0007] ;

[0008] where represents the base learning rate; represents the batch loss value; Indicates the batch Average loss before; Indicates the minimum integer; tanh represents the hyperbolic tangent function; Indicates the batch Comprehensive adaptation index; Indicates the set scaling factor; the comprehensive adaptation index is positively correlated with the entity novelty score and negatively correlated with the prediction confidence; the reciprocal mean of the occurrence frequencies of all entities in the current training batch constitutes the entity novelty score; the arithmetic mean of the maximum prediction probability of all query intents in the current training batch and the label probability of the entity constitutes the prediction confidence.

[0009] The method of the present invention realizes a complete intelligent voice shopping guide process by combining automatic speech recognition, natural language understanding based on a pre-trained model, database query and response generation. Especially in the fine-tuning stage of the natural language understanding model, a specific adaptive learning rate is innovatively adopted. This learning rate not only considers the conventional training loss changes, but also incorporates a comprehensive adaptation index customized for the supermarket shopping guide scenario, which directly correlates with entity novelty and model prediction confidence. Compared with the existing technologies that use a fixed learning rate, a preset decay strategy or a general adaptive algorithm based only on the gradient history, this method can dynamically and finely adjust the learning step according to the actual novelty degree of the current batch of data and the model's understanding of it. When encountering new products or situations where the model is uncertain, it can learn and adapt faster, and when dealing with common and easily understood queries, it can stably optimize, significantly improving the model's understanding accuracy, robustness and convergence efficiency for various queries in the supermarket environment with dynamic changes in commodity information, and overcoming the defect that the learning rate adjustment strategy in the existing technologies lacks scene perception ability.

[0010] Furthermore, the calculation method of the comprehensive adaptation index is specifically as follows:

[0011] ;

[0012] where Indicates the batch Comprehensive adaptation index; Indicates the batch Entity novelty score; Indicates the batch Prediction confidence.

[0013] Furthermore, the calculation method of the entity novelty score is specifically as follows:

[0014] ;

[0015] where Indicates the batch Entity novelty score; Represents the total number of entities included in the batch ; Represents the frequency of occurrence of the entity during the training process; Represents the natural logarithm function.

[0016] The present invention specifically defines the calculation method of the entity novelty score as the average of the inverse logarithm frequency correlation values of the entities within the batch. This method can more smoothly and effectively quantify the novelty degree of the entity compared with the simple inverse frequency or binary new / old judgment, gives the highest novelty weight to extremely rare entities, and can also distinguish the differences between common entities and very common entities, providing a finer-grained signal that is non-linearly related to the entity occurrence frequency for guiding the learning rate adaptation adjustment.

[0017] Furthermore, the calculation method of the prediction confidence is specifically as follows:

[0018] ;

[0019] wherein represents the prediction confidence of the batch ; represents the number of samples in the batch ; represents the average value of the number of tokens corresponding to the samples in the batch ; represents the maximum value of the intent prediction probability of the sample ; represents the entity label probability of the th token in the sample

[0020] The present invention specifically stipulates that the prediction confidence is calculated by comprehensively averaging the maximum intent prediction probability of the samples within the batch and the entity label probabilities of all tokens. This method comprehensively reflects the overall prediction confidence level of the model for the two subtasks of intent classification and sequence labeling on the current batch. Compared with the single index that only considers the intent confidence or only considers the lowest / average entity confidence, it provides a more global and robust confidence evaluation, enabling the learning rate adjustment to be more accurately based on the comprehensive confidence level of the model for the entire prediction task.

[0021] Furthermore, the BERT-base-chinese model is used as the basic natural language understanding model, and the pre-trained language model is obtained by collecting the manually annotated data and performing end-to-end supervised training based on the basic natural language understanding model.

[0022] In the present invention, by specifically selecting the widely verified BERT-base-chinese model as the base model and adopting an end-to-end supervised training method for fine-tuning, it is ensured that the NLU module can effectively utilize the powerful general language representation ability and fully adapt to the specific language patterns and task requirements in the field of supermarket shopping guides through the adjustment of overall parameters. Compared with the method of selecting a weaker base model or only adjusting some parameters for fine-tuning, it is expected to obtain better domain adaptation performance and higher understanding accuracy.

[0023] Furthermore, the manually written simulated queries, supermarket commodity catalogs, promotional activity descriptions, and floor shelf layout information are used as the original text, and the BIO annotation system is used for the original text to obtain the manually annotated data.

[0024] In the present invention, by specifically indicating that the manually annotated data for fine-tuning is derived from highly relevant original texts such as simulated queries, commodity catalogs, promotional descriptions, and shelf layouts, and adopting the standard BIO annotation system, the domain pertinence and annotation standardization of the training data are ensured, enabling the model to directly learn the real language expressions and entity structures in the supermarket scenario. Compared with training using general corpus or data with non-standard annotations, it can significantly improve the accuracy and reliability of the NLU model in practical applications.

[0025] Furthermore, it also includes using a DNN-based denoising model to preprocess the customer's inquiry voice.

[0026] In a second aspect, the present invention provides an intelligent shopping guide voice recognition system based on transfer learning, including: a voice input unit: used to receive the customer's voice instructions; a voice recognition unit: connected to the voice input unit and used to convert the voice instructions into text; a natural language understanding unit: connected to the voice recognition unit and used to process the text based on a pre-trained language model to identify the customer's query intention and key entities; the pre-trained language model included in the natural language understanding unit is trained and configured according to the fine-tuning steps in the intelligent shopping guide voice recognition method based on transfer learning; an information processing and query unit: connected to the natural language understanding unit and a commodity information database, used to construct database queries, execute queries, and retrieve information; a commodity information database: storing detailed information of all commodities in the supermarket, including real-time prices, promotional information, and accurate shelf location information; a reply generation unit: connected to the information processing and query unit and used to generate anthropomorphic reply content according to the retrieval results; an output unit: connected to the reply generation unit and used to output the reply content through voice or / and an interface.

[0027] Further, the output unit includes at least one of the following: a speech synthesizer and a speaker, configured to synthesize the reply content into speech and broadcast it; a display screen, configured to display the text of the reply content, associated images, or a navigation path generated according to the commodity location information.

[0028] Further, the voice input unit includes at least one microphone array for receiving voice commands from customers.

[0029] The technical effects of the present invention are as follows:

[0030] The core innovative content of the present invention is concentrated in a novel multi-stage adaptive learning rate adjustment mechanism adopted when fine-tuning a pre-trained language model for NLU tasks in the supermarket shopping guide field. Different from the existing technologies that adopt a fixed learning rate, a preset decay strategy, or a general adaptive algorithm that only relies on the gradient history, the mechanism proposed by the present invention deeply combines the characteristics of the supermarket shopping guide scenario. First, the relative novelty of entities within a batch is calculated to quantify the emergence degree of new knowledge; secondly, the overall confidence of the model's prediction for the current batch is evaluated to judge the model's mastery of the current input; then, these two scenario-based features are fused into a comprehensive adaptation metric, which reflects the adaptation requirements of the current batch of data in the dimensions of "novelty" and "model certainty"; finally, this scenario adaptation signal is combined with a factor measuring the relative learning difficulty of the current batch, and the base learning rate is jointly modulated through a smooth bounded function. This mechanism enables the learning rate to simultaneously respond to the immediate difficulty during the training process and the domain characteristics of the data itself (such as new and old knowledge, prediction confidence), achieving more intelligent and refined learning step control, thereby significantly improving the learning efficiency, stability, and final performance of the model in a dynamic supermarket environment. Description of the Drawings

[0031] By reading the following detailed description with reference to the accompanying drawings, the above and other purposes, features, and advantages of the exemplary embodiments of the present invention will become readily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0032] Figure 1 is a flowchart schematically showing an intelligent shopping guide voice recognition method based on transfer learning in an embodiment of the present invention;

[0033] Figure 2 is a block diagram schematically showing the structure of an intelligent shopping guide voice recognition system based on transfer learning in an embodiment of the present invention. Detailed Embodiments

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0035] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0036] An embodiment of an intelligent shopping guide speech recognition method based on transfer learning:

[0037] As Figure 1 shown, the intelligent shopping guide speech recognition method based on transfer learning of the present invention includes:

[0038] S1. Implement the complete process of the intelligent voice shopping guide method in large-scale commercial supermarkets, including voice processing, core recognition, database query, intelligent response, and continuous optimization of the model.

[0039] In modern large-scale commercial supermarkets, there are a wide variety of goods, numerous categories, complex layouts, and frequent changes in promotional activities. During the shopping process, customers often face pain points such as time-consuming to find specific goods, inability to obtain accurate price or preferential information in a timely manner, and lack of understanding of new products. Although traditional manual shopping guides can provide help, they are limited by labor costs, service coverage, and response capabilities during peak hours, and it is difficult to meet the immediate and accurate information needs of all customers. Therefore, in this embodiment, intelligent shopping guide services can be provided through natural human-computer voice interaction, accurately understanding customer intentions, and combining real-time dynamic information of the commercial supermarket.

[0040] In one embodiment, the complete intelligent shopping guide process may include: customer voice collection and preprocessing, speech recognition and optimization, information extraction and database, feedback and interaction, and recording interaction data to optimize the model.

[0041] Among them, customer voice collection and preprocessing means that first, the natural voice command issued by the customer near the shopping guide device is captured by the microphone array of the shopping guide device; then, the built-in voice processing system of the shopping guide device enhances the voice of the target speaker; finally, the environmental noise can be suppressed and the reverberation can be eliminated through a DNN-based denoising model, and a digital audio stream of the clear voice of the customer is output. It should be noted that: in this embodiment, the shopping guide device can be an interactive intelligent shopping guide counter or an intelligent shopping guide robot.

[0042] The voice recognition and optimization process is the key process of this embodiment. The recognition effect of this process on the preprocessed digital audio stream directly determines the output accuracy of the final shopping guide device and the human-computer interaction experience. In this embodiment, this process can first convert the audio stream into an initial text through an automatic speech recognition (ASR) engine; then, through a natural language understanding (NLU) module, deeply analyze the recognized text, and finally accurately recognize the customer's query intention (such as querying price, location, promotion, inventory, product details, seeking recommendations, etc.) and extract the key information entities in the query (such as product name, brand, specification, attribute words, etc.). In subsequent steps S2 and S3, the specific implementation of the voice recognition and optimization process of this embodiment will be emphasized. This process can output a structured semantic representation, for example: {intent: "query_price_location", entities: [{"entity_type": "brand", "value": "XX brand"}, {"entity_type": "product", "value": "cucumber"}]}.

[0043] After obtaining the structured semantic representation in the information extraction and database query process, the system parses this structured information and automatically constructs a query instruction that conforms to the background database query language (such as SQL). This instruction will be used to query the real-time updated supermarket product information database. In this embodiment, this database should at least maintain the product category, brand, origin, specification, real-time selling price, current promotion activities, inventory status, associated product information, and the precise physical location in the mall (for example: floor - area - shelf number - tier number). This process can output the detailed information of one or more product records that match the query conditions.

[0044] In the feedback and interaction process, the database first queries and returns the detailed product information. Based on the retrieved information, the system can construct a user-friendly, clear, and accurate response through natural language generation technology (NLG); then the response content will prioritize highlighting the core information directly asked by the user (such as price and location); further, the system can integrate a recommendation algorithm to recommend relevant products or promotions according to the user's query. It should be noted that: in this embodiment, the presentation method of the response of the shopping guide device is diverse. In one embodiment, it can generate a voice broadcast through a text-to-speech (TTS) engine, and in another embodiment, it can display text, pictures, price tags, and promotion details on the display screen of the shopping guide device. If the hardware of the shopping guide device supports it, it can also dynamically render the indoor navigation path map from the current device location to the target product location.

[0045] Finally, in the process of recording and optimizing the interactive data model, the back-end of the shopping guide device can store the real data of human-computer interaction (after user consent and anonymization) for offline analysis and periodic iterative optimization of the model. In particular, the NLU model can be incrementally trained or re-fine-tuned with these new data to continuously improve its performance in real scenarios and adapt to changes in new products, new promotional statements, and customer expression habits. The training and application processes of the models used in the above processes are well-known technologies and will not be elaborated here.

[0046] S2. Detailed Explanation of Speech Recognition and Optimization Core: It involves the whole process of ASR speech-to-text conversion, selection of pre-trained NLP models, construction of domain annotation data, key model fine-tuning and transfer learning, performance verification and iterative optimization, and finally deployment and online launch.

[0047] In the whole process, the performance of "speech recognition and optimization" directly determines the effectiveness of subsequent links and is the key to improving user experience. Specifically, in this embodiment, this process includes the following 6 steps, namely: application of basic acoustic and language models, selection of basic natural language understanding models, collection and preparation of domain-specific data, model fine-tuning - transfer learning, model verification and optimization, and model deployment and (optional) online learning.

[0048] In this embodiment, the step of applying basic acoustic and language models uses an end-to-end ASR system. Exemplarily, in one embodiment, an acoustic model based on the Transformer architecture can be used to directly map the acoustic feature sequence to a subsequence. In another embodiment, a hybrid system combining a deep neural network acoustic model (such as TDNN, LSTM) and a traditional language model (such as N-gram or RNN-LM) can be adopted. The output of this step is an initially recognized text string, for example: "I want to find the XX brand milk".

[0049] In this embodiment, the BERT-base-chinese model can be selected in the step of selecting the basic natural language understanding model. This model has learned rich vocabulary, grammar, and semantic knowledge through unsupervised pre-training (such as Masked Language Model, Next Sentence Prediction tasks) on a large amount of text (such as encyclopedias, news, web pages). Subsequently, domain-specific data is collected. In one embodiment, if there are some traditional text-based shopping guide counters with non-intelligent voice interaction in the supermarket before, the original user query logs can be collected as the original text. In another embodiment, artificially written simulated queries, supermarket product catalogs (including categories, brands, names, specifications, aliases), promotional activity descriptions, floor shelf layout information, etc. can be used as the original text.

[0050] After obtaining the original text, it is necessary to clean, de-duplicate, and standardize the original text, and then perform manual annotation, specifically including: assigning predefined intent categories to each query text, such as: query price, query location, query inventory, query promotion, query product attributes, request recommendation and chat, etc. Perform entity annotation for each query text. In this embodiment, the BIO annotation system can be used to accurately annotate the key information segments and their types involved in the text, such as: B-product name, I-product name, B-brand, I-brand, B-specification, I-specification, B-attribute word (such as "price", "location", "discount"), etc. Preferably, special attention should be paid to annotating various aliases, abbreviations, and colloquial expressions of products and brands. Exemplary description:

[0051] Sample 1:

[0052] Text: "What is the price of cucumbers of XX brand today?";

[0053] Intent: query_price (Query price);

[0054] Entities - Using BIO annotation method:

[0055] Today: O (Outside);

[0056] Day: O;

[0057] X: B-brand (Begin-Brand);

[0058] X: I-brand (Inside-Brand);

[0059] Brand: I-brand;

[0060] Of: O;

[0061] Yellow: B-product (Begin-Product);

[0062] Cucumber: I-product (Inside-Product);

[0063] What: O;

[0064] About: O;

[0065] Price: B-attribute (Begin-Attribute, referring to price);

[0066] Value: I-attribute (Inside-Attribute);

[0067] ?: O.

[0068] Sample 2:

[0069] Text: "Where is the ketchup?";

[0070] Intention: query_location (query location);

[0071] Entities:

[0072] Tomato: B - Product;

[0073] Ketchup: I - Product;

[0074] In: O;

[0075] Which: B - Location Indicator (Begin - Location Indicator);

[0076] One: I - Location Indicator;

[0077] Shelf: I - Location Indicator;

[0078] Shelf: I - Location Indicator;

[0079] Shelf: I - Location Indicator;

[0080] ?: O.

[0081] Finally, this step can output a structured, high - quality labeled dataset that covers the main query types and entities in the supermarket shopping guide scenario , and there are:

[0082] ;

[0083] Among them represents the text of the th labeled sample; represents the intention of the th labeled sample; represents the entity of the th labeled sample; represents the total number of labeled samples.

[0084] After selecting a basic natural language understanding model and obtaining a domain - specific labeled dataset , first, a task - specific output layer can be added on top of the pre - trained model, which is the basic natural language understanding model (for example, a linear layer + Softmax for intention classification, and a linear layer + Softmax for sequence labeling (entity recognition)); then use the dataset to perform end - to - end supervised training on the entire model (including the pre - trained part and the newly added task layer). This step can output a natural language understanding model specifically for the supermarket shopping guide task .

[0085] Further obtain the natural language understanding model After that, the performance of the model can be comprehensively evaluated. The indicators used in this embodiment include the accuracy (Accuracy), precision (Precision), recall (Recall), F1 score of intent recognition, and the accuracy, recall, and F1 score of entity recognition (calculated separately by entity type or macro / micro averaged here). Subsequently, the advantages and disadvantages of the model are analyzed according to the verification results (for example, which types of intent / entity recognition are poor), the model architecture is adjusted, the hyperparameters are modified, the training data is increased or improved, different pre-trained models are tried, etc., and then the fine-tuning and verification process is repeated until the model performance reaches the predetermined goal; after debugging, this step can output an optimal NLU model version for deployment. The above-mentioned evaluation of model performance and adjustment of model parameters are well-known technologies and will not be repeated here.

[0086] Finally, the optimized and finalized NLU model can be deployed to the actual shopping guide device or cloud server to provide a real-time NLU service interface. At the same time, for the continuous evolution of the model, an online learning or regular update mechanism can also be designed. For example, in the process of recording the interactive data optimization model in step S1, user feedback signals (such as whether the user is satisfied with the results, whether the recognition errors have been corrected) and new interactive data are collected, and after screening and labeling, they are used for incremental updates of the model (online fine-tuning) or periodic full retraining.

[0087] S3. Integrate batch entity novelty scores and prediction confidence to obtain a comprehensive adaptation index, and dynamically adjust the learning step size in combination with batch loss to improve the domain adaptability of transfer learning.

[0088] The supermarket shopping guide NLU model may face the following practical situations when fine-tuning transfer learning, such as: New entities emerge: new products, brands, and promotions continue to appear, and the model needs to learn quickly; query complexity varies: simple and direct queries ("Where is the milk") coexist with complex and ambiguous queries ("Is there any promotion for the imported yogurt that is not too sweet for the elderly?"); model confidence changes: the model has different prediction confidence for different inputs, and low confidence may mean that more cautious or stronger adjustments are needed. A single learning rate strategy is difficult to take into account the three requirements of quickly adapting to new knowledge, stably learning known patterns, and adjusting learning intensity according to the model's own state.

[0089] Therefore, in the model fine-tuning-transfer learning step of S2 above, the learning rate The setting is crucial for the convergence speed, stability, and final performance of model training. Traditional learning rate strategies with fixed or pre-determined decay cannot well adapt to the challenges of data diversity and dynamics in the supermarket shopping guide scenario. Therefore, an adaptive learning rate can be set in this embodiment to improve the effect of model training.

[0090] Since new products are the norm in the supermarket scenario, the model should be able to pay more attention to samples containing new entities and accelerate learning. Therefore, an indicator based on the dynamic update of entity frequencies during training can be constructed to quantify the "novelty" degree of named entities (such as products, brands, etc.) appearing in the current training batch relative to the training data that the model has seen.

[0091] Specifically, before or during training, a counter can be maintained first to record the number of occurrences (or frequencies) of each entity encountered in the training data; then for each identified entity in the current batch , which can also be denoted as , calculate its novelty score ; where adding 1 to the logarithmic frequency is to avoid dealing with the first occurrence of an entity as zero, that is, at this time

[0092] , the novelty is the highest, and adding 1 to the denominator ensures that the score is in the range of (0,1]; finally, calculate the average novelty score of all entities in the batch, and there is:

[0093] where represents the entity novelty score of batch , with a value range in (0,1], and the higher its value indicates that the entities included in the batch are more novel or rare; represents the total number of entities included in batch ; represents the frequency of entity appearing during training;

[0094] When the batch contains a large number of entities that appear for the first time or rarely appear, is generally small, at this time is closer to 0, while is closer to 1, that is is closer to 1, and at this time the batch has a high novelty; on the contrary, when most of the entities in the batch are common entities, that is is large, then is also large, and at this time The smaller and closer to 0, the lower the novelty of the batch.

[0095] Further obtain the confidence of the model prediction to measure the overall "confidence" of the model in intent classification and entity annotation for the samples in the current batch The confidence of the model prediction is a direct reflection of its understanding of the current input. Low confidence may mean that the input is ambiguous, the model is insufficient in ability, or it encounters boundary situations and needs to be learned and adjusted; high confidence indicates that the model is more familiar with the current pattern.

[0096] First, for the intent classification task, obtain the highest probability of the model output ; then, for the entity recognition task, obtain the probability of the predicted label for each token ; finally, calculate the arithmetic mean of the average intent confidence and the average entity label confidence for all samples in the batch, and there is:

[0097] ;

[0098] Where represents the prediction confidence of the batch ; represents the number of samples in the batch ; represents the average sequence length (number of tokens) of the samples in the batch ; represents the maximum value of the intent prediction probability of the sample ; represents the entity label probability of the th token in the sample

[0099] In this embodiment, a high novelty ( ) or a low confidence ( ) may both indicate the need for a stronger learning signal, but the combination of the two can better reflect the specific situation. Exemplary illustration: For samples that are novel but have high confidence, only moderate adjustment may be required; for samples that are common but have low confidence, the reasons need to be identified and learning needs to be strengthened. Therefore, in this embodiment, a comprehensive adaptation index can be calculated, specifically as:

[0100] ;

[0101] Where represents the comprehensive adaptation index of the batch , and the value range is [0, 2]; represents the entity novelty score of the batch ; represents the batch The prediction confidence. When the batch is very novel, that is is close to 1, and the model prediction is very unconfident, that is is close to 0, is close to 2, indicating that the batch of data requires more model adaptation; conversely, when the entity is common, that is is close to 0, and the model is very confident, that is is close to 1, is closer to 0, indicating that the batch of data does not require significant model adaptation. Finally, the adaptive learning rate calculation method is obtained as:

[0102] ;

[0103] where represents the adaptive learning rate for fine-tuning the model during the transfer learning process for batch ; represents the base learning rate; represents the loss value of batch ; represents the average loss of all batches during the model learning process; represents a very small integer to avoid the denominator being zero, which can be set to 1e-8 in this embodiment; represents the comprehensive adaptation index of batch ; represents the scaling factor used to control the impact on the learning rate. In this embodiment, it can be set to the empirical value 2 to make it on the same order of magnitude as the loss ratio term and avoid completely dominating the adjustment; tanh represents the hyperbolic tangent function, which is used to smoothly map the comprehensive signal inside the parentheses to (-1, 1), so that the factor (1 + tanh( )) fluctuates within the range of (0, 2).

[0104] The above formula allows the learning rate to respond to information in two dimensions simultaneously, namely the relative difficulty and the scene adaptation requirement. The relative difficulty is represented by the term, which reflects the difficulty level of the current batch compared to the recent average level. The scene adaptation requirement is represented by term, which incorporates the adaptive adjustment signal for the supermarket scene extracted from entity novelty and prediction confidence. After adding the two, they are smoothed and bounded by tanh, and finally modulate the base learning rate , and is introduced to balance the contributions of the two signal sources.

[0105] When the batch is both difficult, that is , and has strong scene adaptation requirements, that is When it is large, the value inside the tanh function's parentheses will be significantly positive, and at this time, tanh approaches 1, will be significantly greater than (close to ), achieving fast learning and adaptation; while when the batch is simple, that is , and has weak scene adaptation requirements, that is When it is small, the value inside the tanh function's parentheses will be significantly positive, and at this time, tanh approaches -1, will be significantly less than (close to 0), achieving fine-tuning and stable learning; at the same time, in other combination cases (such as difficult but weak adaptation requirements, or simple but strong adaptation requirements), the adjustment factor between the two will be obtained according to the relative strength of the two signals, so that the learning rate can more precisely respond to complex data characteristics.

[0106] The adaptive learning rate obtained through calculation can enable the model to more intelligently allocate learning resources during the fine-tuning process, be more sensitive to new knowledge and difficult samples, and be more robust in processing known and simple patterns, thus achieving better overall performance in complex supermarket shopping guide scenarios.

[0107] S4. Review of the complete process of the intelligent shopping guide integrating the adaptive learning rate, the overall workflow and examples after integrating the multi-stage dynamic learning rate adjustment strategy.

[0108] In summary, this embodiment provides an optimized intelligent shopping guide method. This method follows a complete process from voice acquisition, preprocessing, core recognition and optimization, information retrieval, to feedback interaction and data record optimization. During the model fine-tuning process in the "recognition and optimization" stage, it adopts the adaptive learning rate mechanism detailed in step S3. This mechanism enables the NLU model based on transfer learning to more intelligently and efficiently adapt to the characteristics of fast commodity updates and diverse customer queries in a large supermarket environment. The following gives two application examples:

[0109] Suppose the system has been deployed on a shopping guide robot at the entrance of the second floor of a large supermarket, and its NLU model has been fine-tuned using the adaptive learning rate strategy proposed in this embodiment.

[0110] Scenario 1: Common query:

[0111] Customer's voice: "Excuse me, which floor is the pure milk on?"

[0112] First, the robot microphone captures the voice, removes background noise, and obtains clear audio; then the ASR system outputs the text: "Excuse me, which floor is the pure milk on"; further, the NLU model ( ) Process the text:

[0113] During fine-tuning, the learning of such samples: "pure milk" is a common entity, that is low, the query is simple and direct, and the model prediction confidence is high. Therefore the value is low; at the same time, the loss of such samples may also be lower than the recent average . According to the formula in step S3, its comprehensive signal is negative, output a negative value, and finally will be slightly less than , which enables the model to make stable and fine adjustments when processing such familiar samples.

[0114] The current inference output is:

[0115] intent: query_location, entities: [{"entity_type": "product", "value": "pure milk"}];

[0116] Then construct a query statement to find the location information of "pure milk"; subsequently, the database returns: "Pure milk is mainly distributed on the shelves B01 - B05 in the refrigerated area on the second floor"; finally, the robot announces the voice: "Pure milk is mainly in the refrigerated area on the second floor, and it is available on shelves B01 to B05." Optionally, at the same time, text information is displayed on the screen, and the map navigation path of area B in the refrigerated area on the second floor is highlighted.

[0117] Scenario 2: Involving new products or more complex queries:

[0118] Customer's voice: "I'm looking for the newly launched 'Zero Sensation Blueberry Flavor' sparkling water. Is there any discount?"

[0119] First, perform voice collection and preprocessing in the same real-time manner as described above; then the ASR outputs the text: "I'm looking for the newly launched Zero Sensation Blueberry Flavor sparkling water. Is there any discount?"; further, the NLU model ( ) processes the text; during fine-tuning, the learning of such samples: "Zero Sensation Blueberry Flavor sparkling water" may be a newly launched product. At the initial stage of training is low, resulting in being high. At the same time, due to being a new product or a combined name, the initial prediction confidence of the model may not be high. The query also involves the "discount" attribute. Therefore the value will be high. The loss of such samples may also be higher than . According to the formula in step S3, the comprehensive signal is positive, output a positive value, and finally will be significantly greater than . This enables the model to learn this new entity and its associated query patterns at a faster speed.

[0120] The current inference output is:

[0121] intent: query_promotion_availability, entities: [{"entity_type": "product", "value": "Zero Sense Blueberry Flavored Sparkling Water"}, {"entity_type": "attribute", "value": "on discount"}]

[0122] Then, construct a query to find the inventory, location, and promotion information of "Zero Sense Blueberry Flavored Sparkling Water". The database returns: "The 'Zero Sense Blueberry Flavored' sparkling water is in stock, located on Shelf D12 in the Beverage Area on the third floor. It is a new product promotion this week, and you can enjoy a 50% discount on the second item."; Finally, the robot announces: "Found it! The 'Zero Sense Blueberry Flavored' sparkling water you want is on Shelf D12 in the Beverage Area on the third floor. There is an activity now, 50% off on the second item!" At the same time, the screen of the shopping guide device shows: product pictures, prices, promotion details, and the navigation path to Shelf D12 on the third floor.

[0123] Optionally, recording this successful interaction of the new product query in the background helps with subsequent consolidation of learning.

[0124] From the above two comparison examples, it can be seen that the intelligent shopping guide method proposed in this embodiment, which includes a multi-stage adaptive learning rate fine-tuning strategy, can dynamically adjust the learning process according to the commonness, novelty of the query content, and the prediction confidence of the model. Thus, in practical applications, it can not only stably handle regular requests but also quickly adapt to new product information and complex queries, providing a more accurate, efficient, and intelligent shopping guide service.

[0125] An embodiment of an intelligent shopping guide voice recognition system based on transfer learning:

[0126] On the other hand, the present invention also provides an intelligent shopping guide voice recognition system based on transfer learning.

[0127] As Figure 2 shown, the intelligent shopping guide voice recognition method of the present invention includes the following units that are interconnected and work together:

[0128] Voice input unit 101: responsible for receiving the natural voice commands issued by customers. This unit can be configured with at least one microphone array to enhance the voice pickup effect and suppress environmental noise, ensuring clear voice signals are captured.

[0129] Automatic Speech Recognition Unit (ASR) 102: Connected to the voice input unit 101 to receive the collected voice signal. Advanced acoustic models and language models are integrated inside this unit to convert the input voice stream into a text string in real time.

[0130] Natural Language Understanding Unit (NLU) 103: Connected to the Automatic Speech Recognition Unit 102 to receive the text output by it. This unit is based on a powerful pre-trained language model (such as BERT, etc.). The key is that this pre-trained language model is trained and configured according to the steps defined in the aforementioned intelligent shopping guide speech recognition method based on transfer learning of the present invention, which includes a specific adaptive learning rate fine-tuning strategy.

[0131] Therefore, this NLU unit can accurately understand the language in the shopping mall shopping guide scenario, and identify the customer's core query intent (such as querying price) and key entities (such as the commodity name "milk").

[0132] Information Processing and Query Unit 104: Connected to the Natural Language Understanding Unit 103 to receive the structured intent and entity information output by it. This unit is responsible for converting this information into query instructions that can be executed in the background database (for example, generating SQL statements). At the same time, it connects and manages the interaction with the commodity information database 105, and can achieve query and return results through remote API calls.

[0133] Commodity Information Database 105: Stores detailed and real-time updated information of all commodities in this shopping mall. This information includes at least the real-time price of the commodity, details of current promotion activities, and accurate physical location information of the commodity on the shelf.

[0134] Reply Generation Unit 106: Connected to the Information Processing and Query Unit 104 to receive the query results retrieved from the database. This unit is responsible for organizing the retrieved data (such as price, location, promotion description) into one or more paragraphs of natural, fluent, and anthropomorphic reply text. For example, generating the text "The current price of this milk is 10 yuan per bottle, and it is on Shelf A02 in the refrigerated area on the second floor."

[0135] Output Unit 107: Connected to the Reply Generation Unit 106 to receive the generated reply content. This unit is responsible for presenting the information to the customer. This unit may include:

[0136] Speech Synthesizer (TTS) and Speaker: Convert the reply text into voice for broadcast.

[0137] Display Screen: Used to display the reply text, commodity pictures, price tags, promotion posters, or dynamically generate and display a navigation path map according to the retrieved location information.

[0138] Example of Workflow:

[0139] When a customer asks the system (such as a shopping guide robot) "Where is the yogurt of XX brand?",

[0140] The voice input unit 101 (including the microphone array) captures the voice;

[0141] The voice recognition unit 102 converts it into the text "Where is the yogurt of XX brand";

[0142] The natural language understanding unit 103 (using a model fine-tuned by a specific method) analyzes the text, identifies the intention as "querying location", and the entities as "brand: XX brand" and "product: yogurt";

[0143] The information processing and query unit 104 constructs a query based on the NLU result and searches for the location of the yogurt of XX brand in the product information database 105;

[0144] The product information database 105 returns the location information (such as "Shelf C05 in the dairy section on the third floor");

[0145] The reply generation unit 106 generates a reply text;

[0146] The output unit 107 broadcasts "The yogurt of XX brand is on Shelf C05 in the dairy section on the third floor" through the speaker, and may also display a map to the location on the screen.

[0147] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high-bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by an application, module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device. Any application or module described in the present invention can be implemented using computer-readable / executable instructions stored or otherwise held by such a computer-readable medium.

Claims

1. An intelligent shopping guide speech recognition method based on transfer learning, characterized in that the method Including: Obtain the customer's inquiry voice for automatic speech recognition to obtain the initial recognition text; Based on the pre-trained language model, perform natural language understanding on the initial recognition text to identify the customer's query intention and the entities in the query text; construct a query instruction according to the query intention and the entities and retrieve it in the supermarket commodity information database; based on the retrieved relevant information, generate a reply content and output it through voice or / and interface; Adjust the parameters of a pre-trained language model using manually annotated data through backpropagation and parameter updates; the parameter updates include an adaptive learning rate , there is: ; Among them represents the base learning rate; represents the batch loss value; represents the batch average loss before; represents the minimum integer; tanh represents the hyperbolic tangent function; represents the set scaling factor; The comprehensive adaptation index is positively correlated with the entity novelty score and negatively correlated with the prediction confidence; the reciprocal mean of the occurrence frequencies of all entities in the current training batch constitutes the entity novelty score; the arithmetic mean of the maximum prediction probability of all query intentions and the label probability of the entity in the current training batch constitutes the prediction confidence; Indicates the batch The comprehensive adaptation index, and the calculation formula is as follows: ; Among them, represents the entity novelty score of batch, and the calculation formula is: ; where represents the total number of entities contained in the batch; represents the frequency with which the entity appears during the training process; represents the natural logarithm function; Indicates the batch of the prediction confidence, calculated as follows: ; wherein represents the number of samples in the batch ; represents the average value of the number of tokens corresponding to the samples in the batch ; represents the maximum value of the intention prediction probability of the sample ; represents the entity label probability of the th token in the sample 2. The intelligent shopping guide speech recognition method based on transfer learning according to claim 1, characterized in that, Take the BERT-base-chinese model as the basic natural language understanding model, collect the artificial annotation data, and perform end-to-end supervised training based on the basic natural language understanding model to obtain the pre-trained language model.

3. The intelligent shopping guide speech recognition method based on transfer learning according to claim 2, characterized in that Take the artificially written simulated queries, supermarket commodity catalogs, promotional activity descriptions, and floor shelf layout information as the original text, and use the BIO annotation system for the original text to obtain the artificial annotation data.

4. The intelligent shopping guide speech recognition method based on transfer learning according to claim 1, wherein It also includes using a denoising model based on DNN to preprocess the customer's inquiry voice.

5. An intelligent shopping guide speech recognition system based on transfer learning, characterized in that, The system includes: A voice input unit: used to receive the customer's voice command; A speech recognition unit: connected to the voice input unit, used to convert the voice command into text; A natural language understanding unit: connected to the speech recognition unit, processes the text based on the pre-trained language model to identify the customer's query intention and key entities; the pre-trained language model included in the natural language understanding unit is trained and configured according to the fine-tuning step in a method for intelligent shopping guide speech recognition based on transfer learning according to any one of claims 1 to 4; An information processing and query unit: connected to the natural language understanding unit and the commodity information database, used to construct database queries, execute queries, and retrieve information; a commodity information database: stores detailed information of all commodities in the supermarket, including real-time prices, promotional information, and accurate shelf location information; a reply generation unit: connected to the information processing and query unit, used to generate an anthropomorphic reply content according to the retrieval result; an output unit: connected to the reply generation unit, used to output the reply content through voice or / and interface.

6. The intelligent shopping guide speech recognition system based on transfer learning according to claim 5, characterized in that The output unit includes at least one of the following: A speech synthesizer and a speaker, used to synthesize the reply content into voice and broadcast it; A display screen, used to display the text of the reply content, the associated image, or the navigation path generated according to the commodity location information.

7. An intelligent shopping guide voice recognition system based on transfer learning according to claim 5, characterized in that, The voice input unit includes at least one microphone array for receiving the customer's voice command.

Citation Information

Patent Citations

  • Human body behavior recognition method and device based on skeleton points and storage medium

    CN117315770A

  • Method and system for matching domestic sellers and overseas buyers using artificial intelligence

    KR102726001B1