User intention recognition method, device and equipment

By using the intention identification model in the intelligent outbound call system to generate and query the user's delivery intention tags, the problem of inability to effectively accumulate and query user's intentions in the prior art is solved, and the efficiency of express delivery is improved.

CN120260574APending Publication Date: 2025-07-04SHANGHAI DONGPU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510440674.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing intelligent outbound call system cannot effectively accumulate and query the user's delivery intention tags in the express delivery industry, resulting in inefficiency of couriers during multiple outbound call times, increasing user inconvenience.

Method used

By obtaining the voice data between the intelligent outgoing call module and the user in the logistics distribution system, the voice data is converted into recognition text using the intent recognition model, and an intent tag is generated, which is stored in the tag precipitation library to quickly query the intent tag of the target user.

Benefits of technology

Couriers can quickly understand users' delivery preferences, reduce duplicate outgoing calls, and improve delivery efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260574A_ABST
    Figure CN120260574A_ABST
Patent Text Reader

Abstract

The invention discloses a user intention recognition method, device and equipment, and the method is applied to the technical field of logistics, and the method comprises the steps: obtaining voice data of interaction between an intelligent call-out module and a user in a logistics distribution system; inputting the voice data into an intention recognition model, and converting the voice data into a recognition text based on a voice recognition network in the intention recognition model; performing intention recognition on the recognition text based on an intention recognition network in the intention recognition model to obtain an intention label; storing a corresponding relationship between the user identifier of the user and the intention label in a label precipitation library; and in response to an intention query request of a target user identifier, searching a target intention tag matched with the target user identifier in the tag precipitation library. According to the invention, the distribution preference of the user can be quickly known, and the express distribution efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, device and equipment for user intention recognition. Background Art

[0002] In the express delivery industry, couriers usually need to contact users by phone to confirm the delivery time and location. The traditional phone contact method is inefficient and cannot effectively precipitate the delivery preference information of users. With the development of intelligent outbound calling technology, express delivery companies have started to use intelligent outbound calling systems to interact with users and obtain users' delivery intentions. However, the existing intelligent outbound calling systems lack the function of precipitating and querying user intention tags, resulting in couriers being unable to effectively utilize historical data during multiple outbound calls, increasing the redundancy of outbound calls and the inconvenience of users. Summary of the Invention

[0003] The present invention provides a method, device and equipment for user intention recognition, which can generate user intention tags through a machine learning model, enabling couriers to quickly understand users' delivery preferences, reducing repeated outbound calls, and improving delivery efficiency.

[0004] On the one hand, the present invention provides a method for user intention recognition, the method comprising: Obtaining voice data of the interaction between the intelligent outbound calling module and the user in the logistics distribution system; Inputting the voice data into an intention recognition model, and converting the voice data into recognition text based on the voice recognition network in the intention recognition model; Performing intention recognition on the recognition text based on the intention recognition network in the intention recognition model to obtain intention tags; Storing the corresponding relationship between the user identification of the user and the intention tags in a tag precipitation library; Responding to an intention query request of a target user identification, and searching for target intention tags matching the target user identification in the tag precipitation library.

[0005] In an exemplary embodiment, the training method of the intention recognition model comprises: Obtaining sample voice data of the interaction between a sample user and the intelligent outbound calling module; the sample voice data is labeled with sample intention tags; the sample intention tags represent at least one of the express delivery time, delivery location, and delivery method intention of the sample user; Inputting the sample voice data into a machine learning model, and converting the sample voice data into sample recognition text based on the preset voice recognition network in the machine learning model; Performing intent recognition on the sample recognition text based on a preset intent recognition network in the machine learning model to obtain a sample user intent result; Training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model.

[0006] In an exemplary embodiment, after training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model, the method further includes: Obtaining historical voice data of interactions between historical users and the intelligent outbound call module; Inputting the historical voice data into the intent recognition model for intent recognition processing to obtain historical intent labels of the historical users; Obtaining feedback results of the historical users for the historical intent labels, and correcting the historical intent labels according to the feedback results to obtain updated intent labels; Updating and training the intent recognition model based on the historical voice data and the updated intent labels to obtain an updated intent recognition model.

[0007] In an exemplary embodiment, training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model includes: Determining target loss data based on the difference between the sample user intent result and the sample intent label; Adjusting the parameters of the machine learning model based on the target loss data to obtain an initial intent recognition model; Obtaining verification voice data of interactions between verification users and the intelligent outbound call module; the verification voice data is labeled with verification intent labels; Inputting the verification voice data into the initial intent recognition model for intent recognition processing to obtain verification intent prediction results; If the verification intent prediction results match the verification intent labels, determining the initial intent recognition model as the intent recognition model.

[0008] In an exemplary embodiment, before responding to an intent query request of a target user identifier and searching for a target intent label matching the target user identifier in the label precipitation library, the method further includes: Obtaining the sources and types of each intent label in the label precipitation library, and determining character labels corresponding to each intent label; Store the correspondence between each intent tag and character tag in the tag precipitation library; one character tag corresponds to at least one intent tag; Construct a sub-database according to the character tag, and store at least one intent tag corresponding to the character tag and the user identifier corresponding to each intent tag in the sub-database; Determine the user identifier corresponding to each character tag according to the correspondence between the user identifier and the intent tag, and construct the label identifier association relationship between the user identifier and the character tag; Correspondingly, the step of, in response to an intent query request of a target user identifier, searching for a target intent tag matching the target user identifier in the tag precipitation library includes: Query the target label identifier corresponding to the target user identifier based on the label identifier association relationship; In the sub-database corresponding to the target label identifier, search for the target intent tag corresponding to the target user identifier.

[0009] In an exemplary embodiment, the step of obtaining the voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system includes: In response to a user call request, call the intelligent outbound call module in the logistics distribution system to obtain the user voice; Convert the user voice into text data, and input the text data into a question-and-answer model to obtain a reply text corresponding to the text data; Use neural voice synthesis technology to convert the reply text into spectral features, and use the WaveGlow algorithm to synthesize the spectral features into waveform audio; Call the audio device to play the waveform audio; Generate the voice data of the interaction between the intelligent outbound call module and the user according to the user voice and the waveform audio of multiple interactions.

[0010] In an exemplary embodiment, the step of calling the audio device to play the waveform audio includes: Parse the user voice to determine the user's hearing level and the noise information of the environment where the user is located; Determine the playback volume of the waveform audio according to the user's hearing level and the noise information; Call the audio device to play the waveform audio according to the playback volume.

[0011] On the other hand, a user intent recognition device is provided, and the device includes: A voice data acquisition module, configured to acquire voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system; A text recognition module, configured to input the speech data into an intent recognition model, and convert the speech data into recognized text based on a speech recognition network in the intent recognition model; An intent label recognition module, configured to perform intent recognition on the recognized text based on an intent recognition network in the intent recognition model to obtain an intent label; A relationship storage module, configured to store a correspondence relationship between a user identifier of the user and the intent label in a label precipitation library; A target intent search module, configured to, in response to an intent query request of a target user identifier, search for a target intent label that matches the target user identifier in the label precipitation library.

[0012] In an exemplary embodiment, the device further includes: A sample speech acquisition module, configured to acquire sample speech data of an interaction between a sample user and the intelligent outbound call module; the sample speech data is labeled with a sample intent label; the sample intent label represents at least one of an express delivery time, a delivery location, and a delivery method intent of the sample user; A sample text conversion module, configured to input the sample speech data into a machine learning model, and convert the sample speech data into sample recognized text based on a preset speech recognition network in the machine learning model; A sample intent determination module, configured to perform intent recognition on the sample recognized text based on a preset intent recognition network in the machine learning model to obtain a sample user intent result; An intent recognition model training module, configured to train the machine learning model based on a difference between the sample user intent result and the sample intent label to obtain the intent recognition model.

[0013] In an exemplary embodiment, the device further includes: A historical data acquisition module, configured to acquire historical speech data of an interaction between a historical user and the intelligent outbound call module; A historical intent determination module, configured to input the historical speech data into the intent recognition model for intent recognition processing to obtain a historical intent label of the historical user; An intent label update module, configured to obtain a feedback result of the historical user for the historical intent label, and correct the historical intent label according to the feedback result to obtain an updated intent label; An update training module, configured to perform update training on the intent recognition model based on the historical speech data and the updated intent label to obtain an updated intent recognition model.

[0014] In an exemplary embodiment, the intention recognition model training module includes: A target loss determination unit, configured to determine target loss data based on the difference between the sample user intention result and the sample intention label; An initial model determination unit, configured to adjust the parameters of the machine learning model based on the target loss data to obtain an initial intention recognition model; A verification data acquisition unit, configured to acquire verification voice data of the interaction between the verification user and the intelligent outbound call module; the verification voice data is labeled with a verification intention label; A verification result determination unit, configured to input the verification voice data into the initial intention recognition model for intention recognition processing to obtain a verification intention prediction result; A model determination unit, configured to determine the initial intention recognition model as the intention recognition model if the verification intention prediction result matches the verification intention label.

[0015] In an exemplary embodiment, the device further includes: A character label determination module, configured to obtain the source and type of each intention label in the label precipitation library, and determine the character label corresponding to each intention label; A relationship storage module, configured to store the corresponding relationship between each intention label and the character label in the label precipitation library; one character label corresponds to at least one intention label; An identification storage module, configured to construct a sub-database according to the character label, and store at least one intention label corresponding to the character label and the user identification corresponding to each intention label in the sub-database; A relationship construction module, configured to determine the user identification corresponding to each character label according to the corresponding relationship between the user identification and the intention label, and construct a label identification association relationship between the user identification and the character label; Correspondingly, the target intention search module is further configured to: Query the target label identification corresponding to the target user identification based on the label identification association relationship; In the sub-database corresponding to the target label identification, search for the target intention label corresponding to the target user identification.

[0016] In an exemplary embodiment, the voice data acquisition module includes: A user voice acquisition unit, configured to call the intelligent outbound call module in the logistics distribution system to acquire user voice in response to a user call request; A reply text determination unit, configured to convert the user voice into text data, and input the text data into a question and answer model to obtain a reply text corresponding to the text data; An audio synthesis unit, configured to convert the reply text into spectral features by using neural voice synthesis technology, and synthesize waveform audio from the spectral features by using the WaveGlow algorithm; An audio playback unit, configured to call an audio device to play the waveform audio; A voice data generation unit, configured to generate voice data for interaction between the intelligent outbound call module and the user according to the user voice and the waveform audio of multiple interactions.

[0017] In an exemplary embodiment, the audio playback unit includes: A parsing subunit, configured to parse the user voice to determine the user's hearing level and the noise information of the environment where the user is located; A volume determination subunit, configured to determine the playback volume of the waveform audio according to the user's hearing level and the noise information; A playback subunit, configured to call an audio device to play the waveform audio according to the playback volume.

[0018] On the other hand, an electronic device is provided, where the device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the user intention recognition method as described above.

[0019] On the other hand, a computer storage medium is provided, where the computer storage medium stores at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the user intention recognition method as described above.

[0020] On the other hand, a computer program product or a computer program is provided, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the user intention recognition method as described above.

[0021] The user intention recognition method, device and equipment provided by the present invention have the following technical effects: The present invention obtains voice data for interaction between an intelligent outbound call module and a user in a logistics distribution system; inputs the voice data into an intent recognition model, and converts the voice data into recognized text based on a speech recognition network in the intent recognition model; performs intent recognition on the recognized text based on an intent recognition network in the intent recognition model to obtain intent tags; stores the correspondence between the user identification of the user and the intent tags in a tag precipitation library; and in response to an intent query request of a target user identification, searches for a target intent tag matching the target user identification in the tag precipitation library. By generating user intent tags through a machine learning model, couriers can quickly understand the delivery preferences of users, reduce repeated outbound calls, and improve delivery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions and advantages in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 is a schematic diagram of a user intent recognition system provided by an embodiment of this specification; Figure 2 is a flowchart of a user intent recognition method provided by an embodiment of this specification; Figure 3 is a flowchart of a method for obtaining voice data for interaction between an intelligent outbound call module and a user in a logistics distribution system provided by an embodiment of this specification; Figure 4 is a flowchart of a method for calling an audio device to play the waveform audio provided by an embodiment of this specification; Figure 5 is a flowchart of a training method for an intent recognition model provided by an embodiment of this specification; Figure 6 is a schematic diagram of the structure of a user intent recognition device provided by an embodiment of this specification; Figure 7 is a schematic diagram of the structure of a server provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the specification, claims and accompanying drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a user intention recognition system provided by an embodiment of this specification. As Figure 1 shown, the user intention recognition system may at least include a server 01 and a client 02.

[0027] Specifically, in the embodiments of this specification, the server 01 may include an independently operating server, a distributed server, or a server cluster composed of multiple servers, and may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 01 may include a network communication unit, a processor, a memory, and so on. Specifically, the server 01 may be used to train an intention recognition model.

[0028] Specifically, in the embodiments of this specification, the client 02 may include physical devices such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, smart wearable devices, smart speakers, vehicle terminals, smart TVs, etc., and may also include software running on the physical devices, such as web pages provided by some service providers to users, or applications provided by these service providers to users. Specifically, the client 02 may be used to online query the target intention label of the target user.

[0029] The following introduces a method for identifying user intentions of the present invention. Figure 2 It is a schematic flowchart of a method for identifying user intentions provided by an embodiment of this specification. This specification provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiment or the accompanying drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method may include: S201: Obtain the voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system. In the embodiment of this specification, the intelligent outbound call module is used to conduct voice interaction with the user, so as to facilitate obtaining the user's distribution intention.

[0030] S203: Input the voice data into the intention recognition model, and convert the voice data into recognition text based on the speech recognition network in the intention recognition model.

[0031] In the embodiment of this specification, the intention recognition model corresponds to the machine learning model module, which is used to perform natural language processing (NLP) and intention recognition on the user's voice data to generate user intention tags. Speech recognition network: Use a deep learning model (such as a speech recognition model based on Transformer) to convert the user's voice data into text.

[0032] S205: Perform intention recognition on the recognition text based on the intention recognition network in the intention recognition model to obtain intention tags.

[0033] In the embodiment of this specification, the intention tag is the user's intention distribution tag, which may include at least one of the user's express delivery time intention, delivery location intention, and delivery method intention; Intention recognition network: Use a natural language processing model based on BERT to classify the text and generate user intention tags. The intention tags include express delivery attribute tags such as "delivery on weekdays", "delivery now", "call before delivery", "door-to-door delivery", "express delivery cabinet", and "pick-up point".

[0034] S207: Store the corresponding relationship between the user identification of the user and the intention tag in the tag precipitation library.

[0035] In the embodiment of this specification, the tag precipitation module: is used to store the user intention tags generated by the machine learning model in the tag precipitation library.

[0036] S209: In response to the intent query request of the target user identifier, search for the target intent label matching the target user identifier in the label precipitation library.

[0037] In the embodiments of this specification, the intent query request of the target user identifier may be a batch request, that is, a request to obtain the target intent labels of multiple target user identifiers. An efficient batch query interface can be provided to support the courier to query the intent labels of multiple users at one time, improving the operation efficiency. The logistics system may include a query interface module: providing a batch query interface to support batch query of user intent labels according to user IDs. The maximum number of user IDs supported for a single query is 50. The query data source comes from the label precipitation library, and it supports querying data for the most recent 1 month.

[0038] In the embodiments of this specification, as Figure 3 shown, the obtaining of the voice data for interaction between the intelligent outbound call module and the user in the logistics distribution system includes: S20101: In response to the user call request, call the intelligent outbound call module in the logistics distribution system to obtain the user voice; S20103: Convert the user voice into text data, and input the text data into the question-and-answer model to obtain the reply text corresponding to the text data; S20105: Use neural voice synthesis technology to convert the reply text into spectral features, and use the WaveGlow algorithm to synthesize the spectral features into waveform audio; S20107: Call the audio device to play the waveform audio; S20109: Generate the voice data for interaction between the intelligent outbound call module and the user according to the user voice and the waveform audio of multiple interactions.

[0039] In the embodiments of this specification, the intelligent outbound call module is the front-end entry for the system to interact with the user. Its core goal is to conduct efficient and natural voice interactions with the user to accurately obtain the user's delivery intent. This module consists of multiple sub-components working together, including voice synthesis, voice playback, voice collection, and dialogue management, etc. The voice synthesis sub-component is responsible for converting the information to be conveyed by the system into natural and fluent speech. To achieve a highly natural voice effect, advanced neural voice synthesis technology is adopted, such as the architecture combining Tacotron 2 and WaveGlow. Tacotron 2 converts text into spectral features, and WaveGlow then synthesizes the spectral features into waveform audio. This technology can generate speech close to human pronunciation, greatly improving the user's interaction experience. The voice playback sub-component uses professional audio devices and audio processing algorithms to ensure that the voice is clearly and moderately volume-conveyed to the user.

[0040] In the embodiments of this specification, the conversion of the user voice into text data includes: Adopt adaptive filtering and spectral subtraction to filter out the noise information in the user voice, obtain the denoised voice, and convert the denoised voice into text data.

[0041] In the embodiments of this specification, in a complex ambient noise environment, through noise reduction technologies such as adaptive filtering and spectral subtraction, background noise is effectively removed, and the quality of the voice signal is improved.

[0042] In the embodiments of this specification, as Figure 4 shown, the calling of the audio device to play the waveform audio includes: S401: Parse the user voice to determine the hearing level of the user and the noise information of the environment where the user is located; S403: Determine the playback volume of the waveform audio according to the hearing level of the user and the noise information; S405: Call the audio device to play the waveform audio according to the playback volume.

[0043] In the embodiments of this specification, considering the hearing differences of different users and ambient noise, the system supports volume adjustment and voice clarity optimization. The better the hearing, the higher the hearing level. If the hearing level is less than the preset level and the volume of the noise information is greater than the preset volume, the playback volume is set to be greater than the preset volume; the corresponding relationship between the hearing level, the noise volume and the playback volume can also be constructed; for the first hearing level with poor hearing and the first noise volume greater than the first preset volume, the playback volume is set to the highest first volume; for the second hearing level with medium hearing and the second noise volume less than the first preset volume and greater than the second preset volume, the playback volume is set to the medium second volume; for the third hearing level with good hearing and the third noise volume less than the second preset volume, the playback volume is set to the lowest third volume.

[0044] The voice acquisition sub-component uses a high-sensitivity microphone and an advanced audio noise reduction algorithm to accurately acquire the voice data of the user. The dialogue management sub-component is responsible for controlling the process and logic of the dialogue. It guides the user to gradually express the delivery intention according to the preset dialogue strategy. For example, when the user does not clearly express the intention, the system will further obtain information through means such as asking questions and giving prompts. At the same time, the dialogue management sub-component also has the ability to handle exceptions and can handle the user's non-standard expressions and unexpected situations.

[0045] In the embodiments of this specification, as Figure 5 shown, the training method of the intention recognition model includes: S501: Acquire sample voice data of interaction between a sample user and the intelligent outbound call module; the sample voice data is annotated with a sample intention label; the sample intention label represents at least one of the sample user's intention of express delivery time, delivery location, and delivery method; S503: Inputting the sample voice data into a machine learning model, and converting the sample voice data into sample recognition text based on a preset voice recognition network in the machine learning model; S505: Performing intent recognition on the sample recognition text based on a preset intent recognition network in the machine learning model to obtain a sample user intent result; S507: Based on the difference between the sample user intention result and the sample intention label, the machine learning model is trained to obtain the intention recognition model.

[0046] In the embodiments of the present specification, the sample intent label represents at least one of the sample user's intentions of express delivery time, delivery location, and delivery method; illustratively, the sample intent label may include "delivery on working days", "delivery now", "call before delivery", "door-to-door delivery", "express locker", "collection point", etc.

[0047] In the embodiment of this specification, the training of the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model includes: Determining target loss data based on a difference between the sample user intent result and the sample intent label; The parameters of the machine learning model are adjusted based on the target loss data to obtain an intent recognition model.

[0048] In the embodiment of the present specification, after the machine learning model is trained based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model, the method further includes: Acquire historical voice data of interactions between historical users and the intelligent outbound call module; Inputting the historical voice data into the intention recognition model for intention recognition processing to obtain the historical intention label of the historical user; Obtaining feedback results from the historical users regarding the historical intent labels, and modifying the historical intent labels according to the feedback results to obtain updated intent labels; The intention recognition model is updated and trained based on the historical voice data and the updated intention label to obtain an updated intention recognition model.

[0049] In the embodiments of this specification, the model is continuously trained and optimized using historical speech data and user feedback data to improve the accuracy of intent recognition.

[0050] The machine learning model module is the intelligent core of the system, mainly responsible for performing natural language processing (NLP) and intent recognition on the user's speech data to generate accurate user intent labels. It consists of three important parts: a speech recognition model, an intent recognition model, and model training and optimization. Speech Recognition Model A speech recognition model based on Transformer, such as Wav2Vec 2.0, is adopted. This model learns speech features from a large amount of speech data through unsupervised learning and then fine-tunes on a supervised dataset to achieve high-precision speech recognition. Wav2Vec 2.0 has strong context modeling capabilities and can capture long-range dependencies in speech signals, thus more accurately converting speech data into text. In practical applications, model quantization and pruning techniques can also be used to optimize the model to improve recognition speed and efficiency. Intent Recognition Model A natural language processing model based on BERT (Bidirectional Encoder Representations from Transformers) is used to classify the intent of the text obtained from speech recognition. The BERT model learns rich language knowledge and semantic information through pre-training and can better understand the meaning of the text. In the intent classification task, the BERT model is fine-tuned to adapt to specific delivery intent classification scenarios. By adding a fully connected layer and a Softmax activation function to the output layer of the model, the text is mapped to different intent labels. The intent labels cover common delivery requirements such as "delivery on weekdays", "delivery now", "call before delivery", "door-to-door delivery", "express cabinet", "pick-up point", etc. Model Training and Optimization To continuously improve the accuracy of intent recognition, the system continuously trains and optimizes the model using historical speech data and user feedback data. Stochastic Gradient Descent (SGD) and its variant algorithms (such as Adam, Adagrad, etc.) are used to update the model parameters.

[0051] In the embodiments of this specification, training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model includes: Determining target loss data based on the difference between the sample user intent result and the sample intent label; Adjusting the parameters of the machine learning model based on the target loss data to obtain an initial intent recognition model; Obtain verification voice data for the interaction between the verified user and the intelligent outbound call module; the verification voice data is labeled with a verification intent label; Input the verification voice data into the initial intent recognition model for intent recognition processing to obtain a verification intent prediction result; If the verification intent prediction result matches the verification intent label, determine that the initial intent recognition model is the intent recognition model.

[0052] In the embodiments of this specification, if the verification intent prediction result does not match the verification intent label, continue to train the initial intent recognition model and continue to repeatedly verify the trained model until the verification result meets the target condition, where the target condition may be that the verification intent prediction result matches the verification intent label. During the training process, cross-validation and early stopping strategies are used to prevent the model from overfitting. At the same time, a reinforcement learning mechanism is introduced to reward or punish the model according to the actual feedback of the user and the business effect, further optimizing the performance of the model.

[0053] In the embodiments of this specification, before responding to the intent query request of the target user identifier and searching for the target intent label matching the target user identifier in the label precipitation library, the method further includes: Obtain the source and type of each intent label in the label precipitation library, and determine the character label corresponding to each intent label; Store the corresponding relationship between each intent label and the character label in the label precipitation library; one character label corresponds to at least one intent label; Construct a sub-database according to the character label, and store at least one intent label corresponding to the character label and the user identifier corresponding to each intent label in the sub-database; According to the corresponding relationship between the user identifier and the intent label, determine the user identifier corresponding to each character label, and construct the label identification association relationship between the user identifier and the character label.

[0054] Among them, the label precipitation module is responsible for storing and managing the user intention labels generated by the machine learning model. It pushes the generated labels to the delivery summary table and stores them in the label precipitation library. The label precipitation library uses a distributed database (such as Cassandra or HBase) to ensure high availability and scalability. The character label can be tagSourceId, and each intention label is associated with a specific tagSourceId, which is used to identify the source and type of the label. For example, "delivery on weekdays" may be associated with tagSourceId 999901, and "delivery now" is associated with tagSourceId 120101, etc. These tagSourceIds help the system classify and manage the labels, and also facilitate subsequent data analysis and statistics. The data is retained in the label precipitation library for a preset period (such as 3 months) to meet the business's needs for historical data analysis and traceability. To ensure data security and integrity, the data is stored encrypted, and data backup and recovery tests are performed regularly.

[0055] Correspondingly, the step of searching for a target intention label matching the target user identifier in the label precipitation library in response to an intention query request for the target user identifier includes: Querying a target label identifier corresponding to the target user identifier based on the label identifier association relationship; In the sub-database corresponding to the target label identifier, searching for the target intention label corresponding to the target user identifier.

[0056] In the embodiments of this specification, the query interface module provides a function for users such as couriers to batch query user intention labels. It supports batch querying user intention labels according to user IDs, and supports querying up to 50 user IDs at a time. The query data source comes from the label precipitation library and supports querying data for the most recent 1 month. This module adopts the RESTful API design style to improve the scalability and usability of the interface. After receiving a query request, the interface module will perform a legality check on the request, including format verification of the user ID and limit check of the query quantity. Then, it queries the corresponding intention labels from the label precipitation library according to the user ID and returns the results to the client in JSON format. To improve query efficiency, an index is established for the label precipitation library, and a caching technology (such as Redis) is used to cache frequently queried data.

[0057] In the embodiments of this specification, during the intelligent outbound call process, the system converts the user's speech into text through a speech recognition model and generates corresponding intention labels through an intention recognition model.

[0058] The intent tags can include "delivery on weekdays" and "delivery now", and each tag is associated with a specific tagSourceId (such as 999901, 120101, 110101). The tag precipitation module pushes the generated tags to the delivery summary table and stores them in the tag precipitation library, and the data is retained for 3 months.

[0059] In the embodiments of this specification, the courier initiates a query request through the xx logistics system and inputs multiple user IDs.

[0060] The query interface module batch queries the intent tags of users from the tag precipitation library according to the user IDs.

[0061] The returned results include the user's delivery preference tags (such as "delivery on weekdays", "delivery now", etc.) and the update time of the tags.

[0062] As can be seen from the technical solutions provided by the embodiments of this specification above, the embodiments of this specification obtain the voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system; input the voice data into the intent recognition model, and convert the voice data into recognition text based on the voice recognition network in the intent recognition model; perform intent recognition on the recognition text based on the intent recognition network in the intent recognition model to obtain intent tags; store the correspondence between the user identification of the user and the intent tags in the tag precipitation library; in response to the intent query request of the target user identification, search for the target intent tags that match the target user identification in the tag precipitation library. By generating user intent tags through a machine learning model, the courier can quickly understand the user's delivery preferences, reduce repeated outbound calls, and improve the delivery efficiency.

[0063] The embodiments of this specification also provide a user intent recognition device, as Figure 6 shown, the device includes: A voice data acquisition module 610, configured to acquire the voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system; A text recognition module 620, configured to input the voice data into the intent recognition model and convert the voice data into recognition text based on the voice recognition network in the intent recognition model; An intent tag recognition module 630, configured to perform intent recognition on the recognition text based on the intent recognition network in the intent recognition model to obtain intent tags; A relationship storage module 640, configured to store the correspondence between the user identification of the user and the intent tags in the tag precipitation library; A target intent search module 650, configured to search for the target intent tags that match the target user identification in the tag precipitation library in response to the intent query request of the target user identification.

[0064] In an exemplary embodiment, the device further includes: A sample voice acquisition module, configured to acquire sample voice data of an interaction between a sample user and the intelligent outbound call module; the sample voice data is labeled with sample intent tags; the sample intent tags represent at least one of the sample user's express delivery time, delivery location, and delivery method intent; A sample text conversion module, configured to input the sample voice data into a machine learning model, and convert the sample voice data into sample recognition text based on a preset speech recognition network in the machine learning model; A sample intent determination module, configured to perform intent recognition on the sample recognition text based on a preset intent recognition network in the machine learning model to obtain a sample user intent result; An intent recognition model training module, configured to train the machine learning model based on the difference between the sample user intent result and the sample intent tags to obtain the intent recognition model.

[0065] In an exemplary embodiment, the device further includes: A historical data acquisition module, configured to acquire historical voice data of an interaction between a historical user and the intelligent outbound call module; A historical intent determination module, configured to input the historical voice data into the intent recognition model for intent recognition processing to obtain historical intent tags of the historical user; An intent tag update module, configured to acquire a feedback result of the historical user for the historical intent tags, and correct the historical intent tags according to the feedback result to obtain updated intent tags; An update training module, configured to perform update training on the intent recognition model based on the historical voice data and the updated intent tags to obtain an updated intent recognition model.

[0066] In an exemplary embodiment, the intent recognition model training module includes: A target loss determination unit, configured to determine target loss data based on the difference between the sample user intent result and the sample intent tags; An initial model determination unit, configured to adjust parameters of the machine learning model based on the target loss data to obtain an initial intent recognition model; A verification data acquisition unit, configured to acquire verification voice data of an interaction between a verification user and the intelligent outbound call module; the verification voice data is labeled with verification intent tags; A verification result determination unit, configured to input the verification voice data into the initial intent recognition model for intent recognition processing to obtain a verification intent prediction result; A model determination unit, configured to determine the initial intent recognition model as the intent recognition model if the verification intent prediction result matches the verification intent label.

[0067] In an exemplary embodiment, the apparatus further includes: A character label determination module, configured to obtain the source and type of each intent label in the label precipitation library, and determine the character label corresponding to each intent label; A relationship storage module, configured to store the corresponding relationship between each intent label and the character label in the label precipitation library; one character label corresponds to at least one intent label; An identification storage module, configured to construct a sub-database according to the character label, and store at least one intent label corresponding to the character label and the user identification corresponding to each intent label in the sub-database; A relationship construction module, configured to determine the user identification corresponding to each character label according to the corresponding relationship between the user identification and the intent label, and construct a label identification association relationship between the user identification and the character label; Correspondingly, the target intent search module is further configured to: Query the target label identification corresponding to the target user identification based on the label identification association relationship; In the sub-database corresponding to the target label identification, search for the target intent label corresponding to the target user identification.

[0068] In an exemplary embodiment, the voice data acquisition module includes: A user voice acquisition unit, configured to call the intelligent outbound call module in the logistics distribution system to obtain user voice in response to a user call request; A reply text determination unit, configured to convert the user voice into text data, and input the text data into a question-and-answer model to obtain a reply text corresponding to the text data; An audio synthesis unit, configured to convert the reply text into spectral features by using neural voice synthesis technology, and synthesize the spectral features into waveform audio by using the WaveGlow algorithm; An audio playback unit, configured to call an audio device to play the waveform audio; A voice data generation unit, configured to generate voice data for interaction between the intelligent outbound call module and the user according to the user voice and the waveform audio of multiple interactions.

[0069] In an exemplary embodiment, the audio playback unit includes: A parsing subunit, configured to parse the user voice to determine the hearing level of the user and the noise information of the environment where the user is located; A volume determination subunit, configured to determine the playback volume of the waveform audio according to the hearing level of the user and the noise information; A playback subunit, configured to call an audio device to play the waveform audio according to the playback volume.

[0070] The device in the device embodiment and the method embodiment are based on the same inventive concept.

[0071] An embodiment of this specification provides an electronic device, which includes a processor and a memory. At least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or at least one segment of program is loaded and executed by the processor to implement the user intention recognition method provided in the above method embodiment.

[0072] An embodiment of the present invention further provides a computer storage medium, which can be set in a terminal to store at least one instruction or at least one segment of program related to a user intention recognition method in a method embodiment. The at least one instruction or at least one segment of program is loaded and executed by the processor to implement the user intention recognition method provided in the above method embodiment.

[0073] An embodiment of the present invention further provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the user intention recognition method provided in the above method embodiment.

[0074] Optionally, in an embodiment of this specification, the storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as a USB flash drive, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk, or an optical disc.

[0075] The memory described in the embodiments of this specification can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0076] The method embodiments for identifying user intentions provided in the embodiments of this specification can be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 7 is a hardware structure block diagram of a server for the method of identifying user intentions provided in the embodiments of this specification. As Figure 7 shown, the server 700 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 710 (the central processing unit 710 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 730 for storing data, and one or more storage media 720 for storing application programs 723 or data 722 (such as one or more mass storage devices). Among them, the memory 730 and the storage media 720 can be transient storage or persistent storage. The programs stored in the storage media 720 can include one or more modules, and each module can include a series of instruction operations on the server. Further, the central processing unit 710 can be configured to communicate with the storage media 720 and execute a series of instruction operations in the storage media 720 on the server 700. The server 700 can also include one or more power supplies 760, one or more wired or wireless network interfaces 750, one or more input / output interfaces 740, and / or one or more operating systems 721, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0077] The input / output interface 740 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server 700. In one example, the input / output interface 740 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the input / output interface 740 can be a RadioFrequency (RF) module, which is used to communicate with the Internet wirelessly.

[0078] Those of ordinary skill in the art can understand that Figure 7 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the server 700 may also include more or fewer components than those shown Figure 7 in the figure, or have a different configuration from that shown Figure 7 in the figure.

[0079] As can be seen from the embodiments of the user intention recognition method, device, electronic device or storage medium provided by the present invention above, the present invention obtains voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system; inputs the voice data into an intention recognition model, and converts the voice data into recognition text based on the speech recognition network in the intention recognition model; performs intention recognition on the recognition text based on the intention recognition network in the intention recognition model to obtain intention labels; stores the corresponding relationship between the user identification of the user and the intention labels in a label precipitation library; in response to an intention query request of a target user identification, searches for a target intention label matching the target user identification in the label precipitation library. By generating user intention labels through a machine learning model, couriers can quickly understand the delivery preferences of users, reduce repeated outbound calls, and improve delivery efficiency.

[0080] It should be noted that: the above sequence of the embodiments of this specification is only for description and does not represent the superiority or inferiority of the embodiments. And the above-mentioned specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order from that in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0081] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of devices, equipment, and storage media, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0082] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer storage medium, and the storage medium mentioned above can be a read-only memory, a disk, an optical disc, or the like.

[0083] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A user intention recognition method, characterized in that, The method includes: Obtaining voice data of the interaction between the intelligent outbound call module and the user in the logistics distribution system; Inputting the voice data into an intent recognition model, and converting the voice data into a recognized text based on the speech recognition network in the intent recognition model; Performing intent recognition on the recognized text based on the intent recognition network in the intent recognition model to obtain an intent label; Storing the correspondence between the user identification of the user and the intent label in a label precipitation library; In response to an intent query request of a target user identification, searching for a target intent label matching the target user identification in the label precipitation library.

2. The method according to claim 1, wherein The training method of the intent recognition model includes: Obtaining sample voice data of the interaction between a sample user and the intelligent outbound call module; the sample voice data is labeled with a sample intent label; the sample intent label represents at least one of the intent of the delivery time, delivery location, and delivery method of the sample user; Inputting the sample voice data into a machine learning model, and converting the sample voice data into a sample recognized text based on a preset speech recognition network in the machine learning model; Performing intent recognition on the sample recognized text based on a preset intent recognition network in the machine learning model to obtain a sample user intent result; Training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model.

3. The method according to claim 2, wherein After training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model, the method further includes: Obtaining historical voice data of the interaction between a historical user and the intelligent outbound call module; Inputting the historical voice data into the intent recognition model for intent recognition processing to obtain a historical intent label of the historical user; Obtaining a feedback result of the historical user for the historical intent label, and correcting the historical intent label according to the feedback result to obtain an updated intent label; Updating and training the intent recognition model based on the historical voice data and the updated intent label to obtain an updated intent recognition model.

4. The method according to claim 2, wherein Training the machine learning model based on the difference between the sample user intent result and the sample intent label to obtain the intent recognition model, including: Determining target loss data based on the difference between the sample user intent result and the sample intent label; Adjusting the parameters of the machine learning model based on the target loss data to obtain an initial intent recognition model; Obtaining verification voice data of the interaction between a verification user and the intelligent outbound call module; the verification voice data is labeled with a verification intent label; Inputting the verification voice data into the initial intent recognition model for intent recognition processing to obtain a verification intent prediction result; If the verification intent prediction result matches the verification intent label, determining the initial intent recognition model as the intent recognition model.

5. The method according to claim 1, characterized in that, Before querying the target intent tag that matches the target user identifier in the tag precipitation library in response to the intent query request for the target user identifier, the method further includes: Obtaining the source and type of each intent tag in the tag precipitation library, and determining the character tag corresponding to each intent tag; Storing the corresponding relationship between each intent tag and the character tag in the tag precipitation library; one character tag corresponds to at least one intent tag; Constructing a sub-database according to the character tags, and storing at least one intent tag corresponding to the character tags and the user identifier corresponding to each intent tag in the sub-database; Determining the user identifier corresponding to each character tag according to the corresponding relationship between the user identifier and the intent tag, and constructing a tag identifier association relationship between the user identifier and the character tag; Correspondingly, the querying for the target intent tag that matches the target user identifier in the tag precipitation library in response to the intent query request for the target user identifier includes: Querying the target tag identifier corresponding to the target user identifier based on the tag identifier association relationship; In the sub-database corresponding to the target tag identifier, querying the target intent tag corresponding to the target user identifier.

6. The method according to claim 1, characterized in that The obtaining of the voice data for the interaction between the intelligent outbound call module and the user in the logistics distribution system includes: In response to the user call request, invoking the intelligent outbound call module in the logistics distribution system to obtain the user voice; Converting the user voice into text data, and inputting the text data into a question-and-answer model to obtain a reply text corresponding to the text data; Using neural voice synthesis technology to convert the reply text into spectral features, and using the WaveGlow algorithm to synthesize the spectral features into waveform audio; Invoking an audio device to play the waveform audio; Generating the voice data for the interaction between the intelligent outbound call module and the user according to the user voice and the waveform audio of multiple interactions.

7. The method according to claim 6, characterized in that The invoking of the audio device to play the waveform audio includes: Parsing the user voice to determine the hearing level of the user and the noise information of the environment where the user is located; Determining the playback volume of the waveform audio according to the hearing level of the user and the noise information; Invoking an audio device to play the waveform audio according to the playback volume.

8. A user intention recognition device, characterized in that, The device includes: A voice data acquisition module, configured to acquire voice data for the interaction between the intelligent outbound call module and the user in the logistics distribution system; A text recognition module, configured to input the voice data into an intent recognition model, and convert the voice data into recognition text based on the speech recognition network in the intent recognition model; An intent tag recognition module, configured to perform intent recognition on the recognition text based on the intent recognition network in the intent recognition model to obtain intent tags; A relationship storage module, configured to store the corresponding relationship between the user identifier of the user and the intent tags in a tag precipitation library; A target intent search module, configured to search for a target intent tag that matches the target user identifier in the tag precipitation library in response to an intent query request for the target user identifier.

9. An electronic device, characterized in that, The device includes: a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement the user intention recognition method according to any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer storage medium stores at least one instruction or at least one program segment, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement the user intention recognition method according to any one of claims 1-7.