Intention recognition system and multi-level hit response processing method

Through the multi-level hit response processing method of end-side devices working in concert with cloud servers, the problem that the existing intention recognition system cannot work properly without a network is solved, fast intention recognition and stable response are achieved, and user experience is improved.

CN120340474APending Publication Date: 2025-07-18HANGZHOU LINGBAN TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510838107.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing intention identification system is highly dependent on cloud servers, resulting in the inability to work properly without network status or network congestion, poor user experience, unstable data transmission, and long response time.

Method used

The multi-level hit response processing method is adopted that works in collaboration with the cloud server. Audio data conversion and intention recognition are performed through the end-side device, and the end-side and cloud intention recognition modules are processed in parallel under different network connection status to ensure fast response.

Benefits of technology

In the absence of a network or poor network, the end-side device can quickly provide intent identification information, reduce waiting time, improve user experience, and ensure stable system response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340474A_ABST
    Figure CN120340474A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an intention recognition system and a multi-level hit response processing method. A specific embodiment of the system comprises an end-side device and a cloud server, wherein the end-side device is configured to execute the following multi-level hit response processing: receiving user operation demand audio data, and performing identification conversion processing on the user operation demand audio data; in response to determining that the network connection state with the cloud server is an offline state, inputting the user operation demand text information to an end-side intention recognition module group to obtain intention recognition information; in response to determining that the network connection state with the cloud server is a connection state, inputting the user operation demand text information in parallel to an end-side intention recognition module group and a cloud intention recognition module deployed by the cloud server to obtain intention recognition information; response operation task information corresponding to the intention identification information is generated; and executing a response operation corresponding to the response operation task information. According to the embodiment, the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and specifically to an intent recognition system and a multi-level hit response processing method. Background Art

[0002] With the booming development of intelligent interaction technology, intent recognition systems have played a key role in many fields, such as providing various intent understanding and response services for users. An intent recognition system is a system used to recognize user intents and, based on the recognized user intents, provide corresponding services (i.e., response operations or information feedback) to users so that users can quickly obtain operation feedback information. Currently, when recognizing user intents, generating intent recognition information, and performing response operations based on the intent recognition information, the commonly adopted method is that the user-end device sends the voice data input by the user to the cloud server for cloud intent recognition, and only relies on the intent recognition information obtained by the cloud server to perform response operations and feedback operation feedback information.

[0003] However, when generating intent recognition information and performing response operations based on the intent recognition information in the above manner, there are often the following technical problems: Only relying on the cloud server for intent recognition, obtaining the intent recognition information, and performing response operations highly depends on the network environment. When in a network-free state, the user-end device cannot establish a connection with the cloud, and the system cannot work properly, resulting in the user being unable to use the intent recognition function and a poor user experience. At the same time, in the case of network congestion or poor signal, the data transmission speed between the user-end and the cloud becomes slower, and there may even be data loss or transmission errors, resulting in an unstable time-consuming for the entire process from the system receiving user input to returning the recognition result. Sometimes, it may take a long time to perform response operations and feedback operation feedback information, seriously affecting the user experience.

[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] This content part of the present disclosure is used to briefly introduce the concepts, which will be described in detail in the subsequent detailed implementation part. This content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure propose an intent recognition system and a multi-level hit response processing method to solve one or more of the technical problems mentioned in the above background art section.

[0007] In a first aspect, some embodiments of the present disclosure provide an intention recognition system, which includes: an edge device and a cloud server, wherein: the edge device is configured to perform the following multi-level hit response processing: receiving user operation requirement audio data, and performing recognition and conversion processing on the user operation requirement audio data to obtain user operation requirement text information; in response to determining that the network connection status with the cloud server is an offline state, inputting the user operation requirement text information into an edge intention recognition module group to obtain intention recognition information; in response to determining that the network connection status with the cloud server is a connected state, parallelly inputting the user operation requirement text information into the edge intention recognition module group and the cloud intention recognition module deployed on the cloud server to obtain intention recognition information; generating response operation task information corresponding to the intention recognition information; and performing a response operation corresponding to the response operation task information to obtain operation feedback information.

[0008] In a second aspect, some embodiments of the present disclosure provide a multi-level hit response processing method, which includes: receiving user operation requirement audio data, and performing recognition and conversion processing on the user operation requirement audio data to obtain user operation requirement text information; in response to determining that the network connection status with the cloud server is an offline state, inputting the user operation requirement text information into an edge intention recognition module group to obtain intention recognition information; in response to determining that the network connection status with the cloud server is a connected state, parallelly inputting the user operation requirement text information into the edge intention recognition module group and the cloud intention recognition module deployed on the cloud server to obtain intention recognition information; generating response operation task information corresponding to the intention recognition information; and performing a response operation corresponding to the response operation task information to obtain operation feedback information. The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the intention recognition system of some embodiments of the present disclosure, the user experience is improved. Specifically, the reason for the poor user experience is as follows: Only relying on the cloud server for intention recognition, obtaining the intention recognition information and making a response operation highly depends on the network environment. When in a network-free state, the user-end device cannot establish a connection with the cloud, and the system cannot work properly, resulting in the user being unable to use the intention recognition function and a poor user experience. At the same time, in the case of network congestion or poor signal, the data transmission speed between the user end and the cloud becomes slower, and there may even be data loss or transmission errors, resulting in an unstable time-consuming process for the system from receiving user input to returning the recognition result. Sometimes, it may take a long time to make a response operation and feedback operation feedback information, seriously affecting the user experience. Based on this, the intention recognition system of some embodiments of the present disclosure includes: an end-side device and a cloud server, where: the above-mentioned end-side device is configured to perform the following multi-level hit response processing: First, receive the user operation demand audio data, and perform recognition and conversion processing on the above-mentioned user operation demand audio data to obtain the user operation demand text information. Thus, the user operation demand audio data can be converted into text information, that is, the user operation demand text information. Then, in response to determining that the network connection status with the above-mentioned cloud server is an offline state, input the above-mentioned user operation demand text information into the end-side intention recognition module group to obtain the intention recognition information. Thus, when not relying on the network (in a network-free state) and the cloud, it is still possible to perform intention recognition on the user operation demand text information based on the end-side intention recognition module group deployed on the end-side device. In response to determining that the network connection status with the above-mentioned cloud server is a connected state, input the above-mentioned user operation demand text information in parallel into the above-mentioned end-side intention recognition module group and the cloud intention recognition module deployed on the above-mentioned cloud server to obtain the intention recognition information. Thus, the user-end device can process data without relying on the network connection with the cloud server. When the user operation demand text information is input into the end-side module group, it can directly perform intention recognition processing on the local computing resources without transmitting data to the cloud through the network, so it is not affected by network congestion or poor signal. Since the end-side processing does not need to wait for network transmission, it can complete intention recognition and return the result in a short time. In this way, even if the cloud server responds slowly due to network problems, the end-side can promptly provide the user with preliminary intention recognition information, ensuring that the system can at least quickly give a feedback and avoiding the user waiting for a long time, thus improving the user experience. Then, generate the response operation task information corresponding to the above-mentioned intention recognition information. Thus, the response operation task information for performing the response operation can be generated. Finally, execute the response operation corresponding to the above-mentioned response operation task information to obtain the operation feedback information.Also, because an intention recognition system adopting an architecture that coordinates the edge side and the cloud side is used, even when not relying on the network (in a state without network) and the cloud server, it can still identify the intention of the user operation requirement text information based on the edge side intention recognition module group deployed on the edge side device. Moreover, the edge side device does not need to wait for network transmission for processing, and can complete intention recognition and return results in a short time. In this way, even if the cloud server responds slowly due to network problems, the edge side device can provide preliminary intention recognition information to the user in time, ensuring that the system can at least give a quick feedback, avoiding long waiting for the user, and improving the user experience. Brief Description of the Drawings

[0009] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0010] Figure 1 is an architecture diagram of an exemplary system of the intention recognition system according to the present disclosure and shows an application scenario of a multi-level hit response processing method of some embodiments of the present disclosure; Figure 2 A flowchart of some embodiments of the multi-level hit response processing method according to the present disclosure. Detailed Description of the Embodiments

[0011] The embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0012] In addition, it should be noted that for the convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0013] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0014] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0015] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0016] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0017] Figure 1 An exemplary system architecture 100 of an intent recognition system to which some embodiments of the present disclosure can be applied is shown.

[0018] As Figure 1 shown, the system architecture 100 may include: an edge device 101 and a cloud server 102. The edge device 101 and the cloud server 102 are connected through a network. Among them, the edge device 101 may be a device deployed locally by the user and directly interacting with the user in the intent recognition system (for example, a smart phone, a smart speaker, a smart vehicle device, etc.). The above cloud server 102 may be a cloud server deploying a large language model (LLM).

[0019] In some embodiments, the above edge device 101 may be configured to perform the following multi-level hit response processing: First, receive user operation requirement audio data, and perform identification and conversion processing on the user operation requirement audio data to obtain user operation requirement text information. In practice, the above edge device 101 may receive user operation requirement audio data through a built-in microphone. The above edge device 101 may perform identification and conversion processing on the user operation requirement audio data through Automatic Speech Recognition (ASR) technology to obtain user operation requirement text information. Among them, the above user operation requirement audio data may represent voice requirement information (for example, a voice command) issued by the user. The above user operation requirement text information may be information of the user operation requirement presented in text form obtained after the user operation requirement audio data is subjected to identification and conversion processing (for example, querying the weather). For example, the above user operation requirement text information may be "What's the weather like in Area A today?".

[0020] In some optional implementation manners of some embodiments, the above edge device 101 may perform the following steps to perform identification and conversion processing on the user operation requirement audio data to obtain user operation requirement text information: First step, perform frame segmentation and windowing processing on the above-mentioned audio data of user operation requirements to obtain a sequence of audio frames of user operation requirements. Among them, each audio frame of user operation requirements in the above-mentioned sequence of audio frames of user operation requirements corresponds to a serial number. In practice, the above-mentioned edge device 101 can use the frame segmentation and windowing technology of MATLAB speech signal processing to perform frame segmentation and windowing processing on the audio data of user operation requirements to obtain a sequence of audio frames of user operation requirements. The above-mentioned sequence of audio frames of user operation requirements can be a set of a series of discrete audio frames obtained by performing frame segmentation and windowing processing on the audio data of user operation requirements. Each audio frame of user operation requirements in the above-mentioned sequence of audio frames of user operation requirements can represent a section of audio in the audio corresponding to the audio data of user operation requirements.

[0021] Second step, for each audio frame of user operation requirements in the above-mentioned sequence of audio frames of user operation requirements, perform the following processing: First sub-step, perform frequency-domain conversion processing on the above-mentioned audio frame of user operation requirements to obtain amplitude spectrum information. In practice, the above-mentioned edge device 101 can use the short-time Fourier transform (STFT) technology to perform frequency-domain conversion processing on the audio frame of user operation requirements to obtain amplitude spectrum information. The above-mentioned amplitude spectrum information can be a representation of an amplitude spectrum obtained by performing frequency-domain analysis on the audio signal represented by the audio frame of user operation requirements, which reflects the energy distribution of the audio signal at different frequency components.

[0022] Second sub-step, perform audio feature extraction processing on the above-mentioned amplitude spectrum information to obtain audio frequency-domain feature information corresponding to the above-mentioned audio frame of user operation requirements. In practice, the above-mentioned edge device 101 can use a Mel filter bank to perform audio feature extraction processing on the above-mentioned amplitude spectrum information to obtain audio frequency-domain feature information corresponding to the above-mentioned audio frame of user operation requirements. The above-mentioned audio frequency-domain feature information can represent the Fbank feature of the audio signal corresponding to the above-mentioned audio frame of user operation requirements.

[0023] Third step, sort the obtained audio frequency-domain feature information into a preset queue according to its corresponding serial number to obtain an audio frequency-domain feature information queue. In practice, the generated audio frequency-domain feature information can be sequentially entered into the preset queue according to its corresponding serial number to obtain an audio frequency-domain feature information queue. As an example, first put the audio frequency-domain feature information with serial number 1 into the preset queue, then the audio frequency-domain feature information with serial number 2 into the preset queue, and so on.

[0024] Fourth step, based on the audio frequency-domain feature information queue and the initial model state information, perform the following steps of the cache recognition step: The first sub-step is to add the first audio frequency domain feature information in the above audio frequency domain feature information queue to a cache queue with a cache size of a preset number of audio frequency domain feature information, and based on the cache queue, perform the following recognition steps: The first step is to delete the first audio frequency domain feature information from the audio frequency domain feature information queue to update the audio frequency domain feature information queue. In practice, the above-mentioned edge device 101 can perform an enqueue operation on the first audio frequency domain feature information to cache the first audio frequency domain feature information into the cache queue. Among them, the above-mentioned cache queue can be a preset queue. The above-mentioned initial model state information can be empty when performing the first cache recognition step. The above-mentioned preset number can be 2. For example, the cache queue can be {T1}, {T1, T2}, {T2, T3}... {Tn-2, Tn-1}, {Tn-1, Tn}. Among them, T1 represents the first audio frequency domain feature information entering the cache queue, which is the first audio frequency domain feature information in the audio frequency domain feature information queue before update. T2 is the second audio frequency domain feature information in the audio frequency domain feature information queue before update, and so on. The "n" in the above-mentioned "Tn" is the number of audio frequency domain feature information in the audio frequency domain feature information queue before update.

[0025] The second step is to, in response to determining that the updated audio frequency domain feature information queue is not empty, perform the following update steps: Sub-step one: Input the initial model state information and at least one audio frequency domain feature information included in the cache queue into a pre-trained speech recognition model to obtain recognition text information and input feature information. Among them, the above-mentioned speech recognition model can be a Wenet model. The above-mentioned speech recognition model can be a model after pruning, reducing the feature channels from 640 to 320 and performing int8 quantization at the same time. The above-mentioned recognition text information can be the text predicted for the audio features corresponding to the audio frequency domain feature information. At least one audio frequency domain feature information included in the above-mentioned cache queue can refer to the first audio frequency domain feature information in the cache queue. The above-mentioned input feature information can be the downsampled feature of the input at least one audio frequency domain feature information (which can be represented by a two-dimensional matrix, where one dimension is the time frame and the other dimension is the feature dimension). As an example, in the first round, {T1} can be input into the speech recognition model to obtain "Hel" and input feature information: the downsampled feature of "T1". In the second round, {T2}, and the downsampled feature of "T1", can be input into the speech recognition model to obtain recognition text information: "Hello" and input feature information: the downsampled feature of {T1, T2}. In the third round, {T3}, and the downsampled feature of {T1, T2}, can be input into the speech recognition model to obtain recognition text information: "Hello" and input feature information: the downsampled feature of {T1, T2, T3}.

[0026] Sub-step 2: Update the input feature information to the initial model state information to update the initial model state information.

[0027] Sub-step 3: Based on the updated audio frequency domain feature information queue and the initial model state information, execute the above caching recognition step again.

[0028] Step 3: In response to determining that the updated audio frequency domain feature information queue is empty, input the updated initial model state information and at least one audio frequency domain feature information included in the cache queue into the speech recognition model to obtain recognition text information, and determine the obtained recognition text information as the user operation requirement text information corresponding to the above user operation requirement audio data.

[0029] The above technical solution and its related content, as an inventive point of an embodiment of the present disclosure, solve the technical problem of "directly splitting a long audio into independent segments (such as one segment every 2 seconds), with no information transfer between segments, which easily leads to cross-segment semantic breaks during the process of converting audio data into text information (such as the previous segment 'open' and the subsequent segment 'light' being misjudged as 'open etc.'), thereby resulting in relatively low accuracy of the text information of the converted user operation requirements". The factors that lead to relatively low accuracy of the text information of the converted user operation requirements are often as follows: directly splitting a long audio into independent segments (such as one segment every 2 seconds), with no information transfer between segments, which easily leads to cross-segment semantic breaks during the process of converting audio data into text information (such as the previous segment 'open' and the subsequent segment 'light' being misjudged as 'open etc.'), thereby resulting in relatively low accuracy of the text information of the converted user operation requirements. If the above factors are solved, the effect of improving the accuracy of the text information of the converted user operation requirements can be achieved. To achieve this effect, first, perform frame addition and windowing processing on the above user operation requirement audio data to obtain a sequence of user operation requirement audio frames. Among them, each user operation requirement audio frame in the above sequence of user operation requirement audio frames corresponds to a serial number. Thus, the user operation requirement audio data can be split into short-time frames, that is, user operation requirement audio frames, and a serial number is assigned to each user operation requirement audio frame, retaining the timeliness of the audio. Then, for each user operation requirement audio frame in the above sequence of user operation requirement audio frames, perform the following processing: The first step is to perform frequency-domain conversion processing on the above user operation requirement audio frame to obtain amplitude spectrum information. The second step is to perform audio feature extraction processing on the above amplitude spectrum information to obtain audio frequency-domain feature information corresponding to the above user operation requirement audio frame. The third step is to sort the obtained audio frequency-domain feature information into a preset queue according to its corresponding serial number to obtain an audio frequency-domain feature information queue. Thus, through the above steps, the generated audio frequency-domain feature information can enter the preset queue one by one in sequence according to its corresponding serial number. Next, based on the audio frequency-domain feature information queue and the initial model state information, perform the following steps of the cache recognition step: Sub-step one, add the first audio frequency-domain feature information in the above audio frequency-domain feature information queue to a cache queue with a cache size of a preset number of audio frequency-domain feature information, and based on the cache queue, perform the following recognition steps: The first step is to delete the first audio frequency-domain feature information from the audio frequency-domain feature information queue to update the audio frequency-domain feature information queue. The second step, in response to determining that the updated audio frequency-domain feature information queue is not empty, perform the following update steps: Sub-step one, input the initial model state information and at least one audio frequency-domain feature information included in the cache queue into a pre-trained speech recognition model to obtain recognition text information and input feature information. Thus, the audio frequency-domain features in the cache queue and the current model state can be input into the speech recognition model to obtain recognition text and new input feature information.Sub-step 2: Update the input feature information to the initial model state information to update the initial model state information. Thus, the new input feature information (i.e., the model state) can be updated to the initial state for the next inference. This state transfer enables the model to utilize historical information (such as "open" in the previous segment) to understand the content of the current segment. Through the update of the initial model state information, model state transfer is achieved, and the model can capture cross-frame temporal dependencies, avoiding semantic breaks. Sub-step 3: Based on the updated audio frequency domain feature information queue and the initial model state information, perform the above-mentioned cache recognition step again. Step 3: In response to determining that the updated audio frequency domain feature information queue is empty, input the updated initial model state information and at least one audio frequency domain feature information included in the cache queue into the speech recognition model to obtain the recognized text information, and determine the obtained recognized text information as the user operation requirement text information corresponding to the above-mentioned user operation requirement audio data. Thus, the user operation requirement text information corresponding to the above-mentioned user operation requirement audio data can be obtained. Also, because the audio frequency domain feature information queue and the cache queue are adopted, the model can perform inference while the user is speaking through the speech recognition model. The model does not need to wait for the complete audio input but processes it frame by frame, so the recognition result can be output in real time. Moreover, during the recognition conversion process, the cache queue retains the most recent audio frequency domain feature information, i.e., local context information. This enables the model to utilize the local context and the initial model state information that preserves the historical input features (such as "open" in the previous segment) to convert the content of the audio frame. Through the update of the initial model state information, model state transfer is achieved, and the model can capture cross-frame temporal dependencies, avoiding semantic breaks, and improving the accuracy of converting the user operation requirement audio data into text information, i.e., the user operation requirement text information.

[0030] Second, in response to determining that the network connection status with the above cloud server is in an offline state, input the above user operation requirement text information into the edge-side intent recognition module group to obtain intent recognition information. Among them, the above edge-side intent recognition module group includes a rule-driven intent recognition module and a model-based intent recognition module. The above rule-driven intent recognition module can be a software module deployed on the edge-side device that uses a rule engine for intent classification and extraction. The above model-based intent recognition module can be a software module deployed on the edge-side device that relies on machine learning or deep learning models (such as, small language models (LLMs)) to recognize user intents. The intent recognition information output by the above edge-side intent recognition module group includes intent information and slot information. The above intent information can be an abstract category of the user operation requirement text information. The slot information can be parameters or attributes (i.e., keywords) included in the user operation requirement text information. For example, if the user operation requirement text information is "Play Song A by Singer A", the intent recognition information can be "Play music", and the slot information can be "Song name (Slot: Song A), Singer (Slot: Singer A)". In practice, the above edge-side device 101 can input the user operation requirement text information into the rule-driven intent recognition module for the rule-driven intent recognition module to use the rule engine to classify the intent of the user operation requirement text information to obtain intent information. Then, models such as BiLSTM-CRF and BERT-CRF are used to perform sequence annotation on the user operation requirement text information to identify the slot values and obtain slot information. Optionally, the above edge-side device 101 can input the user operation requirement text information into the model-based intent recognition module for the model-based intent recognition module to input the user operation requirement text information into a pre-trained small language model (LLM) to obtain intent recognition information.

[0031] In some optional implementation manners of some embodiments, the above edge-side device 101 can respond to determining that the network connection status with the above cloud server is in an offline state and input the above user operation requirement text information into the edge-side intent recognition module group to obtain intent recognition information through the following steps: First step, in response to determining that the network connection status with the above cloud server is in an offline state, input the above user operation requirement text information into the rule-driven intent recognition module included in the above device-side intent recognition module group, so that the rule-driven intent recognition module inputs the above user operation requirement text information into a preset converter to obtain first device-side intent recognition information, and input the above user operation requirement text information into the model-based intent recognition module included in the above device-side intent recognition module group, so that the model-based intent recognition module performs intent recognition processing on the above user operation requirement text information. In practice, the above device 101 can input the user operation requirement text information into a preset weighted finite state converter to obtain first device-side intent recognition information. The above first device-side intent recognition information can be the intent recognition information output by the rule-driven intent recognition module. The above device 101 can input the user operation requirement text information into the model-based intent recognition module included in the device-side intent recognition module group. The above model-based intent recognition module can use a pre-trained NLU model (for example, the 0.5B model of Alibaba Tongyi Qianwen version 3.0) to perform intent recognition processing on the above user operation requirement text information. It should be noted that the above NLU model can be trained using a preset training sample set, and each preset training sample in the preset training sample set can include user operation requirement text information (input data), intent information, and slot information (output data).

[0032] Second step, in response to determining that the above first device-side intent recognition information meets a preset hit condition, execute a termination task corresponding to the preset intent recognition termination task information to end the intent recognition processing of the above model-based intent recognition module on the above user operation requirement text information, and determine the above first device-side intent recognition information as the intent recognition information. Among them, the above preset hit condition can be that the intent information included in the intent recognition information is the same as the entry information of a rule information in a preset predefined rule information set. As an example, the rule information set can be {{rule information: play [singer name]'s [song name], corresponding entry information is play music}, {query the weather of [location] at [time]}, corresponding entry information is query weather}}. If the first device-side intent recognition information is "play singer A's 'Song A'", then the intent recognition information can be "play music", and the slot information can be "song name (Slot: Song A), singer (Slot: singer A)", then it meets the preset hit condition. The above preset intent recognition termination task information can be a preset instruction for terminating the further processing (such as intent recognition processing, cloud intent recognition processing) of the model-based intent recognition module or the cloud intent recognition module on the user operation requirement text information.

[0033] In some alternative implementation manners of some embodiments, the above-mentioned edge device 101 may be further configured to: First, in response to determining that the above-mentioned first edge intention recognition information does not meet the preset hit condition, determine the information obtained by the model-based intention recognition module for intention recognition processing of the above-mentioned user operation requirement text information as the second edge intention recognition information. Among them, the above-mentioned second edge intention recognition information may be the intention recognition information obtained by the model-based intention recognition module using a pre-trained NLU model (for example, the 0.5B model of Alibaba Tongyi Qianwen version 3.0) to perform intention recognition processing on the above-mentioned user operation requirement text information.

[0034] Second, determine the above-mentioned second edge intention recognition information as the intention recognition information.

[0035] Third, in response to determining that the network connection status with the above-mentioned cloud server is in a connected state, parallelly input the above-mentioned user operation requirement text information into the above-mentioned edge intention recognition module group and the cloud intention recognition module deployed on the above-mentioned cloud server to obtain the intention recognition information. It should be noted that the parallel input here means inputting the user operation requirement text information into the above-mentioned edge intention recognition module group and the cloud intention recognition module deployed on the above-mentioned cloud server simultaneously.

[0036] In some alternative implementation manners of some embodiments, the above-mentioned edge device 101 may be further configured to, in response to determining that the network connection status with the above-mentioned cloud server is in a connected state, parallelly input the above-mentioned user operation requirement text information into the above-mentioned edge intention recognition module group and the cloud intention recognition module deployed on the above-mentioned cloud server through the following steps to obtain the intention recognition information: First step, in response to determining that the network connection status with the above cloud server is in a connected state, input the above user operation requirement text information into the above end-side intent recognition module group to obtain the to-be-verified intent recognition information, and simultaneously send the user operation requirement text information to the above cloud server for the cloud intent recognition module deployed on the above cloud server to perform cloud intent recognition processing on the above user operation requirement text information. In practice, first, the above end-side device 101 can input the above user operation requirement text information into the rule-driven intent recognition module included in the above end-side intent recognition module group, so that the rule-driven intent recognition module inputs the above user operation requirement text information into a preset converter to obtain the first end-side intent recognition information, and inputs the above user operation requirement text information into the model-based intent recognition module included in the above end-side intent recognition module group, so that the model-based intent recognition module performs intent recognition processing on the above user operation requirement text information. Then, in response to determining that the above first end-side intent recognition information meets the preset hit condition, execute the termination task corresponding to the preset intent recognition termination task information to end the intent recognition processing of the above user operation requirement text information by the above model-based intent recognition module, and send the preset intent recognition termination task information to the above cloud server for the above cloud server to execute the termination task corresponding to the above preset intent recognition termination task information to end the above cloud intent recognition processing, and determine the above first end-side intent recognition information as the intent recognition information. Then, in response to determining that the above first end-side intent recognition information does not meet the preset hit condition, determine the information obtained by the model-based intent recognition module performing intent recognition processing on the above user operation requirement text information as the second end-side intent recognition information. Finally, determine the second end-side intent recognition information as the to-be-verified intent recognition information. Among them, the above cloud intent recognition module can be a software module for performing cloud intent recognition processing on the above user operation requirement text information.

[0037] Second step, in response to determining that the above to-be-verified intent recognition information meets the preset hit condition, send the preset intent recognition termination task information to the above cloud server for the above cloud server to execute the termination task corresponding to the above preset intent recognition termination task information to end the above cloud intent recognition processing, and determine the above to-be-verified intent recognition information as the intent recognition information.

[0038] In some optional implementation manners of some embodiments, the above end-side device 101 can be further configured to: First step, in response to determining that the above-mentioned intention recognition information to be verified does not meet the preset hit condition, at least one intention recognition information to be screened sent by the cloud server is received. Among them, the above-mentioned at least one intention recognition information to be screened is the information obtained by the cloud intention recognition module deployed on the cloud server through cloud intention recognition processing of the above-mentioned user operation requirement text information. The above-mentioned cloud intention recognition module can be a software module that performs intention recognition on the user operation requirement text information in the way of function call. In practice, the above-mentioned cloud intention recognition module can use an intention recognition model deployed on the cloud (such as a fine-tuned version of pre-trained models such as BERT and RoBERTa) to perform cloud intention recognition processing on the user operation requirement text information to obtain at least one intention recognition information to be screened. As an example, the above-mentioned user operation requirement text information can be "What's the weather like in Area A tomorrow? Is it suitable for traveling?", then the at least one intention recognition information to be screened can be "Intention information: Weather query, Slot information: Time: tomorrow, Location: Area A", "Intention information: Travel consultation, Slot information: Location: Area A". It should be noted that the parallel here means simultaneous.

[0039] Second step, each intention recognition information to be screened that meets the preset hit condition among the above-mentioned at least one intention recognition information to be screened is determined as each screened intention recognition information.

[0040] Third step, the above-mentioned each screened intention recognition information is determined as intention recognition information. Among them, the intention recognition information includes one of the following: each screened intention recognition information, first end-side intention recognition information, second end-side intention recognition information. The above-mentioned intention recognition information includes at least one recognition information, and each recognition information in the at least one recognition information included in the above-mentioned intention recognition information is one of the following: screened intention recognition information, first end-side intention recognition information, second end-side intention recognition information, and the above-mentioned recognition information includes intention information and slot information.

[0041] Fourth, generate response operation task information corresponding to the above-mentioned intention recognition information.

[0042] In some optional implementation manners of some embodiments, the above-mentioned terminal device 101 can be further configured to generate response operation task information corresponding to the above-mentioned intention recognition information through the following steps: First step, for each recognition information in the at least one recognition information included in the above-mentioned intention recognition information, the following steps are performed: First sub-step, the intention information included in the above-mentioned recognition information is determined as the intention information to be queried.

[0043] The second sub-step is to query the preset intention operation mapping information corresponding to the above-mentioned intention information to be queried from the preset skill structured mapping table, and determine the queried preset intention operation mapping information as the target intention operation mapping information. Among them, the above-mentioned preset skill structured mapping table includes various preset intention operation mapping information, and each preset mapping operation mapping information in the above-mentioned various preset mapping operation mapping information includes intention entry information and intention response operation information. The above-mentioned intention entry information describes the keywords or phrases of the user's intention (such as "weather query", "music playback"). The above-mentioned intention response operation information can represent the specific operations to be performed (such as calling the weather API, starting the music player). The above-mentioned intention response operation information can be a description of a function call. (For example, the description of the function call of the weather API is {operation: "call_api", api_name: "weather_api"}).

[0044] The third sub-step is to determine the intention response operation information included in the above-mentioned target intention operation mapping information as the initial response operation task information.

[0045] The fourth sub-step is to determine the slot information included in the above-mentioned recognition information as the parameter information corresponding to the above-mentioned initial response operation task information.

[0046] The fifth sub-step is to combine the above-mentioned initial response operation task information and the above-mentioned parameter information into response operation task sub-information. As an example, if the initial response operation task information is "{operation: "call_api", api_name: "weather_api"}" and the parameter information is {time: tomorrow, location: Area A}, then the response operation task sub-information can be "{ "operation": "call_api", "api_name": "weather_api", "parameters": {"city": "Area A", "date": "tomorrow"}}".

[0047] The second step is to determine the at least one obtained response operation task sub-information as the response operation task information.

[0048] Fifth, perform the response operation corresponding to the above-mentioned response operation task information to obtain operation feedback information. Among them, the above-mentioned response operation can be the response operation corresponding to at least one response operation task sub-information (for example, query the weather in Area A tomorrow). The above-mentioned operation feedback information can be the data (for example, the weather data in Area A tomorrow) or status (for example, whether the music playback is successful) returned by the system after performing the response operation corresponding to the response operation task information.

[0049] In some alternative implementations of some embodiments, the above-mentioned edge device 101 may be further configured to: In the first step, based on the above-mentioned intent recognition information, perform language re-organization processing on the above-mentioned operation feedback information to obtain normalized operation feedback information. In practice, the above-mentioned edge device 101 may determine a sentence pattern template corresponding to the intent recognition information according to the intent recognition information, and use the pre-defined sentence pattern template to fill in the key information in the operation feedback information into the template to generate normalized operation feedback information. For example, the sentence pattern template corresponding to weather query may be "{Time}{Location}'s weather is {Operation Feedback Information}". Optionally, a lightweight edge language model (such as an LLM with 0.5B parameters) may be used to take the intent recognition information and the operation feedback information as inputs, and the model generates natural and fluent normalized operation feedback information. For example, the intent recognition information may be "Intent Information: Weather Query, Slot Information: Time: Tomorrow, Location: Area A", and the operation feedback information may be "It's raining, temperature is 25 degrees", then the normalized operation feedback information may be "The weather in Area A today is raining, temperature is 25 degrees" or "This afternoon in Hangzhou, the temperature is 25 degrees, remember to bring an umbrella".

[0050] In the second step, convert the above-mentioned normalized operation feedback information into operation feedback voice data, and play the voice corresponding to the above-mentioned operation feedback voice data. In practice, the TTS technology may be used to convert the normalized operation feedback information into operation feedback voice data. The above-mentioned normalized operation feedback information may be text information. The above-mentioned operation feedback voice data may represent the voice corresponding to the normalized operation feedback information. The above-mentioned edge device 101 may play the voice corresponding to the above-mentioned operation feedback voice data through a speaker.

[0051] In some alternative implementations of some embodiments, the above-mentioned edge device 101 may be further configured to convert the above-mentioned normalized operation feedback information into operation feedback voice data through the following steps: In the first step, perform word segmentation processing on the above-mentioned normalized operation feedback information to obtain a word segmentation sequence. In practice, the above-mentioned edge device 101 may use the Jieba word segmentation technology to perform word segmentation processing on the normalized operation feedback information to obtain a word segmentation sequence.

[0052] In the second step, based on the above-mentioned word segmentation sequence, generate a phoneme information sequence corresponding to the above-mentioned normalized operation feedback information. In practice, the above-mentioned edge device 101 may call pypinyin (a Python library) to generate a phoneme information sequence corresponding to the word segmentation sequence. Among them, each word segmentation in the above-mentioned word segmentation sequence corresponds to a phoneme information in the above-mentioned phoneme information sequence. The above-mentioned phoneme information may represent the phoneme of the corresponding word segmentation.

[0053] Step 3: For each phoneme information in the above phoneme information sequence, predict the phoneme duration corresponding to the above phoneme information. In practice, the above edge device 101 can input the above phoneme information into the Tacotron2 model to obtain the phoneme duration corresponding to the above phoneme information.

[0054] Step 4: Based on the predicted phoneme durations, generate a phoneme duration sequence. Each phoneme duration in the above phoneme duration sequence corresponds to one phoneme information in the above phoneme information sequence. In practice, the above edge device 101 can sort the phoneme durations according to the order of the corresponding phoneme information in the phoneme information sequence to obtain the phoneme duration sequence.

[0055] Step 5: Input the above phoneme information sequence and the above phoneme duration sequence into a pre-trained voice model to obtain operation feedback voice data. The above voice model can be a SpeedySpeech model. The above operation feedback voice data can be the mel spectrogram representing the voice corresponding to the normalized operation feedback information. Figure 2 Flow 200 shows the processes of some embodiments of the multi-level hit response processing method of the edge device included in the application of the above intention recognition system according to the present disclosure. The multi-level hit response processing method includes the following steps: Step 201: Receive user operation requirement audio data, and perform recognition and conversion processing on the user operation requirement audio data to obtain user operation requirement text information.

[0056] In some embodiments, the execution subject of the multi-level hit response processing method (such as the edge device included in the intention recognition system) can receive the user operation requirement audio data and perform recognition and conversion processing on the above user operation requirement audio data to obtain user operation requirement text information.

[0057] Step 202: In response to determining that the network connection status with the cloud server is an offline state, input the user operation requirement text information into the edge intention recognition module group to obtain intention recognition information.

[0058] In some embodiments, the above execution subject can, in response to determining that the network connection status with the above cloud server is an offline state, input the above user operation requirement text information into the edge intention recognition module group to obtain intention recognition information.

[0059] Step 203: In response to determining that the network connection status with the cloud server is a connected state, input the user operation requirement text information in parallel into the edge intention recognition module group and the cloud intention recognition module deployed on the cloud server to obtain intention recognition information.

[0060] In some embodiments, the above-mentioned execution entity may, in response to determining that the network connection status with the above-mentioned cloud server is in a connected state, input the above-mentioned user operation requirement text information in parallel to the above-mentioned edge-side intent recognition module group and the cloud intent recognition module deployed on the above-mentioned cloud server to obtain intent recognition information.

[0061] Step 204: Generate response operation task information corresponding to the intent recognition information.

[0062] In some embodiments, the above-mentioned execution entity may generate response operation task information corresponding to the above-mentioned intent recognition information.

[0063] Step 205: Execute a response operation corresponding to the response operation task information to obtain operation feedback information.

[0064] In some embodiments, the above-mentioned execution entity may execute a response operation corresponding to the above-mentioned response operation task information to obtain operation feedback information.

[0065] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the intention recognition system of some embodiments of the present disclosure, the user experience is improved. Specifically, the reasons for the poor user experience are as follows: Relying solely on the cloud server for intention recognition, obtaining the intention recognition information and performing response operations highly depend on the network environment. When in a network-free state, the user device cannot establish a connection with the cloud, and the system cannot work properly, resulting in the user being unable to use the intention recognition function and a poor user experience. At the same time, in the case of network congestion or poor signal, the data transmission speed between the user device and the cloud becomes slower, and there may even be data loss or transmission errors, resulting in an unstable time-consuming for the entire process from the system receiving user input to returning the recognition result. Sometimes, it may take a long time to perform response operations and feedback operation feedback information, seriously affecting the user experience. Based on this, the intention recognition system of some embodiments of the present disclosure includes: an edge device and a cloud server, where: The above-mentioned edge device is configured to perform the following multi-level hit response processing: First, receive the user operation requirement audio data, and perform recognition and conversion processing on the above-mentioned user operation requirement audio data to obtain the user operation requirement text information. Thus, the user operation requirement audio data can be converted into text information, that is, the user operation requirement text information. Then, in response to determining that the network connection status with the above-mentioned cloud server is an offline state, input the above-mentioned user operation requirement text information into the edge intention recognition module group to obtain the intention recognition information. Thus, when not relying on the network (in a network-free state) and the cloud, it is still possible to perform intention recognition on the user operation requirement text information based on the edge intention recognition module group deployed on the edge device. In response to determining that the network connection status with the above-mentioned cloud server is a connected state, input the above-mentioned user operation requirement text information in parallel into the above-mentioned edge intention recognition module group and the cloud intention recognition module deployed on the above-mentioned cloud server to obtain the intention recognition information. Thus, the user device can process data without relying on the network connection with the cloud server. When the user operation requirement text information is input into the edge module group, it can directly perform intention recognition processing on the local computing resources without transmitting data to the cloud through the network, so it is not affected by network congestion or poor signal. Since the edge processing does not need to wait for network transmission, it can complete intention recognition and return the result in a short time. In this way, even if the cloud server responds slowly due to network problems, the edge can promptly provide the user with preliminary intention recognition information, ensuring that the system can at least quickly give a feedback, avoiding the user waiting for a long time, and improving the user experience. Then, generate the response operation task information corresponding to the above-mentioned intention recognition information. Thus, the response operation task information for performing the response operation can be generated. Finally, perform the response operation corresponding to the above-mentioned response operation task information to obtain the operation feedback information.Also, because an intention recognition system adopting an end - side to cloud - side collaborative architecture can still recognize the intention of the text information of the user's operation requirements based on the end - side intention recognition module group deployed on the end - side device without relying on the network (in the state of no network) and the cloud server. Moreover, the end - side device can complete the intention recognition and return the result in a short time without waiting for network transmission. In this way, even if the cloud server responds slowly due to network problems, the end - side device can provide preliminary intention recognition information to the user in time, ensuring that the system can at least give a quick feedback, avoiding long waiting for the user, and improving the user experience.

[0066] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of technical features, and should also cover other technical solutions formed by any combination of technical features or their equivalent features without departing from the inventive concept. For example, a technical solution formed by mutually replacing features with similar functions disclosed in the embodiments of the present disclosure (but not limited to).

Claims

1. An intention recognition system includes: Edge device, cloud server, where: The edge device is configured to perform the following multi-level hit response processing: Receive user operation requirement audio data, and perform identification and conversion processing on the user operation requirement audio data to obtain user operation requirement text information; In response to determining that the network connection status with the cloud server is offline, input the user operation requirement text information into the edge intention recognition module group to obtain intention recognition information; In response to determining that the network connection status with the cloud server is connected, input the user operation requirement text information in parallel into the edge intention recognition module group and the cloud intention recognition module deployed on the cloud server to obtain intention recognition information; Generate response operation task information corresponding to the intention recognition information; Execute a response operation corresponding to the response operation task information to obtain operation feedback information.

2. The intention recognition system according to claim 1, wherein, The edge intention recognition module group includes a rule-driven intention recognition module and a model-based intention recognition module, and the edge device is further configured to: In response to determining that the network connection status with the cloud server is offline, input the user operation requirement text information into the rule-driven intention recognition module included in the edge intention recognition module group, so that the rule-driven intention recognition module inputs the user operation requirement text information into a preset converter to obtain first edge intention recognition information, and input the user operation requirement text information into the model-based intention recognition module included in the edge intention recognition module group, so that the model-based intention recognition module performs intention recognition processing on the user operation requirement text information; In response to determining that the first edge intention recognition information meets a preset hit condition, execute a termination task corresponding to the preset intention recognition termination task information to end the intention recognition processing of the user operation requirement text information by the model-based intention recognition module, and determine the first edge intention recognition information as intention recognition information.

3. The intention recognition system according to claim 2, wherein, The edge device is further configured to: In response to determining that the first edge intention recognition information does not meet the preset hit condition, determine the information obtained by the model-based intention recognition module performing intention recognition processing on the user operation requirement text information as second edge intention recognition information; Determine the second edge intention recognition information as intention recognition information.

4. The intention recognition system according to claim 1, wherein, The edge device is further configured to: In response to determining that the network connection status with the cloud server is connected, input the user operation requirement text information into the edge intention recognition module group to obtain to-be-verified intention recognition information, and send the user operation requirement text information to the cloud server in parallel for the cloud intention recognition module deployed on the cloud server to perform cloud intention recognition processing on the user operation requirement text information; In response to determining that the to-be-verified intent recognition information meets a preset hit condition, send preset intent recognition termination task information to the cloud server for the cloud server to execute a termination task corresponding to the preset intent recognition termination task information to end the cloud intent recognition process, and determine the to-be-verified intent recognition information as intent recognition information.

5. The intention recognition system according to claim 4, wherein, The edge device is further configured to: In response to determining that the to-be-verified intent recognition information does not meet the preset hit condition, receive at least one to-be-screened intent recognition information sent by the cloud server, where the at least one to-be-screened intent recognition information is information obtained by the cloud intent recognition module deployed by the cloud server performing cloud intent recognition processing on the user operation requirement text information; Determine each to-be-screened intent recognition information that meets the preset hit condition among the at least one to-be-screened intent recognition information as each screened intent recognition information; Determine the each screened intent recognition information as intent recognition information.

6. The intention recognition system according to claim 1, wherein, The intent recognition information includes one of the following: each screened intent recognition information, first edge intent recognition information, second edge intent recognition information. Each recognition information in at least one recognition information included in the intent recognition information is one of the following: screened intent recognition information, first edge intent recognition information, second edge intent recognition information. The recognition information includes intent information and slot information. And the edge device is further configured to: For each recognition information in at least one recognition information included in the intent recognition information, perform the following steps: Determine the intent information included in the recognition information as to-be-query intent information; Query preset intent operation mapping information corresponding to the to-be-query intent information from a preset skill structured mapping table, and determine the queried preset intent operation mapping information as target intent operation mapping information, where the preset skill structured mapping table includes each preset intent operation mapping information, and each preset mapping operation mapping information in the each preset mapping operation mapping information includes intent entry information and intent response operation information; Determine the intent response operation information included in the target intent operation mapping information as initial response operation task information; Determine the slot information included in the recognition information as parameter information corresponding to the initial response operation task information; Combine the initial response operation task information and the parameter information into response operation task sub-information; Determine the obtained at least one response operation task sub-information as response operation task information.

7. The intention recognition system according to claim 1, wherein The edge device is further configured to: Based on the intent recognition information, perform language re-organization processing on the operation feedback information to obtain normalized operation feedback information; Convert the normalized operation feedback information into operation feedback voice data, and play the voice corresponding to the operation feedback voice data.

8. The intention recognition system according to claim 7, wherein, The edge device is further configured to: Perform word segmentation processing on the normalized operation feedback information to obtain a word segmentation sequence; Generate a phoneme information sequence corresponding to the normalized operation feedback information based on the word segmentation sequence; For each phoneme information in the phoneme information sequence, predict the phoneme duration corresponding to the phoneme information; Based on the predicted phoneme durations, generate a phoneme duration sequence, where each phoneme duration in the phoneme duration sequence corresponds to a phoneme information in the phoneme information sequence; Input the phoneme information sequence and the phoneme duration sequence into a pre-trained voice model to obtain operation feedback voice data.

9. A multi-level hit response processing method, applied to an edge device included in the intent recognition system according to any one of claims 1-8, the method comprising: Receive user operation requirement audio data, and perform recognition and conversion processing on the user operation requirement audio data to obtain user operation requirement text information; In response to determining that the network connection status with the cloud server is an offline state, input the user operation requirement text information into the edge intent recognition module group to obtain intent recognition information; In response to determining that the network connection status with the cloud server is a connected state, parallelly input the user operation requirement text information into the edge intent recognition module group and the cloud intent recognition module deployed on the cloud server to obtain intent recognition information; Generate response operation task information corresponding to the intent recognition information; Execute a response operation corresponding to the response operation task information to obtain operation feedback information.

Citation Information

Patent Citations

  • Speech synthesis method and device, storage medium and electronic equipment

    CN111653266A

  • Embedded voice interaction system

    CN111833875A

  • Speech recognition method and device, equipment and storage medium

    CN114242067A

  • Voice control instruction recognition method and device, and storage medium

    CN114550719A

  • Voice acquisition method, computer readable storage medium and terminal equipment

    CN114613365A