Voice-based user intent recognition methods, devices, equipment, and media

CN121011179BActive Publication Date: 2026-08-11BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]现有的语音交互技术大都依赖各个意图识别模型,意图识别模型可以基于其输入得到使用人员的意图,然而,现有的语音交互功能仍存在不足,其原因在于对用户的意图的识别准确性不高

Benefits of technology

本申请通过综合多维度信息的方式,能够为模型提供更丰富、全面的数据,相比仅依赖语音文本进行意图识别的现有技术,可减少信息缺失导致的识别误差,从而提高对用户意图识别的准确性。本申请中意图识别模型的参数可以是默认参数、基于该语音设备对应的用户发出的历史指令性语音确定的参数,以及基于中央处理器发送的模型更新参数确定的参数,其中中央处理器发送的模型更新参数是基于多个语音设备发送的信息确定的,可以让模型参考其他模型的参数,防止仅依靠模型本身进行训练导致的训练时间长以及训练准确性低的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121011179B_ABST
    Figure CN121011179B_ABST
Patent Text Reader

Abstract

This application provides a voice-based user intent recognition method, apparatus, device, and medium, belonging to the field of intent recognition technology. The method includes: in response to receiving an initiating voice, acquiring environmental information; determining a first text based on a first voice; wherein the first voice is a commanding voice issued by a user corresponding to the voice device after issuing the initiating voice, and the first text is the text corresponding to the first voice; inputting the environmental information, the first voice, and the first text into a target intent recognition model to obtain the intent of the first voice issued by the user; wherein the parameters of the target intent recognition model are default parameters, or parameters determined based on historical commanding voices issued by the user corresponding to the voice device, or parameters determined based on model update parameters sent by a central processing unit, wherein the model update parameters are determined based on information sent by multiple voice devices. This application can improve the accuracy of intent recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of intent recognition technology, and more specifically, relates to a voice-based user intent recognition method, apparatus, device, and medium. Background Technology

[0002] Most existing voice interaction technologies rely on various intent recognition models, which can obtain the user's intent based on its input. However, existing voice interaction functions still have shortcomings because the accuracy of recognizing the user's intent is not high.

[0003] There is a need for a more intelligent and accurate voice intent recognition method to provide users with a more convenient, efficient, and intuitive interactive experience, and to meet users' needs for intelligent interaction. Summary of the Invention

[0004] The purpose of this application is to provide a voice-based user intent recognition method, apparatus, device, and medium to improve the accuracy of intent recognition.

[0005] A first aspect of this application provides a voice-based user intent recognition method, applied to each voice device in an intent recognition device. The intent recognition device includes: multiple voice devices and a central processing unit (CPU). The multiple voice devices interact with the CPU, including: In response to receiving the initiation voice, obtain environmental information; The first text is determined based on the first voice; wherein, the first voice is the command voice issued by the user corresponding to the voice device after issuing the activation voice, and the first text is the text corresponding to the first voice; The environmental information, the first voice, and the first text are input into the target intent recognition model to obtain the intent of the first voice issued by the user. The parameters of the target intent recognition model are either default parameters, parameters determined based on the historical command voice issued by the user corresponding to the voice device, or parameters determined based on the model update parameters sent by the central processing unit. The model update parameters are determined based on information sent by multiple voice devices.

[0006] A second aspect of this application provides a voice-based user intent recognition device, comprising: The environmental information acquisition module is used to acquire environmental information in response to receiving the initiation voice. The text determination module is used to determine the first text based on the first voice; wherein the first voice is the command voice issued by the user corresponding to the voice device after issuing the initiation voice, and the first text is the text corresponding to the first voice; The intent recognition module is used to input environmental information, first speech, and first text into the target intent recognition model to obtain the intent of the first speech issued by the user; wherein, the parameters of the target intent recognition model are default parameters, or parameters determined based on the historical command speech issued by the user corresponding to the speech device, or parameters determined based on the model update parameters sent by the central processing unit, and the model update parameters are determined based on information sent by multiple speech devices.

[0007] A third aspect of this application provides an intent recognition device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described voice-based user intent recognition method.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described voice-based user intent recognition method.

[0009] The beneficial effects of the voice-based user intent recognition method, apparatus, device, and medium provided in this application are as follows: This application, by integrating multi-dimensional information, provides the model with richer and more comprehensive data. Compared to existing technologies that rely solely on speech and text for intent recognition, it reduces recognition errors caused by missing information, thereby improving the accuracy of user intent recognition. The parameters of the intent recognition model in this application can be default parameters, parameters determined based on historical command speech issued by the user corresponding to the speech device, and parameters determined based on model update parameters sent by the central processing unit (CPU). The model update parameters sent by the CPU are determined based on information sent by multiple speech devices, allowing the model to reference the parameters of other models and preventing the problems of long training time and low training accuracy caused by relying solely on the model itself for training. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a voice-based user intent recognition method provided in an embodiment of this application; Figure 2 This is a structural block diagram of a voice-based user intent recognition device provided in an embodiment of this application; Figure 3 This is a schematic block diagram of an intent recognition device provided in an embodiment of this application. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0014] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a voice-based user intent recognition method provided in an embodiment of this application. The method is applied to each voice device in the intent recognition device, which includes multiple voice devices and a central processing unit. The multiple voice devices interact with the central processing unit. The method is executed by each voice device in the intent recognition device, and each voice device corresponds to one user by default. The method includes steps S101-S103.

[0015] Voice devices can be devices that support voice interaction, such as mobile phones, computers, and smart glasses. A central processing unit (CPU) can be a terminal server, and its essence can be a computer.

[0016] S101: In response to receiving the initiation voice, obtain environmental information.

[0017] In this embodiment, the activation voice can be a specific wake-up word, which can be customized according to user settings or is a default voice. The activation voice is used to activate the device's listening state using keywords or phrases, with high-recognition words prioritized. Environmental information can be time information, spatial information, device status, or environmental perception data, such as ambient noise, temperature, and humidity. Environmental information can be acquired based on the sensors built into the voice device, or it can be obtained by sending information request commands to other devices and acquiring information based on the information sent by those devices.

[0018] In this embodiment, the collection of environmental information only begins after the activation voice is received, taking into account the battery life of the voice device. In particular, when the voice device is a device such as smart glasses that cannot be equipped with a large-capacity battery, real-time acquisition of environmental information will have a significant impact on the battery life of the voice device.

[0019] Secondly, in this embodiment, environmental information is collected because it can affect intent recognition. For example, when a user says "play music," the speaker inputs environmental information (nighttime, home, low noise) along with other features, such as voice features and text content ("play music"), into the target intent recognition model. The final output intent label is "play sleep music," and the speaker automatically adjusts the volume and recommends white noise or light music playlists. However, if the same voice features and text content are used in an environment with daytime, entertainment venues, or medium to high noise levels, the output label will be "play rock music."

[0020] S102: Determine the first text based on the first voice; wherein, the first voice is the command voice issued by the user corresponding to the voice device after issuing the activation voice, and the first text is the text corresponding to the first voice.

[0021] In this embodiment, the first voice is the command-like voice content spoken by the user immediately after issuing a start-up voice (such as a wake word), used to express specific needs, such as setting an alarm, playing music, or checking the weather. The first text can be the corresponding text content converted from the audio signal of the first voice through speech recognition technology.

[0022] S103: Input the environmental information, the first voice, and the first text into the target intent recognition model to obtain the intent of the first voice issued by the user; wherein, the parameters of the target intent recognition model are default parameters, or parameters determined based on the historical command voice issued by the user corresponding to the voice device, or parameters determined based on the model update parameters sent by the central processing unit, and the model update parameters are determined based on information sent by multiple voice devices.

[0023] In this embodiment, the target intent recognition model is a neural network model used to map multimodal inputs to user intent labels. It is the core decision-making module of the system. In this embodiment, the target intent recognition model can be a deep learning model based on the Transformer architecture, a convolutional neural network model, or other similar models. The target intent recognition model can use models commonly used in the art.

[0024] In this embodiment, the parameters of the target intent recognition model should be default parameters, or parameters determined based on the historical command voices issued by the user corresponding to the voice device, or parameters determined based on the model update parameters sent by the central processing unit. The target intent recognition model can only be in one of the three cases mentioned above, but it can change with time or the number of command voices received.

[0025] The parameters of a target intent recognition model can be model structure parameters, such as the number of layers, the number of neurons per layer, the type of activation function, etc., or learning rate, number of iterations, batch size, regularization coefficient, etc., or weight matrix and bias vector in a neural network, etc., which can be set based on the architecture of the target intent model and actual needs.

[0026] As can be seen from the above, this application, by integrating multi-dimensional information, can provide the model with richer and more comprehensive data. Compared with existing technologies that rely solely on speech and text for intent recognition, it can reduce recognition errors caused by missing information, thereby improving the accuracy of user intent recognition. The parameters of the intent recognition model in this application can be default parameters, parameters determined based on historical command speech issued by the user corresponding to the speech device, and parameters determined based on model update parameters sent by the central processing unit. The model update parameters sent by the central processing unit are determined based on information sent by multiple speech devices, allowing the model to reference the parameters of other models and preventing the problems of long training time and low training accuracy caused by relying solely on the model itself for training.

[0027] In one embodiment of this application, the process of determining the parameters of the target intent recognition model includes: In response to the fact that the number of times the voice device receives the instruction voice is less than or equal to the number of times it receives the first voice, the parameters of the target intent recognition model are set to the default parameters; In response to the fact that the number of times the voice device receives command voice is greater than the first time and less than the second time, and the voice device does not receive model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the historical command voice issued by the user corresponding to the voice device and the historical environment information corresponding to the issuance of the historical command voice. In response to the voice device receiving a command voice a number greater than or equal to the second number, and the voice device receiving model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the model update parameters sent by the central processing unit. The first number is less than the second number.

[0028] In this embodiment, the number of times the voice device receives commanding voice can be used as a standard to determine which of the three situations the parameters of the target intent recognition model corresponding to the voice device fall into.

[0029] This can be understood as follows: when the number of times the voice device receives instructional voice commands is less than or equal to the number of times it receives them for the first time, it means that the user has used the device to give commands less often, making effective training impossible. Therefore, when there is insufficient user behavior data, basic intent recognition capabilities are provided to avoid overfitting due to a small amount of data.

[0030] When the number of times the voice device receives commanding voice messages is greater than the first time but less than the second time, it indicates that a certain amount of basic data has been obtained and can be used for iterating the parameters of the target intent recognition model corresponding to the voice device. Simultaneously, during the iteration process of the target intent recognition model's parameters, no model update parameters are sent by the central processing unit. At this point, the parameters of the target intent recognition model can be determined based on the historical commanding voice messages issued by the user corresponding to the voice device and the historical environment information corresponding to the issuance of those messages. In this process, the goal is to minimize the cross-entropy loss of intent classification. The parameter determination process can be understood as an iterative process of model parameters. Those skilled in the art can set the loss function or other parameters during the model parameter iteration process based on actual application scenarios and requirements; this application embodiment does not limit or elaborate on these details.

[0031] When the voice device receives command voice messages a number of times greater than or equal to the second number, and the voice device receives model update parameters sent by the central processing unit, it indicates that the number of user commands is sufficient and the device has received model update parameters from the central processing unit. At this time, the model update parameters sent by the central processing unit can be used as the parameters of its corresponding target intent recognition model.

[0032] In this embodiment, the model update parameters can be understood as an instruction and a set of parameters. The instruction indicates that the voice device receiving this information needs to perform a model update operation. The parameters can be understood as model parameters sent by the central processing unit that may match the user of the voice device.

[0033] In this embodiment, another scenario exists: If the number of times the voice device receives commanding voice messages is greater than or equal to the second count, and the voice device has not received model update parameters from the central processing unit, the parameters of the target intent recognition model are determined based on the historical commanding voice messages issued by the user corresponding to the voice device and the historical environment information corresponding to the issuance of those messages. This can be understood as the voice device having sufficient training data for its target intent recognition model, but not receiving model update parameters from the central server. In this case, it indicates that the target intent recognition model can complete the intent recognition of its corresponding user, and no model parameter update is needed. The first and second counts can be set based on experience or actual application scenarios.

[0034] As can be seen from the above, this application considers the situation where there is insufficient instruction data in the early stages of user device use, avoiding the overfitting problem caused by training the model with a small amount of data. It provides users with basic intent recognition capabilities, ensuring that the system can operate normally in the initial stage and will not fail to effectively recognize user intent due to data scarcity. This application utilizes the central processing unit's comprehensive analysis capabilities of data from multiple voice devices, ensuring that the model can obtain more accurate and optimized parameters in a timely manner, thereby improving the overall intent recognition accuracy of the system.

[0035] In one embodiment of this application, the process of multiple voice devices interacting with a central processing unit includes: multiple voice devices sending information to the central processing unit; The process of sending information from multiple voice devices to a central processing unit includes: For each voice device, in response to the voice device being a target voice device, the usage characteristics of the user corresponding to the voice device are desensitized to obtain the desensitized characteristics corresponding to the voice device; wherein, the target voice device is the device that receives command voice more than three times; The parameters of the target intent recognition model corresponding to the voice device, the de-identification features corresponding to the voice device, and the historical task accuracy corresponding to the voice device are sent to the central processing unit; wherein, the historical task accuracy corresponding to the voice device is determined based on the proportion of times the user corresponding to the voice device makes negative voices, and the command voices include negative voices.

[0036] In this embodiment, the third number should be greater than or equal to the first number. That is, only when the number of command voices received by the voice device is large can the parameters of the corresponding target intent recognition model be sent to the central processing unit. At the same time, the user usage characteristics corresponding to the voice device and the historical task accuracy corresponding to the voice device should also be sent.

[0037] In this embodiment, the user's usage characteristics may include instruction time distribution, frequency of commonly used keywords, environmental characteristics, etc. The desensitization method may be generalization, perturbation, or hashing, etc., with the aim of preventing the leakage of user information and enabling the central processing unit to obtain some data that can represent data characteristics.

[0038] In this embodiment, the accuracy of historical tasks can be determined based on the proportion of negative speech. Negative speech can be words that indicate a denial of the previous answer, such as "not this" or "that's not what I meant." It is essentially a kind of instruction speech because when the target intent recognition model receives negative speech, it will trigger the previous instruction speech again, while adding a limitation to make the direction different from the previous answer.

[0039] In one embodiment of this application, the process of multiple voice devices interacting with a central processing unit further includes: the central processing unit determining model update parameters based on the information it receives; The process by which the central processing unit determines the model update parameters based on the information it receives includes: The federated learning method is determined based on the data distribution of de-identified features received from all target voice devices; In response to the federated learning method being horizontal federated learning, the weights of each target speech device are determined based on the historical task accuracy of each target speech device and the de-identification features of each target speech device. The parameters of the target intent recognition model sent by each target voice device are weighted and calculated based on the weights of each target voice device to determine the model update parameters.

[0040] In this embodiment, the central processing unit (CPU) can determine model update parameters based on the information it receives. Specifically, the CPU analyzes the anonymized features uploaded by all target voice devices and determines the degree of overlap in the feature space, i.e., the data distribution of the anonymized features. For example: The federated learning method is determined based on the data distribution of de-identified features received from all target voice devices, including: If the proportion of desensitized features of the same dimension among all the desensitized features sent by the target voice devices is greater than the preset proportion, the horizontal federated learning method is determined as the federated learning method of the central processing unit.

[0041] In this embodiment, if the proportion of features with the same dimension in the de-identified features of most devices exceeds the preset proportion, it indicates that the feature spaces between devices are similar and only the samples are different (such as the same type of instructions from different users). In this case, horizontal federated learning can be used. Otherwise, vertical federated learning can be used. In this scenario, horizontal federated learning is used in most cases because the differences between voice devices and user instructions are not obvious.

[0042] If the federated learning method is a horizontal federated learning method, the central processing unit should perform weighted calculations on the parameters of all target intent recognition models it receives, and finally determine a model update parameter. The weight setting should be based on the historical task accuracy of each target voice device and the de-identification features of each target voice device. The specific determination process is described below and is not limited in this embodiment.

[0043] As can be seen from the above, this application can prevent the leakage of users' sensitive information through the desensitization process, while ensuring that the central processing unit can obtain data that can represent data features for model optimization. By integrating data from multiple voice devices, this application can make full use of the data resources of different devices and different users, making the model update parameters more comprehensive and accurate, thereby improving the performance and accuracy of the target intent recognition model and providing users with better voice interaction services.

[0044] In one embodiment of this application, the process of multiple voice devices interacting with a central processing unit further includes: the central processing unit sending information to the multiple voice devices; The process of sending information from a central processing unit to multiple voice devices includes: For each target speech device, the system determines whether to send model update parameters to that target speech device based on the historical task accuracy of that target speech device; wherein the model update parameters are parameters determined by the central processing unit based on the information it receives.

[0045] In this embodiment, after obtaining the model update parameters through the aforementioned horizontal federated learning of the central processing unit, the central processing unit will evaluate each target voice device and determine whether to send the model update parameters to the target voice device. For example, if the historical task accuracy of the target voice device is greater than or equal to a preset accuracy, the model update parameters will not be sent to the target voice device. In response to the fact that the historical task accuracy of the target voice device is less than the preset accuracy, model update parameters are sent to the target voice device.

[0046] In this embodiment, when the historical task accuracy of the target voice device is greater than or equal to a preset accuracy, it indicates that the target voice device can currently perform the corresponding user's intent recognition well. In this case, there is no need for the central processing unit to send model update parameters, preventing the model update from actually reducing task accuracy. The preset accuracy can be 0.9, and can be adjusted based on experience or actual needs. Alternatively, the preset accuracy can be determined based on the accuracy of all received historical tasks, for example, it can be set as the average of the historical task accuracies.

[0047] Conversely, if the historical task accuracy of the target voice device is less than the preset accuracy, it means that the target voice device cannot perform the corresponding user intent recognition task well. In this case, it can be updated based on the model update parameters sent by the central processing unit.

[0048] As can be seen from the above, the central processing unit of this application determines whether to send model update parameters to the target voice device based on the historical task accuracy corresponding to the target voice device. This avoids sending unnecessary update parameters to devices that can already perform intent recognition tasks well, saves system resources, including network bandwidth and computing resources, improves the overall operating efficiency of the system, and prevents the risk of performance degradation that may be caused by model updates. This ensures the stability and reliability of the system and provides users with continuous and high-quality intent recognition services.

[0049] In one embodiment of this application, the weight of each target voice device is determined based on the received historical task accuracy of each target voice device and the de-identification features of each target voice device, including: For each target voice device, the historical task accuracy and the de-identified features corresponding to each target voice device are used to determine the first reliability score based on the number of de-identified features sent by the target voice device. A second reliable score is determined based on the historical task accuracy of the target voice device. The first reliability score and the second reliability score are weighted and calculated to obtain the comprehensive reliability score corresponding to the target voice device; The weight of the target voice device is determined based on the comprehensive reliability score corresponding to the target voice device.

[0050] In this embodiment, the more desensitized features the target voice device sends, the larger the training dataset of the target voice device's target intent recognition model is. With more training data, it can be said that the reliability of the parameters of the target voice device's target intent recognition model is higher. Similarly, the higher the accuracy of the historical tasks corresponding to the target voice device, the higher the reliability of the parameters of the target voice device's target intent recognition model is.

[0051] In this embodiment, the first reliability score can be determined based on a preset mapping table or linear relationship. Alternatively, the first reliability score can be calculated based on a first formula, which may be: ,in, Indicates the first reliable rating. Indicates the first The number of de-identification features for each voice device This represents the maximum number of features across all voice devices. Indicates the exponential adjustment factor. Indicates the characteristic diversity factor, Indicates the timeliness factor of the characteristics.

[0052] ,in, Indicates the first The proportion of class features Indicates the total number of feature types.

[0053] ,in This represents the attenuation coefficient, which can be set to 0.2. Indicates the update time of the latest feature. This indicates the half-life, which can be set to 7 (days). This represents the natural constant. It should be noted that the units for the update time and half-life mentioned above are for illustrative purposes only and are calculated numerically during the actual calculation process.

[0054] In the first formula, The number of equipment features is normalized to the 0-1 range, and the curve shape is adjusted using an exponential factor. To assess the balance of feature types, avoid a single type of feature dominating the score. This indicates the value of decaying or obsolete features, ensuring that the score reflects the real-time validity of the data.

[0055] The second reliability score can be determined based on a pre-determined mapping table or a preset linear relationship. The weights of the first and second reliability scores can be determined based on experience, with priority given to the weight of the second reliability score being greater than the weight of the first reliability score.

[0056] As can be seen from the above, this application determines the weight of the device by comprehensively considering two dimensions: the number of anonymized features sent by the target voice device and the accuracy of historical tasks. This allows for a more comprehensive evaluation of the reliability of the target voice device, helps improve the quality of model update parameters determined based on this device data, and ultimately improves the performance of the entire voice interaction system. The first reliability score is determined based on the number of anonymized features sent by the target voice device and considers the feature diversity factor and the feature timeliness factor. The feature diversity factor measures the balance of feature types, avoiding the dominance of a single type of feature in the score, and can more comprehensively reflect the richness of the device's training data. The feature timeliness factor represents the value of decaying old features, ensuring that the score reflects the real-time validity of the data. By considering these data characteristics, the first reliability score can more accurately reflect the quality and reliability of the device's training data, providing a more accurate basis for the calculation of the comprehensive reliability score.

[0057] Corresponding to the voice-based user intent recognition method in the above embodiments, Figure 2 This is a structural block diagram of a voice-based user intent recognition device according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. Reference Figure 2The voice-based user intent recognition device 20 is applied to each voice device in the intent recognition device, which includes: multiple voice devices and a central processing unit. The multiple voice devices interact with the central processing unit. The device includes: an environmental information acquisition module 21, a text determination module 22, and an intent recognition module 23. Among them, the environmental information acquisition module 21 is used to acquire environmental information in response to receiving the initiation voice; The text determination module 22 is used to determine the first text based on the first voice; wherein, the first voice is the command voice issued by the user corresponding to the voice device after issuing the initiation voice, and the first text is the text corresponding to the first voice; The intent recognition module 23 is used to input environmental information, first speech and first text into the target intent recognition model to obtain the intent of the first speech issued by the user; wherein, the parameters of the target intent recognition model are default parameters, or parameters determined based on the historical command speech issued by the user corresponding to the speech device, or parameters determined based on the model update parameters sent by the central processing unit, and the model update parameters are determined based on information sent by multiple speech devices.

[0058] In one embodiment of this application, a voice-based user intent recognition device 20 includes: a parameter determination module for a target intent recognition model, configured to determine the parameters of the target intent recognition model as default parameters in response to the voice device receiving instructional voice less than or equal to the first time. In response to the fact that the number of times the voice device receives command voice is greater than the first time and less than the second time, and the voice device does not receive model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the historical command voice issued by the user corresponding to the voice device and the historical environment information corresponding to the issuance of the historical command voice. In response to the voice device receiving a command voice a number greater than or equal to the second number, and the voice device receiving model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the model update parameters sent by the central processing unit. The first number is less than the second number.

[0059] In one embodiment of this application, the voice-based user intent recognition device 20 includes an information interaction module for sending information from multiple voice devices to a central processing unit; The information interaction module is specifically used to perform desensitization processing on the usage characteristics of the user corresponding to each voice device in response to the voice device being the target voice device, thereby obtaining the desensitized characteristics corresponding to the voice device; wherein, the target voice device is a device that has received command voice messages more than three times. The parameters of the target intent recognition model corresponding to the voice device, the de-identification features corresponding to the voice device, and the historical task accuracy corresponding to the voice device are sent to the central processing unit; wherein, the historical task accuracy corresponding to the voice device is determined based on the proportion of times the user corresponding to the voice device makes negative voices, and the command voices include negative voices.

[0060] In one embodiment of this application, the information interaction module is further configured to determine model update parameters based on the information received by the central processing unit; The information interaction module is also specifically used to determine the federated learning method based on the data distribution of desensitized features sent by all target voice devices. In response to the federated learning method being horizontal federated learning, the weights of each target speech device are determined based on the historical task accuracy of each target speech device and the de-identification features of each target speech device. The parameters of the target intent recognition model sent by each target voice device are weighted and calculated based on the weights of each target voice device to determine the model update parameters.

[0061] In one embodiment of this application, the information interaction module is further configured to send information from the central processing unit to multiple voice devices. The information interaction module is further used to determine whether to send model update parameters to each target voice device based on the historical task accuracy of that target voice device; wherein the model update parameters are parameters determined by the central processing unit based on the information it receives.

[0062] In one embodiment of this application, the information interaction module is further configured to determine the horizontal federated learning method as the federated learning method of the central processing unit in response to the fact that the proportion of desensitized features of the same dimension among the desensitized features sent by all target voice devices is greater than a preset proportion.

[0063] In one embodiment of this application, the information interaction module is further configured to determine a first reliability score based on the number of desensitized features sent by the target voice device, using the historical task accuracy corresponding to each target voice device and the desensitized features corresponding to each target voice device. A second reliable score is determined based on the historical task accuracy of the target voice device. The first reliability score and the second reliability score are weighted and calculated to obtain the comprehensive reliability score corresponding to the target voice device; The weight of the target voice device is determined based on the comprehensive reliability score corresponding to the target voice device.

[0064] See Figure 3 , Figure 3 This is a schematic block diagram of an intent recognition device provided in an embodiment of this application. Figure 3 The intent recognition device 300 shown in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the above-described device embodiments, for example... Figure 2 The functions of the environmental information acquisition module 21, text determination module 22, and intent recognition module 23 are shown.

[0065] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0066] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0067] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store the first number and the second number.

[0068] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the embodiments of the user intent recognition method based on voice provided in the embodiments of this application, or they can execute the implementation methods of the intent recognition device described in the embodiments of this application, which will not be repeated here.

[0069] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0070] The computer-readable storage medium can be an internal storage unit of the intent recognition device in any of the foregoing embodiments, such as a hard disk or memory of the intent recognition device. The computer-readable storage medium can also be an external storage device of the intent recognition device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the intent recognition device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the intent recognition device. The computer-readable storage medium is used to store computer programs and other programs and data required by the intent recognition device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0071] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0072] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the intent recognition device and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0073] In the several embodiments provided in this application, it should be understood that the disclosed intent recognition device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0074] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0075] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0076] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice-based user intent recognition method, characterized in that, The method is applied to each voice device in an intent recognition device, the intent recognition device comprising: multiple voice devices and a central processing unit, wherein each of the multiple voice devices interacts with the central processing unit, the method comprising: In response to receiving the initiation voice, obtain environmental information; First text is determined based on first voice; wherein, the first voice is the command voice issued by the user corresponding to the voice device after issuing the activation voice, and the first text is the text corresponding to the first voice; The environmental information, the first voice, and the first text are input into the target intent recognition model to obtain the intent of the first voice issued by the user; wherein, the parameters of the target intent recognition model are default parameters, or parameters determined based on the historical command voice issued by the user corresponding to the voice device, or parameters determined based on the model update parameters sent by the central processing unit, and the model update parameters are determined based on the information sent by the multiple voice devices; The process of determining the parameters of the target intent recognition model includes: In response to the fact that the number of times the voice device receives the instruction voice is less than or equal to the number of times it receives the first voice, the parameters of the target intent recognition model are determined to be the default parameters; In response to the fact that the number of times the voice device receives command voice is greater than the first number and less than the second number, and the voice device does not receive the model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the historical command voice issued by the user corresponding to the voice device and the historical environment information corresponding to the issuance of the historical command voice. In response to the voice device receiving command voice a number of times greater than or equal to the second number, and the voice device receiving model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the model update parameters sent by the central processing unit; The first number is less than the second number.

2. The user intent recognition method based on voice as described in claim 1, characterized in that, The process of the multiple voice devices interacting with the central processing unit includes: the multiple voice devices sending information to the central processing unit; The process of sending information from the plurality of voice devices to the central processing unit includes: For each voice device, in response to the voice device being a target voice device, the usage characteristics of the user corresponding to the voice device are desensitized to obtain the desensitized characteristics corresponding to the voice device; wherein, the target voice device is a device that receives command voice messages more than three times. The parameters of the target intent recognition model corresponding to the voice device, the de-identification features corresponding to the voice device, and the historical task accuracy corresponding to the voice device are sent to the central processing unit; wherein, the historical task accuracy corresponding to the voice device is determined based on the proportion of times the user corresponding to the voice device issues negative voice, and the instruction voice includes the negative voice.

3. The user intent recognition method based on voice as described in claim 2, characterized in that, The process of the multiple voice devices interacting with the central processing unit also includes: the central processing unit determining model update parameters based on the information it receives; The process by which the central processing unit determines the model update parameters based on the information it receives includes: The federated learning method is determined based on the data distribution of de-identified features received from all target voice devices; In response to the federated learning method being horizontal federated learning, the weights of each target speech device are determined based on the historical task accuracy of each target speech device and the de-identification features of each target speech device. The parameters of the target intent recognition model sent by each target voice device are weighted and calculated based on the weights of each target voice device to determine the model update parameters.

4. The user intent recognition method based on voice as described in claim 2, characterized in that, The process of the multiple voice devices interacting with the central processing unit also includes: the central processing unit sending information to the multiple voice devices; The process of sending information from the central processing unit to the plurality of voice devices includes: For each target speech device, a determination is made based on the historical task accuracy corresponding to that target speech device to determine whether to send model update parameters to that target speech device; wherein the model update parameters are parameters determined by the central processing unit based on the information it receives.

5. The voice-based user intent recognition method as described in claim 3, characterized in that, The method of determining the federated learning approach based on the data distribution of de-identified features received from all target voice devices includes: If the proportion of desensitized features of the same dimension among all the desensitized features sent by the target voice devices is greater than the preset proportion, the horizontal federated learning method is determined as the federated learning method of the central processing unit.

6. The voice-based user intent recognition method as described in claim 3, characterized in that, The step of determining the weight of each target voice device based on the received historical task accuracy of each target voice device and the de-identification features of each target voice device includes: For each target voice device, the historical task accuracy and the de-identified features corresponding to each target voice device are used to determine the first reliability score based on the number of de-identified features sent by the target voice device. A second reliable score is determined based on the historical task accuracy of the target voice device. The first reliability score and the second reliability score are weighted and calculated to obtain the comprehensive reliability score corresponding to the target voice device; The weight of the target voice device is determined based on the comprehensive reliability score corresponding to the target voice device.

7. A voice-based user intent recognition device, characterized in that, An intent recognition device is applied to each voice device in an intent recognition device, the intent recognition device comprising: multiple voice devices and a central processing unit, wherein the multiple voice devices all interact with the central processing unit, the device comprising: The environmental information acquisition module is used to acquire environmental information in response to receiving the initiation voice. A text determination module is used to determine a first text based on a first voice; wherein the first voice is an instruction voice issued by the user corresponding to the voice device after issuing the activation voice, and the first text is the text corresponding to the first voice; An intent recognition module is used to input the environmental information, the first voice, and the first text into a target intent recognition model to obtain the intent of the first voice issued by the user; wherein, the parameters of the target intent recognition model are default parameters, or parameters determined based on historical command voices issued by the user corresponding to the voice device, or parameters determined based on model update parameters sent by the central processing unit, and the model update parameters are determined based on information sent by the multiple voice devices; The parameter determination module of the target intent recognition model is used to determine the parameters of the target intent recognition model as default parameters in response to the fact that the number of times the voice device receives the instruction voice is less than or equal to the first time. In response to the fact that the number of times the voice device receives command voice is greater than the first number and less than the second number, and the voice device does not receive the model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the historical command voice issued by the user corresponding to the voice device and the historical environment information corresponding to the issuance of the historical command voice. In response to the voice device receiving command voice a number of times greater than or equal to the second number, and the voice device receiving model update parameters sent by the central processing unit, the parameters of the target intent recognition model are determined based on the model update parameters sent by the central processing unit; The first number is less than the second number.

8. An intent recognition device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Feature extraction model training method and device, equipment, medium and program product

    CN114912542A

  • Vehicle cabin voice intention recognition method and device and vehicle control method

    CN117854493A

  • Voice interaction method and device, and storage medium

    CN119580707A

  • Dialogue intention recognition method and system, electronic equipment and storage medium

    CN120354860A