Action determination method and device, computer equipment, readable storage medium and program product
By receiving and converting audio signals into text information, combining semantic analysis and contextual understanding, user needs are judged and inferred, so as to accurately match the target execution action from the preset action library when the robot performs actions, solving the problem of inaccurate selection of action libraries in the prior art, and improving the accuracy of robot action execution and personalized response capabilities.
Patent Information
- Application Number
- CN202510473550.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the prior art, it is difficult for robots to accurately select appropriate actions from predefined action databases when performing actions, especially in scenarios where personalized responses and psychic activities are required for execution.
By receiving audio signals, converting them into text information, and using semantic analysis and context understanding to determine whether the text information contains action instructions. If not, corresponding answer information is generated and the target execution action is matched from the preset action library.
It realizes that when the text information does not contain action instructions, user needs are inferred based on semantic analysis and context understanding, and efficiently match the corresponding target execution actions, improving the robot's action execution accuracy and personalized response capabilities in different scenarios.
Smart Images

Figure CN119993151A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large model technology, and in particular to an action determination method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of artificial intelligence technology, the application of robots in different scenarios is gradually increasing, especially in service, education, medical and other scenarios. Robots not only need to meet basic task execution, but also need to have flexible and personalized response behaviors that meet user expectations in order to determine the corresponding user needs.
[0003] Currently, corresponding actions are executed through a predefined action library, so how to select a suitable action from the action library is a problem to be solved. Summary of the invention
[0004] Based on this, it is necessary to provide an action determination method, device, computer equipment, computer-readable storage medium and computer program product that can accurately determine the robot's target execution action from a dialogue to address the above technical problems.
[0005] In a first aspect, the present application provides an action determination method, the method comprising:
[0006] receiving an audio signal;
[0007] Converting the audio signal into text information;
[0008] If the text information does not include an action instruction, generating answer information corresponding to the text information, and determining a target execution action corresponding to the answer information from a preset action library;
[0009] The target execution action is executed, and the response information is output.
[0010] In one embodiment, the process of determining whether the text information includes the action instruction includes:
[0011] If the text information does not conform to a preset sentence structure, analyzing the text information to identify verbs or task-related words in the text information;
[0012] Filtering historical interaction information through a preset window to obtain target interaction information related to the verb or the task-related word;
[0013] The user intention is obtained according to the verb or the task-related word and the target interaction information; the user intention is used to determine whether the text information includes the action instruction.
[0014] In one embodiment, the method further comprises:
[0015] If the text information includes an action instruction, the target execution action is determined by matching the verb in the preset action library.
[0016] In one embodiment, generating answer information corresponding to the text information and determining a target execution action corresponding to the answer information from a preset action library includes:
[0017] generating the answer information according to the user intention;
[0018] Inputting the answer information into a pre-trained multi-classification model to obtain a classification result;
[0019] The target execution action is determined by matching the classification result with the preset action library.
[0020] In one embodiment, the training process of the multi-classification model includes:
[0021] Acquire first sample data; the first sample data carries an action instruction tag;
[0022] Preprocessing the first sample data to obtain multiple word segments;
[0023] The first sample data is input into the first initial model for training to obtain a prediction classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each of the instruction classification heads includes an input layer, an encoding layer and a pooling layer; the input layer converts the word segmentation into a word vector, and transmits the word vector to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and transmits the initial features to the pooling layer; the pooling layer compresses the initial features to obtain the target features; the fully connected layer merges the target features corresponding to each of the classification heads to obtain the prediction classification result;
[0024] The predicted classification result is compared with the action instruction label to obtain a first difference, and the parameters of the first initial model are adjusted according to the first difference until the first initial model is trained to obtain the multi-classification model.
[0025] In one embodiment, the conversion of the audio signal into text information is obtained by a pre-trained speech recognition model; the training process of the speech recognition model includes:
[0026] Acquire second sample data; the second sample data carries a text label;
[0027] adding background noise to the second sample data; the background noise includes one or more of a TV playing sound, a kitchen sound, and a conversation sound;
[0028] Inputting the second sample data carrying the background noise into a second initial model for training to obtain predicted text data;
[0029] A second difference between the predicted text data and the text label is compared, and the parameters of the second initial model are adjusted according to the second difference until the second initial model is trained to obtain the speech recognition model.
[0030] In a second aspect, the present application provides an action determination device, the device comprising:
[0031] A receiving module, used for receiving an audio signal;
[0032] A text conversion module, used for converting the audio signal into text information;
[0033] an action determination module, for generating answer information corresponding to the text information if the text information does not include an action instruction, and determining a target execution action corresponding to the answer information;
[0034] The execution module is used to execute the target execution action and output the response information.
[0035] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0036] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0037] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0038] The above-mentioned action determination method, apparatus, computer device, computer-readable storage medium and computer program product, when determining that the text information does not include action instructions, infer user needs based on semantic analysis and context understanding, and generate corresponding answer information, and efficiently match the corresponding target execution action from the preset action library based on the answer information. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 is a flow chart of an action determination method in one embodiment;
[0041] Figure 2 is a flow chart of an action determination method in another embodiment;
[0042] Figure 3 is a structural block diagram of an action determination device in one embodiment;
[0043] Figure 4 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0045] In one embodiment, Figure 1 As shown, an action determination method is provided. This embodiment takes the method applied to a terminal as an example for illustration. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0046] Step S102: receiving an audio signal.
[0047] Optionally, an audio signal from an external environment may be received through an audio acquisition device, wherein the audio signal may be a user's voice command, background noise, or other environmental sounds.
[0048] Furthermore, in order to ensure that the system can accurately recognize voice commands from mixed information, the system is usually equipped with a noise suppression module that can remove or weaken the interference of environmental noise on voice recognition. Among them, the noise suppression module can focus on receiving sound signals from a specific direction through microphone array technology and beamforming algorithm, thereby reducing the interference of background noise.
[0049] Optionally, the robot can be equipped with a digital signal processor (DSP) or a dedicated speech recognition chip. These hardware modules are responsible for preprocessing and noise reduction to ensure that the quality of the incoming speech signal meets the requirements of speech recognition.
[0050] Optionally, the audio signal is first collected by a microphone array, then amplified by a preamplifier, and converted into a digital signal by an ADC (analog to digital converter). The digital signal then enters a processing unit (such as a DSP chip) for preliminary noise elimination and echo suppression. Finally, the processed signal is transmitted to the speech recognition module to be converted into text information.
[0051] Step S104: convert the audio signal into text information.
[0052] Among them, text information is text data converted from audio signals through voice recognition technology, including the content of voice commands.
[0053] Optionally, the received audio signal is processed by a speech recognition model and converted into text information. The speech recognition model can be trained based on a deep learning model, such as a convolutional neural network, a recurrent neural network, etc., to recognize different pronunciations and speech features.
[0054] Optionally, before the audio signal enters the speech recognition model, a series of preprocessing steps will be performed, such as denoising, speech segmentation, feature extraction, etc.
[0055] Optionally, the speech recognition model can extract the features of the audio signal through methods such as MFCC (Mel-Frequency Cepstral Coefficients), and then use the Long Short-Term Memory (LSTM) or Transformer model for language modeling, and finally output text content.
[0056] Alternatively, assuming that the audio signal is "turn on the lights in the living room", the speech recognition module first converts the frequency and time domain features of the speech signal into Mel-frequency cepstral coefficient (MFCC) features and inputs them into the deep neural network. The neural network extracts high-dimensional features through the convolution layer, then performs sequence modeling through the recurrent neural network (RNN) or Transformer, and finally outputs the corresponding text: "turn on the lights in the living room".
[0057] Step S106: if the text information does not include an action instruction, generate answer information corresponding to the text information, and determine a target execution action corresponding to the answer information.
[0058] Among them, action instructions refer to instructions issued by users through voice, text or other input methods, requiring the robot to perform a specific task. Action instructions not only include the actions that the robot needs to perform, such as "open", "close", "move", etc., but may also include detailed task information related to the action, such as object, location, time, etc.
[0059] Among them, the target execution action refers to the robot determining the specific task or action to be performed based on the analysis of text information, such as controlling a robotic arm, making a sound, etc.
[0060] Optionally, whether the action instruction is included may be judged by a pre-trained judgment model, wherein the judgment model is a binary classification model, and its output result is yes or no.
[0061] Furthermore, the judgment model may judge whether the text information includes action instructions through context semantics, and the details may refer to the description in the following embodiments.
[0062] Step S108, executing the target execution action and outputting answer information.
[0063] Optionally, once the target action is determined, the system will perform specific tasks through a mechanical control system, such as motor drive, sensor feedback, etc. For example, if the target action is "turn on the light", the system will control the power switch through a relay; if it is "carrying an object", it will be completed through a robotic arm.
[0064] Furthermore, when executing the target action, the robot usually contains multiple modules, such as motion control module, voice module, perception module, etc. Each module cooperates with each other to ensure that the task can be executed smoothly. For example, the motion control module adjusts the robot arm movement according to the instruction, and the voice module is responsible for generating and outputting the answer information.
[0065] Optionally, the task may need to be split into multiple subtasks and executed in parallel. For example, if the robot needs to answer questions and perform target execution actions at the same time, the system may need to send the answer part of the task to the language generation module and the action execution part of the task to the control module to complete the tasks in parallel.
[0066] In the above action determination method, when it is determined that the text information does not include action instructions, the user needs are inferred based on semantic analysis and context understanding, and corresponding answer information is generated, and the corresponding target execution action is efficiently matched from the preset action library based on the answer information.
[0067] Furthermore, in one embodiment, the process of determining whether the above text information includes action instructions includes: if the text information does not conform to a preset sentence structure, analyzing the text information to identify verbs or task-related words in the text information; filtering historical interaction information through a preset window to obtain target interaction information; obtaining user intent based on verbs or task-related words, and target interaction information; and using user intent to determine whether the text information includes action instructions.
[0068] The preset sentence structure refers to a predefined sentence structure template based on a previously trained model or rule library, and this template is used to identify core elements and action instructions in the text.
[0069] For example, if the scenario usage scenario is family, the preset sentence structure can be as follows:
[0070] Sentence structure 1: please [verb] [task], such as "please open the door"
[0071] Sentence structure 2: [verb] [object], such as "Close the window"
[0072] Among them, task-related words refer to keywords that are directly related to the execution of the task. They describe the object, target, or background of the operation. For example, in a smart home scenario, if the user issues a command such as "turn up the temperature", the task-related words include "turn up" and "temperature". Since when the user has multiple rounds of dialogue or has established a certain context, the object of turning up the temperature can be inferred by combining the context.
[0073] The preset window refers to the time range in which the system filters historical interaction information when determining whether a text message contains an action instruction. This window helps the system better understand the user's intention, especially in multi-round dialogue scenarios. Based on the time limit, the system only considers historical interaction information within a certain time range. For example, if the user recently requested "turn on the air conditioner" and then said "the temperature is a bit low" or "turn up the temperature", the system will filter out the target interaction information "turn on the air conditioner" from the historical interaction information according to the preset window, and infer that the user may want to adjust the air conditioner temperature. Among them, the target interaction information refers to the key information related to the current task that is filtered out from the historical interaction. The target interaction information can help the system accurately understand the current user's intention and perform appropriate operations. The system records each interaction between the user and the robot and continuously updates the context.
[0074] Finally, combining the verbs, task-related words and historical interaction information in the text, the system can comprehensively judge the user's intentions. In multi-round conversations, the system not only relies on the current instructions, but also infers the user's needs based on previous interactions. For example, if the user says "turn on the air conditioner" in the first interaction, and then mentions "the temperature is a bit low", the system will analyze the "air conditioning" and "low temperature" information in the historical interactions and infer that the user's need is to adjust the temperature. Therefore, the system can accurately judge whether the text information contains action instructions based on this comprehensive information and perform the corresponding actions.
[0075] Furthermore, the judgment process of whether the text information in this embodiment includes action instructions can be obtained through a pre-trained judgment model, and the initial model of the judgment model is a large language model, such as BERT (Bidirectional Encoder Representations from Transformers) and GPT (Generative Pretrained Transformer).
[0076] Exemplarily, after obtaining the text information, the judgment model converts the text into a vector representation through the input layer, and then these vectors are sent to the encoding layer of the model. In the encoding layer, the model can use the self-attention mechanism to match these vectors with the preset sentence vectors to determine whether the text conforms to the predetermined sentence template. The preset sentence structure will be converted into a fixed vector embedding and stored in the parameters of the model. When the vector of the input text is compared with the preset sentence vector, the model can determine whether the input text conforms to a specific sentence structure. If the text information does not conform to the preset sentence structure, the model will proceed to the next step to extract verbs and task-related words. Exemplarily, the model will extract verbs and task-related words from the text through named entity recognition (NER) or dependency parsing. The initial model analyzes the relationship between words in the text through its built-in self-attention mechanism to determine which words are verbs and which words are nouns related to the task. The model not only relies on syntactic rules, but also uses contextual information to enhance the understanding of verbs and task-related words. For example, in "the temperature is a bit low", the model will identify "temperature" and "low" as task-related words, and infer that the user wants to adjust the temperature of the air conditioner. Next, combining verbs, task-related words, and historical interaction information, the model will perform user intent recognition. For multi-round dialogues, the system uses memory networks or context encoders to extract and store key information from historical interactions. This information is filtered through a preset window, and only relevant historical interaction information within the time range is retained. For example, when the user enters "the temperature is a bit low", the model will infer that the user wants to adjust the temperature of the air conditioner based on historical interaction records in the window, such as "turn on the air conditioner". Combined with the current task-related words "temperature" and "low", and the historical information "air conditioner", the model will be able to accurately recognize that the user's intention is to adjust the temperature. Finally, the model uses its deep multi-classification decision layer to determine whether the text information contains action instructions based on the identified intent and task.
[0077] In one embodiment, the method further includes: if the text information includes an action instruction, matching the verb with a preset action library to determine the target execution action.
[0078] The preset action class refers to a predefined database that contains various actions that the robot can perform and their corresponding instructions. For example, the following is an example of a preset action library.
[0079]
[0080] Among them, action description refers to the expression of the robot's action instructions in a standardized form in the system. Each action description corresponds to a specific behavior, such as "robot_wave" for waving, and "robot_move_forward" for moving forward one step. Action description is not only used to clarify the robot's behavior, but also to facilitate the system's scheduling and management of different actions, ensuring the accuracy and consistency of the execution process.
[0081] Among them, the action description is a standardized string or identifier used to refer to a specific behavior of the robot. Each action description should correspond one-to-one to the actual action of the robot to ensure that the system can clearly know how to perform a task.
[0082] For example, taking robot_wave as an example, the action description may actually include target module, action angle, and action amplitude, etc. The target module is used to determine which part performs the action; the action angle can be "from 0° to 45°", which specifies the angle range of the arm from the initial position to the swinging position; the action amplitude is the amplitude of the command hand movement, and the amplitude of the arm swing determines the range of motion of the arm.
[0083] The standardization of action descriptions facilitates the system's scheduling and management of actions, ensuring that there are no conflicts between different tasks and that the system can process different tasks sequentially or in parallel.
[0084] Optionally, in combination with the content in the above embodiment, the judgment model will judge whether the text information includes an action instruction. When the text information includes a text instruction, the judgment model will output "yes" and carry the corresponding verb. After that, the verb is matched with the preset action library to obtain the target execution action.
[0085] Exemplarily, a string comparison method may be used to confirm whether there is an exact match. A match is determined only when the keyword is exactly the same as an instruction in the preset action library or the difference is less than a preset threshold.
[0086] In this way, when the keyword matches the preset action library, the target execution action can be quickly determined.
[0087] Furthermore, in one embodiment, the above-mentioned generating answer information corresponding to the text information and determining the target execution action corresponding to the answer information include: generating the answer information according to the user's intention; inputting the answer information into a pre-trained multi-classification model to obtain a classification result; and matching the classification result with a preset action library to determine the target execution action.
[0088] When the text information does not include action instructions, the judgment model will output "no" and the corresponding user intention.
[0089] After obtaining the user's intention, the system will generate answer information based on the intention, and the answer information includes the system's response to the user's needs.
[0090] Next, the generated answer information is fed into a pre-trained multi-classification model. The model classifies the answer information according to pre-defined categories to determine which type of response the user's request belongs to. The multi-classification model has been extensively trained to identify the correct intent category based on the input information.
[0091] The classification result will indicate the type of response the user expects, such as whether a certain action needs to be performed. The system will then search for actions related to the category from the preset action library based on the classification result and match them. If the classification result indicates that the user's intention is to perform an action, the system will search for the corresponding execution action through the preset action library. For example, if the user's intention is to adjust the air conditioning temperature, the system will find the corresponding action from the action library, such as "increase the air conditioning temperature", and perform the action.
[0092] Furthermore, in one embodiment, the training process of the above-mentioned multi-classification model includes: obtaining first sample data; the first sample data carries an action instruction label; preprocessing the first sample data to obtain multiple word segmentations; inputting the first sample data into the first initial model for training to obtain a predicted classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each instruction classification head includes an input layer, an encoding layer and a pooling layer; the input layer converts the word segmentation into a word vector, and passes the word vector to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and sends the initial features to the pooling layer; the pooling layer compresses the initial features to obtain the target features; the fully connected layer merges the target features corresponding to each classification head to obtain a predicted classification result; the predicted classification result is compared with the action instruction label to obtain a first difference, and the parameters of the first initial model are adjusted according to the first difference until the first initial model is trained to obtain a multi-classification model.
[0093] First, obtain the first sample data, where each sample data contains the text information input by the user and the corresponding action instruction label. The action instruction label is a label that identifies the action type or task category corresponding to the text information, and can provide supervision information for the training process. The first sample data needs to be preprocessed, and the preprocessing step includes word segmentation, which decomposes the text information into multiple word units so that the model can better understand the structure and semantics in the text.
[0094] Next, the segmented data is input into the first initial model for training. The first initial model structure includes multiple instruction classification heads and a fully connected layer. Each instruction classification head mainly includes an input layer, an encoding layer, and a pooling layer. The function of the input layer is to convert the segmented text into word vectors, which represent the semantic information of each word in the text. The word vectors are passed as input to the encoding layer, where the model uses the self-attention mechanism to process the word vectors and obtain the initial feature representation of each word.
[0095] The self-attention mechanism allows the model to dynamically adjust the weights of each word based on the context, thereby better capturing the dependencies between words. For example, when a word depends on the context of another word, the model will enhance the relationship between the words, thereby taking these contextual information into account when generating features. The initial features output by the encoding layer are passed to the pooling layer, which compresses these features and integrates the feature information of each word into a more concise and expressive target feature. This target feature represents the semantic information of the entire sentence or text.
[0096] Then, the fully connected layer merges the target features from each classification head. The merged features are passed to the classification module, and finally a predicted classification result is generated. This result represents the model's predicted category for the input text information, that is, the corresponding action instruction type.
[0097] Next, the model's prediction is compared with the actual action instruction label to obtain the first difference. This difference represents the difference between the model's prediction and the actual label, reflecting the accuracy of the model. Based on this difference, the back propagation algorithm is used to adjust the parameters of the first initial model and optimize the model's weights to reduce the prediction error. The training process will continue until the difference reaches the predetermined tolerance range or the number of model training times reaches the set threshold.
[0098] After the training is completed, the first initial model will go through a preset number of parameter adjustments, and finally a trained multi-classification model will be obtained. The model can accurately predict the corresponding action instructions based on the new text information and provide a basis for subsequent task execution.
[0099] In one embodiment, the above-mentioned conversion of audio signals into text information is obtained through a pre-trained speech recognition model; the training process of the speech recognition model includes: obtaining second sample data; the second sample data carries a text label; adding background noise to the second sample data; the background noise includes one or more of TV playback sound, kitchen stereo sound and conversation sound; inputting the second sample data carrying background noise into the second initial model for training to obtain predicted text data; comparing the second difference between the predicted text data and the text label, and adjusting the parameters of the second initial model according to the second difference, until the second initial model is trained to obtain a speech recognition model.
[0100] In this embodiment, the process of converting audio signals into text information is achieved through a pre-trained speech recognition model. The speech recognition model helps the system understand and process voice commands by converting audio signals into text information. The model training process includes the steps of obtaining sample data, adding noise to enhance robustness, inputting data for training, and adjusting model parameters according to errors.
[0101] First, the training process starts with obtaining the second sample data. The second sample data consists of an audio signal and its corresponding text label, where the text label is an accurate transcription of the audio signal and represents the actual text information corresponding to the audio. These second sample data are the supervised data used to train the model. In order to enhance the adaptability and robustness of the model in real environments, these data are preprocessed during the training process, including adding background noise to the audio signal.
[0102] The purpose of adding background noise is to simulate the noise environment in daily life and enhance the anti-interference ability of the speech recognition model. Background noise can include one or more noise sources such as TV playing sound, kitchen sound, and conversation sound. In this way, the model can still maintain good recognition accuracy in a noisy environment. The process of adding background noise will change the characteristics of the audio signal, allowing the model to learn how to perform speech recognition in a complex environment.
[0103] After adding background noise, the second sample data with noise is input into the second initial model for training. The second initial model usually includes an acoustic feature extraction layer, an acoustic model layer, and a decoding layer. First, the acoustic feature extraction layer extracts useful feature information from the audio signal, such as Mel-frequency cepstral coefficients (MFCC), filter bank features, etc. These feature information are used to characterize the frequency and time domain features of the audio signal, providing basic data for subsequent speech recognition.
[0104] Next, the acoustic model layer processes these feature information and models them through deep neural networks (such as RNN, LSTM, etc.) to capture the timing information and speech regularity in the audio signal. The decoding layer maps the processed features to text output, that is, the predicted text data. This predicted text data represents the model's transcription result of the input audio signal.
[0105] The model then compares the predicted text data with the actual text label and calculates the second difference, which is the error between the predicted result and the actual label. This difference reflects the prediction accuracy of the model. The smaller the error, the better the performance of the model. Based on the second difference, the back propagation algorithm is used to adjust the parameters of the second initial model. Back propagation optimizes the recognition ability of the model by calculating the gradient and updating the weights of the model.
[0106] This training process will continue until the error of the second initial model reaches the predetermined tolerance range, or after a preset number of training times, the model's prediction results gradually approach the actual text labels. Finally, the second initial model will be adjusted multiple times to obtain a trained speech recognition model. The speech recognition model can accurately convert audio signals into text information under different environmental and noise conditions, thereby providing text input for subsequent task execution.
[0107] Exemplary, combined Figure 2 , Figure 2 The present invention is a flowchart of steps of an action determination method in one embodiment.
[0108] After obtaining the text information, the system directly inputs the text information into the binary classification model, that is, the judgment model in the above embodiment, to determine whether the text information includes action instructions. If the action instructions are included, it is directly matched with the preset action library to determine the target execution action. If the action instructions are not included, the reply is understood and output in a free dialogue, and the dialogue is input into the multi-classification model to determine the target execution action in the preset action library.
[0109] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0110] Based on the same inventive concept, the embodiment of the present application also provides an action determination device for implementing the above-mentioned action determination method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in the one or more action determination device embodiments provided below can refer to the limitations of the action determination method above, and will not be repeated here.
[0111] In an exemplary embodiment, Figure 4 As shown, an action determination device is provided, comprising: a receiving module 100, a text conversion module 200, an action determination module 300 and an execution module 400, wherein:
[0112] The receiving module 100 is used to receive an audio signal.
[0113] The text conversion module 200 is used to convert the audio signal into text information.
[0114] The action determination module 300 is used to generate answer information corresponding to the text information if the text information does not include an action instruction, and determine a target execution action corresponding to the answer information.
[0115] The execution module 400 is used to execute the target execution action and output the answer information.
[0116] In one embodiment, the above device further comprises:
[0117] The keyword extraction module is used to extract keywords from the text information if the text information includes action instructions.
[0118] The matching module is used to match keywords with the preset action library to determine the target execution action.
[0119] In one embodiment, the action determination module includes:
[0120] The analysis unit is used to analyze the text information and identify the verbs or task-related words in the text information if the text information does not conform to the preset sentence structure.
[0121] The screening unit is used to screen the historical interaction information through a preset window to obtain target interaction information.
[0122] The intention extraction unit is used to obtain the user intention based on the verb or task-related words and the target interaction information; the user intention is used to determine whether the text information includes action instructions.
[0123] In one embodiment, the above device further includes a first training module, and the first training module includes:
[0124] The first sample acquisition unit is used to acquire first sample data; the first sample data carries an action instruction tag.
[0125] The first preprocessing unit is used to preprocess the first sample data to obtain a plurality of segmented words.
[0126] The first training unit is used to input the first sample data into the first initial model for training to obtain a predicted classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each instruction classification head includes an input layer, an encoding layer and a pooling layer; the input layer converts word segmentation into word vectors, and passes the word vectors to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and sends the initial features to the pooling layer; the pooling layer compresses the initial features to obtain target features; the fully connected layer merges the target features corresponding to each classification head to obtain a predicted classification result.
[0127] The first model adjustment unit is used to compare the predicted classification result with the action instruction label to obtain a first difference, and adjust the parameters of the first initial model according to the first difference until the first initial model training is completed to obtain a multi-classification model.
[0128] In one embodiment, the above device further includes a second training module, and the second training module includes:
[0129] The second sample acquisition unit is used to acquire second sample data; the second sample data carries a text label.
[0130] The second preprocessing unit is used to add background noise to the second sample data; the background noise includes one or more of TV playing sound, kitchen stereo sound and conversation sound.
[0131] The second training unit is used to input the second sample data carrying background noise into the second initial model for training to obtain predicted text data.
[0132] The second model adjustment unit is used to compare the second difference between the predicted text data and the text label, and adjust the parameters of the second initial model according to the second difference until the second initial model training is completed to obtain a speech recognition model.
[0133] Each module in the above-mentioned action determination device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.
[0134] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 4As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store preset actions. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an action determination method is implemented.
[0135] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0136] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the following steps when executing the computer program: receiving an audio signal; converting the audio signal into text information; if the text information does not include an action instruction, generating answer information corresponding to the text information, and determining a target execution action corresponding to the answer information from a preset action library; executing the target execution action, and outputting the answer information.
[0137] In one embodiment, when the processor executes the computer program, the following steps are also implemented: if the text information does not conform to the preset sentence structure, the text information is analyzed to identify the verbs or task-related words in the text information; historical interaction information is filtered through a preset window to obtain target interaction information related to the verbs or task-related words.
[0138] In one embodiment, when the processor executes the computer program, the following steps are also implemented: if the text information includes action instructions, matching is performed in a preset action library according to the verb to determine the target execution action.
[0139] In one embodiment, when the processor executes the computer program, it also implements the following steps: generating answer information according to the user's intention; inputting the answer information into a pre-trained multi-classification model to obtain a classification result; matching the classification result with a preset action library to determine the target execution action.
[0140] In one embodiment, when the processor executes the computer program, the following steps are also implemented: obtaining first sample data; the first sample data carries an action instruction label; preprocessing the first sample data to obtain multiple word segmentations; inputting the first sample data into the first initial model for training to obtain a predicted classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each instruction classification head includes an input layer, an encoding layer and a pooling layer; the input layer converts the word segmentation into a word vector, and passes the word vector to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and sends the initial features to the pooling layer; the pooling layer compresses the initial features to obtain the target features; the fully connected layer merges the target features corresponding to each classification head to obtain a predicted classification result; the predicted classification result is compared with the action instruction label to obtain a first difference, and the parameters of the first initial model are adjusted according to the first difference until the first initial model is trained to obtain a multi-classification model.
[0141] In one embodiment, when the processor executes the computer program, the following steps are also implemented: obtaining second sample data; the second sample data carries a text label; adding background noise to the second sample data; the background noise includes one or more of TV playback sound, kitchen stereo sound, and conversation sound; inputting the second sample data carrying the background noise into the second initial model for training to obtain predicted text data; comparing the second difference between the predicted text data and the text label, and adjusting the parameters of the second initial model according to the second difference, until the second initial model is trained to obtain a speech recognition model.
[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: receiving an audio signal; converting the audio signal into text information; if the text information does not include an action instruction, generating answer information corresponding to the text information, and determining a target execution action corresponding to the answer information from a preset action library; executing the target execution action, and outputting the answer information.
[0143] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: if the text information does not conform to the preset sentence structure, the text information is analyzed to identify the verbs or task-related words in the text information; the historical interaction information is filtered through a preset window to obtain target interaction information related to the verbs or task-related words; based on the verbs or task-related words, and the target interaction information, the user intention is obtained; the user intention is used to determine whether the text information includes action instructions.
[0144] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: if the text information includes action instructions, matching is performed in a preset action library according to the verb to determine the target execution action.
[0145] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: generating answer information according to the user's intention; inputting the answer information into a pre-trained multi-classification model to obtain a classification result; matching the classification result with a preset action library to determine the target execution action.
[0146] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: obtaining first sample data; the first sample data carries an action instruction label; preprocessing the first sample data to obtain multiple word segmentations; inputting the first sample data into the first initial model for training to obtain a predicted classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each instruction classification head includes an input layer, an encoding layer and a pooling layer; the input layer converts the word segmentations into word vectors and passes the word vectors to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and sends the initial features to the pooling layer; the pooling layer compresses the initial features to obtain target features; the fully connected layer merges the target features corresponding to each classification head to obtain a predicted classification result; the predicted classification result is compared with the action instruction label to obtain a first difference, and the parameters of the first initial model are adjusted according to the first difference until the first initial model is trained to obtain a multi-classification model.
[0147] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: obtaining second sample data; the second sample data carries a text label; adding background noise to the second sample data; the background noise includes one or more of TV playback sound, kitchen stereo sound, and conversation sound; inputting the second sample data carrying the background noise into the second initial model for training to obtain predicted text data; comparing the second difference between the predicted text data and the text label, and adjusting the parameters of the second initial model according to the second difference, until the second initial model is trained to obtain a speech recognition model.
[0148] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the following steps: receiving an audio signal; converting the audio signal into text information; if the text information does not include an action instruction, generating answer information corresponding to the text information, and determining a target execution action corresponding to the answer information from a preset action library; executing the target execution action, and outputting the answer information.
[0149] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: if the text information does not conform to the preset sentence structure, the text information is analyzed to identify the verbs or task-related words in the text information; the historical interaction information is filtered through a preset window to obtain target interaction information related to the verbs or task-related words; based on the verbs or task-related words, and the target interaction information, the user intention is obtained; the user intention is used to determine whether the text information includes action instructions.
[0150] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: if the text information includes action instructions, matching is performed in a preset action library according to the verb to determine the target execution action.
[0151] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: generating answer information according to the user's intention; inputting the answer information into a pre-trained multi-classification model to obtain a classification result; matching the classification result with a preset action library to determine the target execution action.
[0152] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: obtaining first sample data; the first sample data carries an action instruction label; preprocessing the first sample data to obtain multiple word segmentations; inputting the first sample data into the first initial model for training to obtain a predicted classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each instruction classification head includes an input layer, an encoding layer and a pooling layer; the input layer converts the word segmentations into word vectors and passes the word vectors to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and sends the initial features to the pooling layer; the pooling layer compresses the initial features to obtain target features; the fully connected layer merges the target features corresponding to each classification head to obtain a predicted classification result; the predicted classification result is compared with the action instruction label to obtain a first difference, and the parameters of the first initial model are adjusted according to the first difference until the first initial model is trained to obtain a multi-classification model.
[0153] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: obtaining second sample data; the second sample data carries a text label; adding background noise to the second sample data; the background noise includes one or more of TV playback sound, kitchen stereo sound, and conversation sound; inputting the second sample data carrying the background noise into the second initial model for training to obtain predicted text data; comparing the second difference between the predicted text data and the text label, and adjusting the parameters of the second initial model according to the second difference, until the second initial model is trained to obtain a speech recognition model.
[0154] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0155] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0156] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. An action determination method, characterized in that: The method comprises: receiving an audio signal; Converting the audio signal into text information; Determine whether the text information includes an action instruction by using a judgment model, and if the text information does not include an action instruction, generate answer information corresponding to the text information, and determine a target execution action corresponding to the answer information from a preset action library; Execute the target execution action and output the answer information; The process of determining whether the text information includes the action instruction includes: If the text information does not conform to a preset sentence structure, analyzing the text information to identify verbs or task-related words in the text information; Filtering historical interaction information through a preset window to obtain target interaction information related to the verb or the task-related word; Obtaining a user intention according to the verb or the task-related word and the target interaction information; the user intention is used to determine whether the text information includes the action instruction; The process of judging whether the text information conforms to the preset sentence structure includes: the judgment model converts the text into a vector representation through an input layer, and sends the vector to the encoding layer of the judgment model; in the encoding layer, the judgment model uses a self-attention mechanism to match the vector with a preset sentence vector to judge whether the text is the preset sentence structure; the preset sentence structure is converted into a fixed vector embedding and stored in the parameters of the judgment model.
2. The method according to claim 1, characterized in that The method further comprises: If the text information includes an action instruction, the target execution action is determined by matching the verb in the preset action library.
3. The method according to claim 1, characterized in that The generating of answer information corresponding to the text information, and determining a target execution action corresponding to the answer information from a preset action library, includes: generating the answer information according to the user intention; Inputting the answer information into a pre-trained multi-classification model to obtain a classification result; The target execution action is determined by matching the classification result with the preset action library.
4. The method according to claim 3, characterized in that The training process of the multi-classification model includes: Acquire first sample data; the first sample data carries an action instruction tag; Preprocessing the first sample data to obtain multiple word segments; The first sample data is input into the first initial model for training to obtain a prediction classification result; the first initial model includes multiple instruction classification heads and a fully connected layer, each of the instruction classification heads includes an input layer, an encoding layer and a pooling layer; the input layer converts the word segmentation into a word vector, and transmits the word vector to the encoding layer; the encoding layer obtains the initial features of each word vector through a self-attention mechanism, and transmits the initial features to the pooling layer; the pooling layer compresses the initial features to obtain the target features; the fully connected layer merges the target features corresponding to each of the classification heads to obtain the prediction classification result; The predicted classification result is compared with the action instruction label to obtain a first difference, and the parameters of the first initial model are adjusted according to the first difference until the first initial model is trained to obtain the multi-classification model.
5. The method according to claim 1, characterized in that The conversion of the audio signal into text information is obtained by a pre-trained speech recognition model; The training process of the speech recognition model includes: Acquire second sample data; the second sample data carries a text label; adding background noise to the second sample data; the background noise includes one or more of a TV playing sound, a kitchen sound, and a conversation sound; Inputting the second sample data carrying the background noise into a second initial model for training to obtain predicted text data; A second difference between the predicted text data and the text label is compared, and the parameters of the second initial model are adjusted according to the second difference until the second initial model is trained to obtain the speech recognition model.
6. An action determination device, characterized in that: The device comprises: A receiving module, used for receiving an audio signal; A text conversion module, used for converting the audio signal into text information; an action determination module, used to determine whether the text information includes an action instruction through a judgment model, and if the text information does not include an action instruction, generate answer information corresponding to the text information, and determine a target execution action corresponding to the answer information; An execution module, used for executing the target execution action and outputting the answer information; The action determination module comprises: An analyzing unit, configured to analyze the text information and identify verbs or task-related words in the text information if the text information does not conform to a preset sentence structure; A screening unit, used to screen the historical interaction information through a preset window to obtain target interaction information related to the verb or the task-related word; An intention extraction unit is used to obtain user intention based on the verb or the task-related word, and the target interaction information; the user intention is used to determine whether the text information includes the action instruction; the judgment process of whether the text information conforms to the preset sentence structure includes: the judgment model converts the text into a vector representation through an input layer, and sends the vector to the encoding layer of the judgment model; at the encoding layer, the judgment model uses a self-attention mechanism to match the vector with a preset sentence vector to determine whether the text is the preset sentence structure; the preset sentence structure is converted into a fixed vector embedding and stored in the parameters of the judgment model.
7. The device according to claim 6, characterized in that The device also includes: The keyword extraction module is used to match the verb in the preset action library to determine the target execution action if the text information includes an action instruction.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Voice passthrough method, apparatus and robot
CN107450367A
Homework checking method and system
CN109346108A
improved text classification method based on TextCNN
CN109918507A
Text classification method and device and model training method
CN111475642A
Robot control method, device and equipment and storage medium
CN114227698A