Motion expression generation method and device and storage medium
By segmenting the reply statements and generating a sequence of target action expressions, the problem of the robot's action expressions not matching the reply content is solved, which improves the naturalness and accuracy of robot interactions, reduces the difficulty of processing and improves efficiency.
Patent Information
- Application Number
- CN202510293704.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-25
Smart Images

Figure CN120374804A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a method, device, and storage medium for generating action expressions. Background Art
[0002] Robots can provide services for users in many scenarios, such as being widely applied in industries such as transportation, medical care, education and training, financial banking, travel, digital film and television, etc. In different scenarios, robots can provide corresponding reply content according to users' questions and interact with users more vividly and naturally by expressing appropriate action expressions, where the action expressions are composed of facial expressions and / or body movements. At present, there is a problem that the action expressions of robots do not match the reply content. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, and storage medium for generating action expressions to solve the problem that the action expressions of robots in the prior art do not match the reply content.
[0004] In a first aspect, a method for generating action expressions provided in the embodiments of the present application includes:
[0005] Obtain a reply statement;
[0006] Use a first large model to segment the reply statement into multiple segmented statements, and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively; wherein, the first large model is trained using sample data and training labels; the training labels include multiple segmented statements obtained by segmenting sample reply statements in the sample data, the arrangement order of the multiple segmented statements, and the corresponding action expression labels respectively;
[0007] According to the action expression labels and playing durations respectively corresponding to the multiple segmented statements, look up the action expression labels and playing durations respectively corresponding to different action expressions pre-configured, and determine the action expressions respectively matched by the multiple segmented statements;
[0008] According to the arrangement order of the multiple segmented statements, splice the action expressions respectively corresponding to the multiple segmented statements to generate a target action expression sequence corresponding to the reply statement.
[0009] In a second aspect, an action expression generation device provided in the embodiments of the present application includes:
[0010] An obtaining module, configured to obtain a reply statement;
[0011] A segmentation module, configured to segment a response statement into multiple segmented statements by using a first large model, and determine the arrangement order of the multiple segmented statements and the corresponding action expression tags respectively; wherein, the first large model is obtained by training with sample data and training labels; the training labels include the multiple segmented statements obtained by segmenting a sample response statement in the sample data, the arrangement order of the multiple segmented statements, and the corresponding action expression tags respectively;
[0012] A lookup module, configured to look up the action expression tags and playing durations respectively corresponding to different action expressions pre-configured according to the action expression tags and playing durations respectively corresponding to the multiple segmented statements, and determine the action expressions respectively matched by the multiple segmented statements;
[0013] An assembly module, configured to assemble the action expressions respectively corresponding to the multiple segmented statements according to the arrangement order of the multiple segmented statements, and generate a target action expression sequence corresponding to the response statement.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory is used to store one or more computer instructions, and wherein, when the one or more computer instructions are executed by the processor, the action expression generation method in the above first aspect is implemented. The electronic device may further include a communication interface for communicating with other devices or communication systems.
[0015] In a fourth aspect, an embodiment of the present application provides a non-transitory machine-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the action expression generation method in the above first aspect.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program or instruction, and when the computer program or instruction is executed by a processor, the processor is caused to implement the action expression generation method in the above first aspect.
[0017] In the embodiments of the present application, the obtained response statement is segmented into multiple segmented statements by using a first large model, and the arrangement order of the multiple segmented statements and the corresponding action expression tags are determined respectively. According to the action expression tags and playing durations corresponding to the multiple segmented statements respectively, the action expression tags and playing durations corresponding to different action expressions configured in advance are searched, and the action expressions respectively matched by the multiple segmented statements are determined; according to the arrangement order of the multiple segmented statements, the action expressions corresponding to the multiple segmented statements are spliced to generate a target action expression sequence corresponding to the response statement. Among them, since the segmented statement fits its corresponding action expression, and the target action expression sequence is generated by splicing the action expressions corresponding to the multiple segmented statements obtained by segmenting the response statement, it can be ensured that the action expressions executed by the robot according to the target action expression sequence fit the response statement.
[0018] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. Brief Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 The flowchart of an embodiment of an action expression generation method provided by the present application is shown;
[0021] Figure 2 The structural schematic diagram of an embodiment of an action expression generation device provided by the present application is shown;
[0022] Figure 3 The structural schematic diagram of a computing device provided by the present application is shown. Detailed Description of the Embodiments
[0023] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.
[0024] In some of the processes described in the specification, claims, and the above-mentioned drawings of this application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear herein or in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations can be executed sequentially or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.
[0025] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0026] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope protected by this application.
[0027] The technical solutions of the embodiments of this application can be applied to scenarios where a robot interacts with a user.
[0028] Currently, there is a problem that the action expressions of the robot do not match the reply content, which makes it difficult for users to understand the intention of the robot and affects the interaction with users. To solve the problem that the action expressions of the robot do not match the reply content, the inventor has proposed the technical solution of this application through a series of studies. In the embodiment of this application, the obtained reply statement is segmented into multiple segmented statements by using a first large model, and the arrangement order of the multiple segmented statements and the corresponding action expression tags are determined respectively. According to the action expression tags and playing durations corresponding to the multiple segmented statements respectively, the action expression tags and playing durations corresponding to different action expressions configured in advance are searched, and the action expressions respectively matched by the multiple segmented statements are determined; according to the arrangement order of the multiple segmented statements, the action expressions corresponding to the multiple segmented statements are spliced to generate a target action expression sequence corresponding to the reply statement. Among them, because the segmented statement matches its corresponding action expression, and the target action expression sequence is generated by splicing the action expressions corresponding to the multiple segmented statements obtained by segmenting the reply statement, it can be ensured that the action expression executed by the robot according to the target action expression sequence matches the reply statement. Secondly, the first large model conducts a detailed analysis of each segmented statement to determine the action expression tag corresponding to the segmented statement, which can avoid misjudgment that may be caused by complex content when analyzing the entire reply statement, and improve the accuracy of statement analysis. In addition, by segmenting the reply statement, the first large model can process the segmented statements one by one, reducing the processing difficulty of the first large model, improving the processing speed, and further improving the overall processing efficiency.
[0029] Figure 1 FIG. is a flowchart of an embodiment of an action expression generation method provided by an embodiment of this application. The method may include the following steps:
[0030] 101: Obtain a reply statement.
[0031] The technical solution of this embodiment may be executed by a server, and the server may control a robot terminal, and the robot terminal may include a physical robot terminal and a virtual robot terminal.
[0032] The response statement can be generated based on the question statement put forward by the user to the robot terminal for answering the question statement. Optionally, the robot terminal can obtain the current question voice of the user through the configured microphone, convert the question voice into a text-based question statement and send it to the server, or the robot terminal can send the question voice to the server, and the server converts the question voice into a question statement. Optionally, the robot terminal can also obtain the text-based question statement currently input by the user through the configured touch screen, etc., and send the question statement to the server. Furthermore, optionally, the server can generate a response statement based on the question statement by using a question-answering model. The question-answering model can be a machine learning model for understanding and answering natural language questions. For example, the question-answering model can be obtained by training using sample question-answering data and using the response statements in the sample question-answering data as training labels.
[0033] 102: Use the first large model to split the response statement into multiple segmented statements, and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
[0034] Among them, the first large model is obtained by training using sample data and training labels; the training labels can include multiple segmented statements obtained by splitting the sample response statement in the sample data, the arrangement order of the multiple segmented statements, and the corresponding action expression labels respectively.
[0035] Among them, the sample data can include sample response statements. The sample response statements can be obtained from an open-source dialogue library or social media, etc., and then the training labels of the sample response statements are manually labeled. Or the response statements in the human dialogue in the video can also be extracted as sample response statements, and the action expression labels are set according to the actions and expressions of the people in the video when expressing the sample response statements.
[0036] The first large model can refer to a machine learning model with a large number of parameters and a complex structure, which can process massive data and complete various complex tasks, and is an AI (Artificial Intelligence) model. The first large model can be implemented by using a large language model (English: Large Language Model, abbreviated as: LLM) or a multimodal large model (English: Multimodal Large Model, abbreviated as: MLM), etc. This application does not limit this. Among them, the first large model and the above-mentioned question-answering model can be the same model or different models.
[0037] Among them, the arrangement order of multiple segmented statements in the training labels can be realized by using the text lengths respectively corresponding to the multiple segmented statements and the starting positions in the sample response statement. Furthermore, the first large model can be used to generate the text lengths of each segmented statement and the starting positions in the response statement to determine the arrangement order of the multiple segmented statements.
[0038] 103: According to the action expression labels and playing durations respectively corresponding to multiple segmented statements, search for the action expression labels and playing durations respectively corresponding to different action expressions pre-configured, and determine the action expressions respectively matched by the multiple segmented statements.
[0039] Among them, an action expression can be a sequence of image frames, and the length of the sequence of image frames can determine the playing duration of the action expression. For example, the longer the sequence of image frames, the longer the playing duration.
[0040] Among them, the action expression label can identify one or more action expressions. For example, a certain action expression label is "happy", and this action expression label can identify action expression A with a playing duration of 3s, and can also identify action expression B with a playing duration of 5s. For example, if the action expression label corresponding to a certain segmented statement determined by the large model is "happy", and the playing duration corresponding to this segmented statement is 5s, then this segmented statement can match action expression B.
[0041] Among them, the playing duration corresponding to the segmented statement can be determined in various ways. As an optional way, the playing duration corresponding to the segmented statement can be determined according to the text length of the segmented statement. For example, a predetermined duration corresponding to one character can be determined, and multiplying the number of characters in the segmented statement by the predetermined duration can obtain the playing duration corresponding to the segmented statement. As another optional way, the segmented statement can be converted into the corresponding audio, and the playing duration of the audio is the playing duration corresponding to the segmented statement. For example, the TTS (Text-to-Speech) technology can be used to convert the segmented statement into the corresponding audio.
[0042] 104: According to the arrangement order of multiple segmented statements, splice the action expressions respectively corresponding to the multiple segmented statements to generate the target action expression sequence corresponding to the response statement.
[0043] In this embodiment, the obtained response statement is segmented into multiple segmented statements by using the first large model, and the arrangement order of the multiple segmented statements and the corresponding action expression labels are determined respectively. According to the action expression labels and the playing durations respectively corresponding to the multiple segmented statements, the action expression labels and the playing durations respectively corresponding to different action expressions pre-configured are searched, and the action expressions respectively matched by the multiple segmented statements are determined; according to the arrangement order of the multiple segmented statements, the action expressions respectively corresponding to the multiple segmented statements are spliced to generate a target action expression sequence corresponding to the response statement. Among them, since the segmented statement fits its corresponding action expression, and the target action expression sequence is generated by splicing the action expressions respectively corresponding to the multiple segmented statements obtained by segmenting the response statement, it can be ensured that the action expression executed by the robot according to the target action expression sequence fits the response statement. Secondly, the first large model analyzes each segmented statement in detail to determine the action expression label corresponding to the segmented statement, which can avoid misjudgment caused by complex content when analyzing the entire response statement, and improves the accuracy. In addition, by segmenting the response statement, the first large model can process the segmented statements one by one, reducing the processing difficulty of the first large model and improving the processing speed, thereby improving the overall processing efficiency.
[0044] In some embodiments, the method may further include: sending the target action expression sequence to the robot terminal to instruct the robot terminal to execute the corresponding action expression according to the target action expression sequence while outputting the response statement. So that the interaction between the robot terminal and the user is more vivid and natural, and the user experience is improved.
[0045] In some embodiments, the sample data may include sample dialogue data, and the sample response statement may refer to the response statement in the sample dialogue data.
[0046] Obtaining the response statement may include: obtaining the response statement and the question statement for which the response statement is directed.
[0047] Using the first large model to segment the response statement into multiple segmented statements and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively may include: based on the question statement and the response statement, using the first large model to segment the response statement into multiple segmented statements and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
[0048] The first large model can be trained by using the sample dialogue data and the training labels. The training labels may include multiple segmented statements obtained by segmenting the response statements in the sample dialogue data, the arrangement order of the multiple segmented statements, and the corresponding action expression labels respectively.
[0049] In this embodiment, the first large model can be trained using sample dialogue data and training labels, which can improve the accuracy of the first large model in dialogue understanding. Based on the question statement and the response statement, the first large model can be used to segment the response statement, which can improve the accuracy of identifying the action expression labels corresponding to the segmented statements, and further improve the degree of fit between the target action expression sequence and the response statement.
[0050] In some embodiments, the sample data may include sample multi-turn dialogue data, and the sample response statement may refer to the response statement in the sample multi-turn dialogue data. Obtaining the response statement may include: obtaining the response statement, the question statement to which the response statement is directed, and the historical dialogue.
[0051] Using the first large model to segment the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively may include: based on the question statement, the response statement, and the historical dialogue, using the first large model to segment the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
[0052] The first large model can be trained using sample multi-turn dialogue data and training labels. The training labels may include multiple segmented statements obtained by segmenting the response statement in the sample multi-turn dialogue data, the arrangement order of the multiple segmented statements, and the corresponding action expression labels respectively. The response statement in the sample multi-turn dialogue data may be one or more of all the response statements in the sample multi-turn dialogue, or the response statement of the last turn of the sample multi-turn dialogue data.
[0053] Optionally, the statement identifiers corresponding to the question statement, the response statement, and the historical dialogue can be set in advance, and the statement identifiers corresponding to the question statement, the response statement, and the historical dialogue are input into the first large model so that the first large model can obtain the question statement, the response statement, and the historical dialogue.
[0054] In this embodiment, the first large model is trained using sample multi-turn dialogue data and training labels, which can further improve the accuracy of the first large model in dialogue understanding. Based on the response statement, the question statement, and the historical dialogue, using the first large model to segment the response statement can improve the accuracy of identifying the action expression labels corresponding to the segmented statements, and further improve the degree of fit between the target action expression sequence and the response statement.
[0055] In some embodiments, using the first large model to segment the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively may include: repeatedly performing using the first large model to segment the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively to obtain multiple segmentation results.
[0056] Based on the action expression tags and playing durations respectively corresponding to multiple segmented statements, search for the action expression tags and playing durations respectively corresponding to different action expressions pre-configured, and determining the action expressions respectively matched by multiple segmented statements may include: Based on the action expression tags and playing durations respectively corresponding to multiple segmented statements in any one segmentation result, search for the action expression tags and playing durations respectively corresponding to different action expressions pre-configured, and determine the action expressions respectively matched by multiple segmented statements in any one segmentation result.
[0057] According to the arrangement order of multiple segmented statements, splice the action expressions respectively corresponding to multiple segmented statements to generate a target action expression sequence corresponding to the reply statement may include: According to the arrangement order of multiple segmented statements in any one segmentation result, splice the action expressions respectively corresponding to multiple segmented statements in any one segmentation result to generate multiple candidate action expression sequences corresponding to multiple segmentation results; Calculate the evaluation scores respectively corresponding to multiple candidate action expression sequences, and use the candidate action expression sequence with the highest evaluation score as the target action expression sequence corresponding to the reply statement.
[0058] In this embodiment, by repeatedly executing the segmentation of the reply statement to obtain multiple segmentation results, and then obtaining multiple candidate action expression sequences corresponding to multiple segmentation results, and selecting the optimal candidate action expression sequence as the target action expression sequence, the accuracy and reliability of the target action expression sequence can be improved.
[0059] In some embodiments, the training labels for training the first large model may further include the scenario identifiers corresponding to the sample reply statements.
[0060] Repeatedly executing using the first large model to segment the reply statement into multiple segmented statements, and determining the arrangement order and the action expression tags respectively corresponding to multiple segmented statements to obtain multiple segmentation results may include: Repeatedly executing using the first large model to segment the reply statement into multiple segmented statements, and determining the arrangement order and the action expression tags respectively corresponding to multiple segmented statements to obtain multiple segmentation results, and generating the scenario identifiers respectively corresponding to multiple segmentation results.
[0061] Calculating the evaluation scores respectively corresponding to multiple candidate action expression sequences, and using the candidate action expression sequence with the highest evaluation score as the target action expression sequence corresponding to the reply statement may include: Generating the scenario compatibility scores respectively corresponding to multiple candidate action expression sequences according to the scenario identifier corresponding to any one segmentation result; Calculating the evaluation scores respectively corresponding to multiple candidate action expression sequences according to the scenario compatibility scores, and using the candidate action expression sequence with the highest evaluation score as the target action expression sequence corresponding to the reply statement.
[0062] Among them, generating the scene compatibility scores corresponding to multiple candidate action-expression sequences according to the scene identifier corresponding to any segmentation result means generating the scene compatibility score of the candidate action-expression sequence corresponding to any segmentation result according to the scene identifier corresponding to any segmentation result, and obtaining the scene compatibility scores corresponding to multiple candidate action-expression sequences respectively.
[0063] Among them, the scene identifier may refer to the scene identifier of the current scene recognized by the first large model according to the reply statement. The scene compatibility score can reflect the matching degree between the candidate action-expression sequence and the current scene. The higher the scene compatibility score, the higher the matching degree between the candidate action-expression sequence and the current scene. Calculating the evaluation score according to the scene compatibility score and taking the candidate action-expression sequence with the highest evaluation score as the target action-expression sequence can ensure that the target action-expression sequence matches the current scene, is appropriately integrated with the atmosphere of the current scene, ensure the naturalness of the robot terminal interacting with the user according to the target action-expression sequence, and improve the user experience.
[0064] In some embodiments, generating the scene compatibility scores corresponding to multiple candidate action-expression sequences according to the scene identifier corresponding to any segmentation result may include: using the second large model to generate the scene compatibility scores corresponding to multiple candidate action-expression sequences according to the scene identifier corresponding to any segmentation result.
[0065] Among them, the second large model can be trained in the following way: training the second large model according to the sample action-expression sequence and the training label of the sample action-expression sequence. Among them, the sample action-expression sequence is composed of multiple sample action-expressions with action-expression labels; the sample action-expression sequence has a corresponding scene identifier; the training label of the sample action-expression sequence is the scene compatibility score corresponding to the sample action-expression sequence.
[0066] In some embodiments, calculating the evaluation scores corresponding to multiple candidate action-expression sequences according to the scene compatibility scores may include: determining the richness scores and smoothness scores of multiple candidate action-expression sequences; calculating the evaluation scores corresponding to multiple candidate action-expression sequences according to the scene compatibility scores, richness scores, and smoothness scores corresponding to multiple candidate action-expression sequences respectively.
[0067] Among them, the weights corresponding to the scene compatibility score, richness score, and smoothness score can be determined based on the positive correlation coefficients corresponding to the scene compatibility score, richness score, and smoothness score respectively with the evaluation score, and the evaluation score is obtained by weighted summation of the scene compatibility score, richness score, and smoothness score. The positive correlation coefficient can be pre-configured or obtained through historical statistical learning. Of course, the specific implementation method for obtaining the evaluation score based on the scene compatibility score, richness score, and smoothness score is not limited to this.
[0068] Among them, the richness score of the candidate action-expression sequence can be determined in a variety of optional ways.
[0069] As an optional way, the richness score of the candidate action-expression sequence can be determined according to the number of types of action-expressions in the candidate action-expression sequence.
[0070] The larger the number of types of action-expressions in the candidate action-expression sequence, the higher the richness of the action-expressions in the candidate action-expression sequence, and the higher the richness score. The specific implementation method for determining the richness score according to the number of types of action-expressions in the candidate action-expression sequence in this application is not limited.
[0071] As another optional way, the richness score of the candidate action-expression sequence can be determined according to the interval between adjacent identical action-expressions in the candidate action-expression sequence.
[0072] The larger the interval between each action-expression in the candidate action-expression sequence, the higher the richness of the action-expressions in the candidate action-expression sequence, and the higher the richness score.
[0073] For example, if the candidate action-expression sequence contains three A action-expressions, the intervals between adjacent identical action-expressions can include: the interval between the first A action-expression and the second A action-expression and the interval between the second A action-expression and the third A action-expression.
[0074] Optionally, determining the richness score of the candidate action expression sequence based on the interval between adjacent identical action expressions in the candidate action expression sequence can be as follows: determining the total number of intervals of each action expression in the candidate action expression sequence based on the interval between adjacent identical action expressions in the candidate action expression sequence, and determining the richness score of the candidate action expression sequence based on the total number of intervals of each action expression. For example, in the above example, the interval between the first A action expression and the second A action expression is 5 action expressions, and the interval between the second A action expression and the third A action expression is 4 action expressions, then the total number of intervals of the A action expression is 9. For example, the total number of action expressions included in the candidate action expression sequence can be used as the total number of intervals of the action expressions that only appear once in the candidate action expression sequence, and the value obtained by adding up the total number of intervals of each action expression in the candidate action expression sequence can be used as the richness score of the candidate action expression sequence, or the value obtained by adding up the total number of intervals of each action expression in the candidate action expression sequence is multiplied by a preset positive correlation coefficient to obtain the richness score. The present application does not limit the specific implementation manner of determining the richness score based on the interval between adjacent identical action expressions in the candidate action expression sequence.
[0075] As another alternative, the richness score of the candidate action expression sequence can be determined based on the number of types of action expressions in the candidate action expression sequence and the interval between adjacent identical action expressions.
[0076] For example, the first richness score can be calculated based on the number of types of action expressions in the candidate action expression sequence, the second richness score can be calculated based on the interval between adjacent identical action expressions, and weights are set for the first richness score and the second richness score respectively to calculate the richness score. The present application does not limit this.
[0077] Among them, the smoothness score of the candidate action expression sequence can be determined according to the transition smoothness between adjacent action expressions in the candidate action expression sequence. The higher the transition smoothness, the higher the smoothness score.
[0078] For example, the distance between the last n frames of the previous action expression and the first n frames of the next action expression among adjacent action expressions can be calculated. The value of n can be pre-configured or determined through historical statistical learning, etc. Furthermore, the transition smoothness score of the adjacent action expressions can be determined according to the corresponding relationship between the preset distance and the transition smoothness score. Furthermore, based on the transition smoothness scores of each pair of adjacent action expressions, the richness score of the candidate action expression sequence can be obtained by means of summation or calculating the average value, etc.
[0079] The smaller the distance between the last n frames of the previous action expression and the first n frames of the next action expression in adjacent action expressions, the smaller the difference degree between the last n frames of the previous action expression and the first n frames of the next action expression, the higher the transition smoothness of the adjacent action expressions, and the higher the transition smoothness score. Therefore, a smaller distance can be preset to correspond to a higher transition smoothness score. Among them, distance calculation formulas such as Euclidean distance, Manhattan distance, or Chebyshev distance can be used to calculate the distance between the last n frames of the previous action expression and the first n frames of the next action expression. In addition, the transition smoothness of adjacent action expressions can be judged by other means in addition to judging according to the distance between adjacent action expressions, and this application does not limit this.
[0080] In this embodiment, an evaluation score is obtained according to the scene compatibility score, richness score, and smoothness score, and the candidate action expression sequence with the highest evaluation score is used as the target action expression sequence, which can ensure that the target action expression sequence can be appropriately integrated with the atmosphere of the current scene, and is rich and colorful, smooth and natural.
[0081] Figure 2 It is a schematic structural diagram of an embodiment of an action expression generation device provided by an embodiment of the present application. The device may include:
[0082] An acquisition module, configured to acquire a reply statement;
[0083] A segmentation module, configured to segment the reply statement into multiple segmented statements by using a first large model, and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively; wherein, the first large model is trained by using sample data and training labels; the training labels include multiple segmented statements obtained by segmenting the sample reply statements in the sample data, the arrangement order of the multiple segmented statements, and the corresponding action expression labels respectively;
[0084] A search module, configured to search for the action expression labels and playing durations respectively corresponding to different action expressions pre-configured according to the action expression labels and playing durations respectively corresponding to the multiple segmented statements, and determine the action expressions respectively matched by the multiple segmented statements;
[0085] A splicing module, configured to splice the action expressions respectively corresponding to the multiple segmented statements according to the arrangement order of the multiple segmented statements, and generate a target action expression sequence corresponding to the reply statement.
[0086] In some embodiments, the sample data may include sample dialogue data, and the sample reply statement may refer to the reply statement in the sample dialogue data.
[0087] The acquisition module acquiring the reply statement may include: acquiring the reply statement, and the question statement targeted by the reply statement.
[0088] The splitting module uses the first large model to split the response statement into multiple segmented statements, and determines the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively, which may include: based on the question statement and the response statement, using the first large model to split the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
[0089] In some embodiments, the sample data may include sample multi-turn conversation data, and the sample response statement may refer to the response statement in the sample multi-turn conversation data.
[0090] The obtaining module obtaining the response statement may include: obtaining the response statement, the question statement that the response statement is directed to, and the historical conversation.
[0091] The splitting module uses the first large model to split the response statement into multiple segmented statements, and determines the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively, which may include: based on the question statement, the response statement, and the historical conversation, using the first large model to split the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
[0092] In some embodiments, the splitting module using the first large model to split the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively includes:
[0093] Executing multiple times the operation of using the first large model to split the response statement into multiple segmented statements, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively, to obtain multiple segmentation results;
[0094] The searching module searches for the action expression labels and playing durations corresponding to different action expressions pre-configured according to the action expression labels and playing durations corresponding to the multiple segmented statements respectively, and determines the action expressions respectively matched by the multiple segmented statements, which may include: according to the action expression labels and playing durations corresponding to the multiple segmented statements in any one segmentation result, searching for the action expression labels and playing durations corresponding to different action expressions pre-configured, and determining the action expressions respectively matched by the multiple segmented statements in any one segmentation result.
[0095] The splicing module splices the action expressions corresponding to multiple segmented statements according to the arrangement order of the multiple segmented statements to generate a target action expression sequence corresponding to the reply statement, which may include: splicing the action expressions corresponding to the multiple segmented statements in any one segmentation result according to the arrangement order of the multiple segmented statements in any one segmentation result to generate multiple candidate action expression sequences corresponding to the multiple segmentation results; calculating the evaluation scores corresponding to the multiple candidate action expression sequences respectively, and using the candidate action expression sequence with the highest evaluation score as the target action expression sequence corresponding to the reply statement.
[0096] In some embodiments, the training label may further include a scene identifier corresponding to the sample reply statement.
[0097] The splitting module repeatedly executes splitting the reply statement into multiple segmented statements by using a first large model, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively to obtain multiple segmentation results, which may include: repeatedly executing splitting the reply statement into multiple segmented statements by using a first large model, and determining the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively to obtain multiple segmentation results, and generating scene identifiers corresponding to the multiple segmentation results respectively.
[0098] The splicing module calculates the evaluation scores corresponding to the multiple candidate action expression sequences respectively, and using the candidate action expression sequence with the highest evaluation score as the target action expression sequence corresponding to the reply statement includes: generating the scene compatibility scores corresponding to the multiple candidate action expression sequences respectively according to the scene identifier corresponding to any one segmentation result; calculating the evaluation scores corresponding to the multiple candidate action expression sequences respectively according to the scene compatibility scores, and using the candidate action expression sequence with the highest evaluation score as the target action expression sequence corresponding to the reply statement.
[0099] In some embodiments, the splicing module generating the scene compatibility scores corresponding to the multiple candidate action expression sequences respectively according to the scene identifier corresponding to any one segmentation result may include: using a second large model to generate the scene compatibility scores corresponding to the multiple candidate action expression sequences respectively according to the scene identifier corresponding to any one segmentation result.
[0100] Among them, the second large model may be trained as follows: training the second large model according to the sample action expression sequence and the training label of the sample action expression sequence; where the sample action expression sequence is spliced by multiple sample action expressions with action expression labels; the sample action expression sequence has a corresponding scene identifier; the training label of the sample action expression sequence is the scene compatibility score corresponding to the sample action expression sequence.
[0101] In some embodiments, the splicing module calculating the evaluation scores corresponding to the multiple candidate action expression sequences respectively according to the scene compatibility scores includes:
[0102] Determine the richness score and smoothness score of multiple candidate action expression sequences;
[0103] Calculate the evaluation scores corresponding to multiple candidate action expression sequences according to the scene compatibility scores, richness scores, and smoothness scores corresponding to the multiple candidate action expression sequences respectively.
[0104] In some embodiments, the apparatus may further include:
[0105] An indication module, configured to send the target action expression sequence to the robot terminal, so as to instruct the robot terminal to execute the corresponding action expression according to the target action expression sequence while outputting a reply statement.
[0106] Figure 2 The described action expression generation apparatus may execute Figure 1 The action expression generation method described in the illustrated embodiment, and its implementation principle and technical effects will not be elaborated. For the action expression generation apparatus in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0107] In a possible design, Figure 2 The action expression generation apparatus in the illustrated embodiment may be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 301 and a processing component 302;
[0108] The storage component 301 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 302 to implement the action expression generation method described in the illustrated embodiment. Figure 1 As shown.
[0109] Of course, the computing device may necessarily further include other components, such as an input / output interface, a communication component, etc.
[0110] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc.
[0111] The communication component is configured to facilitate wired or wireless communication between the electronic device and other devices, etc.
[0112] Among them, the computing device may be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device may refer to a cloud server, and the above processing component, storage component, etc. may be basic server resources leased or purchased from a cloud computing platform.
[0113] The processing components involved in the corresponding embodiments described above may include one or more processors to execute computer instructions to complete all or part of the steps in the above methods. Of course, the processing components may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above methods.
[0114] The storage component 301 is configured to store various types of data to support operations in the device. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0115] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 action expression generation method shown in the embodiments.
[0116] In addition, the embodiments of the present application provide a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor can implement the steps or functions of the action expression generation method shown in the above Figure 1 embodiments.
[0117] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating action expressions, characterized in that, including: Obtain a reply statement; Use a first large model to split the reply statement into multiple segmented statements, and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively; wherein, the first large model is trained using sample data and training labels; the training labels include multiple segmented statements obtained by splitting a sample reply statement in the sample data, the arrangement order of the multiple segmented statements, and the corresponding action expression labels respectively; According to the action expression labels and playing durations respectively corresponding to the multiple segmented statements, look up the action expression labels and playing durations respectively corresponding to different action expressions pre-configured, and determine the action expressions respectively matched by the multiple segmented statements; According to the arrangement order of the multiple segmented statements, splice the action expressions respectively corresponding to the multiple segmented statements to generate a target action expression sequence corresponding to the reply statement.
2. The method according to claim 1, wherein The sample data includes sample dialogue data; The sample reply statement is a reply statement in the sample dialogue data; The obtaining of the reply statement includes: Obtain a reply statement and the question statement targeted by the reply statement; The using the first large model to split the reply statement into multiple segmented statements and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively includes: Based on the question statement and the reply statement, use the first large model to split the reply statement into multiple segmented statements, and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
3. The method according to claim 1, characterized in that, The sample data includes sample multi-turn dialogue data; The sample reply statement is a reply statement in the sample multi-turn dialogue data; The obtaining of the reply statement includes: Obtain a reply statement, the question statement targeted by the reply statement, and the historical dialogue; The using the first large model to split the reply statement into multiple segmented statements and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively includes: Based on the question statement, the reply statement, and the historical dialogue, use the first large model to split the reply statement into multiple segmented statements, and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively.
4. The method according to claim 1, wherein The using the first large model to split the reply statement into multiple segmented statements and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively includes: Execute multiple times the using the first large model to split the reply statement into multiple segmented statements and determine the arrangement order of the multiple segmented statements and the corresponding action expression labels respectively to obtain multiple segmentation results; The according to the action expression labels and playing durations respectively corresponding to the multiple segmented statements, look up the action expression labels and playing durations respectively corresponding to different action expressions pre-configured, and determine the action expressions respectively matched by the multiple segmented statements includes: Based on the action-expression tags and playing durations respectively corresponding to multiple segmented statements in any one of the segmentation results, search for the action-expression tags and playing durations respectively corresponding to different action-expressions pre-configured, and determine the action-expressions respectively matched by the multiple segmented statements in any one of the segmentation results; The generating the target action-expression sequence corresponding to the reply statement by splicing the action-expressions respectively corresponding to the multiple segmented statements according to the arrangement order of the multiple segmented statements includes: According to the arrangement order of the multiple segmented statements in any one of the segmentation results, splice the action-expressions respectively corresponding to the multiple segmented statements in any one of the segmentation results to generate multiple candidate action-expression sequences corresponding to the multiple segmentation results; Calculate the evaluation scores respectively corresponding to the multiple candidate action-expression sequences, and use the candidate action-expression sequence with the highest evaluation score as the target action-expression sequence corresponding to the reply statement.
5. The method according to claim 4, wherein The training label further includes the scene identifier corresponding to the sample reply statement; The multiple executions of using the first large model to segment the reply statement into multiple segmented statements, and determining the arrangement order and the action-expression tags respectively corresponding to the multiple segmented statements to obtain multiple segmentation results include: Multiple executions of using the first large model to segment the reply statement into multiple segmented statements, and determining the arrangement order and the action-expression tags respectively corresponding to the multiple segmented statements to obtain multiple segmentation results, and generating the scene identifiers respectively corresponding to the multiple segmentation results; The calculating the evaluation scores respectively corresponding to the multiple candidate action-expression sequences, and using the candidate action-expression sequence with the highest evaluation score as the target action-expression sequence corresponding to the reply statement includes: Generate the scene compatibility scores respectively corresponding to the multiple candidate action-expression sequences according to the scene identifier corresponding to any one of the segmentation results; Calculate the evaluation scores respectively corresponding to the multiple candidate action-expression sequences according to the scene compatibility scores, and use the candidate action-expression sequence with the highest evaluation score as the target action-expression sequence corresponding to the reply statement.
6. The method according to claim 5, wherein The generating the scene compatibility scores respectively corresponding to the multiple candidate action-expression sequences according to the scene identifier corresponding to any one of the segmentation results includes: Use the second large model to generate the scene compatibility scores respectively corresponding to the multiple candidate action-expression sequences according to the scene identifier corresponding to any one of the segmentation results; The second large model is trained in the following manner: Train the second large model according to the sample action-expression sequence and the training label of the sample action-expression sequence; Wherein, the sample action-expression sequence is spliced by multiple sample action-expressions with action-expression tags; the sample action-expression sequence has a corresponding scene identifier; the training label of the sample action-expression sequence is the scene compatibility score corresponding to the sample action-expression sequence.
7. The method according to claim 5, wherein The calculating the evaluation scores respectively corresponding to the multiple candidate action-expression sequences according to the scene compatibility scores includes: Determine the richness scores and smoothness scores respectively corresponding to the multiple candidate action-expression sequences; Calculate the evaluation scores corresponding to the multiple candidate action expression sequences respectively according to the scene compatibility scores, the richness scores, and the smoothness scores corresponding to the multiple candidate action expression sequences respectively.
8. The method according to claim 1, wherein Further comprising: Sending the target action expression sequence to a robot terminal to instruct the robot terminal to execute corresponding action expressions according to the target action expression sequence while outputting the reply statement.
9. An electronic device, characterized in that, Comprising: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the action expression generation method according to any one of claims 1 to 8.
10. A non-transitory machine-readable storage medium, characterized in that, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor executes the action expression generation method according to any one of claims 1 to 8.