A robot control method based on voice analysis

By acquiring spatial image information of the robot to construct semantic tag groups and guidance corpora, voice control commands are optimized, solving the problem of command recognition under the influence of user language differences and dialects, and improving the accuracy and precision of voice control.

CN121148376BActive Publication Date: 2026-01-13厦门工学院
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511688037.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-01-13
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

In dynamic and complex environments such as homes and offices, differences in user language habits or dialects can lead to low accuracy and poor precision in voice control command recognition.

Method used

By acquiring image information of the spatial area where the target robot is located, marking contour features and identifying semantic labels, constructing a potentially related semantic label group, filtering semantically guided corpus groups, combining the textual keyword relevance of voice control data, determining the tendency of commands to be ambiguous, optimizing ambiguous commands, and generating confident command text.

Benefits of technology

It improves the accuracy of voice control data analysis and control command recognition in dynamic and complex environments, ensuring the accuracy of robot operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148376B_ABST
    Figure CN121148376B_ABST
Patent Text Reader

Abstract

The present application relates to the field of voice analysis, and particularly relates to a robot control method based on voice analysis, the present application determines semantic labels by acquiring image information in a space region where a target robot is located, and determines a potential associated semantic label group, and screens out a semantic-oriented corpus group for the target robot, when receiving voice control data, determines instruction fuzzy variables of the voice control data to determine instruction fuzzy tendency, optimizes voice control data with instruction fuzzy tendency, including determining confidence clusters and semantic discrete clusters, acquires expanded text after semantic expansion for the semantic discrete clusters, screens the expanded text based on the semantic-oriented corpus group, and obtains confidence instruction text, the present application combines image information to construct a semantic-oriented corpus group, guides analysis of voice control data with instruction fuzzy tendency, improves analysis accuracy of voice control data under semantic ambiguity, and ensures control instruction recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech analysis, and more particularly to a robot control method based on speech analysis. Background Technology

[0002] In the current field of human-computer interaction, natural language processing technology has become the core of achieving intelligence. It enables machines to understand, interpret, and respond to language used by humans in daily life, greatly reducing the barrier to entry. At the same time, voice control, as the most instinctive information input method, has been widely used in various scenarios such as smart homes, industrial automation, and service robots due to its natural advantages of freeing up hands and convenient operation.

[0003] For example, Chinese Patent Publication No. CN110136704A discloses a robot voice control method, device, robot, and medium. The robot voice control method includes: receiving control-type voice commands; matching the control-type voice commands with commands stored in the robot's local database; and if a match is successful, executing the control-type voice command. The embodiment described in this document, upon receiving a voice command, first activates the robot's voice recognition system and controls the robot to turn towards the direction of the voice source, placing it in a standby state. Then, within a certain time, it receives control commands to execute operations, performs semantic matching between the control commands and pre-stored commands from the robot's local database and the cloud, and executes the expected action based on the matching result. The technical solution described in this document can accurately operate according to instructions, improve the recognition rate of user-input voice commands, and work more accurately according to user voice commands, while also increasing the fun of human-computer interaction.

[0004] However, the following problems still exist in the existing technology.

[0005] In dynamic and complex environments such as homes and offices, differences in user language habits or dialects, as well as the presence of some ambiguous content in the commands themselves, lead to low command recognition accuracy and poor voice control accuracy. Summary of the Invention

[0006] To address this, the present invention provides a robot control method based on voice analysis, which overcomes the problems in the prior art where, in dynamic and complex environments such as homes and offices, differences in user language habits or dialects, as well as the presence of some unclear references in the commands, result in low command recognition accuracy and poor voice control accuracy.

[0007] To achieve the above objectives, the present invention provides a robot control method based on voice analysis, comprising:

[0008] Acquire image information of the spatial region where the target robot is located, label several contour features in the image information, identify semantic labels of the contour features, and determine potential associated semantic label groups based on all semantic labels in the image information;

[0009] The semantic tags in the potentially related semantic tag group are selected from the sample corpus to form a semantically guided corpus for the target robot;

[0010] In response to receiving voice control data, the voice control data is textualized, and the fuzzy parameters of the voice control data are determined based on the semantic correlation between each keyword in the text data and the semantically guided corpus, so as to determine the fuzzy tendency of the voice control data.

[0011] Optimization is performed on voice control data that tends to have ambiguous commands, including:

[0012] The semantic clusters corresponding to the voice control data are determined, and the confidence clusters are determined in combination with the semantically guided corpus. The semantic discrete clusters in the text data are determined, and the semantic discrete clusters are extended by synonyms to generate several extended texts. The extended texts are filtered based on the confidence clusters to obtain the confidence instruction text.

[0013] The confidence instruction text is identified as the control instruction for the target robot;

[0014] The confidence set cluster contains several keywords.

[0015] Furthermore, the process of parsing the potential associated semantic tag groups for the target robot based on all semantic tags in the image information includes,

[0016] Identify the category corresponding to the contour features, determine the category text, and use the category text as a semantic label;

[0017] Based on the probability of semantic tag combinations appearing simultaneously in a single sample corpus, potential associated semantic tag combinations are determined.

[0018] If the probability is greater than a predetermined probability threshold, the corresponding semantic tag combination is determined to be a potential associated semantic tag group.

[0019] Furthermore, the process of selecting semantically guided corpora for the target robot from the sample corpus based on semantic tags in the potentially related semantic tag group includes,

[0020] Identify each semantic tag in a potentially associated semantic tag group;

[0021] Sample corpora containing the aforementioned semantic tags are selected from the sample corpus;

[0022] The sample corpus is pre-built and stores several sample corpora, each of which is obtained by textualizing the voice control data of the target robot.

[0023] Furthermore, the process of determining the fuzzy parameters of the voice control data based on the semantic correlation between keywords in the text data and the semantically guided corpus includes:

[0024] Calculate the semantic correlation between keywords in the text data;

[0025] Solve for the variance of semantic relevance to obtain discrete semantic features;

[0026] The semantic correlation between keywords and the average semantic correlation between each sample in the semantically oriented corpus is determined to obtain the semantically oriented correlation.

[0027] The semantic discrete features are calculated as a ratio to a preset semantic discrete threshold to obtain the semantic discrete factors;

[0028] The semantic guidance factor is obtained by calculating the ratio of the preset semantic guidance threshold to the semantic guidance relevance.

[0029] The semantic discrete factors and the semantic guiding factors are weighted and summed to obtain the instruction fuzzy parameter.

[0030] Furthermore, the process of determining the ambiguity tendency of the voice control data includes,

[0031] If the instruction ambiguity parameter is greater than or equal to the preset instruction ambiguity threshold, it is determined that there is an instruction ambiguity tendency.

[0032] If the instruction ambiguity parameter is less than the preset instruction ambiguity threshold, it is determined that there is no instruction ambiguity tendency.

[0033] Furthermore, since the voice control data does not tend to be ambiguous, the voice control data is transcribed into text and used as a confidence command text.

[0034] Furthermore, the process of determining the semantic clusters of the text data corresponding to the voice control data includes,

[0035] Calculate the semantic relevance between keywords in the text data, and cluster the keywords based on the semantic relevance to obtain semantic clusters;

[0036] Among them, the semantic correlation between any keywords in the semantic cluster is greater than the preset semantic cluster correlation threshold.

[0037] Furthermore, the process of determining the confidence cluster includes,

[0038] Determine the mean semantic correlation between each keyword in the semantic cluster and the sample corpus in the semantically guided corpus group;

[0039] If the mean of the semantic relevance is greater than the predetermined confidence relevance threshold, then the semantic cluster is determined as a confidence cluster.

[0040] The remaining portion of the text data is identified as a semantic discrete cluster.

[0041] Furthermore, the process of performing synonym expansion on semantic discrete clusters to generate several expanded texts includes,

[0042] The semantic discrete clusters are split into several keywords;

[0043] The keywords were replaced with candidate words to obtain several expanded texts;

[0044] The candidate words are synonyms of the keywords.

[0045] Furthermore, the process of filtering the extended text based on the confidence set cluster to obtain the confidence instruction text includes,

[0046] The keywords in the confidence set cluster are determined, and sample corpora containing the keywords are selected from the semantically guided corpus.

[0047] Each of the extended texts is compared with the sample corpus in the semantically guided corpus to determine the semantic relevance of the extended text, and the extended text corresponding to the maximum semantic relevance is selected as the confidence instruction text.

[0048] Compared with existing technologies, this invention acquires image information of the spatial region where the target robot is located, determines semantic tags, and identifies potentially associated semantic tag groups to filter out a semantically guided corpus for the target robot. When receiving voice control data, it determines the fuzzy parameters of the voice control data to judge the fuzziness tendency of the voice control data. Subsequently, it optimizes the voice control data with fuzzy tendencies, including determining confidence clusters and semantic discrete clusters. After semantic expansion of the semantic discrete clusters, it obtains expanded text. Then, it filters the expanded text based on the semantically guided corpus to obtain the confidence command text. This invention combines image information to construct a semantically guided corpus, guiding the analysis of voice control data with fuzzy tendencies, improving the accuracy of voice control data analysis under semantic fuzziness, and ensuring the accuracy of control command recognition.

[0049] In particular, this invention acquires image information within the spatial region where the target robot is located, determines potential related semantic tag groups, and constructs a semantically guided corpus. In practice, the control commands for the target robot are often related to the current spatial region. Therefore, this invention considers using visual features to guide the parsing of voice control data. First, semantic tags in the image information are identified to characterize the potential interactive objects of the target robot. By constructing potential related semantic tag groups, multiple targets with physical interaction relationships in the spatial region can be identified, such as a cup and a coaster. The command could be to place the cup on the coaster, indicating a physical interaction relationship between the cup and the coaster. Based on this, a semantically guided corpus is selected from the sample corpus based on the semantic tags in the potential related semantic tag groups. The semantically guided corpus consists entirely of sample corpora with semantic tags, and all semantic tags belong to related semantic tag groups. The purpose is to pre-determine several potential commands that may occur based on several targets in the current spatial region, thereby providing data support for the subsequent optimization of voice control data with a tendency for ambiguous commands.

[0050] In particular, this invention calculates fuzzy command parameters by utilizing semantic discrete factors and semantic guidance factors. Semantic discrete factors characterize the discreteness of semantic associations among different parts of the voice control data, while semantic guidance factors characterize the association between the voice control data and the semantic guidance corpus. The calculated fuzzy command parameters reflect the fuzziness of the voice control data. In practice, if the voice control data contains ambiguous or unclear commands, the semantic associations among different parts will be relatively discrete. Furthermore, the description may deviate from the intended meaning and not match the current spatial region. Therefore, by identifying ambiguous or unclear commands in the voice control data in advance, this invention provides data support for subsequent optimization, thereby improving the accuracy of voice control data analysis under semantic fuzziness and ensuring the accuracy of control command recognition.

[0051] In particular, this invention identifies semantic clusters and confidence clusters. Confidence clusters reflect the relatively concentrated parts of voice control data with a tendency for ambiguous instructions, and these parts are associated with semantically guided corpora. Based on this, the extended text is filtered using the high-confidence parts of the voice control data as guidance, and the confident instruction texts are selected, thereby improving the accuracy of voice control data analysis in cases of semantic ambiguity and ensuring the accuracy of control instruction recognition. Attached Figure Description

[0052] Figure 1 This is a schematic diagram illustrating the steps of a robot control method based on voice analysis, as described in an embodiment of the invention.

[0053] Figure 2 This is a logic block diagram for determining potentially associated semantic tag groups according to an embodiment of the invention;

[0054] Figure 3 A logic block diagram for determining the ambiguity tendency of the voice control data in an embodiment of the invention;

[0055] Figure 4 This is a logic block diagram of the cluster in the confidence set for an embodiment of the invention. Detailed Implementation

[0056] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0057] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0058] Please see Figure 1 The diagram illustrates the steps of a voice analysis-based robot control method according to an embodiment of the present invention. The voice analysis-based robot control method of the present invention includes:

[0059] Step S1: Obtain image information of the spatial region where the target robot is located, label several contour features in the image information, identify semantic labels of the contour features, and determine potential associated semantic label groups based on all semantic labels in the image information;

[0060] Step S2: Based on the semantic tags in the potential association semantic tag group, a semantically guided corpus group for the target robot is selected from the sample corpus;

[0061] Step S3: In response to receiving voice control data, the voice control data is textualized, and the fuzzy parameters of the voice control data are determined based on the semantic correlation between each keyword in the text data and the semantically guided corpus, so as to determine the fuzzy tendency of the voice control data.

[0062] Step S4 involves optimizing voice control data that tends to have ambiguous commands, including:

[0063] The semantic clusters corresponding to the voice control data are determined, and the confidence clusters are determined in combination with the semantically guided corpus. The semantic discrete clusters in the text data are determined, and the semantic discrete clusters are extended by synonyms to generate several extended texts. The extended texts are filtered based on the confidence clusters to obtain the confidence instruction text.

[0064] Step S5: Determine the confidence instruction text as the control instruction for the target robot;

[0065] The confidence set cluster contains several keywords.

[0066] Specifically, there are no restrictions on the method of acquiring image information. For example, it can be acquired through image acquisition devices deployed on the robot, as long as it can acquire image information within the target area.

[0067] Please see Figure 2 As shown, Figure 2 This is a logical block diagram illustrating the determination of potential related semantic tag groups according to an embodiment of the invention. The process of parsing potential related semantic tag groups for a target robot based on all semantic tags in image information includes:

[0068] Identify the category corresponding to the contour features, determine the category text, and use the category text as a semantic label;

[0069] Based on the probability of semantic tag combinations appearing simultaneously in a single sample corpus, potential associated semantic tag combinations are determined.

[0070] If the probability is greater than a predetermined probability threshold, the corresponding semantic tag combination is determined to be a potential associated semantic tag group.

[0071] Specifically, there are no restrictions on the method for identifying the category of contour features. Image segmentation algorithms can be used to identify the contour features of objects. Then, existing open-source image processing models that can identify object types can be used to identify the corresponding categories of objects. Alternatively, an image processing model that can identify the categories of objects in an image can be trained to identify the corresponding categories of objects. As long as the corresponding function can be achieved, it will not be elaborated further.

[0072] Specifically, the probability threshold is preset. Several semantic tag combinations are pre-statistically counted, and the probability of each semantic tag combination appearing simultaneously in a single sample corpus is recorded. The mean probability is calculated. To reflect cases with high correlation, the probability threshold is set as the product of the mean probability and the probability offset coefficient. The probability offset coefficient is selected within the interval [1.25, 1.5], preferably 1.35.

[0073] Specifically, the process of selecting semantically guided corpora for the target robot from the sample corpus based on semantic tags in the potentially related semantic tag group includes:

[0074] Identify each semantic tag in a potentially associated semantic tag group;

[0075] Sample corpora containing the aforementioned semantic tags are selected from the sample corpus;

[0076] The sample corpus is pre-built and stores several sample corpora, each of which is obtained by textualizing the voice control data of the target robot.

[0077] Specifically, a number of voice control data for robots can be collected in advance, and the voice control data with clear semantics can be manually screened, transcribed into text, and then stored in the sample corpus.

[0078] This invention acquires image information within the spatial region where the target robot is located, determines potential related semantic tag groups, and constructs a semantically guided corpus. In practice, control commands for the target robot are often related to the current spatial region. Therefore, this invention considers using visual features to guide the parsing of voice control data. First, semantic tags in the image information are identified to characterize the potential interactive objects of the target robot. By constructing potential related semantic tag groups, multiple targets with physical interaction relationships in the spatial region can be identified, such as a cup and a coaster. The command could be to place the cup on the coaster, indicating a physical interaction between the cup and the coaster. Based on this, a semantically guided corpus is selected from the sample corpus based on the semantic tags in the potential related semantic tag groups. The semantically guided corpus consists entirely of sample corpora with semantic tags, and all semantic tags belong to related semantic tag groups. The purpose is to pre-determine several potential commands based on several targets in the current spatial region, thereby providing data support for the subsequent optimization of voice control data with a tendency for ambiguous commands.

[0079] Specifically, the process of determining the fuzzy parameters of the voice control data based on the semantic correlation between keywords in the text data and a semantically guided corpus includes:

[0080] Calculate the semantic correlation between keywords in the text data;

[0081] Solve for the variance of semantic relevance to obtain discrete semantic features;

[0082] The semantic correlation between keywords and the average semantic correlation between each sample in the semantically oriented corpus is determined to obtain the semantically oriented correlation.

[0083] The semantic discrete features are calculated as a ratio to a preset semantic discrete threshold to obtain the semantic discrete factors;

[0084] The semantic guidance factor is obtained by calculating the ratio of the preset semantic guidance threshold to the semantic guidance relevance.

[0085] The semantic discrete factors and the semantic guiding factors are weighted and summed to obtain the instruction fuzzy parameter.

[0086] Specifically, there are no restrictions on the method for determining semantic relevance. For the semantic relevance between keywords, the keywords can be vectorized and the cosine similarity can be calculated. The cosine similarity can then be used as the semantic relevance. For the semantic relevance between keywords and sentences, the cosine similarity between the keywords and each keyword in the sentence can be calculated separately. The mean of the cosine similarity can then be used as the semantic relevance. Of course, other methods can also be used, which will not be elaborated here.

[0087] Specifically, the semantic discrete threshold and the semantic guidance threshold are predetermined. Several voice control data that cause the robot to execute incorrectly are recorded as samples. The semantic discrete features of the text data corresponding to the voice control data are calculated, the semantic guidance correlation degree corresponding to the voice control data is calculated, the mean of the semantic discrete features and the mean of the semantic guidance correlation degree are solved, and the semantic discrete threshold is set as the mean of the semantic discrete features and the semantic guidance threshold is set as the mean of the semantic guidance correlation degree.

[0088] To comprehensively consider both semantic discrete factors and semantic guiding factors, the weight for semantic discrete factors is 0.5 and the weight for semantic guiding factors is 0.5 when performing weighted summation.

[0089] Please see Figure 3 As shown, Figure 3 This is a logic block diagram illustrating the determination of the ambiguity tendency of the voice control data according to an embodiment of the invention. The process of determining the ambiguity tendency of the voice control data includes...

[0090] If the instruction ambiguity parameter is greater than or equal to the preset instruction ambiguity threshold, it is determined that there is an instruction ambiguity tendency.

[0091] If the instruction ambiguity parameter is less than the preset instruction ambiguity threshold, it is determined that there is no instruction ambiguity tendency.

[0092] In practice, the instruction fuzzy parameter calculated when the semantic discrete features are equal to the preset semantic discrete threshold and when the semantic guidance threshold is equal to the semantic guidance correlation degree is used as the instruction fuzzy threshold.

[0093] Specifically, since the voice control data does not tend to be ambiguous, the voice control data is textualized and used as a confidence command text.

[0094] This invention calculates fuzzy parameters for commands, utilizing semantic discreteness and semantic guidance factors. Semantic discreteness factors characterize the discreteness of semantic associations among different parts of the voice control data, while semantic guidance factors characterize the association between the voice control data and the semantic guidance corpus. The calculated fuzzy parameters reflect the fuzziness of the commands in the voice control data. In practice, if the voice control data contains ambiguous or unclear commands, the semantic associations among different parts will be relatively discrete. Furthermore, the description may deviate from the intended meaning and not match the current spatial region. Therefore, by identifying ambiguous or unclear commands in the voice control data in advance, this invention provides data support for subsequent optimization, thereby improving the accuracy of voice control data analysis under semantic fuzziness and ensuring the precision of control command recognition.

[0095] Specifically, the process of determining the semantic clusters of text data corresponding to voice control data includes,

[0096] Calculate the semantic relevance between keywords in the text data, and cluster the keywords based on the semantic relevance to obtain semantic clusters;

[0097] Among them, the semantic correlation between any keywords in the semantic cluster is greater than the preset semantic cluster correlation threshold.

[0098] Specifically, the semantic clustering association threshold is preset. Several target robot command executions without abnormalities are pre-selected for voice control data. The average semantic correlation between keywords after the voice control data is textualized is calculated. The semantic clustering association threshold is set to be the product of the average semantic correlation and the association offset coefficient. The association offset coefficient is selected in the range [0.75, 0.85], preferably 0.8.

[0099] Please see Figure 4 As shown, Figure 4 This is a logic block diagram illustrating the process of determining a cluster within a confidence set according to an embodiment of the invention. The process of determining a cluster within a confidence set includes:

[0100] Determine the mean semantic correlation between each keyword in the semantic cluster and the sample corpus in the semantically guided corpus group;

[0101] If the mean of the semantic relevance is greater than the predetermined confidence relevance threshold, then the semantic cluster is determined as a confidence cluster.

[0102] The remaining portion of the text data is identified as a semantic discrete cluster.

[0103] Specifically, the confidence correlation threshold is preset. Among them, several voice control data corresponding to the execution of target robot instructions without abnormalities are pre-selected, and the semantic correlation mean of each keyword and the sample corpus in the semantically guided corpus is recorded after the voice control data is textualized. The mean of the semantic correlation mean is calculated, and the confidence correlation threshold is set as the product of the mean and the correlation offset coefficient. The correlation offset coefficient is selected in the interval [0.75, 0.85], preferably 0.8.

[0104] Specifically, the process of performing synonym expansion on semantic discrete clusters to generate several expanded texts includes:

[0105] The semantic discrete clusters are split into several keywords;

[0106] The keywords were replaced with candidate words to obtain several expanded texts;

[0107] The candidate words are synonyms of the keywords.

[0108] Specifically, there are no restrictions on the way synonyms are generated. They can be synonyms directly searched from a dictionary or synonyms determined by pre-establishing a mapping relationship between synonyms.

[0109] In some possible implementations, synonyms can be extended across objects of the same type. For example, since plates belong to the tableware type, other tableware can be set as synonyms for plates to improve the applicability of the extended text.

[0110] Specifically, the process of filtering the extended text based on the confidence set cluster to obtain the confidence instruction text includes, as follows:

[0111] The keywords in the confidence set cluster are determined, and sample corpora containing the keywords are selected from the semantically guided corpus.

[0112] Each of the extended texts is compared with the sample corpus in the semantically guided corpus to determine the semantic relevance of the extended text, and the extended text corresponding to the maximum semantic relevance is selected as the confidence instruction text.

[0113] This invention identifies semantic clusters and confidence clusters. Confidence clusters reflect the relatively concentrated parts of voice control data with a tendency for ambiguous instructions, and these parts are associated with semantically guided corpora. Based on this, the extended text is filtered using the high-confidence parts of the voice control data as guidance, and the confident instruction texts are selected, thereby improving the accuracy of voice control data analysis in cases of semantic ambiguity and ensuring the accuracy of control instruction recognition.

[0114] If the robot control method based on voice analysis of the present invention is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A robot control method based on speech analysis, characterized in that, include: Acquire image information of the spatial region where the target robot is located, label several contour features in the image information, identify semantic labels of the contour features, and determine potential associated semantic label groups based on all semantic labels in the image information; The semantic tags in the potentially related semantic tag group are selected from the sample corpus to form a semantically guided corpus for the target robot; In response to receiving voice control data, the voice control data is textualized, and the fuzzy parameters of the voice control data are determined based on the semantic correlation between each keyword in the text data and the semantically guided corpus, so as to determine the fuzzy tendency of the voice control data. Optimize voice control data that tends to have ambiguous commands. include, The semantic clusters corresponding to the voice control data are determined, and the confidence clusters are determined in combination with the semantically guided corpus. The semantic discrete clusters in the text data are determined, and the semantic discrete clusters are extended by synonyms to generate several extended texts. The extended texts are filtered based on the confidence clusters to obtain the confidence instruction text. The confidence instruction text is identified as the control instruction for the target robot; The confidence set cluster contains several keywords; The process of determining the fuzzy parameters of the voice control data based on the semantic correlation between keywords in the text data and the semantically guided corpus includes: Calculate the semantic correlation between keywords in the text data; Solve for the variance of semantic relevance to obtain discrete semantic features; The semantic correlation between keywords and the average semantic correlation between each sample in the semantically oriented corpus is determined to obtain the semantically oriented correlation. The semantic discrete features are calculated as a ratio to a preset semantic discrete threshold to obtain the semantic discrete factors; The semantic guidance factor is obtained by calculating the ratio of the preset semantic guidance threshold to the semantic guidance relevance. The semantic discrete factors and the semantic guiding factors are weighted and summed to obtain the instruction fuzzy parameter.

2. The robot control method based on speech analysis according to claim 1, characterized in that, The process of parsing the potential associated semantic label groups for the target robot based on all semantic labels in the image information includes: Identify the category corresponding to the contour features, determine the category text, and use the category text as a semantic label; Based on the probability of semantic tag combinations appearing simultaneously in a single sample corpus, potential associated semantic tag combinations are determined. If the probability is greater than a predetermined probability threshold, the corresponding semantic tag combination is determined to be a potential associated semantic tag group.

3. The robot control method based on speech analysis according to claim 1, characterized in that, The process of selecting semantically guided corpora for the target robot from the sample corpus based on semantic tags in the potentially related semantic tag group includes: Identify each semantic tag in a potentially associated semantic tag group; Sample corpora containing the aforementioned semantic tags are selected from the sample corpus; The sample corpus is pre-built and stores several sample corpora, each of which is obtained by textualizing the voice control data of the target robot.

4. The robot control method based on speech analysis according to claim 1, characterized in that, The process of determining the ambiguity tendency of the voice control data includes, If the instruction ambiguity parameter is greater than or equal to the preset instruction ambiguity threshold, it is determined that there is an instruction ambiguity tendency. If the instruction ambiguity parameter is less than the preset instruction ambiguity threshold, it is determined that there is no instruction ambiguity tendency.

5. The robot control method based on speech analysis according to claim 1, characterized in that, Since the voice control data does not tend to be ambiguous, the voice control data is transcribed into text and used as a confidence command text.

6. The robot control method based on speech analysis according to claim 1, characterized in that, The process of determining the semantic cluster of the text data corresponding to the voice control data includes, Calculate the semantic relevance between keywords in the text data, and cluster the keywords based on the semantic relevance to obtain semantic clusters; Among them, the semantic correlation between any keywords in the semantic cluster is greater than the preset semantic cluster correlation threshold.

7. The robot control method based on speech analysis according to claim 1, characterized in that, The process of determining the confidence set cluster includes, Determine the mean semantic correlation between each keyword in the semantic cluster and the sample corpus in the semantically guided corpus group; If the mean of the semantic relevance is greater than the predetermined confidence relevance threshold, then the semantic cluster is determined as a confidence cluster. The remaining portion of the text data is identified as a semantic discrete cluster.

8. The robot control method based on speech analysis according to claim 1, characterized in that, The process of performing synonym expansion on semantic discrete clusters to generate several expanded texts includes: The semantic discrete clusters are split into several keywords; The keywords were replaced with candidate words to obtain several expanded texts; The candidate words are synonyms of the keywords.

9. The robot control method based on speech analysis according to claim 1, characterized in that, The process of filtering the extended text based on the aforementioned confidence set clusters to obtain the confidence instruction text includes the following steps: The keywords in the confidence set cluster are determined, and sample corpora containing the keywords are selected from the semantically guided corpus. The extended texts are compared with the sample corpus to determine the semantic relevance of the extended texts; The extended text corresponding to the highest semantic relevance is selected as the confidence instruction text.

Citation Information

Patent Citations

  • Robot voice control method and device, robot and medium

    CN110136704A

  • Fuzzy adaptive robot system capable of recognizing voice demand and working method thereof

    CN106054602A

  • Robot target navigation method, device, equipment and medium

    CN120686826A