Speech recognition methods, devices, equipment and storage media
By identifying the emotional changes, logical relevance, and content relevance between the user's current voice data and the most recent voice interaction data, the misidentification problem of voice interaction devices is solved, and the accuracy of voice interaction is improved.
Patent Information
- Application Number
- CN202210074128.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-01-21
AI Technical Summary
During voice interaction, a user issues a command and wants to end the interaction, but the voice interaction device misidentifies the voice that does not belong to the user's interaction, resulting in a misjudgment.
Collect the user's current voice data and the most recent voice interaction data, identify emotion change values, logical relevance values, and content relevance values, and generate voice data recognition results to determine whether it belongs to voice interaction data.
By analyzing emotional changes, logical relevance, and content relevance, misrecognition by voice interaction devices is avoided, thus improving the accuracy of voice interaction.
Smart Images

Figure CN114420122B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart homes, and more particularly to a voice recognition method, device, equipment, and storage medium. Background Technology
[0002] Smart home devices interact with users through voice interaction devices. Currently, during voice interaction, it frequently happens that after a user issues a command and tries to end the interaction, the voice interaction device misinterprets the voice and attempts to identify a voice that was not intended for the user, leading to misjudgment. Summary of the Invention
[0003] This application provides a speech recognition method, apparatus, device, and storage medium to solve the problem that during speech interaction, when a user issues a command and wants to end the interaction, the speech interaction device may misidentify the speech that was not originally intended for the user, resulting in misjudgment.
[0004] Firstly, a speech recognition method is provided, including:
[0005] Collect the user's current voice data and obtain the most recent voice interaction data, which is used to enable interaction with smart home devices;
[0006] Identify the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data;
[0007] Based on the emotion change value, the logical relevance value, and the content relevance value, a recognition result for the current voice data is generated, and the recognition result indicates whether the current voice data belongs to voice interaction data or not.
[0008] Optionally, identifying the emotional change value of the current voice data relative to the voice interaction data includes:
[0009] Obtain the first emotional state reflected by the current voice data and the second emotional state reflected by the voice interaction data;
[0010] Obtain the first emotional score corresponding to the first emotional state and the second emotional score corresponding to the second emotional state;
[0011] Calculate the difference between the first sentiment score and the second sentiment score;
[0012] The emotional change value is determined based on the score difference.
[0013] Optionally, determining the sentiment change value based on the score difference includes:
[0014] Calculate the product of the score difference and the first weight percentage, where the first weight percentage indicates the importance of the emotional change;
[0015] The product value is used as the emotional change value.
[0016] Optionally, obtaining the first emotional state reflected by the current voice data includes:
[0017] Extract tone features and emotional keywords from the current speech data;
[0018] Obtain the tone score corresponding to the tone feature and the keyword score corresponding to the sentiment keyword;
[0019] The tone score is adjusted using the keyword score to obtain the overall score;
[0020] The first emotional state is obtained based on the comprehensive score.
[0021] Optionally, identifying the logical relevance value of the current voice data relative to the voice interaction data includes:
[0022] Identify the first content category of the current voice data and the second content category of the voice interaction data;
[0023] When the first content category and the second content category are the same, the logical relevance value is determined based on the first score;
[0024] When the first content category and the second content category are different, the logical relevance value is determined based on the second score;
[0025] The first score is greater than the second score.
[0026] Optionally, identifying the content relevance value of the current voice data relative to the voice interaction data includes:
[0027] Identify the first content category and first semantic of the current voice data, and the second content category and second semantic of the voice interaction data;
[0028] When the first content category and the second content category are the same, and both the first semantics and the second semantics refer to the same smart home device, the content relevance value is determined based on the third score;
[0029] When the first content category and the second content category are the same, and the first semantics and the second semantics are for different smart home devices, the content relevance value is determined based on the fourth score;
[0030] When the first content category and the second content category are different, and both the first semantics and the second semantics refer to the same smart home device, the content relevance value is determined based on the fifth score;
[0031] When the first content category and the second content category are different, and the first semantics and the second speech are for different smart home devices, the content relevance value is determined based on the sixth score;
[0032] The third score is greater than the fifth score, the fifth score is greater than the fourth score, and the fourth score is greater than the sixth score.
[0033] Optionally, based on the emotion change value, the logical relevance value, and the content relevance value, a recognition result for the current speech data is generated, including:
[0034] Obtain the score corresponding to the sentiment change value, the score corresponding to the logical relevance value, and the score corresponding to the content relevance value;
[0035] The sum of the scores corresponding to the emotional change value, the logical relevance value, and the content relevance value is calculated to obtain the summed score.
[0036] When the summed score is greater than a preset score threshold, the recognition result indicates that the current voice data belongs to voice interaction data;
[0037] When the summed score is not greater than a preset score threshold, the recognition result indicates that the current voice data does not belong to voice interaction data.
[0038] Optionally, before identifying the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data, the method further includes:
[0039] Extract the voiceprint features of the current voice data and obtain the voiceprint features of the voice interaction data;
[0040] It is determined that the voiceprint features of the current voice data are the same as the voiceprint features of the voice interaction data.
[0041] Secondly, a voice recognition device is provided, comprising:
[0042] The acquisition unit is used to acquire the user's current voice data and the most recent voice interaction data, which is used to realize interaction with smart home devices;
[0043] The recognition unit is used to recognize the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data;
[0044] The generation unit is used to generate a recognition result of the current voice data based on the emotion change value, the logical relevance value, and the content relevance value, wherein the recognition result indicates whether the current voice data belongs to voice interaction data or not.
[0045] Thirdly, a voice interaction device is provided, comprising: a processor, a memory, and a communication bus, wherein the processor and the memory communicate with each other through the communication bus;
[0046] The memory is used to store computer programs;
[0047] The processor is used to execute the program stored in the memory to implement the speech recognition method described in the first aspect.
[0048] Fourthly, a computer-readable storage medium is provided, storing a computer program that, when executed by a processor, implements the speech recognition method described in the first aspect.
[0049] Compared with the prior art, the technical solution provided in this application has the following advantages: In the technical solution provided in this embodiment, the user's current voice data and the most recent voice interaction data are collected. The voice interaction data is used to realize interaction with smart home devices. The emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data are identified. Based on the emotional change value, logical relevance value, and content relevance value, a recognition result for the current voice data is generated. The technical solution provided in this embodiment, by using emotional change value, logical relevance value, and content relevance value, can identify whether the current voice data belongs to voice interaction data, thereby avoiding misidentification by the voice interaction device during the voice interaction process. Attached Figure Description
[0050] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a flowchart illustrating a speech recognition method in an embodiment of this application;
[0053] Figure 2 This is another flowchart illustrating the speech recognition method in the embodiments of this application;
[0054] Figure 3 This is a schematic diagram of the structure of the speech recognition device in the embodiments of this application;
[0055] Figure 4 This is a schematic diagram of the structure of the voice interaction device in the embodiments of this application. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0057] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0058] This application provides a speech recognition method, which can be applied to voice interaction devices in smart home devices; such as... Figure 1 As shown, the method may include the following steps:
[0059] Step 101: Collect the user's current voice data and obtain the most recent voice interaction data. The voice interaction data is used to realize interaction with smart home devices.
[0060] Step 102: Identify the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data.
[0061] The emotional change value indicates the degree of emotional change between the current voice data and the voice interaction data; the logical relevance value indicates the degree of logical relevance between the two voice data; and the content relevance value indicates the degree of content relevance between the two voice data.
[0062] The following describes the specific recognition process for sentiment change value, logical relevance value, and content relevance value:
[0063] Firstly, regarding the emotion change value, in one optional embodiment, the first emotion state reflected by the current voice data and the second emotion state reflected by the voice interaction data are obtained; the first emotion score corresponding to the first emotion state and the second emotion score corresponding to the second emotion state are obtained; the score difference between the first emotion score and the second emotion score is calculated; and the emotion change value is determined based on the score difference.
[0064] The emotional states reflected in the voice data in this embodiment may include requests, gentleness, joy, impatience, commands, etc., and this embodiment does not make specific limitations on them.
[0065] In this embodiment, the emotional state reflected in the speech data is identified through the tone features and emotional keywords of the speech data. Taking the first emotional state reflected in the current speech data as an example, in an optional embodiment, tone features and emotional keywords are extracted from the current speech data; tone scores corresponding to the tone features and keyword scores corresponding to the emotional keywords are obtained; the tone scores are corrected using the keyword scores to obtain a comprehensive score; and the first emotional state is obtained based on the comprehensive score.
[0066] In this embodiment, the tone features include, but are not limited to, volume and pitch. The application obtains the tone features of the current speech data by acquiring the simulated speech signal of the current speech data and extracting tone features from the simulated speech signal.
[0067] In the application, the text data of the current voice data can be obtained, and the text data can be segmented to obtain at least one segment. The at least one segment can be matched with a preset set of keywords to obtain the emotional keywords in the current voice data.
[0068] In this embodiment, a mapping relationship between emotional state and emotional score is pre-set. Therefore, the first emotional state and the second emotional state can be used to query the mapping relationship to obtain the first emotional score and the second emotional score.
[0069] In this embodiment, a pre-set correspondence between scores and emotional states is established, so the first emotional state can be obtained by querying the correspondence based on the comprehensive score.
[0070] It should be understood that the process of acquiring the second emotional state reflected in the voice interaction data is similar to the process of acquiring the first emotional state reflected in the current voice data, and will not be elaborated here.
[0071] In applications, since the influence of emotional changes, logical relevance, and content relevance on speech recognition results varies, this embodiment pre-sets weight percentages for each of these factors. Furthermore, when recognizing emotional change values, logical relevance values, and content relevance values, the corresponding weight percentages are applied.
[0072] In one alternative embodiment, when determining the sentiment change value based on the score difference, the product of the score difference and a first weight percentage is calculated, where the first weight percentage indicates the importance of the sentiment change; the product value is used as the sentiment change value.
[0073] It should be understood that the sum of the pre-set weight percentages for emotional change, logical relevance, and content relevance is 1.
[0074] In applications, depending on actual needs, no weight percentage can be set, meaning that by default, sentiment variation, logical relevance, and content relevance have the same impact on the speech recognition result. In this case, the score difference is the sentiment variation value.
[0075] Secondly, regarding the logical relevance value, in one optional embodiment, a first content category of the current voice data and a second content category of the voice interaction data are identified; when the first content category and the second content category are the same, a logical relevance value is determined based on a first score; when the first content category and the second content category are different, a logical relevance value is determined based on a second score; the first score is greater than the second score.
[0076] In this embodiment, the content categories include control and casual conversation. When identifying the content category of voice data, the semantics of the voice data can be obtained. If the semantics of the voice data belong to the control category, the content category of the voice data is determined to be control; otherwise, the content category of the voice data is determined to be casual conversation.
[0077] It should be understood that when the first content category and the second content category are the same, the logical relevance between the current voice data and the voice interaction data is determined to be high, and therefore a higher score (i.e., the first score) is used to determine the logical relevance value; when the first content category and the second content category are different, the logical relevance between the current voice data and the voice interaction data is determined to be low, and therefore a lower score (i.e., the second score) is used to determine the logical relevance value.
[0078] It should be understood that, similar to the sentiment change value, when a weighted percentage for logical relevance is set, the logical relevance value is the product of the first score and that weighted percentage, or the product of the second score and that weighted percentage. When no weighted percentage is set for logical relevance, either the first score or the second score is used as the logical relevance value.
[0079] Finally, regarding the content relevance value, in one optional embodiment, a first content category and a first semantic meaning of the current voice data, and a second content category and a second semantic meaning of the voice interaction data are identified; when the first content category and the second content category are the same, and both the first semantic meaning and the second semantic meaning refer to the same smart home device, a content relevance value is determined based on a third score; when the first content category and the second content category are the same, and both the first semantic meaning and the second semantic meaning refer to different smart home devices, a content relevance value is determined based on a fourth score; when the first content category and the second content category are different, and both the first semantic meaning and the second semantic meaning refer to the same smart home device, a content relevance value is determined based on a fifth score; when the first content category and the second content category are different, and both the first semantic meaning and the second semantic meaning refer to different smart home devices, a content relevance value is determined based on a sixth score; the third score is greater than the fifth score, the fifth score is greater than the fourth score, and the fourth score is greater than the sixth score.
[0080] In this embodiment, the semantics of voice data can be used for interactive objects and interactive action instructions in the voice data. For example, the semantics of the voice data "It's too hot, please turn on the air conditioner quickly" can be "Turn on the air conditioner".
[0081] It should be understood that when the first content category and the second content category are the same, and the first semantic and the second semantic refer to the same smart home device, it means that the two voice data are control-type or casual-type statements made by the user to the same smart home device in succession, and the content relevance of the two voice data is high. In this case, a higher score is used to determine the content relevance value. When the first content category and the second content category are the same, and the first semantic and the second semantic refer to different smart home devices, it means that the two voice data are likely not directed to the same smart home device, and a lower score is used to determine the content relevance value. When the first content category and the second content category are different, and the first semantic and the second semantic refer to the same smart home device, it means that the two voice data are directed to the same smart home device, but the content types of the interactive voice data are different. In this case, another lower score is used to determine the content relevance value. When the first content category and the second content category are different, and the first semantic and the second voice refer to different smart home devices, it means that the two voice data are neither directed to the same smart home device nor are they the same type of voice. In this case, a lowest score is used to determine the content relevance value.
[0082] It should be understood that regardless of which score from the third to the sixth score is used to determine the content relevance value, when a weighted percentage for content relevance is set, the product of that score and that weighted percentage is used as the content relevance value. When no weighted percentage for content relevance is set, the score is used as the content relevance value.
[0083] In this embodiment, in order to improve recognition efficiency, before recognizing the current voice data, it is also possible to determine whether the speaker of the current voice data and the speaker of the voice interaction data are the same person based on the voiceprint features. If it is determined that they are the same person, then the recognition of the current voice data is performed.
[0084] In one optional embodiment, the voiceprint features of the current voice data are extracted, and the voiceprint features of the voice interaction data are obtained; it is determined that the voiceprint features of the current voice data are the same as the voiceprint features of the voice interaction data.
[0085] It should be understood that when the voiceprint features of the current speech data are different from those of the voice interaction data, it means that the person who generated the current speech data and the person who generated the voice interaction data are not the same person. In this case, no processing is performed on the current speech data.
[0086] The application also allows setting a voice interaction time. If the current voice data collection time is outside the voice interaction time range, no processing will be performed on the current voice data. The voice interaction time refers to the period from when the voice interaction device is activated until it terminates.
[0087] Step 103: Based on the emotion change value, logical relevance value, and content relevance value, generate the recognition result of the current voice data. The recognition result indicates whether the current voice data belongs to voice interaction data or not.
[0088] In one optional embodiment, the scores corresponding to the emotion change value, the logical relevance value, and the content relevance value are obtained; the sum of the scores corresponding to the emotion change value, the logical relevance value, and the content relevance value is calculated to obtain a summed score; when the summed score is greater than a preset score threshold, the recognition result indicates that the current voice data belongs to voice interaction data; when the summed score is not greater than the preset score threshold, the recognition result indicates that the current voice data does not belong to voice interaction data.
[0089] In practice, the score threshold can be preset based on experience, and this embodiment does not impose any specific limitations on it.
[0090] It should be understood that when the recognition result of the current voice data indicates that the current voice data belongs to voice interaction data, the smart home device will perform the interactive operation corresponding to the current voice data.
[0091] In this embodiment, the technical solution collects the user's current voice data and obtains the most recent voice interaction data. The voice interaction data is used to interact with smart home devices. The solution identifies the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data. Based on these values, a recognition result for the current voice data is generated. This embodiment, by utilizing emotional change values, logical relevance values, and content relevance values, can identify whether the current voice data belongs to voice interaction data, thereby avoiding misidentification by the voice interaction device during the voice interaction process.
[0092] The above technical solution will be explained and illustrated with an example below, such as... Figure 2 As shown:
[0093] When the voice recognition of the voice interaction device is finished, wait for voice input;
[0094] The voice interaction device records the user's first voice and extracts voiceprint features as a comparison sample. If a voice signal is recorded while waiting for voice recording, the voiceprint features of the signal are extracted and compared with the voiceprint comparison sample. If the comparison finds that it is the same person, the next step is executed.
[0095] Voice interaction devices use recommendation systems to analyze factors such as the emotion of the voice, the logic between the voice and the previous voice, and the actual content of the voice.
[0096] The speech is classified based on the analysis results of the recommendation system, and relevant weights are calculated.
[0097] The weight ratios of the speech obtained from the analysis are compared with the standard weights.
[0098] When the weight is determined to be greater than the standard weight, the normal voice interaction process will proceed.
[0099] Based on the same concept, this application provides a speech recognition device. The specific implementation of this device can be found in the description of the method embodiments section; repeated details will not be repeated here. Figure 3 As shown, the device mainly includes:
[0100] The acquisition unit 301 is used to acquire the user's current voice data and the most recent voice interaction data. The voice interaction data is used to realize interaction with smart home devices.
[0101] The recognition unit 302 is used to recognize the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data;
[0102] The generation unit 303 is used to generate a recognition result of the current speech data based on the emotion change value, logical relevance value and content relevance value. The recognition result indicates whether the current speech data belongs to speech interaction data or not.
[0103] The identification unit 302 is used for:
[0104] Obtain the first emotional state reflected by the current voice data and the second emotional state reflected by the voice interaction data;
[0105] Obtain the first emotional score corresponding to the first emotional state, and the second emotional score corresponding to the second emotional state;
[0106] Calculate the difference between the first sentiment score and the second sentiment score;
[0107] The emotional change value is determined based on the score difference.
[0108] The identification unit 302 is used for:
[0109] Calculate the product of the score difference and the first weighted percentage, which indicates the importance of the sentiment change;
[0110] The product value is used as the sentiment change value.
[0111] The identification unit 302 is used for:
[0112] Extract intonation features and emotional keywords from the current speech data;
[0113] Obtain the tone score corresponding to the tone features, and the keyword score corresponding to the sentiment keywords;
[0114] The tone score was adjusted using keyword scoring to obtain the overall score.
[0115] The first emotional state is determined based on the overall score.
[0116] The identification unit 302 is used for:
[0117] Identify the first content category of the current voice data and the second content category of the voice interaction data;
[0118] When the first content category and the second content category are the same, the logical relevance value is determined based on the first score;
[0119] When the first content category and the second content category are different, the logical relevance value is determined based on the second score;
[0120] The first score is greater than the second score.
[0121] The identification unit 302 is used for:
[0122] Identify the first content category and first semantics of the current voice data, and the second content category and second semantics of the voice interaction data;
[0123] When the first content category and the second content category are the same, and both the first semantic and the second semantic refer to the same smart home device, the content relevance value is determined based on the third score;
[0124] When the first content category and the second content category are the same, and the first semantics and the second semantics are for different smart home devices, the content relevance value is determined based on the fourth score;
[0125] When the first content category and the second content category are different, and both the first semantic and the second semantic refer to the same smart home device, the content relevance value is determined based on the fifth score;
[0126] When the first content category and the second content category are different, and the first semantic and the second voice are for different smart home devices, the content relevance value is determined based on the sixth score;
[0127] The third score is greater than the fifth score, the fifth score is greater than the fourth score, and the fourth score is greater than the sixth score.
[0128] Production unit 303 is used for:
[0129] Obtain the scores corresponding to the sentiment change value, the logical relevance value, and the content relevance value;
[0130] The sum of the scores corresponding to the sentiment change value, the logical relevance value, and the content relevance value is calculated to obtain the summed score.
[0131] When the sum of the scores is greater than the preset score threshold, the recognition result indicates that the current voice data belongs to voice interaction data.
[0132] When the sum of the scores is not greater than the preset score threshold, the recognition result indicates that the current voice data does not belong to the voice interaction data.
[0133] This device is also used for:
[0134] Before identifying the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data, the voiceprint features of the current voice data and the voiceprint features of the voice interaction data are extracted.
[0135] It was determined that the voiceprint features of the current speech data are the same as those of the voice interaction data.
[0136] Based on the same concept, this application also provides a voice interaction device, such as... Figure 4As shown, the voice interaction device mainly includes a processor 401, a memory 402, and a communication bus 403. The processor 401 and the memory 402 communicate with each other via the communication bus 403. The memory 402 stores programs that can be executed by the processor 401. The processor 401 executes the programs stored in the memory 402 to achieve the following steps:
[0137] Collect the user's current voice data and obtain the most recent voice interaction data. The voice interaction data is used to enable interaction with smart home devices.
[0138] Identify the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data;
[0139] Based on the emotion change value, logical relevance value, and content relevance value, the recognition result of the current voice data is generated, and the recognition result indicates whether the current voice data belongs to voice interaction data or not.
[0140] The communication bus 403 mentioned in the aforementioned voice interaction device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus 403 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0141] The memory 402 may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor 401.
[0142] The processor 401 mentioned above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc., or a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0143] In another embodiment of this application, a computer-readable storage medium is provided, which stores a computer program that, when run on a computer, causes the computer to perform the speech recognition method described in the above embodiments.
[0144] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape, etc.), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0145] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0146] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A speech recognition method, characterized in that, include: Collect the user's current voice data and obtain the most recent voice interaction data, which is used to enable interaction with smart home devices; Identify the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data; Based on the emotion change value, the logical relevance value, and the content relevance value, a recognition result for the current voice data is generated, and the recognition result indicates whether the current voice data belongs to voice interaction data or not. The method of identifying the emotional change value of the current voice data relative to the voice interaction data includes: obtaining a first emotional state reflected by the current voice data and a second emotional state reflected by the voice interaction data; obtaining a first emotional score corresponding to the first emotional state and a second emotional score corresponding to the second emotional state; calculating the score difference between the first emotional score and the second emotional score; and determining the emotional change value based on the score difference. The process of obtaining the first emotional state reflected by the current speech data includes: extracting tone features and emotional keywords from the current speech data; obtaining tone scores corresponding to the tone features and keyword scores corresponding to the emotional keywords; correcting the tone scores using the keyword scores to obtain a comprehensive score; and obtaining the first emotional state based on the comprehensive score. The tone features include the volume and pitch of the current speech data; The method of identifying the logical relevance value of the current voice data relative to the voice interaction data includes: identifying a first content category of the current voice data and a second content category of the voice interaction data; when the first content category and the second content category are the same, determining the logical relevance value based on a first score; when the first content category and the second content category are different, determining the logical relevance value based on a second score; wherein the first score is greater than the second score.
2. The method according to claim 1, characterized in that, Based on the score difference, the emotional change value is determined, including: Calculate the product of the score difference and the first weight percentage, where the first weight percentage indicates the importance of the emotional change; The product value is used as the emotional change value.
3. The method according to claim 1, characterized in that, Identifying the content relevance value of the current voice data relative to the voice interaction data includes: Identify the first content category and first semantic of the current voice data, and the second content category and second semantic of the voice interaction data; When the first content category and the second content category are the same, and both the first semantics and the second semantics refer to the same smart home device, the content relevance value is determined based on the third score; When the first content category and the second content category are the same, and the first semantics and the second semantics are for different smart home devices, the content relevance value is determined based on the fourth score; When the first content category and the second content category are different, and both the first semantics and the second semantics refer to the same smart home device, the content relevance value is determined based on the fifth score; When the first content category and the second content category are different, and the first semantics and the second semantics are for different smart home devices, the content relevance value is determined based on the sixth score; The third score is greater than the fifth score, the fifth score is greater than the fourth score, and the fourth score is greater than the sixth score.
4. The method according to claim 1, characterized in that, Based on the emotion change value, the logical relevance value, and the content relevance value, a recognition result for the current speech data is generated, including: Obtain the score corresponding to the sentiment change value, the score corresponding to the logical relevance value, and the score corresponding to the content relevance value; The sum of the scores corresponding to the emotional change value, the logical relevance value, and the content relevance value is calculated to obtain the summed score. When the summed score is greater than a preset score threshold, the recognition result indicates that the current voice data belongs to voice interaction data; When the summed score is not greater than a preset score threshold, the recognition result indicates that the current voice data does not belong to voice interaction data.
5. The method according to claim 1, characterized in that, Before identifying the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data, the method further includes: Extract the voiceprint features of the current voice data and obtain the voiceprint features of the voice interaction data; It is determined that the voiceprint features of the current voice data are the same as the voiceprint features of the voice interaction data.
6. A voice recognition device, characterized in that, include: The acquisition unit is used to acquire the user's current voice data and the most recent voice interaction data, which is used to realize interaction with smart home devices; The recognition unit is used to recognize the emotional change value, logical relevance value, and content relevance value of the current voice data relative to the voice interaction data; The generation unit is used to generate a recognition result of the current voice data based on the emotion change value, the logical relevance value, and the content relevance value, wherein the recognition result indicates whether the current voice data belongs to voice interaction data or not. The method of identifying the emotional change value of the current voice data relative to the voice interaction data includes: obtaining a first emotional state reflected by the current voice data and a second emotional state reflected by the voice interaction data; obtaining a first emotional score corresponding to the first emotional state and a second emotional score corresponding to the second emotional state; calculating the score difference between the first emotional score and the second emotional score; and determining the emotional change value based on the score difference. The process of obtaining the first emotional state reflected by the current speech data includes: extracting tone features and emotional keywords from the current speech data; obtaining tone scores corresponding to the tone features and keyword scores corresponding to the emotional keywords; correcting the tone scores using the keyword scores to obtain a comprehensive score; and obtaining the first emotional state based on the comprehensive score. The tone features include the volume and pitch of the current speech data; The method of identifying the logical relevance value of the current voice data relative to the voice interaction data includes: identifying a first content category of the current voice data and a second content category of the voice interaction data; when the first content category and the second content category are the same, determining the logical relevance value based on a first score; when the first content category and the second content category are different, determining the logical relevance value based on a second score; wherein the first score is greater than the second score.
7. A voice interaction device, characterized in that, include: The processor, memory, and communication bus are used to communicate with each other. The memory is used to store computer programs; The processor is configured to execute the program stored in the memory to implement the speech recognition method according to any one of claims 1-5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech recognition method according to any one of claims 1-5.
Citation Information
Patent Citations
Session control method of voice robot, session control equipment and storage medium
CN112017629A
Voice interaction method and device, storage medium and electronic device
CN112992137A