Speech recognition and control method, device and system, electronic equipment and storage medium

By identifying and replacing the user's fuzzy voice data, generating the replaced voice data and confirming it with the user, the control error problem caused by the wrong recognition of voice data is solved, and the accuracy of voice recognition and user experience are improved.

CN120236571APending Publication Date: 2025-07-01DAIKIN INDUSTRIES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311850780.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the voice data of the user may be misidentified, causing the air processing device to execute incorrect control instructions, reduce control efficiency and affect the user experience, especially when the voice data of the user is not recognized or cannot be correctly recognized.

Method used

By obtaining the user's voice data, identifying the fuzzy voice data therein, querying the pre-stored replacement data to generate the replaced voice data, broadcasting confirmation information to the user, and executing control instructions when confirming that it is correct.

Benefits of technology

It improves the accuracy of voice recognition, reduces the need for users to issue voice commands again, avoids the execution of incorrect control commands, and improves user experience and device control efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236571A_ABST
    Figure CN120236571A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition and control method, device and system, electronic equipment and a storage medium. The method comprises the following steps: acquiring first voice data of a user; identifying the first voice data, and when first fuzzy voice data exists in the first voice data, determining whether first replacement data, corresponding to the first fuzzy voice data, of the user is pre-stored or not; when the first replacement data exists, generating second voice data according to the first replacement data and the first voice data; generating a first control instruction according to the second voice data; first confirmation information of the first control instruction is broadcasted to the user in a voice mode, and when the first control instruction is confirmed to be correct, operation corresponding to the first control instruction is executed. Therefore, personalized complementation is carried out on the fuzzy part in the voice data of the user, the voice recognition accuracy and the voice control efficiency can be improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition and control, and particularly to a speech recognition and control method, device, system, electronic device and computer-readable storage medium. Background Art

[0002] With the development of science and technology and urban construction, speech recognition technology has been gradually applied to all aspects of people's lives. For example, in places such as office buildings, apartments, schools, shopping malls, etc., devices that can be controlled based on speech, such as air handling devices controlled by speech, have gradually become popular.

[0003] Currently, in the existing technology of device control based on speech, a control instruction is generated based on the speech issued by the user, so as to achieve the control of the device. For example, the speech data of the user is obtained, and the speech data is recognized to obtain a speech instruction corresponding to the control of the device, and then the corresponding control is directly executed on the device based on the speech instruction.

[0004] It should be noted that the above introduction of the technical background is only for the convenience of clearly and completely explaining the technical solution of the present application and facilitating the understanding of those skilled in the art. It cannot be considered that the above technical solutions are well-known to those skilled in the art just because these solutions are described in the background art part of the present application. Summary of the Invention

[0005] The inventors found that in the above-mentioned existing technology, in some cases, the speech data of the user may be misrecognized, and further the control instruction executed on the air handling device may also be incorrect. For example, when the user's Mandarin is not standard, or the dialect is not standard, or the user has an accent, or the user has a unique vocal habit or expression method, some data in the user's speech data cannot be recognized or cannot be correctly recognized. At this time, the user's intention may not be smoothly executed, and the device that the user wants to control cannot perform the corresponding action in time; in addition, the user may need to say the speech instruction again, or the user needs to operate the device in other ways, for example, the user needs to operate the air conditioner again through a remote controller or a wired controller, thereby reducing the control efficiency and affecting the user experience.

[0006] In addition, in the case where the speech data of the user cannot be recognized or cannot be correctly recognized, if the corresponding control instruction is directly executed according to the recognition result, the device may be incorrectly controlled, thereby further reducing the control efficiency and affecting the user experience;

[0007] In addition, in the case where the speech data of the user cannot be recognized all the time or no control instruction can be generated based on the speech recognition result, the user may be troubled, affecting the user experience.

[0008] To solve one or more of the above problems, an embodiment of the present application provides a voice recognition and control method, apparatus, system, electronic device, and computer-readable storage medium.

[0009] According to a first aspect of an embodiment of the present application, there is provided a voice recognition and control method, the method comprising:

[0010] Obtaining first voice data of a user; recognizing the first voice data, and when there is first fuzzy voice data in the first voice data, determining whether there is first replacement data corresponding to the first fuzzy voice data of the user pre-stored; when there is the first replacement data, generating second voice data according to the first replacement data and the first voice data; generating a first control instruction according to the second voice data; voice-broadcasting a first confirmation message of the first control instruction to the user, and when the first control instruction is confirmed to be correct, performing an operation corresponding to the first control instruction.

[0011] According to a second aspect of an embodiment of the present application, there is provided a voice recognition and control system, the system comprising a control device, a voice interaction device, and at least one controlled device; the control device is configured to control the at least one controlled device according to the voice recognition and control method described in the first aspect of the embodiment of the present application.

[0012] According to a third aspect of an embodiment of the present application, there is provided a voice recognition and control apparatus, the apparatus comprising:

[0013] An obtaining unit configured to obtain first voice data of a user; a recognition unit configured to recognize the first voice data; a determination unit configured to determine whether there is first replacement data corresponding to the first fuzzy voice data of the user pre-stored when there is first fuzzy voice data in the first voice data; a voice generation unit configured to generate second voice data according to the first replacement data and the first voice data when there is the first replacement data; a control instruction generation unit configured to generate a first control instruction according to the second voice data; a confirmation unit configured to voice-broadcast a first confirmation message of the first control instruction to the user; and an execution unit configured to perform an operation corresponding to the first control instruction when the first control instruction is confirmed to be correct.

[0014] According to a fourth aspect of an embodiment of the present application, there is provided an electronic device, the electronic device comprising: a memory storing a computer program; and a processor, which when executing the computer program, implements the voice recognition and control method described in the first aspect of the embodiment of the present application.

[0015] According to a fifth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voice recognition and control method described in the first aspect of the embodiments of the present application is implemented.

[0016] One of the beneficial effects of the embodiments of the present application is that when recognizing the voice data of a user, in the case where there is unrecognizable fuzzy voice data in the user's voice data and replacement data corresponding to the fuzzy voice data of the user is pre-stored, replacement voice data is generated based on the replacement data. Thus, it is possible to perform personalized completion of the fuzzy voice data in the user's voice data, improving the accuracy of voice recognition; at the same time, it is possible to reduce or avoid the user from issuing a voice command again or needing to control the device in other ways, improving the user experience and device control efficiency.

[0017] Moreover, a control instruction is generated based on the replacement voice data, and a confirmation message is broadcast to the user by voice. When the user confirms it is correct, the control instruction is executed. Thus, it is possible to avoid executing incorrect control instructions, improve control efficiency, and further improve the user experience.

[0018] In addition, since a confirmation message is broadcast to the user by voice, that is, a feedback on the user's voice control is given, the user can realize that their intention is actively responded to, thereby gaining the user's understanding and favor, and further improving the user experience.

[0019] Referring to the following description and drawings, specific embodiments of the present application are disclosed in detail, indicating the ways in which the principles of the present application can be adopted. It should be understood that the embodiments of the present application are not limited in scope thereby. Within the spirit and terms of the appended claims, the embodiments of the present application include many changes, modifications, and equivalents.

[0020] The characteristic information described and illustrated for one embodiment can be used in the same or similar way in one or more other embodiments, combined with the characteristic information in other embodiments, or replace the characteristic information in other embodiments.

[0021] It should be emphasized that the term "comprising / including" when used herein refers to the presence of characteristic information, whole, step, or component, but does not exclude the presence or addition of one or more other characteristic information, whole, step, or component. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Many aspects of the present application can be better understood with reference to the following drawings. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged or reduced. The elements and feature information described in one drawing or one embodiment of the present application can be combined with the elements and feature information shown in one or more other drawings or embodiments. In addition, in the drawings, like reference numerals denote corresponding components in several drawings and can be used to indicate the corresponding components used in more than one embodiment.

[0023] In the drawings:

[0024] Figure 1 is a block diagram of the speech recognition and control method according to an embodiment of the present application;

[0025] Figure 2 is a block diagram of determining the first replacement data corresponding to the first fuzzy speech data according to an embodiment of the present application;

[0026] Figure 3 is a flowchart of determining the first replacement data corresponding to the first fuzzy speech data according to an embodiment of the present application;

[0027] Figure 4 is a block diagram of generating the first control instruction according to the second speech data according to an embodiment of the present application;

[0028] Figure 5 is a block diagram of generating the third speech data according to an embodiment of the present application;

[0029] Figure 6 is a block diagram of the steps after generating the third speech data according to an embodiment of the present application;

[0030] Figure 7 is a flowchart of one implementation manner of the speech recognition and control method according to an embodiment of the present application;

[0031] Figure 8 is a schematic block diagram of the speech recognition and control system according to an embodiment of the present application;

[0032] Figure 9 is a schematic block diagram of the speech recognition and control device according to an embodiment of the present application;

[0033] Figure 10 is a schematic block diagram of the system composition of the electronic device according to an embodiment of the present application. Detailed Embodiments

[0034] The preferred embodiments of the present application will be described below with reference to the drawings.

[0035] Embodiment 1

[0036] Embodiment 1 of the present application provides a voice recognition and control method.

[0037] Figure 1 It is a block diagram of the voice recognition and control method of the embodiment of the present application. As Figure 1 shown, the method includes:

[0038] Step 101: Obtain the first voice data of the user and recognize the first voice data;

[0039] Step 102: When there is first fuzzy voice data in the first voice data, determine whether there is first replacement data corresponding to the first fuzzy voice data of the user stored in advance;

[0040] Step 103: When there is the first replacement data, generate second voice data according to the first replacement data and the first voice data;

[0041] Step 104: Generate a first control instruction according to the second voice data;

[0042] Step 105: Voice broadcast a first confirmation message of the first control instruction to the user;

[0043] Step 106: When the first control instruction is confirmed to be correct, execute an operation corresponding to the first control instruction.

[0044] Thus, when recognizing the user's voice data, in the case where there is unrecognizable fuzzy voice data in the user's voice data and there is replacement data corresponding to the fuzzy voice data of the user stored in advance, replacement voice data is generated based on the replacement data. Thus, it is possible to perform personalized completion of the fuzzy voice data in the user's voice data, improve the accuracy of voice recognition; at the same time, it is possible to reduce or avoid the user from issuing a voice command again or needing to control the device in other ways, improving the user experience and device control efficiency;

[0045] Moreover, a control instruction is generated based on the replacement voice data, and a confirmation message is voice broadcast to the user, and the control instruction is executed only when the user confirms it is correct. Thus, it is possible to avoid executing incorrect control instructions, improve control efficiency, and further improve the user experience;

[0046] In addition, since a confirmation message is voice broadcast to the user, that is, feedback is given to the user's voice control, the user can realize that their intention is actively responded to, so that the user's understanding and favor can be obtained, further improving the user experience.

[0047] In some embodiments, the voice recognition and control method generates control instructions for a controlled device based on the user's voice data. Among them, the controlled device includes one of an air handling device and a smart device.

[0048] In some embodiments, the air handling device includes at least one of, but is not limited to, an indoor unit, a air supply device, a humidity control device, a floor heating device, and a sensor.

[0049] In some embodiments, the indoor unit is an indoor unit in an air conditioning system, and the air conditioning system can be a commercial air conditioning system or a household air conditioning system.

[0050] In some embodiments, the air conditioning system may include at least one set of outdoor units and at least one indoor unit connected to each set of outdoor units. That is to say, in the air conditioning system, one or more sets of outdoor units can be included, and each set of outdoor units includes at least one outdoor unit; for one set of outdoor units, the set of outdoor units is connected to at least one indoor unit.

[0051] For example, the air conditioning system includes one outdoor unit and one indoor unit connected to the outdoor unit.

[0052] For example, the air conditioning system includes one outdoor unit and at least two indoor units connected to the outdoor unit.

[0053] For example, the air conditioning system includes at least two outdoor units and at least two indoor units connected to the at least two outdoor units.

[0054] For example, one set of outdoor units and at least one indoor unit connected to the one set of outdoor units form a refrigerant system, or multiple sets of outdoor units and at least one indoor unit respectively connected to the multiple sets of outdoor units form multiple refrigerant systems. Thus, the air conditioning system can include one refrigerant system or multiple refrigerant systems.

[0055] In some embodiments, the outdoor unit and the indoor unit can be air conditioning devices of various models, various types, various forms, and various capacities. For example, the form of the indoor unit can be four-way air outlet, two-way air outlet, duct machine, floor air outlet, or skirting air outlet, etc.; the form of the outdoor unit can be single-fan upwind, double-fan upwind, single-fan front wind, or double-fan front wind, etc.

[0056] In some embodiments, the air supply device includes at least one of a fresh air treatment device, a total heat exchanger (for example, with or without an internal circulation function), and a ventilation device (for example, with or without an internal circulation function).

[0057] In some embodiments, the humidity control device has at least one of a humidifying function and a dehumidifying function. For example, the humidity control device is disposed on the output side of the air supply device, and humidifies or dehumidifies the air output by the air supply device, and introduces the humidified or dehumidified air into the indoor space. For example, the humidity control device has a heat exchanger therein, and the heat exchanger can operate as an evaporator.

[0058] In some embodiments, the floor heating device is connected to the outdoor unit through a refrigerant pipeline, and the floor heating device is connected to the floor heating pipeline in the indoor space through a floor heating water pipeline. For example, the outdoor unit and the floor heating device form an air source heat pump water heater, which uses the heat in the air and the operation of the compressor to heat water; the floor heating device may include a water heat exchange unit, and the refrigerant flowing out of the compressor of the outdoor unit enters the water heat exchange unit of the floor heating device through the refrigerant pipeline and exchanges heat with water, so as to achieve the purpose of heating water, and the heated water is introduced into the floor heating pipeline in the indoor space through the floor heating water pipeline to heat the indoor space. In some embodiments, the outdoor unit and the floor heating device are separate or integrated.

[0059] In some embodiments, the intelligent device includes at least one of, but is not limited to, an intelligent socket, an intelligent lighting device, and an intelligent curtain.

[0060] The following specifically describes each step of the voice control method of the air handling device according to the embodiments of the present invention.

[0061] In some scenarios, the user needs to control the device by voice. For example, when cooking in the kitchen, taking a bath in the bathroom, or lying in bed, etc., it is inconvenient for the user to operate with hands, and then the device can be controlled by voice.

[0062] Alternatively, the user is accustomed to controlling the device by voice. For example, even when it is convenient for the user to operate with hands, the user can also control the device by voice.

[0063] In step 101, the first voice data of the user is acquired, and the first voice data is recognized.

[0064] In some embodiments, the first voice data may be a voice signal collected from the user, that is, the first voice data may be audio data; in this case, correspondingly, the first fuzzy voice data, the first replacement data corresponding to the first fuzzy voice data, and the second voice data generated according to the first replacement data and the first voice data are all voice signals, that is, audio data. That is, the recognition and replacement of fuzzy voice are directly based on audio data.

[0065] In some other embodiments, the first voice data may also be text data obtained by converting the collected voice signal into text, that is, the first voice data may also be text data based on the voice signal. In this case, correspondingly, the first fuzzy voice data, the first replacement data corresponding to the first fuzzy voice data, and the second voice data generated according to the first replacement data and the first voice data are all text data. That is, after converting from audio data to text data, fuzzy voice recognition and replacement are performed based on the text data.

[0066] For example, voice-to-text conversion can be performed through an AI model such as a pre-trained neural network model, and this application places no restrictions on this.

[0067] In some embodiments, the user's first voice data can be obtained through various devices with voice collection functions.

[0068] For example, when the controlled device has a voice collection function, the first voice data can be obtained by listening through the controlled device; or

[0069] The first voice data can be obtained by listening through a central controller, a voice control device, a wired controller, a remote controller, an intelligent controller, a user terminal, and an intelligent device connected to the controlled device, etc. Among them, the user terminal can be various terminal devices used by the user, such as a smart phone, a smart wearable device, etc., and the intelligent device includes, for example, a smart TV, a smart screen, etc.

[0070] In some embodiments, the listening method for the user's voice data can be set by the user.

[0071] For example, the user's voice data can be listened to in real time within a preset time after the user says a preset wake-up word, or can be continuously listened to in real time without a wake-up word and with the user's authorization. This application places no restrictions on the acquisition timing of the first voice data.

[0072] In some embodiments, identifying the first voice data includes:

[0073] Identifying whether there is first fuzzy voice data in the first voice data.

[0074] In some embodiments, the first fuzzy voice data includes at least one unrecognizable voice unit in the first voice data. The voice unit includes at least one of a phoneme, a syllable, a character, and a word.

[0075] Among them, the unrecognizable speech units include the speech units that cannot be converted into text. For example, the first speech data of the user is "Adjust the temperature in the bedroom to twenty? degrees". Since the user did not clearly pronounce "?" or "?" is a unique pronunciation of the user, resulting in the text corresponding to "?" not being found in the recognizable language. Therefore, the text data or audio data represented by "?" is an unrecognizable speech unit, that is, the first fuzzy speech data in the first speech data.

[0076] Alternatively, the unrecognizable speech units include the speech units that can be converted into text, but their own meanings are not clear or the speech units that make the overall meaning of the first speech data unclear. For example, the first speech data of the user is "Adjust the temperature in the inner room to twenty-seven degrees". Among them, "inner room" is just the user's habitual expression for their own bedroom. Although it is correctly converted into text, it is impossible to clearly determine which space the user wants to adjust the temperature according to this expression, resulting in the unclear meaning of the first speech data. Therefore, the text data "inner room" is an unrecognizable speech unit, that is, the first fuzzy speech data in the first speech data.

[0077] For example, speech recognition and speech-to-text conversion can be performed through AI models such as pre-trained neural network models. The present application does not limit this.

[0078] In some embodiments, the device for obtaining the first speech data and the device for recognizing the first speech data may be the same, that is, the device for voice monitoring can also perform speech recognition; or, the device for obtaining the first speech data and the device for recognizing the first speech data may be different. For example, the device for voice monitoring sends the monitored first speech data to the server, and the server performs speech recognition and obtains the text data converted from the first speech data from the server.

[0079] In step 102, when there is first fuzzy speech data in the first speech data, it is determined whether there is first replacement data of the user corresponding to the first fuzzy speech data pre-stored.

[0080] In some embodiments, the replacement data is recognizable speech data. The replacement data corresponding to the fuzzy speech data refers to the recognizable speech data that has the same meaning as the fuzzy speech data. The fuzzy speech data and the replacement data are stored in the database correspondingly. Thus, for the unrecognizable fuzzy speech data, its corresponding recognizable replacement data can be found by querying the database, thereby improving the recognition rate of the speech data.

[0081] In some embodiments, the correspondence between the fuzzy voice data and the replacement data includes: one modulus voice data corresponds to one replacement data. For example, a user is accustomed to using fuzzy voice data A1 to represent meaning 1, but the fuzzy voice data A1 cannot be recognized. At this time, there is a replacement data B1 that also represents meaning 1 and can be recognized. Therefore, a correspondence between the fuzzy voice data A1 and the replacement data B1 can be established, for example, expressed as {A1|B1}. Thus, the replacement data B1 can be determined as the replacement data corresponding to the fuzzy voice data A1 through this correspondence.

[0082] Alternatively, the correspondence between the fuzzy voice data and the replacement data includes: one fuzzy voice data corresponds to multiple replacement data. A user is accustomed to using fuzzy voice data A2 to represent meaning 2, but the fuzzy voice data A2 cannot be recognized. At this time, there are replacement data B21 and replacement data B22 that also represent meaning 2 and can be recognized. Therefore, a correspondence between the fuzzy voice data A2, the replacement data B21, and the replacement data B22 can be established, for example, expressed as {A2|B21,B22}. Thus, the replacement data B21 or the replacement data B22 can be determined as the replacement data corresponding to the fuzzy voice data A2 through this correspondence, and the "one-to-many" storage form can reduce the amount of stored data and save the storage space of the database.

[0083] Alternatively, the correspondence between the fuzzy voice data and the replacement data includes: multiple fuzzy voice data correspond to one replacement data. For example, a user is accustomed to using fuzzy voice data A31 and fuzzy voice data A32 to represent meaning 3, but both the fuzzy voice data A31 and the fuzzy voice data A32 cannot be recognized. At this time, there is a replacement data B3 that also represents meaning 3 and can be recognized. Therefore, a correspondence between the fuzzy voice data A31, the fuzzy voice data A32, and the replacement data B3 can be established, for example, expressed as {A31,A32|B3}. Thus, the replacement data B3 can be determined as the replacement data corresponding to the fuzzy voice data A31, or the replacement data B3 can be determined as the replacement data corresponding to the fuzzy voice data A32, and the "many-to-one" storage form can reduce the amount of stored data and save the storage space of the database.

[0084] Alternatively, the correspondence between the fuzzy voice data and the replacement data includes: multiple modulus voice data corresponding to multiple replacement data. For example, a user is accustomed to using fuzzy voice data A41 and fuzzy voice data A42 to represent meaning 4, but neither fuzzy voice data A41 nor fuzzy voice data A42 can be recognized. At this time, there are replacement data B41 and replacement data B42 that also represent meaning 4 and can be recognized. Therefore, a correspondence between fuzzy voice data A41, fuzzy voice data A42 and replacement data B41, replacement data B42 can be established, for example, expressed as {A41, A42|B41, B42}. Thus, the replacement data B41 or replacement data B42 can be determined as the replacement data corresponding to the fuzzy voice data A41 through this correspondence, and the replacement data B41 or replacement data B42 can be determined as the replacement data corresponding to the fuzzy voice data A42. Moreover, the "many-to-many" storage form can reduce the amount of stored data and save the storage space of the database.

[0085] In some embodiments, a plurality of fuzzy voice data sets are pre-stored in the database, and each fuzzy voice data set stores at least one set of fuzzy voice data and its corresponding replacement data. Each fuzzy voice data set corresponds to a user respectively. That is, the fuzzy voice data in the same fuzzy voice data set are all unrecognizable voice data of the same user, and the replacement data corresponding to the fuzzy voice data are voice data that have the same meaning as the fuzzy voice data of the user and can be recognized. Thus, by setting independent fuzzy voice data sets for different users respectively, personalized replacement recognition of the fuzzy voices of different users can be performed. For example, when user 1 and user 2 use the same fuzzy voice data to express different meanings, the replacement data corresponding to the fuzzy voice data in the pre-stored fuzzy voice data sets for user 1 and user 2 are different. Thus, the accuracy of user voice data recognition can be improved.

[0086] In some embodiments, when a fuzzy voice data set corresponding to a user is established in the database, the identification information of the user is associated with its corresponding fuzzy voice data set. Thus, a fuzzy voice data set can be uniquely determined through the identification information of the user, and then the fuzzy voice data of the user and its corresponding replacement data can be queried in the determined fuzzy voice data set. Thus, the accuracy of the user's voice data recognition can be improved.

[0087] In some embodiments, the identification information of the user includes voiceprint information. At this time, the user can be authenticated according to the voiceprint information in the user's voice data to obtain the voiceprint information of the user. And the fuzzy voice data set corresponding to the user can be found in the database according to the voiceprint information of the user. Thus, authenticating the user through the voiceprint information of the voice data can reduce the operation process and improve the authentication efficiency and user experience.

[0088] In addition, the user identification information may further include fingerprint information, face information, iris information, etc., which are not limited in this application.

[0089] Figure 2 It is a block diagram for determining the first replacement data corresponding to the first fuzzy voice data in an embodiment of this application. As Figure 2 shown, determining whether there is pre-stored the first replacement data corresponding to the user for the first fuzzy voice data includes:

[0090] Step S11, perform identity recognition on the user according to the voiceprint information in the first voice data to obtain the user's identification information;

[0091] Step S12, query in the database the fuzzy voice data set corresponding to the user's identification information, where the fuzzy voice data set includes at least one fuzzy voice data and the replacement data corresponding to the fuzzy voice data;

[0092] Step S13, when there is a fuzzy voice data set corresponding to the user's identification information in the database and the first fuzzy voice data exists in the fuzzy voice data set, obtain the first replacement data corresponding to the first fuzzy voice data from the database.

[0093] In step S11, using the voiceprint information of the user's first voice data to perform identity verification on the user can avoid the user re-entering voice data or other verification information, simplify the verification process, and improve the verification efficiency.

[0094] In steps S12 and S13, first query in the database whether there is a fuzzy voice data set associated with the identification information through the user's identification information; in the case of existence, then query the first replacement data corresponding to the first fuzzy voice data from the fuzzy voice data set. Thus, the query efficiency of the first replacement data can be improved, and further the voice recognition and control efficiency can be improved, enhancing the user experience.

[0095] In some embodiments, when there are multiple replacement data corresponding to the first fuzzy voice data, randomly select one of the multiple replacement data to be determined as the first replacement data, or select one of the multiple replacement data according to a preset rule, for example, the first replacement data is the one that was newly added among the multiple replacement data, or the one that was initially added, or the one that was selected the most times within a preset time, which is not limited in this application.

[0096] In some embodiments, the voice recognition and control method further includes:

[0097] When there is no fuzzy voice dataset associated with the identification information of the user obtained in step S11 in the database, it is determined that there is no first replacement data corresponding to the first fuzzy voice data;

[0098] Alternatively, when there is a fuzzy voice dataset associated with the identification information of the user obtained in step S11 in the database, but the first fuzzy voice data does not exist in this fuzzy voice dataset, it is determined that there is no first replacement data corresponding to the first fuzzy voice data.

[0099] Figure 3 It is a flowchart of determining the first replacement data corresponding to the first fuzzy voice data in an embodiment of the present application. As Figure 3 shown, when there is a first fuzzy voice data in the user's first voice data, the process of determining the first replacement data corresponding to the first fuzzy voice data is as follows:

[0100] Step 201, perform identity recognition on the user according to the voiceprint information in the first voice data to obtain the identification information of the user;

[0101] Step 202, determine whether there is a fuzzy voice dataset associated with this identification information in the database; if it exists, execute step 203; if it does not exist, execute step 206;

[0102] Step 203, determine that this fuzzy voice dataset is the fuzzy voice dataset corresponding to this user;

[0103] Step 204, determine whether the first fuzzy voice data exists in this fuzzy voice dataset; if it exists, execute step 205; if it does not exist, execute step 206;

[0104] Step 205, determine the replacement data corresponding to the first fuzzy voice data as the first replacement data;

[0105] Step 206, determine that the first replacement data does not exist.

[0106] In the above embodiment, the corresponding fuzzy voice dataset is queried according to the identification information in the user's first voice data, and the replacement data corresponding to the first fuzzy voice data in the first voice data is queried in this fuzzy voice dataset. Thus, for different users, the replacement data corresponding to their fuzzy voice data can be personalized determined, improving the accuracy of speech recognition.

[0107] The above explains the determination of whether there is the first replacement data in step 102. In step 103, when there is the first replacement data, a second voice data is generated according to the first replacement data and the first voice data.

[0108] In some embodiments, generating second speech data based on the first replacement data and the first speech data includes:

[0109] Replacing the first blurred speech data in the first speech data with the first replacement data to generate the second speech data.

[0110] In some embodiments, for the case where both the first speech data and the first replacement data are audio data, replacing the first blurred speech data in the first speech data with the first replacement data to generate the second speech data includes:

[0111] Segmenting and deleting the audio data corresponding to the first blurred speech data in the first speech data;

[0112] Concatenating the first replacement data to the position where the audio data corresponding to the first blurred speech data was originally in the first speech data to obtain the second speech data.

[0113] In this way, the unrecognizable speech units in the user's audio are directly replaced, and the obtained second speech data is also audio data. At this time, the second speech data can be completely converted into text.

[0114] In some embodiments, for the case where both the first speech data and the first replacement data are text data, replacing the first blurred speech data in the first speech data with the first replacement data to generate the second speech data includes:

[0115] Segmenting and deleting the text data corresponding to the first blurred speech data in the first speech data;

[0116] Concatenating the first replacement data to the position where the text data corresponding to the first blurred speech data was originally in the first speech data to obtain the second speech data.

[0117] Alternatively, for the case where the first speech data is audio data and the first replacement data is text data, replacing the first blurred speech data in the first speech data with the first replacement data to generate the second speech data includes:

[0118] Converting the first speech data into text data;

[0119] Segmenting and deleting the text data corresponding to the first blurred speech data in the text data of the first speech data;

[0120] Concatenating the first replacement data to the position where the text data corresponding to the first blurred speech data was originally in the text data of the first speech data to obtain the second speech data.

[0121] In this way, the unrecognized speech units in the text corresponding to the user's audio are directly replaced, and the obtained second speech data is also text data. At this time, the user's intention can be directly recognized based on the second speech data.

[0122] In some embodiments, generating the second speech data according to the first replacement data and the first speech data further includes:

[0123] After replacing the first fuzzy speech data in the first speech data with the first replacement data, the word order of the replaced speech data is adjusted to obtain the second speech data.

[0124] Thus, the speech units of the second speech data can all be recognized, and the word order of the second speech data is more conducive to recognition, which can improve the speech recognition efficiency.

[0125] In the above embodiments, the first fuzzy speech data is replaced with the recognizable first replacement data to obtain the recognizable second speech data. Thus, based on the second speech data for speech recognition, it is possible to avoid the user from repeating the input of speech, reduce the probability of speech recognition failure, and improve the user experience.

[0126] In step 104, a first control instruction is generated according to the second speech data.

[0127] Figure 4 It is a block diagram of generating the first control instruction according to the second speech data in an embodiment of the present application. As Figure 4 shown, when the second speech data is audio data, generating the first control instruction according to the second speech data includes:

[0128] Step S21, converting the second speech data into text data;

[0129] Step S22, when the text data contains keywords related to control, generating the first control instruction according to the keywords.

[0130] In step S21, speech recognition and speech-to-text conversion can be performed through a pre-trained neural network model or other AI models.

[0131] In some embodiments, the second speech data may be a control instruction issued by the user. For example, the second speech data is "raise the temperature", "turn on the humidification device", "switch to the ventilation mode", etc.;

[0132] Or, the second speech data may be the user's daily conversation. For example, the second speech data is "it's too hot in the room", "the floor is a bit damp", "this smell is so unpleasant", etc.

[0133] In step S22, it is determined whether the text data contains keywords related to "control".

[0134] If it contains, a first control instruction is generated according to the keyword.

[0135] For example, assume the second voice data is "raise the temperature". After converting this voice data into text, it contains the keywords "raise" and "temperature" in the database, so the corresponding first control instruction can be determined as "increase the temperature". Among them, the value of the temperature increase corresponding to this control instruction can be set by the user. For example, every time a control instruction of "increase the temperature" is received, the temperature is raised by 1 degree; or every time a control instruction of "increase the temperature" is received, the temperature is raised by 0.5 degree; or when multiple control instructions of "increase the temperature" are continuously received, as the number of times increases, the adjustment span of the temperature gradually decreases. For example, when the control instruction is received for the first time, the temperature is raised by 2 degrees, when the control instruction is received for the second time, the temperature is raised by 1.5 degrees, when the control instruction is received for the third time, the temperature is raised by 1 degree... and so on. This application does not limit this.

[0136] Another example, assume the second voice data is "the floor is a bit damp". After converting this voice data into text, it contains the keywords "floor" and "damp" in the database, so it can be determined that the user hopes to dehumidify. Based on this, the corresponding first control instruction is determined as "turn on the dehumidification device".

[0137] In some embodiments, when the second voice data is text data, generating a first control instruction according to the second voice data includes:

[0138] When the text data of the second voice data contains keywords related to control, a first control instruction is generated according to the keyword.

[0139] The above steps can be implemented with reference to Figure 4 step S22 of, and will not be repeated here.

[0140] When the second voice data is text data, in the process of generating a first control instruction according to the second voice data, the step of "converting the second voice data into text data" does not need to be executed, that is, when the second voice data is text data, Figure 4 step S21 in can be omitted.

[0141] It should be noted that the above are only examples provided by this application, and the actual application is not limited to this.

[0142] Through the above embodiments, the second voice data is converted into a control instruction that the device can recognize, which is convenient for controlling the device and also convenient for confirming to the user.

[0143] In step 105, the first confirmation information of the first control instruction is voice broadcast to the user.

[0144] In some embodiments, the first confirmation information of the voice broadcast requires a response from the user. For example, the first confirmation information of the voice broadcast includes information that requires a response from the user, or the first confirmation information of the voice broadcast is in a form that requires a response from the user. A form that requires a response from the user is, for example, that the first confirmation information is in the form of an interrogative sentence.

[0145] Taking the corresponding first control instruction of "turn on the indoor unit" as an example, the first confirmation information that requires a response from the user is, for example, "Please confirm whether you want to turn on the indoor unit", "Are you sure you want to turn on the indoor unit?" or other information.

[0146] In some embodiments, the first confirmation information of the voice broadcast does not require a response from the user. For example, the first confirmation information of the voice broadcast does not include information that requires a response from the user, or the first confirmation information of the voice broadcast is in a form that does not require a response from the user.

[0147] Taking the corresponding first control instruction of "turn on the indoor unit" as an example, the first confirmation information that does not require a response from the user is, for example, "The indoor unit will be turned on for you soon".

[0148] Thus, by voice broadcasting the first confirmation information to the user, the execution of the first control instruction can be further confirmed to the user, which is beneficial to improving the user experience.

[0149] In step 106, when the first control instruction is confirmed to be correct, the operation corresponding to the first control instruction is executed.

[0150] In some embodiments, when the first confirmation information of the voice broadcast requires a response from the user, the first control instruction is confirmed to be correct, including: receiving the voice response information of the user within a preset time period, and the voice response information includes information related to confirming the execution of the first control instruction.

[0151] Among them, the "information related to confirming the execution of the first control instruction" is, for example, "confirm the execution of (the first control instruction)", "please execute (the first control instruction)", etc. When the first control instruction is different, the "information related to confirming the execution of the first control instruction" can be adaptively adjusted. For example, assuming the first control instruction is "turn on the indoor unit", the first confirmation information that needs to be replied by the user in the voice broadcast is "please confirm whether you want to turn on the indoor unit", then the information related to confirming the execution of "turn on the indoor unit" included in the user's voice reply information is, for example, "confirm to turn on (the indoor unit)", "please turn on (the indoor unit)", etc.; another example, assuming the first control instruction is "adjust the temperature to 20°C", the first confirmation information that needs to be replied by the user in the voice broadcast is "please confirm whether you want to adjust the temperature to 20°C", then the information related to confirming the execution of "adjust the temperature to 20°C" included in the user's voice reply information is, for example, "confirm to adjust (the temperature)", "please adjust (the temperature)", etc.

[0152] Therefore, in the case where the confirmation information in the voice broadcast requires the user to reply, when the user's voice reply information is received within the preset time period and the voice reply information includes the information related to confirming the execution of the first control instruction, it is considered that the first control instruction is correctly confirmed and the operation corresponding to the first control instruction is executed.

[0153] In some embodiments, when the first confirmation information in the voice broadcast does not require the user to reply, the correct confirmation of the first control instruction includes: not receiving the voice information related to rejecting the execution of the first control instruction from the user within the preset time period.

[0154] Among them, the "voice information related to rejecting the execution of the first control instruction" is, for example, "do not execute (the first control instruction)", "stop executing (the first control instruction)", "cancel the execution of (the first control instruction)", etc. When the first control instruction is different, the "voice information related to rejecting the execution of the first control instruction" can be adaptively adjusted. For example, assuming the first control instruction is "turn on the indoor unit", the first confirmation information that does not require the user to reply in the voice broadcast is "about to turn on the indoor unit for you", then the voice information of the user related to rejecting the execution of "turn on the indoor unit" is, for example, "do not turn on (the indoor unit)", "stop turning on (the indoor unit)", "cancel turning on (the indoor unit)", etc.; another example, assuming the first control instruction is "adjust the temperature to 20°C", the first confirmation information that does not require the user to reply in the voice broadcast is "about to adjust the temperature to 20°C for you", then the voice information of the user related to rejecting the execution of "adjust the temperature to 20°C" is, for example, "do not adjust (the temperature)", "stop adjusting (the temperature)", "cancel adjusting (the temperature)", etc.

[0155] Therefore, in the case where the first confirmation message for voice broadcast does not require a user response, when no voice message related to rejecting the execution of the first control instruction from the user is received within a preset duration, it is considered that the first control instruction is correctly confirmed, and the operation corresponding to the first control instruction is executed.

[0156] In the above embodiment, after generating a control instruction corresponding to the user's voice data, a confirmation message is broadcast to the user's voice to confirm the execution of the control instruction, and the operation corresponding to the control instruction is executed after the user confirms that the control instruction is correct. Thereby, the execution of incorrect control instructions can be avoided, and the user experience can be improved.

[0157] In some embodiments, the voice recognition and control method further includes:

[0158] When the first control instruction is confirmed to be incorrect, third voice data for speculating the first voice data is generated according to the first voice data and reference data.

[0159] In some embodiments, in the case where the first confirmation message for voice broadcast requires a user response, if no voice response message from the user is received within a preset duration, or the voice response message from the user does not include information related to confirming the execution of the first control instruction, it is determined that the first control instruction is incorrect.

[0160] In the case where the confirmation message for voice broadcast does not require a user response, if a voice message related to rejecting the execution of the first control instruction from the user is received within a preset duration, it is determined that the first control instruction is incorrect.

[0161] In some embodiments, the preset duration is set by the user. For example, the preset duration is 30 seconds.

[0162] In some embodiments, the reference data refers to data used to assist in generating the third voice data. The reference data includes at least one of the user's usage habit data, the user's current health status data, the current weather data, and the current indoor environment data.

[0163] Figure 5 It is a block diagram of generating the third voice data in an embodiment of the present application. As Figure 5 shown, generating third voice data for speculating the first voice data according to the first voice data and reference data includes:

[0164] Step S31, determining target reference data according to the keywords that can be recognized in the first voice data;

[0165] Step S32, generating the third voice data according to the target reference data and the first voice data.

[0166] In step S31, the target reference data refers to the reference data related to the keywords that can be recognized in the first voice data. The corresponding relationship between the target reference data and the keywords can be preset.

[0167] For example, assume that the first voice data is "The indoor temperature is too?". Here, "?" is the fuzzy voice data that cannot be recognized. The keywords that can be recognized in the first voice data include "temperature". Thus, the "current weather data", "current indoor environment data", and "user usage habit data" related to "temperature" are determined as the target reference data.

[0168] In step S32, the fuzzy voice data that cannot be recognized in the first voice data is speculated based on the target reference data, and the third voice data is generated based on the speculation result and the first voice data.

[0169] For example, as described in the previous example, assume that the outdoor temperature in the "current weather data" is 32 degrees, the indoor temperature in the "current indoor environment data" is 29 degrees, and the indoor temperature setting in the "user usage habit data" for the current season and current time period is between 22 and 25 degrees. Thus, it can be speculated that the meaning of the fuzzy voice data in the user's first voice data may be "high", and then the "third voice data for speculating the first voice data" is determined as "The indoor temperature is too high".

[0170] Thus, in the case where the first control instruction is incorrect, speculating on the first voice data in combination with the reference data can reduce user operations and improve the user experience.

[0171] In some embodiments, multiple third voice data are generated.

[0172] For example, as described in the previous example, the first voice data is "The indoor temperature is too?". The outdoor temperature in the "current weather data" is 25 degrees, the indoor temperature in the "current indoor environment data" is 24 degrees, and the indoor temperature setting in the "user usage habit data" for the current season and current time period is between 22 and 25 degrees. Thus, it may not be possible to speculate the meaning of the fuzzy voice data in the user's first voice data. At this time, multiple meanings that the user may want to express can be respectively generated into the third voice data. For example, the user may think that the temperature is on the high side, and the corresponding third voice data is "The indoor temperature is too high"; the user may also think that the temperature is on the low side, and the corresponding third voice data is "The indoor temperature is too low".

[0173] Thus, in the case where the first control instruction is incorrect, speculating on the first voice data in combination with the reference data to obtain multiple third voice data can provide multiple choices for the user, improve the accuracy of speculation, and further improve the user experience.

[0174] In some embodiments, after Figure 3 step 206 shown, that is, after determining that the first replacement data does not exist, the speech recognition and control method further includes:

[0175] When it is determined that the first replacement data does not exist, third speech data for speculating the first speech data is generated according to the first speech data and reference data.

[0176] In this embodiment, the content related to the third speech data can be referred to the steps executed in the foregoing "when the first control instruction is confirmed to be incorrect", which will not be repeated here.

[0177] Figure 6 is a block diagram of the steps after generating the third speech data in an embodiment of the present application. As Figure 6 shown, after generating the third speech data, the speech recognition and control method further includes:

[0178] Step S41, confirming the third speech data to the user through voice interaction;

[0179] Step S42, generating a second control instruction according to the result of the user's confirmation; and

[0180] Step S43, performing an operation corresponding to the second control instruction.

[0181] In step S41, when generating one piece of third speech data, the third speech data is broadcast by voice to confirm to the user whether the third speech data is correct; when generating multiple pieces of third speech data, the multiple pieces of third speech data are broadcast by voice to confirm to the user whether the correct third speech data is included in the multiple pieces of third speech data, and which one is the correct third speech data.

[0182] In step S42, when it is confirmed by the user that there is correct third speech data among one or more pieces of third speech data, a second control instruction is generated according to the correct third speech data confirmed by the user. The specific manner of generating the second control instruction according to the correct third speech data can refer to the manner of generating the first control instruction according to the second speech data in the foregoing embodiments, which will not be elaborated here.

[0183] In step S43, an operation corresponding to the second control instruction generated according to the correct third speech data is performed.

[0184] Thus, when the first control instruction is incorrect, for one or more third voice data inferred from the first voice data, confirm with the user, generate a second control instruction based on the third voice data confirmed by the user, and execute the corresponding operation. Thus, by confirming the third voice data with the user, it is possible to avoid executing incorrect control instructions and improve control efficiency.

[0185] In some embodiments, the voice recognition and control method further includes:

[0186] Establish or update the fuzzy voice data set corresponding to the user in the database according to the result of the user's confirmation.

[0187] Among them, the result of the user's confirmation includes the third voice data confirmed by the user.

[0188] In some embodiments, when the fuzzy voice data set corresponding to the user is stored in the database, that is, the identification information of the user determined according to the voiceprint information of the user's first voice data, in the case of querying the associated fuzzy voice data set in the data, update the fuzzy voice data set according to the result of the user's confirmation.

[0189] In some embodiments, when there is a first fuzzy voice data in the first voice data and its corresponding first replacement data in the fuzzy voice data set corresponding to the user, updating the fuzzy voice data set includes: using the voice unit corresponding to the first fuzzy voice data in the third voice data confirmed by the user to update the first replacement data.

[0190] In some embodiments, when there is no first fuzzy voice data in the first voice data and its corresponding first replacement data in the fuzzy voice data set corresponding to the user, updating the fuzzy voice data set includes: correspondingly storing the first fuzzy voice data and the voice unit corresponding to the first fuzzy voice data in the third voice data confirmed by the user into the fuzzy voice data set of the user.

[0191] In some embodiments, when the fuzzy voice data set corresponding to the user is not stored in the database, that is, the identification information of the user determined according to the voiceprint information of the user's first voice data, in the case of being unable to query the associated fuzzy voice data set in the data, establish a fuzzy voice data set according to the result of the user's confirmation.

[0192] In some embodiments, establishing a fuzzy voice data set includes: determining the identification information of the user according to the voiceprint information of the user's first voice data, establishing a fuzzy voice data set associated with the identification information, and correspondingly storing the first fuzzy voice data and the voice unit corresponding to the first fuzzy voice data in the third voice data confirmed by the user into the established fuzzy voice data set.

[0193] Through the above embodiments, by establishing or updating the user's fuzzy voice data set according to the result confirmed by the user, the accuracy of user voice recognition can be further improved, the control efficiency can be improved, and thus the user experience can be improved.

[0194] In some embodiments, the voice recognition and control method further includes:

[0195] When the user does not confirm the third voice data, at least one of the voice recognition failure, the reason for the failure, and the suggestion for voice control is announced by voice.

[0196] For example, the reason for the failure may be "unable to recognize the first voice data", "there are unrecognizable voice units"; the suggestions for voice control may be "please re-enter the voice data", "please slow down the voice speed", "please enter the voice data according to the regulations", etc.

[0197] In some embodiments, after Figure 1 step 101 shown, the voice recognition and control method further includes:

[0198] When there is no first fuzzy voice data in the first voice data, a third control instruction is generated according to the first voice data; and an operation corresponding to the third control instruction is executed.

[0199] Thus, when there is no first fuzzy voice data, the first voice data can be recognized. At this time, a control instruction is generated according to the first voice data, and the operation corresponding to the control instruction is directly executed without confirmation from the user, which can improve the control efficiency and thus improve the user experience.

[0200] In some embodiments, the voice recognition and control method further includes:

[0201] Providing instructions related to control to the user through at least one of a centralized controller, a line controller, a remote controller, a user terminal, and a smart device. The smart device includes, but is not limited to, a smart TV, a smart speaker, a smart screen, etc.

[0202] For example, a device with a voice announcement function can be used to provide instructions related to control to the user by voice announcement, such as through the centralized controller for voice announcement; a device with a display function can also be used to provide instructions related to control to the user by means of text and / or picture push, such as through the smart screen for push, through the APP of the user terminal for push, etc.

[0203] In some embodiments, the instructions related to control include: keywords related to the control instruction.

[0204] For example, the names of the controlled devices, such as the first indoor unit, the second indoor unit, and the humidifying device, facilitate users to distinguish different controlled devices; the action expressions executed by the controlled devices, such as "turn on", "start", "open", "close", "switch", etc., facilitate users to issue accurate instructions.

[0205] In some embodiments, the description related to control includes: suggestions related to voice control.

[0206] For example, the expression mode of voice data, such as "When you want to adjust the indoor temperature, you can try saying: 'Increase the room temperature' or 'Adjust the temperature to 20 degrees'", which is beneficial to improving the accuracy rate of voice recognition.

[0207] Thus, it is convenient for users to understand and learn the voice control function, can guide users to accurately issue control instructions for specific devices, is beneficial to improving the efficiency of voice recognition and control, and improves the user experience.

[0208] Figure 7 is a flowchart of an implementation manner of the voice recognition and control method according to an embodiment of the present application. As Figure 7 shown, the voice recognition and control method includes:

[0209] Step 301: Obtain the first voice data of the user;

[0210] Step 302: Determine whether there is first fuzzy voice data in the first voice data; if so, execute Step 305; if not, execute Step 303;

[0211] Step 303, generate a third control instruction according to the first voice data;

[0212] Step 304, execute the third control instruction and end.

[0213] Step 305: Determine whether there is first replacement data corresponding to the first fuzzy voice data of the user stored in advance; if so, execute Step 306; if not, execute Step 311;

[0214] Step 306: Generate second voice data according to the first replacement data and the first voice data;

[0215] Step 307, generate a first control instruction according to the second voice data;

[0216] Step 308, voice broadcast the first confirmation information of the first control instruction to the user;

[0217] Step 309, determine whether the first control instruction is correct; if so, execute Step 310; if not, execute Step 311;

[0218] Step 310, execute the first control instruction, and end.

[0219] Step 311, generate third voice data for speculating the first voice data according to the first voice data and reference data;

[0220] Step 312, confirm the third voice data to the user through voice interaction;

[0221] Step 313, determine whether the user confirms the third voice data; if so, execute Step 314; if not, execute Step 316;

[0222] Step 314, generate a second control instruction according to the third voice data confirmed by the user;

[0223] Step 315, execute the second control instruction;

[0224] Step 316, establish or update the fuzzy voice data set corresponding to the user in the database according to the confirmation result of the user, and end.

[0225] Step 317, voice broadcast that the voice recognition fails, as well as the reasons and suggestions for the failure, and end.

[0226] For the specific implementation of Steps 301 - 317, reference can be made to Figure 1 and Figure 2 the relevant steps in and the foregoing embodiments, which will not be repeated here. Among them, the execution order of Step 315 and Step 316 is not limited to Figure 3 as shown, that is, Step 316 can also be executed first, followed by Step 315, or Step 315 and Step 316 can also be executed simultaneously.

[0227] According to the above embodiments, when recognizing the voice data of a user, in the case that there is fuzzy voice data that cannot be recognized in the user's voice data and there is replacement data corresponding to the fuzzy voice data of the user stored in advance, generate the replaced voice data based on the replacement data. Thus, it is possible to perform personalized completion of the fuzzy voice data in the user's voice data, improve the accuracy of voice recognition; at the same time, it can reduce or avoid the user from issuing voice commands again or needing to control the device in other ways, improving the user experience and device control efficiency;

[0228] Moreover, generate a control instruction based on the replaced voice data, and voice broadcast a confirmation message to the user, and then execute the control instruction when the user confirms correctly. Thus, it is possible to avoid executing incorrect control instructions, improve control efficiency, and further improve the user experience;

[0229] In addition, since the confirmation information is spoken to the user, that is, a feedback is given to the user's voice control, the user can realize that their intention is actively responded to, thereby gaining the user's understanding and favor, and further improving the user experience.

[0230] Embodiment 2

[0231] Embodiment 2 of the present application provides a voice recognition and control system, which corresponds to the voice recognition and control method described in Embodiment 1. For specific content, reference can be made to the description in Embodiment 1.

[0232] Figure 8 is a schematic block diagram of the voice recognition and control system of the embodiment of the present application. As Figure 8 shown, the voice recognition and control 10 includes at least one controlled device 11, a control device 12, and a voice interaction device 13;

[0233] The control device 12 is used to obtain the user's first voice data; recognize the first voice data, and when there is first fuzzy voice data in the first voice data, determine whether there is first replacement data corresponding to the first fuzzy voice data of the user pre-stored; when there is the first replacement data, generate second voice data according to the first replacement data and the first voice data; generate a first control instruction according to the second voice data; speak the first confirmation information of the first control instruction to the user, and when the first control instruction is confirmed to be correct, execute the operation corresponding to the first control instruction.

[0234] In some embodiments, the controlled device 11 includes at least one of an air handling device and a smart device.

[0235] In some embodiments, the air handling device includes at least one of an indoor unit, a ventilation device, a humidity control device, and a floor heating device.

[0236] In some embodiments, the smart device includes at least one of a smart socket, a smart lighting device, and a smart curtain.

[0237] In some embodiments, the control device 12 includes at least one of a server, a gateway device, a centralized controller, a line controller, a remote controller, a user terminal, and a control board.

[0238] In some embodiments, the voice interaction device 13 includes at least one of a centralized controller, a line controller, a remote controller, a user terminal, a smart speaker, and a smart display device.

[0239] In some embodiments, the control device 12 has the function of receiving the user's voice data and making announcements through voice. At this time, the voice interaction device 13 and the control device 12 can be the same device, and the functions implemented by the voice interaction device 13 in this application are implemented by the control device 12.

[0240] For example, both the control device 12 and the voice interaction device 13 are centralized controllers.

[0241] For the specific processing procedures of the above-mentioned various devices, reference can be made to the relevant records in Embodiment 1, which will not be elaborated here.

[0242] According to the above embodiments, when identifying the user's voice data, if there is unrecognizable fuzzy voice data in the user's voice data and replacement data corresponding to the fuzzy voice data of the user is pre-stored, replacement voice data is generated based on the replacement data. Thus, it is possible to perform personalized completion of the fuzzy voice data in the user's voice data, improving the accuracy of voice recognition; at the same time, it can reduce or avoid the user from issuing voice commands again or needing to control the device in other ways, improving the user's usage experience and device control efficiency;

[0243] And, a control instruction is generated based on the replacement voice data, and a confirmation message is announced to the user by voice. The control instruction is executed only when the user confirms it is correct. Thus, it is possible to avoid executing incorrect control instructions, improve control efficiency, and further improve the user's usage experience;

[0244] In addition, since a confirmation message is announced to the user by voice, that is, a feedback on the user's voice control, the user can realize that their intention is actively responded to, thereby gaining the user's understanding and favor, and further improving the user experience.

[0245] Embodiment 3

[0246] Embodiment 3 of the present application provides a voice recognition and control device. The voice recognition and control device corresponds to the voice recognition and control method described in Embodiment 1, and the specific content can be referred to the records in Embodiment 1.

[0247] Figure 9 is a schematic block diagram of the voice recognition and control device of the embodiment of the present application. As Figure 9 shown, the device 20 includes:

[0248] An acquisition unit 21, which is used to acquire the user's first voice data;

[0249] An identification unit 22, which is used to identify the first voice data;

[0250] A determination unit 23 configured to determine whether first replacement data corresponding to the user for the first ambiguous speech data is pre-stored when there is first ambiguous speech data in the first speech data;

[0251] A first generation unit 24 configured to generate second speech data according to the first replacement data and the first speech data when there is the first replacement data;

[0252] A second generation unit 25 configured to generate a first control instruction according to the second speech data;

[0253] A confirmation unit 26 configured to voice broadcast a first confirmation message of the first control instruction to the user;

[0254] An execution unit 27 configured to execute an operation corresponding to the first control instruction when the first control instruction is confirmed to be correct.

[0255] For the specific functions of the above units, reference may be made to the relevant steps in Embodiment 1, which will not be elaborated here.

[0256] According to the above embodiment, when recognizing the user's speech data, in the case where there is unrecognizable ambiguous speech data in the user's speech data and the replacement data corresponding to the user for the ambiguous speech data is pre-stored, the replaced speech data is generated based on the replacement data. Thus, the ambiguous speech data in the user's speech data can be personalized and completed, improving the accuracy of speech recognition; at the same time, it can reduce or avoid the user from issuing a voice command again or needing to control the device in other ways, improving the user experience and the device control efficiency;

[0257] Moreover, a control instruction is generated based on the replaced speech data, and a confirmation message is voice broadcast to the user, and the control instruction is executed only when the user confirms it is correct. Thus, it can avoid executing incorrect control instructions, improve the control efficiency, and further improve the user experience;

[0258] In addition, since a confirmation message is voice broadcast to the user, that is, a feedback on the user's voice control, the user can realize that their intention is actively responded to, thereby obtaining the understanding and favor of the user, and further improving the user experience.

[0259] Embodiment 4

[0260] Embodiment 4 of the present application provides an electronic device. The steps executed by the processor of the electronic device correspond to all or part of the steps of the speech recognition and control method described in Embodiment 1, and the specific content can be referred to the description in Embodiment 1.

[0261] Figure 10It is a schematic block diagram of the system composition of the electronic device according to an embodiment of the present application. As Figure 10 shown, the electronic device 300 may include a processor 310 and a memory 320; the memory 320 is coupled to the processor 310. It should be noted that this figure is exemplary; other types of structures can also be used to supplement or replace this structure to achieve telecommunication functions or other functions.

[0262] In one embodiment, the processor 310 may be configured to: obtain the first voice data of the user; identify the first voice data, and when there is first ambiguous voice data in the first voice data, determine whether there is first replacement data corresponding to the first ambiguous voice data of the user pre-stored; when there is the first replacement data, generate second voice data according to the first replacement data and the first voice data; generate a first control instruction according to the second voice data; voice broadcast a first confirmation message of the first control instruction to the user, and when the first control instruction is confirmed to be correct, execute an operation corresponding to the first control instruction.

[0263] As Figure 10 shown, the electronic device 300 may further include: a communication module 330, an input unit 340, a display 350, a speaker 360, a microphone 370, and a power supply 380. It should be noted that the electronic device 300 does not necessarily have to include all the components Figure 10 shown; in addition, the electronic device 300 may further include components not shown in Figure 10 which can refer to related technologies.

[0264] As Figure 10 shown, the processor 310 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The processor 310 receives inputs and controls the operations of the various components of the electronic device 300.

[0265] Among them, the memory 320 may be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. Various data can be stored, and in addition, programs for executing relevant information can also be stored. And the processor 310 can execute the programs stored in the memory 320 to achieve information storage or processing, etc. The functions of other components are similar to those in the prior art and will not be elaborated here. The various components of the electronic device 300 can be implemented by dedicated hardware, firmware, software, or a combination thereof without departing from the scope of the present invention.

[0266] According to the above embodiments, when recognizing the user's voice data, if there is unclear voice data in the user's voice data that cannot be recognized and replacement data corresponding to the unclear voice data of the user is pre-stored, the replaced voice data is generated based on the replacement data. Thus, it is possible to perform personalized completion of the unclear voice data in the user's voice data, improving the accuracy of voice recognition; at the same time, it is possible to reduce or avoid the user from issuing voice commands again or needing to control the device in other ways, improving the user experience and device control efficiency;

[0267] Moreover, a control instruction is generated based on the replaced voice data, and a confirmation message is broadcast to the user by voice. When the user confirms it is correct, the control instruction is executed. Thus, it is possible to avoid executing incorrect control instructions, improve control efficiency, and further improve the user experience;

[0268] In addition, since a confirmation message is broadcast to the user by voice, that is, feedback is provided for the user's voice control, the user can realize that their intention is actively responded to, thereby gaining the user's understanding and favor, and further improving the user experience.

[0269] The embodiment of the present application also provides a computer-readable program, where when the program is executed, the program causes the computer to execute the voice recognition and control method described in Embodiment 1 of the present application.

[0270] The embodiment of the present application also provides a computer-readable storage medium, where the storage medium stores a computer program, and the computer-readable program causes the computer to execute the voice recognition and control method described in Embodiment 1 of the present application.

[0271] The above-mentioned devices and methods in the embodiments of the present application can be implemented by hardware or by a combination of hardware and software. The present application relates to such a computer-readable program that when the program is executed by a logic component, it can cause the logic component to implement the above-mentioned device or component, or cause the logic component to implement the above-mentioned various methods or steps.

[0272] The embodiment of the present application also relates to a storage medium for storing the above program, such as a hard disk, a magnetic disk, an optical disk, a DVD, a flash memory, etc.

[0273] It should be noted that the limitations on the steps involved in the present application, without affecting the implementation of the specific solution, do not constitute a limitation on the order of the steps. The steps written in the front can be executed first, or can be executed later, or even can be executed simultaneously. As long as the solution can be implemented, it should be regarded as falling within the protection scope of the present application.

[0274] The present application has been described in conjunction with specific embodiments, but those skilled in the art should understand that these descriptions are exemplary and not limitations on the scope of protection of the present application. Those skilled in the art can make various variations and modifications to the present application according to the spirit and principle of the present application, and these variations and modifications are also within the scope of the present application.

Claims

1. A voice recognition and control method, characterized in that The method comprises: Acquire first voice data of a user, and recognize the first voice data; When there is first ambiguous voice data in the first voice data, determining whether first replacement data of the user corresponding to the first ambiguous voice data is pre-stored; When the first replacement data exists, generating second voice data according to the first replacement data and the first voice data; generating a first control instruction according to the second voice data; First confirmation information of the first control instruction is voice broadcasted to the user, and when the first control instruction is confirmed to be correct, an operation corresponding to the first control instruction is executed.

2. The method according to claim 1, characterized in that: The first ambiguous speech data includes at least one speech unit that cannot be recognized in the first speech data; The speech unit includes at least one of a phoneme, a syllable, a character, and a word.

3. The method according to claim 1, wherein Determining whether first replacement data corresponding to the first ambiguous speech data of the user is pre-stored includes: Identify the user according to the voiceprint information in the first voice data to obtain identification information of the user; Searching a database for a fuzzy speech data set corresponding to the identification information of the user, wherein the fuzzy speech data set includes at least one fuzzy speech data and replacement data corresponding to the fuzzy speech data; When an ambiguous speech data set corresponding to the identification information of the user exists in a database and the first ambiguous speech data exists in the ambiguous speech data set, the first replacement data corresponding to the first ambiguous speech data is acquired from the database.

4. The method according to claim 3, characterized in that In the ambiguous speech data set, the correspondence between the ambiguous speech data and the replacement data includes at least one of the following: One analog voice data corresponds to one replacement data; One ambiguous speech data corresponds to multiple replacement data; A plurality of ambiguous speech data corresponds to one replacement data; The plurality of analog-to-digital voice data corresponds to the plurality of replacement data.

5. The method according to claim 3, wherein When there is no ambiguous speech data set corresponding to the identification information of the user in the database, or the first ambiguous speech data does not exist in the ambiguous speech data set corresponding to the identification information of the user, it is determined that there is no first replacement data corresponding to the first ambiguous speech data.

6. The method according to claim 1 or 5, characterized in that The method further comprises: When the first replacement data does not exist, or when the first control instruction is confirmed to be incorrect, third voice data for estimating the first voice data is generated based on the first voice data and reference data.

7. The method according to claim 6, wherein The method further comprises: confirming the third voice data to the user through voice interaction; generating a second control instruction according to the result of the user confirmation; and Execute an operation corresponding to the second control instruction.

8. The method according to claim 7, characterized in that, The method further comprises: A fuzzy speech data set corresponding to the user in a database is established or updated according to the result of the user confirmation.

9. The method according to claim 1, characterized in that The method further comprises: When the first voice data does not contain first ambiguous voice data, generating a third control instruction according to the first voice data; and Perform an operation corresponding to the third control instruction.

10. The method according to claim 1, wherein The method further includes: Providing control-related instructions to the user through at least one of a centralized controller, a line controller, a remote controller, and a user terminal.

11. A voice recognition and control system, characterized in that, The system includes a control device, a voice interaction device, and at least one controlled device; The control device is configured to control the at least one controlled device according to the voice recognition and control method according to any one of claims 1 to 10.

12. The system according to claim 11, wherein The controlled device includes at least one of an air handling device and a smart device.

13. The system according to claim 12, wherein The air handling device includes at least one of an indoor unit, a ventilation device, a humidity control device, a floor heating device, and a sensor; The smart device includes at least one of a smart socket, a smart lighting device, and a smart curtain.

14. The system according to claim 11, wherein The control device includes at least one of a server, a gateway device, a centralized controller, a line controller, a remote controller, a user terminal, and a control board; The voice interaction device includes at least one of a centralized controller, a line controller, a remote controller, a user terminal, a smart speaker, and a smart display device.

15. A voice recognition and control device, characterized in that, The apparatus includes: An acquisition unit configured to acquire first voice data of a user; An identification unit configured to identify the first voice data; A determination unit configured to determine whether first replacement data corresponding to the first fuzzy voice data of the user is pre-stored when there is first fuzzy voice data in the first voice data; A voice generation unit configured to generate second voice data according to the first replacement data and the first voice data when there is the first replacement data; A control instruction generation unit configured to generate a first control instruction according to the second voice data; A confirmation unit configured to voice broadcast a first confirmation message of the first control instruction to the user; An execution unit configured to perform an operation corresponding to the first control instruction when the first control instruction is confirmed to be correct.

16. An electronic device, characterized in that, The electronic device includes: A memory storing a computer program; and A processor that implements the voice recognition and control method according to any one of claims 1-10 when executing the computer program.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the voice recognition and control method according to any one of claims 1-10.