Voice output method and air conditioner

By acquiring the target voice data of the air conditioner's current scene and preset tone, personalized response information is output, solving the problem of the monotonous voice broadcast format of the air conditioner and improving the user experience.

CN115540260BActive Publication Date: 2026-02-17FOSHAN SHUNDE MIDEA ELECTRONICS TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110736641.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2026-02-17
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

The existing air conditioners have a limited range of voice announcements, resulting in a poor user experience.

Method used

By acquiring voice control commands, the current scene of the air conditioner is determined, and target voice data with preset timbre is acquired. Response information is then output using the target voice data with preset timbre.

Benefits of technology

It improves the user experience, enables personalized voice responses, and overcomes the problem of poor user experience caused by a single way of outputting response information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115540260B_ABST
    Figure CN115540260B_ABST
Patent Text Reader

Abstract

The application discloses a voice output method and an air conditioner. The method is applied to the air conditioner, and the method comprises the following steps: obtaining a voice control command, wherein the voice control command comprises control information used for controlling the running state of a household appliance; determining a current scene of the air conditioner according to the voice control command; obtaining target voice data corresponding to the current scene, wherein the target voice data is voice data with a preset tone; and outputting response information of the control information by using the target voice data. The scheme provided by the application aims to solve the technical problem that the voice broadcast form of the air conditioner in the prior art is single.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electrical equipment, and more specifically, to a voice output method and an air conditioner. Background Technology

[0002] With the maturity of voice recognition technology, more and more voice-controlled air conditioners are entering users' homes. The differentiation in air conditioner voice recognition technology has shifted from performance differences to user experience. Voice technology needs to adopt a user-centric approach, proposing more human-centered new scenarios. Users' needs for voice-controlled air conditioners have evolved from simple control to a growing demand for more personalized features. Currently, the content and tone of voice-controlled air conditioners are relatively limited, resulting in a poor user experience. Summary of the Invention

[0003] The main objective of this invention is to provide a voice output method and an air conditioner, aiming to solve the technical problem of the limited voice broadcast format in existing air conditioners.

[0004] To achieve the above objectives, this invention proposes a voice output method applied to an air conditioner, the method comprising:

[0005] Acquire voice control commands, wherein the voice command control includes control information for controlling the operating status of home appliances;

[0006] The current scene of the air conditioner is determined based on the voice control command;

[0007] Acquire the target speech data corresponding to the current scene, wherein the target speech data is speech data with a preset timbre;

[0008] Using the target speech data, output the response information of the control information.

[0009] Preferably, the method further includes:

[0010] Obtain the amount of voice data corresponding to the current scene;

[0011] Determine whether the data volume meets the preset data volume judgment condition;

[0012] If the data volume does not meet the data volume judgment condition, then the acquisition of the target speech data corresponding to the current scene is not allowed.

[0013] Preferably, obtaining the target speech data corresponding to the current scene information includes:

[0014] Obtain the first voiceprint information of the voice control command;

[0015] Determine the second voiceprint information corresponding to the first voiceprint information;

[0016] From the voice data corresponding to the current scene, select the voice data containing the second voiceprint information as the target voice data.

[0017] Preferably, the step of using the target speech data to output the response information of the control information includes:

[0018] Obtain the first voiceprint information of the voice control command;

[0019] Based on the first voiceprint information and the control information, determine the response information corresponding to the control information;

[0020] The response information is output using the target speech data.

[0021] Preferably, before determining the target speech data corresponding to the current scene information, the method further includes:

[0022] Acquisition strategy for voice data in each scenario;

[0023] Voice data is collected for each scenario according to the voice data collection strategy for each scenario.

[0024] Preferably, the acquisition strategy includes at least one of the voiceprint information of the voice data corresponding to each scene and the start condition of the acquisition operation.

[0025] Preferably, after collecting voice data for each scenario according to the acquisition strategy corresponding to each scenario, the method further includes:

[0026] Perform the following operations for each scenario, including:

[0027] The collected voice data is parsed to obtain content information;

[0028] Determine whether the content information matches the scenario;

[0029] If the content information does not conform to the scenario, then the voice data is not allowed to be used as the voice data corresponding to the scenario.

[0030] A storage medium storing a computer program, wherein the computer program is configured to execute the method described above when run.

[0031] An electronic device includes a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the methods described above.

[0032] An air conditioner includes the electronic device described above.

[0033] In the technical solution of this invention, target speech data with preset timbre is used to output response information, which overcomes the problem of poor user experience caused by outputting response information in a single way and improves the user experience. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0035] Figure 1 A flowchart of the voice output method provided in Embodiment 1 of the present invention;

[0036] Figure 2 A flowchart of the voice output method provided in Embodiment 2 of the present invention;

[0037] Figure 3 A flowchart of the voice output method provided in Embodiment 3 of the present invention;

[0038] Figure 4 A flowchart of the voice data acquisition method provided in Embodiment 4 of the present invention;

[0039] Figure 5(a) is a flowchart of the voice data acquisition method provided in Embodiment 5 of the present invention;

[0040] Figure 5(b) is a flowchart of the voice data output method provided in Embodiment 5 of the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this utility model clearer, the embodiments of this utility model will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0043] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0044] Furthermore, in this invention, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0045] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0046] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0047] Example 1

[0048] Figure 1 This is a flowchart of a voice output method provided in Embodiment 1 of the present invention. Figure 1 As shown, this method is applied to an air conditioner, and the method includes:

[0049] Step 101: Obtain a voice control command. This voice control command includes control information, which is used to control the operating status of home appliances.

[0050] In one exemplary embodiment, the voice control command can be used to manage the operating parameters of the air conditioner, such as setting the working mode, adjusting the temperature value, and controlling the airflow direction; or, it can be used to invoke additional intelligent functions of the air conditioner, such as weather inquiries, story readings, and music playback.

[0051] Step 102: Determine the current scene corresponding to the air conditioner based on the voice control command;

[0052] In an exemplary embodiment, the control information in the voice control command is parsed to determine the control object to be controlled, and the current scene corresponding to the air conditioner is determined based on the control object.

[0053] For example, if the voice control command is "temperature is 25℃", based on semantic parsing, it can be determined that the function to be invoked by the voice control command is the temperature setting function, and the current scenario is determined to be the setting scenario of running parameters.

[0054] For example, if the voice control command is "What's the weather like today?", based on semantic parsing, it can be determined that the function to be invoked by the voice control command is the weather query function, and thus the current scenario is determined to be a weather query scenario.

[0055] Step 103: Obtain the target speech data with preset timbre corresponding to the current scene;

[0056] The target voice data is the voice data collected by the air conditioner in the preset space, and the preset timbre is a voice with emotion.

[0057] For example, when the scenario is reading a children's encyclopedia, the preset tone can be the voice data of a parent; when the scenario is a weather report, the preset tone can be the voice data of a couple; when an elder operates the air conditioner, the preset tone can be the dialect data of the children, so as to achieve the purpose of operation guidance.

[0058] The preset timbre can further include voiceprint information.

[0059] If the first voiceprint information of the voice control command belongs to a child in the family, then the target voice data with the preset timbre is the voice data of the mother in the family; if the first voiceprint information belongs to the voiceprint information of the male head of the household, then the target voice data with the preset timbre is the voice data of the female head of the household.

[0060] By selecting target voice information with preset timbres for the current scene, the user experience can be effectively improved, and the problem of monotonous single broadcast format can be solved.

[0061] Step 104: Based on the target voice data, output the response information of the control information;

[0062] In an exemplary embodiment, response information can be determined based on control information and first voiceprint information, and then the response information can be output based on target speech data.

[0063] For example, if the control message is "Adjust the temperature to 25℃", and the first voiceprint is from an elder in the family, the response message could be a preset dialect voice message played, indicating successful setup. If the first voiceprint is from a child in the family, the response message could be a warning message such as "Do not operate, there is danger", output in a parent's tone or voice.

[0064] As can be seen from the example above, when the control information is the same, the response information will be different if the voiceprint information is different. Therefore, determining the response information based on both the voiceprint information and the control information can ensure the accuracy of the response information.

[0065] Unlike existing technologies that output response information in a single way, the solution provided in this invention uses target speech data with preset timbre to output response information, which can improve the user experience.

[0066] The method provided in Embodiment 1 of the present invention obtains a voice control command, which includes control information. Based on the voice control command, the current scene corresponding to the air conditioner is determined, target voice data with a preset timbre corresponding to the current scene is obtained, and response information of the control information is output based on the target voice data. By using target voice data with a preset timbre to output the response information, the problem of low user experience caused by outputting response information in a single way is overcome, thereby improving the user experience.

[0067] Example 2

[0068] Figure 2 This is a flowchart of the voice output method provided in Embodiment 2 of the present invention. Figure 2 As shown, this method is applied to an air conditioner, and the method includes:

[0069] Step 201: Obtain a voice control command. The voice control command includes control information, which is used to control the operating status of the home appliance.

[0070] In one exemplary embodiment, the voice control command can be used to manage the operating parameters of the air conditioner, such as setting the working mode, adjusting the temperature value, and controlling the airflow direction; or, it can be used to invoke additional intelligent functions of the air conditioner, such as weather inquiries, story readings, and music playback.

[0071] Step 202: Determine the current scene corresponding to the air conditioner based on the voice control command;

[0072] In an exemplary embodiment, the control information in the voice control command is parsed to determine the control object to be controlled, and the current scene corresponding to the air conditioner is determined based on the control object.

[0073] For example, if the voice control command is "temperature is 25℃", based on semantic parsing, it can be determined that the function to be invoked by the voice control command is the temperature setting function, and the current scenario is determined to be the setting scenario of running parameters.

[0074] For example, if the voice control command is "What's the weather like today?", based on semantic parsing, it can be determined that the function to be invoked by the voice control command is the weather query function, and thus the current scenario is determined to be a weather query scenario.

[0075] Step 203: Obtain the amount of voice data in the current scene;

[0076] Step 204: Determine whether the data volume meets the preset data volume judgment conditions;

[0077] If the data volume meets the data volume judgment condition, it means that the voice data of the current scene stored locally is sufficient and can be used as the material required for response information output, then proceed to step 205.

[0078] If the data volume does not meet the data volume judgment condition, it means that the voice data of the current scene stored locally is insufficient and cannot be used as the material required for response information output. The voice data corresponding to the current scene cannot be used as the target voice data, and the process ends.

[0079] By determining whether the amount of voice data is sufficient, we can determine whether the voice data corresponding to the current scenario can be used for the output of response information, so as to ensure that the response information can be output completely and smoothly, and to ensure the broadcast effect of voice output.

[0080] Step 205: Determine the current scene's voice data as the target voice data;

[0081] The target voice data is the voice data collected by the air conditioner in the preset space, and it has a preset timbre, which is a voice with emotion.

[0082] For example, when the scenario is reading a children's encyclopedia, the preset tone can be the voice data of a parent; when the scenario is a weather report, the preset tone can be the voice data of a couple; when an elder operates the air conditioner, the preset tone can be the dialect data of the children, so as to achieve the purpose of operation guidance.

[0083] The preset timbre can further include voiceprint information.

[0084] If the first voiceprint information of the voice control command belongs to a child in the family, then the target voice data with the preset timbre is the voice data of the mother in the family; if the first voiceprint information belongs to the voiceprint information of the male head of the household, then the target voice data with the preset timbre is the voice data of the female head of the household.

[0085] By selecting target voice information with preset timbres for the current scene, the user experience can be effectively improved, and the problem of monotonous single broadcast format can be solved.

[0086] Step 206: Based on the target speech data, output the response information of the control information;

[0087] In an exemplary embodiment, response information can be determined based on control information and first voiceprint information, and then the response information can be output based on target speech data.

[0088] For example, if the control message is "Adjust the temperature to 25℃", and the first voiceprint is from an elder in the family, the response message could be a preset dialect voice message played, indicating successful setup. If the first voiceprint is from a child in the family, the response message could be a warning message such as "Do not operate, there is danger", output in a parent's tone or voice.

[0089] As can be seen from the example above, when the control information is the same, the response information will be different if the voiceprint information is different. Therefore, determining the response information based on both the voiceprint information and the control information can ensure the accuracy of the response information.

[0090] Unlike existing technologies that output response information in a single way, the solution provided in this invention uses target speech data with preset timbre to output response information, which can improve the user experience.

[0091] The method provided in Embodiment 2 of this invention obtains a voice control command, which includes control information. Based on the voice control command, the current scene corresponding to the air conditioner is determined, and target voice data with a preset timbre corresponding to the current scene is obtained. Based on the target voice data, response information for the control information is output. By using target voice data with a preset timbre to output the response information, the method overcomes the problem of poor user experience caused by outputting response information in a single way, thus improving the user experience. By judging whether the amount of voice data is sufficient, it is determined whether the voice data corresponding to the current scene can be used for the output of response information, so as to ensure that the response information can be output completely and smoothly, and to ensure the broadcast effect of the voice output.

[0092] Example 3

[0093] Figure 3 This is a flowchart of the voice output method provided in Embodiment 3 of the present invention. Figure 3 As shown, this method is applied to an air conditioner, and the method includes:

[0094] Step 301: Obtain a voice control command. The voice control command includes control information, which is used to control the operating status of the home appliance.

[0095] In one exemplary embodiment, the voice control command can be used to manage the operating parameters of the air conditioner, such as setting the working mode, adjusting the temperature value, and controlling the airflow direction; or, it can be used to invoke additional intelligent functions of the air conditioner, such as weather inquiries, story readings, and music playback.

[0096] Step 302: Determine the current scene corresponding to the air conditioner based on the voice control command;

[0097] In an exemplary embodiment, the control information in the voice control command is parsed to determine the control object to be controlled, and the current scene corresponding to the air conditioner is determined based on the control object.

[0098] For example, if the voice control command is "temperature is 25℃", based on semantic parsing, it can be determined that the function to be invoked by the voice control command is the temperature setting function, and the current scenario is determined to be the setting scenario of running parameters.

[0099] For example, if the voice control command is "What's the weather like today?", based on semantic parsing, it can be determined that the function to be invoked by the voice control command is the weather query function, and thus the current scenario is determined to be a weather query scenario.

[0100] Step 303: Obtain the amount of voice data for the current scene;

[0101] Step 304: Determine whether the data volume meets the preset data volume judgment conditions;

[0102] If the data volume meets the data volume judgment condition, it means that the voice data of the current scene stored locally is sufficient and can be used as the material required for response information output, then proceed to step 305.

[0103] If the data volume does not meet the data volume judgment condition, it means that the voice data of the current scene stored locally is insufficient and cannot be used as the material required for response information output. The voice data corresponding to the current scene cannot be used as the target voice data, and the process ends.

[0104] By determining whether the amount of voice data is sufficient, we can determine whether the voice data corresponding to the current scenario can be used for the output of response information, so as to ensure that the response information can be output completely and smoothly, and to ensure the broadcast effect of voice output.

[0105] Step 305: Obtain the first voiceprint information corresponding to the voice control command;

[0106] Step 306: Determine the second voiceprint information corresponding to the first voiceprint information;

[0107] A correspondence between voiceprint information can be established in advance, and the second voiceprint information can be determined through this correspondence.

[0108] For example, a household has three users: user A, user B, and user C. The corresponding voiceprint information can be established as follows, where the first voiceprint information is the first voiceprint information and the second voiceprint information is the second voiceprint information:

[0109] User A corresponds to User B, User B corresponds to User A, and User C corresponds to User A.

[0110] For example, if the first voiceprint information is the voiceprint information of a child in the family, then the target speech data can be the voiceprint information of the mother in the family; if the first voiceprint information is the voiceprint information of the male head of the household, then the second voiceprint information can be the voiceprint information of the female head of the household.

[0111] Step 307: Select the voice data corresponding to the second voiceprint information as the target voice data with a preset timbre from the voice data of the current scene.

[0112] The target voice data is the voice data collected by the air conditioner in a preset space, and the preset timbre is a voice with emotion.

[0113] For example, when the scenario is reading a children's encyclopedia, the preset tone can be the voice data of a parent; when the scenario is a weather report, the preset tone can be the voice data of a couple; when an elder operates the air conditioner, the preset tone can be the dialect data of the children, so as to achieve the purpose of operation guidance.

[0114] The preset timbre can further include voiceprint information.

[0115] If the first voiceprint information of the voice control command belongs to a child in the family, then the target voice data with the preset timbre is the voice data of the mother in the family; if the first voiceprint information belongs to the voiceprint information of the male head of the household, then the target voice data with the preset timbre is the voice data of the female head of the household.

[0116] By selecting target voice information with preset timbres for the current scene, the user experience can be effectively improved, and the problem of monotonous single broadcast format can be solved.

[0117] By performing identity recognition on voice control commands, the system identifies the target voice data that matches that identity, enabling personalized responses based on each user and improving user experience.

[0118] Step 308: Based on the target voice data, output the response information of the control information;

[0119] In an exemplary embodiment, response information can be determined based on control information and first voiceprint information, and then the response information can be output based on target speech data.

[0120] For example, if the control message is "Adjust the temperature to 25℃", and the first voiceprint is from an elder in the family, the response message could be a preset dialect voice message played, indicating successful setup. If the first voiceprint is from a child in the family, the response message could be a warning message such as "Do not operate, there is danger", output in a parent's tone or voice.

[0121] As can be seen from the example above, when the control information is the same, the response information will be different if the voiceprint information is different. Therefore, determining the response information based on both the voiceprint information and the control information can ensure the accuracy of the response information.

[0122] Unlike existing technologies that output response information in a single way, the solution provided in this invention uses target speech data with preset timbre to output response information, which can improve the user experience.

[0123] The method provided in Embodiment 3 of this invention obtains a voice control command, which includes control information. Based on the voice control command, the current scene corresponding to the air conditioner is determined, and target voice data with a preset timbre corresponding to the current scene is obtained. Based on the target voice data, response information for the control information is output. Using target voice data with a preset timbre for the output of response information overcomes the problem of poor user experience caused by a single method of outputting response information, thus improving the user experience. By judging whether the amount of voice data is sufficient, it is determined whether the voice data corresponding to the current scene can be used for the output of response information, ensuring that the response information can be output completely and smoothly, and guaranteeing the broadcast effect of the voice output. By performing identity recognition on the voice control command, target voice data matching the identity is determined, realizing personalized responses based on each user and improving the user experience.

[0124] In the above embodiments one to three, the air conditioner stores corresponding voice data for each scenario, and the stored voice data is all collected by the air conditioner, which does not require manual operation by the user and simplifies the user operation.

[0125] The following describes the voice data acquisition process:

[0126] Example 4

[0127] Figure 4This is a flowchart of a method for acquiring voice data provided in Embodiment 4 of the present invention. Figure 4 As shown, this method is applied to an air conditioner, and the method includes:

[0128] Step 401: Obtain and save the voice data acquisition strategy set for each scenario;

[0129] In one exemplary embodiment, a scene identifier for each scenario can be pre-set for the air conditioner, and a corresponding acquisition strategy can be set based on each scene identifier. The acquisition strategy includes the activation conditions for the acquisition operation and the voiceprint information corresponding to the acquired voice data. The activation conditions can be preset time information.

[0130] Step 402: Collect voice data for each scenario according to the voice data collection strategy for each scenario;

[0131] In one exemplary embodiment, the voice data of a scene can be collected after the conditions for initiating the scene's collection operation are met.

[0132] By controlling the initiation of acquisition operations based on activation conditions, targeted activation can be achieved, reducing the execution of unnecessary acquisition operations and improving the success rate of voice data acquisition.

[0133] In one exemplary embodiment, the voiceprint information of the received voice data is detected, and if the voiceprint information is the voiceprint information corresponding to the scene, the voice data of the voiceprint information is collected as the voice data of the scene.

[0134] By detecting voiceprint information, targeted voice data collection can be achieved, reducing the amount of data collected and improving the efficiency of the collection operation.

[0135] For example, when the scenario is story reading, you can set 9 pm every night as the start time for voice collection, and the voiceprint information collected will be the voiceprint information of the female homeowner.

[0136] For each scenario, the following operations are performed, including:

[0137] Step 403: Analyze the content of the voice data obtained through the acquisition operation to obtain content information;

[0138] Step 404: Determine whether the content information matches the scenario;

[0139] If the content matches the scenario, proceed to step 405;

[0140] If the content does not match the scenario, it means that the obtained voice data is irrelevant to the scenario, and the process ends.

[0141] Step 405: Save the collected voice data as the voice data for this scene.

[0142] The method provided in Embodiment 4 of the present invention sets a voice data acquisition strategy for each scenario and acquires voice data for each scenario according to the voice data acquisition strategy for each scenario, which can effectively improve the voice data acquisition efficiency and reduce manual operation by the user; in addition, by judging whether the acquired voice data matches the scenario, the accuracy of the acquired voice data is guaranteed and the user experience is improved.

[0143] Example 5

[0144] The method provided in Embodiment 5 of the present invention is applied to an air conditioner and includes two processes: a voice data acquisition process and a voice data output process.

[0145] Figure 5(a) is a flowchart of the speech data acquisition method provided in Embodiment 5 of the present invention. As shown in Figure 5(a), this method is an emotion-based TTS (Text To Speech) learning process, including the following steps:

[0146] (1) Pre-set the scene of telling stories to the daughter from 21:00 to 22:00 every night; after the voice collection start conditions of the scene are met at the current time, the voice module enters scene recognition;

[0147] (2) The voice module starts scanning and listening, and collects voiceprint information once every time T1.

[0148] (3) If the voiceprint identity is the set owner, proceed to step 5; otherwise, proceed to step 4.

[0149] (4) Has the scene end time been reached? If yes, exit the scene; otherwise, return to step 2.

[0150] (5) Start sound acquisition, acquisition time is S seconds, interval T2;

[0151] (6) The voice module uploads the collected audio to the cloud;

[0152] (7) Analyze audio content in the cloud and perform scene content analysis;

[0153] (8) If the audio content matches the set scene content, the collected audio data will be stored in the material library;

[0154] (9) Determine whether the quantity of materials meets the training requirement N. If it does, then conduct sound training. If the quantity is insufficient, continue to collect materials.

[0155] (10) After the voice model training is completed, save it to the emotional TTS database.

[0156] After the above steps, user emotional TTS data already exists in the database without any manual operation by the user.

[0157] Figure 5(b) is a flowchart of the voice data output method provided in Embodiment 5 of the present invention. As shown in Figure 5(b), when a user interacts with the air conditioner, the air conditioner enters the emotional TTS broadcast process, and the method flow is as follows:

[0158] a) Voiceprint identity locking, assuming the current user's identity is that of the daughter;

[0159] b) Dialogue scene recognition, including but not limited to air conditioning control, storytelling, children's encyclopedia, warning prompts, etc.

[0160] c) After obtaining the identity information and current scene information, search the database for matching emotional TTS broadcast content. If a match is found, perform emotional TTS broadcast; otherwise, perform regular TTS broadcast.

[0161] After setting the desired learning scenarios, the voice-activated air conditioner automatically completes the voice replication process without user intervention. Through scenario and identity matching, it performs emotional TTS broadcasts, solving the problem of monotonous broadcasts in current air conditioners and improving the user experience.

[0162] This invention provides a storage medium storing a computer program that, when run, can execute any of the methods described above.

[0163] This invention provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor, when running the computer program in the memory, can execute any of the methods described above.

[0164] This invention provides an air conditioner for implementing the electronic device described above.

[0165] The aforementioned electronic device can be installed as a separate module in the air conditioner, or integrated into the air conditioner's processor; it is used to enable voice output operation.

[0166] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made under the concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A voice output method, characterized in that, Applied to air conditioners, the method includes: Acquire voice control commands, wherein the voice control commands include control information for controlling the operating status of home appliances; The current scene of the air conditioner is determined based on the voice control command; Obtain the amount of voice data corresponding to the current scene; Determine whether the data volume meets the preset data volume judgment conditions; If the data volume meets the data volume judgment condition, then the target voice data corresponding to the current scene is obtained, wherein the target voice data is voice data with a preset timbre, and the preset timbre is the voice of another family member that matches the family member corresponding to the voice control command; Using the target speech data, output the response information of the control information.

2. The method according to claim 1, characterized in that, The method further includes: If the data volume does not meet the data volume judgment condition, then the acquisition of the target speech data corresponding to the current scene is not allowed.

3. The method according to claim 1, characterized in that, The acquisition of the target speech data corresponding to the current scene includes: Obtain the first voiceprint information of the voice control command; Determine the second voiceprint information corresponding to the first voiceprint information; From the voice data corresponding to the current scene, select the voice data containing the second voiceprint information as the target voice data.

4. The method according to any one of claims 1 to 3, characterized in that, The step of outputting response information for the control information using the target speech data includes: Obtain the first voiceprint information of the voice control command; Based on the first voiceprint information and the control information, determine the response information corresponding to the control information; The response information is output using the target speech data.

5. The method according to claim 1, characterized in that, Before acquiring the target speech data corresponding to the current scene, the method further includes: Acquisition strategy for voice data in each scenario; Voice data is collected for each scenario according to the voice data collection strategy for each scenario.

6. The method according to claim 5, characterized in that, The acquisition strategy includes at least one of the following: voiceprint information of the voice data corresponding to each scene and the start condition for the acquisition operation.

7. The method according to claim 5, characterized in that, After collecting voice data for each scenario according to the acquisition strategy corresponding to each scenario, the method further includes: Perform the following operations for each scenario, including: The collected voice data is parsed to obtain content information; Determine whether the content information matches the scenario; If the content information does not conform to the scenario, then the voice data is not allowed to be used as the voice data corresponding to the scenario.

8. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method described in any one of claims 1 to 7 when it is run.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method as described in any one of claims 1 to 7.

10. An air conditioner, characterized in that, Includes the electronic device as described in claim 9.

Citation Information

Patent Citations

  • Audio broadcast method, device, compute device and storage medium

    CN109273001A

  • Voice control method and computer storage medium

    CN110164426A

  • Voice processing method and device, electronic equipment and storage medium

    CN110930986A

  • Voice interaction method and device, computer readable storage medium and processor

    CN112185344A