Voice-based interaction method and apparatus, intelligent device, and computer-readable storage medium

By receiving home scene sound data from smart devices, determining the scene type, generating predictive acoustic features, and synthesizing Lombard speech, the problem of low voice broadcast adaptability of smart devices in noisy environments is solved, achieving clearer and more natural voice interaction.

CN114863909BActive Publication Date: 2026-02-13MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110071583.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-19
Publication Date
2026-02-13
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

Existing smart devices have poor adaptability to voice broadcasting in noisy acoustic environments, which makes it difficult for users to receive voice information clearly and reliably, affecting the smoothness of human-computer interaction.

Method used

Smart devices receive sound data from the home space, determine the type of home scene, generate predictive acoustic features based on the speech prediction strategy of that type, synthesize and output response speech data, mimicking how humans actively change their vocalization under the Lombard effect, and synthesize Lombard speech with the corresponding acoustic style of the scene.

Benefits of technology

It improves the clarity and naturalness of voice broadcasts from smart devices in noisy acoustic environments, ensuring that users can reliably receive voice information and enhancing the smoothness of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863909B_ABST
    Figure CN114863909B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a voice-based interaction method and device, smart device and computer readable storage medium, relating to the technical field of intelligent interaction, the method comprising: receiving sound data of a home space, obtaining a home scene type corresponding to the sound data, when an interaction instruction is obtained, generating predicted acoustic characteristics through a voice prediction strategy corresponding to the home scene type, synthesizing the predicted acoustic characteristics, and obtaining output response voice data corresponding to the interaction instruction, thereby improving the adaptability of the output response voice data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent interaction, in particular to a voice-based interaction method and device, intelligent device and computer readable storage medium. BACKGROUND

[0002] In the existing smart home scene, the intelligent device needs to perform voice synthesis and broadcast the synthesized voice during voice interaction with the user. Research has found that the adaptability of the voice broadcast by the existing intelligent device needs to be improved. SUMMARY

[0003] French ear, nose and throat doctor Etienne Lombard discovered in 1909 that when communicating in a noisy environment, the speaker has to actively change the way of making sound and increase the effect of the sound in order to make the other party hear clearly. Research has found that even the same person speaking the same voice has different voice characteristics in different environments, and the changed characteristics include increasing the pitch, tone, loudness and formant characteristics of the sound. This phenomenon is called Lombard effect. In view of this, the inventors have researched how to improve the intelligibility of the voice broadcast by the intelligent device in a noisy acoustic environment, and have proposed an intelligent device that "imitates" (i.e., applies) the changes in the way of making sound by the human being under Lombard effect, so that in the case of a noisy home scene type, the synthesized voice data with the corresponding scene acoustic style is broadcast, and through the synthesis of voice with better recognition, naturalness and intelligibility, which we call Lombard speech, to ensure the smoothness of voice interaction with the user in a noisy home environment.

[0004] One of the purposes of the present application includes, for example, to provide a voice-based interaction method and device, intelligent device and computer readable storage medium to at least partially improve the adaptability of the output response voice data.

[0005] Embodiments of the present application can be implemented as follows:

[0006] In a first aspect, embodiments of the present application provide a voice-based interaction method applied to an intelligent device, the method comprising:

[0007] receiving sound data of a home space;

[0008] obtaining a home scene type corresponding to the sound data;

[0009] when an interaction instruction is obtained, generating a predicted acoustic feature through a voice prediction strategy corresponding to the home scene type;

[0010] synthesize the predicted acoustic features to obtain output response voice data corresponding to the interaction indication.

[0011] By obtaining the home scene type, generating predicted acoustic features based on a voice prediction strategy corresponding to the home scene type, and synthesizing output response voice data, the adaptability of the output response voice data is ensured.

[0012] In a second aspect, an embodiment of the present application provides a voice-based interaction method, comprising:

[0013] inputting the response text content information into a Lombard voice generation model to obtain output response voice data corresponding to the home scene, the Lombard voice generation model being obtained based on Lombard voice learning.

[0014] In a third aspect, an embodiment of the present application provides a voice-based interaction device applied to a smart device, comprising:

[0015] The information determination module is configured to receive sound data of a home space, obtain a home scene type corresponding to the sound data, and generate predicted acoustic features based on a voice prediction strategy corresponding to the home scene type when an interaction indication is obtained.

[0016] The response voice data synthesis module is configured to synthesize the predicted acoustic features to obtain output response voice data.

[0017] By obtaining the home scene type, generating predicted acoustic features based on a voice prediction strategy corresponding to the home scene type, and synthesizing output response voice data, the adaptability of the output response voice data is ensured.

[0018] In a fourth aspect, an embodiment of the present application provides a smart device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the voice-based interaction method of any one of the preceding embodiments when executing the program. Accordingly, the smart device includes the beneficial effects of the voice-based interaction method.

[0019] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium comprising a computer program, wherein the computer program controls a smart device where the computer-readable storage medium is located to execute the voice-based interaction method of any one of the preceding embodiments when running. Accordingly, the computer-readable storage medium includes the beneficial effects of the voice-based interaction method.

[0020] To make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the following describes a preferred embodiment in detail, and the accompanying drawings are referred to as follows. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as a limitation to the scope, and other related drawings can also be obtained by those of ordinary skill in the art without any creative effort.

[0022] Figure 1 An interaction architecture of an intelligent device provided by the embodiments of the present application is shown.

[0023] Figure 2 An application scenario diagram provided by the embodiments of the present application is shown.

[0024] Figure 3 A flow diagram of a voice-based interaction method provided by the embodiments of the present application is shown.

[0025] Figure 4 A quiet type of scenario diagram provided by the embodiments of the present application is shown.

[0026] Figure 5 A noisy type of scenario diagram provided by the embodiments of the present application is shown.

[0027] Figure 6 A sound source type of diagram provided by the embodiments of the present application is shown.

[0028] Figure 7 A scenario diagram of configuring a voice prediction strategy provided by the embodiments of the present application is shown.

[0029] Figure 8 An interaction interface diagram of configuring a voice prediction strategy provided by the embodiments of the present application is shown.

[0030] Figure 9 An implementation principle diagram of synthesizing output response voice data provided by the embodiments of the present application is shown.

[0031] Figure 10 Another implementation principle diagram of synthesizing output response voice data provided by the embodiments of the present application is shown.

[0032] Figure 11 A flow diagram of a scenario acoustic style extractor training method provided by the embodiments of the present application is shown.

[0033] Figure 12 A training architecture diagram of a scenario acoustic style extractor provided by the embodiments of the present application is shown.

[0034] Figure 13 Fig. 6 shows another flowchart of a scene acoustic style extractor training method according to an embodiment of the present application.

[0035] Figure 14 Fig. 7 shows another training architecture diagram of a scene acoustic style extractor according to an embodiment of the present application.

[0036] Figure 15 Fig. 8 shows another flowchart of a scene acoustic style extractor training method according to an embodiment of the present application.

[0037] Figure 16 Fig. 9 shows another training architecture diagram of a scene acoustic style extractor according to an embodiment of the present application.

[0038] Figure 17 Fig. 10 shows another flowchart of a scene acoustic style extractor training method according to an embodiment of the present application.

[0039] Figure 18 Fig. 11 shows a schematic diagram of a reference encoder according to an embodiment of the present application.

[0040] Figure 19 Fig. 12 shows a schematic diagram of a reference encoder according to an embodiment of the present application.

[0041] Figure 20 Fig. 13 shows a schematic diagram of a reference encoder according to an embodiment of the present application.

[0042] Figure 21 Fig. 14 shows another training architecture diagram of a scene acoustic style extractor according to an embodiment of the present application.

[0043] Figure 22 Fig. 15 shows a flowchart of a first scene classification model training method according to an embodiment of the present application.

[0044] Figure 23 Fig. 16 shows another flowchart of a first scene classification model training method according to an embodiment of the present application.

[0045] Figure 24 Fig. 17 shows a flowchart of an acoustic feature prediction model training method according to an embodiment of the present application.

[0046] Figure 25 Fig. 18 shows a training architecture diagram of an acoustic feature prediction model according to an embodiment of the present application.

[0047] Figure 26 Fig. 19 shows an implementation principle diagram of a scene acoustic style extractor according to an embodiment of the present application.

[0048] Figure 27 Another implementation principle schematic diagram of a scene acoustic style extractor provided by an embodiment of the present application is shown.

[0049] Figure 28 An implementation principle schematic diagram of an acoustic feature prediction model provided by an embodiment of the present application is shown.

[0050] Figure 29 An implementation principle schematic diagram of synthesized output response speech data provided by an embodiment of the present application is shown.

[0051] Figure 30 Another implementation principle schematic diagram of synthesized output response speech data provided by an embodiment of the present application is shown.

[0052] Figure 31 A flow schematic diagram of a speech synthesis model training method provided by an embodiment of the present application is shown.

[0053] Figure 32 A training architecture schematic diagram of a speech synthesis model provided by an embodiment of the present application is shown.

[0054] Figure 33 An exemplary structural block diagram of a speech-based interaction device provided by an embodiment of the present application is shown.

[0055] Figure 34 An exemplary structural block diagram of a scene classification model training device provided by an embodiment of the present application is shown.

[0056] Figure 35 An exemplary structural block diagram of a scene acoustic style extractor training device provided by an embodiment of the present application is shown.

[0057] Figure 36 An exemplary structural block diagram of an acoustic feature prediction model training device provided by an embodiment of the present application is shown.

[0058] Figure 37 An exemplary structural block diagram of a speech synthesis model training device provided by an embodiment of the present application is shown.

[0059] Icon: 100 - smart device; 101 - sound emitting unit; 102 - microphone; 103 - processor; 104 - storage module; 105 - communication module; 106 - I / O interface; 110 - TV set; 121 - stove; 122 - microwave oven; 130 - refrigerator; 140 - voice-based interaction device; 141 - information determining module; 142 - response voice data synthesizing module; 160 - scene classification model training device; 161 - ambient acoustic feature obtaining module; 162 - scene classification network training module; 170 - scene acoustic style extractor training device; 171 - noisy ambient acoustic feature obtaining module; 172 - scene acoustic style extractor training module; 180 - acoustic feature prediction model training device; 181 - data obtaining module; 182 - acoustic feature prediction model training module; 190 - voice synthesis model training device; 191 - information obtaining module; 192 - to-be-output voice data synthesizing module; 200 - background server system; 300 - user terminal. DETAILED DESCRIPTION

[0060] In order to make the objects, technical solutions, and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings for the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations.

[0061] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.

[0062] It should be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article, or equipment.

[0063] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it will not be repeatedly defined and explained in subsequent drawings without special instructions.

[0064] It should be noted that the features in the embodiments of the present application can be combined with each other without conflict.

[0065] Please refer toFigure 1 This is a schematic diagram of the architecture of a smart device 100 provided in this embodiment. Figure 1 As shown, the smart device 100 can be any type of electronic device used to respond to detected sound data and thus realize interaction. The smart device 100 can detect sound data containing specific content, such as trigger words or sounds with specific audio characteristics, and then identify the interactive content within the sound data, such as function enable / disable commands, query commands, configuration commands, etc., and then respond to the interactive content to perform one or more operations. The smart device 100 can include, but is not limited to: computers, laptops, smartphones, tablets, televisions, hotspot access devices, smart wearable devices, smart furniture, smart home devices, or smart in-vehicle devices, etc.

[0066] In some possible implementations, the main function of the smart device 100 is to perform corresponding functions through the input / output of voice data. Optionally, the smart device 100 may also have other I / O functions, such as a keyboard / mouse, touch screen, buttons, etc. The I / O functions can be externally connected to the smart device 100 through the I / O interface, or they can be integrated into the smart device 100.

[0067] See Figure 1 The smart device 100 includes: one or more sound-emitting units 101, one or more microphones 102, one or more processors 103, a storage module 104, a communication module 105, and an I / O interface 106. Obviously, the smart device 100 may also include other components to achieve the corresponding functions, or reduce the number of components. For example, the smart device 100 may also include environmental sensors, such as temperature and humidity sensors. In another possible implementation, the smart device 100 may not include the I / O interface 106. Furthermore, the... Figure 1 The smart device 100 shown is only an example, and there are no restrictions on the combination, number, and connection method of the various components.

[0068] One or more processors 103 may include processing devices for control operations, interactive functions, and communication between various devices. Depending on different processing requirements, the one or more processors 103 may include: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a microprocessor, a signal processor, or any other possible type of processor or combination thereof. Furthermore, each processor may have its own local memory used to store programs, threads, or the operating system. Furthermore, the one or more processors 103 may be used solely to run the operating system, applications, or other background programs / tasks.

[0069] The storage module 104 can include one or more storage media for storing data, such as non-volatile memory or volatile memory, including but not limited to: a redundant array of independent disks (RAID) storage module, an electrically erasable programmable read-only memory (EEPROM), a flash memory, a solid state disk, a compact disc read-only memory (CD-ROM), a digital video disc (DVD), a magnetic disk, and the like. The storage module 104 can be accessed by one or more processors 103 to execute programs, data, and the like stored in the storage module 104 to implement corresponding operational functions.

[0070] Further, the storage module 104 can include a voice prediction strategy related to the voice-based interaction scheme in the embodiments of the present application. The voice prediction strategy can be various, for example, the voice prediction strategy can include user pre-defined content, a pre-set model structure, such as one or more models capable of acoustic feature prediction. For another example, the voice prediction strategy can also only include output results of one or more models capable of acoustic feature prediction.

[0071] Each of the one or more sound generating units 101 can include a speaker unit, a transducer unit, and the like, and the plurality of sound generating units can constitute a sound generating array, or can be configured as a sound generating group corresponding to different sound channels based on sound channel requirements. Of course, in some scenarios, when the smart device 100 is connected to an external audio reproduction device in a wired or wireless manner through an I / O interface or a communication module, the external audio reproduction device can be used to output sound. The external audio reproduction device can be a wired earphone, a wired earphone, a wireless earphone, a wireless earphone, a sound system with an external connection function, or other sound output devices with an external connection function, and the like.

[0072] The one or more microphones 102 are used to detect sound signals in the external environment to obtain corresponding sound data. The microphone 102 can be divided into one or more microphone arrays. Different microphones can be configured to detect sound signals of different frequency levels. Moreover, the one or more microphones 102 can be configured to work or partially work based on different functions and power consumption requirements. Obviously, each microphone can also be configured to listen to sound in different spatial relationships, such as being configured to capture distant sound, or being configured to capture close-range sound, and the like.

[0073] The communication module 105 can have any circuitry for communicating with other devices in wired or wireless form. The communication module 105 can use any communication protocol, such as Transmission Control Protocol / Internet Protocol (TCP / IP), Wi-Fi, Near Field Communication (NFC), Bluetooth, cellular network communication, High Definition Multimedia Interface (HDMI), etc. The smart device 100 can use the communication module 105 to network with a network access device to form a communication group with other devices, such as the communication module 105 establishing communication with a smart router to network with other smart home devices, so that the smart device 100 transmits data, operates interaction, updates operating system / firmware / policy, etc. through the networking. For another example, the smart device 100 can communicate with the background server system 200 through the communication module 105 to implement functions such as cloud storage, cloud computing, etc.

[0074] The I / O interface 106 can include, but is not limited to, a camera module, a keyboard / mouse, a handle, a touch screen, a Light Emitting Diode (LED), etc. possible interactive mechanisms, which can be used to realize different forms of interaction with the user. Of course, the interactive mechanism connected with the I / O interface 106 can be an external device, or its function can be integrated into the smart device 100, which is not limited here.

[0075] It should be understood that, Figure 1 The structure shown is only a structural schematic diagram of the smart device 100, and the smart device 100 can further include more or less components than those shown in the Figure 1 or have a different configuration from that shown in the Figure 1 . Figure 1 The components shown in the can be realized by hardware, software or a combination thereof.

[0076] Please refer to Figure 2At present, the intelligent device 100 in the home scene can perceive the change of the acoustic environment state such as receiving user interaction voice data, and reasonably answers the user's question through output response voice data. For example, the intelligent device 100 monitors the user to issue the instruction information "turn on the microwave oven", controls the microwave oven to be turned on and voice broadcasts "the microwave oven has been turned on", and completes the response to the instruction information. It is found through research that the adaptation degree of the voice broadcast by the intelligent device is poor at present. For example, at present, the research is basically on the voice synthesis technology of the home scene in the quiet type. However, the application scene is more of the noisy type, and in the existing technology, the same voice data is output regardless of the home scene, which may affect the interaction reliability.

[0077] Exemplarily, there are often various acoustic events in the home scene, such as water flow, extractor hood, blender, cooking, talking and the like in the kitchen, water flow, electric toothbrush, shaver, hair dryer and the like in the bathroom, and television, vacuum cleaner, talking and the like in the living room. The intelligent device will broadcast the voice synthesized based on the quiet scene in the noisy acoustic environment, so that the voice broadcast by the intelligent device is affected by the environment sound, and the user may not clearly and reliably obtain the content of the voice broadcast.

[0078] Therefore, the inventors have researched how to improve the adaptation degree of the voice broadcast by the intelligent device to the home scene, and further put forward an interaction method based on voice. The intelligent device obtains the home scene type, generates predicted acoustic features based on the voice prediction strategy corresponding to the home scene type when obtaining the interaction instruction, and synthesizes and outputs the response voice data, so as to ensure the adaptation degree of the output response voice data to the home scene, and further ensure the reliability of the voice interaction.

[0079] Please refer to Figure 3 The flowchart of the interaction method based on voice provided by the embodiment of the present application can be executed by Figure 1 The intelligent device 100, for example, the processor 103 in the intelligent device 100. The interaction method based on voice includes S110, S120, S130 and S140.

[0080] S110, receiving sound data of a home space.

[0081] S120, obtaining a home scene type corresponding to the sound data.

[0082] S130, when obtaining an interaction instruction, generating predicted acoustic features through a voice prediction strategy corresponding to the home scene type.

[0083] S140, synthesizing the predicted acoustic features to obtain output response voice data corresponding to the interaction instruction.

[0084] The home scene type can be defined flexibly. For example, the home scene type can include a quiet type and a noisy type. The smart device can determine the home scene type as the quiet type or the noisy type in multiple ways based on the sound data of the home space where the smart device is located. For example, the smart device can determine the home scene type corresponding to the sound data based on the decibel of the sound data, the sound source type, and the like.

[0085] In an implementation manner, the smart device can determine whether the volume decibel of the sound data of the home space where the smart device is located exceeds a set volume threshold. In a case where the volume decibel of the sound data does not exceed the set volume threshold, the smart device determines that the home scene type corresponding to the sound data is the quiet type. In a case where the volume decibel of the sound data exceeds the set volume threshold, the smart device determines that the home scene type corresponding to the sound data is the noisy type. For example, referring to Figure 4 , the smart device 100 is located in a living room, and the main sound source in the living room comes from the television 110. In a case where the television 110 is not turned on, the smart device 100 analyzes the received sound data and determines that the volume decibel of the sound data does not exceed the set volume threshold. Then, the smart device 100 determines that the home scene type corresponding to the sound data is the quiet type. For another example, referring to Figure 5 , in a case where the television 110 is turned on, there is television noise in the living room. The smart device 100 analyzes the received sound data and determines that the volume decibel of the sound data exceeds the set volume threshold. Then, the smart device 100 determines that the home scene type corresponding to the sound data is the noisy type.

[0086] In another implementation manner, the smart device can determine the home scene type according to the sound source type of the home space where the smart device is located. The sound source type represents the acoustic characteristics of the sound source contained in the home space where the smart device is located. The sound source types of the smart device in different home spaces are different, and the home scene types corresponding to different sound source types or combinations thereof can be different. For example, referring to Figure 6, the smart device 100 is in a kitchen, the sound sources in the kitchen include a stove 121, a microwave oven 122, and a refrigerator 130, the stove 121 corresponds to a sound source type 1, the microwave oven 122 corresponds to a sound source type 2, and the refrigerator 130 corresponds to a sound source type 3. The smart device 100 can obtain the sound source types in the kitchen where the smart device 100 is located, i.e., the sound source type 1, the sound source type 2, and the sound source type 3, when the smart device 100 is in an initial state. The smart device 100 can further determine the home scene type according to the sound source types. For example, the smart device 100 can determine, according to a user configuration, that the home scene type corresponding to one or a combination of at least two of the sound source type 1, the sound source type 2, and the sound source type 3 is a quiet type or a noisy type. For another example, the smart device 100 can analyze the sound data of one or a combination of at least two of the sound source type 1, the sound source type 2, and the sound source type 3, respectively, and automatically determine, according to the analysis result, that the home scene type corresponding to one or a combination of at least two of the sound source type 1, the sound source type 2, and the sound source type 3 is a quiet type or a noisy type. For example, the smart device 100 can determine, according to the sound source type 1, the sound source type 2, and the sound source type 3, that the home scene type corresponding to the sound source type 1 is a noisy type, the home scene type corresponding to the sound source type 2 is a quiet type, and the home scene type corresponding to a combination of the sound source type 2 and the sound source type 3 is a noisy type.

[0087] To improve the adaptation degree of the output response voice data to the home scene type and ensure that the output response voice data broadcast by the smart device in the noisy type can be reliably received by the user, the inventors have studied the voice interaction of the user in the noisy environment. As mentioned above, it is found through research that when communicating in a noisy environment, the speaker has to actively change the way of making sound and improve the effect of the sound in order to make the other party hear clearly. This phenomenon is called Lombard effect. With the rise of smart home, how to improve the smoothness of human-computer interaction and achieve the purpose of better serving humans has become a problem concerned in the field. However, it is found through research that the content of voice broadcast by the smart device cannot be accurately received by the user many times, such as the broadcast voice in the noisy environment is drowned by the environmental sound, so that the user cannot clearly receive the broadcast voice, which affects the smoothness of human-computer interaction in the smart home scene and cannot meet the actual application requirements.

[0088] Based on the above findings, the inventors concluded that the premise for the user to accurately receive the content of the voice broadcast such as "the microwave oven has been turned on" is that the smart device can change the way of making sound in different acoustic environments and actively improve the intelligibility and naturalness of the synthesized speech. However, in existing research, how the smart device actively changes the way of making sound like humans under the Lombard effect to improve speech intelligibility and let the user receive accurate information has not been fully considered, resulting in that the voice broadcast feedback by the smart device may not be received by the user.

[0089] Therefore, the inventors have researched how to improve the intelligibility of the voice broadcast by the smart device in a noisy acoustic environment, and further proposed a communication mode in which the smart device "imitates" the way of actively changing the way of making sound by humans under the Lombard effect, synthesizes voice data with the corresponding scene acoustic style to broadcast when the home scene type is a noisy type, and ensures the fluency of voice interaction with the user in a noisy home environment by synthesizing Lombard speech with good naturalness and intelligibility.

[0090] Based on the above research, the embodiment designs a voice prediction strategy for different home scene types. When the home scene type is a noisy type, the voice prediction strategy is to determine the scene style embedding information corresponding to the noisy type, and determine the second predicted acoustic feature corresponding to the noisy type according to the response text content information and the scene style embedding information. The scene style embedding information represents the scene acoustic style corresponding to the noisy home scene. When the home scene type is a quiet type, the voice prediction strategy is to determine the first predicted acoustic feature corresponding to the quiet type according to the response text content information.

[0091] When the home scene type is a noisy type, the scene style embedding information is introduced to determine the second predicted acoustic feature corresponding to the noisy type, so that the second predicted acoustic feature includes the acoustic style of the noisy type, so that the synthesized output response voice data is Lombard speech with Lombard effect, and the smart device outputs the output response voice data in a noisy home scene to be clearly delivered to the user, ensuring the fluency of interaction.

[0092] The above description of the home scene type and the voice prediction strategy is only an example, and the home scene type and the voice prediction strategy can also be other. For example, in order to further improve the adaptability of the synthesized output response voice data and improve the interaction experience, the noisy home scene can also be further classified in more detail.

[0093] In an implementation manner, multiple volume ranges can be divided, and the noisy type can be subdivided according to different volume ranges. Correspondingly, the voice prediction strategy is configured with scene style embedding information corresponding to different volume ranges. For example, if four volume ranges are divided, namely, volume range A1, volume range A2, volume range A3 and volume range A4, wherein the volume range A1 corresponds to the quiet type, the volume range A2 corresponds to the noisy type one, the volume range A3 corresponds to the noisy type two, and the volume range A4 corresponds to the noisy type three. Correspondingly, when the home scene type is the quiet type, the voice prediction strategy can be to determine the first predicted acoustic feature corresponding to the quiet type according to the response text content information. When the home scene type is the noisy type one, the noisy type two or the noisy type three, the voice prediction strategy can be to determine the scene style embedding information corresponding to the noisy type one, the noisy type two or the noisy type three, and to determine the second predicted acoustic feature corresponding to the corresponding noisy type according to the response text content information and the corresponding scene style embedding information. The scene style embedding information corresponding to the noisy type one, the noisy type two and the noisy type three is different.

[0094] In another implementation manner, the noisy type can be subdivided in combination with the home space where the intelligent device is located. For example, the noisy type can be subdivided by subdividing the home scene. For example, if the home scene is divided into the kitchen, the living room and the washroom, then the kitchen, the living room and the washroom where the noisy sound occurs can correspond to different noisy types, such as the kitchen where the noisy sound occurs corresponding to the noisy type four, the living room where the noisy sound occurs corresponding to the noisy type five, and the washroom where the noisy sound occurs corresponding to the noisy type six. When the home scene type is the noisy type four, the noisy type five or the noisy type six, the voice prediction strategy can be to determine the scene style embedding information corresponding to the noisy type four, the noisy type five or the noisy type six, and to determine the second predicted acoustic feature corresponding to the corresponding noisy type according to the response text content information and the corresponding scene style embedding information. The scene style embedding information corresponding to the noisy type four, the noisy type five and the noisy type six is different. Similarly, the kitchen, the living room and the washroom can also correspond to different quiet types, and the voice prediction strategy for different quiet types can refer to the voice synthesis scheme in the quiet environment in the prior art, or can be set by the user, which is not limited in the embodiment.

[0095] By subdividing the home scene type and the voice prediction strategy, the adaptation degree of the output response voice data to each home scene type can be further improved.

[0096] In the embodiment, the voice prediction strategy can be configured by the user or automatically recognized by the intelligent device. For example, please refer to Figure 7The user can use the user terminal 300 to be communicatively connected with the smart device 100, and configure the voice prediction strategy based on the user terminal 300.

[0097] Please refer to Figure 8 In an implementation manner, a configuration main interface can be displayed on the user terminal 300, and the configuration main interface includes device names, states, a home scene selection interface, and a networking interface corresponding to each home scene. The home scene selection interface includes divided home scenes, such as a living room, a dining room, a kitchen, a bedroom, and a bathroom. The networking interface includes networking information of the home scene, such as divided groups and included home devices. The user performs a selection operation in the home scene selection interface to determine a current home scene, and selects home devices in the networking interface. After the user completes the selection operation, the user terminal 300 jumps to a voice prediction strategy determination interface, and the current home scene and the voice prediction strategy are displayed on the voice prediction strategy determination interface. Exemplarily, the voice prediction strategy displayed on the voice prediction strategy determination interface can include: voice prediction strategy A: the home scene type is a quiet type, and the scene acoustic style is gentle; voice prediction strategy B: the home scene type is a noisy type, and the scene acoustic style is standard clear; and voice prediction strategy C: the home scene type is automatically identified, and the scene acoustic style is automatically identified. The user can configure the smart device to use voice prediction strategy A, voice prediction strategy B, or voice prediction strategy C for voice synthesis through a selection operation. Correspondingly, when the smart device obtains an interaction instruction, the smart device generates a predicted acoustic feature through the voice prediction strategy selected by the user, and then synthesizes output response voice data. For example, if the user selects voice prediction strategy B, the smart device receives the voice prediction strategy B configured by the user, the voice prediction strategy B includes scene style embedding information corresponding to the noisy type as standard clear, and the smart device further determines a second predicted acoustic feature corresponding to the noisy type according to response text content information and the scene style embedding information. For another example, if the user selects voice prediction strategy C, the smart device automatically identifies the home scene type and generates scene acoustic style information, and determines a predicted acoustic feature corresponding to the identified home scene type according to response text content information and the generated scene style embedding information.

[0098] In the above, the scene acoustic style of gentle and the scene acoustic style of standard clear can be any one or a combination of more than two of the following attributes: tone, sound intensity, speech rate, and emotion. The attributes corresponding to different scene acoustic styles are different.

[0099] In the embodiment of the present application, the interaction indication can be various, for example, the interaction indication can be user interaction demand content information, such as user interaction demand content information issued by the user through voice. Correspondingly, if the smart device analyzes the received sound data and obtains that the sound data contains user interaction demand content information, it is determined that the interaction indication is obtained, and the content information corresponding to the user interaction demand content information is taken as the response text content information. For another example, the interaction indication can be a pre-defined indication set by the user, such as the user sets to automatically ask whether to start cooking at 6 pm, and the smart device determines that the interaction indication is obtained at 6 pm, and takes the pre-defined content information such as whether to start cooking as the response text content information.

[0100] In view of the fact that there can be a certain time interval from the time when the smart device receives the sound data, obtains the home scene type to the time when the interaction indication is obtained, in order to ensure the adaptation degree of the obtained output response voice data to the current home scene type, when the interaction indication is obtained, it can also be determined whether the current home scene type has changed, such as by re-obtaining the sound data and analyzing to obtain the current home scene type, and then judging whether the current home scene type has changed. If it changes, the predicted acoustic feature is generated by the voice prediction strategy corresponding to the changed home scene type.

[0101] In the embodiment, the voice prediction strategy can have various implementations, for example, it can be obtained by user self-defined setting. For another example, it can be obtained by big data collection and analysis. For another example, the smart device can obtain it by calling a trained model.

[0102] Please refer to Figure 9 In an implementation mode, the smart device can call a trained scene acoustic style extractor to obtain the acoustic feature of the home scene in a noisy acoustic environment, input the acoustic feature of the home scene in the noisy acoustic environment into the scene acoustic style extractor to obtain the scene style embedding information corresponding to the home scene. The smart device can call a trained acoustic feature prediction model, take the response text content information and the scene style embedding information as the input of the acoustic feature prediction model, obtain the predicted acoustic feature corresponding to the home scene, and then synthesize the output response voice data.

[0103] Please refer to Figure 10In another implementation, the intelligent device can call the trained scene acoustic style extractor and the first scene classification model, input the environmental acoustic features of the home scene in the noisy acoustic environment into the first scene classification model, obtain the scene type information corresponding to the home scene, input the scene type information into the scene acoustic style extractor for fusion, such as inputting the scene type information and the acoustic features of the home scene in the noisy acoustic environment as inputs of the scene acoustic style extractor, and the scene acoustic style extractor outputs the scene style embedding information corresponding to the home scene. The intelligent device can call the trained acoustic feature prediction model, input the response text content information and the scene style embedding information as inputs of the acoustic feature prediction model, obtain the predicted acoustic features corresponding to the home scene, and then synthesize the output response speech data.

[0104] It should be noted that for the above embodiments, the intelligent device can also directly use the output data obtained by training the above scene acoustic style extractor, acoustic feature prediction model, first scene classification model or combination as part of the input for speech synthesis. Obviously, for some possible implementations, the above scene acoustic style extractor, acoustic feature prediction model, first scene classification model or combination needs to be trained in different scenes in order to obtain output data that can represent different situations.

[0105] In this embodiment, the training process of the scene acoustic style extractor, acoustic feature prediction model and scene classification model can be flexibly selected, and the following examples are given in this embodiment.

[0106] Please refer to Figure 11 A flowchart of a scene acoustic style extractor training method provided by an embodiment of the present application can be executed by Figure 1 the intelligent device 100, for example, can be executed by the processor 103 in the intelligent device 100. The scene acoustic style extractor training method includes S210 and S220.

[0107] S210, obtaining acoustic features of a home scene in a noisy acoustic environment.

[0108] S210, inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain training style embedding information corresponding to the home scene. The training style embedding information represents the scene acoustic style corresponding to the home scene.

[0109] The scene acoustic style in this embodiment can include one or more attributes, for example, can include any one or a combination of two or more of tone, sound intensity, speech rate, emotion and the like.

[0110] The home scenes can be flexibly divided, and acoustic features of the home scenes in a noisy acoustic environment can be collected. For example, the kitchen, the living room, and the bathroom where noisy sounds often occur can be selected as the home scenes and acoustic features in a noisy acoustic environment can be collected. For another example, the home gym, the multimedia audio-visual room, and the study can be selected as the home scenes and acoustic features in a noisy acoustic environment can be collected.

[0111] It can be understood that the more detailed the division of the home scenes is, the more speech can be prepared according to each subdivided home scene, more acoustic features can be selected for classification, the acoustic features in a noisy acoustic environment can be used to train the scene acoustic style extractor, the training style embedding information obtained is more detailed, and the corresponding scene acoustic style of each home scene is more delicate. Correspondingly, the scene style embedding information corresponding to the home scene determined by the trained scene acoustic style extractor is more delicate, and the response speech data synthesized by the intelligent device using the scene style embedding information is clearer and more natural.

[0112] The acoustic features of the home scenes in a noisy acoustic environment are obtained based on Lombard speech in a noisy home scene. For example, Lombard speech in a noisy home scene can be collected and acoustic features can be extracted, such as using an acoustic feature extraction module to extract acoustic features, to obtain acoustic features of the home scenes in a noisy acoustic environment.

[0113] The Lombard speech in a noisy home scene can be obtained in various ways. For example, the speech of a user can be recorded under the interference of environmental sound in various home scenes as training data of the scene acoustic style extractor. For example, volunteers can be made to wear portable earphones, the earphones can play environmental sound in different home scenes, and the environmental noise can induce the volunteers to change their voices, so that Lombard speech in different home scenes can be captured as training data of the scene acoustic style extractor.

[0114] In order to improve the training effect, the acoustic features in a noisy home scene obtained can be as rich as possible. For example, the number of volunteers, the amount of speech of each volunteer in each noisy home scene, and the like can be increased to increase the Lombard speech captured in each noisy home scene, and thus the richness of the acoustic features in a noisy home scene obtained can be improved.

[0115] By obtaining the Lombard speech corresponding to the home scene as rich and comprehensive as possible, acoustic features are extracted to obtain acoustic features of the home scene in a noisy acoustic environment, and the scene acoustic style extractor is input for training, which can effectively improve the accuracy of the training style embedding information extracted by the scene acoustic style extractor. Correspondingly, the more accurate and reliable the scene style embedding information corresponding to the home scene determined by the trained scene acoustic style extractor is.

[0116] In this embodiment, the acoustic features in a noisy acoustic environment can be various. For example, the acoustic features can be spectral envelope, fundamental frequency (F0) of sound, acoustic energy, acoustic spectrum diagram such as linear spectrogram, Mel Frequency Cepstrum Coefficient (MFCC), Mel spectrogram, etc.

[0117] In an implementation manner, the training of the scene acoustic style extractor can be based on the Lombard speech only. For example, for each home scene, Lombard speech in a noisy home scene can be collected and acoustic features are extracted to obtain acoustic features of the home scene in a noisy acoustic environment, and the acoustic features in the noisy acoustic environment are input into the scene acoustic style extractor to extract the embedding representation of the corresponding scene acoustic style features (training style embedding information).

[0118] The scene acoustic style extractor can have various implementation manners. For example, please refer to Figure 12 and Figure 13 The scene acoustic style extractor can include a reference encoder, a first attention module and a fully connected layer. Wherein, the step of inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene can be realized by S221, S222 and S223.

[0119] S221, input the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain the reference embedding information.

[0120] S222, input the reference embedding information into the first attention module to obtain the attention weight.

[0121] S223, input the attention weight into the fully connected layer to obtain the training style embedding information.

[0122] Based on this training method, the training style embedding information is obtained by inputting the as rich and comprehensive acoustic features of the home scene in the noisy acoustic environment into the scene acoustic style extractor, so as to realize the training of the scene acoustic style extractor.

[0123] In another implementation, the training of the scene acoustic style extractor can be based on Lombard speech and the scene type information output by the first scene classification model.

[0124] For example, in the case of inputting the acoustic features of the home scene in the noisy acoustic environment into the scene acoustic style extractor, the environmental acoustic features of the home scene can also be obtained, the environmental acoustic features of the home scene are input into the first scene classification model to obtain the scene type information corresponding to the home scene, and the scene type information is input into the scene acoustic style extractor for fusion, that is, the scene type information is introduced for learning of the scene acoustic style extractor.

[0125] In order to improve the convenience and efficiency of fusion, the scene type information can be input into the scene acoustic style extractor in the form of weight for fusion.

[0126] For example, in the case of the first scene classification model being a VGG16 network, the step of inputting the environmental acoustic features of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene can include: inputting the environmental acoustic features of the home scene into the VGG16 network to obtain a Softmax probability value; determining a first scene type weight corresponding to the Softmax probability value; and taking the first scene type weight as the scene type information.

[0127] The corresponding relationship between the Softmax probability value and the first scene type weight can be set in advance, and accordingly, the step of determining the first scene type weight corresponding to the Softmax probability value can include: determining the first scene type weight corresponding to the Softmax probability value according to the corresponding relationship between the Softmax probability value and the first scene type weight.

[0128] Of course, in the case of the scene classification model being a VGG16 network, a label value can also be obtained.

[0129] Please refer to Figure 14 and Figure 15 In the case of the scene acoustic style extractor including a reference encoder, a first attention module and a full connection layer, Figure 11 The step S220 shown in FIG. 2, that is, the step of inputting the acoustic features in the noisy acoustic environment into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene, can be implemented by S224, S225 and S226.

[0130] S224, input the acoustic feature of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information.

[0131] S225, input the reference embedding information into the first attention module to obtain attention weight.

[0132] In the case of taking the first scene type weight as the scene type information, correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion is implemented through S226.

[0133] S226, input the attention weight and the first scene type weight into the full connection layer for weighting to obtain training style embedding information.

[0134] For example, in the case of taking the first scene classification model as a ResNet network, the step of inputting the environmental acoustic feature of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene can include: inputting the environmental acoustic feature of the home scene into the ResNet network to obtain a label value of the home scene; determining a second scene type weight corresponding to the label value; and taking the second scene type weight as the scene type information.

[0135] The corresponding relationship between the label value and the second scene type weight can be set in advance, and correspondingly, the step of determining the second scene type weight corresponding to the label value can include: determining the second scene type weight corresponding to the label value according to the corresponding relationship between the label value and the second scene type weight.

[0136] Please refer to Figure 16 and Figure 17 In the case of taking the first scene classification model as a ResNet network, the step of inputting the environmental acoustic feature of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene can include: inputting the environmental acoustic feature of the home scene into the ResNet network to obtain a label value of the home scene; determining a second scene type weight corresponding to the label value; and taking the second scene type weight as the scene type information.

[0137] S227, input the acoustic feature of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information.

[0138] S228, input the reference embedding information into the first attention module to obtain attention weight.

[0139] In the case of taking the second scene type weight as the scene type information, correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion is implemented through S229.

[0140] S229, input the attention weight and the second scene type weight into the full connection layer for weighting to obtain training style embedding information.

[0141] The training process of the scene acoustic style extractor and the implementation structure of the scene acoustic style extractor are only examples, and the training process of the scene acoustic style extractor and the implementation structure of the scene acoustic style extractor can be other, which will not be illustrated one by one in the embodiment.

[0142] In the above example of the training process of the scene acoustic style extractor, the step of inputting the acoustic features of the home scene in the noisy acoustic environment into the reference encoder to obtain the reference embedding information can include: inputting the acoustic features of the variable length Lombard speech (the acoustic features of the home scene in the noisy acoustic environment) into the reference encoder, and the reference encoder compresses the acoustic features of the variable length Lombard speech into fixed length reference embedding information such as a fixed size vector. The reference embedding information is used to encode the acoustic style of the whole piece of audio.

[0143] The reference encoder can have various implementation structures, and the following examples are illustrated in the embodiment.

[0144] For example, as shown in Figure 18 The reference encoder can include a convolutional neural network (CNN), a bidirectional long short-term memory recurrent neural network (including a backward recurrent neural network (backward LSTM) and a forward recurrent neural network (forward LSTM)), and a mapping layer.

[0145] The acoustic features of the home scene in the noisy acoustic environment are input into the CNN to extract further acoustic features, the further acoustic features are input into the bidirectional long short-term memory recurrent neural network to obtain context-related acoustic features, and the context-related acoustic features are input into the mapping layer to output fixed length reference embedding information, which encodes the acoustic style of the whole piece of audio. For example, the Lombard speech acoustic style of the user in each home scene.

[0146] It can be understood that if different home scenes with similar environmental acoustic features are distinguished, the reference encoder processes the acoustic features of different home scenes with similar environmental acoustic features in the noisy acoustic environment, and the acoustic style corresponding to the different home scenes with similar environmental acoustic features can be refined.

[0147] For another example, as shown in Figure 19As shown, the reference encoder may include a first pre-trained model output layer, a CNN, and a mapping layer. The first pre-trained model can refer to the implementation of Audio word2vec, which is mainly implemented by a sequence-to-sequence autoencoder. Audio word2vec mainly uses a recurrent neural network (RNN) structure. The output layer of the first pre-trained model outputs pre-trained vectors. The purpose of using Audio word2vec is to obtain a better representation of acoustic features. Further acoustic features are then extracted through two to three convolutional neural networks, and then mapped to predefined dimensions through two to three fully connected networks as mapping layers.

[0148] For example, such as Figure 20 As shown, the reference encoder may include a second pre-trained model output layer, a bidirectional long short-term memory recurrent neural network, and a mapping layer. The second pre-trained model can refer to the implementation of an unsupervised pre-trained model (wav2vec), mainly implemented by a CNN-based encoder. The output layer of the second pre-trained model outputs pre-trained vectors. The purpose of using Wave2Vec is to obtain a better representation of acoustic features. Further context-dependent acoustic features are then extracted through the bidirectional long short-term memory recurrent neural network, and then mapped to a predefined dimension through two to three fully connected networks as mapping layers.

[0149] Using the first pre-trained model and the second pre-trained model can yield more representations of acoustic features.

[0150] The first attention module can obtain the attention weight in multiple ways. Illustratively, a set of acoustic style markers can be formed based on a permutation and combination of attributes such as pitch, sound intensity, speech rate, emotion, etc., and in the process of obtaining the attention weight by the first attention module, a group of acoustic style markers are randomly called from the set of acoustic style markers for embedding. Each acoustic style marker can include one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. The number of embedded acoustic style markers can be flexibly set, for example, it can be set to one, three, five, ten, etc., to represent a small number of different acoustic dimensions in the training data of the scene acoustic style extractor, such as one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. The reference embedding information is taken as the query information of the first attention module, and the first attention module learns the similarity measure between the reference embedding information and each acoustic style marker in the embedded group of acoustic style markers. The first attention module further outputs a group of combination weights (attention weights), which represent the contribution of each acoustic style marker to the reference embedding information. Accordingly, the training style embedding information output by the scene acoustic style extractor can be information related to any one or more of the attributes such as pitch, sound intensity, speech rate, emotion, etc. For example, in the case of a living room in a home scene, the training style embedding information corresponding to the living room can include increasing the pitch, increasing the sound intensity, and increasing the speech rate.

[0151] In this embodiment, the training style embedding information can have multiple presentation modes. For example, the training style embedding information can be a specific assignment of attributes such as pitch, sound intensity, speech rate, emotion, etc. For another example, the training style embedding information can be an adjustment value for the attributes such as pitch, sound intensity, speech rate, emotion, etc. based on a certain normal voice. The normal voice can be a voice emitted in a quiet scene.

[0152] Please refer to Figure 21 In the case of introducing scene type information such as Softmax probability value or label value for learning by the scene acoustic style extractor, the Softmax probability value and the label value can be processed into information that can be recognized and used by the fully connected layer through the style embedding adjuster. For example, the first scene type weight corresponding to the Softmax probability value is determined through the style embedding adjuster, and the second scene type weight corresponding to the label value is determined. In this case, the home scene to which the acoustic feature in the noisy acoustic environment belongs is known.

[0153] In the case that the scene type information is a label value, the full connection layer can multiply the second scene type weight corresponding to the label value with the attention weight, only retain one acoustic style label as the training style embedding information corresponding to the corresponding home scene, and can determine which attribute a certain acoustic style label in the home scene has greater influence on. Manual definition of multiple groups of weight parameters can be supported during training, and the training efficiency can be improved by directly calling during inference.

[0154] In the case that the scene type information is a Softmax probability value, the full connection layer can multiply the first scene type weight corresponding to the Softmax probability value with the attention weight, and balance the contribution of each embedded acoustic style label to the reference embedding information again, so that the scene acoustic style extractor can not only learn the acoustic style of Lombard speech in the corresponding home scene, but also learn the influence of different home scenes on the acoustic style.

[0155] In order to improve the training efficiency of the scene acoustic style extractor and accelerate the convergence, in another implementation manner, a second scene classification model can also be called, which can indicate the home scene type to which the acoustic feature in the noisy acoustic environment belongs. By inputting the acoustic feature of the home scene in the noisy acoustic environment into the second scene classification model, the scene type indication information corresponding to the home scene is obtained, so that the scene acoustic style extractor can be trained according to the scene type indication information. For example, in the process of inputting the acoustic feature in the noisy acoustic environment into the scene acoustic style extractor and training the scene acoustic style extractor, the acoustic feature in the noisy acoustic environment is inputted into the second scene classification model, the second scene classification model outputs the scene type indication information, and the scene type indication information is transmitted to the scene acoustic style extractor, such as the full connection layer of the scene acoustic style extractor, to train the full connection layer and assist the convergence and improve the training efficiency. In the embodiment, in the case that the second scene classification model is called to train the scene acoustic style extractor, the acoustic feature in the noisy acoustic environment inputted into the second scene classification model and the acoustic feature in the noisy acoustic environment inputted into the scene acoustic style extractor can be the same or different, as long as the acoustic feature of the home scene in the noisy acoustic environment inputted into the scene acoustic style extractor and the acoustic feature of the home scene in the noisy acoustic environment inputted into the second scene classification model correspond to the same home scene.

[0156] Please refer to Figure 22 A flowchart of a first scene classification model training method provided by the embodiment of the present application can be executed by Figure 1 the intelligent device 100, for example, can be executed by the processor 103 in the intelligent device 100. The first scene classification model training method includes S310 and S320.

[0157] S310, obtaining environment acoustic features of a plurality of home scenes.

[0158] S320, inputting the environment acoustic features of the plurality of home scenes into a scene classification network for training to obtain scene type information corresponding to each home scene.

[0159] Among them, the home scenes can be flexibly divided, and the environment acoustic features of the home scenes can be collected. For example, the kitchen, the living room, and the bathroom where noisy sounds often occur can be selected as home scenes respectively, and the environment acoustic features of the home scenes can be collected. For another example, the home gym, the multimedia audio-visual room, and the study can also be selected as home scenes respectively, and the environment acoustic features of the home scenes can be collected. Through fine division of the home scenes, the environment acoustic features of different home scenes are used to train the scene classification network, and the scene type information corresponding to each home scene is obtained, and based on the scene type information, the judgment of the home scene type can be realized.

[0160] In order to improve the training effect, the environment acoustic features of each home scene obtained can be as rich as possible. For example, the environment acoustic features of each home scene can include the environment acoustic features of each home device in the home scene alone, and the environment acoustic features of each home device in the home scene arranged in at least two combinations. For example, in the case of a kitchen as a home scene, the environment acoustic features of the kitchen can include the environment acoustic features of each home device in the kitchen, such as an exhaust hood, a dishwasher, a gas stove, and a pot, bowl, and ladle, and the environment acoustic features of two, three, four, etc. combinations of the exhaust hood, the dishwasher, the gas stove, and the pot, bowl, and ladle in the kitchen.

[0161] By obtaining the environment acoustic features of the home scenes as rich and comprehensive as possible, inputting the scene classification network for training, the sensitivity and reliability of the first scene classification model obtained by training can be effectively improved. For example, by inputting the environment acoustic features of the home devices alone into the scene classification network for training, the type of the home scene can be directly determined based on the environment acoustic features of certain specific home devices. For example, based on the acoustic features of the toilet flushing sound, the home scene can be directly determined as the bathroom. For another example, by inputting the environment acoustic features of various combinations of home devices in each home scene into the scene classification network for training, reliable differentiation of different home scenes with similar environment acoustic features can be achieved. For example, the kitchen and the bathroom may both contain the acoustic features of water flow, but in combination with the acoustic features of the frying pan cooking sound, the home scene can be determined as the kitchen.

[0162] Please refer to Figure 23In an implementation manner, S310, obtaining the environmental acoustic features of the plurality of home scenes can be implemented through S311 and S312.

[0163] S311, obtaining the environmental acoustic data of the plurality of home scenes.

[0164] S312, performing acoustic feature extraction on the environmental acoustic data to obtain the environmental acoustic features of the plurality of home scenes.

[0165] The acoustic feature extraction can be performed in various manners, for example, the acoustic feature extraction can be performed on the environmental acoustic data by an acoustic feature extraction module. The environmental acoustic data in the embodiment is acoustic data generated by each home device in use in the home scene, and does not include human voice.

[0166] The scene classification network can be flexibly selected, for example, the scene classification network can be a set deep learning network. Illustratively, the scene classification network can be a deep convolutional neural network, such as a VGG16 network, a ResNet network, etc. According to different scene classification networks, different scene type information can be obtained. Illustratively, in the case where the scene classification network is a VGG16 network, in S320, the environmental acoustic features are input into the VGG16 network, and the obtained scene type information is a Softmax probability value (of course, a label value such as one-hot encoding can also be obtained). In the case where the scene classification network is a ResNet network, in S320, the environmental acoustic features are input into the ResNet network, and the obtained scene type information is a label value such as one-hot encoding (of course, the obtained scene type information can also be a Softmax probability value).

[0167] Based on the above design, according to the label value representing the home scene type, the home scene type to which the environmental acoustic features of the home scene belong can be determined, or according to the Softmax probability value representing that the home scene belongs to each type, different weight combinations of the home scene style can be determined, so as to realize the refinement of the home scene recognition granularity, improve the home scene recognition accuracy, and better "guide" the extraction of the scene style embedding.

[0168] Based on the refined classification of the home scene by the first scene classification model, the scene acoustic style under the home scene can be more finely divided. And the extraction of the scene style embedding can be better "guided". Illustratively, in the case where the home scene includes a kitchen, a living room and a washroom, the scene acoustic style extractor can combine the first scene classification model to perform scene acoustic style extraction for the kitchen, the living room and the washroom, respectively.

[0169] In order to more clearly illustrate the first scene classification model training method in the embodiment of the present application, the following scene is taken as an example for illustration.

[0170] In the case where the home scenarios include three kinds of kitchen, living room and washroom, the environmental acoustic data generated by each home device in the kitchen alone, the environmental acoustic data generated by two permutations and combinations, three permutations and combinations, four permutations and combinations, etc. are collected, and all the environmental acoustic data collected in the kitchen are taken as a first acoustic data set. The environmental acoustic data generated by each home device in the living room alone, the environmental acoustic data generated by two permutations and combinations, three permutations and combinations, four permutations and combinations, etc. are collected, and all the environmental acoustic data collected in the living room are taken as a second acoustic data set. The environmental acoustic data generated by each home device in the washroom alone, the environmental acoustic data generated by two permutations and combinations, three permutations and combinations, four permutations and combinations, etc. are collected, and all the environmental acoustic data collected in the washroom are taken as a third acoustic data set.

[0171] The acoustic features of the environmental acoustic data in the first acoustic data set are extracted to obtain a first acoustic feature set. The acoustic features of the environmental acoustic data in the second acoustic data set are extracted to obtain a second acoustic feature set. The acoustic features of the environmental acoustic data in the third acoustic feature set are extracted to obtain a third acoustic feature set.

[0172] The scene type information corresponding to the kitchen is set as a first type, the scene type information corresponding to the living room is set as a second type, and the scene type information corresponding to the washroom is set as a third type.

[0173] In the case where the acoustic features in the first acoustic feature set are input into the scene classification network, the first type is taken as the target output of the scene classification network. In the case where the acoustic features in the second acoustic feature set are input into the scene classification network, the second type is taken as the target output of the scene classification network. In the case where the acoustic features in the third acoustic feature set are input into the scene classification network, the third type is taken as the target output of the scene classification network. The scene classification network is trained based on this until a convergence condition is reached, and then a required first scene classification model is obtained.

[0174] The convergence condition can be flexibly set. For example, the accuracy of obtaining the scene type information corresponding to each home scenario can be set to a preset value. Illustratively, the acoustic features in the first acoustic feature set, the second acoustic feature set and the third acoustic feature set can be divided into training data and test data. The scene classification network is trained based on the training data until the accuracy of the target output obtained after inputting the test data reaches the preset value, it is determined that the convergence condition is reached, and the required first scene classification model is obtained. For another example, the number of training times can be set to a certain amount. This embodiment does not limit this.

[0175] After the first scene classification model is obtained by using the above scene classification model training method, the first scene classification model can be used to identify the weight combination of the home scene or the home scene style to which the environmental acoustic feature belongs. For example, if the environmental acoustic feature of a home scene is input into the scene classification model, the scene classification model outputs a first type, and according to the first type, it can be determined that the environmental acoustic feature belongs to a kitchen.

[0176] After the environmental acoustic features of multiple home scenes are input into the trained first scene classification model, the first scene classification model outputs scene type related information, such as a label value or a softmax probability value. The scene type information output by the first scene classification model can be used as input of the scene acoustic style extractor, and be combined with the training and use of the scene acoustic style extractor.

[0177] The embodiment of the present application also provides a second scene classification model training method, which comprises: obtaining acoustic features of multiple home scenes in a noisy acoustic environment, inputting the acoustic features of the multiple home scenes in the noisy acoustic environment into a scene classification network for training, and obtaining scene type indication information corresponding to each home scene.

[0178] The second scene classification model uses acoustic features of multiple home scenes in a noisy acoustic environment as input, and accordingly, the second scene classification model can indicate the home scene type to which the acoustic feature in the noisy acoustic environment belongs.

[0179] Since the difference between the training processes of the second scene classification model and the first scene classification model is only that the input data of the two models is different, and since the output data is different due to the different input data, the scene type indication information output by the second scene classification model is used as feedback for training of the scene acoustic style extractor, such as training feedback of the fully connected layer of the scene acoustic style extractor according to the scene type indication information, so as to accelerate convergence. The scene type information output by the first scene classification model is used for weighting, such as weighting of the fully connected layer of the scene acoustic style extractor by inputting the scene type information and attention weight. The optional structure and training principle of the second scene classification model can be referred to the above description of the first scene classification model, and thus will not be repeated here.

[0180] During the training and application of the scene acoustic style extractor, the first scene classification model can be selected to be called, the second scene classification model can be selected to be called, the first scene classification model and the second scene classification model can be selected to be called, or neither the first scene classification model nor the second scene classification model can be called.

[0181] Please refer to Figure 24 A flowchart of an acoustic feature prediction model training method provided by the embodiment of the present application can be obtained by Figure 1The intelligent device 100 shown performs, for example, can be performed by the processor 103 in the intelligent device 100. The acoustic feature prediction model training method includes S410 and S420.

[0182] S410, obtain training style embedding information corresponding to the to-be-output text information and the home scene.

[0183] S420, input the to-be-output text information and the training style embedding information into an acoustic feature prediction network to obtain predicted features of the to-be-synthesized sound corresponding to the home scene.

[0184] The training style embedding information corresponding to the home scene can be obtained by the aforementioned scene acoustic style extractor, which is not described here.

[0185] In this embodiment, the training style embedding information corresponding to the home scene is ingeniously introduced as a new voice synthesis consideration dimension, and the to-be-output text information and the training style embedding information are both input into the acoustic feature prediction network to obtain predicted features of the to-be-synthesized sound corresponding to the home scene, which has the Lombard speech acoustic style. Based on the two dimensions of the to-be-output text information and the training style embedding information, reliable prediction of the acoustic features corresponding to the home scene can be realized, and predicted acoustic features with the Lombard speech acoustic style are output.

[0186] In one implementation, the voice uttered by the user in a quiet scene can be collected as normal voice, and the acoustic feature prediction network is jointly trained using the normal voice and the training style embedding information to obtain the predicted features of the to-be-synthesized sound corresponding to the home scene. For example, in the case of increasing pitch, increasing sound intensity, and increasing speech rate, the acoustic feature prediction network is jointly trained using the normal voice and the training style embedding information to obtain predicted features of the to-be-synthesized sound, which can include increasing the volume, sound intensity, and speech rate by a set value based on the normal voice.

[0187] The acoustic feature prediction network can have various implementations. In order to improve the naturalness of the predicted features of the to-be-synthesized sound, Tacotron can be selected as the acoustic feature prediction network.

[0188] Please refer to Figure 25In an implementation manner, the acoustic feature prediction network can include an encoder, a second attention module, and a decoder. Accordingly, the step of inputting the to-be-output text information and the training style embedding information into the acoustic feature prediction network to obtain the to-be-synthesized sound prediction feature corresponding to the home scene in S420 can be implemented by the following manner: inputting the to-be-output text information into the encoder to obtain character embedding information with a fixed length; inputting the character embedding information with the fixed length and the training style embedding information into the second attention module to obtain aligned acoustic features and character information; and inputting the aligned acoustic features and character information into the decoder to obtain the to-be-synthesized sound prediction feature.

[0189] In the case of selecting Tacotron as the acoustic feature prediction network, the to-be-synthesized sound prediction feature output by the decoder can be linear spectrum (Linear Spectrum) or mel spectrum (Mel Spectrum) or other vocoder applicable acoustic features after the aligned acoustic features and character information are input into the decoder.

[0190] The above lists the training processes of the scene acoustic style extractor, the acoustic feature prediction model, the first scene classification model, and the second scene classification model. Based on the above training processes, the trained scene acoustic style extractor, the acoustic feature prediction model, the first scene classification model, and the second scene classification model can be obtained. The above first scene classification model training method, second scene classification model training method, scene acoustic style extractor training method, and acoustic feature prediction model training method can be executed in the same intelligent device or in different intelligent devices. The first scene classification model training method, second scene classification model training method, scene acoustic style extractor training method, and acoustic feature prediction model training method can be run individually or combined and run according to different collocation modes according to requirements. The trained scene acoustic style extractor, acoustic feature prediction model, first scene classification model, and second scene classification model can be run individually or combined and run according to different collocation modes according to requirements.

[0191] For example, the intelligent device can call the trained scene acoustic style extractor and acoustic feature prediction model in the process of synthesizing and outputting response voice data. Please refer to Figure 26, the scene acoustic style extractor can include a reference encoder, a first attention module and a full connection layer. In a case where only the acoustic feature of the home scene in the noisy acoustic environment is taken as the input of the scene acoustic style extractor, the scene style embedding information corresponding to the home scene can be obtained by: inputting the acoustic feature of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; inputting the reference embedding information into the first attention module to obtain attention weights; and inputting the attention weights into the full connection layer to obtain the scene style embedding information. The response text content information and the scene style embedding information are taken as the input of the acoustic feature prediction model, so as to obtain the predicted acoustic feature corresponding to the home scene. The predicted acoustic feature is synthesized, and then the output response voice data is obtained. It can be understood that, when it is analyzed according to the received sound data that the home scene type is the noisy type, the acoustic feature of the home scene in the noisy acoustic environment can be obtained by performing feature extraction on the sound data.

[0192] For example, the intelligent device can call the trained scene acoustic style extractor, the first scene classification model and the acoustic feature prediction model in the process of synthesizing the output response voice data. Please refer to Figure 27 , the scene acoustic style extractor can include a reference encoder, a first attention module and a full connection layer. In a case where only the acoustic feature of the home scene in the noisy acoustic environment is taken as the input of the scene acoustic style extractor, the scene style embedding information corresponding to the home scene can be obtained by: inputting the acoustic feature of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; inputting the reference embedding information into the first attention module to obtain attention weights; and inputting the attention weights into the full connection layer to obtain the scene style embedding information. The response text content information and the scene style embedding information are taken as the input of the acoustic feature prediction model, so as to obtain the predicted acoustic feature corresponding to the home scene. The predicted acoustic feature is synthesized, and then the output response voice data is obtained. It can be understood that, when it is analyzed according to the received sound data that the home scene type is the noisy type, the acoustic feature of the home scene in the noisy acoustic environment can be obtained by performing feature extraction on the sound data.

[0193] When the first scene classification model is a VGG16 network, the environmental acoustic feature of the home scene can be input into the VGG16 network to obtain a Softmax probability value; a first scene type weight corresponding to the Softmax probability value is determined; and the first scene type weight is taken as the scene type information. Wherein, the first scene type weight corresponding to the Softmax probability value can be determined according to the corresponding relationship between the Softmax probability value and the first scene type weight.

[0194] Correspondingly, the attention weights and the first scene type weight are input into the full connection layer for weighting to obtain the scene style embedding information.

[0195] When the first scene classification model is a ResNet network, the environmental acoustic feature of the home scene can be input into the ResNet network to obtain a label value of the home scene; a second scene type weight corresponding to the label value is determined; and the second scene type weight is taken as the scene type information. The second scene type weight corresponding to the label value can be determined according to the correspondence between the label value and the second scene type weight.

[0196] Correspondingly, the attention weight and the second scene type weight are input into the full connection layer for weighting to obtain the scene style embedding information.

[0197] In this embodiment, the acoustic feature of the home scene in a noisy acoustic environment is taken as the input of the scene acoustic style extractor, or the acoustic feature of the home scene in a noisy acoustic environment and the scene type information output by the first scene classification model are both taken as the input of the scene acoustic style extractor, and the scene acoustic style extractor thereby outputs the scene style embedding information. The scene style embedding information output by the scene acoustic style extractor can be taken as the input of the acoustic feature prediction model.

[0198] Please refer to Figure 28 The acoustic feature prediction model can include an encoder, a second attention module and a decoder. Correspondingly, the response text content information and the scene style embedding information are taken as the input of the acoustic feature prediction model, and the predicted acoustic feature corresponding to the home scene can be obtained by the following manner: the response text content information is input into the encoder to obtain fixed-length character embedding information; the fixed-length character embedding information and the scene style embedding information are input into the second attention module to obtain aligned acoustic feature and character information; and the aligned acoustic feature and character information are input into the decoder to obtain the predicted acoustic feature corresponding to the home scene.

[0199] In this embodiment, the response text content information and the scene style embedding information are taken as the input of the acoustic feature prediction model, and the acoustic feature prediction model thereby outputs the predicted acoustic feature. Based on the predicted acoustic feature, the output response speech data with Lombard effect can be synthesized, so that the intelligent device can play the output response speech data in a noisy home scene, which can be reliably received by the user, thereby ensuring the smoothness of human-computer interaction in the smart home scene and better meeting the actual application requirements.

[0200] In an implementation manner, a speech synthesis device can be used to synthesize the output response speech data. For example, Figure 29As shown, the vocoder can be used as a speech synthesis device to synthesize the predicted acoustic features output by the acoustic feature prediction model. The predicted acoustic features are input into the vocoder to obtain output response speech data containing information about the content of the response text. In this embodiment, the vocoder can be flexibly selected. For example, when the predicted acoustic features are linear spectrum, Griffin-Lim can be selected as the vocoder for converting the spectrum to waveform, or WaveNet can be selected as the vocoder. For another example, when the predicted acoustic features are Mel spectrum, WaveNet can be selected as the vocoder.

[0201] The above describes optional embodiments of the method for synthesizing output response speech data in the embodiments of the present application. In other implementations, the same design concept is used: by "imitating" the communication mode of human beings actively changing the way of making sounds under the Lombard effect, speech data with corresponding scene acoustic features is synthesized for different home scenes to report, and Lombard speech with good naturalness and clarity is synthesized to ensure smoothness of voice interaction with the user in a noisy home environment. The method for synthesizing output response speech data in the embodiments of the present application can also have other implementations.

[0202] In another implementation, please refer to Figure 30 The Lombard speech can be input into the Lombard speech generation model, and the Lombard speech generation model can be trained based on the Lombard speech so that the Lombard speech generation model can directly obtain output response speech data according to the content information of the response text. Based on this kind of Lombard speech generation model, the intelligent device does not need to train and call the first scene classification model, the second scene classification model, the scene acoustic style extractor and the acoustic feature prediction model, and directly input the content information of the response text into the Lombard speech generation model to obtain output response speech data corresponding to the home scene.

[0203] In this embodiment, a speech synthesis device can be used to synthesize the predicted acoustic features to obtain output response speech data. In one implementation, the speech synthesis device can be a vocoder, and the predicted acoustic features are input into the vocoder to obtain the output response speech data.

[0204] In another implementation, the intelligent device can also directly train a speech synthesis model to obtain output response speech data based on the trained speech synthesis model.

[0205] Please refer to Figure 31 A flowchart of a speech synthesis model training method provided by an exemplary embodiment of the present application can be performed by Figure 1The execution is performed by the smart device 100, for example, by the processor 103 within the smart device 100. The speech synthesis model training method includes steps S510, S520, and S530.

[0206] S510 obtains the text information to be output and the training style embedding information corresponding to the home scene.

[0207] S520: Based on the text information to be output and the training style embedding information, the predicted features of the synthesized sound corresponding to the home scene are obtained.

[0208] The S530 synthesizes the predicted features of the synthesized speech to obtain the speech data to be output.

[0209] The training style embedding information corresponding to the home scene can be obtained through a scene acoustic style extractor. For example, the acoustic features of a noisy acoustic environment corresponding to the home scene can be input into the scene acoustic style extractor to obtain the training style embedding information corresponding to the home scene. The predicted features of the synthesized sound can be obtained through an acoustic feature prediction network. For example, the training style embedding information and the text information to be output can be used as input to the acoustic feature prediction network to obtain the predicted features of the synthesized sound corresponding to the home scene.

[0210] Accordingly, please refer to Figure 32 The acoustic features of a home scene in a noisy acoustic environment can be input into a scene acoustic style extractor to obtain training style embedding information corresponding to the home scene. This training style embedding information represents the scene acoustic style corresponding to the home scene. The training style embedding information and the text information to be output are then used as input to an acoustic feature prediction network to obtain the predicted features for synthesized sound corresponding to the home scene. These predicted features are then synthesized to obtain the output speech data.

[0211] In this embodiment, the scene acoustic style extractor may include a reference encoder, a first attention module, and a fully connected layer. The training style embedding information can be obtained through the following steps: inputting the acoustic features of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; inputting the reference embedding information into the first attention module to obtain attention weights; and inputting the attention weights into the fully connected layer to obtain training style embedding information. Please return to the reference section. Figure 12 and Figure 13 One training process for the scene acoustic style extractor can be found in the descriptions of S210 and S221 to S223 above, and will not be repeated here.

[0212] The acoustic feature prediction network can include an encoder, a second attention module, and a decoder. Accordingly, the predicted acoustic feature of the to-be-synthesized sound can be obtained by the following steps: inputting the to-be-output text information into the encoder to obtain character embedding information of a fixed length; inputting the character embedding information of the fixed length and the training style embedding information into the second attention module to obtain aligned acoustic features and character information; and inputting the aligned acoustic features and character information into the decoder to obtain the predicted acoustic feature of the to-be-synthesized sound. Please refer back to Figure 24 and Figure 25 The training process of the acoustic feature prediction model can be referred to the related descriptions of S410 and S420 above, which will not be repeated here.

[0213] On the basis of training the speech synthesis model in combination with the scene acoustic style extractor training method and the acoustic feature prediction model training method, the first scene classification model training method can also be introduced to train the speech synthesis model. The training process of the first scene classification model can be referred to the related descriptions of S310 and S320 above, which will not be repeated here. Accordingly, the speech synthesis model training method can further include: inputting the environmental acoustic feature of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene, and inputting the scene type information into the scene acoustic style extractor for fusion.

[0214] In an implementation manner, when the first scene classification model is a VGG16 network, the step of inputting the environmental acoustic feature of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene can include: inputting the environmental acoustic feature of the home scene into the VGG16 network to obtain a Softmax probability value; determining a first scene type weight corresponding to the Softmax probability value; and taking the first scene type weight as the scene type information. Illustratively, the first scene type weight corresponding to the Softmax probability value can be determined according to the corresponding relationship between the Softmax probability value and the first scene type weight.

[0215] The scene acoustic style extractor can include a reference encoder, a first attention module, and a full connection layer. The training style embedding information can be obtained by the following way: inputting the acoustic feature of the home scene in a noisy acoustic environment into the reference encoder to obtain reference embedding information; and inputting the reference embedding information into the first attention module to obtain an attention weight. Accordingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion includes: inputting the attention weight and the first scene type weight into the full connection layer for weighting to obtain the training style embedding information. Please refer back to Figure 14 and Figure 15 One of the training processes of the scene acoustic style extractor can be referred to the related descriptions of S210 and S224 to S226 above, which will not be repeated here.

[0216] In another implementation manner, when the scene classification model is a ResNet network, the step of inputting the environmental acoustic feature of the home scene into the first scene classification model to obtain the scene type information corresponding to the home scene can comprise: inputting the environmental acoustic feature of the home scene into the ResNet network to obtain a label value of the home scene; determining a second scene type weight corresponding to the label value; and taking the second scene type weight as the scene type information. Illustratively, the second scene type weight corresponding to the label value can be determined according to the corresponding relationship between the label value and the second scene type weight.

[0217] The scene acoustic style extractor can comprise a reference encoder, a first attention module and a full connection layer. The training style embedding information can be obtained by: inputting the acoustic feature of the home scene in the noisy acoustic environment into the reference encoder to obtain reference embedding information; and inputting the reference embedding information into the first attention module to obtain an attention weight. Correspondingly, the step of inputting the scene type information into the scene acoustic style extractor for fusion can comprise: inputting the attention weight and the second scene type weight into the full connection layer for weighting to obtain the training style embedding information. Please refer back to Figure 16 and Figure 17 One of the training processes of the scene acoustic style extractor can refer to the related descriptions of S210 and S227 to S229 above, which will not be repeated here.

[0218] In this embodiment, a second scene classification model training method can also be introduced. The acoustic feature of the home scene in the noisy acoustic environment is inputted into the above-mentioned second scene classification model to obtain the scene type indication information corresponding to the home scene, and the full connection layer is trained and fed back according to the scene type indication information to accelerate the convergence.

[0219] The voice synthesis device can be used to synthesize the to-be-synthesized sound prediction feature to obtain the to-be-output voice data. The voice synthesis device can be a vocoder. Correspondingly, the vocoder can be used as the voice synthesis device to synthesize the to-be-synthesized sound prediction feature output by the acoustic feature prediction network, and the to-be-synthesized sound prediction feature is inputted into the vocoder to obtain the to-be-output voice data.

[0220] Based on the above, the speech synthesis model in the embodiment can be trained based on both the aforementioned scene acoustic style extractor and acoustic feature prediction network. The speech synthesis model can also be trained based on the aforementioned first scene classification model, second scene classification model, scene acoustic style extractor, and acoustic feature prediction network. The training process and implementation principle of the first scene classification model, second scene classification model, scene acoustic style extractor, and acoustic feature prediction model applied in the speech synthesis model training method can be directly referred to the corresponding description in the aforementioned scene classification model training method, scene acoustic style extractor training method, and acoustic feature prediction model training method. Similar content will not be repeated here.

[0221] Through the training of the speech synthesis model, when the home scene type is a noisy type, the intelligent device can directly call the speech synthesis model. After inputting the acoustic features of the home scene in the noisy acoustic environment and the response text content information into the trained speech synthesis model, the speech synthesis model can output output response speech data containing the response text content information. The output response speech data is Lombard speech with good naturalness and clarity.

[0222] In order to more clearly set forth the scheme of synthesizing output response speech data in a noisy type in the embodiment of the application, the following scenario is taken as an example for illustration.

[0223] The intelligent device is located in a home scene and can perform voice interaction with the user, such as playing output response speech data. The intelligent device is loaded with a first scene classification model, a second scene classification model, a scene acoustic style extractor, and an acoustic feature prediction model.

[0224] The intelligent device determines whether the user has performed pre-configuration. If it is determined that the user has performed pre-configuration, the output response speech data is generated according to the pre-configuration. If it is determined that the user has not performed pre-configuration, the intelligent device receives the surrounding sound and determines whether the received sound is only the environmental sound of the home scene or the user's voice in a noisy acoustic environment.

[0225] If it is determined that the received sound is only the environmental sound of the home scene, the first scene classification model is called, the environmental sound of the home scene is input into the first scene classification model, the scene type information corresponding to the home scene such as a label value or a softmax probability value is obtained, and the home scene is classified.

[0226] If it is determined that the received sound is the user voice in the noisy acoustic environment, the smart device calls the scene acoustic style extractor and the second scene classification model. In a case where the first scene classification model does not obtain the scene type information corresponding to the home scene, the smart device inputs the user voice in the noisy acoustic environment into the scene acoustic style extractor and the second scene classification model to generate the corresponding scene style embedding information. In a case where the first scene classification model has obtained the scene type information corresponding to the home scene, the smart device inputs the user voice in the noisy acoustic environment and the scene type information corresponding to the home scene into the scene acoustic style extractor to generate the corresponding scene style embedding information.

[0227] Before the user voice in the noisy acoustic environment is input into the scene acoustic style extractor, a de-noising process can also be performed, and the de-noised sound is input into the scene acoustic style extractor to further improve the accuracy of the scene style embedding information extraction.

[0228] In a case where the smart device monitors that the user asks a question, it is determined whether the scene style embedding information corresponding to the home scene has been obtained. If it is determined that the scene style embedding information corresponding to the home scene has been obtained, an acoustic feature prediction model trained based on the to-be-output text information and the scene style embedding information is called, the response text content information and the scene style embedding information corresponding to the home scene are input into the acoustic feature prediction model to generate the predicted acoustic features of the Lombard speech in the corresponding home scene, and the output response voice data is synthesized through a vocoder. If it is determined that the scene style embedding information corresponding to the home scene has not been obtained, a Lombard speech generation model is called, the response text content information is input into the Lombard speech generation model to generate the predicted acoustic features of the Lombard speech in the corresponding home scene, and the output response voice data is synthesized through a vocoder. The synthesized output response voice data is broadcast by the smart device, so that the user can interact with the voice with the acoustic style in the corresponding home scene. The broadcast voice is the Lombard speech with good naturalness and clarity, so that the fluency of voice interaction can be ensured. It can be understood that the above voice interaction process can be repeatedly executed according to the user demand.

[0229] In order to perform the corresponding steps in the above embodiments and various possible manners, an implementation manner of a voice-based interaction device is given below. Please refer to Figure 33 , Figure 33 A functional module diagram of a voice-based interaction device 140 provided by the embodiment of the present application, which can be applied to Figure 1The intelligent device 100 is shown. It should be noted that the basic principle and technical effects of the voice-based interaction apparatus 140 provided in this embodiment are the same as those of the voice-based interaction method described above. For brief description, the voice-based interaction apparatus 140 provided in this embodiment is not described in some parts, and the corresponding content can be referred to the voice-based interaction method described above.

[0230] The information determination module 141 is configured to receive sound data of a home space, obtain a home scene type corresponding to the sound data, and generate a predicted acoustic feature through a voice prediction strategy corresponding to the home scene type when an interaction instruction is obtained, the predicted acoustic feature representing an acoustic style of the home scene type.

[0231] The response voice data synthesis module 142 is configured to synthesize the predicted acoustic feature to obtain output response voice data.

[0232] On the basis described above, the embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium comprising a computer program, the computer program being configured to control the intelligent device where the computer readable storage medium is located to execute the voice-based interaction method described above when the computer program is running.

[0233] In order to execute the corresponding steps in the above-described embodiments and various possible manners, an implementation manner of a scene classification model training apparatus is given below. Please refer to Figure 34 , Figure 34 A functional module diagram of a scene classification model training apparatus 160 provided in the embodiment of the present application is shown, and the scene classification model training apparatus 160 can be applied to Figure 1 The intelligent device 100 is shown. It should be noted that the basic principle and technical effects of the voice-based interaction apparatus 140 provided in this embodiment are the same as those of the voice-based interaction method described above. For brief description, the voice-based interaction apparatus 140 provided in this embodiment is not described in some parts, and the corresponding content can be referred to the voice-based interaction method described above. The scene classification model training apparatus 160 comprises an environment acoustic feature obtaining module 161 and a scene classification network training module 162.

[0234] The environment acoustic feature obtaining module 161 is configured to obtain environment acoustic features of a plurality of home scenes.

[0235] The scene classification network training module 162 is configured to input the environment acoustic features of the plurality of home scenes into a scene classification network for training to obtain scene type information corresponding to each of the home scenes.

[0236] To perform the corresponding steps in the above-mentioned embodiments and various possible manners, an implementation of a scene acoustic style extractor training apparatus is given below. Please refer to Figure 35 , Figure 35 A functional module diagram of a scene acoustic style extractor training apparatus 170 provided by an embodiment of the present application is provided. The scene acoustic style extractor training apparatus 170 can be applied to the intelligent device 100 as shown in Figure 1 It should be noted that the scene acoustic style extractor training apparatus 170 provided by the present embodiment has the same basic principles and technical effects as the above-mentioned scene acoustic style extractor training method embodiments. For brief description, the corresponding content in the above-mentioned scene acoustic style extractor training method embodiments can be referred to for the part not mentioned in the present embodiment. The scene acoustic style extractor training apparatus 170 includes a noisy environment acoustic feature obtaining module 171 and a scene acoustic style extractor training module 172.

[0237] The noisy environment acoustic feature obtaining module 171 is configured to obtain the acoustic features of the home scene in a noisy acoustic environment.

[0238] The scene acoustic style extractor training module 172 is configured to input the acoustic features of the home scene in the noisy acoustic environment into a scene acoustic style extractor to obtain training style embedding information corresponding to the home scene. The training style embedding information represents the scene acoustic style corresponding to the home scene.

[0239] To perform the corresponding steps in the above-mentioned embodiments and various possible manners, an implementation of an acoustic feature prediction model training apparatus is given below. Please refer to Figure 36 , Figure 36 A functional module diagram of an acoustic feature prediction model training apparatus 180 provided by an embodiment of the present application is provided. The acoustic feature prediction model training apparatus 180 can be applied to the intelligent device 100 as shown in Figure 1 It should be noted that the acoustic feature prediction model training apparatus 180 provided by the present embodiment has the same basic principles and technical effects as the above-mentioned acoustic feature prediction model training method embodiments. For brief description, the corresponding content in the above-mentioned acoustic feature prediction model training method embodiments can be referred to for the part not mentioned in the present embodiment. The acoustic feature prediction model training apparatus 180 includes a data obtaining module 181 and an acoustic feature prediction model training module 182.

[0240] The data obtaining module 181 is configured to obtain text information to be output and scene style embedding information corresponding to a home scene.

[0241] The acoustic feature prediction model training module 182 is configured to input the to-be-output text information and the training style embedding information into an acoustic feature prediction network, and obtain predicted features of to-be-synthesized sound corresponding to the home scene.

[0242] To perform the corresponding steps in the above-mentioned embodiments and various possible manners, an implementation of a speech synthesis model training apparatus is given below. Please refer to Figure 37 , Figure 37 A functional module diagram of a speech synthesis model training apparatus 190 provided by an embodiment of the present application is provided. The speech synthesis model training apparatus 190 can be applied to the intelligent device 100 as shown in the figure. It should be noted that the speech synthesis model training apparatus 190 provided by the present embodiment has the same basic principles and technical effects as the speech synthesis model training method embodiments described above. For brevity, some parts of the present embodiment are not mentioned in the above-mentioned speech synthesis model training method embodiments, and the corresponding content in the above-mentioned speech synthesis model training method embodiments can be referred to. The speech synthesis model training apparatus 190 includes an information obtaining module 191 and a to-be-output speech data synthesis module 192. Figure 1

[0243] The information obtaining module 191 is configured to obtain to-be-output text information and scene style embedding information corresponding to a home scene, and obtain predicted features of to-be-synthesized sound corresponding to the home scene according to the to-be-output text information and the scene style embedding information.

[0244] The to-be-output speech data synthesis module 192 is configured to synthesize the predicted features of to-be-synthesized sound, and obtain to-be-output speech data.

[0245] In the embodiment of the present application, the synthesized output response speech data is Lombard speech with good naturalness and clarity, which can ensure the smoothness of voice interaction and improve the user experience.

[0246] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can also be implemented in other manners. The above-described apparatus embodiments are only illustrative. For example, the flowchart and block diagram in the accompanying drawings show the possible implementation architectures, functions and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a segment or a portion of code which comprises one or more executable instructions for implementing the specified logic function. It should also be noted that each block in the flowchart or block diagram, as well as a combination of blocks in the flowchart or block diagram, can be implemented by dedicated hardware-based systems, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.

[0247] ​In addition, each functional module in various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0248] If the functions are realized in the form of software functional modules and sold or used as independent products, the functions can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0249] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A voice-based interaction method applied to smart devices, characterized in that, The method includes: Receives sound data from the surrounding home space; Obtain the home scene type corresponding to the sound data; When an interaction instruction is received, predictive acoustic features are generated using a voice prediction strategy corresponding to the home scene type. The step of generating predictive acoustic features using a voice prediction strategy corresponding to the home scene type includes: When the home scene type is noisy, the scene acoustic style extractor is invoked to determine the scene style embedding information corresponding to the noisy type; the scene style embedding information includes any one or a combination of any two or more of the following: pitch, sound intensity, speech rate, and emotion. Based on the response text content information and the scene style embedding information, a second predicted acoustic feature corresponding to the noise type is determined; The predicted acoustic features are synthesized to obtain output response speech data corresponding to the interaction instruction, wherein the predicted acoustic features include the second predicted acoustic features; The training method for the scene acoustic style extractor includes: Obtain the acoustic characteristics of a home environment in a noisy acoustic setting; The acoustic features of the home scene in a noisy acoustic environment are input into the reference encoder to obtain reference embedding information; The reference embedding information is input into the first attention module to obtain the attention weights; The attention weights are input into the fully connected layer to obtain training style embedding information.

2. The voice-based interaction method according to claim 1, characterized in that, The step of generating predictive acoustic features using a voice prediction strategy corresponding to the home scene type includes: When the home scene type is quiet, the first predicted acoustic feature corresponding to the quiet type is determined based on the response text content information.

3. The voice-based interaction method according to claim 1, characterized in that, The scene style embedding information represents the scene acoustic style corresponding to a noisy home scene.

4. The voice-based interaction method according to claim 1, characterized in that, The process of obtaining the home scene type corresponding to the sound data includes: When the smart device is in its initial state, the sound source type of the sound data of the home space where the smart device is located is obtained; the sound source type represents the acoustic characteristics of the sound sources contained in the home space. The home scene type is determined based on the sound source type.

5. The voice-based interaction method according to claim 1, characterized in that, The step of generating predicted acoustic features using a voice prediction strategy corresponding to the home scene type when an interaction instruction is received includes: When an interaction instruction is received, it is determined whether the home scene type has changed. If it has changed, predictive acoustic features are generated using a voice prediction strategy corresponding to the changed home scene type.

6. The voice-based interaction method according to claim 1 or 2, characterized in that, When the interaction instruction is user interaction request content information, the response text content information is the content information corresponding to the user interaction request content information; when the interaction instruction is a predefined instruction, the response text content information is the predefined content information.

7. A voice-based interactive device, applied to smart devices, characterized in that, The voice-based interactive device includes: The information determination module is used to receive sound data from the current home space, obtain the home scene type corresponding to the sound data, and when an interaction instruction is received, if the home scene type is noisy, call the scene acoustic style extractor to determine the scene style embedding information corresponding to the noisy type; the scene style embedding information includes any one or a combination of two or more of pitch, sound intensity, speech rate, and emotion; and determine the second predicted acoustic feature corresponding to the noisy type based on the response text content information and the scene style embedding information. A response speech data synthesis module is used to synthesize the predicted acoustic features to obtain output response speech data, wherein the predicted acoustic features include the second predicted acoustic features; The scene acoustic style extractor was trained in the following manner: Obtain the acoustic characteristics of a home environment in a noisy acoustic setting; The acoustic features of the home scene in a noisy acoustic environment are input into the reference encoder to obtain reference embedding information; The reference embedding information is input into the first attention module to obtain the attention weights; The attention weights are input into the fully connected layer to obtain training style embedding information.

8. A smart device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the voice-based interaction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program that, when executed, controls the smart device containing the computer-readable storage medium to perform the voice-based interaction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice outputting method, and terminal equipment

    CN109413277A

  • Voice synthesis method and device based on self-defined voice library

    CN109903748A

  • Terminal processing method and device based on sound analysis, storage medium and terminal

    CN111081275A

  • Phrase-based end-to-end text-to-speech (TTS) synthesis

    CN111681641A