Simulated voice playback method, device, electronic device and storage medium

By obtaining environmental sound information, determining the target user and using simulated voice playback that is different from their voice features, the problem of low interaction efficiency caused by simulated voice playback of smart home appliances is solved and the user experience is improved.

CN114203148BActive Publication Date: 2025-08-19QINGDAO HAIER WASHING MASCH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010899170.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-31
Publication Date
2025-08-19
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

When existing smart home appliances simulate voice playback, they cause the user to experience less human-computer interaction efficiency and affect the user experience.

Method used

By obtaining environmental sound information, the target user is determined, the target simulated voice is determined based on the target user's voice content, and the simulated voice that is different from the target user's voice characteristics are played to avoid voice conflicts.

Benefits of technology

Improves human-computer interaction efficiency and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114203148B_ABST
    Figure CN114203148B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of smart home appliances, and specifically relates to a method, device, electronic device and storage medium for playing simulated voice. The present invention aims to solve the problem of low efficiency of human-computer interaction between users and smart home appliances in existing simulated voice playing methods, which affects the user experience. The present invention obtains environmental sound information and determines the target user based on the environmental sound information; determines the target simulated voice based on the voice content emitted by the target user, wherein the target simulated voice has different voice characteristics from the target user; and uses the target simulated voice for voice playback. Since the target simulated voice is determined by the voice content emitted by the target user, the simulated voice is determined as a target simulated voice with different voice characteristics from the target user based on the voice content and played, which can avoid conflicts with the target user's voice, affect communication between users, and improve the efficiency of human-computer interaction and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart home appliances, and in particular relates to a simulated voice playback method, device, electronic equipment and storage medium. Background Art

[0002] At present, with the increasing maturity of voice recognition technology, voice control and voice interaction functions have become more and more common functions in smart home appliances. In order to further improve the user experience and meet the user's personalized needs, smart home appliance manufacturers have set up simulated voice functions on smart home appliances. Smart home appliances can imitate the user's pronunciation and play voice, making the voice interaction process between users and smart home appliances more interesting and meeting the diverse needs of users.

[0003] However, in the actual use of such smart home appliances, when smart home appliances use simulated pronunciation to play voice, it often affects the user's auditory judgment, resulting in reduced efficiency of human-computer interaction between the user and the smart home appliances, affecting the user experience.

[0004] Accordingly, the art requires a new analog voice playback method, device, electronic device and storage medium to solve the above problems. Summary of the Invention

[0005] In order to solve the above-mentioned problems in the prior art, that is, to solve the problem that the human-computer interaction efficiency between existing users and smart home appliances is reduced, affecting the user experience, the present invention provides a simulated voice playback method, device, electronic device and storage medium.

[0006] According to a first aspect of an embodiment of the present invention, the present invention provides a simulated voice playback method, which is applied to an electronic device, comprising:

[0007] Acquire environmental sound information, and determine a target user based on the environmental sound information; determine a target simulated voice based on the voice content emitted by the target user; and use the target simulated voice for voice playback.

[0008] In the preferred technical solution of the above-mentioned simulated voice playback method, determining the target user according to the environmental sound information includes: obtaining the user voiceprint feature according to the environmental sound information; and determining the target user according to the user voiceprint feature.

[0009] In the preferred technical solution of the above-mentioned simulated voice playback method, the method further includes: obtaining environmental image information, and obtaining user location information based on the environmental image information; and extracting the voice content emitted by the target user from the environmental sound information based on the user location information.

[0010] In the preferred technical solution of the above-mentioned simulated voice playback method, the target simulated voice is determined according to the voice content emitted by the target user, including: performing semantic analysis according to the voice content emitted by the target user to determine the semantic type corresponding to the voice content; and determining the target simulated voice according to the semantic type.

[0011] In the preferred technical solution of the above-mentioned simulated voice playback method, the target simulated voice is determined according to the semantic type, including: if the semantic type is the first type, the alternative voice is determined as the target simulated voice; otherwise, the currently used simulated voice is determined as the target simulated voice; wherein the voice content corresponding to the first type is used to control the electronic device to execute instructions.

[0012] In the preferred technical solution of the above-mentioned simulated voice playing method, the alternative voice is a preset robot voice or a preset alternative simulated voice, wherein the voice characteristics of the alternative simulated voice are different from the voice characteristics of the target user.

[0013] In the preferred technical solution of the above-mentioned simulated voice playing method, after obtaining the environmental information, it also includes: determining a specific target user according to the environmental information; and setting a preset specific voice as the currently used simulated voice according to the specific target user.

[0014] According to a second aspect of an embodiment of the present invention, the present invention provides a simulated voice playback device, which is applied to an electronic device, and the device includes:

[0015] An acquisition module is used to acquire environmental sound information and determine a target user based on the environmental sound information;

[0016] A determination module, configured to determine a target simulated voice according to the voice content of the target user;

[0017] The playing module is used to play the target simulated voice.

[0018] In the preferred technical solution of the above-mentioned simulated voice playback device, when determining the target user based on the environmental sound information, the acquisition module is specifically used to: obtain the user voiceprint characteristics based on the environmental sound information; and determine the target user based on the user voiceprint characteristics.

[0019] In the preferred technical solution of the above-mentioned simulated voice playback device, the device further includes:

[0020] A positioning module is used to obtain environmental image information and obtain user location information based on the environmental image information;

[0021] The extraction module is used to extract the voice content uttered by the target user from the ambient sound information according to the user location information.

[0022] In the preferred technical solution of the above-mentioned simulated voice playback device, the determination module is specifically used to: perform semantic analysis based on the voice content emitted by the target user to determine the semantic type corresponding to the voice content; and determine the target simulated voice based on the semantic type.

[0023] In the preferred technical solution of the above-mentioned simulated voice playback device, when the determination module determines the target simulated voice according to the semantic type, it is specifically used to: if the semantic type is the first type, the alternative voice is determined as the target simulated voice; otherwise, the currently used simulated voice is determined as the target simulated voice; wherein the voice content corresponding to the first type is used to control the electronic device to execute instructions.

[0024] In a preferred technical solution of the above-mentioned simulated voice playback device, the candidate voice is a preset robot voice or a preset candidate simulated voice, wherein the voice characteristics of the candidate simulated voice are different from the voice characteristics of the target user.

[0025] In the preferred technical solution of the above-mentioned simulated voice playback device, the device further includes:

[0026] The setting module is used to determine a specific target user according to the environmental information; and set a preset specific voice as the currently used simulated voice according to the specific target user.

[0027] According to a third aspect of an embodiment of the present invention, the present invention provides an electronic device, comprising: a memory, a processor, and a computer program;

[0028] The computer program is stored in the memory and is configured to be executed by the processor to perform the simulated voice playback method as described in any one of the first aspects of the embodiments of the present invention.

[0029] According to the fourth aspect of the embodiments of the present invention, the present invention provides a computer-readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, they are used to implement the simulated voice playback method as described in any one of the first aspects of the embodiments of the present invention.

[0030] Those skilled in the art will understand that the simulated voice playback method of the present invention obtains environmental sound information, and determines the target user based on the environmental sound information; determines the target simulated voice based on the voice content emitted by the target user, wherein the target simulated voice is different from the voice characteristics of the target user; and uses the target simulated voice for voice playback. Since the target simulated voice is determined by the voice content emitted by the target user, when the voice content emitted by the target user is related to the content of the simulated voice emitted by the electronic device, the simulated voice is determined to be a target simulated voice that is different from the voice characteristics of the target user. Voice playback through the target simulated voice can avoid conflicts with the target user's voice, affect communication between users, and improve the efficiency of human-computer interaction and enhance user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The following describes preferred embodiments of the simulated voice playback method, device, and electronic device of the present invention with reference to the accompanying drawings.

[0032] Figure 1 A diagram of an application scenario of the simulated voice playback method provided in an embodiment of the present application;

[0033] Figure 2 A flowchart of a simulated voice playback method provided in one embodiment of the present application;

[0034] Figure 3 A schematic diagram of voice interaction between a user and an electronic device provided in an embodiment of the present application;

[0035] Figure 4 A flowchart of a simulated voice playback method provided in another embodiment of the present application;

[0036] Figure 5 A schematic diagram of voice content classification provided for one embodiment of the present application;

[0037] Figure 6 A schematic diagram of the structure of a simulated voice playback device provided in one embodiment of the present application;

[0038] Figure 7 A schematic structural diagram of a simulated voice playback device provided in another embodiment of the present application;

[0039] Figure 8 A schematic diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0040] First of all, it should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make adjustments as needed to adapt to specific applications. For example, although the simulated voice playback method of the present invention is described in conjunction with a smart washing machine, this is not limiting. Other devices with voice interaction requirements can be configured with the simulated voice playback method of the present invention, such as smart refrigerators, smart TVs and other devices.

[0041] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "connected" and "connection" should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integral connection; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0042] First, let’s explain the terms involved in this application:

[0043] 1) Smart home appliances refer to home appliances that are formed by introducing microprocessors, sensor technology, and network communication technology into home appliances. They have the characteristics of intelligent control, intelligent perception, and intelligent application. The operation of smart home appliances often relies on the application and processing of modern technologies such as the Internet of Things, the Internet, and electronic chips. For example, smart home appliances can be connected to cloud servers to enable users to remotely control and manage smart home appliances.

[0044] 2) "Multiple" refers to two or more, and other quantifiers are similar. "And / or" describes the relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.

[0045] 3) “Correspondence” may refer to an association relationship or a binding relationship. A and B correspond to each other, which means that there is an association relationship or a binding relationship between A and B.

[0046] The following explains the application scenarios of the embodiments of the present application:

[0047] Figure 1 An application scenario diagram of the simulated voice playback method provided in the embodiment of the present application, such as Figure 1As shown, the simulated voice playback method provided in the embodiment of the present application can be applied to an electronic device, such as a smart washing machine. In the scenario provided in this embodiment, the smart washing machine is placed in a multi-user environment including a target user and other users. After pre-setting, the smart washing machine can simulate the target user's voice for voice playback.

[0048] In this application scenario, a smart washing machine simulates the target user's pronunciation to play control commands or prompts, creating a sense of familiarity and interest for users, thereby improving their user experience. However, because the simulated pronunciation used by the smart appliance is similar to that of the target user, when the target user is present, other users cannot distinguish whether the voice they hear is from the smart appliance or the target user. This reduces the efficiency of interaction between other users and the smart appliance, affecting the user experience.

[0049] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0050] Figure 2 A flowchart of a method for simulating voice playback provided in one embodiment of the present application is applied to electronic devices, such as Figure 2 As shown, the simulated voice playback method provided in this embodiment includes the following steps:

[0051] Step S101: Acquire environmental sound information.

[0052] Exemplarily, environmental sound information refers to sound information obtained from the environment in which an electronic device that applies the simulated voice playback method provided in an embodiment of the present application is located. The environmental information can be collected and obtained from the environment by the electronic device through its own sensors. For example, the electronic device obtains the sound signal in the environment through a sound sensor or a vibration sensor. Of course, in another possible implementation method, the electronic device can obtain the collected environmental sound information from other devices, such as other terminal devices, server devices, network devices, etc. The specific implementation method of obtaining the environmental sound information is not limited here.

[0053] Furthermore, ambient sound information can be obtained by one or more directional sound sensors installed on or in communication with the electronic device. In one possible implementation, there are multiple sound sensors, each pointing in different directions and / or installed in different locations. By collecting ambient sound information through these multiple sound sensors, the sound source can be located and identified.

[0054] Step S102: Determine the target user based on the environmental information.

[0055] Since the user and the electronic device are in the same environment, such as in the same room or in a set of rooms, voiceprint recognition can be performed based on the ambient sound information to determine the specific user making the sound in the environment where the electronic device is located, that is, the target user. The target user can be one user or multiple users, which is not specifically limited here. Specifically, based on the collected ambient sound information, the sound information is parsed and classified to determine the corresponding one or more users, and one or more of all users can be determined as target users, which will not be described in detail here.

[0056] Step S103: determining a target simulated voice according to the voice content of the target user, wherein the target simulated voice has different voice features from the target user.

[0057] For example, by parsing and recognizing the collected speech of the target user, the specific speech content of the target user can be obtained. The speech content refers to the content of the voice conversation between the target user and the electronic device, or between target users. For example, a conversation between target users might be something like, "Hi, why are you back so early today?" or a conversation between a target user and an electronic device might be something like, "Set the temperature to 26 degrees Celsius." The electronic device then performs semantic analysis and classification on this content to identify differences between conversations between target users and between target users and the electronic device.

[0058] Then, the electronic device can determine the corresponding target simulated voice according to the difference. Figure 3 A schematic diagram of voice interaction between a user and an electronic device provided in an embodiment of the present application is shown in FIG. Figure 3As shown, for example, when the electronic device determines that the target user is nearby and the target user is having a conversation with another user, since there is no need for voice interaction between the user and the electronic device at this time, it can be considered that the target user and the other user are in a communication state. Users in this communication state are not easily interfered with by information from external electronic devices, and the user can easily distinguish whether the voice heard is from the target user or from the electronic device. Therefore, the current simulated voice can continue to be set to the target user's voice; when the electronic device determines that the current target user is present and the target user's voice content is an instruction to instruct the electronic device to work, it can be considered that the target user and the other user are in a non-communication state, for example, the target user and the other user are doing their own things. The attention of users in the non-communication state is not as focused as in the communication state. Therefore, when they suddenly hear voice information from an external electronic device, they cannot distinguish whether the response sound they hear is from the target user or from the electronic device. In order to avoid the response sound characteristics of the electronic device being the same as the target user's voice characteristics, which affects the auditory understanding of other users, the target simulated voice is set to a simulated voice different from the target user, such as a robot voice, to improve the recognition of the voice and improve the efficiency of the interaction between the user and the electronic device.

[0059] Specifically, there are many methods to determine the target simulated voice. For example, based on a preset user identification mapping relationship table and the identification of the target user, a simulated voice with different voice characteristics from the target user is selected from the alternative simulated voice library as the target simulated voice, or the robot voice is directly determined as the target simulated voice. It can be set according to specific needs and scenarios, and is not specifically limited here.

[0060] Step S104: Use the target simulated voice to play the voice.

[0061] Exemplarily, the electronic device plays the target simulated voice through the sound playback unit to achieve human-computer interaction with the user. Since the target simulated voice is different from the voice characteristics of the target user, it will not cause voice interference with the target user, thereby improving the efficiency of human-computer interaction. Specifically, the simulated voice can be a preset voice library with user voice characteristics, and the target simulated voice can be a voice library with the target user voice. The voice library is preset with several voice information. Under different trigger conditions, different voice information is played to achieve the purpose of voice playback. Of course, it is understandable that the simulated voice can also be obtained through a conversion model that can realize voice feature transfer, that is, the model can convert specific text content or information into sound with user voice characteristics. Therefore, the target simulated voice can be obtained through a conversion model corresponding to the target user. The model can be designed and trained by deep learning and other methods. This process will not be described in detail.

[0062] In this embodiment, the target user is determined based on the environmental sound information obtained and the target user is determined based on the environmental sound information; the target simulated voice is determined based on the voice content emitted by the target user, wherein the target simulated voice has different voice characteristics from the target user; and the target simulated voice is used for voice playback. Since the target simulated voice is determined by the voice content emitted by the target user, when the voice content emitted by the target user is related to the content of the simulated voice emitted by the electronic device, the simulated voice is determined to be the target simulated voice with different voice characteristics from the target user. Voice playback using the target simulated voice can avoid conflicts with the target user's voice, affect communication between users, improve the efficiency of human-computer interaction, and enhance the user experience.

[0063] Figure 4 A flowchart of a method for simulating voice playback provided in another embodiment of the present application is shown in FIG. Figure 4 As shown, the simulated voice playing method provided by this embodiment is Figure 2 Based on the simulated voice playback method provided in the illustrated embodiment, steps S102-S103 are further refined. The simulated voice playback method provided in this embodiment includes the following steps:

[0064] Step S201: Acquire user voiceprint features based on environmental sound information.

[0065] Exemplarily, ambient sound information refers to the sound information in the environment of the electronic device to which the method provided in this embodiment is applied. This ambient sound information can be acquired by a sound sensor, such as a microphone. More specifically, the ambient sound information can be acquired by one or more sound sensors acquiring sound signals. If the ambient sound information is acquired by multiple sound sensors, signal processing steps such as sound superposition, mixing, and denoising may also be included, which will not be detailed here.

[0066] Ambient sound information can describe the sound conditions in the environment, such as the sounds of user conversations, the sounds of moving tables and chairs, and the sound of the television. These sounds have different voiceprint characteristics, and they can be distinguished based on their voiceprint characteristics. Among them, the user's voice, as the sound of interest, can be distinguished based on the user's voiceprint characteristics.

[0067] Furthermore, there are many ways to obtain user voiceprint features through environmental sound information, such as by parsing the environmental sound information and identifying the frequency domain components and / or time domain waveform features corresponding to different sounds, thereby obtaining user voiceprint features. The specific method of performing time domain and frequency domain analysis on sound signals is the existing technology in this field and will not be repeated here.

[0068] Step S202: Determine the target user based on the user's voiceprint features.

[0069] The target user may refer to a user whose pronunciation may be interfered with by the simulated voice of the electronic device. For example, the pronunciation characteristics of user A are recorded into a smart washing machine, and the smart washing machine can emit simulated voice with the pronunciation characteristics of user A. At this time, user A is the target user. Taking the execution subject of the method provided in the embodiment of the present application as an example of a smart washing machine, the smart washing machine compares the user's voiceprint characteristics with the voiceprint characteristics of the preset target user. If the similarity between the two is higher than the preset similarity threshold, it is considered that the user's voiceprint characteristics match the preset target user's voiceprint characteristics, that is, the user corresponding to the collected user voiceprint characteristics is the target user. Otherwise, it is considered that the user is not the target user.

[0070] Step S203: Acquire environmental image information, and acquire user location information based on the environmental image information.

[0071] Exemplarily, environmental image information refers to image information of the environment in which the electronic device to which the method provided by this embodiment is applied is located. This environmental image information can be acquired by, for example, an image sensor such as a camera. More specifically, the environmental image information can be acquired by acquiring image signals from one or more image sensors. If the environmental image information is acquired by multiple image sensors, image processing steps such as image overlay, splicing, and denoising may also be included, which will not be detailed here.

[0072] Based on the environmental image information, the user's location information within the coverage range of the image sensor of the electronic device can be determined. Specifically, by performing image feature recognition on the environmental image information, the "user" in the environmental image information can be distinguished from other objects and extracted, and then the user location information corresponding to each user can be determined. Exemplarily, the location information can be the distance and direction between the user and the electronic device, which can be represented by a vector. The location information can also be the position coordinate information of the user within the coverage range of the image sensor of the entire electronic device. Here, the specific implementation form of the user location information is not limited, as long as the user location information can realize the positioning of the user.

[0073] In one possible implementation, while obtaining user location information based on environmental image information, the target user can be identified from multiple users based on the preset appearance information of the target user, and then the user location information corresponding to the target user can be determined.

[0074] Step S204: extracting the voice content of the target user from the ambient sound information according to the user location information.

[0075] As described in the above steps, ambient sound information includes not only the target user's voice but also the voices of other users. Directly extracting the voice from the ambient sound information is ineffective. Using user location information, we can determine the location of the target user and other users. Then, we weight the ambient sounds received by multiple sound sensors from different directions based on the user's location, giving a higher weight to the ambient sound information at the target user's location. This allows us to extract the target user's voice information. Then, we perform sound recognition on the target user's voice information to obtain the target user's voice content.

[0076] In the embodiments of the present application, all voice information within the ambient sound information is filtered using user location information to extract the target user's voice information and the corresponding voice content. Due to the complex ambient sound environment in which electronic devices operate, directly extracting voice content based on voiceprint information is ineffective and difficult to meet the needs of subsequent voice content analysis. However, by weighting the ambient sound information collected by different sound sensors based on user location information, the target user's voice energy level can be amplified, thereby improving the accuracy of extracting the target user's voice content.

[0077] Step S205 : performing semantic analysis based on the speech content uttered by the target user to determine the semantic type corresponding to the speech content.

[0078] According to the meaning of the speech content uttered by the target user, the speech content can be classified to determine the corresponding semantic type. Figure 5 A schematic diagram of voice content classification provided in one embodiment of the present application is shown in FIG. Figure 5 As shown, the voice content sent by the target user is classified. When the target user sends a voice command to the electronic device, such voice content is, for example, "adjust the temperature to 25 degrees" or "dry the clothes", and the semantic type corresponding to such voice content is the first type. When the target user is communicating with other users, there is no voice interaction with the electronic device. Such voice content is, for example, "Hi, why are you back so early today?" or "He didn't tell me about this yesterday", etc. The semantic type corresponding to such voice content is the second type. Among them, the voice content corresponding to the first type is used to control the electronic device to execute instructions.

[0079] Specifically, there are multiple ways to perform semantic analysis based on the speech content and determine the semantic type corresponding to the speech content. For example, the speech content can be matched based on preset instruction information or keywords. If the speech content can match the instruction information or keywords, the corresponding semantic type is determined to be the first type; if not, the corresponding semantic type is determined to be the second type. For another example, the speech content can be recognized based on a semantic recognition model based on a neural network that has been trained to convergence, and the semantic type is determined based on the output of the semantic recognition model. This can be set according to specific needs and is not specifically limited here.

[0080] Step S206: determining the target simulated speech according to the semantic type.

[0081] For example, if the semantic type is the first type, the candidate voice is determined as the target simulated voice; otherwise, the currently used simulated voice is determined as the target simulated voice. The candidate voice is a preset robot voice or a preset candidate simulated voice, and the voice characteristics of the candidate simulated voice are different from the voice characteristics of the target user.

[0082] Specifically, if the semantic type is the first type, it means that the electronic device determines that the current target user is present, and the target user's voice content is an instruction to instruct the electronic device to work. At this time, the target user and other users are in a non-communication state, and the attention of other users is not focused on the target user. Therefore, if the electronic device emits the same sound as the target user, it is likely to cause other users to mishear. In actual use, when other users mishear and think that the target user is talking to them, they will immediately make a voice confirmation to the target user. At this time, the target user is interacting with the electronic device by voice and has no time to respond to other users, or responds to other users in a hurry, causing the electronic device to receive erroneous voice instructions, affecting the interaction efficiency and user experience between the user and the electronic device. Therefore, if the semantic type is the first type, the alternative voice with different voice characteristics from the target user is confirmed as the target simulated voice, thereby avoiding the problem caused by the simulated voice emitted by the electronic device having the same voice characteristics as the target user's voice.

[0083] Correspondingly, if the semantic type is the second type, this means that the electronic device determines that the target user is nearby and the target user is having a conversation with other users. At this time, the target user and other users are in a communication state. Users in this communication state are not easily interfered with by information from external electronic devices. The user can easily distinguish whether the voice heard is emitted by the target user or by the electronic device. Therefore, the target user's voice can continue to be set as the target simulated voice, which makes the user feel friendly and interesting and improves the user's usage experience.

[0084] Step S207: Use the target simulated voice to play the voice.

[0085] In this embodiment, the implementation of step S207 is the same as Figure 2 The implementation method of step S104 in the illustrated embodiments is the same. Please refer to the description of the specific implementation method and related technical effects of Figure S104, which will not be repeated here.

[0086] Figure 6 A schematic diagram of the structure of a simulated voice playback device provided in one embodiment of the present application is shown as follows: Figure 6 As shown, the simulated voice playback device 3 provided in this embodiment includes:

[0087] The acquisition module 31 is used to acquire environmental sound information and determine the target user based on the environmental sound information.

[0088] The determination module 32 is configured to determine a target simulated voice according to the voice content uttered by the target user.

[0089] The playing module 33 is used to play the target simulated voice.

[0090] The acquisition module 31, the determination module 32, and the determination module 33 are connected in sequence. The simulated voice playback device 3 provided in this embodiment can execute the following steps: Figure 2 The technical solution of the method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.

[0091] Figure 7 A structural diagram of a simulated voice playback device provided in another embodiment of the present application is shown as follows: Figure 7 As shown, the simulated voice playing device 4 provided in this embodiment is Figure 6 The simulated voice playback device 3 shown in the figure further includes a positioning module 41, an extraction module 42, and a setting module 43, wherein:

[0092] In the preferred technical solution of the above-mentioned simulated voice playback device, when determining the target user according to the ambient sound information, the acquisition module 31 is specifically configured to: obtain the user voiceprint feature according to the ambient sound information; and determine the target user according to the user voiceprint feature.

[0093] In the preferred technical solution of the above-mentioned simulated voice playback device, the device further includes:

[0094] The positioning module 41 is used to obtain environmental image information and obtain user location information based on the environmental image information;

[0095] The extraction module 42 is configured to extract the speech content uttered by the target user from the ambient sound information according to the user location information.

[0096] In the preferred technical solution of the above-mentioned simulated voice playback device, the determination module 32 is specifically used to: perform semantic analysis based on the voice content emitted by the target user to determine the semantic type corresponding to the voice content; and determine the target simulated voice based on the semantic type.

[0097] In the preferred technical solution of the above-mentioned simulated voice playback device, the determination module 32 is specifically used to determine the target simulated voice according to the semantic type: if the semantic type is the first type, the alternative voice is determined as the target simulated voice; otherwise, the currently used simulated voice is determined as the target simulated voice; wherein the voice content corresponding to the first type is used to control the electronic device to execute instructions.

[0098] In the preferred technical solution of the above-mentioned simulated voice playback device, the alternative voice is a preset robot voice or a preset alternative simulated voice, wherein the voice characteristics of the alternative simulated voice are different from the voice characteristics of the target user.

[0099] In the preferred technical solution of the above-mentioned simulated voice playback device, the device further includes:

[0100] The setting module 43 is configured to determine a specific target user according to the environmental information; and set a preset specific voice as the currently used simulated voice according to the specific target user.

[0101] The simulated voice playing device 4 provided in this embodiment can perform the following operations: Figure 3-Figure 5 The technical solution of the method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.

[0102] Figure 8 FIG8 is a schematic diagram of an electronic device provided in accordance with an embodiment of the present application. As shown in FIG8 , the electronic device 5 provided in accordance with the present embodiment includes a memory 51 , a processor 52 and a computer program.

[0103] The computer program is stored in the memory 51 and is configured to be executed by the processor 52 to implement the present application. Figure 2-Figure 5 The simulated voice playback method provided in any one of the corresponding embodiments.

[0104] The memory 51 and the processor 52 are connected via a bus 53 .

[0105] For related instructions, please refer to Figure 2-Figure 5 The relevant descriptions and effects corresponding to the steps in the corresponding embodiments can be understood, and no further details are given here.

[0106] One embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the present application. Figure 2-Figure 5 The simulated voice playback method provided in any one of the corresponding embodiments.

[0107] Among them, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0108] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0109] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0110] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

[0111] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. A method for simulating voice playback, applied to electronic equipment, characterized in that: include: Acquiring environmental sound information, and determining a target user based on the environmental sound information; Performing semantic analysis on the speech content sent by the target user to determine the semantic type corresponding to the speech content; If the semantic type is the first type, determining the candidate voice as the target simulated voice, the voice characteristics of the candidate voice are different from the voice characteristics of the target user, and the voice content corresponding to the first type is used to control the electronic device to execute instructions; If the semantic type is the second type, a voice having the same characteristics as the voice of the target user is determined as the target simulated voice, and the voice content corresponding to the second type indicates that the target user is having a conversation with another user; The target simulated voice is used for voice playback.

2. The method according to claim 1, characterized in that Determining a target user according to the ambient sound information includes: Acquiring user voiceprint features based on the environmental sound information; Determine the target user based on the user's voiceprint features.

3. The method according to claim 1, characterized in that The method further comprises: Acquire environmental image information, and acquire user location information based on the environmental image information; The voice content uttered by the target user is extracted from the ambient sound information according to the user location information.

4. The method according to claim 1, wherein The alternative voice is a preset robot voice or a preset alternative simulated voice.

5. The method according to any one of claims 1 to 4, characterized in that After obtaining the environment information, it also includes: determining a specific target user based on the environmental information; According to the specific target user, the preset specific voice is set as the currently used simulated voice.

6. A simulated voice playback device, applied to electronic equipment, characterized in that: The device comprises: An acquisition module is used to acquire environmental sound information and determine a target user based on the environmental sound information; Determine the module for Perform semantic analysis on the voice content emitted by the target user to determine the semantic type corresponding to the voice content; if the semantic type is the first type, determine the candidate voice as the target simulated voice, the voice characteristics of the candidate voice are different from the voice characteristics of the target user, and the voice content corresponding to the first type is used to control the electronic device to execute instructions; if the semantic type is the second type, determine the voice with the same voice characteristics as the target user as the target simulated voice, and the voice content corresponding to the second type indicates that the target user is having a conversation with another user; The playing module is used to play the target simulated voice.

7. An electronic device, characterized in that: include: memory, processors and computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the simulated voice playback method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the simulated voice playback method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice feedback method and device, storage medium and electronic device

    CN108806699A

  • Voice broadcast method and device for electric equipment, air conditioner and storage medium

    CN111415642A