Voice broadcast method, device, electronic device and storage medium
By obtaining environmental information in smart home appliances, and selecting simulated voices with different voice characteristics from the target user for broadcasting, the interference problem of analog pronunciation of smart home appliances is solved, and the human-computer interaction efficiency and user experience are improved.
Patent Information
- Application Number
- CN202010787023.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-08-07
AI Technical Summary
When smart home appliances use analog pronunciations to the same user pronunciations, they may affect the language communication between users and the clarity of voice broadcasts of smart home appliances, reducing human-computer interaction efficiency and user experience.
By obtaining environmental information, determining the target user, and selecting the target simulated voice for voice broadcast based on the different voice characteristics of the target user, thereby avoiding voice conflicts with the target user.
It improves the efficiency of human-computer interaction and user experience, and avoids voice broadcasts interfering with user communication.
Smart Images

Figure CN114093351B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart home appliances, and particularly relates to a voice broadcast method, device, electronic device, and storage medium. Background Art
[0002] Currently, the interaction functions of smart home appliances are becoming increasingly rich. In order to better achieve voice interaction between smart home appliances and users, many manufacturers have set up a voice function that simulates the pronunciation of users on smart home appliances, that is, the smart home appliances use the specified pronunciation of the user for voice broadcast to complete the voice interaction between the smart home appliances and the user, meeting the diverse needs of users.
[0003] However, in the actual use of such smart home appliances, when the simulated pronunciation used by the smart home appliances is the same as that of the user, the voice broadcast of the smart home appliances may affect the normal language communication between users, or affect the conversation between users and the clarity of the voice broadcast of the smart home appliances, resulting in the problem of reducing the human-computer interaction efficiency of the smart home appliances and affecting the user experience.
[0004] Correspondingly, there is a need in the art for a new voice broadcast method, device, electronic device, and storage medium to solve the above problems. Summary of the Invention
[0005] In order to solve the above problems in the prior art, that is, to solve the problem of reducing the human-computer interaction efficiency of smart home appliances and affecting the user experience, the present invention provides a voice broadcast method, device, electronic device, and storage medium.
[0006] According to the first aspect of the embodiments of the present invention, the present invention provides a voice broadcast method, which is applied to an electronic device and includes:
[0007] Obtain environmental information, and determine a target user according to the environmental information; determine a target simulated voice according to the target user, wherein the voice characteristics of the target simulated voice are different from those of the target user; use the target simulated voice for voice broadcast.
[0008] In a preferred technical solution of the above voice broadcast method, the environmental information includes environmental sound information. Determining a target user according to the environmental information includes:
[0009] Obtain the user voiceprint characteristics according to the environmental sound information; determine the target user according to the user voiceprint characteristics.
[0010] In a preferred technical solution of the above voice broadcast method, the environmental information further includes environmental image information. Determining a target user according to the environmental information includes:
[0011] Obtain the user's appearance features based on the environmental image information; determine the target user according to the user's appearance features.
[0012] In a preferred technical solution of the above voice broadcast method, determining the target simulated voice according to the target user includes:
[0013] Obtain the current simulated voice information, which is used to characterize the voice features of the currently used simulated voice; if the voice features corresponding to the current simulated voice information are the same as the voice features of the target user, determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
[0014] In a preferred technical solution of the above voice broadcast method, the alternative voice is a preset robot voice or a preset alternative simulated voice, where the voice features of the alternative simulated voice are different from the voice features of the target user.
[0015] In a preferred technical solution of the above voice broadcast method, the method further includes:
[0016] Obtain the user distance of the target user, where the user distance is used to characterize the distance between the target user and the electronic device; if the voice features corresponding to the current simulated voice information are the same as the voice features of the target user, and the user distance is less than a preset distance threshold, determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
[0017] In a preferred technical solution of the above voice broadcast method, after obtaining the environmental information, it further includes:
[0018] Determine a specific target user according to the environmental information; set a preset specific voice as the currently used simulated voice according to the specific target user.
[0019] According to the second aspect of the embodiments of the present invention, the present invention provides a voice broadcast device, which is applied to an electronic device, and the device includes:
[0020] An acquisition module, configured to acquire environmental information;
[0021] A first determination module, configured to determine a target user according to the environmental information;
[0022] A second determination module, configured to determine a target simulated voice according to the target user;
[0023] A voice broadcast module, configured to perform voice broadcast using the target simulated voice.
[0024] In the preferred technical solution of the above voice broadcast device, the environmental information includes environmental sound information, and the first determination module is specifically configured to:
[0025] Obtain the user voiceprint feature according to the environmental sound information; determine the target user according to the user voiceprint feature.
[0026] In the preferred technical solution of the above voice broadcast device, the environmental information further includes environmental image information, and the first determination module is specifically configured to:
[0027] Obtain the user appearance feature according to the environmental image information; determine the target user according to the user appearance feature.
[0028] In the preferred technical solution of the above voice broadcast device, the second determination module is specifically configured to:
[0029] Obtain the current simulated voice information, where the current simulated voice information is used to characterize the voice feature of the currently used simulated voice; if the voice feature corresponding to the current simulated voice information is the same as the voice feature of the target user, determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
[0030] In the preferred technical solution of the above voice broadcast device, the alternative voice is a preset robot voice or a preset alternative simulated voice, where the voice feature of the alternative simulated voice is different from the voice feature of the target user.
[0031] In the preferred technical solution of the above voice broadcast device, the second determination module is specifically configured to:
[0032] Obtain the user distance of the target user, where the user distance is used to characterize the distance between the target user and the electronic device; if the voice feature corresponding to the current simulated voice information is the same as the voice feature of the target user, and the user distance is less than a preset distance threshold, determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
[0033] In the preferred technical solution of the above voice broadcast device, the device further includes: a setting module, configured to:
[0034] After obtaining the environmental information, determine a specific target user according to the environmental information; set a preset specific voice as the currently used simulated voice according to the specific target user.
[0035] According to the third aspect of the embodiments of the present invention, the present invention provides an electronic device, including: a memory, a processor, and a computer program;
[0036] Wherein, the computer program is stored in the memory and is configured to be executed by the processor to perform the voice broadcast method according to any one of the first aspects of the embodiments of the present invention.
[0037] According to the fourth aspect of the embodiments of the present invention, the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the voice broadcast method according to any one of the first aspects of the embodiments of the present invention.
[0038] Those skilled in the art can understand that the voice broadcast method of the present invention obtains environmental information, determines a target user according to the environmental information, determines a target simulated voice according to the target user, and uses the target simulated voice for voice broadcast. Since the target simulated voice is determined by the target user and is different from the voice characteristics of the target user, therefore, by using the target simulated voice for voice broadcast, it is possible to avoid conflicts with the voice of the target user, affect the communication between users, and improve the efficiency of human-computer interaction and the user experience. Description of the Drawings
[0039] The preferred embodiments of the voice broadcast method, device, and electronic device of the present invention will be described below with reference to the accompanying drawings. The drawings are as follows:
[0040] Figure 1 It is an application scenario diagram of the voice broadcast method provided by an embodiment of the present application;
[0041] Figure 2 It is a flowchart of the voice broadcast method provided by an embodiment of the present application;
[0042] Figure 3 It is a flowchart of the voice broadcast method provided by another embodiment of the present application;
[0043] Figure 4 It is a schematic structural diagram of the voice broadcast device provided by an embodiment of the present application;
[0044] Figure 5 It is a schematic structural diagram of the voice broadcast device provided by another embodiment of the present application;
[0045] Figure 6 It is a schematic diagram of the electronic device provided by an embodiment of the present application. Detailed Embodiments
[0046] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention. Those skilled in the art can make adjustments according to needs to adapt to specific application scenarios. For example, although the voice broadcast method of the present invention is described in combination with an intelligent washing machine, this is not restrictive, and the voice broadcast method of the present invention can be configured for other devices with voice interaction requirements, such as intelligent refrigerators, smart TVs, and other devices.
[0047] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0048] First, the nouns involved in this application are explained:
[0049] 1) An intelligent household appliance device refers to a household appliance product formed after introducing microprocessor, sensor technology, and network communication technology into household appliance devices, with the characteristics of intelligent control, intelligent perception, and intelligent application. The operation process of intelligent household appliance devices often depends on the application and processing of modern technologies such as the Internet of Things, the Internet, and electronic chips. For example, an intelligent household appliance device can be connected to a cloud server to realize remote control and management of the intelligent household appliance device by users.
[0050] 2) "Multiple" means two or more, and other quantifiers are similar. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0051] 3) "Corresponding" can refer to an association relationship or a binding relationship. A corresponding to B means that there is an association relationship or a binding relationship between A and B.
[0052] Next, the application scenarios of the embodiments of this application are explained:
[0053] Figure 1 It is an application scenario diagram of the voice broadcast method provided by the embodiments of this application, as Figure 1As shown in the figure, the voice broadcast method provided by the embodiment of the present application can be applied to an electronic device, such as a smart washing machine. In the scenario provided by this embodiment, the smart washing machine is placed in an environment with multiple users. After being preset, the smart washing machine can simulate the voices of one or more users for voice broadcast, so as to realize the human-computer interaction between the user and the smart washing machine.
[0054] In the prior art, a general household smart washing machine is usually set in locations such as the living room and the bathroom. When the smart washing machine simulates the pronunciation of a certain family member to broadcast control instructions or prompt messages, it will make the user feel kind and interesting, thus improving the user experience. However, when there are multiple users in this environment, such as multiple family members, since the simulated voice of the smart washing machine is very similar to the voices of the users, during the daily conversations among family members or the interactive voices between the user and the smart washing machine, mutual interference will occur, resulting in the user being unable to clearly hear the voice broadcast sent by the smart washing machine; or the conversation and communication among the users will be affected, which will reduce the efficiency and experience of human-computer interaction.
[0055] The following will specifically describe the technical solution of the present application and how the technical solution of the present application solves the above technical problems with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0056] Figure 2 It is a flowchart of the voice broadcast method provided by an embodiment of the present application, applied to an electronic device, such as Figure 2 As shown in the figure, the voice broadcast method provided by this embodiment includes the following steps:
[0057] Step S101, obtain environmental information.
[0058] Exemplarily, the environmental information refers to the information obtained from the environment where the electronic device applying the voice broadcast method provided by the embodiment of the present application is located. The environmental information can be collected by the electronic device from the environment through its own set sensing units. For example, the electronic device obtains the sound and image in the environment through a sound sensor and an image sensor. Of course, in another possible implementation manner, the electronic device can obtain the already collected environmental information from other devices, such as other terminal devices, server devices, network devices, etc. The specific implementation manner of obtaining environmental information is not limited herein.
[0059] Further, the environmental information may be environmental sound information, i.e., information obtained by collecting sound signals in the environment where the electronic device is located, or environmental image information, i.e., information obtained by collecting image signals in the environment where the electronic device is located. Of course, it may also be other information, such as environmental vibration information, which can be set according to specific needs and will not be elaborated here one by one.
[0060] Step S102: Determine the target user according to the environmental information.
[0061] Since the user and the electronic device are in the same environment, for example, in the same room or a set of rooms, according to the environmental information, the specific user in the environment where the electronic device is located, i.e., the target user, can be determined. Here, the target user may be one user or multiple users, and no specific limitation is made here. Specifically, in a possible implementation manner, the environmental information includes environmental sound information. By parsing and classifying the collected environmental sound information, one or more corresponding users can be determined, and one or more of all users can be determined as the target user. Of course, in another possible implementation manner, image recognition can also be performed through environmental image information to confirm the target user, which will not be elaborated here.
[0062] Step S103: Determine the target simulated speech according to the target user; wherein, the speech feature of the target simulated speech is different from that of the target user.
[0063] Exemplarily, after determining the target user, the speech feature corresponding to the target user can be determined. Specifically, in a possible implementation manner, there is a fixed mapping relationship between the target user and its corresponding speech feature. According to this mapping relationship, the speech feature of the target user can be determined. For example, according to the identifier of the target user and a preset identifier mapping table, the speech feature corresponding to the identifier of the target user can be determined. In another possible implementation manner, the target user is determined through environmental sound information. Therefore, the environmental sound information includes the sound data corresponding to the target user, and the sound data can be parsed to correspondingly determine the speech feature of the target user, which will not be elaborated here.
[0064] After determining the speech feature of the target user, in order to avoid interference between the sound of the target user and the simulated speech emitted by the electronic device, the speech feature of the simulated speech emitted by the electronic device can be controlled to be different from that of the target user, so as to achieve this purpose. And this simulated speech is the target simulated speech.
[0065] Specifically, there are multiple methods for determining the target simulated voice. For example, according to a preset user identification mapping relationship table and the identification of the target user, a simulated voice different from the voice characteristics of the target user is selected from the alternative simulated voice libraries as the target simulated voice, or the robot voice can be directly determined as the target simulated voice, which can be set according to specific needs and scenarios and will not be specifically limited here.
[0066] Step S104, perform voice broadcast using the target simulated voice.
[0067] Exemplarily, the electronic device broadcasts the target simulated voice through a sound playback unit to achieve human-computer interaction with the user. Since the target simulated voice is different from the voice characteristics of the target user, it will not cause voice interference with the target user, thus improving the efficiency of human-computer interaction. Specifically, the simulated voice can be a preset voice library with user voice characteristics, and the target simulated voice can be a voice library with the voice of the target user. A number of voice messages are preset in this voice library, and different voice messages are played under different trigger conditions to complete the purpose of voice playback. Of course, it can be understood that the simulated voice can also be obtained through a conversion model capable of realizing voice feature migration, that is, this model can convert specific text content or information into a voice with user voice characteristics. Therefore, the target simulated voice can be obtained through the conversion model corresponding to the target user, and this model can be designed and trained through deep learning and other methods, and this process will not be elaborated here.
[0068] In this embodiment, by obtaining environmental information, determining the target user according to the environmental information, determining the target simulated voice according to the target user, and performing voice broadcast using the target simulated voice. Since the target simulated voice is determined by the target user and is different from the voice characteristics of the target user, voice broadcast through this target simulated voice can avoid conflicts with the voice of the target user, affect the communication between users, and improve the efficiency of human-computer interaction and the user experience.
[0069] Figure 3 The flowchart of the voice broadcast method provided by another embodiment of the present application is as Figure 3 shown. The voice broadcast method provided in this embodiment further refines steps S102 - S103 on the basis of the voice broadcast method provided in the embodiment shown in Figure 2 Therefore, the XXX provided in this embodiment includes the following steps:
[0070] Step S201, obtain the user's voiceprint feature according to the environmental sound information.
[0071] Exemplarily, the environmental sound information refers to the sound information in the environment where the electronic device applying the method provided in this embodiment is located. The environmental sound information can be collected by a sound sensor such as a microphone. More specifically, the environmental sound information can be obtained by one or more sound sensors collecting sound signals. Among them, if the environmental sound information is obtained after being collected by multiple sound sensors, it may also include signal processing steps such as sound superposition, mixing, and noise reduction, which will not be elaborated here one by one.
[0072] The environmental sound information can describe the sound situation in the environment, such as the sound generated by user conversations, the sound emitted by moving tables and chairs, the sound emitted by a television, etc. These sounds have different voiceprint characteristics, and they can be distinguished according to the voiceprint characteristics of different sounds. Among them, the user's voice, as the voice of interest, can be distinguished by the user's voiceprint characteristics.
[0073] Furthermore, there are many implementation methods for obtaining the user's voiceprint characteristics through the environmental sound information. For example, by analyzing the environmental sound information and identifying it through the frequency-domain components and / or time-domain waveform characteristics corresponding to different sounds, the user's voiceprint characteristics can be obtained. Among them, the specific methods for performing time-domain and frequency-domain analysis on sound signals are existing technologies in this field and will not be elaborated here.
[0074] Step S202, obtain the user's external characteristics according to the environmental image information.
[0075] Exemplarily, the environmental image information refers to the image information in the environment where the electronic device applying the method provided in this embodiment is located. The environmental image information can be collected by an image sensor such as a camera. More specifically, the environmental image information can be obtained by one or more image sensors collecting image signals. Among them, if the environmental image information is obtained after being collected by multiple image sensors, it may also include image processing steps such as image superposition, stitching, and noise reduction, which will not be elaborated here one by one.
[0076] The environmental image information can describe the specific objects in the environment, as well as the position and movement of objects relative to each other. The environmental image information is an image restoration of the external characteristics of specific objects. Since different objects have different external characteristics, they can be distinguished according to the external characteristics of different objects. Specifically, for example, by processing the environmental image information, the user's external characteristics can be obtained. Furthermore, based on the user's external characteristics, image recognition can be performed to distinguish the users in the room from furniture such as tables, chairs, and sofas. Among them, there are various implementation methods for processing the image information to obtain the user's external characteristics. For example, feature extraction and classification can be performed through a trained classifier or neural network model, which will not be elaborated here one by one.
[0077] Step S203: Determine the target user according to the user's voiceprint feature and the user's appearance feature.
[0078] Specifically, the target user may refer to a user who will be interfered by the simulated voice of the electronic device. For example, the pronunciation feature of user A is recorded in the intelligent washing machine, and the intelligent washing machine can emit a simulated voice with the pronunciation feature of user A. At this time, user A is the target user.
[0079] Exemplarily, the target user can be determined according to either the user's voiceprint feature or the user's appearance feature. Specifically, taking the intelligent washing machine as an example of the execution subject of the method provided in the embodiment of the present application, the intelligent washing machine compares the user's voiceprint feature with the voiceprint feature of the preset target user. For example, if the similarity between the two is higher than the preset similarity threshold, it is considered that the user's voiceprint feature matches the preset target user's voiceprint feature, that is, the user corresponding to the collected user's voiceprint feature is the target user. Another example is that the intelligent washing machine has previously recorded the appearance information of the target user. The appearance feature of the previously recorded appearance information is compared with the collected user's appearance feature. If the similarity between the two is higher than the preset similarity threshold, it is considered that the user's appearance feature matches the preset target user's appearance feature, that is, the user corresponding to the collected user's appearance feature is the target user.
[0080] However, since there are certain limitations in using the user's voiceprint feature or the user's appearance feature to identify the target user. For example, when the user is far away from the intelligent washing machine or there are other interfering sound sources, the accuracy of using the user's voiceprint feature to identify the target user is affected; when the user's body is blocked by an object, the accuracy of using the user's appearance feature to identify the target user is affected. Therefore, using both the user's voiceprint feature and the user's appearance feature to identify the target user can further improve the accuracy of target user identification. Of course, it can be understood that according to the specific usage scenario, the user's voiceprint feature or the user's appearance feature can also be used alone to identify the target user, which is not specifically limited here.
[0081] Step S204: Obtain the current simulated voice information, where the current simulated voice information is used to represent the voice feature of the currently used simulated voice.
[0082] Specifically, in the electronic device applying the method provided in the embodiment of the present invention, one or more preset simulated voices are stored, such as the simulated voice of user A, the simulated voice of user B, etc. The simulated voice can be realized through preset recording or through a simulated voice model, which is not specifically limited here. Different simulated voices have different voice features to distinguish the simulated voices.
[0083] Among them, the analog voice currently used by the electronic device is the current analog voice. Correspondingly, the voice features possessed by the current analog voice are the current analog voice information. There are various ways to obtain the current analog voice information. For example, the voiceprint features of the analog voice can be extracted, or the time and frequency domain features of the analog voice can be extracted, etc. It can be specifically set according to needs and will not be specifically limited here.
[0084] Step S205: Obtain the user distance of the target user. The user distance is used to represent the distance between the target user and the electronic device.
[0085] Specifically, according to the locations of different users, there is a certain distance between the electronic device and the user, that is, the user distance. The distance between the target user and the electronic device is the user distance of the target user.
[0086] There are various ways to obtain the user distance. In one possible implementation, based on the environmental image information, image recognition is performed on the target user, and according to the description information of the target user's location in the environmental image information, the distance between the target user and the electronic device is measured, so as to obtain the user distance of the target user. In another possible implementation, based on the environmental sound information, the sound of the target user is located, and according to the positioning result of the target user, the user distance of the target user is determined. More specifically, for example, the amplitude-phase difference of the user's sound can be obtained through multiple sound sensors to determine the user's location, and the specific implementation method will not be elaborated here.
[0087] Step S206A: If the voice features corresponding to the current analog voice information are the same as the voice features of the target user, and the user distance is less than the preset distance threshold, then determine the alternative voice as the target analog voice.
[0088] Step S206B: Otherwise, determine the currently used analog voice as the target analog voice.
[0089] The following will explain steps S206A and S206B in combination with a more specific usage scenario.
[0090] The execution subject of the method provided by the embodiments of the present application is an electronic device, such as a smart washing machine. The smart washing machine has a simulated voice function and can simulate the pronunciation of users to achieve more diverse functions. For example, in a specific application scenario, when the smart washing machine installed in the bedroom makes a voice announcement, for a young child in the bedroom, it will affect the child's sleep and rest, causing the child to be woken up or frightened and cry. By simulating the pronunciation characteristics of the child's mother, the smart washing machine can achieve the effect of soothing the child. However, at this time, if there are other users in the bedroom including the child's mother herself, due to the conflict between the simulated voice of the smart washing machine and the pronunciation of the child's mother herself, it will affect the normal communication between the child's mother herself and other users, as well as the voice interaction between other users and the smart washing machine.
[0091] In this scenario, exemplarily, according to the environmental information, a specific target user is determined, and according to the specific target user, a preset specific voice is set as the currently used simulated voice. As in the above example, according to the environmental information, the specific target user is determined to be the child. When a child is detected in the environment, then the specific voice, that is, the voice of the child's mother, is set as the currently used simulated voice to achieve the purpose of soothing the child.
[0092] In a possible implementation manner, if the voice characteristics corresponding to the current simulated voice information are the same as those of the target user, that is, in the above example, the smart washing machine determines that the currently used simulated voice is the same as the voice of the child's mother, then the alternative voice is determined as the target simulated voice, where the voice characteristics of the alternative simulated voice are different from those of the target user, such as a robot voice.
[0093] In another possible implementation manner, if the voice characteristics corresponding to the current simulated voice information are the same as those of the target user, and the distance of the user is less than a preset distance threshold, that is, in the above example, the smart washing machine determines that the currently used simulated voice is the same as the voice of the child's mother, and the child's mother is near the smart washing machine, then the alternative voice is determined as the target simulated voice, where the voice characteristics of the alternative simulated voice are different from those of the target user, such as a robot voice.
[0094] In this embodiment, by the voice characteristics of the target user and the user's location, it is judged whether there is a target user with the same voice characteristics in the environment, and whether the target user is nearby, so as to determine the best target simulated voice. Since the target simulated voice can be switched according to the specific location of the target user, the target simulated voice can, on the premise of realizing personalized functions such as soothing the child, avoid mutual interference caused by the same voice characteristics between the target user and the smart washing machine, and improve the efficiency of human-machine interaction and the user experience.
[0095] Figure 4The structural schematic diagram of the voice broadcast device provided by an embodiment of the present application is as follows: Figure 4 As shown, the voice broadcast device 3 provided by this embodiment includes:
[0096] An acquisition module 31, configured to acquire environmental information.
[0097] A first determination module 32, configured to determine a target user according to the environmental information.
[0098] A second determination module 33, configured to determine a target simulated voice according to the target user.
[0099] A voice broadcast module 34, configured to perform voice broadcast using the target simulated voice.
[0100] Among them, the acquisition module 31, the first determination module 32, the second determination module 33, and the voice broadcast module 34 are connected in sequence. The voice broadcast device 3 provided by this embodiment can execute the technical solutions of the method embodiment as shown in Figure 2 As shown. The implementation principle and technical effects are similar, and will not be elaborated here.
[0101] Figure 5 The structural schematic diagram of the voice broadcast device provided by another embodiment of the present application is as follows: Figure 5 As shown, on the basis of the voice broadcast device 3 shown, the voice broadcast device 4 provided by this embodiment further includes a setting module 41, where: Figure 4 In the preferred technical solution of the above voice broadcast device 4, the environmental information includes environmental sound information. The first determination module 32 is specifically configured to:
[0102] Acquire the user voiceprint feature according to the environmental sound information; determine the target user according to the user voiceprint feature.
[0103] In the preferred technical solution of the above voice broadcast device, the environmental information further includes environmental image information. The first determination module 32 is specifically configured to:
[0104] Acquire the user appearance feature according to the environmental image information; determine the target user according to the user appearance feature.
[0105] In the preferred technical solution of the above voice broadcast device, the second determination module 33 is specifically configured to:
[0106] Acquire the current simulated voice information, where the current simulated voice information is used to characterize the voice feature of the currently used simulated voice; if the voice feature corresponding to the current simulated voice information is the same as the voice feature of the target user, determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
[0107] In the preferred technical solution of the above voice broadcast device, the second determination module 33 is specifically configured to:
[0108] In the preferred technical solution of the above voice broadcast device, the alternative voice is a preset robot voice or a preset alternative simulated voice, wherein the voice characteristics of the alternative simulated voice are different from those of the target user.
[0109] In the preferred technical solution of the above voice broadcast device, the second determination module 33 is specifically configured to:
[0110] Obtain the user distance of the target user, where the user distance is used to represent the distance between the target user and the electronic device; if the voice characteristics corresponding to the current simulated voice information are the same as those of the target user, and the user distance is less than a preset distance threshold, then determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
[0111] In the preferred technical solution of the above voice broadcast device, the device further includes: a setting module 41, which is used to:
[0112] After obtaining the environmental information, determine a specific target user according to the environmental information; and set a preset specific voice as the currently used simulated voice according to the specific target user.
[0113] Among them, the acquisition module 31, the setting module 41, the first determination module 32, the second determination module 33, and the voice broadcast module 34 are connected in sequence. The voice broadcast device 4 provided in this embodiment can execute the technical solution of the method embodiment as Figure 3 shown, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0114] Figure 6 It is a schematic diagram of an electronic device provided in an embodiment of the present application. As Figure 6 shown, the electronic device 5 provided in this embodiment includes: a memory 51, a processor 52, and a computer program.
[0115] Among them, the computer program is stored in the memory 51 and is configured to be executed by the processor 52 to implement the Figures 2 - 3 voice broadcast method provided in any one of the embodiments corresponding to the present application.
[0116] Among them, the memory 51 and the processor 52 are connected through a bus 53.
[0117] For relevant descriptions, reference can be made to the relevant descriptions and effects corresponding to the steps in the Figures 2 - 3 corresponding embodiments, and no more elaboration will be made here.
[0118] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the Figures 2 - 3The voice broadcast method provided by any one of the corresponding embodiments.
[0119] Among them, the computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0120] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0121] Those skilled in the art will readily think of other implementation schemes of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0122] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0123] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A voice broadcast method, applied to an electronic device, characterized in that, Including: Obtain environmental information and determine a target user according to the environmental information; Determine a target simulated voice according to the target user; wherein, the voice feature of the target simulated voice is different from the voice feature of the target user; Use the target simulated voice for voice broadcast; Determine a target simulated voice according to the target user, including: Obtain current simulated voice information, where the current simulated voice information is used to characterize the voice feature of the currently used simulated voice; If the voice feature corresponding to the current simulated voice information is the same as the voice feature of the target user, determine the alternative voice as the target simulated voice; Otherwise, determine the currently used simulated voice as the target simulated voice.
2. The method according to claim 1, characterized in that, The environmental information includes environmental sound information. Determining a target user according to the environmental information includes: Obtain the user voiceprint feature according to the environmental sound information; Determine the target user according to the user voiceprint feature.
3. The method according to claim 2, characterized in that, The environmental information further includes environmental image information. Determining a target user according to the environmental information includes: Obtain the user appearance feature according to the environmental image information; Determine the target user according to the user appearance feature.
4. The method according to claim 1, characterized in that, The alternative voice is a preset robot voice or a preset alternative simulated voice, wherein the voice feature of the alternative simulated voice is different from the voice feature of the target user.
5. The method according to claim 1, characterized in that, The method further includes: Obtain the user distance of the target user, where the user distance is used to characterize the distance between the target user and the electronic device; If the voice feature corresponding to the current simulated voice information is the same as the voice feature of the target user and the user distance is less than a preset distance threshold, determine the alternative voice as the target simulated voice; Otherwise, determine the currently used simulated voice as the target simulated voice.
6. The method according to any one of claims 1-5, characterized in that, After obtaining the environmental information, it further includes: Determine a specific target user according to the environmental information; Set a preset specific voice as the currently used simulated voice according to the specific target user.
7. A voice broadcast device, characterized in that, The device includes: An obtaining module, configured to obtain environmental information; A first determining module, configured to determine a target user according to the environmental information; A second determining module, configured to determine a target simulated voice according to the target user; A voice broadcast module, configured to use the target simulated voice for voice broadcast; The second determining module is specifically configured to obtain current simulated voice information, where the current simulated voice information is used to characterize the voice feature of the currently used simulated voice; if the voice feature corresponding to the current simulated voice information is the same as the voice feature of the target user, determine the alternative voice as the target simulated voice; otherwise, determine the currently used simulated voice as the target simulated voice.
8. An electronic device, characterized in that, Including: A memory, a processor, and a computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the voice broadcast method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the voice broadcast method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the voice broadcast method according to any one of claims 1 to 6 above.
Citation Information
Patent Citations
Voice interaction method and device, storage medium and computer device
CN108470567A
Method for processing dialogue based on multiple user and apparatus for performing the same
KR1020150066882A