A control method and apparatus of an electronic device, and an electronic device

By automatically judging environmental noise and user voice characteristics, and combining sound and image acquisition, intelligent voice separation of electronic devices is realized, which solves the problem of inaccurate voice assistant response in noisy environments, reduces power consumption and improves user experience.

CN116189659BActive Publication Date: 2026-03-20HUAWEI DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In noisy environments, voice assistants on electronic devices struggle to respond correctly to user voice commands, leading to a degraded user experience. Furthermore, existing deep learning models for speech separation technologies consume excessive memory resources, increasing power consumption and impacting the response speed of other applications.

Method used

The system collects sound using sound acquisition equipment, determines the decibel value and signal energy ratio of the sound, and automatically turns the voice separation function on or off. Combined with vehicle status and image acquisition devices, it adopts a multi-mode voice separation method based on voiceprint and audio-visual data to achieve the separation of the target user's voice.

Benefits of technology

It reduces the power consumption of electronic devices, avoids excessive consumption of memory resources, improves user experience and driving safety, and simplifies user operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189659B_ABST
    Figure CN116189659B_ABST
Patent Text Reader

Abstract

The application provides a control method and device of an electronic device and the electronic device, and relates to the technical field of terminals. In the method, the electronic device is configured with a voice separation function for separating out a target user's voice. The voice separation function can be automatically selected to be turned on or off on the electronic device according to noise conditions in an environment and the number of users who speak in the environment, thereby avoiding excessive consumption of memory resources caused by the voice separation function being always turned on on the electronic device, reducing power consumption of the electronic device, and avoiding cumbersome operations of manually turning on or off the voice separation function, and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of terminal, and particularly relates to a control method and device of electronic equipment and electronic equipment. BACKGROUND

[0002] At present, a voice assistant is often configured on an electronic equipment. Through the voice assistant, a user can interact with the electronic equipment in a voice interaction mode, such as starting a navigation service, playing a song and the like. However, in a case where the environment is noisy, the voice assistant on the electronic equipment often cannot correctly respond to the voice of the user, such as not responding or responding incorrectly. For example, when the electronic equipment is a car machine on a vehicle, under the influence of in-vehicle sound, loud speaking of in-vehicle personnel, tire noise, wind noise and the like, the voice assistant on the car machine often cannot correctly respond to the voice of the user, which affects the user experience.

[0003] In the related art, the electronic equipment can separate the voice of the user by using a voice separation technology, so that the voice of the user can be correctly responded. However, when the voice separation technology is used, a deep learning scheme is often used, and the model required in the deep learning scheme is relatively large, which causes the memory resource of the electronic equipment to be excessively consumed when the voice separation function is always in an open state, increases the power consumption of the electronic equipment, and affects the response speed of other application services on the electronic equipment. If the voice separation function on the electronic equipment needs to be manually opened or closed, the operation is relatively cumbersome, which affects the user experience. SUMMARY

[0004] The present application provides a control method and device of electronic equipment, electronic equipment, computer storage readable storage medium and computer program product, which can automatically open or close the voice separation function on the electronic equipment, thereby avoiding the situation that the memory resource of the electronic equipment is excessively consumed due to the voice separation function on the electronic equipment being always on, reducing the power consumption of the electronic equipment, and avoiding the cumbersome operation of manual opening or closing by the user, and improving the user experience.

[0005] In a first aspect, the present application provides a control method of an electronic device, the method comprising: collecting sound of a preset time length through a sound collecting device; determining a decibel value of the sound of the preset time length; when the decibel value of the sound of the preset time length obtained is greater than a first threshold value, determining whether a proportion of energy of a first type of signal in the sound is less than a second threshold value; when the proportion of energy of the first type of signal in the sound is less than the second threshold value, determining whether multiple fundamental frequencies exist in the sound; when the multiple fundamental frequencies exist in the sound, starting a target function on the electronic device, the target function being used to separate sound of a target user from the sound obtained. Thus, whether to start a voice separation function on the electronic device is automatically selected according to a noise condition in the environment, thereby avoiding a situation that a memory resource is excessively consumed due to the voice separation function on the electronic device being always started, reducing power consumption of the electronic device, and avoiding a cumbersome operation of manually starting or stopping, thereby improving user experience.

[0006] For example, the sound collecting device can be a sound pickup device (such as a microphone), which can be integrated on the electronic device or arranged independently of the electronic device. The sound obtained can include a first type of signal and a second type of signal, wherein the first type of signal and the second type of signal are in different frequency ranges, or the first type of signal and the second type of signal are in different frequency ranges and have different amplitude values of sound intensity. For example, the first type of signal can include a noise signal, and the second type of signal can include a noise signal and / or a sound signal emitted by a user.

[0007] In a possible implementation, the electronic device is configured in a vehicle, and before determining whether the proportion of energy of the first type of signal in the sound is less than the second threshold value, the method further comprises: determining whether a sound playing device on the vehicle is in a closed state; if the sound playing device on the vehicle is in the closed state, determining whether the proportion of energy of the first type of signal in the sound is less than the second threshold value; and if the sound playing device on the vehicle is in an open state, starting the target function on the electronic device. Thus, whether to determine the proportion of energy of the first type of signal is determined based on sound emitted by the sound playing device in the vehicle, thereby improving the efficiency of the determination. Meanwhile, the target function is directly started when the sound playing device in the vehicle is in the open state, thereby reducing the influence of the sound emitted by the sound playing device in the vehicle on the recognition of the voice of the user, and improving the efficiency of starting the target function.

[0008] In a possible implementation, the method further includes: the electronic device is configured in the vehicle, and before determining whether the proportion of the energy of the first type of signal in the environmental sound is less than the second threshold, the method further includes: determining whether a window on the vehicle is in an open state; if the window on the vehicle is in the open state, determining whether the proportion of the energy of the first type of signal in the sound is less than the second threshold; if the window on the vehicle is in a closed state, determining whether the sound has multiple base frequencies, and when the sound has multiple base frequencies, starting the target function on the electronic device. In this way, based on the state of the window on the vehicle, it is determined whether to judge the proportion of the energy of the first type of signal, thereby improving the efficiency of the judgment. Meanwhile, when the window is in the closed state, the large decibel value of the sound in the vehicle is often caused by internal factors of the vehicle, such as someone speaking in the vehicle, or other electronic devices in the vehicle are outputting sound, and the like, and therefore, the number of base frequencies in the sound can be judged, and when the sound has multiple base frequencies, the target function is directly started, thereby improving the efficiency of starting the target function.

[0009] In a possible implementation, the method further includes: when the decibel value of the sound is less than or equal to the first threshold, or when the sound has one base frequency, controlling the target function on the electronic device to be in a closed state. In this way, the condition for starting the target function is set, which can avoid the target function being continuously in the open state, thereby avoiding the situation that the memory resource is excessively consumed due to the voice separation function of the electronic device being always on, and reducing the power consumption of the electronic device.

[0010] In a possible implementation, a sound collection device associated with the electronic device is in an open state, and the sound collection device is at least used to acquire the sound. In this way, the sound can be continuously collected through the sound collection device.

[0011] In a possible implementation, starting the target function on the electronic device specifically includes: controlling an image collection device associated with the electronic device to be started, and the image collection device is at least used to acquire a facial image of the user. In this way, the multi-modal voice separation manner based on audio and video data can be used to separate the voice of the target user.

[0012] In a possible implementation, before the image acquisition device associated with the electronic device is controlled to be turned on, the method further includes: determining whether a voiceprint feature of the target user is stored; determining whether the voiceprint feature is stored; if it is determined that the voiceprint feature is not stored, controlling the image acquisition device associated with the electronic device to be turned on; and if it is determined that the voiceprint feature is stored, separating the voice of the target user from the acquired sound based on the voiceprint feature. In this way, when the voiceprint feature of the user is pre-stored, the voice separation can be implemented by using the voiceprint-based voice separation manner, so that the increase in power consumption of the electronic device or a device matched with the electronic device caused by turning on the image acquisition device is avoided. When the voiceprint feature of the user is not stored, the image acquisition device can be turned on, and then the voice separation of the target user can be implemented by using the multi-mode voice separation manner based on audio and video data.

[0013] In a possible implementation, before the voice of the target user is separated from the acquired sound based on the voiceprint feature, the method further includes: determining whether the stored voiceprint feature matches the voiceprint feature of the user contained in the acquired sound; if the stored voiceprint feature matches the voiceprint feature of the user contained in the acquired sound, separating the voice of the target user from the acquired sound based on the voiceprint feature; and if the stored voiceprint feature does not match the voiceprint feature of the user contained in the acquired sound, controlling the image acquisition device associated with the electronic device to be turned on. In this way, when the voiceprint feature of the user pre-stored matches the voiceprint feature of the user in the acquired sound, the voice separation can be implemented by using the voiceprint-based voice separation manner, so that the increase in power consumption of the electronic device or a device matched with the electronic device caused by turning on the image acquisition device is avoided. When the voiceprint feature of the user pre-stored does not match the voiceprint feature of the user in the acquired sound, the image acquisition device can be turned on, and then the voice separation of the target user can be implemented by using the multi-mode voice separation manner based on audio and video data. In this way, the voice of the user is separated to the greatest extent, and the user experience is improved.

[0014] In a possible implementation, before it is determined whether the proportion of the energy of the first type of signal in the sound is less than the second threshold, the method further includes: determining the energy of the first type of signal in the sound according to a pre-set frequency range, or determining the energy of the first type of signal in the sound according to the pre-set frequency range and the amplitude value of the sound intensity.

[0015] In a possible implementation, the first type of signal is a signal in a first frequency range, or the first type of signal is a signal in a second frequency range and the amplitude value of the sound intensity is in a first amplitude value range.

[0016] In a possible implementation, after the target function on the electronic device is turned on, the method further includes: continuing to collect sound through the sound collection device, and turning off the target function on the electronic device when a time length during which the decibel value of the sound collected by the sound collection device is less than or equal to the first threshold value reaches a preset time. In this way, the phenomenon of multiple turning on and turning off within a short time due to the environment becoming quiet within a short time can be avoided.

[0017] In a second aspect, the present application provides a control device of an electronic device, comprising:

[0018] at least one processor and an interface;

[0019] The at least one processor acquires program instructions or data through the interface;

[0020] The at least one processor is configured to execute the program instructions to implement the method in the first aspect.

[0021] In a third aspect, the present application provides an electronic device, comprising:

[0022] at least one memory configured to store a program;

[0023] at least one processor configured to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method in the first aspect.

[0024] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program (also referred to as instructions or code) for implementing the method in the first aspect or the second aspect.

[0025] For example, when the computer program is executed by a computer, the computer can execute the method in the first aspect.

[0026] In a fifth aspect, the present application provides a chip comprising a processor. The processor is configured to read and execute a computer program stored in a memory to execute the method in the first aspect.

[0027] Optionally, the chip further comprises the memory, and the memory is connected to the processor through a circuit or a wire.

[0028] In a sixth aspect, the present application provides a computer program product, which comprises a computer program (also referred to as instructions or code). When the computer program is executed by a computer, the computer program makes the computer implement the method in the first aspect.

[0029] It can be understood that the beneficial effects of the second aspect to the sixth aspect described above can be referred to the related description of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 FIG. 1 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0031] Figure 2 FIG. 2 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;

[0032] Figure 3 FIG. 3 is a display schematic diagram of an electronic device provided by an embodiment of the present application;

[0033] Figure 4 FIG. 4 is a flow schematic diagram of a control method of an electronic device provided by an embodiment of the present application;

[0034] Figure 5 FIG. 5 is a schematic diagram of a spectrum graph of a sound provided by an embodiment of the present application;

[0035] Figure 6 FIG. 6 is a flow schematic diagram of another control method of an electronic device provided by an embodiment of the present application;

[0036] Figure 7 FIG. 7 is a flow schematic diagram of still another control method of an electronic device provided by an embodiment of the present application;

[0037] Figure 8 FIG. 8 is a step schematic diagram of a multi-modal speech separation method based on audio and video data provided by an embodiment of the present application;

[0038] Figure 9 FIG. 9 is a step schematic diagram of a speech separation method based on a voiceprint provided by an embodiment of the present application;

[0039] Figure 10 FIG. 10 is a structural schematic diagram of a control device of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0041] The term “and / or” in the present document is a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The symbol “ / ” in the present document represents an or relationship of the associated objects, for example, A / B represents A or B.

[0042] The terms "first" and "second" and the like in the description and claims of this application are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. For example, the first response message and the second response message are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order.

[0043] In the embodiments of this application, the words "exemplary" and "for example" are used to mean serving as an example, instance, or illustration, at 99 2 / 27 / 2017 11:45 AM not necessarily as preferable or advantageous over other embodiments or implementations. The words "exemplary" and "for example" are used herein to mean serving as an example, instance, or illustration. Any embodiment or implementation described as "exemplary" or "for example" is not necessarily to be construed as preferred or advantageous over other embodiments or implementations.

[0044] In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more, for example, a plurality of processing units means two or more processing units, and the like; a plurality of elements means two or more elements, and the like.

[0045] For ease of understanding, the technical terms involved in the present solution are introduced first.

[0046] (1) Speech separation

[0047] Speech separation mainly refers to separating the speech of a target user from the complex sound obtained, and / or separating the speech of the target person from other interference sounds to obtain clear speech of the target user. The target user can be one or multiple. When the target user is multiple, the speech of each target user can be separated, but is not limited thereto.

[0048] (2) Speech separation function

[0049] The speech separation function refers to a function configured on an electronic device and capable of enabling the electronic device to realize speech separation. For example, the electronic device can use a multi-modal speech separation method based on audio and video data and / or use a speech separation method based on voiceprints for speech separation.

[0050] Next, the technical solutions in the embodiments of this application are described.

[0051] For example, for ease of illustration, the electronic device is taken as an example of a car machine in a vehicle, and the application scenario of this application is illustrated in combination with Figure 1 For example, for ease of illustration, the electronic device is taken as an example of a car machine in a vehicle, and the application scenario of this application is illustrated in combination with For example, for ease of illustration, the electronic device is taken as an example of a car machine in a vehicle, and the application scenario of this application is illustrated in combination with

[0052] Exemplarily, Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present application is shown. As shown in the figure, Figure 1 Users A1 and A2 are in vehicle B, and user A1 is the driver. The electronic device 100 is arranged in vehicle B. When user A1 is interacting with the electronic device 100, if user A2 is speaking in a high decibel voice, and / or the noise inside or outside the vehicle is large, the electronic device 100 can actively start the voice separation function to extract the voice of user A1 from the noisy environment and respond. In this way, the electronic device 100 automatically selects whether to start the voice separation function based on the noise in the environment when user A1 speaks, thereby avoiding the situation that the memory resource is excessively consumed due to the voice separation function of the electronic device 100 being always on, reducing the power consumption of the electronic device 100, and also avoiding the cumbersome operation of manual starting or stopping, improving the user experience. In addition, it also avoids the safety problem caused by the driver manually operating during driving, improves the driving safety, and improves the driving safety.

[0053] It can be understood that in the present scheme, the electronic device 100 can be a mobile phone, a tablet computer, a wearable device, a smart television, a Huawei smart screen, a smart speaker, a car machine, etc. Exemplary embodiments of the electronic device 100 include but are not limited to electronic devices running iOS, android, Windows, Harmony OS or other operating systems. The type of electronic device 100 is not limited in the embodiments of the present application.

[0054] Exemplarily, Figure 2 A structural schematic diagram of an electronic device 100 provided by an embodiment of the present application is shown. As shown in the figure, Figure 2 The electronic device 100 can include a processor 110, a memory 120, a sound pickup device 130, an image acquisition device 140, a clock synchronization module 150, a transceiver unit 160 and a display screen 170.

[0055] The processor 110 can be a general-purpose processor or a special-purpose processor. For example, the processor 110 can include a central processing unit (CPU) and / or a baseband processor. The baseband processor can be configured to process communication data, and the CPU can be configured to implement corresponding control and processing functions, execute software programs, and process data of the software programs. For example, the processor 110 can determine the size of the sound (i.e., the decibel value of the sound) in a preset time period based on the audio signal collected by the microphone 130, determine whether to enable the voice separation function on the electronic device 100, and separate the sound of the target user (i.e., perform voice separation). For example, the processor 110 can sample the audio signal collected by the microphone 130, and then calculate the size of the sound based on a sound pressure level calculation formula. For example, the processor 110 can perform automatic speech recognition (ASR), natural language understanding (NLU), dialogue management (DM), natural language generation (NLG), text-to-speech (TTS), and the like on the audio signal collected by the microphone 130 to identify the voice of the target user.

[0056] In one example, the processor 110 can include one or more processing units. For example, the processor 110 can include one or more of an application processor (AP), a modem, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). In some embodiments, the electronic device 100 can include one or more processors 110. Different processing units can be independent devices or integrated into one or more processors.

[0057] The memory 120 can store a program, which can be run by the processor 110, so that the processor 110 performs the method provided in the present application. The memory 120 can also store data. The processor 110 can read the data stored in the memory 120 (for example, the voiceprint feature data of the target user, etc.). The memory 120 and the processor 110 can be separately arranged. Alternatively, the memory 120 can also be integrated in the processor 110.

[0058] The pickup 130 can also be referred to as a "microphone", a "sound receiver", and is used to convert a sound signal into an electrical signal. The electronic device 100 can include one or more pickups 130. The pickup 130 can collect sound in the environment where the electronic device 100 is located, and transmit the collected sound to the processor 110 for processing by the processor 110. The pickup 130 can be a microphone, for example. The pickup 130 can be built-in or external to the electronic device 100, which is not limited here.

[0059] Optionally, the electronic device 100 can also include an image acquisition device 140. The image acquisition device 140 can be used to acquire images of the target user or other users. The image acquisition device 140 can be an RGB camera, an infrared camera, etc., for example. The image acquisition device 140 can be built-in or external to the electronic device 100, which is not limited here.

[0060] Optionally, the electronic device 100 can also include a clock synchronization module 150. The clock synchronization module 150 can be used to synchronize the time between various components on the electronic device 100, such as synchronizing the time of the pickup 130 and the image acquisition device 140, etc.

[0061] Optionally, the electronic device 100 can also include a transceiver unit 160. The transceiver unit 160 can realize the input (reception) and output (transmission) of signals. For example, the transceiver unit 160 can include a transceiver or a radio frequency chip. The transceiver unit 160 can also include a communication interface. The electronic device 100 can obtain the state signal of other devices (such as a vehicle, etc.) through the transceiver unit 160, for example. For example, when the electronic device 100 is a car machine on a vehicle, the electronic device 100 obtains the signal of the central control playing device on the vehicle, the opening and closing signal of the vehicle window, the seat belt fastening signal, the sensor (such as a pressure sensor on the seat, etc.) signal, etc. through the transceiver unit 160.

[0062] Optionally, the electronic device 100 can also include a display screen 170. The display screen 170 can be used to display a user interface (UI). The display screen 170 can be a liquid crystal display (LCD) screen, an organic light-emitting diode (OLED) screen, etc., for example.Figure 3 As shown, the display screen 170 of the electronic device 100 can display the on / off state of "intelligent opening of voice separation", and the electronic device 100 can be provided with a button for opening / closing the "intelligent opening of voice separation" function. The button can be a mechanical button or a virtual button, which is not limited herein. In addition, the user can also issue a voice instruction to open / close the "intelligent opening of voice separation" function. In an example, after the user opens the "intelligent opening of voice separation" function, the electronic device 100 can automatically select to open or close the voice separation function, that is, to execute the method provided by the embodiments of the present application. For example, when the user opens / closes the "intelligent opening of voice separation" function, the electronic device 100 can output prompt information, such as one or more of text prompts, voice prompts, and image prompts.

[0063] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0064] Next, based on the above description, a control method of an electronic device provided by the embodiments of the present application is introduced.

[0065] For example, Figure 4 A flowchart of a control method of an electronic device provided by the embodiments of the present application is shown. As shown in the flowchart, Figure 4 The electronic device involved in the embodiments of the present application can be the electronic device 100 described in the embodiments of the present application. Figure 1 or 2. In addition, the electronic device 100 can be but not limited to be configured with an application related to voice interaction with the user (such as a voice assistant, etc.), which can be in an open state or in a closed state. As shown in the flowchart, Figure 4 The method can include the following steps:

[0066] S401, determining the decibel value of the sound obtained for a preset time length.

[0067] Specifically, after the sound collector on the electronic device 100 obtains the sound for a preset time length, the processor in the electronic device 100 can sample the sound for the preset time length, and then calculate the decibel value of the sound for the preset time length according to the sound pressure level calculation formula.

[0068] S402, determining whether the decibel value of the obtained sound is greater than a preset sound threshold.

[0069] Specifically, the obtained decibel value of the preset duration of sound is compared with the preset sound threshold, i.e., the relationship between the two can be determined. When the obtained decibel value of the sound is greater than the preset sound threshold, S403 can be performed. When the obtained decibel value of the sound is less than or equal to the preset sound threshold, S406 can be performed. In an example, the preset sound threshold can also be referred to as a first threshold.

[0070] In an example, when the user does not speak and the ambient noise around the user is small, the sound obtained by the electronic device is generally lower than the preset sound threshold, at this time it can be considered that there is no possibility of the user interacting with the electronic device 100, and therefore the voice separation function on the electronic device 100 does not need to be started, i.e., S406 is performed. When the user speaks and / or the ambient noise around the user is large, the sound obtained by the electronic device is generally greater than the preset sound threshold, at this time it can be considered that there is a possibility of the user interacting with the electronic device 100, and therefore it can be further determined whether the voice separation function on the electronic device 100 needs to be started, i.e., S403 is performed. For example, the user's voice is generally around 40-60 decibels, and therefore the preset sound threshold can be greater than 40 decibels or 60 decibels, such as 65 decibels.

[0071] S403, determining the proportion of energy of the first type of signal in the obtained sound.

[0072] Specifically, the obtained sound of the preset duration can be subjected to spectral energy analysis to obtain the proportion of energy of the first type of signal in the obtained sound. In an embodiment of the present application, the obtained sound can include the first type of signal and the second type of signal, wherein the frequency ranges of the first type of signal and the second type of signal are different, or the frequency ranges of the first type of signal and the second type of signal are different and the amplitude values of the sound intensities corresponding to the two are also different. For example, the first type of signal includes noise signals, and the second type of signal includes noise signals and / or voice signals emitted by the user.

[0073] As a possible implementation manner, the obtained sound of the preset duration can be subjected to short-time Fourier transform (STFT) to obtain a frequency spectrum diagram of the obtained sound, and the proportion of energy of the first type of signal in the obtained sound can be obtained through the frequency spectrum diagram. For example, in the frequency spectrum diagram, the voice signal generated by the user when speaking is generally distributed in a low frequency band, such as 100 Hz to 400 Hz, and noise signals are generally distributed in a medium-high frequency band (such as a frequency band greater than 400 Hz), and therefore the energy proportion of the medium-high frequency band in the frequency spectrum diagram can be taken as the proportion of energy of the first type of signal, and the energy proportion of the low frequency band can be taken as the proportion of the second type of signal.

[0074] For example, as shown in Figure 5 , the figure is a frequency spectrum diagram obtained by performing short-time Fourier transform on the obtained sound of the preset time length, the horizontal axis is frequency, and the vertical axis is sound intensity; wherein the energy value of the obtained sound of the preset time length can be obtained by performing square summation on the amplitude of the sound intensity in the frequency domain. In Figure 5 , the energy of the obtained sound in the medium-high frequency band (i.e. the area surrounded by the dashed box m in the figure) can be taken as the energy of the first type signal, and similarly, the energy value of the obtained sound in the medium-high frequency band can be obtained by performing square summation on the amplitude of the sound intensity in the frequency domain, that is, the energy value of the first type signal in the obtained sound. Then, the ratio of the energy value of the first type signal to the energy value of the sound of the preset time length can be obtained, that is, the proportion of the energy of the first type signal in the obtained sound of the preset time length.

[0075] As another possible implementation, since the decibel value of the sound signal emitted by the user is generally within a preset decibel range, such as around 40 to 60 decibels. Therefore, the signal within the preset decibel range can be taken as the second type signal, and the signal outside the second type signal can be taken as the first type signal. For example, referring to Figure 5 , if the pre-set decibel value of the second type signal is 40-60 dB, the energy of the signal whose amplitude value is within 40-60 dB in Figure 5 can be taken as the energy of the second type signal, and the energy of the signal other than the second type signal can be taken as the energy of the first type signal. For example, to determine the energy of the signal within the preset decibel range, the amplitude value corresponding to each sampling point within the preset decibel range can be squared and summed to obtain the energy of the signal within the preset decibel range.

[0076] As another possible implementation, since the frequency of the sound signal emitted by the user is generally within a preset frequency range, such as 100 Hz to 400 Hz, and the decibel value is generally within a preset decibel range, such as around 40 to 60 decibels. Therefore, in order to improve the accuracy of determining the energy proportion of the first type signal, the signal within the preset frequency range and within the preset decibel range can be taken as the second type signal, and the signal outside the second type signal can be taken as the first type signal.

[0077] For example, referring to Figure 5 , if the pre-set frequency of the second type signal is 100-400 Hz and the decibel value is 40-60 dB, the energy of the signal whose amplitude value is within 40-60 dB in Figure 5The energy of the signal with the frequency in the range of 100-400 Hz and the amplitude value in the range of 40-60 dB is regarded as the energy of the second type of signal, and the energy of the signal other than the second type of signal is regarded as the energy of the first type of signal. For example, to determine the energy of the signal in the preset frequency range and in the preset decibel range, the amplitude values corresponding to each sampling point of the signal in the preset frequency range and in the preset decibel range are squared and summed to obtain the energy of the signal in the preset frequency range and in the preset decibel range.

[0078] In some embodiments of the present application, the decibel value can also be understood as the amplitude value of the sound intensity.

[0079] S404, determining whether the proportion of the energy of the first type of signal in the obtained sound is greater than a preset proportion threshold.

[0080] Specifically, the proportion of the energy of the first type of signal in the obtained sound is compared with the preset proportion threshold, i.e., the relationship between the two is determined. When the proportion of the energy of the first type of signal in the obtained sound is greater than the preset proportion threshold, S405 is performed. When the proportion of the energy of the first type of signal in the obtained sound is less than or equal to the preset proportion threshold, S406 is performed.

[0081] In one example, when the proportion of the energy of the first type of signal in the obtained sound is greater than the preset proportion threshold, it indicates that the current user is in a noisy environment, and the proportion of the energy of the useful signal in the current obtained sound is small, and the ratio of the energy of the second type of signal to the energy of the first type of signal in the obtained sound is small. Therefore, at this time, in order to extract the voice of the user from the noisy environment and respond, the voice separation function on the electronic device 100 can be started, i.e., S405 is performed. When the proportion of the energy of the first type of signal in the obtained sound is less than or equal to the preset proportion threshold, it indicates that the current user is in a quiet environment, and the proportion of the energy of the useful signal in the current obtained sound is large, and the ratio of the energy of the second type of signal to the energy of the first type of signal in the obtained sound is high. The voice of the user can be accurately recognized and responded, so at this time the voice separation function on the electronic device 100 can not be started, i.e., S406 is performed.

[0082] S405, starting the voice separation function.

[0083] Specifically, the voice separation function on the electronic device 100 can be started at this time to extract the voice of the user and respond.

[0084] S406, controlling the voice separation function to be in an off state.

[0085] Specifically, at this time, the voice separation function on the electronic device 100 can be controlled to continue in the off state to reduce the power consumption of the electronic device 100.

[0086] In this way, by the noise condition in the environment, it is automatically selected whether to turn on the voice separation function on the electronic device 100, thereby avoiding the case that the memory resource consumption is too large due to the voice separation function on the electronic device 100 being always on, reducing the power consumption of the electronic device 100, and also avoiding the cumbersome operation of manual turning on or turning off, thereby improving the user experience.

[0087] Exemplarily, Figure 6 A flowchart of another control method of an electronic device provided by an embodiment of the present application is shown. Wherein, Figure 6 The electronic device involved in the above embodiments can be Figure 1 or the electronic device 100 described in the above embodiments. In addition, the electronic device 100 can be, but is not limited to, configured with an application related to voice interaction with the user (such as a voice assistant, etc.), which can be in an on state or in an off state. In Figure 6 , S601 to S604, and S606 and S607 are the same as the contents in S401 to S406 in the above Figure 4 embodiments, and can refer to the description of S401 to S406 in the above Figure 4 embodiments, which will not be repeated here. As Figure 6 shown, the method can include the following steps:

[0088] S601, determine the decibel value of the sound obtained within the preset time length.

[0089] S602, determine whether the decibel value of the obtained sound is greater than a preset sound threshold.

[0090] S603, determine the proportion of the energy of the first type of signal in the obtained sound.

[0091] S604, determine whether the proportion of the energy of the first type of signal in the obtained sound is greater than a preset proportion threshold.

[0092] Specifically, when the proportion of the energy of the first type of signal in the obtained sound is greater than the preset proportion threshold, S606 can be performed. When the proportion of the energy of the first type of signal in the obtained sound is less than or equal to the preset proportion threshold, S605 can be performed. In one example, the preset proportion threshold can also be referred to as a second threshold.

[0093] When the proportion of the energy of the first type of signal in the acquired sound is less than or equal to the preset proportion threshold, it indicates that the current user is in a quiet environment, and the proportion of the energy of the useful signal in the current acquired sound is large, and the ratio of the energy of the second type of signal to the energy of the first type of signal of the acquired sound is high, which can accurately identify the voice of the user, but at this time there may be multiple users speaking at the same time, and therefore it can be determined whether there are multiple users speaking in the current environment, that is, S605 is performed to determine whether the current acquired sound is mixed voice, and further determine whether the voice separation function of the electronic device 100 needs to be started.

[0094] S605, determining whether multiple base frequencies exist in the acquired sound.

[0095] Specifically, the number of base frequencies in the acquired sound can be determined by using a short-time Fourier transform (STFT), an autocorrelation algorithm, or a machine learning algorithm, to determine whether multiple base frequencies exist in the acquired sound. If multiple base frequencies exist in the acquired sound, S606 is performed; otherwise, S607 is performed.

[0096] For example, when the number of base frequencies in the acquired sound is determined by using the short-time Fourier transform, the frequency spectrum graph can be obtained by performing the short-time Fourier transform on the acquired sound of the preset time length, and the number of base frequencies in the sound of the preset time length can be determined by using a pitch detection algorithm, and then it can be determined whether multiple base frequencies exist in the sound of the preset time length. In addition, after performing the short-time Fourier transform on the acquired sound of the preset time length, the frequencies corresponding to the more obvious peaks in the low frequency band of the frequency spectrum graph can be used as the base frequencies. For example, continuing to refer to Figure 5 In Figure 5 In the low frequency band, the peaks n1 and n2 are more obvious, so the frequencies corresponding to the peaks n1 and n2 can be used as the base frequencies. At this time, it can be determined that two base frequencies exist in the acquired sound.

[0097] When the autocorrelation algorithm is used to determine the number of base frequencies in the acquired sound, the sound signal can be shifted and convolution operation can be performed. If the result reaches the maximum value, the time length corresponding to the result can be used as the period of the signal corresponding to the corresponding base frequency, and the inverse of the period is the base frequency.

[0098] When the machine learning method is used to determine the number of base frequencies in the acquired sound, the sound of the preset time length can be input into the pre-trained machine learning model, and then the number of base frequencies can be obtained by analyzing the machine learning model.

[0099] In one example, when the user is speaking, the vocal cords vibrate to generate a pulse. Among them, the fundamental frequency can correspond to the speed of the user's vocal cord vibration when the user is speaking. The higher the fundamental frequency, the faster the vocal cord vibration, and the sharper the sound emitted by the user. Because the vocal cord vibration speed of different people is different, that is, the frequency of the vocal cord vibration is different, therefore, the fundamental frequency of the voice signal generated by different people when speaking is also different. Among them, when multiple fundamental frequencies exist in the acquired sound, it indicates that multiple users are speaking at this time, and therefore the voice separation function on the electronic device 100 can be started at this time, that is, S606 is executed. When there is no multiple fundamental frequencies in the acquired sound, it indicates that there is no multiple users speaking at this time, and therefore the voice separation function on the electronic device 100 does not need to be started at this time, that is, S607 is executed.

[0100] S606, starting the voice separation function.

[0101] S607, controlling the voice separation function to be in a closed state.

[0102] Therefore, by the noise condition in the environment and the number of users in the environment who emit voice, it is automatically selected whether to start the voice separation function on the electronic device 100, thereby avoiding the case that the memory resource consumption is too large caused by the voice separation function on the electronic device 100 being always on, reducing the power consumption of the electronic device 100, and also avoiding the cumbersome operation of manual starting or closing, thereby improving the user experience.

[0103] Exemplarily, Figure 7 A flowchart of another control method of an electronic device provided by an embodiment of the present application is shown. Among them, Figure 7 The electronic device involved in the above Figure 1 The electronic device 100 described in the above Figure 7 application (such as a voice assistant, etc.) related to voice interaction with the user, which can be in an open state or in a closed state. In addition, Figure 7 The electronic device involved in the above Figure 4 The contents in S401-S406 in the above Figure 4 The description in the above Figure 6 The contents in S605 in the above Figure 6 The description in the above Figure 7As shown, the method can include the following steps:

[0104] S701, determining the decibel value of the obtained sound of the preset time length.

[0105] S702, determining whether the decibel value of the obtained sound is greater than a preset sound threshold.

[0106] Specifically, the decibel value of the obtained sound of the preset time length is compared with the preset sound threshold, that is, the relationship between the two can be determined. When the decibel value of the obtained sound is greater than the preset sound threshold, S703 can be performed. When the decibel value of the obtained sound is less than or equal to the preset sound threshold, S711 can be performed.

[0107] In one example, when the decibel value of the obtained sound is greater than the preset sound threshold, it can be considered that there is a possibility of user interaction with the electronic device 100. At this time, since the scene is in the vehicle riding scene, when the sound of the sound system in the vehicle is emitted (that is, the sound system is in an open state), the electronic device 100 is also difficult to recognize the voice of the user, and therefore the state of the sound system in the vehicle can be determined to further determine whether the voice separation function on the electronic device 100 needs to be turned on, that is., S703 is executed.

[0108] S703, determining whether the sound system in the vehicle is in an open state.

[0109] Specifically, whether the sound system in the vehicle is in an open state can be determined according to the state signal of the vehicle obtained from the vehicle, such as the signal of the central control playback device (such as a sound system, etc.) on the vehicle, etc. When the sound system in the vehicle is in an open state, S707 is executed; when the sound system in the vehicle is not in an open state, S704 is executed.

[0110] When the sound system in the vehicle is in an open state, the sound emitted by the sound system will affect the recognition of the voice of the user, so at this time the voice separation function on the electronic device 100 can be turned on, that is., S707 is executed. When the sound system in the vehicle is not in an open state, the sound system will not affect the recognition of the voice of the user, so the state of the vehicle window in the vehicle can be determined to further determine whether the voice separation function on the electronic device 100 needs to be turned on, that is., S704 is executed.

[0111] In one example, the electronic device 100 can communicate with the vehicle, so that the electronic device 100 can obtain the state signal of the vehicle, such as the signal of the central control playback device (such as a sound system, etc.) on the vehicle, the opening and closing signal of the vehicle window, the seat belt fastening signal, the sensor (such as a pressure sensor on the seat, etc.) signal, etc.

[0112] S704, determine whether the vehicle window is in an open state.

[0113] Specifically, whether the vehicle window is in an open state can be determined according to a state signal of the vehicle obtained from the vehicle, such as an on-off signal of the vehicle window, etc. When the vehicle window in the vehicle is in an open state, S705 is executed; when the sound in the vehicle is not in an open state, S708 is executed.

[0114] When the vehicle window in the vehicle is in an open state, the wind noise, tire noise and other noises outside the vehicle can cause the decibel value of the obtained sound to increase, and these noises can affect the recognition of the voice uttered by the user. Therefore, in order to determine whether the current sound decibel value is large due to the noise outside the vehicle, the noise proportion in the obtained sound can be analyzed, that is, S705 is executed.

[0115] When the vehicle window in the vehicle is not in an open state, the large decibel value of the sound in the vehicle is often caused by internal factors in the vehicle, such as someone speaking in the vehicle, or other electronic devices in the vehicle are outputting sound, etc. Therefore, at this time, in order to determine whether the current obtained sound decibel value is large due to internal factors in the vehicle, the internal factors in the vehicle can be analyzed, that is, S708 is executed.

[0116] S705, determine the proportion of the energy of the first type of signal in the obtained sound.

[0117] S706, determine whether the proportion of the energy of the first type of signal in the obtained sound is greater than a preset proportion threshold.

[0118] Specifically, when the proportion of the energy of the first type of signal in the obtained sound is greater than the preset proportion threshold, S707 can be executed. When the proportion of the energy of the first type of signal in the obtained sound is less than or equal to the preset proportion threshold, S708 can be executed.

[0119] When the proportion of the energy of the first type of signal in the obtained sound is less than or equal to the preset proportion threshold, it indicates that the current user is in a quiet environment, and the energy of the useful signal in the current obtained sound is large, the ratio of the energy of the second type of signal to the energy of the first type of signal in the obtained sound is high, and the voice uttered by the user can be accurately recognized. However, at this time, the large decibel value of the sound in the vehicle is often caused by the user speaking in the vehicle, but at this time, it may be caused by one person speaking, or it may be caused by multiple people speaking. Therefore, at this time, in order to determine the number of users speaking in the vehicle, S708 can be executed.

[0120] S707, turn on the voice separation function.

[0121] S708, determining whether multiple base frequencies exist in the acquired sound.

[0122] Specifically, the STFT, autocorrelation algorithm or machine learning algorithm can be used to find the base frequency in the acquired sound to determine whether multiple base frequencies exist in the acquired sound. If multiple base frequencies exist in the acquired sound, S707 is performed. If multiple base frequencies do not exist in the sound, S709 is performed.

[0123] When multiple base frequencies exist in the acquired sound, it indicates that multiple users are speaking in the current environment, and / or other electronic devices exist in the current environment and are outputting sound, etc. Therefore, at this time, the voice separation function on the electronic device 100 can be turned on, so that the electronic device 100 can accurately identify the voice of the target user (such as the driver) and respond accordingly, that is, S707 is performed.

[0124] When multiple base frequencies do not exist in the acquired sound, and the decibel value of the acquired sound is large, it indicates that there may be only one user speaking at the moment, and since the energy of the first type of signal in the acquired sound at this time is small, the voice of the user in the acquired sound can be accurately identified. Therefore, at this time, the voice separation function can continue to be controlled in the off state, that is, S709 is performed.

[0125] S709, control the voice separation function in the off state.

[0126] Therefore, when the electronic device 100 is configured in the vehicle, the voice separation function on the electronic device 100 can be automatically selected to be turned on or off according to the noise in the environment, the running state of the vehicle, the number of users speaking in the environment, etc. Therefore, the problem of excessive memory resource consumption caused by the voice separation function on the electronic device 100 being always on is avoided, the power consumption of the electronic device 100 is reduced, and the user experience is improved. In addition, the tedious operation of manually turning on or off is avoided, the safety problem caused by the driver manually operating during driving is avoided, and the driving safety is improved.

[0127] In one example, after turning on the voice separation function on the electronic device 100 each time, at least a certain period of time (such as 3 minutes, etc.) can be set to avoid the phenomenon of turning on or off multiple times in a short period of time due to the environment becoming quiet in a short period of time. Further, when the decibel value of the sound acquired by the electronic device continuously falls below the preset sound threshold for a preset time, that is, when the time during which the environment is relatively quiet reaches the preset time, the electronic device 100 can automatically turn off the voice separation function.

[0128] In one example, after enabling the voice separation function on electronic device 100, electronic device 100 or its components (such as a processor) can perform voice separation on the acquired sound to extract the target user's voice. For instance, electronic device 100 or its components can use a multi-modal voice separation method based on audio and video data and / or a voiceprint-based voice separation method to perform voice separation on the acquired sound. These two voice separation methods are described below.

[0129] (I) Multimodal speech separation method based on audio and video data

[0130] When employing a multimodal speech separation method based on audio and video data, the electronic device 100 needs to be equipped with a microphone and / or an image acquisition device (e.g., a camera). Alternatively, when the electronic device 100 is configured on other devices (e.g., vehicles), those other devices need to be equipped with microphones and / or cameras so that the electronic device 100 can acquire the user's speech and facial image in the current environment. The following section provides a detailed introduction to the multimodal speech separation method based on audio and video data, using an electronic device 100 equipped with a microphone and an image acquisition device (e.g., a camera) as an example.

[0131] For example, such as Figure 8 As shown, after enabling the voice separation function on electronic device 100, in S801, electronic device 100 can control its microphone and camera to be both on and synchronize their times so that the time of the sound in the environment acquired by the microphone is consistent with the time of the image in the environment captured by the camera. This allows the camera to accurately record the user's facial image while the microphone is acquiring the user's voice, reducing the probability of inconsistency between the user's voice and facial image, and improving the accuracy of voice separation. For example, both the microphone and camera on electronic device 100 can obtain a standard time from the clock synchronization module on electronic device 100. Then, the microphone and camera can synchronize their respective data acquisition (i.e., acquiring the user's voice or facial image) start time to this standard time; alternatively, the clock synchronization module can provide a timestamp for the microphone and camera, which can then be used as the standard for data acquisition. For example, the microphone on electronic device 100 can remain on continuously to acquire the user's voice promptly; the camera on electronic device 100 can be off before enabling the voice separation function and on after enabling the voice separation function to reduce power consumption.

[0132] Next, in S802, the microphone on the electronic device 100 can acquire ambient sound in real time to obtain audio data, and the camera on the electronic device 100 can capture images of the environment in real time to obtain video data, thereby recording changes in the user's facial image. For example, the audio recorded by the microphone and the images captured by the camera can be sampled in real time according to a predetermined sampling rate to obtain audio data and video data.

[0133] Next, in S803, audio and video data can be fed into a pre-trained speech separation model for speech separation. For example, the electronic device 100 may be equipped with a face recognition module. This face recognition model module can process images of the environment captured by a camera and extract user facial data from the images. Then, the electronic device 100 can feed the audio data and the user facial data extracted by the face recognition module into the pre-trained speech separation model, which then separates the user's relatively clear speech. The speech separation model may include a multi-modal speech separation network, which may, but is not limited to, networks such as residual networks (ResNets) and / or fully convolutional temporal audio separation networks (Conv-TasNet). It is understood that facial movements during speech should be associated with the sound produced during speech; that is, sound is associated with the user's facial movements. Therefore, in the case of mixed speech, the user's speech can be separated by combining the user's facial movements with the sound.

[0134] Finally, in S804, the separated user voice can be sent to the voice assistant on the electronic device 100 so that the electronic device 100 can respond to the user's voice. Of course, when the voice assistant on the electronic device 100 is turned off, the electronic device 100 may not respond.

[0135] In one example, when the electronic device 100 is configured in a vehicle, if the camera on the electronic device 100 only captures the facial image of the user in the driver's seat and does not capture the facial images of users in other seats in the vehicle, then when performing voice separation, if the selected voice separation model requires facial data of multiple users, the facial image data of other seats can be zero-padding processed in order to separate the voice of the user in the driver's seat.

[0136] When the electronic device 100 is configured in the vehicle, if the camera on the electronic device 100 captures the facial image of the user on the main driving position and the facial image of the user on the other seat in the vehicle, the speech separation can be implemented by the multi-mode speech separation network when the speech separation is performed. Among them, the single channel can only select the output of the main driver's voice, and the multi-channel can output the voice of the main driver and the user on the other seat (such as the co-driver), so that the user on the other seat can also realize voice interaction.

[0137] (II) Speech separation mode based on voiceprint

[0138] When the speech separation mode based on voiceprint is adopted, a sound pickup device (such as a microphone) needs to be configured on the electronic device 100, or when the electronic device 100 is configured in other devices (such as a vehicle), a sound pickup device (such as a microphone) needs to be configured on the other devices, so that the electronic device 100 can obtain the voice of the user in the current environment. In addition, the electronic device 100 can pre-acquire the registered voice of at least one target user, and deliver the registered voice to the voiceprint encoder on the electronic device 100 to extract the voiceprint encoding features of the target user. For example, the voiceprint encoder on the electronic device 100 can be integrated on the electronic device 100, or can be separately arranged from the electronic device 100, which is not limited here.

[0139] Next, taking the electronic device 100 configured with a sound pickup device and the sound pickup device being a microphone as an example, the speech separation mode based on voiceprint is described in detail.

[0140] For example, as shown in FIG. 10, after the speech separation function on the electronic device 100 is started, in S901, the electronic device 100 can obtain the sound in the environment through the microphone thereon. Among them, the microphone can include at least one of the voice of the user and other sounds in the environment in the obtained sound. Figure 9 Next, in S902, the electronic device 100 can deliver the sound in the environment obtained by the microphone to the mixed speech encoder to extract the sound encoding features. For example, the mixed speech encoder can be used to encode the sound to obtain the encoding features of the sound, and the mixed speech encoder can include a pre-trained neural network-based mixed sound encoding model.

[0141] Next, in S903, the electronic device 100 can deliver the sound encoding features and the pre-acquired voiceprint encoding features of the target user to the speech separation network based on deep learning, and then separate the voice of the target user by the speech separation network.

[0142]

[0143] ​Finally, in S904, the separated voice of the target user can be sent to the voice assistant on the electronic device 100 so that the electronic device 100 can respond to the user's voice.

[0144] In some embodiments, before employing a multimodal speech separation method based on audio and video data, it can be determined whether the user's voiceprint features (or voiceprint coding features) are pre-stored. If the user's voiceprint features are not stored, the multimodal speech separation method based on audio and video data can be used to separate the target user's speech. If the user's voiceprint features are stored, a voiceprint-based method can be used to separate the target user's speech. This avoids the increased power consumption of the electronic device or its associated devices caused by activating the image acquisition device when employing the multimodal speech separation method based on audio and video data.

[0145] Furthermore, when user voiceprint features are stored, it can be determined whether the user's voiceprint features contained in the acquired audio match the stored user voiceprint features. If they match, the voiceprint-based method is used to separate the target user's speech; if they do not match, a multi-modal speech separation method based on audio and video data can be used to separate the target user's speech. This maximizes the separation of user speech from the acquired audio, improving the user experience.

[0146] Understandably, when using a voiceprint-based voice separation method, the audio within a preset duration before determining whether to enable the voice separation function can be stored first. After determining that the voice separation function needs to be enabled, the target user's voice can be extracted from the pre-stored audio and then concatenated with the subsequently separated target user's voice to obtain the target user's voice before and after enabling the voice separation function, thereby improving the accuracy of voice interaction and enhancing the user experience.

[0147] The above is a description of the control method for the electronic device provided in the embodiments of this application. For ease of understanding, examples will be given below for each scenario.

[0148] Scene 1

[0149] This scenario is Figure 1 The scenario shown is a driving scenario. In this scenario, the electronic device 100 in vehicle B can determine whether the voice separation function needs to be activated based on the method described above. After determining that the voice separation function needs to be activated, the electronic device 100 in vehicle B can activate the voice separation function, separate the voices of user A1 and / or user A2 from the acquired sound, and respond accordingly.

[0150] Scene 2

[0151] In this scenario, the electronic device 100 can be a robot, i.e., the scenario is a scenario in which the robot serves a user.

[0152] For example, the robot can use infrared detection, ultrasonic detection, camera detection, etc. to determine whether a user is approaching. When it is determined that a user is approaching, the robot can determine whether the voice separation function needs to be turned on based on the method described above. For example: whether the decibel value of the surrounding sound exceeds the threshold value, if so, by analyzing and / or judging whether there are multiple base frequencies in the same time period, to determine whether to turn on the voice separation function. After determining that the voice separation function needs to be turned on, the robot can turn on the voice separation, separate the voice of the user from the obtained sound, and respond. In this way, the intelligent voice separation function is applied to the robot, improving the accuracy of the voice interaction of the robot and improving the user experience.

[0153] Scenario three

[0154] In this scenario, the electronic device 100 can be an electronic device with camera function, such as a mobile phone, etc., i.e., the scenario is a scenario of shooting a video, and needs to be a scenario of shooting a person video, for example: a selfie scenario, a live scenario, etc.

[0155] For example, taking the user m's selfie video as an example, there may be interference such as horn, human voice, wind sound, etc. during the shooting of the user m's personal video. Among them, during the shooting of the user m's personal video, the electronic device used by the user m can determine whether the voice separation function needs to be turned on based on the method described above. For example: whether the decibel value of the surrounding sound exceeds the threshold value, if so, by analyzing and / or judging whether there are multiple base frequencies in the same time period, to determine whether to turn on the voice separation function. After determining that the voice separation function needs to be turned on, the electronic device used by the user m can turn on the voice separation, separate the voice of the user m from the obtained sound, suppress the interference noise in the environment, and then highlight the voice information of the user m. Thus, to improve the user's video recording experience.

[0156] It can be understood that in this scenario, the user can actively turn on the voice separation function on the electronic device, or the electronic device can determine whether to turn on the voice separation function by itself. In addition, when performing voice separation, real-time voice separation can be performed to suppress the interference noise in the environment in real time; in addition, the video after shooting can also be input to the pre-trained video processing model after the user finishes shooting the video, to obtain a video containing clear user voice.

[0157] It should be understood that the above three scenarios are only exemplary descriptions, and the specific scenarios can be determined according to actual conditions without departing from the principles of the present application, and will not be repeated here.

[0158] It can be understood that the method provided in any embodiment of the present application can be executed by any device, apparatus, platform, device cluster with computing and processing capability, in addition to the electronic device 100, and the execution manner of the execution subject other than the electronic device 100 is within the protection scope of the present application. In addition, the execution order of each step in any embodiment of the present application can be adjusted according to actual conditions without contradiction, and the adjusted technical solution is within the scope of the present application. In addition, each step in any embodiment of the present application can be selectively executed, which is not limited here. In addition, all or part of any features of any embodiment of the present application can be freely and arbitrarily combined without contradiction, and the combined technical solution is within the protection scope of the present application.

[0159] Based on the method in the above embodiment, the present embodiment further provides a control device of an electronic device. Please refer to Figure 10 , Figure 10 for a structural schematic diagram of the control device of the electronic device provided by the present embodiment. As shown in Figure 10 , the control device 1000 of the electronic device comprises one or more processors 1001 and interface circuits 1002. Optionally, the control device 1000 of the electronic device can also contain a bus 1003. Wherein:

[0160] The processor 1001 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 1001 or the instruction in the form of software. The processor 1001 mentioned above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method and step disclosed in the present embodiment can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The interface circuit 1002 can be used for sending or receiving data, instructions or information, and the processor 1001 can process the data, instructions or other information received by the interface circuit 1002, and can send the processed information through the interface circuit 1002.

[0161] Optionally, the control device of the electronic device further includes a memory, which can include a read-only memory and a random access memory, and provides operation instructions and data for the processor. A part of the memory can further include a non-volatile random access memory (NVRAM). Optionally, the memory stores executable software modules or data structures, and the processor can perform corresponding operations by calling operation instructions stored in the memory (which can be stored in an operating system).

[0162] Optionally, the interface circuit 1002 can be used to output the execution result of the processor 1001.

[0163] It should be noted that the functions of the processor 1001 and the interface circuit 1002 respectively can be realized by hardware design, software design, or a combination of hardware and software, which is not limited here.

[0164] It should be understood that each step of the above method embodiments can be completed by a logic circuit in the form of hardware in the processor or instructions in the form of software. The control device of the electronic device can be applied to the electronic device 100 in the above Figure 2 to realize the method provided in the embodiments of the present application.

[0165] It can be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.

[0166] The method steps in the embodiments of the present application can be implemented by hardware, or by a combination of software and hardware executed by a processor. The software instructions can be composed of a corresponding software module, which can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0167] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted by the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0168] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of differentiation, and do not limit the scope of the embodiments of the present application.

Claims

1. A control method for an electronic device, characterized in that, The method includes: Collect sound of a preset duration using a sound acquisition device; Determine the decibel value of the sound for the preset duration; When the decibel value of the sound of a preset duration is greater than the first threshold, it is determined whether the proportion of the energy of the first type of signal in the sound is less than the second threshold. When the proportion of the energy of the first type of signal in the sound is less than the second threshold, it is determined whether the sound has multiple fundamental frequencies; the first type of signal is a signal outside the preset decibel range; When multiple fundamental frequencies exist in the sound, the target function on the electronic device is activated. The target function is used to separate the target user's voice from the acquired sound.

2. The method according to claim 1, characterized in that, The electronic device is configured in a vehicle, and before determining whether the proportion of the energy of the first type of signal in the sound is less than a second threshold, the method further includes: Determine whether the audio playback device on the vehicle is turned off; If the sound playback device on the vehicle is turned off, then determine whether the proportion of the energy of the first type of signal in the sound is less than the second threshold. If the audio playback device on the vehicle is turned on, then the target function on the electronic device is activated.

3. The method according to claim 1 or 2, characterized in that, The electronic device is configured in a vehicle, and before determining whether the proportion of the energy of the first type of signal in the sound is less than a second threshold, the method further includes: Determine whether the windows of the vehicle are open; If the vehicle window is open, determine whether the proportion of the energy of the first type of signal in the sound is less than the second threshold. If the vehicle windows are closed, determine whether the sound contains multiple fundamental frequencies, and if multiple fundamental frequencies are present in the sound, activate the target function on the electronic device.

4. The method according to any one of claims 1-2, characterized in that, The method further includes: When the decibel value of the sound is less than or equal to the first threshold, or when there is a fundamental frequency in the sound, the target function on the electronic device is turned off.

5. The method according to any one of claims 1-2, characterized in that, The sound acquisition device associated with the electronic device is turned on, and the sound acquisition device is used at least to acquire sound.

6. The method according to any one of claims 1-2, characterized in that, Enabling the target function on the electronic device specifically includes: The system controls the activation of an image acquisition device associated with the electronic device, the image acquisition device being used at least to acquire a user's facial image.

7. The method according to claim 6, characterized in that, Before the control activates the image acquisition device associated with the electronic device, the method further includes: Determine whether voiceprint features are stored; If it is determined that no voiceprint features are stored, then the image acquisition device associated with the electronic device is activated. If it is determined that voiceprint features are stored, the voice of the target user is separated from the acquired sound based on the voiceprint features.

8. The method according to claim 7, characterized in that, Before separating the target user's voice from the acquired voice based on the voiceprint features, the method further includes: Determine whether the stored voiceprint features match the user's voiceprint features contained in the acquired audio; If the stored voiceprint features match the user's voiceprint features contained in the acquired sound, then the target user's voice is separated from the acquired sound based on the voiceprint features. If the stored voiceprint features do not match the user's voiceprint features contained in the acquired sound, then the image acquisition device associated with the electronic device is activated.

9. The method according to any one of claims 1-2 or 7-8, characterized in that, Before determining whether the proportion of energy of the first type of signal in the sound is less than the second threshold, the method further includes: The energy of the first type of signal in the sound is determined according to a preset frequency range, or the energy of the first type of signal in the sound is determined according to a preset frequency range and the amplitude value of the sound intensity.

10. The method according to any one of claims 1-2 or 7-8, characterized in that, The first type of signal is a signal within a first frequency range, or the first type of signal is a signal within a second frequency range and whose sound intensity amplitude value is within a first amplitude value range.

11. The method according to any one of claims 1-2 or 7-8, characterized in that, After activating the target function on the electronic device, the method further includes: Continue to collect sound through the sound acquisition device, and after a preset time has elapsed since the decibel value of the sound collected by the sound acquisition device is less than or equal to the first threshold, shut down the target function on the electronic device.

12. A control device for an electronic device, characterized in that, include: At least one processor and interface; The at least one processor obtains program instructions or data through the interface; The at least one processor is used to execute the program instructions to implement the method as described in any one of claims 1-11.

13. An electronic device, characterized in that, include: At least one memory for storing programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-11.

14. A computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-11.

15. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Voice speaker separation method and device

    CN111429935A

  • Method for extracting table tennis instruction of target human voice in complex scene

    CN112992131A