Signal processing method, device and electronic equipment
By monitoring the user's voice status and adjusting the playback parameters of the audio player, the problem of the audio player signal covering the user's audio is solved, the voice call quality is improved, and reliable audio recognition in electronic devices is achieved.
Patent Information
- Application Number
- CN202210042604.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-01-14
AI Technical Summary
During the voice call of the electronic device, the volume played by the audio player is too high, causing the collected user's voice signal to be overwritten, causing the communication party to be unable to receive the user's speech content, which reduces the quality of the voice call.
By monitoring the user's voice state, the playback parameters of the audio player are flexibly adjusted to ensure that the user's audio can be reliably recognized, including controlling the audio player to adopt different playback parameters in different states, such as reducing the volume in the voice input state or switching to a mute state, and performing echo cancellation processing during the audio acquisition process.
It effectively avoids the audio player signal covering user audio, ensures the improvement of voice communication quality, and realizes reliable audio recognition in multi-party voice call scenarios.
Smart Images

Figure CN114171039B_ABST
Abstract
Description
Technical Field
[0001] The present application mainly relates to the field of communication technology, and more specifically to a signal processing method, device and electronic equipment. Background Art
[0002] In the voice call application scenario of electronic devices, in order to improve the quality of voice calls, voice processing technologies in artificial intelligence (AI), such as echo cancellation technology, can be used to eliminate noise from the voice signals collected by the electronic devices to ensure that the other party can reliably receive the voice content. Summary of the Invention
[0003] In view of this, the present application proposes a signal processing method, comprising:
[0004] When the audio player of the electronic device is in a playing state, obtaining the voice state of the user of the electronic device;
[0005] controlling playback parameters of the audio player based at least on the voice state;
[0006] The playback parameters are at least used by the electronic device to perform corresponding processing on the audio collected by its audio collector.
[0007] Optionally, controlling the playback parameters of the audio player at least based on the voice state includes:
[0008] If the user is in a voice input state, controlling the audio player to be in a first playback parameter; and / or,
[0009] If the user is in a state of not inputting a voice, controlling the audio player to be in a second playing parameter;
[0010] The signal energy value output by the audio player under the second playback parameter is higher than the signal energy value output by the audio player under the first playback parameter.
[0011] Optionally, controlling the playback parameters of the audio player at least based on the voice state includes:
[0012] If the user is in a voice input state and the electronic device is in a first state, controlling the audio player to be in a first playback parameter; or,
[0013] If the user is in a voice input state and the electronic device is in the second state, controlling the audio player to be in a third playback parameter; or,
[0014] If the user is in a voice input state and the user and the electronic device are in a first positional relationship, controlling the audio player to be in a fourth playback parameter; or
[0015] If the user is in a voice input state and the user and the electronic device are in a second positional relationship, controlling the audio player to be in a fifth playback parameter;
[0016] Among them, the signal energy value output by the audio player under the third playback parameter is higher than the signal energy value under the first playback parameter, and the signal energy value output by the audio player under the fifth playback parameter is higher than the signal energy value under the fourth playback parameter.
[0017] Optionally, it also includes:
[0018] Performing corresponding processing on the audio collected by the audio collector so that the electronic device outputs the first audio to the communication terminal, or when the user is in a state of not inputting voice, the electronic device does not output the audio collected by the audio collector;
[0019] The first audio does not include the audio played by the audio player and collected by the audio collector.
[0020] Optionally, obtaining the voice status of the user of the electronic device includes:
[0021] obtaining, based at least on parameter information collected by a target sensor of the electronic device, information on changes in the user's mouth contour, and determining the user's voice state using the information on changes in the mouth contour; or
[0022] Determining the voice status of a user of the electronic device based on the operation or status of a control acting on the electronic device; or
[0023] The voice status of the user is determined based on a comparison result of the audio collected by the audio collector of the electronic device with the preset voiceprint information of the user of the electronic device.
[0024] Optionally, also include:
[0025] If the user is in a voice input state, controlling the audio player to be muted, and converting the audio to be played by the audio player into text information;
[0026] outputting the text information;
[0027] If the user is in a state of not inputting a voice, the audio player is controlled to switch from the mute state to the play state.
[0028] Optionally, when the user is in a voice input state, the implementation process of controlling the playback parameters of the audio player includes:
[0029] Obtaining a parameter threshold for a current playback parameter of the audio player; wherein the parameter threshold is a pre-configured value or is determined based on an audio attribute value of audio collected by the user when the audio player is in a muted state;
[0030] If the current playback parameter reaches the parameter threshold, the current playback parameter of the audio player is adjusted to a preset playback parameter; the preset playback parameter is the first playback parameter, the third playback parameter, the fourth playback parameter, or the fifth playback parameter; and / or,
[0031] If the current playback parameter does not reach the parameter threshold, the current playback parameter is determined as the preset playback parameter, and the audio player is controlled to maintain the preset playback parameter unchanged.
[0032] Optionally, after adjusting the current playback parameters of the audio player to preset playback parameters, the process of controlling the playback parameters of the audio player further includes:
[0033] If the user switches from the voice input state to the non-voice input state, the preset playing parameters of the audio player are restored to the playing parameters before adjustment.
[0034] The present application also proposes a signal processing device, comprising:
[0035] A voice status obtaining module, configured to obtain the voice status of a user of the electronic device when the audio player of the electronic device is in a playing state;
[0036] a playback parameter control module, configured to control playback parameters of the audio player based at least on the voice state;
[0037] The playback parameters are at least used by the electronic device to perform corresponding processing on the audio collected by its audio collector.
[0038] The present application also proposes an electronic device, comprising:
[0039] Audio collector; audio player; communication interface;
[0040] A memory for storing a program for implementing the above-mentioned signal processing method;
[0041] A processor is used to load and execute the program stored in the memory to implement the signal processing method as described above.
[0042] It can be seen that the present application proposes a signal processing method, device and electronic device. The present application will obtain the voice status of the user of the electronic device when the audio player of the electronic device is in the playing state, and control the playback parameters of the audio player based on at least the voice status, so as to avoid the audio played by the audio player covering the user's audio. In this way, when the audio collected by the audio collector of the electronic device (which may be a mixture of the user's audio and the played audio) is processed accordingly, the user's audio can be reliably identified, thereby ensuring the voice communication quality in the voice communication scenario and improving the voice communication efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0044] Figure 1 A schematic diagram of an optional scenario applicable to the signal processing method proposed in this application;
[0045] Figure 2 A schematic diagram of the hardware structure of an optional example of an electronic device applicable to the signal processing method proposed in this application;
[0046] Figure 3 A schematic diagram of the hardware structure of another optional example of an electronic device applicable to the signal processing method proposed in this application;
[0047] Figure 4 A flowchart of an optional example of the signal processing method proposed in this application;
[0048] Figure 5 A flowchart of another optional example of the signal processing method proposed in this application;
[0049] Figure 6 A flowchart of another optional example of the signal processing method proposed in this application;
[0050] Figure 7 A schematic diagram of a scenario in which the electronic device form is transformed in accordance with the signal processing method proposed in this application;
[0051] Figure 8 A flowchart of another optional example of the signal processing method proposed in this application;
[0052] Figure 9 A flowchart of another optional example of the signal processing method proposed in this application;
[0053] Figure 10 Schematic diagram of a flow chart of another optional scenario applicable to the signal processing method proposed in this application;
[0054] Figure 11 This is a schematic structural diagram of an optional example of the signal processing device proposed in this application. DETAILED DESCRIPTION
[0055] Regarding the description of the background technology, in application scenarios such as multi-person online conferences and network calls, when a user of a participating electronic device is speaking, the volume of the audio played by the audio player of the electronic device (i.e., the audio from the communication end of the electronic device, such as the audio sent by other participating electronic devices, which can be used as a reference signal for echo cancellation processing) is very high, resulting in the signal energy of the played audio collected by the audio collector of the electronic device being higher than the signal energy of the audio of the user's speech content, such as Figure 1 The processing flow is shown in the first row of the attached figure. In this way, the corresponding reference signal is subsequently used to perform echo cancellation on the audio currently actually collected by the audio collector. Due to the fusion of at least part of the signal between the audio of the user's speech and the audio played by the audio player, the reference signal is directly filtered from the collected audio. This may result in filtering all audio and making it impossible to output the user's audio, thereby causing the communication end to be unable to receive the local user's speech content, greatly reducing the quality of the voice call.
[0056] To improve the above problems, refer to Figure 1 The processing flow is shown in the second row of the accompanying drawings. This application proposes that the playback parameters of the audio player of the electronic device can be flexibly adjusted according to the local user's speaking situation, so that the local user can reliably hear the voice communication content output by the communication end played by the audio player. It can also reduce the signal energy of the audio played by the audio player when the local user speaks, ensuring that the user's audio can be reliably identified during subsequent processing, thereby ensuring the call quality in the multi-party voice call scenario.
[0057] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0058] Reference Figure 2, is a hardware structure diagram of an optional example of an electronic device suitable for the signal processing method proposed in this application. The electronic device may include but is not limited to: mobile phones, laptops, tablet computers, desktop computers, wearable devices, all-in-one computers, smart speakers, smart transportation equipment, smart medical equipment, etc., which can be determined according to the application scenario requirements. This application does not limit the product type of electronic devices. Figure 2 As shown, the electronic device may include but is not limited to: an audio collector 210, an audio player 220, a communication interface 230, a memory 240 and a processor 250, wherein:
[0059] The number of each of the audio collector 210, the audio player 220, the communication interface 230, the memory 240 and the processor 250 can be at least one; and the audio collector 210, the audio player 220, the communication interface 230, the memory 240 and the processor 250 can be connected to the communication bus in the electronic device to realize the communication connection between different components and meet the data transmission requirements between different components. This application does not elaborate on the method for realizing the communication connection between the various components contained in the electronic device, which can be determined as the case may be.
[0060] The audio collector 210 can be used to collect audio in the environment where the electronic device is located, such as the audio generated by the electronic device user speaking, the audio played by the electronic device audio player, and of course other noise audio in the environment where the electronic device is located. The audio collected by the audio collector 210 may be different in different scenarios, and this application does not give examples and details here. In the embodiment of the present application, the audio collector 210 can be a microphone installed in each electronic device listed above. The installation position and installation quantity of the audio collector 210 in the electronic device (such as a specific microphone array, etc.) can be flexibly determined according to different application requirements. This application does not impose any restrictions on this.
[0061] The audio player 220 can be a speaker installed in an electronic device, for example, and is used to play various audio signals obtained by the electronic device. The present embodiment of the application does not limit the number of audio players 220 installed in the electronic device or their respective installation locations. These locations can be determined based on a variety of factors, such as the type of electronic device, its structure, and audio playback requirements. In the present embodiment of the application, playback parameters of the audio player 220, such as playback volume and playback speed, can be adjusted based on different audio playback requirements. The implementation process is not described in detail in the present embodiment of the application.
[0062] The communication interface 230 may be a data interface of a corresponding communication module in an electronic device. For different types of communication modules, the type of the corresponding communication interface 230 and its communication protocol requirements may be different, depending on the circumstances. Among them, the communication module may include a communication module that can use a wireless communication network to implement data interaction, such as a WIFI module, a 5G / 6G (fifth generation mobile communication network / sixth generation mobile communication network) module, a GPRS module, a GMS module, a near-field communication module, etc. Therefore, the communication interface 230 may include a network interface that supports wireless communication; it is understood that the communication interface 230 may also include an interface for implementing data interaction between components within the electronic device, such as a USB interface, a serial / parallel port, etc., as well as a data interface such as a multimedia interface for implementing communication with a local device. This application does not limit the type and number of interfaces included in the communication interface 230, which may be determined as the circumstances require.
[0063] The memory 240 can be used to store programs for implementing the signal processing methods described in the above-mentioned method embodiments; the processor 250 can load and execute the program stored in the memory 240 to implement the various steps of the signal processing methods described in the corresponding method embodiments below. The specific implementation process can refer to the description of the corresponding parts of the embodiments below, and this embodiment will not be described in detail here.
[0064] It is understood that the memory 240 may include a program storage area and a data storage area. The program storage area may store the operating system of the electronic device and the application programs required for at least one function implemented by it (such as voice communication applications for voice communication functions, such as social software and phone calls), as well as programs for implementing the signal processing methods proposed in this application. The data storage area may store various data generated during the operation of the electronic device, such as collected audio, audio obtained from external devices, and audio after corresponding processing of the collected audio.
[0065] In an embodiment of the present application, the memory 240 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device. The processor 250 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices. The present application does not limit the structure and model of the above-mentioned memory 240 and processor 250, and they can be flexibly adjusted according to actual needs.
[0066] It should be understood that Figure 2The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiment of the present application. In actual applications, the electronic device may include Figure 2 More or fewer components as shown, or combinations of certain components, such as Figure 3 As shown, the electronic device may further include a sensor module composed of various sensors such as a temperature sensor, a pressure sensor, a gyroscope, and a distance sensor; input components such as a camera and a touch sensing unit for sensing touch events on a touch display panel; at least one output component such as a display, a vibration mechanism, and a lamp; an antenna; a power supply module, etc. Figure 3 The input components and output components listed are not shown in the figure. The hardware structure can be determined according to the type of electronic device and its functional requirements. This application does not list them one by one here.
[0067] Reference Figure 4 , is a flow chart of an optional example of a signal processing method proposed in this application, which can be executed by an electronic device, such as Figure 4 As shown, the method may include but is not limited to the following steps:
[0068] Step S41, when the audio player of the electronic device is in a playing state, obtaining the voice state of the user of the electronic device;
[0069] In combination with the above description of the technical solution of the present application, in order to avoid the situation where the signal energy of the audio played by the audio player is too high when the audio player is in the playback state and covers the audio signal energy of the local speaker (i.e., the user of the electronic device), resulting in the inability to recognize the audio of the local speaker, the present application proposes to monitor the voice status of the user of the electronic device when the audio player is in the playback state, and determine whether the user is speaking in the playback state, that is, whether the user is in the voice input state. The present application does not limit the method for obtaining the user's voice status.
[0070] In some embodiments, the present application can combine image recognition algorithms in artificial intelligence (AI) technology to monitor changes in the mouth shape of the user of the electronic device to determine whether the user is speaking; it can also use multiple distance sensors configured in the electronic device to sense changes in the distance of multiple consecutive positions of the user's mouth, thereby analyzing whether the user is speaking, etc. In some other embodiments, the present application can also pre-record the audio of the user of the electronic device and extract the user's voiceprint features. In this way, in actual applications, the audio collected by the audio collector can be used for voiceprint recognition to determine whether the user of the electronic device is speaking, etc. This can be determined based on scenario requirements, including but not limited to the methods for obtaining the voice status of these users described in this embodiment.
[0071] Step S42: Control the playback parameters of the audio player based at least on the voice state.
[0072] As analyzed above, when the audio player is in the playback state, if its playback parameters are inappropriate, such as the playback volume is too loud, the playback speed is close to the user's speaking speed, etc., the audio played by the audio player will cover the user's audio. In order to solve this problem, when it is determined that the user is speaking during the audio player playing audio, that is, the user is in the voice input state, it is necessary to control the playback parameters of the audio player so that they are at appropriate parameter values for the current scenario. The implementation method will not be described in detail in the embodiment of this application.
[0073] Based on this, the audio collector of the electronic device collects audio, and the mixed audio collected, that is, the audio played by the audio player and the audio of the user speaking in the same space and collected, can ensure that the user's audio is reliably identified and meet the application requirements. Therefore, the playback parameters in this application can at least be used by the electronic device to perform corresponding processing on the audio collected by its audio collector, such as performing echo cancellation processing on the collected audio based on different playback parameters to improve the voice call quality in the voice call scenario. This application does not limit the implementation method of the processing of the collected audio, which can be determined according to the situation.
[0074] According to the above analysis method, it is determined that the user is in a state of not inputting voice, that is, when the audio player of the electronic device is in the playing state, the user is not speaking. The user usually wants to listen to the content of the audio played by the audio player. In order to avoid interference of the playback content with the communication counterpart, the audio collector can be temporarily controlled to be in a mute state, or the electronic device can be prohibited from sending audio to the outside, or the audio collected during this period can be eliminated, etc. The implementation method is not limited in this application and can be determined according to the situation.
[0075] Reference Figure 5 , which is a flow chart of another optional example of the signal processing method proposed in this application. This embodiment can be an optional refined implementation method of the signal processing method described above, but is not limited to the refined implementation method described in this embodiment, and the method can still be executed by an electronic device, such as Figure 5 As shown, the method may include but is not limited to the following steps:
[0076] Step S51, when the audio player of the electronic device is in a playing state, obtaining the voice state of the user of the electronic device;
[0077] Regarding the implementation process of step S51, reference may be made to the description of the corresponding part above, which will not be elaborated here in this embodiment.
[0078] Step S52: If the user's voice state is a voice input state, control the audio player to be in a first playback parameter;
[0079] Step S53: If the user's voice status is no voice input, control the audio player to the second playback parameter;
[0080] Continuing the description of the above embodiment, according to the above processing method, during the process of the audio player playing audio, it is determined that the user is in a voice input state. In order to avoid the played audio covering the user's audio, it is possible to detect whether the current playback parameter of the audio player is the first playback parameter. If not, it can be adjusted to the first playback parameter; if the currently configured playback parameter is the first playback parameter, no adjustment is required. Afterwards, the audio generated by the user's speech is collected by the audio collector. Although the audio played by the audio player is also collected at the same time, the electronic device can perform corresponding processing on the actually collected mixed audio to meet actual application requirements, such as using echo cancellation technology to reliably separate the user's audio from the mixed audio. This application does not limit the processing method of the mixed audio collected by the audio collector.
[0081] If the user is not inputting any voice during the audio player's audio playback, it means that the user's audio does not need to be collected in the current scenario. In order to ensure that the user can reliably hear the played audio content, the audio player's playback parameters can be controlled to be at the second playback parameters, so that the signal energy value output by the audio player under the second playback parameters is higher than the signal energy value under the first playback parameters.
[0082] Therefore, in the actual application of the present application, during the audio player playing audio, since the audio collector of the electronic device is in working state, if the user speaks during this process, in order to reduce the interference of the playing audio, the audio player can be controlled to be in the first playback parameter, such as lowering the playback volume of the audio player, adjusting the playback speed, etc.; after the user finishes speaking and temporarily stops speaking, the previously adjusted playback parameter can be adjusted to the second playback parameter, such as increasing the playback volume, to ensure that the user can clearly and reliably hear the played audio content; after that, the user speaks again and enters the voice input state again, and the audio player is still in the playback state, and the playback parameters can continue to be adjusted as described above, so as to ensure the signal processing requirements for audio collection and audio playback at different stages in the entire voice communication environment.
[0083] It should be noted that the first and second playback parameters mentioned above can be determined based on the relative positional relationship between the audio player and the audio collector of the electronic device, the relative positional relationship between the user and the audio collector, and / or one or more combinations of voice input parameters such as the audio volume and speech speed of the user each time they speak (i.e., in the voice input state). Therefore, changes in the influencing factors listed herein, including but not limited to, may result in differences in the first and / or second playback parameters controlled above. This application does not limit the adjustment method of the playback parameters or the adjusted playback parameter values under different circumstances.
[0084] Exemplarily, for any type of electronic device, the present application can pre-count the voice input parameters (such as volume, speaking speed, etc.) adopted by different users when using such electronic devices for voice communication. Afterwards, the average value of the voice input parameters can be used to determine what playback parameters should be configured when the audio player is in the playback state so as not to affect the audio collected by the electronic device from the audio collector (that is, the mixed audio of the playback audio and the user's audio), and reliably identify the user's audio, that is, determine the above-mentioned first playback parameter, such as adjusting the playback volume to 50%, etc. The present application does not limit the value of the first playback parameter.
[0085] It can be understood that for different types of electronic devices, the way users use the electronic devices and the positional relationship between the audio player and audio collector of the electronic devices may be different, and the first playback parameters determined in the manner described above may be different. This requires pre-configuring the corresponding first playback parameters for different types of electronic devices.
[0086] When a user is not inputting a voice, the positional relationship between the user and the electronic device can be counted when different users use a certain type of electronic device. Based on this, the playback parameters of the electronic device's audio player can be determined to ensure that the user can clearly hear the audio content without reducing the user's experience due to excessive playback parameters such as playback volume and / or speech speed. This playback parameter is determined as the second playback parameter. Similarly, for electronic devices with different performance and different ways of using electronic devices by users, the second playback parameters configured according to the method described above may be different and can be determined depending on the situation.
[0087] It should be noted that in the process of configuring the playback parameters according to the method described above, the configured first or second playback parameter can be a certain parameter value or a parameter value range. In this way, the actual voice input parameters of the electronic device user can be used in actual applications to flexibly select the appropriate first playback parameter from the parameter value range of the preset first playback parameter, thereby improving the reliability of subsequent signal processing.
[0088] In some other embodiments proposed in the present application, according to the method described above, it is determined that the user of the electronic device is in a voice input state, and after controlling the audio player to be in the first playback parameter, if the audio is played according to the first playback parameter, the user can hear the audio content, and the audio player can be controlled to maintain the first playback parameter. In other words, when the user is no longer speaking, that is, in a state where no voice is input, there is no need to adjust the first playback parameter of the audio player to the second playback parameter. Of course, if the user cannot hear the played audio content clearly, the playback parameter of the audio player can also be adjusted to the second playback parameter or other playback parameter by means of buttons, voice control, etc., and is not limited to the processing method of step S53 above.
[0089] In some other embodiments, if the default audio player plays audio according to the second playback parameter and the user does not speak, this processing method can be maintained; if the user starts speaking, the audio played according to the second playback parameter is detected and does not interfere with the user's speech content, and the playback parameters of the audio player do not need to be adjusted. Or, if the current application scenario has strict requirements on the user's audio content, the audio player can be directly controlled to enter a silent state, that is, the parameter value of the first playback parameter is zero. Therefore, when the user speaks during the audio playback process of the electronic device, step S52 is not necessarily executed. Other processing methods can also be used to meet specific application requirements. This application will not be described in detail here.
[0090] Step S54 : performing corresponding processing on the audio collected by the audio collector of the electronic device so that the electronic device outputs a first audio to the communication terminal. The first audio does not include the audio played by the audio player collected by the audio collector.
[0091] Continuing from the above description, if the audio player of an electronic device plays audio and the user is in a voice input state, the audio collector of the electronic device will collect audio and will collect a mixed audio consisting of two types of audio. Afterwards, the collected audio can be processed according to the current application requirements. For example, in a voice communication scenario, the audio played by the audio player is usually the audio sent by the communication end (i.e., other devices that communicate with this electronic device for voice, such as electronic devices participating in voice communication, and / or communication servers that support voice communication functions, etc., as the case may be) received by the electronic device. This electronic device can use the received audio as a reference signal, perform echo cancellation processing on the collected audio, and send the first audio obtained after processing (which is usually the audio generated by the user's speech, excluding the audio played by the audio player collected by the audio collector) to the communication end.
[0092] Of course, in other application scenarios, such as audio recording, this noise reduction technology or other speech recognition technology can also be used to process the collected audio to obtain the desired target audio. For example, using speech synthesis technology, the user audio obtained after noise reduction can be voice-changed, and the resulting synthesized audio with the timbre of another specified user can be determined as the first audio. This is not limited to the scenario processing method of step S54. The implementation process can be determined in combination with the processing principle of the executed speech recognition technology. This embodiment will not be described in detail here.
[0093] It can be understood that when the user of the electronic device is in a voice input state, the audio player is in a silent state, and the electronic device can directly send the collected audio as the first audio to the communication end; or as described above, according to the specific application scenario requirements, the user audio is processed accordingly; and when the audio player is in a playing state, if the user is in a non-voice input state, in order to avoid interference of the played audio to the communication end, the electronic device may not output the audio collected by the audio collector (that is, the audio played by the audio player). To this end, the electronic device can still perform echo cancellation on the collected audio, thereby filtering out the audio played by the audio player, that is, the currently collected audio, so that the electronic device does not output audio; the audio collector can also be controlled to be in a silent state; or not respond to the audio output instruction and delete the collected audio.
[0094] In some other embodiments, in a multi-terminal electronic device voice communication scenario, after receiving the audio sent by the electronic device, the communication terminal of the electronic device can also compare it with the audio sent by the communication terminal, and not output audio with the same content as the audio sent by itself. In this way, the communication terminal avoids the situation where the audio collected and sent by itself is played. Therefore, after the electronic device receives the audio sent by the communication terminal, it can compare it with the historical audio collected by the electronic device within a specific time period before and sent to the communication terminal, and determine that the received audio contains the historical audio collected by the electronic device. The historical audio in the received audio can be filtered and sent to the audio player for playback.
[0095] Reference Figure 6 , is a flow chart of another optional example of the signal processing method proposed in this application. This embodiment can be another optional detailed implementation method of the signal processing method described above. Different from the playback parameter control implementation method described in the detailed embodiment above, this method can still be executed by an electronic device, such as Figure 6 As shown, the method may include but is not limited to the following steps:
[0096] Step S61, when the audio player of the electronic device is in a playing state, obtaining the voice state of the user of the electronic device and the form of the electronic device;
[0097] Regarding the method for obtaining the voice status of the user of the electronic device, it can be combined with the detailed description of the corresponding part of the context embodiment, and will not be described in detail in this embodiment.
[0098] In the embodiments of the present application, as analyzed above, the relative positional relationship between the audio collector and the audio player in the electronic device directly affects the echo cancellation effect, and the audio player and the audio collector may be located on different main body structures of the electronic device. As the form of the electronic device changes, the relative positional relationship between the audio player and the audio collector changes. That is to say, under different electronic device forms, the relative positional relationship between the audio collector and the audio player is different. If the playback parameter control method described in the above embodiment is still followed, it may affect the subsequent echo cancellation effect of the collected mixed audio.
[0099] Therefore, when the audio player is in the playback state, the embodiment of the present application can monitor the shape of the electronic device, thereby determining the relative positional relationship between the audio player and the audio collector of the electronic device. The present application does not limit the method for obtaining the shape of the electronic device. Optionally, the current shape of the electronic device can be determined based on the parameters sensed by sensor modules such as gyroscopes and attitude sensors configured in the electronic device; of course, if a corresponding conversion instruction is generated when the shape of the electronic device changes, the current shape of the electronic device can be determined based on the conversion instruction, etc. This application does not provide detailed examples one by one.
[0100] Step S62, determining that the user is in a voice input state and the electronic device is in a first state, and controlling the audio player to a first playback parameter;
[0101] Step S63, determining that the user is in a voice input state and the electronic device is in the second state, and controlling the audio player to a third playback parameter;
[0102] In the embodiment of the present application, the electronic device is taken as an example to illustrate that it is a terminal with a display screen. Figure 7 The all-in-one computer shown in the figure can have an audio player on the base and an audio collector on its display unit; or the audio collector on the base and the audio player on the side and / or back of the display unit. Figure 7 In the first state (i.e. vertical screen state) shown on the right, Figure 7 In the second form (i.e., horizontal screen state) shown on the left, the relative position relationship between the audio collector and the audio player will change, and the effect of echo cancellation on the collected audio containing the playback audio under the same playback parameters is often different.
[0103] Assume that compared to an electronic device in the second form, if it is in the first form, the distance between the audio collector and the audio player of the electronic device is smaller, that is, when the electronic device switches from a horizontal screen state to a vertical screen state, the distance between its audio collector and the audio player will be reduced, which can increase the echo interference to a certain extent. When adjusting the playback parameters of the audio player, the signal energy value output by the audio player under the third playback parameter can be made higher than the signal energy value under the first playback parameter. Taking the playback parameter volume as an example, if the user speaks during audio playback, the playback volume configured for the audio player controlled by the electronic device in the horizontal screen state is greater than the playback volume configured for the audio player controlled by the electronic device in the vertical screen state. However, this application does not limit the values of the first playback parameter and the second playback parameter in these two cases.
[0104] It is understandable that if the distance between the audio collector and the audio player of the electronic device is greater when it is in the first form relative to the distance between the audio collector and the audio player of the electronic device when it is in the second form, that is, when the electronic device switches from a landscape mode to a portrait mode, the distance between the audio collector and the audio player increases. Then, the signal energy value output by the audio player under the third playback parameter is lower than the signal energy value under the first playback parameter. Therefore, the numerical relationship between the first playback parameter and the second playback parameter can be determined based on the relative positional relationship between the audio collector and the audio player represented by the first and second forms, which here refers to the relative distance between the two devices.
[0105] Among them, regarding the specific method of obtaining the above-mentioned first playback parameters and third playback parameters, you can refer to the description of the method of obtaining the first playback parameters and the second playback parameters in the above embodiment, and in the acquisition process, in addition to considering the user's voice input parameters, the embodiment of the present application can also consider the form of the electronic device (that is, the relative position relationship between the audio collector and the audio player) to determine, that is, based on the user's voice input parameters and the form of the electronic device, determine when the user of the electronic device is in a voice input state and the audio player is in a playback state, the first playback parameters and the third playback parameters required by the audio player. The implementation process is not described in detail in this application.
[0106] In some other embodiments proposed in the present application, for applications of a type of electronic device in which a change in the form of the electronic device will cause a change in the relative positional relationship between the audio player and the audio collector, when the audio player is in the playing state and the playback parameters of the audio player are controlled, the above-mentioned steps S62 and S63 are not constrained to be executed in the same application scenario. That is, when the user is in the voice input state and the step S62 or step S63 is executed, after the form of the electronic device changes, the other of the two steps is not necessarily executed according to the method described in this embodiment. As described in the corresponding parts of step S52 and step S53 in the above embodiment, other control methods can also be used to achieve control of the audio player, such as controlling the audio player to be in a mute state.
[0107] In some other embodiments, during the entire voice communication process, the user may not speak all the time. According to the voice state acquisition method described above, it is determined that the user is in a state of no voice input, that is, the user is no longer speaking. The playback parameters of the audio player can be adjusted according to the processing method described in step S53 above; or, based on the processing method described in step S53, new second playback parameters can be determined in combination with the form of the electronic device to ensure that the user can reliably hear the content of the audio played by the audio player. The implementation process is not described in detail in the embodiments of this application.
[0108] In order to avoid frequent adjustments to the playback parameters of the audio player, that is, when the user pauses in speaking (such as a shorter time such as 2s), the audio player is controlled to be in the second playback parameter according to the method described above; after the pause, the audio player is controlled to enter the first playback parameter or the third playback parameter, etc., resulting in a waste of resources and a reduced user experience. The present application can also configure the electronic device to switch from the voice input state to the non-voice input state, and the preset time length that needs to be maintained in the non-voice state. If the statistical time length of the user in the non-voice input state reaches the preset time length, it can be considered that the user is no longer speaking at the current stage, and then the audio player is controlled to be in the second playback parameter according to the method described above. The present application does not limit the value of the preset time length, and it can be determined based on the user's speaking speed, etc.
[0109] Step S64, performing corresponding processing on the audio collected by the audio collector of the electronic device to obtain a first audio;
[0110] Step S65: Send the first audio to the communication terminal of the electronic device.
[0111] Regarding the implementation process of step S64 and step S65, please refer to the description of the corresponding parts of the above embodiment, and this embodiment will not be repeated. It is understood that the audio processing method involved in step S64 includes but is not limited to signal processing technologies such as echo cancellation and speech synthesis, which can be determined according to the application scenario requirements.
[0112] It should be understood that for different playback parameters, in the above-mentioned application scenarios, when the audio collected by the audio collector (i.e., the mixed audio generated by the simultaneous existence of multiple types of audio) is subjected to echo cancellation processing, due to the different signal energy and / or signal energy changes of the audio played by the collected audio player contained in the audio, when it is compared with a known reference signal, the judgment standard for echo noise can be adjusted accordingly based on the comparison result. The specific criteria can be determined in combination with the working principle of echo cancellation, which will not be described in detail in this application.
[0113] Optionally, after the electronic device obtains the audio to be played by the audio player, it can configure a corresponding reference signal based on the different playback parameters of the audio player. In this way, after controlling the playback parameters of the audio player according to the actual situation according to the method described above, when performing echo elimination on the audio collected by the audio collector, the reference signal corresponding to the playback parameter can be called to implement it. The implementation process is not described in detail.
[0114] In the actual application of this application, combined with the above analysis, in the process of controlling the playback parameters of the audio player, the audio properties such as the collected user's audio signal energy will also affect the processing effect of the collected audio. Therefore, in addition to considering the user's voice status and the form of the electronic device as described in the above embodiments, this application can also consider other factors, such as the positional relationship between the user and the electronic device, or even the positional relationship between the user and the audio collector of the electronic device, such as the relative distance.
[0115] Based on this, refer to Figure 8 As shown, it is a flow chart of another optional example of the signal processing method proposed in this application. This embodiment can be another optional refined implementation method of the signal processing method described above. Different from the playback parameter control implementation method described in the refined embodiment above, this method can still be executed by an electronic device, such as Figure 8 As shown, the method may include:
[0116] Step S81, when the audio player of the electronic device is in a playing state, obtaining the voice state of the user of the electronic device and the positional relationship between the user and the electronic device;
[0117] Regarding the method for obtaining the user's voice status, please refer to the description of the corresponding embodiment in the context. The positional relationship between the user and the electronic device can include the relative distance between the two. It can be based on the user's video data (which can be obtained by an image collector (such as a camera) configured by the electronic device, or obtained by an independent image acquisition device locally configured by the electronic device and sent to the electronic device, etc. This application does not limit the method for implementing the acquisition of video data), and the positional relationship between the user and the electronic device can be determined by image analysis. For example, distance detection is achieved by a monocular or binocular camera, etc. This application does not elaborate on how to use an image collector to achieve distance detection.
[0118] Optionally, the present application may also determine the positional relationship between the user and the electronic device based on the parameters sensed by the distance sensor (such as infrared or ultrasonic or Tof (Time of flight) sensor, etc.) in the electronic device. In order to improve the detection accuracy, multiple distance sensors may be configured, such as multiple distance sensors arranged in an array. When the user is within the distance sensing range of the distance sensor, the positional relationship between the user and the electronic device is determined by analyzing the distance between each distance sensor and the corresponding position point (i.e., the position point on the user's body). Changes in the positional relationship between the user and the electronic device, such as changes in distance, may be monitored as needed. The implementation process will not be described in detail.
[0119] It should be noted that the method for detecting the positional relationship between the user and the electronic device includes but is not limited to the image analysis and distance sensing implementation methods described above. In some other embodiments, the sound source (i.e., user) positioning method or other positioning devices carried by the user can also be used to determine the positional relationship between the user and the electronic device. This application does not provide detailed examples one by one. It is understandable that the analysis and determination process of the positional relationship between the user and the electronic device can be performed by the electronic device, or it can be obtained by other devices and sent to the electronic device in real time. This application does not limit this.
[0120] Among them, according to the position relationship acquisition method described above, this application can obtain the position relationship between the user and the audio collector of the electronic device to represent the position relationship between the above-mentioned user and the electronic device, but it is not limited to the audio collector representing the electronic device, and it may depend on the situation.
[0121] Step S82, determining that the user is in a voice input state and the user and the electronic device are in a first positional relationship, and controlling the audio player to be in a fourth playback parameter;
[0122] Step S83, determining that the user is in a voice input state and the user and the electronic device are in a second positional relationship, and controlling the audio player to be in a fifth playback parameter;
[0123] In the embodiment of the present application, it is assumed that the first position relationship indicates that the distance between the user and the electronic device is farther than the second position relationship described above. That is, the user is close to the electronic device, indicating that the position relationship between the user and the electronic device changes from the first position relationship to the second position relationship. The user is in a voice input state, and the signal energy of the user's audio collected by the audio collector of the electronic device will become higher and higher. When the relative position relationship between the audio collector of the electronic device and the audio player remains unchanged, the user's audio will be less disturbed by the echo noise of the audio played with the same playback parameters. Therefore, while ensuring that the user's audio can be reliably identified from the mixed audio later, the requirements for the playback parameters of the audio player are relatively low, and the signal energy value output under the fifth playback parameter of the audio player can be controlled to be higher than the signal energy value output under the fourth playback parameter.
[0124] Based on this, taking the volume as an example of the above-mentioned playback parameter, when the audio player is in the playback state and the user is in the voice input state, if the distance between the user and the electronic device is relatively far (the actual detected distance value can be compared with the preset distance threshold. If it is greater than or equal to the distance threshold, it is considered to be in the first position relationship; conversely, if it is less than the distance threshold, it can be considered to be in the second position relationship, but it is not limited to this detection method), the audio player can be controlled to a lower volume; if the distance between the user and the electronic device is relatively close (that is, in the second position relationship), the audio player can be controlled to a relatively high volume (which is usually less than the volume of the audio played by the audio player when the user is not inputting voice), so as to avoid the audio played by the player covering the user's audio, resulting in the subsequent inability to identify the user's audio from the actually collected audio.
[0125] It should be understood that when an audio player plays the same audio, the higher the volume configured, the higher the signal energy of the played audio; conversely, the lower the volume, the lower the signal energy of the played audio. Furthermore, the aforementioned playback parameters include, but are not limited to, audio volume and may also include other audio attribute parameters as needed, which are not detailed in this application.
[0126] In some other embodiments proposed in the present application, before controlling the playback parameters of the audio player, the present application can also comprehensively analyze the three influencing factors of the user's voice state, the form of the electronic device (i.e., the relative position relationship between the audio player and the audio collector), and the position relationship between the user and the electronic device (such as the audio collector), to determine the user's voice input state, the relative distance between the user and the audio player and the audio collector, etc., to control the playback parameters of the audio player. The implementation process can be implemented in combination with the description of the corresponding parts of the above two embodiments, and this application will not give detailed examples.
[0127] Among them, based on one or more combined influencing factors listed above, the playback parameters of the audio player can be controlled according to the correspondence between the playback parameters pre-configured based on the corresponding influencing factors, that is, based on different voice input parameters of the user, different distances between the user and / or the audio player and the audio collector, etc., the audio playback parameters corresponding to the audio player when the echo cancellation effect meets the preset requirements are determined. The implementation process can be combined with the description of the corresponding part of the above embodiment, and this embodiment will not be described in detail here.
[0128] Step S84: performing corresponding processing on the audio collected by the audio collector so that the electronic device outputs the first audio to the communication end.
[0129] The implementation process of step S84 can refer to the description of the corresponding part of the above embodiment, and will not be repeated here in this embodiment.
[0130] In summary, in the embodiment of the present application, when the audio player of the electronic device is in the playback state, if the user is in the voice input state, the positional relationship between the user and the electronic device will be taken into consideration, that is, the audio generated by sound sources at different distances will be interfered with by the same played audio, and the playback parameters of the audio player will be adaptively adjusted, thereby reliably ensuring the processing efficiency of the collected audio in this scenario and better meeting application requirements.
[0131] For the implementation method of obtaining the voice status of the user of the electronic device in each of the above embodiments, the present application can be implemented through any one or more combinations of image analysis, distance sensing, operation or status of the controls acting on the electronic device, etc. The implementation process can refer to but is not limited to the description of the corresponding part of the embodiment below. Regarding the implementation method of controlling the playback parameters of the audio player in the signal processing method, reference can be made to the description of the corresponding part of the embodiment above, and the following embodiment will not be repeated.
[0132] In some embodiments, the present application can determine the user's voice state by monitoring changes in the mouth contour of the user of the electronic device. Therefore, the user's mouth contour change information can be obtained based on at least parameter information collected by a target sensor of the electronic device. The target sensor can be an image collector, an infrared sensor, a Tof sensor, etc. The parameter information collected by different types of target sensors may be represented in different forms, but it can be used to characterize changes in the user's mouth contour. The present application does not elaborate on the implementation process of how the target sensor collects parameter information.
[0133] Reference Figure 9 , is a flow chart of another optional example of the signal processing method proposed in this application. This embodiment can be another optional refinement of the signal processing method described above, and can be a refinement of the method for acquiring the user's voice status. This embodiment is described by taking the above-mentioned target sensor as an image collector as an example, and is not limited to the refinement of the implementation method described in this application. The method can still be executed by an electronic device, such as Figure 9 As shown, the method may include:
[0134] Step S91: When the audio player of the electronic device is in a playing state, obtain video data of the user of the electronic device; the video data at least includes video data of the user's mouth;
[0135] Step S92, obtaining the user's mouth contour change information based at least on the user's mouth video data;
[0136] Step S93, using the mouth contour change information to determine the user's voice state;
[0137] In embodiments of the present application, the parameter information collected by the target sensor may be video data, and the user's video data may be collected by an image collector in the electronic device, or collected by an independent image collector distinct from the electronic device and then transmitted to the electronic device. This application does not restrict the method for acquiring the user's video data. To ensure that the collected video data includes at least video data of the user's mouth, tracking detection technology may be combined to dynamically control the image acquisition range of the image collector. This implementation process is not described in detail in this application.
[0138] In some embodiments, such as in voice communication application scenarios such as video conferencing, refer to Figure 10 As shown in the signal processing flow diagram, the image collector of the electronic device is in a shooting state, and obtains video data within the shooting range in real time, such as the video data of the user, and sends it or the processed audio of the user to the communication server, which forwards it to other electronic devices participating in the meeting for output, so that each electronic device participating in the meeting outputs Figure 10The conference interface shown in the figure outputs the audio of the current speaking user. This application does not elaborate on the communication principle of multi-party video conferencing.
[0139] Based on this, for any electronic device participating in a video conference, to determine whether its audio player is playing audio, in order to detect whether the user of the electronic device is speaking, the video data captured by the image collector can be analyzed to determine the user's mouth video data (i.e., multiple consecutive frames of mouth images) so that the user's mouth contour change information (i.e., mouth shape change) can be subsequently analyzed based on this. This application does not elaborate on how to implement the method of determining the user's mouth shape change through image analysis.
[0140] It should be noted that for other voice communication application scenarios different from video conferencing, the implementation process of obtaining the user's mouth video data is similar, and the embodiments of this application will not be described in detail here. It can be understood that if the audio player is in the playback state and the image collector of the electronic device is in the off state, an image acquisition instruction can be sent to the image collector to control the image collector to enter the shooting state (i.e., the image acquisition state), and after controlling the image acquisition direction of the image collector to face the user's face, the user is imaged and the user's video data is obtained.
[0141] Subsequently, by analyzing the user's mouth contour change information, it is possible to determine whether the user is speaking, that is, whether the user is in a voice input state or a non-voice input state. The implementation process is not described in detail. Generally, if the mouth contour change information is determined to meet the mouth contour change conditions required for emitting valid audio, the user can be considered to be in a voice input state. Conversely, if the mouth contour does not change or the shape of the change is fixed, resulting in the mouth contour change conditions not being met for emitting valid audio, the user can be considered to be in a non-voice input state. However, this analysis implementation method is not limited to this.
[0142] In some other embodiments proposed in the present application, since the user's facial expressions usually change with the changes in the content of the speech during the speech, in order to improve the detection accuracy of the user's voice state (i.e., voice input state or no voice input state), in addition to obtaining the mouth contour change information, the embodiments of the present application can also obtain the expression changes in the user's facial area, thereby comprehensively determining whether the user is in the voice input state. The implementation process will not be described in detail.
[0143] Step S94: Control the playback parameters of the audio player based at least on the voice state.
[0144] The implementation process of step S94 can refer to the description of the corresponding part of the above embodiment, and will not be repeated in this embodiment.
[0145] For example, Figure 10 As shown, according to the method described above, the driver of the audio player of the electronic device determines that the user is speaking, that is, in the voice input state, and the audio collector is in the collection state. In order to avoid the playback audio covering the user audio and affecting the subsequent processing effect, the driver can control the audio player to be in the first playback parameter, such as lowering the volume of the audio player; conversely, if it is determined that the user is in the state of not inputting voice, even if the audio collector is in the collection state, the playback parameters of the audio player do not need to be adjusted, so that it is maintained at the second playback parameter, ensuring that the user can reliably hear the played audio content.
[0146] Optionally, according to the method described above, when it is determined that the playback parameters of the audio player need to be adjusted, the electronic device may also output corresponding playback parameter adjustment prompt information through output methods such as text or indicator lights to remind the user to lower or increase the volume of the audio player. This application does not limit this prompt implementation method.
[0147] In some other embodiments proposed in the present application, in order to determine the user's voice status, when detecting the user's mouth contour change information, the parameter information collected by the target sensor configured by the electronic device, such as infrared, ultrasonic or Tof sensor, can also be used. For the convenience of description, such sensors can be recorded as distance sensors. In actual applications, as needed, multiple distance sensors can be configured in an array or other regular manner. When it is determined that the audio player is in the playback state, the parameter information collected by each of the multiple distance sensors (such as the sensing distance parameter) can be obtained, that is, the distance value between the corresponding distance sensor and the position point in its detection direction (such as the position point on the user's body, which can be the position point in the mouth area) can be obtained. Afterwards, the user's mouth contour change information can be obtained at least based on the obtained parameter information. The implementation process is not described in detail in this application.
[0148] It should be noted that in order to improve the reliability and accuracy of mouth movement detection, the user can be prompted to adjust the relative position between his or her mouth and the distance sensor based on the parameter information or other position identifiers sensed by the distance sensor to ensure that the distance sensing range of the distance sensor at least includes the user's mouth area. The method of implementing the prompt adjustment is not limited in this application and can be determined according to the circumstances.
[0149] Regarding determining the user's voice state after obtaining mouth contour change information, and even controlling the playback parameters of the audio player based on this information, please refer to the description of the corresponding part of the above embodiment, and this embodiment will not be repeated here. In some embodiments, this application can also combine the above two mouth contour change information detection results to determine the user's voice state, depending on the situation.
[0150] In some other embodiments proposed in this application, in order to determine the user's voice status, this application can also determine the user's voice status based on the comparison result of the audio collected by the audio collector of the electronic device and the preset voiceprint information of the user of the electronic device. That is to say, before the user uses the electronic device for voice communication, the audio player can be put in muted state, and the audio of the user can be collected by the audio collector, and the voiceprint feature can be extracted to obtain the preset voiceprint information of the user and then stored. Of course, the preset voiceprint information of the user can also be obtained from other channels, and this application does not limit its acquisition method.
[0151] Afterwards, when it is determined that the audio player of the electronic device is in the playback state, the audio collected by the audio collector can be obtained, the voiceprint information contained therein can be extracted, and it can be compared with the user's preset voiceprint information. If the similarity between the voiceprint information of the collected audio and the preset voiceprint information is greater than the similarity threshold, it can be considered that the user is in the voice input state; conversely, if the similarity is less than or equal to the similarity threshold, it can be considered that the user is in the non-voice input state. For the above-mentioned voiceprint feature extraction and voiceprint comparison implementation methods, appropriate artificial intelligence technology can be selected for implementation, and this application does not impose any restrictions on this.
[0152] Optionally, the present application can pre-build a voiceprint recognition model, input the audio collected by the audio collector into the voiceprint recognition model, and output whether the collected audio contains the user's audio, thereby determining the user's voice status. Among them, the voiceprint recognition model can be obtained by training sample audio based on voiceprint recognition algorithms, machine learning algorithms / deep learning algorithms, etc. in artificial intelligence technology. The present application does not limit the training implementation method of the voiceprint recognition model. In order to improve the reliability of the output results of the voiceprint recognition model, during the training process, sample audio obtained by the same user using voice input information such as different volume, timbre, and speaking speed can be considered to more accurately identify each speaker. The implementation process is not described in detail.
[0153] In some other implementations, the audio collector of the electronic device can be turned on only when the user needs to speak and is about to enter the voice input state; if the user does not need to explain, the audio collector can be turned off, thereby avoiding the resource consumption caused by the electronic device filtering the noise collected by the audio collector when the user is not speaking. Based on this, in order to determine the user's voice state, the present application can determine whether the user is in the voice input state based on the operation (such as on, off, input, not input, etc.) or state of the controls acting on the electronic device (such as control icons or physical control keys for adjusting the working state of the audio collector, etc.), that is, determine the voice state of the user of the electronic device.
[0154] In the actual application of this embodiment, in various voice communication scenarios such as video conferencing, the user of the electronic device needs to speak, which can trigger the above control and control the audio collector to enter the audio collection state by generating a start instruction or recording instruction for the audio collector. At the same time, it can be considered that the user is in the voice input state, and the playback parameters of the audio player can be controlled in the manner described above; conversely, if the user stops speaking, the above control can be triggered to generate a shutdown instruction or a mute instruction (i.e., no recording instruction) for the audio collector, thereby controlling the audio collector to be in a non-audio collection state, which can be a mute state. At this time, it can be considered that the user is in a non-voice input state, and the playback parameters of the audio player can be controlled accordingly according to the above method. In this case, the audio collector will not collect the audio played by the audio player, and the electronic device will not output audio to its communication end.
[0155] The method for detecting the operation of the above-mentioned control can be determined by a trigger signal generated based on the operation, or by detecting the operating status of the audio collector or the status of the control, etc., to determine the operation of the control, that is, to determine the control operation of the audio collector, thereby determining the user's voice status. This application does not limit the implementation method of detecting the control operation or the above-mentioned status.
[0156] Based on the description of each embodiment above, in the process of executing the signal processing method provided by each embodiment, according to the corresponding method described in each embodiment above, when it is determined that the user is in a voice input state, in order to fundamentally solve the problem that the audio player plays audio at this time, and the played audio interferes with the audio output by the user, the present application can control the audio player to be in a silent state. Of course, if the user needs to know the content of the played audio at this time, the audio to be played of the audio player can be converted into text information, such as by using artificial intelligence technologies such as speech recognition and machine learning. Optionally, a pre-trained audio-to-text conversion model is retrieved, and the obtained audio to be played of the audio player (such as the audio sent by the communication terminal) is input into the conversion model to obtain the corresponding text information, that is, the audio content.
[0157] Afterwards, the electronic device can output the text information through the display screen, such as presenting the text information in a pop-up text prompt window, or presenting the text information on the interface corresponding to the source of the audio to be played (such as the conference interface corresponding to the speaker in a multi-party video conference, etc.), so that the user of the electronic device can make feedback by viewing the text information, thereby improving the efficiency and quality of voice communication. The implementation process is not described in detail in this application.
[0158] According to the detection method described above, if it is determined that the user switches from a voice input state to a non-voice input state, that is, the current user is in a non-voice input state, the audio player can be controlled to switch from a silent state to a play state to meet the user's communication needs for listening to and playing audio. In this case, the audio player can be controlled to be in the second play parameter, or other default play parameters to ensure that the user can reliably hear the content of the audio played by the audio player. Regarding the control implementation process of the audio player, you can refer to the description of the above embodiment, which will not be repeated in this embodiment.
[0159] In some other embodiments proposed in the present application, for the control process of the audio playback parameters of the audio player described in the above embodiments, if it is determined that the user is in a voice input state, the control process of the playback parameters of the audio player can be a refined processing of the control process of the above-mentioned first playback parameter, third playback parameter, fourth playback parameter and / or fifth playback parameter.
[0160] When the audio player is in the playback state and the user is in the voice input state, this embodiment can first detect whether the current playback parameters of the audio player need to be adjusted before adjusting the playback parameters of the audio player to the corresponding first playback parameter, third playback parameter, fourth playback parameter, or fifth playback parameter. Therefore, the present application can determine the parameter threshold that the playback parameters of the audio player at least reach when the audio played by the audio player is used to indicate that the audio will interfere with the user's audio, resulting in the subsequent processing results of the collected audio (such as the echo cancellation result) failing to meet the processing requirements.
[0161] Based on this, in actual applications, the parameter threshold of the current playback parameter of the audio player can be directly obtained, and the current playback parameter can be compared with the parameter threshold. If the current playback parameter reaches the parameter threshold, the current playback parameter of the audio player can be adjusted to the corresponding preset playback parameter (in the control implementation process of other playback parameters as described in different embodiments above, the preset playback parameter can correspond to the first playback parameter or the third playback parameter or the fourth playback parameter or the fifth playback parameter, etc., which can be determined as the situation); if the current playback parameter does not reach the parameter threshold, the current playback parameter can be determined as the preset playback parameter, and the audio player can be controlled to maintain the preset playback parameter unchanged. That is to say, when the user is in the voice input state, the audio player plays the audio according to the current playback parameter, and then performs corresponding processing on the audio collected by the audio collector, and the processing effect can meet the corresponding processing requirements, such as Figure 1 In the scenario shown in the figure below, there is no need to adjust the current playback parameters of the audio player, which reduces the processing steps.
[0162] Among them, the above-mentioned parameter threshold can be a pre-configured value (the size of which can be determined through experiments and is not limited in this application), which can be called directly; it can also be configured online to improve processing reliability, such as determining the parameter threshold based on the audio attribute value of the audio collected by the user when the audio player is in silent state. In this way, the parameter threshold can be adaptively configured for the audio attribute values such as the volume, speaking speed, and timbre of the user's speech at the current stage, so that the parameter threshold is more in line with the audio processing effect monitoring at the current stage, thereby improving signal processing reliability. The method for obtaining the above-mentioned parameter threshold includes but is not limited to the two implementation methods described above, which can be determined according to the scenario requirements.
[0163] Optionally, after adjusting the current playback parameters of the audio player to the preset playback parameters according to the method described above, if it is detected that the user switches from the voice input state to the non-voice input state, the preset playback parameters of the audio player are restored to the playback parameters before adjustment, such as adjusting from the first playback parameter to the second playback parameter or the default playback parameter, etc. The specific control method of the playback parameters can be determined according to the application scenario, and this application does not give detailed examples one by one.
[0164] To sum up, taking the multi-party video conferencing scenario as an example, for any electronic device and its user participating in the video conference, during the entire video conference, the electronic device can synchronously play the audio of the video conference. If multiple users speak, for any of the users' electronic devices, it will not only play the audio of each user's speech, but also need to collect the audio of the local user's speech. In the audio collection process, Figure 1 In the scenario shown in the figure above, if the speaker volume of the local electronic device is too loud, the user's voice will be drowned out by the sound played by the speaker. When the electronic device performs echo cancellation, the local user's audio and echo will be eliminated together, making it impossible for other users participating in the meeting to hear what the user is saying.
[0165] In this case, this application will use any one or more combinations of image analysis, voiceprint recognition, control monitoring, etc. to monitor the local user's mouth contour changes during the speaker playing audio to determine whether the local user is speaking. If the local user is speaking, the speaker driver can be notified to automatically lower the speaker volume, such as Figure 1The scenario shown in the accompanying figure below ensures that the local user's voice is not drowned out by the sound played by the speaker. During the subsequent echo cancellation process, the speaker audio can be reliably filtered, retaining the audio of the local user's speech, and sending it to other electronic devices participating in the video conference for playback, ensuring that other users can hear the local user's speech. It can be understood that for any electronic device participating in the video conference, the signal processing method proposed in this application can be executed to ensure the quality of voice communication. Of course, for other voice communication application scenarios, the implementation process of the signal processing method is similar, and this application will not further explain it in detail.
[0166] Reference Figure 11 , is a schematic structural diagram of an optional example of a signal processing device proposed in this application, the device may include:
[0167] The voice status obtaining module 111 is used to obtain the voice status of the user of the electronic device when the audio player of the electronic device is in the playing state;
[0168] a playback parameter control module 112, configured to control playback parameters of the audio player based at least on the voice state;
[0169] The playback parameters are at least used by the electronic device to perform corresponding processing on the audio collected by its audio collector.
[0170] In some embodiments, the playback parameter control module 112 may include:
[0171] A first control unit is configured to control the audio player to be in a first playback parameter if the user is in a voice input state; and / or,
[0172] a second control unit, configured to control the audio player to a second playback parameter if the user is in a state of not inputting a voice;
[0173] The signal energy value output by the audio player under the second playback parameter is higher than the signal energy value output by the audio player under the first playback parameter.
[0174] In some further embodiments, the playback parameter control module 112 may include:
[0175] A third control unit is configured to control the audio player to be in a first playback parameter if the user is in a voice input state and the electronic device is in a first state; or
[0176] a fourth control unit, configured to control the audio player to a third playback parameter if the user is in a voice input state and the electronic device is in a second state; or
[0177] a fifth control unit, configured to control the audio player to be in a fourth playback parameter if the user is in a voice input state and the user and the electronic device are in a first positional relationship; or
[0178] a sixth control unit, configured to control the audio player to a fifth playback parameter if the user is in a voice input state and the user and the electronic device are in a second positional relationship;
[0179] Among them, the signal energy value output by the audio player under the third playback parameter is higher than the signal energy value under the first playback parameter, and the signal energy value output by the audio player under the fifth playback parameter is higher than the signal energy value under the fourth playback parameter.
[0180] Based on the description of the above embodiments, the signal processing device may further include:
[0181] an audio processing module, configured to perform corresponding processing on the audio collected by the audio collector so that the electronic device outputs the first audio to the communication terminal, or, when the user is not inputting a voice, the electronic device does not output the audio collected by the audio collector;
[0182] The first audio does not include the audio played by the audio player and collected by the audio collector.
[0183] In some further embodiments, the voice status obtaining module 111 may include:
[0184] a parameter information acquisition unit, configured to acquire parameter information collected by a target sensor of the electronic device; and a mouth contour change information acquisition unit, configured to acquire mouth contour change information of the user based at least on the parameter information;
[0185] A first determining unit is configured to determine the user's voice state by using mouth contour change information; or
[0186] A second determining unit is configured to determine a voice state of a user of the electronic device based on an operation or state of a control acting on the electronic device; or
[0187] a voiceprint information comparison unit, configured to compare the audio collected by the audio collector of the electronic device with the preset voiceprint information of the user of the electronic device to obtain a comparison result;
[0188] The third determining unit is configured to determine the user's voice status based on the comparison result.
[0189] Based on the description of the above embodiments, the above apparatus may further include:
[0190] A text information acquisition module is used to control the audio player to be in a mute state if the user is in a voice input state, and convert the audio to be played by the audio player into text information;
[0191] A text information output module, configured to output the text information;
[0192] The playing state switching module is used to control the audio player to switch from the mute state to the playing state if the user is in a state of not inputting voice.
[0193] In some further embodiments, the playback parameter control module 112 may include:
[0194] a parameter threshold acquisition unit, configured to acquire a parameter threshold for a current playback parameter of the audio player when the user is in a voice input state; wherein the parameter threshold is a pre-configured value; or is determined based on an audio attribute value of audio collected by the user when the audio player is in a muted state;
[0195] a playback parameter adjustment unit, configured to adjust the current playback parameter of the audio player to a preset playback parameter if the current playback parameter reaches a parameter threshold; the preset playback parameter being the first playback parameter, the third playback parameter, the fourth playback parameter, or the fifth playback parameter; and / or,
[0196] The playback parameter maintaining unit is used to determine the current playback parameter as the preset playback parameter if the current playback parameter does not reach the parameter threshold, and control the audio player to maintain the preset playback parameter unchanged.
[0197] Optionally, the playback parameter control module 112 may further include:
[0198] The playback parameter recovery control unit is used to restore the preset playback parameters of the audio player to the playback parameters before adjustment if the user switches from the voice input state to the non-voice input state.
[0199] It should be noted that the various modules, units, etc. in the above-mentioned device embodiments can be stored in the memory as program modules, and the processor executes the above-mentioned program modules stored in the memory to implement the corresponding functions. Regarding the functions implemented by each program module and its combination, as well as the technical effects achieved, please refer to the description of the corresponding parts of the above-mentioned method embodiments, which will not be repeated in this embodiment.
[0200] The present application also provides a computer-readable storage medium on which computer-readable instructions can be stored. The computer-readable instructions can be called and loaded by a processor to implement the various steps of the signal processing method described in the above embodiment.
[0201] Finally, it should be noted that, in the above embodiments, unless the context clearly indicates an exception, the terms "a," "an," "an," and / or "the" do not specifically refer to the singular but also include the plural. Generally speaking, the terms "comprise" and "include" only indicate the inclusion of the steps and elements explicitly identified, and these steps and elements do not constitute an exclusive list. A method or device may also include other steps or elements. The phrase "comprises a..." does not preclude the presence of other identical elements in the process, method, product, or device that includes the elements.
[0202] In the description of the embodiments of this application, unless otherwise specified, " / " represents or. For example, A / B can represent A or B. "And / or" in this article is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "plurality" means two or more than two.
[0203] Terms such as "first" and "second" used in this application are used for descriptive purposes only to distinguish one operation, unit, or module from another operation, unit, or module, and do not necessarily require or imply any actual relationship or order between these units, operations, or modules. They should not be understood as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of such features.
[0204] In addition, the various embodiments in this specification are described in a progressive or parallel manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices and electronic devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0205] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A signal processing method, comprising: When the audio player of the electronic device is in a playing state, obtaining the voice state of the user of the electronic device; controlling playback parameters of the audio player based at least on the voice state; The playback parameters are at least used by the electronic device to perform corresponding processing on the audio collected by its audio collector; The controlling of the playback parameters of the audio player based at least on the voice state includes at least one of the following: If the user is in a voice input state and the electronic device is in a first state, controlling the audio player to be in a first playback parameter; If the user is in a voice input state and the electronic device is in a second state different from the first state, controlling the audio player to be in a third playback parameter, wherein the signal energy value output by the audio player under the third playback parameter is higher than the signal energy value under the first playback parameter; If the user is in a voice input state and the user and the electronic device are in a first positional relationship, controlling the audio player to be in a fourth playback parameter; or If the user is in a voice input state and the user and the electronic device are in a second position relationship different from the first position relationship, the audio player is controlled to be in a fifth playback parameter, and the signal energy value output by the audio player under the fifth playback parameter is higher than the signal energy value under the fourth playback parameter.
2. The method according to claim 1, wherein Also includes: Performing corresponding processing on the audio collected by the audio collector so that the electronic device outputs the first audio to the communication terminal, or when the user is in a state of not inputting voice, the electronic device does not output the audio collected by the audio collector; The first audio does not include the audio played by the audio player and collected by the audio collector.
3. The method according to claim 1, wherein obtaining the voice status of the user of the electronic device comprises: obtaining, based at least on parameter information collected by a target sensor of the electronic device, mouth contour change information of the user, and determining the user's voice state using the mouth contour change information; or, determining a voice state of a user of the electronic device based on an operation or state of a control acting on the electronic device; or, The voice status of the user is determined based on a comparison result of the audio collected by the audio collector of the electronic device with the preset voiceprint information of the user of the electronic device.
4. The method according to claim 1, further comprising: If the user is in a voice input state, controlling the audio player to be muted, and converting the audio to be played by the audio player into text information; outputting the text information; If the user is in a state of not inputting a voice, the audio player is controlled to switch from the mute state to the play state.
5. The method according to claim 1, wherein when the user is in a voice input state, the process of controlling the playback parameters of the audio player comprises: Obtaining a parameter threshold for a current playback parameter of the audio player; wherein the parameter threshold is a pre-configured value or is determined based on an audio attribute value of audio collected by the user when the audio player is in a muted state; If the current playback parameter reaches the parameter threshold, the current playback parameter of the audio player is adjusted to a preset playback parameter; the preset playback parameter is the first playback parameter, the third playback parameter, the fourth playback parameter, or the fifth playback parameter; and / or, If the current playback parameter does not reach the parameter threshold, the current playback parameter is determined as the preset playback parameter, and the audio player is controlled to maintain the preset playback parameter unchanged.
6. The method according to claim 5, after adjusting the current playback parameters of the audio player to the preset playback parameters, the step of controlling the playback parameters of the audio player further comprises: If the user switches from the voice input state to the non-voice input state, the preset playing parameters of the audio player are restored to the playing parameters before adjustment.
7. A signal processing device, comprising: a voice status obtaining module, configured to obtain the voice status of a user of the electronic device through at least one of image recognition, distance sensor, or voiceprint recognition when the audio player of the electronic device is in a playing state; a playback parameter control module, configured to control playback parameters of the audio player based at least on the voice state; The playback parameters are at least used by the electronic device to perform corresponding processing on the audio collected by its audio collector; The controlling of the playback parameters of the audio player based at least on the voice state includes at least one of the following: If the user is in a voice input state and the electronic device is in a first state, controlling the audio player to be in a first playback parameter; If the user is in a voice input state and the electronic device is in a second state different from the first state, controlling the audio player to be in a third playback parameter, wherein the signal energy value output by the audio player under the third playback parameter is higher than the signal energy value under the first playback parameter; If the user is in a voice input state and the user and the electronic device are in a first positional relationship, controlling the audio player to be in a fourth playback parameter; or If the user is in a voice input state and the user and the electronic device are in a second position relationship different from the first position relationship, the audio player is controlled to be in a fifth playback parameter, and the signal energy value output by the audio player under the fifth playback parameter is higher than the signal energy value under the fourth playback parameter.
8. An electronic device, comprising: Audio collector; Audio player; Communication interface; A memory, configured to store a program for implementing the signal processing method according to any one of claims 1 to 6; A processor, configured to load and execute the program stored in the memory to implement the signal processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Input / output mode control for audio processing
US20180277133A1