Voice tone area switching method and device, equipment and storage medium

By dynamically switching the working mode of the voice client using sound source localization technology, the problem of low switching efficiency between single-zone and dual-zone modes of the voice client is solved, realizing the rational use of resources and optimization of user interaction.

CN114550717BActive Publication Date: 2026-01-06BEIJING PHOENIX AUTO INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210139939.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2026-01-06
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

Existing voice clients are inefficient when switching between single-zone and dual-zone modes, resulting in wasted resources and poor interaction.

Method used

By using sound source localization technology to determine the direction of the voice audio source, the working mode of the voice client is dynamically switched to match the number and location of the target objects, thereby achieving the rational use of resources.

Benefits of technology

It improves the resource utilization efficiency of the voice client, reduces resource consumption caused by unnecessary mode switching, and optimizes the user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550717B_ABST
    Figure CN114550717B_ABST
Patent Text Reader

Abstract

The application discloses a voice sound area switching method and device, equipment and storage medium, and belongs to the voice recognition field. The method comprises the following steps: acquiring voice audio of a target object, the voice audio being audio emitted by the target object when using a voice client at a current time; performing sound source positioning on the voice audio to determine a sound source positioning value corresponding to the voice audio; and switching a voice sound area based on sound source positioning values corresponding to each voice audio acquired within a first time period. Since the sound source positioning value corresponding to the voice audio can indicate the source direction of the voice audio, after the sound source positioning value corresponding to the voice audio is determined, whether the target object using the voice client at the current time is from the same direction can be determined based on the sound source positioning value, and then the voice sound area is switched based on the sound source positioning values corresponding to each voice audio within the first time period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition, and in particular to a method, apparatus, device and storage medium for switching speech regions. Background Technology

[0002] Currently, voice clients, such as voice assistants, typically operate in single-zone and dual-zone modes. In single-zone mode, the voice client captures audio through a single channel. In dual-zone mode, it captures audio through both channels. However, in some situations, voice clients need to switch between single-zone and dual-zone modes. Therefore, how to switch between voice zones has become a pressing issue. Summary of the Invention

[0003] This application provides a method, apparatus, device, and storage medium for switching voice regions. The technical solution is as follows:

[0004] On the one hand, a method for switching voice registers is provided, the method comprising:

[0005] Acquire the voice audio of the target object, wherein the voice audio is the audio emitted by the target object when using the voice client at the current time;

[0006] The speech audio is localized to determine the localization value of the speech audio.

[0007] The speech region is switched based on the sound source localization values ​​corresponding to each speech audio obtained within the first time period, where the first time period is a time period that includes the current time and is located before the current time.

[0008] Optionally, the step of performing sound source localization on the speech audio to determine the sound source localization value corresponding to the speech audio includes:

[0009] The voice audio is localized to determine whether the target is the driver or the passenger.

[0010] In the case where the target object is the main driver, the sound source localization value corresponding to the voice audio is determined to be a first value;

[0011] If the target is a passenger in the front seat, the sound source localization value corresponding to the voice audio is determined to be a second value.

[0012] Optionally, the step of switching speech regions based on the sound source localization values ​​corresponding to each speech audio obtained within the first time period includes:

[0013] If the sound source localization values ​​corresponding to each of the voice audios are the same, and the voice client is in mono-zone mode, then the voice client's working mode remains unchanged.

[0014] When the sound source localization values ​​corresponding to each of the voice audios are the same, and the voice client is in dual-zone mode, switch the voice client to single-zone mode.

[0015] If the sound source localization values ​​corresponding to the various voice audios are different, and the voice client is in single-zone mode, then switch the voice client to dual-zone mode.

[0016] If the sound source localization values ​​corresponding to the various voice audios are different, and the voice client is operating in dual-zone mode, the operating mode of the voice client shall remain unchanged.

[0017] Optionally, the step of switching speech regions based on the sound source localization values ​​corresponding to each speech audio obtained within the first time period includes:

[0018] Obtain the sound source localization value corresponding to each voice audio in the second time period, where the second time period is the time period before the first time period and the closest to the first time period.

[0019] Based on the sound source localization values ​​corresponding to each audio audio in the first time period and the sound source localization values ​​corresponding to each audio audio in the second time period, the voice region is switched.

[0020] Optionally, the step of switching speech registers based on the sound source localization values ​​corresponding to each speech audio in the first time period and the sound source localization values ​​corresponding to each speech audio in the second time period includes:

[0021] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are the same, and the working mode of the voice client is a single-zone mode, the working mode of the voice client shall remain unchanged.

[0022] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are the same, and the voice client is in dual-zone mode, then switch the voice client to single-zone mode.

[0023] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are different, and the voice client is in single-zone mode, then switch the voice client to dual-zone mode.

[0024] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are different, and the voice client is in dual-zone mode, the voice client's working mode remains unchanged.

[0025] Optionally, the time interval between the first time period and the second time period is a reference duration.

[0026] Optionally, the method further includes:

[0027] Determine the frequency of the voice audio emitted by the target object when using the voice client;

[0028] When the frequency increases, the reference duration is reduced;

[0029] When the frequency decreases, the reference duration is increased.

[0030] On the other hand, a voice zone switching device is provided, the device comprising:

[0031] The acquisition module is used to acquire the voice audio of the target object, wherein the voice audio is the audio emitted by the target object when using the voice client at the current time;

[0032] A sound source localization module is used to locate the sound source of the speech audio to determine the sound source localization value corresponding to the speech audio.

[0033] The switching module is used to switch speech regions based on the sound source localization values ​​corresponding to each speech audio obtained within a first time period, wherein the first time period is a time period that includes the current time and is located before the current time.

[0034] Optionally, the sound source localization module is specifically used for:

[0035] The voice audio is localized to determine whether the target is the driver or the passenger.

[0036] In the case where the target object is the main driver, the sound source localization value corresponding to the voice audio is determined to be a first value;

[0037] If the target is a passenger in the front seat, the sound source localization value corresponding to the voice audio is determined to be a second value.

[0038] Optionally, the switching module is specifically used for:

[0039] If the sound source localization values ​​corresponding to each of the voice audios are the same, and the voice client is in mono-zone mode, then the voice client's working mode remains unchanged.

[0040] When the sound source localization values ​​corresponding to each of the voice audios are the same, and the voice client is in dual-zone mode, switch the voice client to single-zone mode.

[0041] If the sound source localization values ​​corresponding to the various voice audios are different, and the voice client is in single-zone mode, then switch the voice client to dual-zone mode.

[0042] If the sound source localization values ​​corresponding to the various voice audios are different, and the voice client is operating in dual-zone mode, the operating mode of the voice client shall remain unchanged.

[0043] Optionally, the switching module includes:

[0044] The acquisition unit is used to acquire the sound source localization value corresponding to each speech audio in the second time period, where the second time period is the time period before the first time period and the closest time period to the first time period.

[0045] The switching unit is used to switch speech regions based on the sound source localization values ​​corresponding to each speech audio in the first time period and the sound source localization values ​​corresponding to each speech audio in the second time period.

[0046] Optionally, the switching unit is specifically used for:

[0047] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are the same, and the working mode of the voice client is a single-zone mode, the working mode of the voice client shall remain unchanged.

[0048] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are the same, and the voice client is in dual-zone mode, then switch the voice client to single-zone mode.

[0049] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are different, and the voice client is in single-zone mode, then switch the voice client to dual-zone mode.

[0050] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are different, and the voice client is in dual-zone mode, the voice client's working mode remains unchanged.

[0051] Optionally, the time interval between the first time period and the second time period is a reference duration.

[0052] Optionally, the device further includes:

[0053] The determining module is used to determine the frequency of the voice audio emitted by the target object when using the voice client;

[0054] A reduction module is used to reduce the reference duration when the frequency increases;

[0055] An amplification module is used to increase the reference duration when the frequency decreases.

[0056] On the other hand, a computer device is provided, the computer device including a memory and a processor, the memory for storing computer programs, and the processor for executing the computer programs stored in the memory to implement the steps of the above-described voice zone switching method.

[0057] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of the above-described voice zone switching method.

[0058] On the other hand, a computer program product containing instructions is provided, which, when run on a computer, cause the computer to perform the steps of the above-described voice zone switching method.

[0059] The technical solution provided in this application can bring at least the following beneficial effects:

[0060] Since the sound source localization value corresponding to the audio source can indicate the direction of the audio source, after determining the sound source localization value, it is possible to determine whether the target object using the voice client at the current time is coming from the same direction. Then, based on the sound source localization values ​​corresponding to each audio source within the first time period, the voice register can be switched. That is, the sound source localization value corresponding to the audio source is dynamically determined, and based on the sound source localization value, it is dynamically determined whether the target object using the voice client is coming from the same direction, thereby dynamically switching the voice register and achieving rational utilization of voice client resources. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a flowchart of a voice zone switching method provided in an embodiment of this application;

[0063] Figure 2 This is a schematic diagram of a voice zone switching process provided in an embodiment of this application;

[0064] Figure 3 This is a schematic diagram of the structure of a voice zone switching device provided in an embodiment of this application;

[0065] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0067] Before providing a detailed explanation of the voice zone switching method provided in the embodiments of this application, the application scenarios provided in the embodiments of this application will be introduced first.

[0068] The voice zone switching method provided in this application can be applied to various scenarios. For example, in a vehicle driving scenario, users can use a voice client on the in-vehicle terminal for voice interaction. When only one user is using the voice client, if the voice client operates in dual-zone mode, it will lead to prolonged high performance consumption. When at least two users are using the voice client, if the voice client operates in single-zone mode, it will affect the interaction between the voice client and the user. Therefore, the voice zone switching method provided in this application can be used to switch voice zones, thereby achieving reasonable utilization of voice client resources.

[0069] The voice zone switching method provided in this application can be executed by a computer device. This computer device can be any electronic product capable of human-computer interaction with the user via voice, such as a PC (Personal Computer), mobile phone, smartphone, PDA (Personal Digital Assistant), PPC (Pocket PC), tablet computer, smart TV, in-vehicle terminal, etc. Furthermore, the computer device can also interact with the user via one or more methods such as a keyboard, touchpad, touchscreen, remote control, or handwriting device.

[0070] Those skilled in the art should understand that the above application scenarios and computer devices are merely examples. Other existing or future application scenarios or computer devices that are applicable to the embodiments of this application should also be included within the scope of protection of the embodiments of this application, and are hereby incorporated by reference.

[0071] It should be noted that the application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0072] The speech region switching method provided in the embodiments of this application will be explained in detail below.

[0073] Figure 1 This is a flowchart of a voice register switching method provided in an embodiment of this application. Please refer to it. Figure 1 The method includes the following steps.

[0074] Step 101: Obtain the target object's voice audio, which is the audio emitted by the target object when using the voice client at the current time.

[0075] Based on the above description, the computer equipment provided in the embodiments of this application can be electronic products such as PCs, mobile phones, smartphones, PDAs, PPCs, tablets, smart TVs, and vehicle terminals. For ease of description, the following will take a vehicle terminal as an example.

[0076] The target audience is the user currently using the voice client on the in-vehicle terminal. Since the voice client is used to collect and recognize the target audience's voice audio, and then provides voice interaction services based on that audio, the in-vehicle terminal can acquire the target audience's voice audio when they use the voice client.

[0077] In some embodiments, two microphones are installed on the roof of the vehicle, one near the driver's side and the other near the passenger's side, with a very small distance between them. When the target uses a voice client, the in-vehicle terminal can acquire two audio streams with equal energy, and these two audio streams correspond one-to-one with the two microphones installed on the roof of the vehicle. At this time, the in-vehicle terminal identifies these two audio streams as the target's audio.

[0078] Step 102: Perform sound source localization on the speech audio to determine the corresponding sound source localization value.

[0079] The voice audio is localized to determine whether the target is the driver or the passenger. If the target is the driver, the localization value of the voice audio is determined to be the first value. If the target is the passenger, the localization value of the voice audio is determined to be the second value.

[0080] In some embodiments, the in-vehicle terminal determines the angle of the voice audio relative to the midpoint of the line connecting the center points of the two microphones. This angle is the angle between a ray originating from the midpoint of the line connecting the center points of the two microphones and the reference line, and the reference line. After determining this angle, the in-vehicle terminal can match it with multiple stored angle ranges to determine the target angle range. Then, based on the target angle range, it determines whether the target is the driver or the passenger.

[0081] For example, the vehicle terminal stores two angle ranges: [0°-90°) and [90°-180°]. Suppose the vehicle terminal determines that the angle of the voice audio relative to the midpoint of the line connecting the center points of the two microphones is 60°. Since this angle 60° falls within the angle range [0°-90°), the angle range [0°-90°) is then defined as the target angle range.

[0082] The vehicle-mounted terminal stores a correspondence between angle ranges and user locations, indicating whether the user is the driver or a passenger. Therefore, after determining the target angle range, the vehicle-mounted terminal retrieves the corresponding user location from the stored correspondence, and then determines whether the target is the driver or a passenger based on the retrieved user location.

[0083] In some embodiments, the vehicle terminal also stores a correspondence between user location and sound source localization values. Therefore, after determining the corresponding user location, the vehicle terminal can obtain the corresponding sound source localization value from the stored correspondence between user location and sound source localization values ​​based on the user location. Then, the obtained sound source localization value is determined as the sound source localization value corresponding to the voice audio.

[0084] Optionally, the vehicle terminal includes a sound source localization module, an angle range matching module, a target object determination module, and a sound source localization value determination module. The vehicle terminal can determine the sound source localization value corresponding to the voice audio through these modules. Specifically, the sound source localization module determines the angle of the voice audio relative to the midpoint of the line connecting the center points of the two microphones and sends this angle to the angle range matching module. Upon receiving the angle, the angle range matching module matches it with multiple stored angle ranges to obtain a target angle range. The angle range matching module sends the target angle range to the target object determination module. Upon receiving the target angle range, the target object determination module retrieves the corresponding user position from the stored correspondence between angle ranges and user positions. The target object determination module then sends the user position to the sound source localization value determination module. Upon receiving the user position, the sound source localization value determination module retrieves the corresponding sound source localization value from the stored correspondence between user positions and sound source localization values. Finally, the retrieved sound source localization value is determined as the sound source localization value corresponding to the voice audio.

[0085] The sound source localization value can be represented by a DOA (Direction of Arrival) value. Of course, it can also be represented by other values, and this application embodiment does not limit this. When the sound source localization value is represented by a DOA value, the first value is 0 and the second value is 1. Of course, the first and second values ​​can also be reversed, or other values ​​can be used.

[0086] In some embodiments, before the vehicle-mounted terminal performs sound source localization on the voice audio to determine the corresponding sound source localization value, it may also preprocess the voice audio to improve the accuracy of sound source localization, thereby improving the accuracy of voice zone switching. Preprocessing may include noise reduction and echo cancellation; of course, preprocessing may also include other processing methods, which are not limited in this application embodiment.

[0087] Step 103: Switch the speech region based on the sound source localization values ​​corresponding to each speech audio obtained within the first time period. The first time period is a time period that includes the current time and is located before the current time.

[0088] In some embodiments, if the sound source localization values ​​corresponding to each audio audio within a first time period are the same, and the voice client's operating mode is a single-zone mode, the operating mode of the voice client remains unchanged. If the sound source localization values ​​corresponding to each audio audio within a first time period are the same, and the voice client's operating mode is a dual-zone mode, the operating mode of the voice client is switched to a single-zone mode. If the sound source localization values ​​corresponding to each audio audio within a first time period are different, and the voice client's operating mode is a single-zone mode, the operating mode of the voice client is switched to a dual-zone mode. If the sound source localization values ​​corresponding to each audio audio within a first time period are different, and the voice client's operating mode is a dual-zone mode, the operating mode of the voice client remains unchanged.

[0089] It should be noted that the vehicle-mounted terminal can acquire the voice audio of the target object in real time, and then switch the voice register based on the sound source localization value corresponding to the voice audio of the target object. Of course, in practical applications, the vehicle-mounted terminal can also acquire voice audio in real time within a first time period, and determine the sound source localization value corresponding to each acquired voice audio according to 102 above, and then switch the voice register based on the sound source localization value corresponding to each voice audio within the first time period.

[0090] Since the sound source localization value corresponding to the voice audio can indicate the direction of the voice audio's source, if the sound source localization values ​​corresponding to all voice audio within the first time period are the same, it indicates that the target object using the voice client within the first time period originates from the same direction. That is, only one target object uses the voice client within the first time period. In this case, the target operating mode of the voice client should be monophonic mode, and the vehicle terminal uses the acoustic front-end algorithm to process the voice audio according to monophonic mode. In this situation, the vehicle terminal can determine whether the current operating mode of the voice client is monophonic mode. If the current operating mode of the voice client is monophonic mode, the operating mode of the voice client remains unchanged. If the current operating mode of the voice client is dual-zone mode, the operating mode of the voice client is switched to monophonic mode.

[0091] If the sound source localization values ​​corresponding to different audio recordings within the first time period are different, it indicates that the target objects using the voice client within the first time period originate from different directions. That is, two target objects are simultaneously using the voice client within the first time period. In this case, the target operating mode of the voice client should be dual-zone mode, and the vehicle terminal will use the acoustic front-end algorithm to process the audio recordings according to dual-zone mode. In this situation, the vehicle terminal can determine whether the current operating mode of the voice client is dual-zone mode. If the current operating mode of the voice client is single-zone mode, switch the operating mode of the voice client to dual-zone mode. If the current operating mode of the voice client is dual-zone mode, keep the operating mode of the voice client unchanged.

[0092] After the vehicle terminal determines the sound source localization value corresponding to each voice audio obtained in the first time period, it directly switches the voice zone based on whether the sound source localization values ​​corresponding to each voice audio obtained in the first time period are the same, without performing other calculations. This simplifies the calculation process and improves the efficiency of voice zone switching.

[0093] The vehicle-mounted terminal switching voice audio zones in the manner described above is one example. In other embodiments, the vehicle-mounted terminal may also switch voice audio zones in other ways. For example, the vehicle-mounted terminal obtains the sound source localization values ​​corresponding to each voice audio in a second time period, where the second time period is the time period preceding and closest to the first time period. Based on the sound source localization values ​​corresponding to each voice audio in the first time period and the sound source localization values ​​corresponding to each voice audio in the second time period, the vehicle-mounted terminal switches the voice audio zones. In other words, if the sound source localization values ​​corresponding to each audio audio segment in the first and second time periods are the same, and the voice client's working mode is monophonic, the working mode of the voice client remains unchanged. If the sound source localization values ​​corresponding to each audio audio segment in the first and second time periods are the same, and the voice client's working mode is dual-zone, the working mode of the voice client is switched to monophonic. If the sound source localization values ​​corresponding to each audio audio segment in the first and second time periods are different, and the voice client's working mode is monophonic, the working mode of the voice client is switched to dual-zone. If the sound source localization values ​​corresponding to each audio audio segment in the first and second time periods are different, and the voice client's working mode is dual-zone, the working mode of the voice client remains unchanged.

[0094] If the sound source localization values ​​corresponding to all voice audio recordings are the same in both the first and second time periods, it indicates that only one target object is using the voice client in both time periods, and the target object using the voice client in the first time period is the same as the target object using the voice client in the second time period. In this case, the target operating mode of the voice client should be monophonic mode, and the in-vehicle terminal will process the voice audio according to monophonic mode using the acoustic front-end algorithm. In this situation, the in-vehicle terminal can determine whether the current operating mode of the voice client is monophonic mode. If the current operating mode of the voice client is monophonic mode, the operating mode of the voice client remains unchanged. If the current operating mode of the voice client is dual-zone mode, the operating mode of the voice client is switched to monophonic mode.

[0095] Since the sound source localization values ​​corresponding to each voice audio in the first and second time periods are different and include multiple cases, the following will be explained in four different cases.

[0096] In the first scenario, the sound source localization values ​​for all audio audio segments within the first time period are the same, and the sound source localization values ​​for all audio audio segments within the second time period are also the same. However, the sound source localization values ​​for the audio audio segments within the first time period are different from those within the second time period. This indicates that only one target object is using the voice client in both the first and second time periods, but the target object using the voice client in the first time period is different from the target object using the voice client in the second time period. In this case, the target operating mode of the voice client should be dual-zone mode, so the in-vehicle terminal uses an acoustic front-end algorithm to process the audio audio according to the dual-zone mode.

[0097] In the second scenario, if the sound source localization values ​​are the same for all audio recordings within the first time period, but different for each audio recording within the second time period, it indicates that only one target was using the voice client during the first time period, while two target individuals were using the voice client simultaneously during the second time period. In this case, the target operating mode of the voice client should be dual-zone mode, and the in-vehicle terminal will use an acoustic front-end algorithm to process the audio recordings according to the dual-zone mode.

[0098] In the third scenario, if the sound source localization values ​​for each audio segment within the first time period are different, but the sound source localization values ​​for each audio segment within the second time period are the same, it indicates that two target objects were simultaneously using the voice client within the first time period, while only one target object was using the voice client within the second time period. In this case, the target operating mode of the voice client should be dual-zone mode, and the in-vehicle terminal will use the acoustic front-end algorithm to process the audio and voice according to the dual-zone mode.

[0099] In the fourth scenario, if the source localization values ​​for each audio segment within the first time period are different, and the source localization values ​​for each audio segment within the second time period are also different, it indicates that two target objects are simultaneously using the voice client in both the first and second time periods. In this case, the target operating mode of the voice client should be dual-zone mode, and the in-vehicle terminal will use an acoustic front-end algorithm to process the audio in dual-zone mode.

[0100] In the four scenarios described above, after the in-vehicle terminal determines that the target operating mode of the voice client should be dual-zone mode, it can then determine whether the current operating mode of the voice client is dual-zone mode. If the current operating mode of the voice client is mono-zone mode, the operating mode of the voice client is switched to dual-zone mode. If the current operating mode of the voice client is dual-zone mode, the operating mode of the voice client remains unchanged.

[0101] The vehicle terminal compares the sound source localization values ​​corresponding to each voice audio in the first time period with the sound source localization values ​​corresponding to each voice audio in the second time period to switch the voice audio region. This avoids frequent switching of the voice audio region and reduces the resource consumption caused by the switching of the voice audio region.

[0102] The first and second time periods can have the same or different durations; for example, both could be 5 minutes long. The time interval between the first and second time periods is a reference duration, which is preset, such as 2 minutes. Furthermore, the reference duration can be adjusted based on the frequency of voice audio emitted by the target user using the voice client. That is, the reference duration decreases as the frequency increases and increases as the frequency decreases. Of course, in practical applications, the reference duration can also be set to 0. In this case, the first and second time periods are adjacent.

[0103] In some embodiments, the vehicle terminal determines the frequency of the voice audio emitted by the target object when using the voice client. If the frequency of the voice audio is between a first frequency and a second frequency, the reference duration remains unchanged. If the frequency of the voice audio is less than the first frequency, the reference duration is increased; if the frequency of the voice audio is greater than the second frequency, the reference duration is decreased. The second frequency is greater than the first frequency.

[0104] When the frequency of the voice audio is less than a first frequency, the vehicle terminal determines the difference between the frequency of the voice audio and the first frequency to obtain a first difference. Then, based on the first difference, it retrieves the corresponding adjustment duration from the stored correspondence between frequency differences and adjustment durations, and increases this adjustment duration on the current reference duration to obtain an increased reference duration. Similarly, when the frequency of the voice audio is greater than a second frequency, the vehicle terminal determines the difference between the frequency of the voice audio and the second frequency to obtain a second difference. Then, based on the second difference, it retrieves the corresponding adjustment duration from the stored correspondence between frequency differences and adjustment durations, and decreases this adjustment duration on the current reference duration to obtain a decreased reference duration.

[0105] Of course, adjusting the reference duration using the method described above is just one example. In practical applications, other methods can also be used to adjust the reference duration.

[0106] Next Figure 2 Taking this as an example, the voice zone switching process provided in the embodiments of this application will be fully described. Figure 2 In this process, the in-vehicle terminal acquires and preprocesses the audio, saving the preprocessed audio. Then, it performs sound source localization on the preprocessed audio, determining the corresponding sound source localization value, and comparing the sound source localization values ​​of each audio segment acquired within a first time period. If the sound source localization values ​​are all the same, the target operating mode for the voice client is determined to be single-zone mode. If the sound source localization values ​​are different, the target operating mode for the voice client is determined to be dual-zone mode. The voice zone is then switched according to the target operating mode of the voice client.

[0107] Since the sound source localization value corresponding to the voice audio can indicate the direction of the voice audio's source, after determining the sound source localization value, it is possible to determine whether the target object using the voice client at the current time originates from the same direction. Then, based on the sound source localization values ​​corresponding to each voice audio within a first time period, the voice register can be switched. That is, the sound source localization value corresponding to the voice audio is dynamically determined, and based on this value, it is dynamically determined whether the target object using the voice client originates from the same direction, thus dynamically switching the voice register, thereby achieving rational utilization of voice client resources. Furthermore, during the voice register switching process, the voice register can be switched directly based on the sound source localization values ​​corresponding to each voice audio obtained within the first time period, thereby improving the efficiency of voice register switching. Alternatively, the voice register can be switched based on the sound source localization values ​​corresponding to each voice audio within the first time period and the sound source localization values ​​corresponding to each voice audio within a second time period, thereby avoiding frequent voice register switching.

[0108] Figure 3 This is a schematic diagram of a voice region switching device provided in an embodiment of this application. This voice region switching device can be implemented as part or all of a computer device by software, hardware, or a combination of both. Please refer to... Figure 3 The device includes: an acquisition module 301, a sound source localization module 302, and a switching module 303.

[0109] The acquisition module 301 is used to acquire the voice audio of the target object, which is the audio emitted by the target object when using the voice client at the current time. For detailed implementation details, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0110] The sound source localization module 302 is used to locate the sound source of the speech audio to determine the corresponding sound source localization value. For detailed implementation process, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0111] The switching module 303 is used to switch speech regions based on the sound source localization values ​​corresponding to each speech audio obtained within a first time period. The first time period is a time period that includes the current time and is located before the current time. For detailed implementation process, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0112] Optionally, the sound source localization module 302 is specifically used for:

[0113] The source of the voice audio is located to determine whether the target is the driver or the passenger.

[0114] When the target is the main driver, the sound source localization value corresponding to the voice audio is determined to be the first value.

[0115] When the target is the passenger in the front seat, the sound source localization value corresponding to the voice audio is determined to be the second value.

[0116] Optionally, the switching module 303 is specifically used for:

[0117] If the sound source localization values ​​corresponding to each voice audio are the same, and the voice client is in monophonic mode, then the voice client's working mode remains unchanged.

[0118] If the sound source localization values ​​corresponding to each voice audio are the same, and the voice client is in dual-zone mode, then switch the voice client to single-zone mode.

[0119] If the sound source localization values ​​corresponding to each voice audio are different, and the voice client is in single-zone mode, switch the voice client to dual-zone mode.

[0120] If the sound source localization values ​​corresponding to each voice audio are different, and the voice client is operating in dual-zone mode, the operating mode of the voice client should remain unchanged.

[0121] Optionally, the switching module 303 includes:

[0122] The acquisition unit is used to acquire the sound source localization value corresponding to each speech audio in the second time period, which is the time period before the first time period and the closest to the first time period.

[0123] The switching unit is used to switch speech regions based on the sound source localization values ​​corresponding to each speech audio in the first time period and the sound source localization values ​​corresponding to each speech audio in the second time period.

[0124] Optionally, the switching unit is specifically used for:

[0125] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are the same, and the working mode of the voice client is monophonic mode, the working mode of the voice client shall remain unchanged.

[0126] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are the same, and the voice client is in dual-zone mode, then switch the voice client to single-zone mode.

[0127] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are different, and the voice client is in single-zone mode, then switch the voice client to dual-zone mode.

[0128] If the sound source localization values ​​corresponding to each voice audio in the first time period and the second time period are different, and the voice client is in dual-zone mode, the voice client's working mode remains unchanged.

[0129] Optionally, the time interval between the first time period and the second time period is a reference duration.

[0130] Optionally, the device further includes:

[0131] The determination module is used to determine the frequency of the voice audio emitted by the target object when using the voice client.

[0132] The reduction module is used to reduce the reference duration when the frequency increases.

[0133] The increase module is used to increase the reference duration when the frequency decreases.

[0134] Since the sound source localization value corresponding to the voice audio can indicate the direction of the voice audio's source, after determining the sound source localization value, it is possible to determine whether the target object using the voice client at the current time originates from the same direction. Then, based on the sound source localization values ​​corresponding to each voice audio within a first time period, the voice register can be switched. That is, the sound source localization value corresponding to the voice audio is dynamically determined, and based on this value, it is dynamically determined whether the target object using the voice client originates from the same direction, thus dynamically switching the voice register, thereby achieving rational utilization of voice client resources. Furthermore, during the voice register switching process, the voice register can be switched directly based on the sound source localization values ​​corresponding to each voice audio obtained within the first time period, thereby improving the efficiency of voice register switching. Alternatively, the voice register can be switched based on the sound source localization values ​​corresponding to each voice audio within the first time period and the sound source localization values ​​corresponding to each voice audio within a second time period, thereby avoiding frequent voice register switching.

[0135] It should be noted that the speech region switching device provided in the above embodiments is only illustrated by the division of the above functional modules when performing speech region switching. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech region switching device and the speech region switching method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0136] Figure 4 This is a structural block diagram of a computer device 400 provided in an embodiment of this application. The computer device 400 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The computer device 400 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0137] Typically, computer device 400 includes a processor 401 and a memory 402.

[0138] Processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0139] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, which is executed by the processor 401 to implement the voice zone switching method provided in the method embodiments of this application.

[0140] In some embodiments, the computer device 400 may also optionally include a peripheral device interface 403 and at least one peripheral device. The processor 401, memory 402, and peripheral device interface 403 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 403 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 404, a touch display screen 405, a camera 406, an audio circuit 407, a positioning component 408, and a power supply 409.

[0141] Peripheral device interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 401 and memory 402. In some embodiments, processor 401, memory 402 and peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 401, memory 402 and peripheral device interface 403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0142] The radio frequency (RF) circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 404 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 404 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application embodiment.

[0143] Display screen 405 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 405 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 401 for processing. In this case, display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 405, which is disposed on the front panel of the computer device 400; in other embodiments, there may be at least two display screens 405, respectively disposed on different surfaces of the computer device 400 or in a folded design; in still other embodiments, display screen 405 may be a flexible display screen, disposed on a curved or folded surface of the computer device 400. Furthermore, display screen 405 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 405 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0144] The camera assembly 406 is used to acquire images or videos. Optionally, the camera assembly 406 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the computer device, and the rear-facing camera is located on the back of the computer device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0145] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 401 for processing, or input to the radio frequency circuit 404 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located in a different part of the computer device 400. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 407 may also include a headphone jack.

[0146] The positioning component 408 is used to locate the current geographical location of the computer device 400 in order to enable navigation or LBS (Location Based Service). The positioning component 408 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.

[0147] Power supply 409 is used to supply power to the various components in computer device 400. Power supply 409 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When power supply 409 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0148] In some embodiments, the computer device 400 further includes one or more sensors 410. The one or more sensors 410 include, but are not limited to: an accelerometer 411, a gyroscope 412, a pressure sensor 413, a fingerprint sensor 414, an optical sensor 415, and a proximity sensor 416.

[0149] Accelerometer 411 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 400. For example, accelerometer 411 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 401 can control touchscreen 405 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 411. Accelerometer 411 can also be used for games or for acquiring user motion data.

[0150] The gyroscope sensor 412 can detect the orientation and rotation angle of the computer device 400. The gyroscope sensor 412, in conjunction with the accelerometer sensor 411, can collect 3D motion data from the user on the computer device 400. Based on the data collected by the gyroscope sensor 412, the processor 401 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0151] The pressure sensor 413 can be disposed on the side bezel of the computer device 400 and / or on the lower layer of the touch display screen 405. When the pressure sensor 413 is disposed on the side bezel of the computer device 400, it can detect the user's grip signal on the computer device 400, and the processor 401 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 413. When the pressure sensor 413 is disposed on the lower layer of the touch display screen 405, the processor 401 can control the operable controls on the UI interface based on the user's pressure operation on the touch display screen 405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0152] The fingerprint sensor 414 is used to collect a user's fingerprint. The processor 401 identifies the user based on the fingerprint collected by the fingerprint sensor 414, or vice versa. When the user's identity is identified as trusted, the processor 401 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 414 can be located on the front, back, or side of the computer device 400. When the computer device 400 has physical buttons or a manufacturer's logo, the fingerprint sensor 414 can be integrated with the physical buttons or the manufacturer's logo.

[0153] An optical sensor 415 is used to collect ambient light intensity. In one embodiment, the processor 401 can control the display brightness of the touch screen 405 based on the ambient light intensity collected by the optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of the touch screen 405 is increased; when the ambient light intensity is low, the display brightness of the touch screen 405 is decreased. In another embodiment, the processor 401 can also dynamically adjust the shooting parameters of the camera assembly 406 based on the ambient light intensity collected by the optical sensor 415.

[0154] A proximity sensor 416, also known as a distance sensor, is typically located on the front panel of a computer device 400. The proximity sensor 416 is used to detect the distance between the user and the front of the computer device 400. In one embodiment, when the proximity sensor 416 detects that the distance between the user and the front of the computer device 400 is gradually decreasing, the processor 401 controls the touchscreen display 405 to switch from a screen-on state to a screen-off state; when the proximity sensor 416 detects that the distance between the user and the front of the computer device 400 is gradually increasing, the processor 401 controls the touchscreen display 405 to switch from a screen-off state to a screen-on state.

[0155] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on computer device 400, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0156] In some embodiments, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of the voice zone switching method described in the above embodiments. For example, the computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0157] It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium, in other words, it can be a non-transient storage medium.

[0158] It should be understood that all or part of the steps of the above embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-described computer-readable storage medium.

[0159] That is, in some embodiments, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of the above-described voice zone switching method.

[0160] It should be understood that "at least one" as mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., are not necessarily different.

[0161] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the voice and audio involved in the embodiments of this application were all obtained under full authorization.

[0162] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice region switching method, characterized by, The method comprises: acquiring voice audio of a target object, the voice audio being audio emitted by the target object when using a voice client at a current time; performing sound source positioning on the voice audio to determine a sound source positioning value corresponding to the voice audio; switching a voice audio zone based on sound source positioning values corresponding to each voice audio acquired within a first time period, the first time period being a time period containing the current time and located before the current time, the voice audio zone being a working mode of the voice client, the working mode comprising a single audio zone mode and a double audio zone mode, the voice client collecting voice audio through a single channel when the working mode of the voice client is the single audio zone mode, the voice client collecting voice audio through double channels when the working mode of the voice client is the double audio zone mode.

2. The method of claim 1, wherein, The sound source positioning on the voice audio to determine a sound source positioning value corresponding to the voice audio comprises: performing sound source positioning on the voice audio to determine whether the target object is a primary driver or a secondary driver; in the case where the target object is a primary driver, determining that the sound source positioning value corresponding to the voice audio is a first numerical value; in the case where the target object is a secondary driver, determining that the sound source positioning value corresponding to the voice audio is a second numerical value.

3. The method of claim 1 or 2, wherein, The switching of the voice audio zone based on sound source positioning values corresponding to each voice audio acquired within a first time period comprises: in the case where the sound source positioning values corresponding to each voice audio are all the same and the working mode of the voice client is the single audio zone mode, keeping the working mode of the voice client unchanged; in the case where the sound source positioning values corresponding to each voice audio are all the same and the working mode of the voice client is the double audio zone mode, switching the working mode of the voice client to the single audio zone mode; in the case where the sound source positioning values corresponding to each voice audio are different and the working mode of the voice client is the single audio zone mode, switching the working mode of the voice client to the double audio zone mode; in the case where the sound source positioning values corresponding to each voice audio are different and the working mode of the voice client is the double audio zone mode, keeping the working mode of the voice client unchanged.

4. The method of claim 1 or 2, wherein, The switching of the voice audio zone based on sound source positioning values corresponding to each voice audio acquired within a first time period comprises: acquiring sound source positioning values corresponding to each voice audio within a second time period, the second time period being a time period before the first time period and closest to the first time period; switching the voice audio zone based on the sound source positioning values corresponding to each voice audio within the first time period and the sound source positioning values corresponding to each voice audio within the second time period.

5. The method of claim 4, wherein, The switching of the voice audio zone based on the sound source positioning values corresponding to each voice audio within the first time period and the sound source positioning values corresponding to each voice audio within the second time period comprises: In a case where the sound source positioning values corresponding to each voice audio in the first time period and the second time period are all the same and the working mode of the voice client is the single sound area mode, the working mode of the voice client is kept unchanged; In a case where the sound source positioning values corresponding to each voice audio in the first time period and the second time period are all the same and the working mode of the voice client is the double sound area mode, the working mode of the voice client is switched to the single sound area mode; In a case where the sound source positioning values corresponding to each voice audio in the first time period and the second time period are different and the working mode of the voice client is the single sound area mode, the working mode of the voice client is switched to the double sound area mode; In a case where the sound source positioning values corresponding to each voice audio in the first time period and the second time period are different and the working mode of the voice client is the double sound area mode, the working mode of the voice client is kept unchanged.

6. The method of claim 4, wherein, The time interval between the first time period and the second time period is a reference duration.

7. The method of claim 6, wherein, The method further comprises: determining a frequency at which the target object emits voice audio when using the voice client; in a case where the frequency increases, decreasing the reference duration; in a case where the frequency decreases, increasing the reference duration.

8. A voice region switching apparatus characterized by comprising: The apparatus comprises: an acquisition module configured to acquire voice audio of a target object, the voice audio being audio emitted by the target object when using a voice client at a current time; a sound source positioning module configured to perform sound source positioning on the voice audio to determine a sound source positioning value corresponding to the voice audio; a switching module configured to switch a voice sound area based on sound source positioning values corresponding to each voice audio acquired in a first time period, the first time period being a time period containing the current time and located before the current time, the voice sound area being a working mode of the voice client, the working mode including a single sound area mode and a double sound area mode, in a case where the working mode of the voice client is the single sound area mode, the voice client acquires voice audio through a single channel, and in a case where the working mode of the voice client is the double sound area mode, the voice client acquires voice audio through double channels.

9. A computer device, comprising: The computer device comprises a memory and a processor, the memory is configured to store a computer program, and the processor is configured to execute the computer program stored on the memory to implement the steps of the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium has a computer program stored therein, and the computer program is executed by a processor to implement the steps of the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Pickup method and device for intelligent rearview mirror

    CN111688580A

  • Multi-sound-zone voice interaction method for vehicle and electronic equipment

    CN111816189A