Pitch Range Adjustment Method and Related Devices, Equipment, Systems, and Media
By acquiring and analyzing audio data in the on-board control system, performing sound area judgment and frequency domain analysis, and adjusting the parameters of the on-board sound pickup device, the problem that the existing system cannot adaptively adjust the sound pickup area, and achieving more efficient voice interaction quality.
Patent Information
- Application Number
- CN202210253992.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-15
AI Technical Summary
The existing vehicle-mounted control system cannot adaptively adjust the sound pickup area, which leads to the sound pickup area being unable to be effectively adjusted when factors such as the vehicle model and the interior environment change, affecting the quality of voice interaction.
By obtaining the audio collected by the vehicle-mounted sound pickup device based on the preset sound pickup area, perform sound area judgment and frequency domain analysis, and adjusting the device parameters to optimize the sound pickup area. The specific steps include: obtaining the first audio, performing preliminary sound area judgment, obtaining the second audio for frequency domain analysis, and adjusting the device parameters according to the analysis results to optimize the sound pickup area.
The sound pickup area is adaptively adjusted, which can optimize the sound pickup area according to changes in the personnel in the car and the car usage scene, and improve the accuracy and quality of voice interaction.
Smart Images

Figure CN114885273B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular, to a method for adjusting sound zones and related devices, equipment, systems, and media. Background Art
[0002] With the development of automotive technology, more and more vehicles are configured with in-vehicle control systems to enhance the user interaction experience based on technologies such as speech recognition.
[0003] Currently, in-vehicle control systems usually divide the in-vehicle space into different sound pickup areas to enhance the processing of voice commands in the target sound zone and reduce the interference of sound signals in non-target areas. However, the existing division of sound pickup areas is strongly related to factors such as vehicle models and in-vehicle environments. Once locked, it cannot be adaptively adjusted according to the actual in-vehicle occupants and actual driving scenarios. In view of this, how to adaptively adjust the sound pickup area has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a method for adjusting sound zones and related devices, equipment, systems, and media that can adaptively adjust the sound pickup area.
[0005] To solve the above technical problem, in the first aspect of this application, a method for adjusting sound zones is provided, including: obtaining a first audio collected by an in-vehicle sound pickup device based on a preset sound pickup area; wherein the first audio includes a preset voice command issued by a voice interaction person in the vehicle; performing a sound zone judgment on the voice interaction person based on the first audio to obtain a first sound pickup area, and obtaining a second audio collected by the in-vehicle sound pickup device based on the first sound pickup area; performing a frequency domain analysis on the second audio to obtain the frequency domain distribution of the second audio; wherein the frequency domain distribution includes: the sound frequencies of each sound source forming the second audio; in response to the frequency domain distribution satisfying a first condition, adjusting the device parameters of the in-vehicle sound pickup device to obtain an optimized sound pickup area of the in-vehicle sound pickup device.
[0006] To solve the above technical problems, a second aspect of the present application provides a sound zone adjustment device, including: a first acquisition module, a sound zone judgment module, a second acquisition module, a frequency domain analysis module, and a sound zone optimization module. The first acquisition module is configured to acquire a first audio collected by an in-vehicle sound pickup device based on a preset sound pickup area; wherein, the first audio includes a preset voice command issued by a voice interaction person in the vehicle. The sound zone judgment module is configured to perform a sound zone judgment on the voice interaction person based on the first audio to obtain a first sound pickup area. The second acquisition module is configured to acquire a second audio collected by the in-vehicle sound pickup device based on the first sound pickup area. The frequency domain analysis module is configured to perform a frequency domain analysis on the second audio to obtain the frequency domain distribution of the second audio; wherein, the frequency domain distribution includes: the sound frequencies of each sound source forming the second audio. The sound zone optimization module is configured to adjust the device parameters of the in-vehicle sound pickup device in response to the frequency domain distribution satisfying a first condition to obtain an optimized sound pickup area of the in-vehicle sound pickup device.
[0007] To solve the above technical problems, a third aspect of the present application provides a control device, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the sound zone adjustment method in the first aspect above.
[0008] To solve the above technical problems, a fourth aspect of the present application provides an in-vehicle control system, including an in-vehicle sound pickup device and the control device in the third aspect above. The in-vehicle sound pickup device is communicatively connected to the control device, and the in-vehicle sound pickup device is configured to collect audio data based on the currently configured sound pickup area of the in-vehicle sound pickup device.
[0009] To solve the above technical problems, a fifth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are configured to implement the sound zone adjustment method in the first aspect above.
[0010] In the above solution, the first audio collected by the vehicle-mounted sound pickup device based on a preset sound pickup area is obtained, and the first audio includes a preset voice command issued by a voice interaction person in the vehicle. Then, based on the first audio, a sound area judgment is performed on the voice interaction person to obtain the first sound pickup area, and the second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area is obtained, and a frequency domain analysis is performed on the second audio to obtain the frequency domain distribution of the second audio, and the frequency domain distribution includes the sound frequencies of each sound source forming the second audio. Thus, in response to the frequency domain distribution satisfying the first condition, the device parameters of the vehicle-mounted sound pickup device are adjusted to obtain the optimized sound pickup area of the vehicle-mounted sound pickup device. Therefore, during the voice interaction process of the vehicle occupants, once a preset voice command containing the voice interaction person is collected, the sound area adjustment can be started, and during the adjustment process, a preliminary sound area judgment is first performed according to the first audio, and then further adjusted according to the frequency domain distribution of the second audio. On the one hand, the sound pickup area can be adjusted according to the actual situation, and on the other hand, the optimized sound pickup area is determined through two stages, which can also improve the accuracy of the sound pickup area. Therefore, the sound pickup area can be adjusted adaptively. Description of the Drawings
[0011] Figure 1 is a schematic flowchart of an embodiment of the sound area adjustment method of the present application;
[0012] Figure 2 is a schematic diagram of an embodiment of the vehicle interior situation;
[0013] Figure 3 is a schematic diagram of an embodiment of the second audio in the time domain;
[0014] Figure 4 is a schematic diagram of an embodiment of the second audio in the frequency domain;
[0015] Figure 5 is a schematic diagram of another embodiment of the vehicle interior situation;
[0016] Figure 6 is a schematic flowchart of another embodiment of the sound area adjustment method of the present application;
[0017] Figure 7 is a schematic framework diagram of an embodiment of the sound area adjustment device of the present application;
[0018] Figure 8 is a schematic framework diagram of an embodiment of the control device of the present application;
[0019] Figure 9 is a schematic framework diagram of an embodiment of the vehicle-mounted control system of the present application;
[0020] Figure 10 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments
[0021] The solution of the embodiment of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0022] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0023] In this article, the terms "system" and "network" are often used interchangeably. The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0024] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the audio zone adjustment method of the present application.
[0025] Specifically, it may include the following steps:
[0026] Step S11: Obtain a first audio collected by an in-vehicle sound pickup device based on a preset sound pickup area.
[0027] In the embodiment of the present disclosure, the first audio includes a preset voice command issued by a voice interaction person in the vehicle. It should be noted that the number of people in the vehicle can be one, two, three, four, etc., which is not limited herein, and the voice interaction person is the person in the vehicle who issues the preset voice command. Exemplarily, please refer to Figure 2 , Figure 2 which is a schematic diagram of an embodiment of the in-vehicle situation. As Figure 2 shown, the people in the vehicle can include two people, A and B, located in the driver's seat and the back seat respectively, and A in the driver's seat can be the voice interaction person who issues the preset voice command. Of course, it can also be B in the back seat who is the voice interaction person who issues the preset voice command, which is not limited herein. In addition, Figure 2 only an exemplary possible implementation manner of the in-vehicle situation in the actual application process is given, and the in-vehicle situation in the actual application process is not limited thereby.
[0028] In an implementation scenario, the preset voice command can be used to wake up the in-vehicle voice assistant. Exemplarily, the preset voice command can specifically be a voice signal including a wake-up word. For example, the preset voice command can be "Mofei Mofei", "Xiaofei Xiaofei", etc., which is not limited herein.
[0029] In an implementation scenario, the preset sound pickup area can be regarded as the sound pickup area default-configured by the vehicle-mounted sound pickup device. Specifically, the preset sound pickup area can cover as many vehicle seats in the vehicle as possible. Exemplarily, the preset sound pickup area can cover all vehicle seats in the vehicle. In addition, the preset sound pickup area can specifically be a fan-shaped area, and the center line of the fan-shaped area has a preset included angle with the normal line of the vehicle-mounted sound pickup device. For example, the preset included angle can be set to zero degrees, that is, the center line of the preset sound pickup area can coincide with the normal line of the vehicle-mounted sound pickup device. Of course, the preset included angle can also be greater than zero degrees, such as it can be set to 1 degree, 2 degrees, 3 degrees, etc., which is not limited herein. In addition, the opening angle (i.e., the central angle) of the fan-shaped area can be set according to the vehicle seats to be covered. For example, in the case where all vehicle seats in the vehicle need to be covered, the opening angle (i.e., the central angle) of the fan-shaped area can be set to 160 degrees, 170 degrees, etc., which is not limited herein. Other situations can be deduced by analogy and will not be exemplified one by one herein. In addition, in addition to the above parameters such as the included angle and the opening angle characterizing the sound pickup area, the parameters characterizing the sound pickup area can further include the tilt angle (such as the upward tilt angle, the downward tilt angle, etc.), which is not limited herein.
[0030] In an implementation scenario, please continue to refer to Figure 2 , the vehicle-mounted sound pickup device can be a microphone array, and the microphone array can include several microphones. It should be noted that the vehicle-mounted sound pickup device can be arranged in the interior trim panel above the front windshield. Of course, the vehicle-mounted sound pickup device can also be arranged in other positions such as the center console area, etc., which is not limited herein. In addition, the distribution positions of each microphone can also be adjusted to support adaptively adjusting the sound pickup area of the vehicle-mounted sound pickup device. Exemplarily, each microphone can be respectively installed on a mechanical mechanism such as a lead screw, so that the distribution position of the microphone can be changed by driving the lead screw and other mechanical mechanisms.
[0031] Step S12: Based on the first audio, perform a sound area judgment on the voice interaction person to obtain the first sound pickup area, and acquire the second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area.
[0032] In an implementation scenario, to improve the accuracy of determining the first sound pickup area, the first audio can be processed based on beamforming to obtain a local judgment result regarding the first sound pickup area, and the first audio can be uploaded to the cloud, and a cloud judgment result regarding the first sound pickup area analyzed by the cloud based on the first audio can be received. On this basis, the first sound pickup area can be determined based on the local judgment result and the cloud judgment result. It should be noted that a mobile communication device (such as a 4G communication module, a 5G communication module, etc.) can also be integrated in the vehicle, so that the first audio can be uploaded to the cloud through the mobile communication device. The above method, combining the sound area judgment result of the first audio at the local end and the sound area judgment result of the first audio in the cloud, jointly decides the first sound pickup area, and can improve the accuracy of determining the first sound pickup area through the fusion decision-making method of cloud + local end.
[0033] In a specific implementation scenario, it should be noted that in the case where a mobile communication device is not integrated in the vehicle, the sound area of the first audio can also be judged only at the local end to obtain the first sound pickup area; or, in the case where the vehicle does not have the local judgment function, the sound area of the first audio can also be judged only in the cloud to obtain the first sound pickup area, which is not limited here.
[0034] In a specific implementation scenario, the first audio can include sub-audios respectively collected by each microphone. Then, based on the technical idea of beamforming, at the local end, each sub-audio can be processed such as weighted, time-delayed, and summed to judge the sound area of the voice interaction person who issues the preset voice command, and a local judgment result can be obtained. For the specific process of local judgment, the technical details of beamforming can be referred to and will not be elaborated here.
[0035] In a specific implementation scenario, a sound area judgment algorithm (such as a pre-trained neural network, etc.) can be integrated in the cloud, and then the algorithm in the cloud can be used for processing and partition recognition to obtain a cloud judgment result. Specifically, taking the audio judgment algorithm implemented by a neural network as an example, sample audio collected by in-vehicle sound pickup devices in the vehicle can be obtained in advance. The sample audio includes preset voice commands issued by sample voice interaction persons, and the sample audio is also marked with sample sound pickup areas covering the sample voice interaction persons. Exemplarily, the sample sound pickup area can be a fan-shaped area, and then the included angle between the center line of the fan-shaped area and the normal line of the in-vehicle sound pickup device, and the opening angle (i.e., the central angle) of the fan-shaped area can be used to represent the sample sound pickup area. That is, the sample audio can be marked with two values of the included angle and the opening angle to represent the sample sound pickup area through the marked included angle and opening angle. On this basis, the sample audio can be input into the neural network for prediction to obtain the included angle and the opening angle, and then based on the difference between the marked included angle and the predicted included angle, and the difference between the marked opening angle and the predicted opening angle, the network parameters of the neural network can be adjusted.
[0036] In a specific implementation scenario, the local judgment result may specifically include the local judged sound area, and the cloud judgment result may specifically include the cloud judged sound area. After obtaining the local judgment result and the cloud judgment result, the overlap rate of the local judged sound area and the cloud judged sound area can be obtained. Exemplarily, the first area of the intersection of the local judged sound area and the cloud judged sound area in the vehicle can be obtained first, and the second area of the union of the local judged sound area and the cloud judged sound area in the vehicle can be obtained, and then the ratio of the first area to the second area can be used as the overlap rate. On this basis, in response to the overlap rate being lower than a preset threshold, either the local judged sound area or the cloud judged sound area can be selected as the first sound pickup area. It should be noted that the preset threshold can be set according to the actual situation. For example, in the case where the judgment of the sound area is relatively strict, the preset threshold can be set higher, such as 0.8, 0.85, etc., or in the case where the judgment of the audio is relatively loose, the preset threshold can be set slightly lower, such as 0.5, 0.6, etc., which is not limited here. That is to say, in the case of a low overlap rate, either the local judged sound area or the cloud judged sound area can be directly selected as the first sound pickup area. Specifically, if the confidence of the local algorithm is high, the local judged sound area can be selected as the first sound pickup area, and if the confidence of the cloud algorithm is high, the cloud judged sound area can be selected as the first sound pickup area. In addition, in response to the overlap rate not being lower than the preset threshold, the union of the local judged sound area and the cloud judged sound area can be used as the first sound pickup area. Of course, in order to simplify the cloud + local fusion process, the overlap rate of the local judged sound area and the cloud judged sound area can no longer be obtained, but it can be first judged whether the two coincide. If they do not coincide, either the local judged sound area or the cloud judged sound area can be directly selected as the first sound pickup area, and the specific selection method can refer to the foregoing related description, which will not be elaborated here. On the contrary, if the two coincide, since the local judged sound area and the cloud judged sound area are the same, either of them can be directly used as the first sound pickup area. In the above manner, the local judgment result includes the local judged sound area, and the cloud judgment result includes the cloud judged sound area. Based on this, the overlap rate of the local judged sound area and the cloud judged sound area is obtained, and then in response to the overlap rate being lower than the preset threshold, either the local judged sound area or the cloud judged sound area is selected as the first sound pickup area, and in response to the overlap rate not being lower than the preset threshold, the union of the local judged sound area and the cloud judged sound area is used as the first sound pickup area. Therefore, by detecting the overlap rate of the local judged sound area and the cloud judged sound area, different fusion methods can be decided, which is beneficial to improving the accuracy of the first sound pickup area.
[0037] In an implementation scenario, in order to further improve the accuracy of the first sound pickup area, vehicle-mounted sensing devices can also be integrated in the vehicle. After determining the first sound pickup area through the cloud + terminal, or only through the cloud, or only through the local terminal, the sensing data of the vehicle-mounted sensing devices on the vehicle occupants can be further obtained. On this basis, personnel analysis can be performed based on the sensing data to obtain the seating position of the voice interaction personnel in the vehicle, and a dual sound area judgment can be performed based on the first sound pickup area and the seating position to obtain a new first sound pickup area. It should be noted that the vehicle-mounted sensing devices can include, but are not limited to, vehicle-mounted cameras, pressure sensors on the seats, etc., which are not limited here. Correspondingly, the sensing data can include, but are not limited to, image data of the vehicle interior, pressure data measured by the pressure sensors, etc., which are not limited here. In the above manner, after obtaining the first sound pickup area, further combining the seating position of the voice interaction personnel in the vehicle analyzed from the sensing data of the vehicle-mounted sensing devices to perform a dual sound area judgment is beneficial to avoiding, as much as possible, the misjudgment of the first sound pickup area due to microphone failures, sudden noises in the vehicle, algorithm deviations, etc. Therefore, the accuracy of the first sound pickup area can be further improved.
[0038] In a specific implementation scenario, when the vehicle-mounted sensing device includes a vehicle-mounted camera, speaker recognition can be performed on the image data captured when the first audio is collected to determine the in-vehicle speaker when the first audio is collected and use it as the voice interaction person, so that the seating position of the voice interaction person in the vehicle can be located; or, when the vehicle-mounted sensing device includes a pressure sensor on the seat, the pressure data can be analyzed. If the pressure value included in the pressure data is higher than the preset threshold, it can be considered that there is a passenger on the corresponding seat, and the in-vehicle position of the seat with a passenger can be directly used as the seating position of the voice interaction person; or, in order to further improve the accuracy of the seating position, the seating position of the voice interaction person can be jointly determined by combining both the image data and the pressure data. For example, the person on the corresponding seat can be used as a candidate, and the candidate who is speaking when the first audio is collected can be selected as the voice interaction person, so that the seating position of the voice interaction person in the vehicle can be determined.
[0039] In a specific implementation scenario, after obtaining the seating position of the voice interaction person in the vehicle, in response to the seating position being within the first sound pickup area, the first sound pickup area can be maintained unchanged. That is to say, the new first sound pickup area is the aforementioned first sound pickup area obtained by judging the sound area based on the first audio. Of course, in response to the seating position being outside the first sound pickup area, the first sound pickup area can be adjusted to cover the seating position to obtain a new first sound pickup area. Exemplarily, as described above, the first sound pickup area can be a fan-shaped area. In this case, the angle between the center line of the fan-shaped area and the normal line of the vehicle-mounted sound pickup device can be adjusted, that is, rotated with the vertex of the fan-shaped area as the base point. At the same time, the opening angle (i.e., the central angle) of the fan-shaped area is maintained unchanged until the fan-shaped area covers the seating position, and then the rotation can be stopped. At this time, the fan-shaped area can be regarded as the new first sound pickup area. In the above manner, in response to the seating position being within the first sound pickup area, the first sound pickup area is maintained unchanged, and / or in response to the seating position being outside the first sound pickup area, the first sound pickup area is adjusted to cover the seating position to obtain a new first sound pickup area. The first sound pickup area can be adjusted according to the relative position relationship between the first sound pickup area and the seating position, so as to effectively correct the sound area judgment errors caused by microphone failures, sudden noises in the vehicle, algorithm deviations, etc., and further improve the accuracy of the first sound pickup area.
[0040] In a specific implementation scenario, as described above, the first sound pickup area can be a fan-shaped area. To further improve the accuracy of the first sound pickup area, the line connecting the seating position and the vertex of the fan-shaped area can be obtained first, and the angle between the line and the center line of the fan-shaped area can be obtained. On this basis, in response to the angle being lower than the preset threshold, the first sound pickup area can be maintained unchanged. That is to say, the new first sound pickup area is the aforementioned first sound pickup area obtained by judging the sound area based on the first audio. In addition, in response to the angle not being lower than the preset threshold, rotation can be performed with the vertex of the fan-shaped area as the base point. At the same time, the opening angle (i.e., the central angle) of the fan-shaped area is maintained unchanged until the angle between the center line of the fan-shaped area and the line is lower than the preset threshold (for example, the center line coincides with the line, that is, the angle is zero degree), and then the rotation can be stopped. At this time, the fan-shaped area can be regarded as the new first sound pickup area.
[0041] It should be noted that after obtaining the first sound pickup area by judging the sound area based on the first audio, or after further combining the sensing data on the basis of the first sound pickup area to perform a dual sound area judgment to obtain a new first sound pickup area, the vehicle-mounted sound pickup device can collect the second audio based on the latest first sound pickup area. Exemplarily, the vehicle-mounted sound pickup device can continuously collect audio signals based on the latest first sound pickup area at this time to obtain the second audio.
[0042] Step S13: Perform frequency-domain analysis on the second audio to obtain the frequency-domain distribution of the second audio.
[0043] In the embodiments of the present disclosure, the frequency-domain distribution includes: the sound frequencies of each sound source that forms the second audio. It should be noted that in addition to the voice of the voice interaction person, the first pickup area may also pick up sounds such as music and videos played on a mobile phone, conversation sounds, etc., which are not limited herein. Please refer to Figure 3 and Figure 4 , Figure 3 is a schematic diagram of an embodiment of the second audio in the time domain, Figure 4 is a schematic diagram of an embodiment of the second audio in the frequency domain. As Figure 3 shown and Figure 4 shown, different sound sources cannot be distinguished in the time domain, while different sound sources can be clearly distinguished in the frequency domain. For example, it can be clearly seen from the spectrogram of the second audio shown in Figure 4 that it contains components such as 80 Hz, 400 Hz, 1 kHz, 1.5 kHz, 2 kHz, 2.5 kHz, 3 kHz, etc. Thus, it can be seen that there are many interfering sounds other than the voice interaction person in the second audio. It should be noted that Figure 3 and Figure 4 are only a possible implementation manner of the second audio in the actual application process, and do not limit the actual spectrum of the second audio accordingly. In addition, please continue to refer to Figure 2 , generally speaking, the larger the opening angle of the first pickup area, the higher the possibility that the second audio contains many interfering sounds other than the voice interaction person. For example, for Figure 2 shown, the second audio may contain the voice of "A" sitting in the driver's seat, or may contain the voice of "B" sitting in the back seat. Other situations can be deduced by analogy, and no further examples will be given here. In addition, the second audio can be subjected to frequency-domain analysis through algorithms such as Fourier transform. The specific process can refer to the technical details of frequency-domain analysis algorithms such as Fourier transform, which will not be elaborated here.
[0044] Step S14: In response to the frequency-domain distribution satisfying the first condition, adjust the device parameters of the vehicle-mounted pickup device to obtain an optimized pickup area of the vehicle-mounted pickup device.
[0045] In one implementation scenario, the first condition can be set to include: the second sound frequency exists in the frequency domain distribution after the first sound frequency is removed. It should be noted that the first sound frequency is the sound frequency of the person interacting with the voice, and the second sound frequency is a sound frequency different from the first sound frequency. That is to say, if there is a second sound frequency in addition to the first sound frequency removed from the frequency domain distribution, the device parameters of the vehicle-mounted sound pickup device can continue to be adjusted to obtain the optimized sound pickup area of the vehicle-mounted sound pickup device. In the above manner, the first condition is set to include: the second sound frequency exists in the frequency domain distribution after the first sound frequency is removed. When the second sound frequency exists in the frequency distribution, the device parameters of the vehicle-mounted sound pickup device can be continuously adjusted, which is conducive to adaptively optimizing the sound pickup area of the vehicle-mounted sound pickup device as much as possible according to the actual situation in the vehicle.
[0046] In one implementation scenario, the center line of the optimized sound pickup area coincides with the center line of the first sound pickup area, and the opening angle of the optimized sound pickup area is smaller than the opening angle of the first sound pickup area. Exemplarily, the optimized sound pickup area and the first sound pickup area can both be fan-shaped areas, and the above-mentioned opening angle is the center angle of the circle. In other words, by adjusting the device parameters of the vehicle-mounted sound pickup device, the opening angle of the first sound pickup area can be continuously shrunk without changing the direction of the first sound pickup area, so that the shrunk first sound pickup area excludes interfering sound sources other than the voice interacting person as much as possible, until the shrunk first sound pickup area only covers the sound source of the voice interacting person, and then the current first sound pickup area can be used as the optimized sound pickup area.
[0047] In one implementation scenario, during the optimization process, the device parameters can be adjusted based on a preset adjustment step length to obtain a second sound pickup area, and the center line of the second sound pickup area coincides with the center line of the first sound pickup area, and the first sound pickup area and the second sound pickup area have different opening angles. On this basis, the second sound pickup area can be used as a new first sound pickup area, and the step of obtaining the second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area and the subsequent steps are re-executed until the frequency domain distribution does not meet the first condition. That is to say, when the frequency domain distribution does not exist after the first sound frequency is eliminated, it can be considered that the device parameters have been adjusted in place, and the latest first sound pickup area can be used as the optimized sound pickup area. In the above manner, the device parameters are adjusted based on the preset adjustment step length to obtain the second sound pickup area, and the center lines of the first sound pickup area and the second sound pickup area coincide with each other but have different opening angles. Based on this, the second sound pickup area is used as a new first sound pickup area, and the aforementioned second audio collection and subsequent steps are re-executed until the frequency domain distribution does not meet the first condition, so that the device parameters can be adjusted in real time according to the actual situation in the car to adaptively adjust the sound pickup area, which helps to improve the accuracy of the sound pickup area.
[0048] In a specific implementation scenario, as described above, the in-vehicle sound pickup device is a microphone array, and the microphone array includes several microphones. The device parameters include the distribution positions of several microphones in the vehicle. Then, the preset adjustment step size can be regarded as the distance for moving each microphone each time. In addition, as described above, by adjusting the device parameters based on the preset adjustment step size, it is possible to change the opening angle of the sound pickup area without changing the center line of the sound pickup area. For example, when the sound pickup area is a fan-shaped area, the opening angle of the sound pickup area is the central angle of the fan-shaped area. Exemplarily, the opening angle A of the sound pickup area can be expressed as:
[0049] A = 2cos -1 (c / Mdf)……(1)
[0050] In the above formula (1), c is the speed of sound, M is the total number of microphones included in the microphone array, d is the distance between microphones, and f is the frequency of the incident sound wave (i.e., the voice frequency of the voice interaction person, i.e., the aforementioned first sound frequency). As described above, the larger the opening angle A, the larger the recognition sound area of the voice command, and the more likely it is to introduce noise and interference sounds from the surrounding environment (such as, Figure 2 as shown). Therefore, through the aforementioned adjustment process, the beam sound pickup range can be accurately shrunk to exclude the noise and interference sounds from the surrounding environment as much as possible while ensuring that the sound pickup area covers the voice interaction person.
[0051] In a specific implementation scenario, please refer to Figure 2 and Figure 5 , Figure 5 which is a schematic diagram of another embodiment of the in-vehicle situation. As Figure 2 shown, through frequency domain analysis, it can be determined that there is still a second sound frequency "B" in the frequency domain distribution after removing the first sound frequency of the voice interaction person "A". Then, the device parameters of the in-vehicle sound pickup device can be adjusted to shrink the opening angle of the first sound pickup area (such as the shaded part in Figure 2 ) as shown, to obtain the second sound pickup area (such as the shaded part in Figure 2 ) as shown. At this time, since the second sound pickup area no longer covers "B", after using the second sound pickup area as the new first sound pickup area and obtaining the second audio collected by the in-vehicle sound pickup device based on the first sound pickup area, and then performing frequency domain analysis on the second audio, it can be found that there is no second sound frequency in the frequency domain distribution after removing the first sound frequency. Then, the latest first sound pickup area (i.e., the shaded part in Figure 5 ) as shown can be used as the optimized sound pickup area of the sound pickup device. Of course, if the sound pickup device is based on Figure 5 Figure 5 shown Figure 5 in the shaded partThe second audio is collected from the pickup area shown. After frequency-domain analysis, it is found that there is a second sound frequency in the frequency-domain distribution after removing the first sound frequency. Then, the above steps of adjusting the device parameters according to the preset adjustment step can be continued to dynamically adjust the pickup area.
[0052] In an implementation scenario, in the real world, it is possible that the included angle of the optimized pickup area is small, and the voice interaction person moves within a certain range in the vehicle, resulting in the situation that the optimized pickup area may no longer cover the voice interaction person after the movement. Therefore, in order to further improve the adaptive adjustment ability of the pickup area, after determining the optimized pickup area, the optimized pickup area can be further used as the new first pickup area, and the second audio collected by the vehicle-mounted pickup device based on the first pickup area can be obtained, and the frequency-domain analysis is performed on the second audio to obtain the frequency-domain distribution of the second audio. If the frequency-domain distribution does not include the first sound frequency, the above steps of obtaining the first audio collected by the vehicle-mounted pickup device based on the preset pickup area and the subsequent steps can be re-executed to restart the two-stage sound area adjustment process from the beginning.
[0053] In an implementation scenario, the pickup areas set by the vehicle and the setting times of each pickup area can be counted within a preset time period (such as one month, one quarter, one year, etc.). Exemplarily, if the vehicle is mainly used for commuting, among the pickup areas set by the vehicle, the setting times corresponding to the pickup area covering the driver's seat will theoretically be much higher than those of other pickup areas. In addition, since the inclination angle of the pickup area usually also varies more or less with the height of the voice interaction person. On this basis, the vehicle can upload the pickup area with the most setting times and the corresponding height of the voice interaction person to the background service system. The background service system can further analyze each vehicle model. For example, for each vehicle model, the distribution of the pickup area with respect to height can be counted. For example, the proportion of each different height in the total data volume can be counted, and the pickup area corresponding to the height with the largest proportion can be selected as the default pickup area for this vehicle model. Thus, when the vehicle model leaves the factory, its vehicle-mounted pickup device can be directly configured as the default pickup area, and then when it is delivered for use, it can be used without adjusting the sound area or only need to make a small adjustment to the sound area.
[0054] In the above solution, the first audio collected by the vehicle-mounted sound pickup device based on a preset sound pickup area is obtained, and the first audio includes a preset voice command issued by a voice interaction person in the vehicle. Then, based on the first audio, the sound area of the voice interaction person is judged to obtain the first sound pickup area. The second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area is obtained, and frequency domain analysis is performed on the second audio to obtain the frequency domain distribution of the second audio. The frequency domain distribution includes the sound frequencies of each sound source forming the second audio. Thus, in response to the frequency domain distribution satisfying the first condition, the device parameters of the vehicle-mounted sound pickup device are adjusted to obtain the optimized sound pickup area of the vehicle-mounted sound pickup device. Therefore, during the voice interaction of the vehicle occupants, once a preset voice command containing the voice interaction person is collected, the sound area adjustment can be started. During the adjustment process, the initial sound area judgment is first performed according to the first audio, and then further adjusted according to the frequency domain distribution of the second audio. On the one hand, the sound pickup area can be adjusted according to the actual situation, and on the other hand, the optimized sound pickup area can be determined through two stages, which can also improve the accuracy of the sound pickup area. Therefore, the sound pickup area can be adjusted adaptively.
[0055] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of another embodiment of the sound area adjustment method of the present application.
[0056] Specifically, it may include the following steps:
[0057] Step S601: Obtain the first audio collected by the vehicle-mounted sound pickup device based on a preset sound pickup area.
[0058] In the embodiments of the present disclosure, the first audio includes a preset voice command issued by a voice interaction person in the vehicle. Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein.
[0059] Step S602: Judge the sound area of the voice interaction person based on the first audio to obtain the first sound pickup area.
[0060] Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein.
[0061] Step S603: Obtain the second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area.
[0062] Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein.
[0063] Step S604: Perform frequency domain analysis on the second audio to obtain the frequency domain distribution of the second audio.
[0064] In the embodiments of the present disclosure, the frequency domain distribution includes: the sound frequencies of each sound source forming the second audio. Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein.
[0065] Step S605: Determine whether the frequency domain distribution satisfies the first condition. If so, execute Step S606; otherwise, execute Step S609.
[0066] Specifically, the first condition can be set to include that there is a second voice frequency in the frequency domain distribution after removing the first voice frequency. It should be noted that the first voice frequency is the voice frequency of the voice interaction person, and the second voice frequency is a voice frequency different from the first voice frequency. For details, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein.
[0067] Step S606: Adjust the device parameters based on a preset adjustment step size to obtain a second sound pickup area.
[0068] In the embodiments of the present disclosure, the center line of the second sound pickup area coincides with the center line of the first sound pickup area, and the opening angles of the first sound pickup area and the second sound pickup area are different. For details, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein.
[0069] Step S607: Use the second sound pickup area as the new first sound pickup area.
[0070] Step S608: Re-execute the steps of Step S603 and subsequent steps.
[0071] Step S609: Use the latest first sound pickup area as the optimized sound pickup area.
[0072] Step S610: Determine whether the frequency domain distribution satisfies the second condition. If so, execute Step S611; otherwise, execute Step S613.
[0073] Specifically, the second condition can be set to include that the frequency domain distribution does not have the voice frequency of the voice interaction person, that is, the foregoing first voice frequency. In addition, the operation of determining whether the frequency domain distribution satisfies the second condition can be executed simultaneously with Step S605 after Step S604. Of course, it can also be executed before or after Step S604, which is not limited herein. Exemplarily, Step S610 can be executed first, and if it is determined that the frequency domain distribution satisfies the second condition, Step S604 can be continued; or, as Figure 6 shown, after executing Step S609, Step S610 can be executed. In this case, to obtain the latest frequency domain distribution, the optimized sound pickup area can be used as the new first sound pickup area, and the foregoing Steps S603 and S604 can be re-executed to obtain the latest frequency domain distribution. It should be noted that Figure 6 this is only an exemplary illustration of the sound area adjustment process, and for the sake of simplifying the flowchart, Figure 6 these two steps have been omitted.
[0074] Step S611: Determine that the voice interaction person has changed.
[0075] It should be noted that when the frequency domain distribution satisfies the second condition, the current voice interaction person may still be the original voice interaction person, but only the position of the voice interaction person has changed, resulting in the frequency domain distribution no longer containing the first sound frequency; or, the current voice interaction person is no longer the original voice interaction person, that is, the identity of the voice interaction person has changed (for example, the original voice interaction person is a passenger in the driver's seat, and the current voice interaction person is a passenger in the back seat), resulting in the frequency domain distribution no longer containing the first sound frequency. In the embodiments of the present disclosure, both of the above situations can be regarded as a change in the voice interaction person.
[0076] Step S612: Re-execute the steps of Step S601 and subsequent steps.
[0077] Specifically, when it is determined that the voice interaction person has changed, the two-stage sound zone adjustment process can be restarted from the beginning to determine the optimized sound pickup area.
[0078] Step S613: Re-execute Step S603 and subsequent steps.
[0079] Specifically, when it is determined that the voice interaction person has not changed, the second audio can be collected again in the latest first sound pickup area, and frequency domain analysis and subsequent operations can be performed to adjust the sound pickup area in real time according to the vehicle interior situation.
[0080] In the above solution, during the sound zone adjustment process, it is further determined whether the frequency domain distribution satisfies the second condition, and the second condition is set to include: the frequency domain distribution does not contain the sound frequency of the voice interaction person, and when the second condition is satisfied, it is determined that the voice interaction person has changed, and the steps of obtaining the first audio collected by the vehicle-mounted sound pickup device based on the preset sound pickup area and subsequent steps are re-executed, so as to be able to timely sense the change of the voice interaction personnel and adjust the sound pickup sound zone in real time according to the change situation, thereby improving the adaptive ability of the sound zone adjustment.
[0081] Please refer to Figure 7 , Figure 7It is a schematic framework diagram of an embodiment of the voice range adjustment device 70 of the present application. The voice range adjustment device 70 includes: a first acquisition module 71, a voice range judgment module 72, a second acquisition module 73, a frequency domain analysis module 74, and a voice range optimization module 75. The first acquisition module 71 is configured to acquire a first audio collected by an in-vehicle sound pickup device based on a preset sound pickup area; wherein, the first audio includes a preset voice command issued by a voice interaction person in the vehicle. The voice range judgment module 72 is configured to perform a voice range judgment on the voice interaction person based on the first audio to obtain a first sound pickup area. The second acquisition module 73 is configured to acquire a second audio collected by the in-vehicle sound pickup device based on the first sound pickup area. The frequency domain analysis module 74 is configured to perform a frequency domain analysis on the second audio to obtain the frequency domain distribution of the second audio; wherein, the frequency domain distribution includes: the sound frequencies of each sound source forming the second audio. The voice range optimization module 75 is configured to, in response to the frequency domain distribution satisfying a first condition, adjust the device parameters of the in-vehicle sound pickup device to obtain an optimized sound pickup area of the in-vehicle sound pickup device.
[0082] In the above solution, the first audio collected by the in-vehicle sound pickup device based on the preset sound pickup area is acquired, and the first audio includes a preset voice command issued by the voice interaction person in the vehicle. Then, a voice range judgment is performed on the voice interaction person based on the first audio to obtain the first sound pickup area, and the second audio collected by the in-vehicle sound pickup device based on the first sound pickup area is acquired, and a frequency domain analysis is performed on the second audio to obtain the frequency domain distribution of the second audio. The frequency domain distribution includes the sound frequencies of each sound source forming the second audio. Thus, in response to the frequency domain distribution satisfying the first condition, the device parameters of the in-vehicle sound pickup device are adjusted to obtain the optimized sound pickup area of the in-vehicle sound pickup device. Therefore, during the voice interaction process of the vehicle occupants, once a preset voice command including the voice interaction person is collected, the voice range adjustment can be started, and during the adjustment process, a preliminary voice range judgment is first made according to the first audio, and then further adjusted according to the frequency domain distribution of the second audio. On the one hand, the sound pickup area can be adjusted according to the actual situation, and on the other hand, the optimized sound pickup area is determined through two stages, which can also improve the accuracy of the sound pickup area. Therefore, the sound pickup area can be adaptively adjusted.
[0083] In some disclosed embodiments, the voice zone optimization module 75 includes a parameter adjustment sub-module for adjusting device parameters based on a preset adjustment step size to obtain a second pickup area; wherein, the center line of the second pickup area coincides with the center line of the first pickup area, and the opening angles of the first pickup area and the second pickup area are different; the voice zone optimization module 75 includes a first selection sub-module for using the second pickup area as the new first pickup area, the voice zone optimization module 75 includes a loop execution sub-module for re-executing the steps of obtaining the second audio collected by the vehicle-mounted pickup device based on the first pickup area and subsequent steps in combination with the second acquisition module 73 and the frequency domain analysis module 74, and the voice zone optimization module 75 includes a second selection sub-module for using the latest first pickup area as the optimized pickup area when the frequency domain distribution does not meet the first condition.
[0084] Therefore, device parameters are adjusted based on a preset adjustment step size to obtain a second pickup area, and the center lines of the first pickup area and the second pickup area coincide while the opening angles are different. Based on this, the second pickup area is used as the new first pickup area, and the foregoing steps of collecting the second audio and subsequent steps are re-executed until the frequency domain distribution does not meet the first condition, so that the device parameters can be adjusted in real time according to the actual situation in the vehicle to adaptively adjust the pickup area, which helps to improve the accuracy of the pickup area.
[0085] In some disclosed embodiments, the vehicle-mounted pickup device is a microphone array, and the microphone array includes a plurality of microphones, and the device parameters include the distribution positions of the plurality of microphones in the vehicle.
[0086] Therefore, the vehicle-mounted pickup device is set as a microphone array, and the microphone array includes a plurality of microphones, and the device parameters include the distribution positions of the plurality of microphones in the vehicle, so that the adaptive ability of the voice zone adjustment can be improved by changing the distribution positions of the microphones in the vehicle in real time according to the actual situation in the vehicle.
[0087] In some disclosed embodiments, the first condition includes: there is a second sound frequency in the frequency domain distribution after removing the first sound frequency; wherein, the first sound frequency is the sound frequency of the voice interaction person, and the second sound frequency is a sound frequency different from the first sound frequency.
[0088] Therefore, the first condition is set to include: there is a second sound frequency in the frequency domain distribution after removing the first sound frequency, which can continuously adjust the device parameters of the vehicle-mounted pickup device when there is a second sound frequency in the frequency distribution, and is beneficial to adaptively optimize the pickup area of the vehicle-mounted pickup device as much as possible according to the actual situation in the vehicle.
[0089] In some disclosed embodiments, the voice zone adjustment device 70 further includes a personnel determination module, configured to determine that the voice interaction person has changed in response to the frequency domain distribution satisfying a second condition. The voice zone adjustment device 70 further includes a loop adjustment module, configured to re-execute the step of obtaining the first audio collected by the vehicle-mounted sound pickup device based on a preset sound pickup area and subsequent steps in combination with the first acquisition module 71, the voice zone determination module 72, the second acquisition module 73, the frequency domain analysis module 74, and the voice zone optimization module 75; the second condition includes: the frequency domain distribution does not have the voice frequency of the voice interaction person.
[0090] Therefore, during the voice zone adjustment process, it is further determined whether the frequency domain distribution satisfies the second condition, and the second condition is set to include: the frequency domain distribution does not have the voice frequency of the voice interaction person. When the second condition is satisfied, it is determined that the voice interaction person has changed, and the step of obtaining the first audio collected by the vehicle-mounted sound pickup device based on a preset sound pickup area and subsequent steps are re-executed, so as to be able to timely sense the change of the voice interaction personnel and adjust the sound pickup voice zone in real time according to the change situation, thereby improving the adaptive ability of the voice zone adjustment.
[0091] In some disclosed embodiments, the voice zone determination module 72 includes a local determination sub-module, configured to process the first audio based on beamforming to obtain a local determination result regarding the first sound pickup area; the voice zone determination module 72 includes a cloud determination sub-module, configured to upload the first audio to the cloud and receive a cloud determination result regarding the first sound pickup area analyzed by the cloud based on the first audio; the voice zone determination module 72 includes a determination combination sub-module, configured to determine the first sound pickup area based on the local determination result and the cloud determination result.
[0092] Therefore, by combining the voice zone determination result of the first audio at the local end and the voice zone determination result of the first audio in the cloud, and jointly making a decision on the first sound pickup area, the accuracy of determining the first sound pickup area can be improved through the fusion decision-making method of cloud + local end.
[0093] In some disclosed embodiments, the local determination result includes a local determination voice zone, and the cloud determination result includes a cloud determination voice zone; the determination combination sub-module includes an overlap calculation unit, configured to obtain the overlap rate of the local determination voice zone and the cloud determination voice zone; the determination combination sub-module includes a first response unit, configured to select the local determination voice zone or the cloud determination voice zone as the first sound pickup area in response to the overlap rate not satisfying a third condition; the determination combination sub-module includes a second response unit, configured to use the union of the local determination voice zone and the cloud determination voice zone as the first sound pickup area in response to the overlap rate satisfying the third condition; wherein, the third condition includes: the overlap rate is not lower than a preset threshold.
[0094] Therefore, the judgment result of the local end includes the judgment sound area of the local end, and the judgment result of the cloud end includes the judgment sound area of the cloud end. Based on this, the overlap rate of the judgment sound area of the local end and the judgment sound area of the cloud end is obtained, and then in response to the overlap rate being lower than the preset threshold, the judgment sound area of the local end or the judgment sound area of the cloud end is selected as the first sound pickup area. In response to the overlap rate not being lower than the preset threshold, the union of the judgment sound area of the local end and the judgment sound area of the cloud end is used as the first sound pickup area. Therefore, by detecting the overlap rate of the judgment sound area of the local end and the judgment sound area of the cloud end, it is possible to decide to adopt different fusion methods, which is conducive to improving the accuracy of the first sound pickup area.
[0095] In some disclosed embodiments, the sound zone adjustment device 70 includes a personnel analysis module, which is used to perform personnel analysis on the sensing data of the people in the car based on the vehicle-mounted sensing equipment, and obtain the seating position of the person interacting with the voice in the car; the sound zone adjustment device 70 includes a double judgment module, which is used to perform a double sound zone judgment based on the first sound pickup area and the seating position, and obtain a new first sound pickup area.
[0096] Therefore, after obtaining the first sound pickup area, the double sound zone judgment is performed by further combining the seating position of the voice interaction person in the car obtained by analyzing the sensing data of the on-board sensing device. This is beneficial to avoid as much as possible the first sound pickup area judgment error due to microphone failure, sudden noise in the car, algorithm deviation, etc., so the accuracy of the first sound pickup area can be further improved.
[0097] In some disclosed embodiments, the double determination module includes a first response submodule for maintaining the first sound pickup area unchanged in response to the seating position being located in the first sound pickup area; the double determination module includes a second response submodule for adjusting the first sound pickup area to cover the seating position in response to the seating position being located outside the first sound pickup area to obtain a new first sound pickup area.
[0098] Therefore, in response to the seating position being located in the first sound pickup area, the first sound pickup area is maintained unchanged, and / or, in response to the seating position being located outside the first sound pickup area, the first sound pickup area is adjusted to cover the seating position to obtain a new first sound pickup area. The first sound pickup area can be adjusted according to the relative position relationship between the first sound pickup area and the seating position, thereby effectively correcting sound zone judgment errors caused by microphone failure, sudden noise in the car, algorithm deviation, etc., and further improving the accuracy of the first sound pickup area.
[0099] See also Figure 8 , Figure 8It is a schematic framework diagram of an embodiment of the electronic device 80 of the present application. The electronic device 80 includes a memory 81 and a processor 82 which are coupled to each other. Program instructions are stored in the memory 81, and the processor 82 is configured to execute the program instructions to implement the steps in any of the above-described embodiments of the sound zone adjustment method. Specifically, the electronic device 80 may include, but is not limited to, an industrial control host, a vehicle-mounted central control unit, etc., which are not limited herein.
[0100] Specifically, the processor 82 is configured to control itself and the memory 81 to implement the steps in any of the above-described embodiments of the sound zone adjustment method. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with the ability to process signals. The processor 82 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.
[0101] In the above solution, during the in-vehicle personnel voice interaction process, once a preset voice command including the voice interaction person is collected, the sound zone adjustment can be started. And during the adjustment process, the initial sound zone judgment is first made according to the first audio, and then further adjusted according to the frequency domain distribution of the second audio. On the one hand, it can adjust the sound pickup area according to the changes of the actual situation. On the other hand, by determining the optimized sound pickup area in two stages, it can also improve the accuracy of the sound pickup area. Therefore, it can adaptively adjust the sound pickup area.
[0102] Please refer to Figure 9 , Figure 9 It is a schematic framework diagram of an embodiment of the vehicle-mounted control system 90 of the present application. The vehicle-mounted control system 90 includes a vehicle-mounted sound pickup device 91 and a control device 92 as described in the foregoing disclosed embodiments. The vehicle-mounted sound pickup device 91 is communicatively connected to the control device 92, and the vehicle-mounted sound pickup device 91 is configured to collect audio data based on the sound pickup area currently configured by the vehicle-mounted sound pickup device 91. Specifically, the vehicle-mounted sound pickup device 91 may include, but is not limited to, a microphone array. For specific details, reference may be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0103] In an implementation scenario, the vehicle-mounted control system 90 may further include vehicle-mounted sensing devices 93. As described in the foregoing disclosed embodiments, the vehicle-mounted sensing devices 93 may include, but are not limited to, vehicle-mounted cameras, pressure sensors for vehicle seats, etc., which are not limited herein. The vehicle-mounted sensing devices 93 are communicatively connected to the control device 92, and the vehicle-mounted sensing devices 93 are configured to collect information about the vehicle occupants to obtain sensing data. For the specific content of the sensing data, reference may be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0104] In an implementation scenario, the vehicle-mounted control system 90 may further include a vehicle-mounted multimedia device (not shown), and the vehicle-mounted multimedia device is communicatively connected to the control device 92. Specifically, the vehicle-mounted multimedia device may include, but is not limited to, speakers, screens, etc., which are not limited herein.
[0105] In the above solution, during the voice interaction process of the vehicle occupants, once a preset voice command including the voice interaction person is collected, the sound zone adjustment can be started. And during the adjustment process, the initial sound zone judgment is first made according to the first audio, and then the adjustment is further made according to the frequency domain distribution of the second audio. On the one hand, the sound pickup area can be adjusted according to the actual situation, and on the other hand, the optimized sound pickup area can be determined through two stages, which can also improve the accuracy of the sound pickup area. Therefore, the sound pickup area can be adaptively adjusted.
[0106] Please refer to Figure 10 , Figure 10 FIG. is a schematic framework diagram of an embodiment of the computer-readable storage medium 100 of the present application. The computer-readable storage medium 100 stores program instructions 101 that can be run by a processor, and the program instructions 101 are used to implement the steps in any of the above-described method embodiments of the sound zone adjustment method.
[0107] In the above solution, during the voice interaction process of the vehicle occupants, once a preset voice command including the voice interaction person is collected, the sound zone adjustment can be started. And during the adjustment process, the initial sound zone judgment is first made according to the first audio, and then the adjustment is further made according to the frequency domain distribution of the second audio. On the one hand, the sound pickup area can be adjusted according to the actual situation, and on the other hand, the optimized sound pickup area can be determined through two stages, which can also improve the accuracy of the sound pickup area. Therefore, the sound pickup area can be adaptively adjusted.
[0108] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the foregoing method embodiments. The specific implementation can refer to the description of the foregoing method embodiments. For the sake of brevity, it will not be elaborated herein.
[0109] The descriptions of the foregoing embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0110] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical or other forms.
[0111] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0112] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.
Claims
1. A method for voice range adjustment, characterized in that, it includes: Obtain a first audio collected by an in-vehicle sound pickup device based on a preset sound pickup area; wherein, the first audio includes a preset voice command issued by a voice interaction person in the vehicle; Based on the first audio, perform a voice range judgment on the voice interaction person to obtain a first sound pickup area, and obtain a second audio collected by the in-vehicle sound pickup device based on the first sound pickup area; Perform frequency domain analysis based on the second audio to obtain the frequency domain distribution of the second audio; wherein, the frequency domain distribution includes: the sound frequencies of each sound source forming the second audio; In response to the frequency domain distribution satisfying a first condition, adjust the device parameters of the in-vehicle sound pickup device to obtain an optimized sound pickup area of the in-vehicle sound pickup device; wherein, the first condition includes: there is a second sound frequency after the first sound frequency is removed from the frequency domain distribution, the first sound frequency is the sound frequency of the voice interaction person, and the second sound frequency is a sound frequency different from the first sound frequency.
2. The method according to claim 1, characterized in that, The adjusting the device parameters of the in-vehicle sound pickup device to obtain the optimized sound pickup area of the in-vehicle sound pickup device includes: Adjust the device parameters based on a preset adjustment step size to obtain a second sound pickup area; wherein, the center line of the second sound pickup area coincides with the center line of the first sound pickup area, and the opening angles of the first sound pickup area and the second sound pickup area are different; Take the second sound pickup area as the new first sound pickup area, and re-execute the step of obtaining the second audio collected by the in-vehicle sound pickup device based on the first sound pickup area and subsequent steps until the frequency domain distribution does not satisfy the first condition, and take the latest first sound pickup area as the optimized sound pickup area.
3. The method according to claim 1, characterized in that, The in-vehicle sound pickup device is a microphone array, and the microphone array includes a plurality of microphones, and the device parameters include the distribution positions of the plurality of microphones in the vehicle.
4. The method according to claim 1, characterized in that, The method further includes: In response to the frequency domain distribution satisfying a second condition, determine that the voice interaction person has changed, and re-execute the step of obtaining the first audio collected by the in-vehicle sound pickup device based on the preset sound pickup area and subsequent steps; wherein, the second condition includes: the frequency domain distribution does not have the sound frequency of the voice interaction person.
5. The method according to claim 1, characterized in that, The performing a voice range judgment on the voice interaction person based on the first audio to obtain a first sound pickup area includes: Process the first audio based on beamforming to obtain a local judgment result regarding the first sound pickup area; and, Upload the first audio to the cloud, and receive a cloud judgment result regarding the first sound pickup area analyzed by the cloud based on the first audio; Based on the local judgment result and the cloud judgment result, determine the first sound pickup area.
6. The method according to claim 5, characterized in that, The local judgment result includes the local judged sound area, and the cloud judgment result includes the cloud judged sound area; determining the first sound pickup area based on the local judgment result and the cloud judgment result includes: Obtaining the overlap rate of the local judged sound area and the cloud judged sound area; In response to the overlap rate not satisfying the third condition, selecting the local judged sound area or the cloud judged sound area as the first sound pickup area, and / or, in response to the overlap rate satisfying the third condition, taking the union of the local judged sound area and the cloud judged sound area as the first sound pickup area; Wherein, the third condition includes: the overlap rate is not lower than a preset threshold.
7. The method according to claim 1, wherein, after judging the sound area of the voice interaction person based on the first audio to obtain the first sound pickup area, and before obtaining the second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area, the method further includes: Performing personnel analysis on the sensing data of the vehicle occupants by a vehicle-mounted sensing device to obtain the seating position of the voice interaction person in the vehicle; Performing a dual sound area judgment based on the first sound pickup area and the seating position to obtain a new first sound pickup area.
8. The method according to claim 7, wherein, performing a dual sound area judgment based on the first sound pickup area and the seating position to obtain a new first sound pickup area includes: In response to the seating position being within the first sound pickup area, keeping the first sound pickup area unchanged; and / or, in response to the seating position being outside the first sound pickup area, adjusting the first sound pickup area to cover the seating position to obtain a new first sound pickup area.
9. A sound area adjustment device, wherein, comprising: A first acquisition module, configured to acquire a first audio collected by a vehicle-mounted sound pickup device based on a preset sound pickup area; wherein, the first audio includes a preset voice command issued by a voice interaction person in the vehicle; A sound area judgment module, configured to judge the sound area of the voice interaction person based on the first audio to obtain a first sound pickup area; A second acquisition module, configured to acquire a second audio collected by the vehicle-mounted sound pickup device based on the first sound pickup area; A frequency domain analysis module, configured to perform frequency domain analysis on the second audio to obtain the frequency domain distribution of the second audio; wherein, the frequency domain distribution includes: the sound frequencies of each sound source forming the second audio; A sound area optimization module, configured to adjust the device parameters of the vehicle-mounted sound pickup device in response to the frequency domain distribution satisfying the first condition to obtain an optimized sound pickup area of the vehicle-mounted sound pickup device; wherein, the first condition includes: there is a second sound frequency after removing the first sound frequency from the frequency domain distribution, the first sound frequency is the sound frequency of the voice interaction person, and the second sound frequency is a sound frequency different from the first sound frequency.
10. A control device, wherein, Comprising a mutually coupled memory and a processor, program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the sound zone adjustment method according to any one of claims 1 to 8.
11. A vehicle-mounted control system, Characterized in that it comprises a vehicle-mounted sound pickup device and the control device according to claim 10, the vehicle-mounted sound pickup device is communicatively connected to the control device, and the vehicle-mounted sound pickup device is configured to collect audio data based on the sound pickup area currently configured by the vehicle-mounted sound pickup device.
12. The vehicle-mounted control system according to claim 11, Characterized in that the vehicle-mounted control system further comprises a vehicle-mounted sensing device, the vehicle-mounted sensing device is communicatively connected to the control device, and the vehicle-mounted sensing device is configured to collect sensing data by performing information collection on the vehicle occupants.
13. A computer-readable storage medium, Characterized in that it stores program instructions that can be run by a processor, and the program instructions are used to implement the sound zone adjustment method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Pickup method and device
CN107277699A
System and method for sound direction detection in vehicle
CN107852179A