A vehicle-mounted multi-sound-zone sound pickup device and a vehicle-mounted control system

By combining a single microphone array and a voice processing module, the problem of inaccurate sound zone recognition in vehicle voice control is solved, achieving efficient voice interaction with a simple structure and low cost.

CN114401470BActive Publication Date: 2025-12-23BEIJING HEXINDA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210035413.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-13
Publication Date
2025-12-23
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

In existing in-vehicle voice control technology, single microphone arrays cannot accurately identify the speaker's voice in each vocal range, resulting in complex system structure and high cost.

Method used

The system employs a combination of a single microphone array, a transceiver module, and a voice processing module. The single microphone array collects the raw sound signal, and the voice processing module determines the conventional beam and end-fire beam for each sound zone in the vehicle. Voice recognition is then performed and the signal is sent to the vehicle control equipment.

Benefits of technology

It achieves accurate recognition of speakers in different vocal registers in complex environments, improves the quality of voice interaction, and has the advantages of simple structure, convenient wiring and low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114401470B_ABST
    Figure CN114401470B_ABST
Patent Text Reader

Abstract

The present application relates to the field of vehicle voice, and provides a vehicle-mounted multi-sound-zone sound pickup device and a vehicle-mounted control system, wherein the vehicle-mounted multi-sound-zone sound pickup device comprises a single microphone array, a transceiving module and a voice processing module; the single microphone array is arranged in front of a cabin and is used for collecting a plurality of original sound signals; the transceiving module is connected with the voice processing module and a vehicle-mounted control device; the voice processing module is connected with the single microphone array and is used for receiving the plurality of original sound signals collected by the single microphone array; a normal beam of each sound zone in the vehicle and an end-on beam of a main driver sound zone and a deputy driver sound zone are determined according to the plurality of original sound signals; voice recognition is performed on the end-on beam and the normal beam of a rear seat sound zone to obtain a voice recognition result; and the normal beam and the voice recognition result are sent to the vehicle-mounted control device through the transceiving module. The present application can accurately recognize the sound of each sound zone, thereby improving the voice interaction quality in a complex environment and having the advantages of simple structure and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present document relates to the field of vehicle voice, and particularly relates to a vehicle-mounted multi-sound-zone sound pickup device and a vehicle-mounted control system. BACKGROUND

[0002] With the development of automobile intelligence, vehicle-mounted voice control technology appears, and vehicle-mounted voice control relies on in-vehicle sound pickup scheme, therefore, whether the accuracy of sound pickup will affect the vehicle control effect. In the prior art, the existing vehicle-mounted sound pickup scheme includes:

[0003] 1. Single-sound-zone sound pickup scheme: this scheme cannot distinguish speakers, therefore, it cannot provide more personalized services for each speaker, and when multiple speakers speak at the same time, there will be a problem of response error.

[0004] 2. Non-microphone array four-sound-zone scheme: this scheme arranges one microphone module at each door, and each microphone module corresponds to each sound zone. This scheme can obtain the voice of each sound zone speaker and provide personalized services, but in this scheme, each microphone module needs a separate interface chip and wiring, and there is a problem of complex system structure and high cost. SUMMARY

[0005] The present document is used to solve the problem in the prior art that the voice of each sound zone speaker cannot be accurately recognized by a single microphone array in vehicle voice control technology, and the sound of each sound zone speaker is collected by a microphone module, and there is a problem of complex system structure and high cost.

[0006] To solve the above technical problems, the first aspect of the present document provides a vehicle-mounted multi-sound-zone sound pickup device, comprising: a single microphone array, a transceiver module, and a voice processing module.

[0007] The single microphone array is arranged at the front of the cabin and includes a plurality of array elements for collecting a plurality of original sound signals.

[0008] The transceiver module is connected to the voice processing module and a vehicle-mounted control device, and is used to realize communication between the vehicle-mounted multi-sound-zone sound pickup device and the vehicle-mounted control device.

[0009] The voice processing module is connected to the single microphone array and is used to receive the plurality of original sound signals collected by the single microphone array; determine the normal beam of each sound zone in the vehicle and the end-fire beam of the main and vice driver sound zone according to the plurality of original sound signals; perform voice recognition on the end-fire beam and the normal beam of the rear seat sound zone to obtain a voice recognition result; and send the normal beam and the voice recognition result to the vehicle-mounted control device through the transceiver module.

[0010] As a further embodiment of the present document, the transceiver module is also connected to a vehicle-mounted entertainment device.

[0011] The voice processing module is further configured to receive a reference sound signal transmitted by the in-vehicle entertainment device through the transceiving module; and perform echo cancellation processing on the plurality of original sound signals by using the reference sound signal to obtain a sound signal to be analyzed.

[0012] The determination of the conventional beam of each sound zone and the end-fire beam of the driver and passenger sound zones according to the plurality of original sound signals is further based on the sound signal to be analyzed.

[0013] As a further embodiment herein, the voice processing module comprises an echo cancellation unit, a beam forming unit, a voice recognition unit and a processing unit connected in sequence.

[0014] The echo cancellation unit is configured to perform echo cancellation processing on the plurality of original sound signals by using the reference sound signal to obtain a sound signal to be analyzed.

[0015] The beam forming unit is configured to generate the conventional beam of each sound zone and the end-fire beam of the driver and passenger sound zones according to the sound signal to be analyzed.

[0016] The voice recognition unit is configured to perform voice recognition on the end-fire beam and the conventional beam of the rear seat sound zone to obtain a voice recognition result.

[0017] The processing unit is configured to transmit the conventional beam and the voice recognition result to the in-vehicle control device through the transceiving module.

[0018] As a further embodiment herein, the beam forming unit generates the end-fire beam of the driver and passenger sound zones according to the sound signal to be analyzed, comprising:

[0019] A first end of a straight line where the array elements are located is set as a first direction, and a second end of the straight line is set as a second direction, wherein the first direction points to the driver sound zone, and the second direction points to the passenger sound zone.

[0020] One of the array elements is selected as a reference array element, and the phase compensation is performed on the sound signal to be analyzed of each of the other array elements in a manner of forming the end-fire beam of the first direction, and the end-fire beam of the driver sound zone is obtained by weighted summation of the phase-compensated signals.

[0021] One of the array elements is selected as a reference array element, and the phase compensation is performed on the sound signal to be analyzed of each of the other array elements in a manner of forming the end-fire beam of the second direction, and the end-fire beam of the passenger sound zone is obtained by weighted summation of the phase-compensated signals.

[0022] As a further embodiment herein, the beam forming unit generates the conventional beam of each sound zone according to the sound signal to be analyzed, comprising:

[0023] According to the first direction and the second direction, a desired direction of each sound region is set;

[0024] One of the elements is selected as a reference element, and phase compensation is performed on each of the other elements to make the phases of the signals consistent;

[0025] For the desired direction of each sound region, the following operations are performed:

[0026] In the desired direction of each sound region other than the desired direction of the sound region, sound wave suppression is performed on the phase-compensated signals;

[0027] According to the sound wave-suppressed signals, a conventional beam in the desired direction of the sound region is obtained.

[0028] As a further embodiment herein, the speech processing module further comprises a noise suppression unit connected to the beamforming unit and the processing unit, configured to perform noise suppression processing on the conventional beam of each sound region, and the processing unit sends the noise-suppressed beam to the vehicle-mounted control device through the transceiver module.

[0029] As a further embodiment herein, the speech processing module further comprises a speech enhancement unit connected to the noise suppression unit and the processing unit, configured to perform speech enhancement processing on the noise-suppressed beam, and the processing unit sends the speech-enhanced sound region beam to the vehicle-mounted control device through the transceiver module.

[0030] As a further embodiment herein, the speech recognition unit performs speech recognition on the end-fire beam and the conventional beam of the backseat sound region to obtain a speech recognition result, which includes:

[0031] Speech endpoint detection is performed on the end-fire beam and the conventional beam of the backseat sound region respectively to determine whether there is speech in each sound region;

[0032] For the sound region with speech, the energy of the sound region is calculated;

[0033] The sound region with the maximum energy is set as an effective sound region;

[0034] The speech recognition result is composed of the speech endpoint detection result and the effective sound region.

[0035] In an embodiment herein, a vehicle-mounted control system is also provided, which comprises the vehicle-mounted multi-sound region sound pickup device of any of the preceding embodiments and a vehicle-mounted control device.

[0036] The embodiment designs the vehicle-mounted multi-sound area recognition device to include a single microphone array, a transceiver module, and a voice processing module, collects multiple original sound signals by using the single microphone array, realizes communication between the vehicle-mounted multi-sound area pickup device and the vehicle-mounted control device by using the transceiver module, receives the multiple original sound signals collected by the single microphone array by using the voice processing module, determines the regular beam of each sound area in the vehicle and the end-fire beam of the main and vice driver sound area according to the multiple original sound signals, performs voice recognition on the end-fire beam and the regular beam of the rear seat sound area to obtain a voice recognition result, and sends the regular beam and the voice recognition result to the vehicle-mounted control device through the transceiver module, so that the sound of each person in each sound area (for example, the independent pickup of four speakers in the main and vice driver sound area, the left rear sound area, and the right rear sound area) can be accurately distinguished by using the single microphone array, and the voice interaction quality in a complex environment is improved, and the structure is simple and the cost is low.

[0037] In order to make the above and other purposes, features and advantages of the present text more obvious and easy to understand, the following describes a preferred embodiment, and the accompanying drawings are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present text, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0039] Figure 1 A first structure diagram of the vehicle-mounted multi-sound area recognition device of the embodiment is shown;

[0040] Figure 2 A second structure diagram of the vehicle-mounted multi-sound area recognition device of the embodiment is shown;

[0041] Figure 3 A first structure diagram of the voice processing module of the embodiment is shown;

[0042] Figure 4 A flowchart of the end-fire beam forming process of the embodiment is shown;

[0043] Figure 5 A flowchart of the regular beam forming process of the embodiment is shown;

[0044] Figure 6 A second structure diagram of the voice processing module of the embodiment is shown;

[0045] Figure 7 A third structure diagram of the voice processing module of the embodiment is shown;

[0046] Figure 8 A flow chart of a voice recognition process of embodiments herein is shown;

[0047] Figure 9 A flow chart of a vehicle control system control of embodiments herein is shown;

[0048] Figure 10A An end-fire beam pattern of embodiments herein is shown;

[0049] Figure 10B A conventional beam pattern of embodiments herein is shown.

[0050] List of Symbols:

[0051] 100, a vehicle multi-zone sound pickup device;

[0052] 110, a single microphone array;

[0053] 120, a transceiver module;

[0054] 130, a voice processing module;

[0055] 200, a vehicle control device;

[0056] 300, a vehicle entertainment device;

[0057] 131, an echo cancellation unit;

[0058] 132, a beamforming unit;

[0059] 133, a voice recognition unit;

[0060] 134, a processing unit;

[0061] 135, a noise suppression unit;

[0062] 136, a voice enhancement unit;

[0063] 1001, an end-fire beam;

[0064] 1002, a conventional beam. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments herein will be described clearly and completely below with reference to the drawings in the embodiments herein. Obviously, the described embodiments are only part of the embodiments herein, rather than all the embodiments. Based on the embodiments herein, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present disclosure.

[0066] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, unless the contrary is indicated. For example, the description and drawings merely prefer certain arrangements, embodiments and procedures, which can vary, unless otherwise specifically stated. It is to be understood that such alterations can be practiced within the scope of the embodiments and / or methods described herein. In addition, it should be understood that the words "comprise", "have", "contain", and "include", and variations thereof, are used herein to indicate the presence of the stated feature but do not preclude the presence or addition of one or more other features. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so indicated by the applicant.

[0067] The specification provides method operational steps as described in the embodiments or flow charts, but can include more or less steps than those shown in the embodiments or flow charts. The order of the steps in the embodiments can be different than those shown in the embodiments or flow charts. The steps in the embodiments can be performed in parallel, in series, or in any order.

[0068] It should be noted that the sound signals involved in the present application are all information and data authorized by the user or fully authorized by all parties.

[0069] In an embodiment, a vehicle-mounted multi-sound area sound pickup device is provided, which is used in the existing vehicle voice control technology. The single microphone array cannot accurately identify the sound of the speaker in each sound area. The system structure is complex and the cost is high. The specific implementation is as shown in the Figure 1 The vehicle-mounted multi-sound area sound pickup device 100 includes a single microphone array 110, a transceiver module 120, and a voice processing module 130.

[0070] The single microphone array 110 is arranged in front of the cabin and includes a plurality of array elements for collecting a plurality of original sound signals. Each array element collects an original sound signal.

[0071] The transceiver module 120 is connected to the voice processing module 130 and the vehicle-mounted control device 200, and is used to realize the communication between the vehicle-mounted multi-sound area sound pickup device 100 and the vehicle-mounted control device 200.

[0072] The voice processing module 130 is connected with the single-microphone array 110, configured to receive a plurality of original sound signals collected by the single-microphone array 110; determine a regular beam of each sound area in the vehicle and an end-fire beam of the main driver sound area and the assistant driver sound area according to the plurality of original sound signals; perform voice recognition on the end-fire beam and the regular beam of the backseat sound area to obtain a voice recognition result; and transmit the regular beam and the voice recognition result to the vehicle-mounted control device 200 through the transceiver module 120.

[0073] The embodiment can separate the end-fire beam of the main driver sound area and the assistant driver sound area and the regular beam of each sound area by analyzing the original sound signals collected by the single-microphone array, and accurately recognize the speaker voice of each sound area by performing voice recognition on the end-fire beam and the regular beam of the backseat sound area, thereby improving the voice interaction quality in a complex environment and having the advantages of simple structure, convenient wiring and low cost.

[0074] In detail, the single-microphone array 110 described herein includes at least two array elements, which can be arranged in a straight line, a circle, etc., and the arrangement of the array elements is not limited herein. Each array element obtains an original sound signal, and each original sound signal is a set of sound signals generated by all speakers. The front of the cabin described herein is, for example, the front part of the sunroof, and the installation position of the single-microphone array is not limited herein.

[0075] The transceiver module 120 described herein is the same as the vehicle-mounted audio bus transceiver chip (for example, an A2B chip) in the vehicle-mounted system, configured to transmit the regular beam and the voice recognition result to the vehicle-mounted control device, so that the vehicle-mounted control device recognizes the regular beam with sound according to the regular beam and the voice recognition result, and then obtains the control command of each sound area to control the devices on the vehicle according to the preset control logic, for example, opening or closing the window, opening or closing the air purification system, opening or closing the sunroof, etc.

[0076] The voice processing module 130 described herein can be a central processing unit (CPU), or other programmable general-purpose or special-purpose microprocessors (Microprocessor), digital signal processors (Digital Signal Processor, DSP), programmable controllers, application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), or other similar elements or combinations of the above elements.

[0077] The end-fire beam described herein refers to a beam in the direction of extension of both ends of a single microphone array. Compared with a conventional beam, the end-fire beam has the advantages of strong directivity, narrow main lobe, large attenuation outside the main lobe, strong suppression of sound in the back of the main lobe, and small mutual interference between the main driver and the vice driver. The end-fire beam close to the main driver can accurately identify whether the main driver position has sound, and the end-fire beam close to the vice driver can accurately identify whether the vice driver position has sound. Through voice recognition of the end-fire beam, it can be accurately identified whether the main and vice drivers have sound, thereby avoiding recognition errors caused by voice interference of the main and vice drivers.

[0078] The conventional beam described herein refers to a beam pointing to the expected direction of an audio zone. Each audio zone expected direction is the center position of each audio zone, which can be determined according to the center of the seat, and can also be determined according to the user's riding habit. Compared with the end-fire beam, the conventional beam has a higher signal-to-noise ratio, and thus can better identify the voice content. By sending the conventional beam and the voice recognition result to the vehicle-mounted control device, the vehicle-mounted control device can determine the audio zone where the voice exists according to the voice recognition result, identify the conventional beam of the audio zone where the voice exists, and then obtain the specific voice content. If the voice content contains a control instruction, the control instruction is responded to.

[0079] In an embodiment, in order to avoid the voice played by the vehicle-mounted entertainment device from interfering with the speaker voice recognition, the transceiver module 120 is further connected to the vehicle-mounted entertainment device 300. Figure 2

[0080] The voice processing module 130 is further configured to receive a reference sound signal sent by the vehicle-mounted entertainment device 300 through the transceiver module 120; perform echo cancellation processing on the plurality of original sound signals to obtain a to-be-analyzed sound signal; determine the conventional beam of each audio zone in the vehicle and the end-fire beam of the main and vice driver audio zones according to the plurality of original sound signals; and further determine the conventional beam of each audio zone in the vehicle and the end-fire beam of the main and vice driver audio zones according to the to-be-analyzed sound signal.

[0081] The vehicle-mounted entertainment device 300 described herein includes a vehicle-mounted playback device, a vehicle-mounted loudspeaker, etc. The vehicle-mounted playback device can play audio in a radio station or a removable storage device. The number of reference sound signals is the same as the number of sound channels played by the vehicle-mounted playback device. For example, if the vehicle-mounted playback device is N-channel, the reference sound signal is N. In specific implementation, the transceiver module 120 can also be wirelessly connected to a smart terminal (such as a smart phone, a PAD, etc.) for receiving a reference sound signal sent by the smart terminal.

[0082] ​The voice processing module 130 can utilize a band-pass filter to perform echo cancellation on the plurality of original sound signals using the reference sound signal, and the specific implementation process can refer to the prior art, which is not limited herein. The sound signal to be analyzed is the sound signal after echo cancellation, i.e., the signal after eliminating the noise interference of the in-vehicle entertainment device.

[0083] The embodiment can effectively eliminate the sound generated by the in-vehicle entertainment device, thereby improving the signal-to-noise ratio of the speaker voice signal.

[0084] In an embodiment, as shown in Figure 3 The voice processing module 130 includes an echo cancellation unit 131, a beamforming unit 132, a voice recognition unit 133, and a processing unit 134 connected in sequence.

[0085] The echo cancellation unit 131 is configured to perform echo cancellation on the plurality of original sound signals using the reference sound signal to obtain a sound signal to be analyzed.

[0086] The beamforming unit 132 is configured to generate a conventional beam for each sound zone and an end-fire beam for the main driver sound zone and the copilot sound zone according to the sound signal to be analyzed.

[0087] The voice recognition unit 133 is configured to perform voice recognition on the end-fire beam and the conventional beam corresponding to the backseat sound zone to obtain a voice recognition result.

[0088] The processing unit 134 is configured to send the conventional beam and the voice recognition result to the in-vehicle control device through the transceiver module.

[0089] In detail, the echo cancellation unit 131 can implement echo cancellation through hardware or software, and the specific echo cancellation process can refer to the prior art, which is not described herein.

[0090] The beamforming unit 132 and the voice recognition unit 133 can refer to the subsequent embodiments for implementation.

[0091] The processing unit 134 is connected to the voice recognition unit 133 and the beamforming unit 132 and is configured to control the specific content sent to the in-vehicle control device. In specific implementation, the function implemented by the processing unit 134 can also be integrated into the voice recognition unit 133.

[0092] In an embodiment, as shown in Figure 4 The beamforming unit 132 generates an end-fire beam for the main driver sound zone and the copilot sound zone according to the sound signal to be analyzed, and the process includes:

[0093] Step 411: Set a first end of a straight line where an array element is located as a first direction and a second end as a second direction, wherein the first direction is close to the main driver sound zone, and the second direction is close to the copilot sound zone.

[0094] Step 412, one of the array elements is selected as a reference array element, and the other array elements each have a sound signal to be analyzed. The sound signals are phase compensated in a manner of forming a first direction end-fire beam, and the phase-compensated signals are weighted and summed to obtain an end-fire beam of the main driver sound area;

[0095] Step 413, one of the array elements is selected as a reference array element, and the other array elements each have a sound signal to be analyzed. The sound signals are phase compensated in a manner of forming a second direction end-fire beam, and the phase-compensated signals are weighted and summed to obtain an end-fire beam 1001 of the co-driver sound area, as shown in Figure 10A .

[0096] In detail, the first direction and the second direction described in the present text are 180° apart. In a specific implementation, the first direction is, for example, 0°, and the second direction is, for example, 180°, or the first direction is 180°, and the second direction is 0°. In subsequent implementations, for example, no special description is made, and both refer to the first direction being 0° and the second direction being 180°.

[0097] In the implementation of step 412 and step 413, the noise ratio between the signals is calculated by using the following formula:

[0098]

[0099] wherein B is the signal-to-noise ratio, ω is the angular frequency, α is proportional to, and θ is the angle between the suppression direction and the line on which the array elements are located, is the phase compensation value.

[0100] When θ is 0°, the sound wave in the 0° direction can be suppressed, and the 180° direction has the maximum signal-to-noise ratio, that is, the 0° direction has the maximum signal. The end-fire beam of the co-driver sound area can be calculated by weighting and summing the sound wave suppressed signals.

[0101] When θ is 180°, the sound wave in the 180° direction can be suppressed, and the 0° direction has the maximum signal-to-noise ratio, that is, the 0° direction has the maximum signal. The end-fire beam of the main driver sound area can be calculated by weighting and summing the sound wave suppressed signals.

[0102] In an embodiment of the present text, as shown in Figure 5 , the beam forming unit 132 generates a conventional beam for each sound area according to the sound signals to be analyzed, including:

[0103] Step 511, according to the first direction and the second direction, set the expected direction of each sound area;

[0104] Step 512, one of the array elements is selected as a reference array element, and the other array elements each have a sound signal to be analyzed. The sound signals are phase compensated to make the phases of the sound signals to be analyzed consistent;

[0105] Step 513, for the expected direction of each sound area, the following operations are performed:

[0106] Step 5131, sound wave suppression is performed on the phase-compensated signals in the desired direction of the other sound zones except the desired direction of the sound zone;

[0107] Step 5132, the sound wave-suppressed signals are weighted and summed to obtain the conventional beam 1002 of the sound zone in the desired direction, as shown in the following formula: Figure 10B As shown, there is no mutual overlap between the conventional beams.

[0108] When step 511 is implemented, the desired direction of each sound zone can be set according to the included angle between the line connecting the seat center and a point on the single microphone array and the line on which the array elements are located.

[0109] When step 5131 is implemented, sound wave suppression in the desired direction of each sound zone is achieved by the signal-to-noise ratio between the phase-compensated signals and the reference signal.

[0110] When step 5132 is implemented, the conventional beam of each sound zone in the desired direction can be calculated by the following formula:

[0111]

[0112] wherein S is the conventional beam, s0 is the sound wave signal in the desired direction, s i is the sound wave signal after sound wave suppression in the other desired direction, and N is the number of array elements.

[0113] In an embodiment, as shown in the following figure, the voice processing module 130 further includes a noise suppression unit 135 connected to the beam forming unit 132 and the processing unit 134, for performing noise suppression processing on the conventional beam of each sound zone, and sending the noise-suppressed conventional beam to the vehicle control device 200 by the transceiver module 120 under the control of the processing unit 134. Figure 6 In detail, the noise suppressed by the noise suppression unit 135 includes but is not limited to wind noise, tire noise, fan noise, etc.

[0114] By further noise suppression on the conventional beam, the embodiment can reduce environmental noise, improve signal-to-noise ratio, and improve the accuracy of voice recognition of the vehicle control device.

[0115] In an embodiment, as shown in the following figure, the voice processing module 130 further includes a voice enhancement unit 136 connected to the noise suppression unit, for performing voice enhancement processing on the noise-suppressed beam, and sending the voice-enhanced sound zone beam to the vehicle control device by the transceiver module under the control of the processing unit 134.

[0116] Figure 7

[0117] ​​The sound zone beam determined by the embodiment is a clean independent sound, and by identifying the sound zone beam, the command of each sound zone can be determined.

[0118] In an embodiment, as shown in Figure 8 The voice recognition unit 133 performs voice recognition on the end-fire beam of the main and co-driver sound zone and the regular beam of the rear seat sound zone, and obtains a voice recognition result including:

[0119] Step 811, voice endpoint detection (VAD) is performed on the end-fire beam of the main and co-driver sound zone and the regular beam of the rear seat sound zone, to determine whether there is voice in each sound zone.

[0120] Step 812, for the sound zone with voice, the energy of the sound zone is calculated.

[0121] Step 813, the energy of each sound zone is evaluated, and the sound zone with the maximum energy is set as the effective sound zone.

[0122] Step 814, the voice endpoint detection result and the effective sound zone constitute the voice recognition result.

[0123] When steps 811 and 812 are implemented, the existing voice endpoint detection process and the existing beam energy calculation process can be referred to, and details are not described here.

[0124] In a specific implementation, in order to save unnecessary calculation, step 811 further includes: determining whether there is voice in several sound zones in the result, if the detection result is that there is no voice in all sound zones, determining that the voice recognition result is "there is no effective voice in all sound zones", if the detection result is that there is voice in only one beam, determining that "the corresponding sound zone of the beam is effective voice", and if the detection result is that there is voice in two or more beams, executing step 812 and the steps thereafter. Whether there is voice in each sound zone can be determined by the following configuration, for example, 1 represents that there is voice, and 0 represents that there is no voice.

[0125] In an embodiment, a vehicle-mounted control system is also provided, which includes the vehicle-mounted multi-sound zone sound pickup device of any of the preceding embodiments and a vehicle-mounted control device.

[0126] In order to more clearly illustrate the technical solutions of the present application, a specific embodiment is described in detail below. Specifically, as shown in Figure 9 The vehicle-mounted control system includes a vehicle-mounted multi-sound zone sound pickup device 100 and a vehicle-mounted control device 200, and the vehicle-mounted multi-sound zone sound pickup device 100 and the vehicle-mounted control device 200 communicate with each other through an A2B chip.

[0127] Specifically, the voice vehicle-mounted control process includes:

[0128] Step 911, the single microphone array collects original sound signals, i.e. the sound uttered by the user, and sends the original sound signals to the voice processing module.

[0129] Step 912, the vehicle-mounted entertainment device 300 sends the reference sound signals to the voice processing module.

[0130] Step 913, the voice processing module performs the following processing process:

[0131] (1) Perform echo cancellation processing on the multiple original sound signals based on the reference sound signals to obtain a to-be-analyzed sound signal;

[0132] (2) Determine the regular beams of each sound zone in the vehicle and the end-fire beams of the main and assistant driver sound zones based on the to-be-analyzed sound signal;

[0133] (3) Perform voice recognition on the end-fire beams and the regular beams of the backseat sound zone to obtain a voice recognition result;

[0134] (4) Send the regular beams and the voice recognition result to the vehicle-mounted control device through the transceiver module.

[0135] Step 914, the vehicle-mounted control device performs control operations based on the regular beams, the voice recognition result, and the control mode selected by the user.

[0136] Specifically, the control mode includes a preemptive mode and a non-preemptive mode.

[0137] When the control mode is the preemptive mode, the effective sound zone corresponding to the regular beam is identified based on the effective sound zone in the voice recognition result, the control content is obtained, the control instruction is sent to the corresponding controller based on the control content, and the control is completed by the corresponding controller. For example, if the control content is to open the window, an open control instruction is sent to the window controller, and the window controller opens the window.

[0138] When the control mode is the non-preemptive mode, the regular beam of the sound zone with sound is identified based on the voice endpoint detection result in the voice recognition result, the control content of each sound zone is obtained, the control instruction is sent to the corresponding controller based on the control content of each sound zone, and the control is completed by the corresponding controller. For example, if the left backseat sound zone control content is to open the window, an open left back window instruction is sent to the window controller, and the window controller opens the left back window. If the right backseat sound zone control content is to open the air conditioner, an open right backseat air conditioner instruction is sent to the air conditioner controller, and the air conditioner controller opens the right backseat air conditioner.

[0139] The vehicle-mounted multi-sound area pickup device and the vehicle-mounted control system provided in the present application are characterized in that the vehicle-mounted multi-sound area recognition device is designed to include a single microphone array, a transceiving module, and a voice processing module; the single microphone array is used to collect a plurality of original sound signals; the transceiving module is used to realize communication between the vehicle-mounted multi-sound area pickup device and the vehicle-mounted control device; the voice processing module is used to receive the plurality of original sound signals collected by the single microphone array; the regular beam of each sound area in the vehicle and the end-on beam of the main and vice driver sound area are determined according to the plurality of original sound signals; voice recognition is performed on the end-on beam and the regular beam of the rear seat sound area to obtain a voice recognition result; and the regular beam and the voice recognition result are sent to the vehicle-mounted control device through the transceiving module. The single microphone array can accurately distinguish the sound of each sound area (for example, the independent pickup of the sound of four speakers in the main and vice driver sound area, the left rear sound area, and the right rear sound area), thereby improving the voice interaction quality in a complex environment, and having the advantages of simple structure, convenient wiring, and low cost.

[0140] It should be understood that, in various embodiments herein, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments herein.

[0141] It should also be understood that, in the embodiments herein, the term "and / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships. For example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " in the present application generally represents an "or" relationship between the front and rear associated objects.

[0142] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0143] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0144] In several embodiments provided herein, it should be understood that the disclosed system, apparatus and method can be implemented in other manners. For example, the apparatus embodiments described above are merely schematic. For example, the division of the units is only a logical function division. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electric, mechanical or other form.

[0145] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e. may be located in one place, or may be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments provided herein.

[0146] In addition, each functional unit in each embodiment herein can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0147] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions herein, essentially or the part of the prior art that contributes to the technical solutions, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment herein. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0148] The specific embodiments are applied herein to describe the principles and implementation manners of the embodiments herein. The above description of the embodiments is only to help understand the methods and core ideas thereof; meanwhile, for those skilled in the art, according to the ideas herein, the specific implementation manners and application ranges can be changed; and in view of the above, the content of the specification should not be understood as limiting the embodiments herein.

Claims

1. A vehicle-mounted multi-sound-zone sound pickup device characterized by comprising: The application relates to a vehicle multi-sound-zone sound pickup device. The single microphone array is arranged in front of a cabin and comprises two array elements for collecting multiple original sound signals. The transceiver module is connected with the voice processing module and a vehicle control device, and is used for realizing communication between the vehicle multi-sound-zone sound pickup device and the vehicle control device. The voice processing module is connected with the single microphone array and is used for receiving the multiple original sound signals collected by the single microphone array. The transceiver module is also connected with a vehicle entertainment device, and the voice processing module is also used for receiving a reference sound signal sent by the vehicle entertainment device through the transceiver module. The voice processing module comprises an echo cancellation unit, a beam forming unit, a voice recognition unit and a processing unit which are connected in sequence. The echo cancellation unit is used for performing echo cancellation processing on the multiple original sound signals by using the reference sound signal to obtain sound signals to be analyzed. The beam forming unit is used for generating a normal beam of a rear seat sound zone and end-on beams of a main driver sound zone according to the sound signals to be analyzed. The voice recognition unit is used for performing voice recognition on the end-on beams of the main driver sound zone and the normal beam of the rear seat sound zone to obtain voice recognition results. The processing unit is used for sending the normal beam and the voice recognition results to the vehicle control device through the transceiver module. A first end of a straight line where the array elements are located is defined as a first direction, and a second end is defined as a second direction, wherein the first direction points to the main driver sound zone, the second direction points to the co-driver sound zone, and the first direction and the second direction are 180 degrees apart. One of the array elements is selected as a reference array element, and phase compensation is performed on the sound signals to be analyzed of the other array elements in a manner of forming a first direction end-on beam, and the phase-compensated signals are weighted and summed to obtain an end-on beam of the main driver sound zone. One of the array elements is selected as a reference array element, and phase compensation is performed on the sound signals to be analyzed of the other array elements in a manner of forming a second direction end-on beam, and the phase-compensated signals are weighted and summed to obtain an end-on beam of the co-driver sound zone. The voice recognition unit performs voice recognition on the end-on beams of the main driver sound zone and the normal beam of the rear seat sound zone to obtain voice recognition results, which comprises the following steps. Voice endpoint detection is performed on the end-on beam of the main driver sound zone and the end-on beam of the co-driver sound zone respectively to determine whether there is voice in the main driver sound zone or the co-driver sound zone. Voice endpoint detection is performed on the normal beam of the rear seat sound zone to determine whether there is voice in the rear seat sound zone. For the sound zone with voice, the energy of the sound zone is calculated. The sound zone with the maximum energy is set as an effective sound zone. The voice recognition results are formed by the voice endpoint detection results and the effective sound zone. The beam forming unit generates normal beams of each sound zone according to the sound signals to be analyzed, which comprises the following steps.

2. The vehicle-mounted multi-zoned sound pickup device according to claim 1, wherein Desired directions of each sound zone are set according to the first direction and the second direction. One of the array elements is selected as a reference array element, and phase compensation is performed on the sound signals to be analyzed of the other array elements to make the phases of the sound signals to be analyzed consistent. For the desired direction of each sound zone, the following operations are performed. ​ The phase-compensated signals are subjected to sound wave suppression in the desired direction of other sound zones except the desired direction of the sound zone; and a normal beam of the sound zone in the desired direction is obtained according to the sound wave-suppressed signals.

3. The vehicle-mounted multi-zonal sound pickup device according to claim 1, wherein The voice processing module further comprises: A noise suppression unit connected to the beam forming unit and the processing unit, for performing noise suppression processing on the normal beam of each sound zone, and sending the noise-suppressed beam to the vehicle-mounted control device by the processing unit through the transceiving module.

4. The vehicle-mounted multi-zoned sound pickup device of claim 3, wherein The voice processing module further comprises: A voice enhancement unit connected to the noise suppression unit and the processing unit, for performing voice enhancement processing on the noise-suppressed beam, and sending the voice-enhanced sound zone beam to the vehicle-mounted control device by the processing unit through the transceiving module.

5. An in-vehicle control system characterized by comprising: Comprise: A vehicle-mounted control device and a vehicle-mounted multi-sound zone sound pickup device according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • In-vehicle voice frequency signal processing method and device

    CN109545230A