A vehicle-mounted voice interaction method, device and electronic device
By removing noise noise and determining the intersection of sound information in the sensors in the cabin of the car, the problem of noise interference affecting voice communication is solved, and a clearer voice interaction effect is achieved.
Patent Information
- Application Number
- CN202110998594.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-08-27
AI Technical Summary
The noise interference in the car cockpit causes the quality of voice communication to decline, especially when the front row staff is low, it is difficult for the rear row staff to hear the sound in the front row.
By obtaining the sound information collected by various sensors, deleting noise noises that are not earlier than other sensors, determining the intersection of sound information to determine the target voice, and playing through the target speaker.
It effectively reduces noise interference, improves the quality of voice interaction, and makes communication between users and their voice interaction objects clearer.
Smart Images

Figure CN115731939B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automotive electronics and electrical appliances, and particularly to a vehicle voice interaction method, device, and electronic device. Background Art
[0002] The cockpit of a vehicle is an enclosed space. Due to relatively common noise interference caused by sound refraction in the cockpit, voice communication between people in the cockpit is affected by noise interference. For example, if the voices of the front-row passengers in the vehicle are relatively low, the large noise interference in the vehicle can cause the rear-row passengers to be unable to hear the voices of the front-row passengers when the front-row and rear-row passengers are having a voice conversation. Therefore, how to reduce the noise interference during voice interaction between people in the vehicle has become an urgent problem to be solved. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a vehicle voice interaction method, device, and electronic device to reduce the noise interference during voice interaction between people in the vehicle.
[0004] To achieve the above purpose, the embodiments of the present invention provide a vehicle voice interaction method, including:
[0005] When receiving a voice interaction instruction issued by a user, obtaining a set of first sound information collected by each first type of sensor and a set of second sound information collected by a second type of sensor; wherein, the voice interaction instruction includes: identification information of a target speaker corresponding to the voice interaction object of the user.
[0006] For each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, then deleting the first sound information in the set of first sound information corresponding to the first type of sensor to obtain a target first sound information set corresponding to the first type of sensor.
[0007] For each target first sound information set, determining an intersection of sound information between the target first sound information set and the set of second sound information.
[0008] Determining a target voice according to the intersection of target sound information, and playing the target voice through the target speaker based on the identification information; wherein, the intersection of target sound information is: the intersection of sound information corresponding to the first type of sensor corresponding to the seat where the user is located.
[0009] Optionally, before obtaining the set of first sound information collected by each first type of sensor and the set of second sound information collected by the second type of sensor, it further includes:
[0010] Obtain the window state information and door state information of the vehicle;
[0011] Based on the window state information and the door state information, adjust the sensitivity of each first-type sensor in the vehicle.
[0012] Optionally, the adjusting the sensitivity of each first-type sensor in the vehicle based on the window state information and the door state information includes:
[0013] For each first-type sensor, calculate the product of the preset door influence coefficient of the door closest to the seat corresponding to the first-type sensor and the opening / closing degree of the door as the first product;
[0014] Calculate the product of the preset window influence coefficient of the window closest to the seat corresponding to the first-type sensor and the opening / closing degree of the window as the second product;
[0015] Calculate the sum of the first product, the second product and the theoretical sensitivity of the first-type sensor as the adjusted sensitivity of the first-type sensor.
[0016] Optionally, before determining the target voice according to the intersection of the target voice information, it further includes:
[0017] Based on the speaker identifier, eliminate the speaker voice information in each of the voice information intersections to obtain the corresponding eliminated voice information intersections;
[0018] Determine the eliminated voice information intersection corresponding to the first-type sensor corresponding to the seat where the user is located as the target voice information intersection.
[0019] Optionally, each of the first-type sensors is a non-omnidirectional sensor corresponding to a seat in the vehicle, and the second-type sensor is an omnidirectional sensor capable of collecting all the voice information in the vehicle.
[0020] To achieve the above object, an embodiment of the present invention provides an in-vehicle voice interaction device, including:
[0021] An information receiving module, configured to obtain a set of first voice information collected by each first-type sensor and a set of second voice information collected by a second-type sensor when receiving a voice interaction instruction issued by a user; wherein, the voice interaction instruction includes: identification information of a target speaker corresponding to the voice interaction object of the user;
[0022] The first information elimination module is used to, for each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, delete the first sound information from the set of first sound information corresponding to the first type of sensor, so as to obtain the target set of first sound information corresponding to the first type of sensor;
[0023] The intersection determination module is used to, for each target set of first sound information, determine the intersection of sound information between the target set of first sound information and the set of second sound information;
[0024] The target voice determination module is used to determine the target voice according to the intersection of target sound information and play the target voice through the target speaker based on the identification information; wherein, the intersection of target sound information is: the intersection of sound information corresponding to the first type of sensor corresponding to the seat where the user is located.
[0025] Optionally, the device further includes:
[0026] The sensitivity adjustment module is used to obtain the window state information and door state information of the vehicle; based on the window state information and the door state information, adjust the sensitivity of each first type of sensor in the vehicle.
[0027] Optionally, the sensitivity adjustment module is specifically used to, for each first type of sensor, calculate the product of the preset door influence coefficient of the door closest to the seat corresponding to the first type of sensor and the opening and closing degree of the door as the first product; calculate the product of the preset window influence coefficient of the window closest to the seat corresponding to the first type of sensor and the opening and closing degree of the window as the second product; calculate the sum value of the first product, the second product and the theoretical sensitivity of the first type of sensor as the adjusted sensitivity of the first type of sensor.
[0028] Optionally, the device further includes: a second information elimination module, which is used to eliminate the speaker sound information in each intersection of sound information based on the speaker identification, so as to obtain each corresponding intersection of sound information after elimination; determine the intersection of sound information after elimination corresponding to the first type of sensor corresponding to the seat where the user is located as the intersection of target sound information.
[0029] Optionally, each of the first type of sensors is a non-omnidirectional sensor corresponding to a seat in the vehicle, and the second type of sensor is an omnidirectional sensor capable of collecting all sound information in the vehicle.
[0030] To achieve the above object, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0031] The memory is used to store a computer program;
[0032] When the processor is used to execute the program stored on the memory, it implements the steps of any one of the above-mentioned vehicle-mounted voice interaction methods.
[0033] To achieve the above object, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any one of the above-mentioned vehicle-mounted voice interaction methods.
[0034] To achieve the above object, an embodiment of the present invention further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the steps of any one of the above-mentioned vehicle-mounted voice interaction methods.
[0035] Beneficial effects of the embodiments of the present invention:
[0036] By using the method provided by the embodiment of the present invention, when a voice interaction instruction sent by a user is received, a set of first sound information collected by each first type of sensor and a set of second sound information collected by a second type of sensor are obtained; for each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, then the first sound information in the set of first sound information corresponding to the first type of sensor is deleted, and a target set of first sound information corresponding to the first type of sensor is obtained; for each target set of first sound information, a sound information intersection between the target set of first sound information and the set of second sound information is determined; a target voice is determined according to the target sound information intersection, and the target voice is played through a target speaker based on identification information. By deleting the first sound information in the set of first sound information corresponding to the first type of sensor whose collection time is not earlier than the collection time of other first type of sensors, the noise and noise collected by each first type of sensor are further deleted, and a sound information intersection is obtained with the set of second sound information, further eliminating the noise and noise collected by the first type of sensor. Then, according to the target sound information intersection, the target voice is determined and sent to the target speaker corresponding to the user's voice interaction object for playback, reducing the noise interference during the voice interaction between the user and their voice interaction object and improving the quality of the voice interaction.
[0037] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.
[0039] Figure 1 A flowchart of a vehicle-mounted voice interaction method provided by an embodiment of the present invention;
[0040] Figure 2 A layout diagram of the settings of sensors in a vehicle;
[0041] Figure 3 A schematic diagram of the sound information collected by each sensor;
[0042] Figure 4 A circuit diagram corresponding to the bias voltage of the first type of sensor;
[0043] Figure 5 A schematic diagram of the distance from the sound information to each sensor;
[0044] Figure 6 A schematic diagram of the acquisition range of the sensor;
[0045] Figure 7 A schematic diagram of the intersection of sound information;
[0046] Figure 8 A structural diagram of a vehicle-mounted voice interaction device provided by an embodiment of the present invention;
[0047] Figure 9 Another structural diagram of a vehicle-mounted voice interaction device provided by an embodiment of the present invention;
[0048] Figure 10 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope of protection of the present invention.
[0050] The in-vehicle voice interaction method provided by the embodiments of the present invention can be applied to in-vehicle electronic devices, edge servers, cloud servers, etc., without specific limitation here. Among them, the in-vehicle electronic device can be any in-vehicle information interaction terminal such as an in-vehicle central computer, a central domain controller, an integrated ECU (Electronic Control Unit), a driving brain, a car machine, a DHU (Drilling Head Unit, an integrated machine of an entertainment host and an instrument), an IHU (Infotainment Head Unit), or an IVI (In-Vehicle Infotainment system). Taking the in-vehicle central domain controller as an example of the in-vehicle electronic device, the in-vehicle central domain controller can control various devices in the vehicle, such as speakers, sensors, etc., and integrate the functions of various devices in the vehicle. For example, the in-vehicle central domain controller can obtain the data collected by sensors, such as the sound information in the vehicle, and process it to control the speaker to play the processed sound information. That is, various information in the vehicle can be processed by the in-vehicle central domain controller that can control various devices in the vehicle.
[0051] Figure 1 is a flowchart of the in-vehicle voice interaction method provided by the embodiments of the present invention. As Figure 1 shown, the method includes the following steps:
[0052] Step 101, when receiving a voice interaction instruction issued by a user, obtain a set of first sound information collected by each first type of sensor and a set of second sound information collected by a second type of sensor.
[0053] Among them, the voice interaction instruction includes: identification information of a target speaker corresponding to the user's voice interaction object.
[0054] Specifically, each first type of sensor can be a non-omnidirectional sensor corresponding to a seat in the vehicle, and the second type of sensor can be an omnidirectional sensor capable of collecting all sound information in the vehicle.
[0055] Step 102, for each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, delete the first sound information in the set of first sound information corresponding to the first type of sensor to obtain a target set of first sound information corresponding to the first type of sensor.
[0056] Step 103, for each target set of first sound information, determine the intersection of sound information between the target set of first sound information and the set of second sound information.
[0057] Step 104: Determine the target voice according to the intersection of the target voice information, and play the target voice through the target speaker based on the identification information.
[0058] Among them, the intersection of the target voice information is: the intersection of the voice information corresponding to the first type of sensor corresponding to the seat where the user is located.
[0059] In the embodiment of the present invention, the voice interaction instruction may include the identification information of the target speaker corresponding to the voice interaction object of the user. Specifically, multiple buttons can be set beside each seat, and each button identifies the speakers corresponding to other seats. For example, in a 4-seat vehicle, there are buttons B, C, and D set at seat A, which respectively identify the speakers at seats B, C, and D; there are buttons A, C, and D set at seat B, which respectively identify the speakers at seats A, C, and D; there are buttons A, B, and D set at seat C, which respectively identify the speakers at seats A, B, and D; there are buttons B, C, and A set at seat D, which respectively identify the speakers at seats B, C, and A. If the user at seat A presses button B, it means that the user needs to have a voice conversation with their voice interaction object (the user at seat B).
[0060] In the embodiment of the present invention, the speakers can be set on the front, rear, left, and right doors of the vehicle, and each seat corresponds to one speaker.
[0061] If the target voice is determined for the user at seat A, the target speaker for playing the target voice can be determined as the speaker corresponding to seat B based on the identification information, and then the target voice can be played through the speaker corresponding to seat B.
[0062] Using the method provided by the embodiments of the present invention, when a voice interaction instruction sent by a user is received, a set of first sound information collected by each first type of sensor and a set of second sound information collected by a second type of sensor are obtained; for each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, then the first sound information in the set of first sound information corresponding to the first type of sensor is deleted, and a target set of first sound information corresponding to the first type of sensor is obtained; for each target set of first sound information, a sound information intersection between the target set of first sound information and the set of second sound information is determined; a target voice is determined according to the target sound information intersection, and the target voice is played through a target speaker based on identification information. By deleting the first sound information in the set of first sound information corresponding to the first type of sensor whose collection time is not earlier than the collection time of other first type of sensors, the noise and miscellaneous sounds collected by each first type of sensor are further deleted, and a sound information intersection is calculated with the set of second sound information, further eliminating the noise and miscellaneous sounds collected by the first type of sensors. Then, the target voice is determined according to the target sound information intersection and sent to the target speaker corresponding to the voice interaction object of the user for playing, reducing the noise interference during the voice interaction between the user and their voice interaction object and improving the quality of the voice interaction.
[0063] The method provided by the embodiments of the present invention can be applied to any vehicle, such as a 4-seater vehicle and a 6-seater vehicle. In each vehicle, a plurality of first type of sensors equal in number to the number of seats can be set, and a second type of sensor can be set. Figure 2 It is a layout diagram of the sensors inside the vehicle. Figure 2 The layout diagram of the first type of sensor and the second type of sensor in the 4-seater vehicle shown, such as Figure 2 As shown, the 4-seater vehicle includes 4 seats: Seat 1, Seat 2, Seat 3, and Seat 4. A first type of sensor MIC1 can be set near Seat 1, a second type of sensor MIC2 can be set near Seat 2, a first type of sensor MIC3 can be set near Seat 3, a first type of sensor MIC4 can be set near Seat 4, and a second type of sensor MIC5 can be set at the central position in the vehicle. Among them, each first type of sensor is a non-omnidirectional sensor corresponding to a seat inside the vehicle. Generally, the non-omnidirectional sensor can be set to have a collection range less than 180°, such as Figure 2The collection ranges a1 of the first type of sensor MIC1, a2 of the first type of sensor MIC2, a3 of the first type of sensor MIC3, and a4 of the first type of sensor MIC2 are all less than 180°. Moreover, each first type of sensor is biased towards its corresponding seat. For example, the first type of sensor MIC1 is biased towards its corresponding seat 1. Such a setting facilitates the first type of sensor to collect the voice of the user at its corresponding seat.
[0064] The second type of sensor at the vehicle center position is an omnidirectional sensor. Generally, the omnidirectional sensor can be set with a collection range of 360°, such as Figure 2 In [the relevant context], the collection range a5 of the second type of sensor MIC5 is 360°, that is, the second type of sensor can collect all sound information inside the vehicle, including background noise, noise, etc.
[0065] In the embodiments of the present invention, the second type of vehicle sensor and each first type of sensor can both include performance indicators: sensitivity, signal-to-noise ratio, frequency response, and directivity. The definitions of each performance indicator are specifically as described below:
[0066] The sensitivity of the sensor is: when a sound pressure of 1 pa (94 dB) is fed back to the sensor, the voltage (dBV) at the output end of the sensor. In the embodiments of the present invention, generally, the sensitivity of the first type of sensor and the second type of sensor inside the vehicle is required to be ≥ -35 dB.
[0067] The signal-to-noise ratio of the sensor is: the ratio of the signal to the noise. In the embodiments of the present invention, generally, the signal-to-noise ratio of the first type of sensor and the second type of sensor inside the vehicle is required to be ≥ -40 dB; moreover, since the first type of sensor is usually closer to the windows and doors, the signal-to-noise ratio of the first type of sensor is greater than that of the second type of sensor.
[0068] The frequency response of the sensor is: the different values of the sensitivity corresponding to different frequencies are the frequency response. Representing the dependence of the sensitivity on the frequency by a curve is called the frequency response curve. In the embodiments of the present invention, the frequency response requirement of the first type of sensor is usually: high sensitivity in the frequency response within the voice signal frequency range, where the voice signal frequency range is: + / - 3 dB (300 Hz - 3 kHz). The frequency range of the second type of sensor can be: 300 Hz to 3 kHz.
[0069] The directivity of the sensor is: the pointing range of the first type of sensor is less than <180°, and the pointing range of the second type of sensor can be equal to 360°.
[0070] In the embodiments of the present invention, the first type of sensor can collect the voice information of users within its pointing range, some environmental sounds inside the vehicle (noise inside the vehicle, noise), and the sounds played by the speaker; for example, Figure 2 As shown, the first type of sensor MC1 can collect the voice information of the user on seat 1 within its pointing range a1, some environmental sounds within its pointing range, and the sounds played by the speaker. Since the pointing range of the second type of sensor is 360°, it can collect all the sounds inside the vehicle, including: user voice information, environmental sounds, and the sounds played by the speaker. Figure 3 Figure 4 is a schematic diagram of the sound information collected by each sensor. As Figure 3 shown, both the first type of sensor E and the second type of sensor F will collect user voice point A, user voice point B, and other user voices, as well as environmental sounds and the sounds played by the speaker. And, Figure 3 in Figure 4, the distances between user voice point A and the first type of sensor E and the second type of sensor F are S1 and S2 respectively, Figure 3 and in Figure 4, the distances between user voice point B and the first type of sensor E and the second type of sensor F are S3 and S4 respectively.
[0071] In a possible implementation manner, before performing the above step 101, the following steps A1 - A2 may further be included:
[0072] Step A1, obtain the window state information and door state information of the vehicle.
[0073] Step A2, adjust the sensitivity of each first type of sensor inside the vehicle based on the window state information and door state information.
[0074] Specifically, the opening states of the windows and doors in the vehicle can be obtained through the vehicle body bus. If it is determined according to the obtained window state information and door state information that the left front door and / or the left front window of the vehicle are in the open state, the sensitivity of the first type of sensor corresponding to the seat close to the driver can be increased; as Figure 2 shown, seat 1 is the driver's seat. If it is determined that the left front door and / or the left front window of the vehicle are in the open state, the sensitivity of the first type of sensor MIC1 can be increased. If it is determined that the right front door and / or the right front window of the vehicle are in the open state, the sensitivity of the first type of sensor corresponding to the seat close to the co - driver can be increased. If it is determined that the left rear door and / or the left rear window of the vehicle are in the open state, the sensitivity of the first type of sensor corresponding to the seat close to the rear of the driver can be increased. If it is determined that the right rear door and / or the right rear window of the vehicle are in the open state, the sensitivity of the first type of sensor corresponding to the seat close to the rear of the co - driver can be increased.
[0075] Specifically, adjusting the sensitivities of various first-type sensors in the vehicle based on the window state information and the door state information may include the following steps B1 - B3:
[0076] Step B1: For each first-type sensor, calculate the product of the preset door influence coefficient of the door closest to the seat corresponding to the first-type sensor and the opening / closing degree of the door as the first product.
[0077] In the embodiments of the present invention, the preset door influence coefficient of the door can be set for a specific vehicle and is not specifically limited herein. The value range of the opening / closing degree of the door can be 0 - 100%. For example, when the door is in the unopened state, its opening / closing degree is 0, and when the door is in the fully opened state, its opening / closing degree is 100%.
[0078] Step B2: Calculate the product of the preset window influence coefficient of the window closest to the seat corresponding to the first-type sensor and the opening / closing degree of the window as the second product.
[0079] In the embodiments of the present invention, the preset door influence coefficient of the window can be set for a specific vehicle and is not specifically limited herein. The value range of the opening / closing degree of the window can be 0 - 100%. For example, when the window is in the unopened state, its opening / closing degree is 0, and when the window is in the fully opened state, its opening / closing degree is 100%.
[0080] Among them, the order of execution of the step of calculating the product of the preset door influence coefficient of the door closest to the seat corresponding to the first-type sensor and the opening / closing degree of the door as the first product in step B1 and step B2 is not limited.
[0081] Step B3: Calculate the sum of the first product, the second product and the theoretical sensitivity of the first-type sensor as the adjusted sensitivity of the first-type sensor.
[0082] In the embodiments of the present invention, each first-type sensor can set a theoretical sensitivity, which is: the best sensitivity for the first-type sensor to collect sound information when the door and window closest to the seat corresponding to the first-type sensor are both in the closed state.
[0083] Specifically, in the embodiments of the present invention, for each first-type sensor, the following formula can be used to adjust the sensitivity of the first-type sensor:
[0084] S = S0 + K1 * X1 + K2 * X2
[0085] Wherein, S is the adjusted sensitivity of the first type of sensor, and S0 is the theoretical sensitivity of the first type of sensor; KI is the preset door influence coefficient corresponding to the door closest to the seat corresponding to the first type of sensor, and X1 is the opening / closing degree of the door closest to the seat corresponding to the first type of sensor; K2 is the preset window influence coefficient corresponding to the window closest to the seat corresponding to the first type of sensor, and X2 is the opening / closing degree of the window closest to the seat corresponding to the first type of sensor. Moreover, S0 is a preset value, which can be specifically set according to different vehicles and is not limited herein; X1 is the opening / closing degree of the door. When the door is fully open, X1 takes a value of 100%, and when the door is not opened, X1 takes a value of 0; X2 is the opening / closing degree of the window. When the window is fully open, X2 takes a value of 100%, and when the window is not opened, X2 takes a value of 0.
[0086] Figure 4 It is a circuit diagram corresponding to the offset voltage of the first type of sensor. As Figure 4 shown, the circuit principle is as follows: MICP2 and MICN2 are a pair of differential signals, which can be input to both ends of the first type of sensor MIC after being filtered by the capacitor C1. The two pins of the first type of sensor MIC are grounded and powered respectively, Figure 4 and the resistor R1 in it is related to the sensitivity of the first type of sensor MIC. The offset voltage of the first type of sensor MIC is V 总 for powering the first type of sensor MIC. The voltage corresponding to the resistor R3 is V3. If the first type of sensor MIC receives sound information, it will generate a voltage V4, and the voltage corresponding to R1 is V1. The capacitor C1 is equivalent to a short circuit in the DC circuit and plays a role in filtering low-frequency waves. The resistance corresponding to the resistor R2 is V2. The resistor R2 is a variable resistor, and the sensitivity of the first type of sensor MIC can be adjusted by adjusting the size of R2. The greater the sensitivity of the first type of sensor MIC, the more weakly the sound signal the first type of sensor MIC can collect. The specific method for adjusting the variable resistor R2 of the first type of sensor MIC is as follows:
[0087] V4 = V 总 - V1 - V3 - V2 = V 总 - V1 - V3 - I 总 * R2 = V 总 - V1 - V3 - ((V 总 - V4) / (R1 + R2 + R3))
[0088] * R2. Assume: VA = V 总 – V1 - V3; RA = R1 + R3; then V4 = (VARA – V 总 R2 + VAR2) / RA.
[0089] Assume B = VARA + VAR2; then (S0 + K1*X1 + K2*X2) / 20 = lg(B - VR2) –
[0090] g(RA*V0); S = S0 + K1*X1 + K2*X2 = 20Lg((B - V 总 R2) / (RA*V0));
[0091] Assume C = (S0 + K1*X1 + K2*X2) / 20 + lg(RA*V0); then R2 can be adjusted to R2 = V 总 /
[0092] (B - 10^C). Wherein, the values of V 总 , V2 and V3 are all known, and V0 usually takes the value of 1.
[0093] That is, for each first-type sensor, if the window and the door closest to the seat corresponding to the first-type sensor are not fully closed, the adjusted sensitivity S of the first-type sensor can be determined by using the formula S = S0 + K1*X1 + K2*X2 according to the opening and closing degree X1 of the door, the preset door influence coefficient K1 corresponding to the door, the opening and closing degree X2 of the window, the preset window influence coefficient K2 corresponding to the window, and the theoretical sensitivity S0. Specifically, by adjusting the size of the variable resistor R2 of the first-type sensor, R2 can be adjusted to "R2 = V 总 / (B - 10^C)" to make the first-type sensor reach the adjusted sensitivity S.
[0094] If the door and the window are already closed, the R2 of the first-type sensor can be restored to the resistance value corresponding to its theoretical sensitivity state, because if the sensitivity of the first-type sensor MIC is too high and the duration is too long, the heat of the first-type sensor MIC will be too high, which will cause great damage to the first-type sensor MIC and affect its performance of collecting sound information.
[0095] In a possible implementation manner, as Figure 2 shown, the first-type sensors MIC1, MIC2, MIC3, and MIC4 are respectively provided with corresponding collection areas (sector areas facing the vehicle seat), but the installation position deviation of the vehicle-mounted first-type sensors will cause deviation in the collection areas, and the jitter during driving will also cause deviation in the collection areas of the first-type sensors. Based on this, in the embodiments of the present invention, the first sound information collected can be sorted by comparing the closest distances from the sound sources of the collected first sound information to each first-type sensor. Specifically, Figure 2Taking the 4-seater vehicle shown as an example, the sets of the first sound information collected by the first type of sensors MIC1, MIC2, MIC3, and MIC4 are E1, E2, E3, and E4 respectively.
[0096] The sets E1 - E4 are all the collected data of each first type of sensor. The sets E1 - E4 include: a timestamp and the first sound information. The first sound information is a voice data segment of a certain frequency point, and the timestamp is the moment when the first type of sensor collects the first sound information. The sets D1 - D4 are the target first sound information sets corresponding to each first type of sensor after performing the above step 102.
[0097] Figure 5 As a schematic diagram of the distance from the sound information to each sensor, the following takes Figure 5 as an example to illustrate the step of, for each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, then deleting the first sound information in the set of the first sound information corresponding to the first type of sensor to obtain the target first sound information set corresponding to the first type of sensor as follows:
[0098] As Figure 5 shown, the distances from the first sound information A to each of the first type of sensors MIC1, MIC2, MIC3, MIC4, and to the second type of sensor MIC5 are x1, x2, x3, x4, and x5 respectively; the distances from the first sound information B to each of the first type of sensors MIC1, MIC2, MIC3, MIC4, and to the second type of sensor MIC5 are y1, y2, y3, y4, and y5 respectively.
[0099] For example, for the first sound information A, if each of the first type of sensors and the second type of sensor collects the first sound information A, then the magnitudes of x1, x2, x3, x4, x5 can be compared to determine the minimum value xA among them. If the sensor corresponding to the minimum value xA is the first type of sensor, then only save the first sound information A in the first type of sensor corresponding to the minimum distance xA; if the sensor corresponding to the minimum value xA is the second type of sensor, then the first sound information A may not be saved in any sensor.
[0100] Since the distance from the sound to the sensor is equal to the speed of sound * propagation time, it is possible to compare the timestamps of the sound information collected by each sensor to achieve the purpose of comparing the distances from the sound information to each sensor. The nearest timestamp when the first sound information A reaches the sensor is the nearest distance from the first sound information A to the sensor. Therefore, for each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, then delete the first sound information in the set of the first sound information corresponding to the first type of sensor to obtain the target first sound information set corresponding to the first type of sensor.
[0101] When the sound source position is in the middle area of the vehicle (i.e., the distances from the sound source position to two of the first type of sensors MIC1-4 are equal), the first sound information is usually background noise and can be directly deleted from each of the first type of sensors MIC1-4. For example, if the sound source position of the first sound information B in the sets D1 and D2 is the same distance from the first type of sensor MIC1 and the first type of sensor MIC2, then the first sound information B in the sets D1 and D2 can be directly deleted.
[0102] Figure 6 As a schematic diagram of the acquisition range of the sensor, the second type of sensor MIC5 collects the second sound information S1, S2, and S3 in an omnidirectional manner. Then, the set D[5] of the second sound information collected by the second type of sensor MIC5 includes S1, S2, and S3. As Figure 6 shown, the first type of sensor MIC1 collects the first sound information S1. Then, the set E[1] of the first sound information collected by the first type of sensor MIC1 includes: S1. The first type of sensor MIC2 collects the first sound information S2 and S2. Then, the set E[2] of the first sound information collected by the first type of sensor MIC2 includes: S1 and S2.
[0103] Sound information S1 has been collected by all of MIC1, MIC2, and MIC5. If the distances from the sound information S1 to MIC1 and MIC2 are the same, then the sound information S1 in the sets of the first sound information corresponding to MIC1 and MIC2 will be deleted, and the resulting target first sound information sets D[1] and D[2] corresponding to MIC1 and MIC2 do not include the sound information S1. If the distance from the sound information S1 to MIC1 is less than the distance to MIC2, then the sound information S1 in the set of the first sound information corresponding to MIC2 will be deleted, and the sound information S1 will be saved in the set of the first sound information corresponding to MIC1. The resulting target first sound information set D[1] corresponding to MIC1 includes the sound information S1, while the resulting target first sound information set D[2] corresponding to MIC2 does not include the sound information S1.
[0104] In an embodiment of the present invention, for each target first sound information set, a sound information intersection between the target first sound information set and a set of second sound information is determined. Since the sensitivity of the second type of sensor can collect all the sound information inside the vehicle, and the first type of sensor collects all the sound information within its directivity range. Therefore, the sound information collected by all the first type of sensors but not collected by the second type of sensor can be determined as the sound outside the vehicle, the sound of the window, the self-interference of circuit filtering, the electromagnetic interference, and the misjudged sound caused by the wind blowing the first type of sensor, that is, it can be determined as noise or interference sound. Therefore, by finding the sound information intersection between each target first sound information set and the set of second sound information, this type of noise and interference sound can be eliminated.
[0105] Figure 7 It is a schematic diagram of a sound information intersection. As Figure 7 shown, the intersection of the target first sound information set D[1] corresponding to the first type of sensor MIC1 and the set of second sound information D[5] corresponding to the second type of sensor MIC5 is F[1]. The sound information Z1 is in the set D[1] of MIC1 but not in the set D[5] of MIC5, and the sound information Z2 is in both the set D[1] of MIC1 and the set D[5] of MIC5. Then F[1] only retains the data Z2 that exists in both D[1] and D[5], and Z1 is not retained in F[1].
[0106] In a possible implementation manner, before determining the target voice according to the target sound information intersection, it further includes steps C1 - C2:
[0107] Step C1, based on the speaker identifier, eliminate the speaker sound information in each of the sound information intersections to obtain the corresponding eliminated sound information intersections;
[0108] Step C2: Determine the intersection of the filtered sound information corresponding to the first type of sensors corresponding to the seat where the user is located as the target sound information intersection.
[0109] For example, when the sound information Y of the speaker is being played, the played sound information Y can be transmitted to the processor of the in-vehicle electronic device through a connection line, and the speaker sound information Y in each intersection of sound information can be directly filtered out by the in-vehicle electronic device.
[0110] By using the method provided in the embodiment of the present invention, by deleting the first sound information in the set of the first sound information corresponding to the first type of sensors whose acquisition time is not earlier than the acquisition time by other first type of sensors, the noise and miscellaneous sounds collected by each first type of sensor are further deleted, and the intersection of sound information is calculated with the set of the second sound information, further eliminating the noise and miscellaneous sounds collected by the first type of sensors. Moreover, by arranging multiple sensors in the vehicle, the sensitivity of the corresponding influence area sensors is automatically adjusted according to the window and door states; by comparing with the second type of sensors, the direction of the sound source is judged, and according to the sound of the loop of the speaker corresponding to the acquisition area, the interference of the speaker is eliminated, and the determined target voice is sent to the target speaker corresponding to the user's voice interaction object for playback according to the user's key operation, reducing the noise interference during the voice interaction between the user and his / her voice interaction object and improving the quality of the voice interaction.
[0111] Based on the same inventive concept, according to the in-vehicle voice interaction method provided in the above embodiment of the present invention, correspondingly, another embodiment of the present invention further provides an in-vehicle voice interaction device, and its structural schematic diagram is as Figure 8 shown, specifically including:
[0112] An information receiving module 801, configured to obtain a set of first sound information collected by each first type of sensor and a set of second sound information collected by a second type of sensor when receiving a voice interaction instruction issued by a user; wherein, the voice interaction instruction includes: identification information of a target speaker corresponding to the user's voice interaction object;
[0113] A first information filtering module 802, configured to, for each first sound information collected by each first type of sensor, if the acquisition time of the first sound information by the first type of sensor is not earlier than the acquisition time by other first type of sensors, delete the first sound information in the set of the first sound information corresponding to the first type of sensor to obtain a target first sound information set corresponding to the first type of sensor;
[0114] An intersection determination module 803, configured to, for each target first sound information set, determine the intersection of sound information between the target first sound information set and the set of the second sound information;
[0115] A target voice determination module 804, configured to determine a target voice according to an intersection of target voice information, and play the target voice through the target speaker based on the identification information; wherein, the intersection of the target voice information is: an intersection of voice information corresponding to a first type of sensor corresponding to the seat where the user is located.
[0116] It can be seen that when using the device provided by the embodiment of the present invention, in the case of receiving a voice interaction instruction issued by a user, a set of first voice information collected by each first type of sensor and a set of second voice information collected by a second type of sensor are obtained; for each first voice information collected by each first type of sensor, if the time when the first voice information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, then the first voice information in the set of first voice information corresponding to the first type of sensor is deleted to obtain a target first voice information set corresponding to the first type of sensor; for each target first voice information set, an intersection of voice information between the target first voice information set and the set of second voice information is determined; a target voice is determined according to the intersection of the target voice information, and the target voice is played through the target speaker based on the identification information. By deleting the first voice information in the set of first voice information corresponding to the first type of sensor whose collection time is not earlier than the collection time of other first type of sensors, the noise collected by each first type of sensor is further deleted, and an intersection of voice information is obtained with the set of second voice information, further eliminating the noise collected by the first type of sensor, and then a target voice is determined according to the intersection of the target voice information and sent to the target speaker corresponding to the voice interaction object of the user for playing, reducing the noise interference during the voice interaction between the user and his voice interaction object, and improving the quality of the voice interaction.
[0117] Optionally, referring to Figure 9 the device further includes:
[0118] A sensitivity adjustment module 901, configured to obtain window state information and door state information of the vehicle; based on the window state information and the door state information, adjust the sensitivity of each first type of sensor in the vehicle.
[0119] Optionally, the sensitivity adjustment module 901 is specifically configured to, for each first-type sensor, calculate the product of the preset door influence coefficient of the door closest to the seat corresponding to the first-type sensor and the opening / closing degree of the door as the first product; calculate the product of the preset window influence coefficient of the window closest to the seat corresponding to the first-type sensor and the opening / closing degree of the window as the second product; calculate the sum of the first product, the second product and the theoretical sensitivity of the first-type sensor as the adjusted sensitivity of the first-type sensor.
[0120] Optionally, referring to Figure 9 the device further includes: a second information elimination module 902, configured to eliminate the speaker sound information in each of the sound information intersections based on the speaker identifier, to obtain the corresponding eliminated sound information intersections; and determine the eliminated sound information intersection corresponding to the first-type sensor corresponding to the seat where the user is located as the target sound information intersection.
[0121] Optionally, each of the first-type sensors is a non-omnidirectional sensor corresponding to a seat in the vehicle, and the second-type sensor is an omnidirectional sensor capable of collecting all the sound information in the vehicle.
[0122] It can be seen that by using the device provided in the embodiment of the present invention, by deleting the first sound information in the set of the first sound information corresponding to the first-type sensor whose acquisition time is not earlier than the acquisition time of other first-type sensors, the noise collected by each first-type sensor is further deleted, and the sound information intersection is obtained with the set of the second sound information, further eliminating the noise collected by the first-type sensor. Moreover, by arranging multiple sensors in the vehicle, the sensitivity of the sensors in the corresponding influence area is automatically adjusted according to the window and door states; by comparing with the second-type sensor, the direction of the sound source is judged, and according to the sound of the loop of the speaker corresponding to the acquisition area, the interference of the speaker is eliminated, and according to the user's key operation, the determined target voice is sent to the target speaker corresponding to the user's voice interaction object for playing, reducing the noise interference during the voice interaction between the user and his / her voice interaction object and improving the quality of the voice interaction.
[0123] In the embodiment of the present invention, after determining the target voice according to the target sound information intersection, the target voice can be played on the target speaker based on the identification information. And the playing volume of the target speaker also affects the voice interaction effect between the user and his / her voice interaction object. Usually, the volume of the speaker can be set to multiple levels such as 0-39, and the higher the level, the greater the volume of the speaker. For the human ear, the volume level of the speaker suitable for the human ear is usually 5-20. Therefore, in the embodiment of the present invention, the following limitations can be made on the volume of the target speaker for playing the target voice:
[0124] If the level of the set volume of the target speaker is higher than 20, the target speaker plays the target voice at a volume level of 20; if the level of the set volume of the target speaker is lower than 5, the target speaker plays the target voice at a volume level of 5; if the level of the set volume of the target speaker is not less than 5 and not greater than 20, the target speaker plays the target voice at its set volume level.
[0125] In the embodiments of the present invention, the playback volume of the target speaker can be adjusted to be more suitable for the human ear, so as to ensure that the user's ears will not be stimulated by the excessive speaker volume during the voice interaction with the voice interaction object, nor will they not hear the target voice due to the too low speaker volume, improving the effect of the voice interaction between the user and the voice interaction object.
[0126] The embodiments of the present invention also provide an electronic device, as Figure 10 shown, including a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete the communication with each other through the communication bus 1004,
[0127] The memory 1003 is used to store computer programs;
[0128] When the processor 1001 is used to execute the program stored in the memory 1003, the following steps are implemented:
[0129] In the case of receiving a voice interaction instruction issued by the user, obtain a set of first sound information collected by each first type of sensor and a set of second sound information collected by the second type of sensor; wherein, the voice interaction instruction includes: identification information of the target speaker corresponding to the voice interaction object of the user;
[0130] For each first sound information collected by each first type of sensor, if the time when the first sound information is collected by the first type of sensor is not earlier than the time when it is collected by other first type of sensors, delete the first sound information in the set of first sound information corresponding to the first type of sensor to obtain the target first sound information set corresponding to the first type of sensor;
[0131] For each target first sound information set, determine the intersection of the sound information between the target first sound information set and the set of second sound information;
[0132] Determine a target voice according to the intersection of target voice information, and play the target voice through the target speaker based on the identification information; wherein, the intersection of the target voice information is: the intersection of the voice information corresponding to the first type of sensor corresponding to the seat where the user is located.
[0133] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0134] The communication interface is used for communication between the above electronic device and other devices.
[0135] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0136] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0137] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above vehicle-mounted voice interaction methods are implemented.
[0138] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute any of the vehicle-mounted voice interaction methods in the above embodiments.
[0139] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0140] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.
[0141] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for devices, electronic devices, and storage media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.
Claims
1. A vehicle-mounted voice interaction method, characterized in that, Including: When receiving a voice interaction instruction issued by a user, obtaining a set of first sound information collected by each first-type sensor and a set of second sound information collected by a second-type sensor; wherein, the voice interaction instruction includes: identification information of a target speaker corresponding to the voice interaction object of the user; each of the first-type sensors is a non-omnidirectional sensor corresponding to a seat in the vehicle, and the second-type sensor is an omnidirectional sensor capable of collecting all sound information in the vehicle; For each first sound information collected by each first-type sensor, if the time when the first sound information is collected by the first-type sensor is not earlier than the time when it is collected by other first-type sensors, then delete the first sound information in the set of first sound information corresponding to the first-type sensor to obtain a target first sound information set corresponding to the first-type sensor; For each target first sound information set, determine the sound information intersection between the target first sound information set and the set of second sound information; Determine a target voice according to the target sound information intersection and play the target voice through the target speaker based on the identification information; wherein, the target sound information intersection is: the sound information intersection corresponding to the first-type sensor of the seat where the user is located.
2. The method according to claim 1, characterized in that, Before obtaining the set of first sound information collected by each first-type sensor and the set of second sound information collected by the second-type sensor, it further includes: Obtaining the window state information and door state information of the vehicle; Based on the window state information and the door state information, adjusting the sensitivity of each first-type sensor in the vehicle.
3. The method according to claim 2, characterized in that, The adjusting the sensitivity of each first-type sensor in the vehicle based on the window state information and the door state information includes: For each first-type sensor, calculate the product of a preset door influence coefficient of the door closest to the seat corresponding to the first-type sensor and the opening and closing degree of the door as a first product; Calculate the product of a preset window influence coefficient of the window closest to the seat corresponding to the first-type sensor and the opening and closing degree of the window as a second product; Calculate the sum of the first product, the second product and the theoretical sensitivity of the first-type sensor as the adjusted sensitivity of the first-type sensor.
4. The method according to claim 1, characterized in that, Before determining the target voice according to the target sound information intersection, it further includes: Based on the speaker identification, removing the speaker sound information in each of the sound information intersections to obtain corresponding removed sound information intersections; Determine the removed sound information intersection corresponding to the first-type sensor of the seat where the user is located as the target sound information intersection.
5. A vehicle-mounted voice interaction device, characterized in that, Including: An information receiving module, configured to obtain a set of first sound information collected by each first-type sensor and a set of second sound information collected by second-type sensors when receiving a voice interaction instruction issued by a user; wherein, the voice interaction instruction includes: identification information of a target speaker corresponding to the voice interaction object of the user; each of the first-type sensors is a non-omnidirectional sensor corresponding to a seat in the vehicle, and the second-type sensors are omnidirectional sensors capable of collecting all sound information in the vehicle; A first information elimination module, configured to, for each first sound information collected by each first-type sensor, if the time when the first sound information is collected by the first-type sensor is not earlier than the time when it is collected by other first-type sensors, delete the first sound information from the set of first sound information corresponding to the first-type sensor, to obtain a target first sound information set corresponding to the first-type sensor; An intersection determination module, configured to, for each target first sound information set, determine a sound information intersection between the target first sound information set and the set of second sound information; A target voice determination module, configured to determine a target voice according to the target sound information intersection and play the target voice through the target speaker based on the identification information; wherein, the target sound information intersection is: the sound information intersection corresponding to the first-type sensor corresponding to the seat where the user is located.
6. The device according to claim 5, characterized in that, Further comprising: A sensitivity adjustment module, configured to obtain window state information and door state information of the vehicle; Based on the window state information and the door state information, adjust the sensitivity of each first-type sensor in the vehicle.
7. The device according to claim 6, characterized in that, Specifically, the sensitivity adjustment module is configured to, for each first-type sensor, calculate the product of a preset door influence coefficient of the door closest to the seat corresponding to the first-type sensor and the opening and closing degree of the door as a first product; calculate the product of a preset window influence coefficient of the window closest to the seat corresponding to the first-type sensor and the opening and closing degree of the window as a second product; calculate the sum value of the first product, the second product and the theoretical sensitivity of the first-type sensor as the adjusted sensitivity of the first-type sensor.
8. The device according to claim 5, characterized in that, Further comprising: A second information elimination module, configured to eliminate speaker sound information from each of the sound information intersections based on the speaker identification, to obtain corresponding eliminated sound information intersections; Determine the eliminated sound information intersection corresponding to the first-type sensor corresponding to the seat where the user is located as the target sound information intersection.
9. An electronic device, characterized in that, Comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; The memory is used for storing a computer program; When the processor is configured to execute the program stored on the memory, implement the method steps described in any one of claims 1-4.
Citation Information
Patent Citations
Speech recognition device, speech recognition system, and speech recognition method
CN111556826A
Method and device for controlling a plurality loudspeakers to play audio and electronic device
CN111629301A