Method, device, processor and vehicle for processing voice information in vehicle
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2023-05-11
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本发明实施例提供了一种车辆中语音信息的处理方法、装置、处理器和车辆,以至少解决对语音信息处理的效果差的技术问题
[0016] In this embodiment of the invention, the noise signal collected from the vehicle, as well as the number and location information of the passengers, can be analyzed. The system determines whether the volume of the noise signal and the number of passengers meet the activation conditions for the sound acquisition device. If the activation conditions are met, the sound acquisition device at the location of the passengers can be activated. The activated sound acquisition device can then collect the original voice information of the corresponding passengers. The volume and/or clarity of the collected original voice information can be enhanced to obtain the target voice information. The system can also control the corresponding audio playback device to play the target voice. By automating the activation of the sound acquisition device and the enhancement of the original voice information, the system enables smooth communication between passengers in a noisy environment, thus solving the technical problem of poor voice information processing and improving the overall quality of voice information processing.
Smart Images

Figure CN116798443B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle-related technologies, and more specifically, to a method, apparatus, processor, and vehicle for processing voice information in a vehicle. Background Technology
[0002] Currently, excessive ambient noise during vehicle operation can hinder smooth communication between passengers. Current technologies only allow passengers to manually activate their microphones to amplify their voices, without addressing the noise signals from the surrounding environment. Therefore, the technology still suffers from poor voice information processing capabilities within the vehicle.
[0003] There is currently no effective solution to the technical problem of poor performance in voice information processing in the aforementioned related technologies. Summary of the Invention
[0004] This invention provides a method, apparatus, processor, and vehicle for processing voice information in a vehicle, in order to at least solve the technical problem of poor performance in voice information processing.
[0005] According to one aspect of the present invention, a method for processing voice information in a vehicle is provided. The method may include: acquiring volume information of noise signals in the vehicle, and the number of occupants in the vehicle and their location information within the vehicle; activating a sound acquisition device in response to the volume information and number satisfying the activation conditions for starting a sound acquisition device in the vehicle, wherein the sound acquisition device corresponds to the location information; enhancing the original voice information acquired by the activated sound acquisition device to obtain target voice information, wherein the volume information of the target voice information is higher than the volume information of the original voice information, and / or the clarity of the target voice information is higher than the clarity of the original voice information; and controlling a sound playback device to play the target voice information.
[0006] Optionally, in response to the volume information and quantity meeting the activation conditions of the sound acquisition device in the vehicle, the sound acquisition device is activated, including: in response to the volume information exceeding a first volume threshold and the quantity being greater than or equal to a quantity threshold, detecting the status data of the auxiliary voice system in the vehicle; and in response to the detected status data indicating that the auxiliary voice system is not activated, activating the sound acquisition device.
[0007] Optionally, acquiring the volume information of the noise signal in the vehicle includes: determining the noise signal based on the vibration amplitude of the electret film in the condenser electret microphone in the vehicle, wherein the vibration amplitude is used to characterize the change in capacitance of the condenser electret microphone; converting the detected noise signal into a corresponding electrical signal; and determining the volume information based on the electrical signal.
[0008] Optionally, obtaining the number of passengers and the location information of the passengers in the vehicle includes: controlling the seat occupancy sensor in the vehicle to identify the passengers and obtain the number and location information; and / or controlling the image acquisition device in the vehicle to acquire images of the passengers and obtain image information; and determining the number and location information based on the image information.
[0009] Optionally, the vehicle controls the seat occupancy sensor to identify the occupants and obtain quantity and location information, including: connecting the circuit where the seat occupancy sensor is located in response to the pressure of the pressure-type safety configuration of the seat occupancy sensor; and outputting the quantity and location information after the circuit is connected.
[0010] Optionally, determining the quantity and location information based on image information includes: detecting the image information to obtain detection results; matching the detection results with seats in the vehicle to obtain matching results; and determining the quantity and location information based on the matching results.
[0011] Optionally, the raw speech information acquired by the activated sound acquisition device is enhanced to obtain target speech information, including: performing echo cancellation and noise reduction on the raw speech information to obtain a processing result; determining the speech volume of the processing result based on the average amplitude level of the processing result within the target time range; and enhancing the processing result in response to the speech volume being less than a second volume threshold to obtain target speech information.
[0012] According to another aspect of the present invention, a device for processing voice information in a vehicle is also provided. The device may include: an acquisition unit for acquiring volume information of noise signals in the vehicle, and the number of occupants in the vehicle and their location information within the vehicle; an activation unit for activating a sound acquisition device in response to the volume information and number satisfying the activation conditions for activating a sound acquisition device in the vehicle, wherein the sound acquisition device corresponds to the location information; a processing unit for enhancing the original voice information acquired by the activated sound acquisition device to obtain target voice information, wherein the volume information of the target voice information is higher than the volume information of the original voice information, and / or the clarity of the target voice information is higher than the clarity of the original voice information; and a playback unit for controlling a sound playback device to play the target voice information.
[0013] According to another aspect of the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the vehicle voice information processing method of the present invention.
[0014] According to another aspect of the present invention, a processor is also provided. The processor is configured to run a program, wherein the program, when running, executes the method for processing voice information in a vehicle according to the embodiments of the present invention.
[0015] According to another aspect of the present invention, a vehicle is also provided. This vehicle is used to execute the vehicle voice information processing method of the present invention.
[0016] In this embodiment of the invention, the noise signal collected from the vehicle, as well as the number and location information of the passengers, can be analyzed. The system determines whether the volume of the noise signal and the number of passengers meet the activation conditions for the sound acquisition device. If the activation conditions are met, the sound acquisition device at the location of the passengers can be activated. The activated sound acquisition device can then collect the original voice information of the corresponding passengers. The volume and / or clarity of the collected original voice information can be enhanced to obtain the target voice information. The system can also control the corresponding audio playback device to play the target voice. By automating the activation of the sound acquisition device and the enhancement of the original voice information, the system enables smooth communication between passengers in a noisy environment, thus solving the technical problem of poor voice information processing and improving the overall quality of voice information processing. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0018] Figure 1 This is a flowchart of a method for processing voice information in a vehicle according to an embodiment of the present invention;
[0019] Figure 2 This is a flowchart of a method for activating a sound acquisition device according to an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of the microphone deployment location in a vehicle according to an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram illustrating the activation of a microphone corresponding to the location of a driver or passenger in a vehicle according to an embodiment of the present invention.
[0022] Figure 5 This is a flowchart of a method for enhancing speech information according to an embodiment of the present invention;
[0023] Figure 6 This is a schematic diagram of a speaker playing music corresponding to the position of a driver or passenger in a vehicle according to an embodiment of the present invention;
[0024] Figure 7 This is a flowchart of a vehicle voice information processing device according to an embodiment of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] Example 1
[0028] According to an embodiment of the present invention, a method for processing voice information in a vehicle is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0029] Figure 1 This is a flowchart illustrating the processing of voice information in a vehicle according to an embodiment of the present invention, such as... Figure 1 As shown, the method may include the following steps:
[0030] Step S102: Obtain the volume information of the noise signal in the vehicle, as well as the number of passengers in the vehicle and the location information of the passengers in the vehicle.
[0031] In the technical solution provided in step S102 of the present invention, the volume information of the noise signal in the vehicle, the number of occupants in the vehicle, and the position information of the occupants in the vehicle can be collected in real time. The noise signal can be the noise inside the vehicle's cabin. The volume information can be used to represent the decibel level of the noise signal. The occupants can be the drivers and passengers in the vehicle. The position information can be used to represent the positions of the seats in the vehicle, which may include the driver's seat, the front passenger seat, the seat behind the driver, and the seat behind the front passenger. This is merely an example and does not impose specific limitations on the number and names of the seat positions.
[0032] Optionally, a noise sensor module can be deployed in the vehicle's cabin. Once activated, the noise sensor can detect noise signals inside the cabin in real time and determine the volume of the noise signal.
[0033] Optionally, a occupant quantity determination module can be deployed in the vehicle, or a occupancy sensor module can be deployed in each seat of the vehicle, with the occupancy sensor module connected to the occupant quantity determination module. Once activated, the occupancy sensor module can collect the quantity and location information of the passengers and transmit this information to the occupant quantity determination module. The occupant quantity determination module then uses this information to determine whether the activation conditions for the sound acquisition device are met.
[0034] Step S104: In response to the volume information and quantity satisfying the activation conditions of the sound acquisition device in the starting vehicle, the sound acquisition device is activated, wherein the sound acquisition device corresponds to the location information.
[0035] In the technical solution provided by step S104 of the present invention, the volume and quantity of the collected noise information can be judged to determine whether the volume and quantity meet the activation conditions of the sound acquisition device. If the volume and quantity meet the activation conditions, the sound acquisition device corresponding to the location of the driver / passenger can be activated, wherein the sound acquisition device corresponds to the location information. The sound acquisition device can be a microphone.
[0036] Optionally, after acquiring the noise signal and its volume information via the noise sensor module, it can be determined whether the volume information meets the volume activation conditions in the sound acquisition device activation conditions. If the volume information does not meet the volume activation conditions, it is not necessary to send a signal to the personnel judgment module. That is, the personnel judgment module and the occupancy sensor do not need to collect the number and location information of the driving and riding objects until the volume information meets the volume activation conditions before sending a signal. If the volume information meets the volume activation conditions, a signal can be sent to the personnel quantity judgment module via the noise sensor module. After receiving the signal, the personnel quantity judgment module can control the occupancy sensor to detect the location and number of driving and riding objects, and send the collected location information and quantity to the personnel quantity judgment module. The personnel quantity judgment module can then determine whether the quantity meets the quantity activation conditions for activating the sound acquisition device. If it does, the location information can be sent to the control unit in the vehicle, which can then control the sound acquisition device corresponding to the location information to activate the corresponding sound acquisition device.
[0037] For example, it can be determined whether the volume of the noise signal is greater than 55 decibels. If it is, the volume information meets the volume activation condition, thus allowing the collection of the number and location information of the occupants. It can also be determined whether the number of occupants is greater than or equal to two. If the number is greater than or equal to two, the number meets the quantity activation condition, thus activating the sound acquisition device corresponding to the location information of the occupants. It should be noted that this is merely an example and does not impose specific limitations on the process and method of determining whether a sound acquisition device should be activated based on volume information and quantity. Any process or method that activates a sound acquisition device based on quantity information and quantity should be within the protection scope of this invention.
[0038] Currently, voice enhancement can only be achieved manually by the driver or passenger when they need it, by activating the microphone at their seat. However, this method doesn't automatically activate the sound acquisition device, leaving a technical problem. In this invention, however, the volume information of the noise signal can be collected in real-time to determine if it meets the volume activation conditions for the sound acquisition device. If the volume activation conditions are met, the number of passengers and their positions in the vehicle can be further collected to determine if the number meets the quantity activation conditions for the sound acquisition device. If both activation conditions are met, the sound acquisition device corresponding to the position information can be controlled to activate it. By automatically collecting the required data and determining whether the activation conditions are met using a noise sensor module and a personnel quantity judgment module, the technical effect of automatically activating the sound acquisition device is achieved.
[0039] For example, the vehicle's sound zones can be divided based on location information. Four location information points could divide the sound zones into four: the driver's seat zone, the passenger seat zone, the zone behind the driver's seat, and the zone behind the passenger seat. A corresponding sound acquisition device can be placed in each zone. These devices can be deployed horizontally and symmetrically in positions such as the front dome lights or the center console. The spacing between adjacent sound acquisition devices can be consistent, for example, between 8cm and 20cm. The sound holes of the sound acquisition devices can be selected from arrayed multi-hole or grille-type holes. It should be noted that this is only an example and does not impose specific restrictions on the number of sound acquisition devices, their deployment locations, the number of sound holes, or the spacing between adjacent sound acquisition devices.
[0040] Step S106: Enhance the original speech information acquired by the activated sound acquisition device to obtain target speech information, wherein the volume of the target speech information is higher than that of the original speech information, and / or the clarity of the target speech information is higher than that of the original speech information.
[0041] In the technical solution provided by step S106 of the present invention, the activated sound acquisition device can enhance the volume and / or clarity of the original speech information of the corresponding driver / passenger object to obtain enhanced target speech information. The volume of the target speech information is higher than that of the original speech information, and / or the clarity of the target speech information is higher than that of the original speech information. The enhancement processing may include echo cancellation, noise reduction, and volume enhancement.
[0042] Optionally, after the personnel quantity judgment module determines that the quantity meets the activation conditions for activating the sound acquisition device, the control unit in the vehicle can send a signal to the sound acquisition device at the corresponding location to activate the sound acquisition device. The activated sound acquisition device can then collect the raw voice information of the driver and passengers at the corresponding location in real time. The real-time raw voice information can be converted into electrical signals and sent to the control unit.
[0043] Optionally, after receiving the electrical signal, the control unit can process the original speech information through the vehicle's speech processing module for echo cancellation and noise reduction (ECNR) to obtain processed speech information. It can determine whether the volume of the processed speech information needs to be increased. If so, volume enhancement processing can be performed to obtain the enhanced target speech information. If not, the speech information after echo cancellation and noise reduction processing can be directly used as the final target speech information.
[0044] In this embodiment of the invention, the need for volume enhancement processing can be determined based on the volume level of the original speech information. If the volume of the original speech information is high, no volume enhancement processing is required; only echo cancellation and noise reduction processing can be performed to obtain the final target speech information. Conversely, if the volume is low, echo cancellation and noise reduction processing are performed, and the volume of the original speech information is adjusted to a suitable level to obtain the final target speech information. Because targeted processing can be applied based on the actual situation of the speech, the technical problem of low accuracy in speech processing is solved.
[0045] Step S108: Control the sound playback device to play the target voice information.
[0046] In the technical solution provided by step S108 of the present invention, after the target voice information is obtained through enhancement processing, the sound playback device can be controlled to play the target voice information. The sound playback device can be a speaker and can correspond to the location information.
[0047] Optionally, after obtaining the target voice information, the control unit can synchronously transmit the target voice information to a sound playback device corresponding to the location information of other locations besides the location where the target voice information is located, and convert the electrical signal of the target voice into a sound signal for playback.
[0048] For example, if it is determined that there are two people in the vehicle, their positions can be the driver's seat and the passenger's seat. If it is determined that the sound acquisition devices can be activated, then the sound acquisition devices corresponding to the driver's seat and the passenger's seat can be started. The sound acquisition device at the driver's seat can collect the driver's original voice information, and the sound acquisition device at the passenger's seat can collect the passenger's original voice information.
[0049] In the embodiments of the present invention, steps S102 to S108 above can be used to collect noise signals in the vehicle and information on the number and location of the occupants. By determining whether the volume of the noise signal and the number of occupants meet the activation conditions for the sound acquisition device, if the conditions are met, the sound acquisition device at the location of the occupants can be activated. The activated sound acquisition device can then collect the original voice information of the corresponding occupants. The volume and / or clarity of the collected original voice information can be enhanced to obtain target voice information, and the corresponding audio playback device can be controlled to play the target voice. Because the activation of the sound acquisition device and the enhancement of the original voice information are automated, the goal of enabling smooth communication between occupants in a noisy environment is achieved, thus solving the technical problem of poor voice information processing and improving the overall quality of voice information processing.
[0050] The method described in this embodiment will be further described below.
[0051] As an optional embodiment, step S104, in response to the volume information and quantity meeting the activation conditions of the sound acquisition device in the vehicle, activates the sound acquisition device, including: in response to the volume information exceeding a first volume threshold and the quantity being greater than or equal to a quantity threshold, detecting the status data of the auxiliary voice system in the vehicle; in response to the detected status data indicating that the auxiliary voice system is not activated, activating the sound acquisition device.
[0052] In this embodiment, the relationship between the volume information of the noise signal and a first volume threshold, and the relationship between the number of occupants and a quantity threshold, can be determined. Based on these two relationships and the status data of the auxiliary voice system, it can be determined whether to activate the sound acquisition device at the location of the occupants. If the volume information of the noise signal exceeds the first volume threshold and the number of occupants is greater than or equal to the quantity threshold, the status data of the auxiliary voice system can be determined. Conversely, except for the cases where the volume information is greater than the first volume threshold and the number is greater than or equal to the quantity threshold, in other cases, it is not necessary to detect the status data of the auxiliary voice system or to activate the sound acquisition device. If it is determined that the status data of the auxiliary voice system indicates that the auxiliary voice system is not activated, the sound acquisition device can be activated. Conversely, if it is determined that the status data indicates that the auxiliary voice system is activated, it is not necessary to activate the sound acquisition device. The first volume threshold and the quantity threshold can be preset values or values set according to the actual noise environment. It should be noted that the setting method of the first volume threshold and the quantity threshold here is only for illustrative purposes and is not a specific limitation. The auxiliary voice system can be the vehicle's main voice assistant.
[0053] For example, a noise sensor module can compare the detected noise signal's volume information of 60 dB with a first volume threshold (e.g., 55 dB) to determine if the volume exceeds the first volume threshold. A personnel quantity judgment module can determine if the current number of passengers (three) is greater than or equal to a quantity threshold (e.g., two). If the comparison confirms the number exceeds the threshold, a signal can be sent to the control unit to obtain information on whether the host voice assistant is activated. If the host voice assistant is activated, it can send a suppression signal to the control unit to suppress the activation of the sound acquisition device. If the host voice assistant is not activated, it can send an activation signal to the control unit to activate the sound acquisition device. In other words, if the control unit does not receive a suppression signal indicating the voice assistant is activated, it can normally send signals to the sound acquisition device to control its activation.
[0054] As an optional embodiment, step S102, obtaining the volume information of the noise signal in the vehicle, includes: determining the noise signal based on the vibration amplitude of the electret film in the condenser electret microphone in the vehicle, wherein the vibration amplitude is used to characterize the change in capacitance of the condenser electret microphone; converting the detected noise signal into a corresponding electrical signal; and determining the volume information based on the electrical signal.
[0055] In this embodiment, the noise signal can be determined based on the vibration amplitude of the electret film in the condenser electret microphone in the vehicle, and the detected noise signal can be converted into a corresponding electrical signal. The volume information of the noise signal can be determined through the electrical signal. The vibration amplitude can be used to characterize the change in capacitance of the condenser electret microphone.
[0056] Optionally, the noise sensor in the vehicle uses a built-in electret condenser microphone to identify and collect noise signals. Since the sound waves of the noise signal cause the electret diaphragm inside the microphone to vibrate, the magnitude of the vibration amplitude determines the change in capacitance caused by the noise signal. This change in capacitance generates a tiny voltage corresponding to the capacitance value, thus converting the sound signal of the noise signal into an electrical signal. The magnitude of this electrical signal determines the volume of the noise signal.
[0057] In this embodiment of the invention, a sound-sensitive electret microphone can be deployed inside the noise sensor. The vibrations generated by the sound waves of the noise signal on the electret diaphragm inside the microphone allow for the sensitive acquisition of the noise signal. The magnitude of the vibration amplitude determines the change in capacitance corresponding to the noise signal. This capacitance change generates a minute voltage, converting the noise signal into a corresponding electrical signal. The magnitude of this electrical signal determines the volume of the noise signal. By considering the sensitivity of the electret microphone to noise signals, the accuracy of noise signal acquisition is improved.
[0058] As an optional embodiment, step S102, obtaining the number of passengers in the vehicle and the location information of the passengers in the vehicle, includes: controlling the seat occupancy sensor in the vehicle to identify the passengers and obtain the number and location information; and / or controlling the image acquisition device in the vehicle to acquire images of the passengers and obtain image information of the images; and determining the number and location information based on the image information.
[0059] In this embodiment, seat occupancy sensors in the vehicle can be controlled to identify occupants and determine their number and location. Alternatively, images of occupants can be acquired using image acquisition devices in the vehicle, obtaining image information. Based on this image information, the number and location of occupants can be determined. The seat occupancy sensors are used to determine the location of occupants, and the number of seat occupants corresponds to the number of seats in the vehicle. The image acquisition device can be a camera deployed inside the vehicle. The image information can be used to display the location of occupants. The seat occupancy sensors can also be referred to as occupancy sensors.
[0060] In this embodiment of the invention, a seat occupancy sensor deployed in each seat can detect whether an object or a passenger is present in that seat, further determining the number of such objects and their location within the seat. However, since objects may occupy seats in this situation, the determination of the number and location of passengers may be flawed. Therefore, an image acquisition device can be used to acquire images of the vehicle's cabin interior and identify the number and location of passengers, supplementing the results collected by the seat occupancy sensor. Alternatively, the number and location of passengers can be determined solely by the image acquisition device, thus achieving the goal of identifying passengers using both seat occupancy sensors and / or image acquisition devices, thereby improving the accuracy of passenger identification.
[0061] As an optional embodiment, step S102, controlling the seat occupancy sensor in the vehicle to identify the occupants and obtain quantity and location information, includes: in response to the pressure of the pressure-type safety configuration of the seat occupancy sensor being squeezed, connecting the circuit where the seat occupancy sensor is located; and after the circuit is connected, outputting the quantity and location information.
[0062] In this embodiment, it is possible to detect whether the pressure-type safety configuration in the seat occupancy sensor is being squeezed. If it is squeezed, the circuit containing the seat occupancy sensor can be connected, and the circuit after connection can output quantity and position information.
[0063] Optionally, the seat occupancy sensor is a magnetically operated button-type sensor with a simple clamp-on mounting configuration. The presence of an object or passenger on the corresponding seat can be determined by whether the clamp-on mounting configuration is compressed. When the clamp-on mounting configuration is compressed, the circuit of the seat occupancy sensor can be connected to activate the Hall effect sensor, transmitting the quantity and position information along the signal.
[0064] As an optional embodiment, step S102, determining the quantity and location information based on image information, includes: detecting the image information to obtain a detection result; matching the detection result with the seats in the vehicle to obtain a matching result; and determining the quantity and location information based on the matching result.
[0065] In this embodiment, image information can be detected to obtain detection results. By matching the detection results with the seats in the vehicle, a matching result can be determined. Based on the matching result, the quantity and location information corresponding to the detection result can be determined.
[0066] Optionally, the vehicle occupant quantity determination module can also be implemented by a visual recognition module. After the occupant quantity determination module is activated, it can control the activation of the visual recognition module and the image acquisition device. The image acquisition device can acquire images of the occupants inside the vehicle's cabin and transmit these images to the visual recognition module. The visual recognition module can analyze the acquired images in real time, preprocess the images using detection algorithms stored in the module, extract and classify features from the image information, obtain detection results, and match and identify the seats in the vehicle to determine the number and location of the occupants inside the vehicle.
[0067] For example, image acquisition devices can be deployed in locations such as the dashboard or rearview mirror inside a vehicle. It should be noted that this is only an example and no specific restrictions are placed on the deployment location of the image acquisition devices.
[0068] As an optional embodiment, step S106 involves enhancing the original speech information acquired by the activated sound acquisition device to obtain target speech information, including: performing echo cancellation and noise reduction processing on the original speech information to obtain a processing result; determining the speech volume of the processing result based on the average amplitude level of the processing result within the target time range; and enhancing the processing result in response to the speech volume being less than a second volume threshold to obtain the target speech information.
[0069] In this embodiment, echo cancellation and noise reduction processing can be performed on the original speech information to obtain the processing result. The speech volume of the vehicle result can be determined by the average amplitude level of the processing result within a target time range. The speech volume can be compared with a second volume threshold. If the speech volume is less than the second volume threshold, the processing result can be enhanced to determine the target speech information. Conversely, if the speech volume is greater than or equal to the second volume threshold, no enhancement processing is needed, and the processing result can be directly determined as the target speech information. The second threshold can be a preset value or a value set according to the actual noise environment or speech volume. The target time range can be a preset time or a time set according to the actual situation. It should be noted that this is only an illustrative example and does not impose specific limitations on the values of the second threshold and the target time range.
[0070] Optionally, after the vehicle's control unit receives the electrical signal of the original voice information, it can synchronously perform echo cancellation and noise reduction processing on the original voice information through the voice processing module to suppress environmental noise, ensure voice quality, and obtain the processing result. The processing result can be used to determine the relationship between the voice volume of each driver and passenger and the second volume threshold based on the average amplitude level of a target time range (e.g., three seconds). If the voice volume is less than the second volume threshold, the processing result can be enhanced to a specified preset level through enhancement processing.
[0071] For example, based on a preset three-second average amplitude level, the volume of the processed result is determined to be 45 dB. By comparing it with a preset second volume threshold of 65 dB, it is found that the volume of the voice is less than the second volume threshold. Therefore, the processed result can be enhanced to increase the volume of the voice to 65 dB.
[0072] In this embodiment of the invention, the noise signal collected from the vehicle, along with the number and location information of the occupants, can be analyzed. The system determines whether the volume of the noise signal and the number of occupants meet the activation conditions for the sound acquisition device. If the conditions are met, the sound acquisition device at the location of the occupants can be activated. The activated sound acquisition device can then collect the original speech information of the corresponding occupants. The volume and / or clarity of the collected original speech information can be enhanced to obtain target speech information. The system can then control the corresponding audio playback device to play the target speech. By automating the activation of the sound acquisition device and the enhancement of the original speech information, the system enables smooth communication between occupants in noisy environments, thus solving the technical problem of poor speech information processing and improving the overall effectiveness of speech information processing.
[0073] Example 2
[0074] The technical solutions of the embodiments of the present invention will be illustrated below with reference to preferred embodiments.
[0075] Currently, excessive ambient noise during vehicle operation can hinder smooth communication between occupants. Current technologies only allow occupants to manually activate their microphones to amplify their voices; they do not process the noise signals from the ambient environment. Therefore, the technical problem of passengers being unable to communicate effectively in noisy environments persists, resulting in a poor driving experience.
[0076] In one related technology, a vehicle-mounted occupant dialogue enhancer is proposed. The vehicle-mounted occupant dialogue enhancer includes: at least one sound acquisition device for acquiring initial audio information from one or more occupants; a noise reduction module for identifying noise information in the initial audio information and performing noise reduction processing to obtain the actual dialogue audio of one or more occupants; and a control module for controlling the vehicle to reduce the volume of the currently playing audio to a first preset volume or stop playing the currently playing audio and play the actual dialogue audio at a second preset volume, wherein the second preset volume is greater than the first preset volume. This solves the problem in related technologies where, because the voice acquisition microphone is located in the front of the vehicle, if either the voice enhancement function or the voice interaction function occupies the voice acquisition channel, the other function may not function properly. However, the above method does not consider automating the processes of activating the sound acquisition device and enhancing the original voice information, and therefore still suffers from the technical problem that occupants cannot communicate smoothly in noisy environments, resulting in a poor driving experience.
[0077] However, this invention proposes a method for enhancing the voice of conversations between drivers and passengers. This method may include: collecting noise signals from the vehicle and determining whether the volume of the noise signal and the number of drivers and passengers meet the activation conditions for a sound acquisition device. If the activation conditions are met, the sound acquisition device at the location of the drivers and passengers can be activated. The activated sound acquisition device can then collect the original voice information of the corresponding drivers and passengers. The volume and / or clarity of the collected original voice information can be enhanced to obtain target voice information. The corresponding audio playback device can be controlled to play the target voice. By automating the activation of the sound acquisition device and the enhancement of the original voice information, the method achieves the goal of enabling smooth communication between drivers and passengers in noisy environments, thus solving the technical problem of poor voice information processing performance and improving the overall quality of voice information processing.
[0078] The method described in this embodiment will be further described below.
[0079] Figure 2 This is a flowchart of a method for activating a sound acquisition device according to an embodiment of the present invention, such as... Figure 2 As shown, the method may include the following steps:
[0080] Step S201: Obtain the volume information of the noise signal through the noise sensor module.
[0081] In the technical solution provided by step S201 of the present invention, a noise sensor module can be deployed in the vehicle cabin. After the noise sensor is activated, it can detect the noise signal inside the cabin in real time and determine the volume information of the noise signal.
[0082] Optionally, the noise signal can be determined based on the vibration amplitude of the electret film in the condenser electret microphone in the vehicle, and the detected noise signal can be converted into a corresponding electrical signal. The volume information of the noise signal can be determined through the electrical signal.
[0083] In this embodiment of the invention, a sound-sensitive electret microphone can be deployed inside the noise sensor. The vibrations generated by the sound waves of the noise signal on the electret diaphragm inside the microphone allow for the sensitive acquisition of the noise signal. The magnitude of the vibration amplitude determines the change in capacitance corresponding to the noise signal. This capacitance change generates a minute voltage, converting the noise signal into a corresponding electrical signal. The magnitude of this electrical signal determines the volume of the noise signal. By considering the sensitivity of the electret microphone to noise signals, the accuracy of noise signal acquisition is improved.
[0084] Step S202: Determine whether the volume information exceeds the first volume threshold.
[0085] In the technical solution provided by step S202 in the above embodiment of the present invention, the relationship between the volume information and the first volume threshold can be determined. If the volume information exceeds the first volume threshold, step S203 can be executed; otherwise, step S201 can be executed.
[0086] Step S203: Obtain the number of drivers and passengers through the personnel quantity judgment module.
[0087] In the technical solution provided in step S203 of this embodiment of the invention, a personnel quantity judgment module can be deployed in the vehicle, or a occupancy sensor module can be deployed in each seat of the vehicle, and the occupancy sensor module can be connected to the personnel quantity judgment module. After activation, the occupancy sensor module can collect the quantity and location information of the passengers and transmit this information to the personnel quantity judgment module. The personnel quantity judgment module then judges the quantity and location information to determine whether the activation conditions for activating the sound acquisition device are met.
[0088] Optionally, the vehicle's seat occupancy sensors can be controlled to identify the number of occupants and their location. Alternatively, images of the occupants can be acquired using the vehicle's image acquisition equipment to obtain image information, and the number and location of the occupants can be determined based on this image information.
[0089] In this embodiment of the invention, a seat occupancy sensor deployed in each seat can detect whether an object or passenger is present in that seat, further determining the number of objects or passengers and their location within the seat. However, since objects may occupy seats in this situation, the determination of the number and location of passengers may be flawed. Therefore, an image acquisition device can be used to acquire images of the vehicle's cabin interior and identify the number and location of passengers, supplementing the results collected by the seat occupancy sensor. Alternatively, the number and location of passengers can be determined solely by the image acquisition device, thus achieving the goal of identifying passengers through seat occupancy sensors and / or image acquisition devices, thereby improving the accuracy of passenger identification.
[0090] Step S204: Determine whether the number of passengers is greater than or equal to the number threshold.
[0091] In the technical solution provided by step S204 of the present invention, the relationship between the number of passengers and a quantity threshold can be determined. If the number is greater than or equal to the quantity threshold, step S205 can be executed; otherwise, step S201 can be executed.
[0092] Step S205: Detect whether the auxiliary voice system is activated.
[0093] In the technical solution provided by step S205 of the present invention, the status data of the auxiliary voice system can be determined.
[0094] Optionally, if the status data of the auxiliary voice system indicates that the auxiliary voice system is not activated, the sound acquisition device can be activated. Conversely, if the status data indicates that the auxiliary voice system is activated, the sound acquisition device does not need to be activated.
[0095] Step S206: Activate the sound acquisition device.
[0096] In the technical solution provided by step S206 of the present invention, the sound acquisition device corresponding to the location information can be activated.
[0097] Optionally, the vehicle's audio zones can be divided based on location information. For example, four location information points can divide the audio zones into four: the driver's seat audio zone, the passenger's seat audio zone, the driver's rear seat audio zone, and the passenger's rear seat audio zone. Figure 3 This is a schematic diagram of the microphone deployment location in a vehicle according to an embodiment of the present invention, such as... Figure 3 As shown, each sound zone can correspond to one seat, and a corresponding microphone can be placed near each seat. The microphones can be deployed horizontally and symmetrically in positions such as the front dome lights or center console. The spacing between adjacent microphones can be consistent, for example, between 8cm and 20cm. The microphone apertures can be arrayed or grille-style. It should be noted that this is only an example and does not impose specific limitations on the number of microphones, their deployment locations, apertures, or the spacing between adjacent microphones.
[0098] Optionally, Figure 4 This is a schematic diagram illustrating the activation of a microphone corresponding to the location of a driver or passenger in a vehicle according to an embodiment of the present invention. Figure 4 As shown, microphones corresponding to seats with occupants can be activated. An ellipse can be used to represent seats in the vehicle's cabin, a hollow circle can be used to represent an inactive microphone, a black circle can be used to represent an activated microphone, and an inverted triangle can be used to represent occupants on the seats.
[0099] Figure 5 This is a flowchart of a method for enhancing speech information according to an embodiment of the present invention, such as... Figure 5 As shown, the method may include the following steps:
[0100] Step S501: Collect the original voice information of the corresponding location information using a sound acquisition device.
[0101] In the technical solution provided by step S501 of the present invention, after the sound acquisition device is activated, the original voice information of the location of the driver or passenger can be acquired through the sound acquisition device.
[0102] Optionally, after the personnel quantity judgment module determines that the quantity meets the activation conditions for activating the sound acquisition device, the control unit in the vehicle can send a signal to the sound acquisition device at the corresponding location to activate the sound acquisition device. The activated sound acquisition device can then collect the raw voice information of the driver and passengers at the corresponding location in real time. The real-time raw voice information can be converted into electrical signals and sent to the control unit.
[0103] In step S502, the original voice information is processed by the control unit to perform echo cancellation and noise reduction to obtain the processing result.
[0104] In the technical solution provided by step S502 of the present invention, the control unit can perform echo cancellation and noise reduction processing on the original voice information to obtain the processing result.
[0105] Optionally, after receiving the electrical signal, the control unit can process the original voice information through the voice processing module in the vehicle used for echo cancellation and noise reduction to obtain processed voice information. It can determine whether the volume of the processed voice information needs to be increased. If so, volume enhancement processing can be performed to obtain the target voice information with increased volume. If not, the voice information after echo cancellation and noise reduction processing can be directly used as the final target voice information.
[0106] In this embodiment of the invention, the need for volume enhancement processing can be determined based on the volume of the original speech information. If the volume of the original speech information is high, no volume enhancement processing is needed; echo cancellation and noise reduction processing can be performed only to obtain the final target speech information. Conversely, if the volume is low, echo cancellation and noise reduction processing are performed, and the volume of the original speech information is adjusted to a suitable level to obtain the final target speech information. Because targeted processing can be performed based on the actual situation of the speech, the technical problem of low accuracy in speech processing is solved.
[0107] Step S503: Determine whether the volume information of the processing result needs to be enhanced.
[0108] In the technical solution provided by step S503 in the above embodiment of the present invention, it can be determined whether the volume information of the processing result needs to be enhanced. If it needs to be enhanced, step S504 can be executed; otherwise, step S505 can be executed.
[0109] Optionally, after the vehicle's control unit receives the electrical signal of the original voice information, it can synchronously perform echo cancellation and noise reduction processing on the original voice information through the voice processing module to suppress environmental noise, ensure voice quality, and obtain the processing result. The processing result can be used to determine the relationship between the voice volume of each driver and passenger and the second volume threshold based on the average amplitude level of a target time range (e.g., within three seconds). If the voice volume is less than the second volume threshold, the processing result can be enhanced to a specified preset level through enhancement processing.
[0110] Step S504: Enhance the processing result.
[0111] In the technical solution provided by step S504 of the present invention, the processing result that needs to be enhanced can be enhanced.
[0112] For example, based on a preset three-second average amplitude level, the volume of the processed result is determined to be 45 dB. By comparing it with a preset second volume threshold of 65 dB, it is found that the volume of the voice is less than the second volume threshold. Therefore, the processed result can be enhanced to increase the volume of the voice to 65 dB.
[0113] Step S505: Play the voice through a speaker.
[0114] In the technical solution provided by step S505 of the present invention, the final processed voice can be played through a speaker.
[0115] Optionally, the processed voice can be transmitted as an electrical signal to a speaker at a location other than the location where the voice originated, and the speaker can convert the electrical signal into an audio signal to play the processed voice.
[0116] Optionally, Figure 6 This is a schematic diagram illustrating the speaker playback corresponding to the position of a driver or passenger in a vehicle according to an embodiment of the present invention, such as... Figure 6 As shown, an ellipse can be used to represent a seat in a vehicle's cabin, a hollow circle can be used to represent a speaker that is not playing music, a black circle can be used to represent a speaker that is playing music, and an inverted triangle can be used to represent the driver or passenger on the seat.
[0117] This invention collects noise signals from a vehicle, along with the number and location information of the occupants. It determines whether the volume of the noise signal and the number of occupants meet the activation conditions for a sound acquisition device. If the conditions are met, the sound acquisition device at the location of the occupants is activated. The activated device then collects the original speech information of the corresponding occupants. The volume and / or clarity of the collected original speech information can be enhanced to obtain target speech information. The corresponding audio playback device can be controlled to play the target speech. By automating the activation of the sound acquisition device and the enhancement of the original speech information, the invention enables smooth communication between occupants in noisy environments, thus solving the technical problem of poor speech information processing and improving the overall effectiveness of speech information processing.
[0118] Example 3
[0119] According to embodiments of the present invention, a vehicle voice information processing device is also provided. It should be noted that this vehicle voice information processing device can be used to execute a vehicle voice information processing method as described in Embodiment 1.
[0120] Figure 7 This is a schematic diagram of a vehicle voice information processing device according to an embodiment of the present invention, such as... Figure 7 As shown, the voice information processing device 700 in the vehicle may include: an acquisition unit 702, an activation unit 704, a processing unit 706, and a playback unit 708.
[0121] The acquisition unit 702 is used to acquire the volume information of the noise signal in the vehicle, as well as the number of the driving and riding objects in the vehicle and the position information of the driving and riding objects in the vehicle.
[0122] The activation unit 704 is used to activate the sound acquisition device in response to the volume information and quantity meeting the activation conditions of the sound acquisition device in the vehicle, wherein the sound acquisition device corresponds to the location information.
[0123] The processing unit 706 is used to enhance the original speech information acquired by the activated sound acquisition device to obtain target speech information, wherein the volume of the target speech information is higher than that of the original speech information, and / or the clarity of the target speech information is higher than that of the original speech information.
[0124] The playback unit 708 is used to control the sound playback device to play the target voice information.
[0125] Optionally, the activation unit 704 may include: a detection module, used to detect the status data of the auxiliary voice system in the vehicle in response to the volume information exceeding a first volume threshold and the number being greater than or equal to a quantity threshold; and an activation module, used to activate the sound acquisition device in response to the detected status data indicating that the auxiliary voice system is not activated.
[0126] Optionally, the acquisition unit 702 may include: a first determining module, used to determine a noise signal based on the vibration amplitude of the electret film in the condenser electret microphone in the vehicle, wherein the vibration amplitude is used to characterize the change in capacitance of the condenser electret microphone; a conversion module, used to convert the detected noise signal into a corresponding electrical signal; and a second determining module, used to determine volume information based on the electrical signal.
[0127] Optionally, the acquisition unit 702 may include: a first control module for controlling the seat occupancy sensor in the vehicle to identify the driving and riding objects and obtain quantity and location information; and / or a second control module for controlling the image acquisition device in the vehicle to acquire images of the driving and riding objects and obtain image information of the images; and a third determination module for determining the quantity and location information based on the image information.
[0128] Optionally, the first control module may include: a connection submodule for connecting the circuit containing the seat occupancy sensor in response to compression of the press-fit safety configuration of the seat occupancy sensor; and an output submodule for outputting quantity and position information after the circuit connection.
[0129] Optionally, the third determining module may include: a detection submodule for detecting image information and obtaining detection results; a matching submodule for matching the detection results with seats in the vehicle and obtaining matching results; and a determining submodule for determining quantity and location information based on the matching results.
[0130] Optionally, the processing unit 706 may include: a first processing module for performing echo cancellation and noise reduction processing on the original speech information to obtain a processing result; a fourth determining module for determining the speech volume of the processing result based on the average amplitude level of the processing result within a target time range; and a second processing module for performing enhancement processing on the processing result in response to the speech volume being less than a second volume threshold to obtain target speech information.
[0131] According to an embodiment of the present invention, an acquisition unit acquires the volume information of noise signals in a vehicle, as well as the number of occupants in the vehicle and their location information within the vehicle; an activation unit activates a sound acquisition device in response to the volume information and number satisfying the activation conditions for starting the sound acquisition device in the vehicle, wherein the sound acquisition device corresponds to the location information; a processing unit enhances the original speech information acquired by the activated sound acquisition device to obtain target speech information, wherein the volume of the target speech information is higher than that of the original speech information, and / or the clarity of the target speech information is higher than that of the original speech information; and a playback unit controls a sound playback device to play the target speech information, thereby solving the technical problem of poor speech information processing effect and achieving the technical effect of improving the speech information processing effect.
[0132] Example 4
[0133] According to an embodiment of the present invention, a computer-readable storage medium is also provided, the storage medium including a stored program, wherein the program executes the method for processing voice information in a vehicle as described in Embodiment 1.
[0134] Example 5
[0135] According to an embodiment of the present invention, a processor is also provided for running a program, wherein the program executes the method for processing voice information in a vehicle as described in Embodiment 1.
[0136] Example 6
[0137] According to an embodiment of the present invention, a vehicle is also provided, which is used to perform any of the vehicle voice information processing methods in Embodiment 1.
[0138] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0139] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0141] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0142] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0143] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0144] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for processing voice information in a vehicle, characterized in that, include: Obtain the volume information of noise signals in the vehicle; In response to pressure being applied to the press-fit mounting configuration of the seat occupancy sensor in the vehicle, the circuit containing the seat occupancy sensor is connected; After the circuit is connected, the results collected by the seat occupancy sensor are obtained; Control the image acquisition device in the vehicle to acquire images of the driver and passengers, and obtain image information of the images; The image information is then detected to obtain the detection results; The detection results are matched with the seats in the vehicle to obtain matching results; Based on the matching results, the results collected by the seat occupancy sensor are supplemented to determine the number of passengers in the vehicle and the location information of the passengers in the vehicle. In response to the volume information exceeding a first volume threshold and the quantity being greater than or equal to a quantity threshold, the status data of the auxiliary voice system in the vehicle is detected; In response to the detected status data indicating that the auxiliary voice system is not activated, the sound acquisition device is activated, wherein the sound acquisition device corresponds to the location information; The raw speech information acquired by the sound acquisition device is subjected to echo cancellation and noise reduction processing to obtain the processing result; The voice volume of the processed result is determined based on the average amplitude level of the processed result within the target time range; In response to the voice volume being less than a second volume threshold, the processing result is enhanced to obtain target voice information, wherein the volume information of the target voice information is higher than the volume information of the original voice information, and / or the clarity of the target voice information is higher than the clarity of the original voice information; Control the sound playback device to play the target voice information.
2. The method according to claim 1, characterized in that, Obtain the volume information of noise signals in the vehicle, including: The noise signal is determined based on the vibration amplitude of the electret film in the electret microphone in the vehicle, wherein the vibration amplitude is used to characterize the change in capacitance of the electret microphone. The detected noise signal is converted into a corresponding electrical signal; The volume information is determined based on the electrical signal.
3. A device for processing voice information in a vehicle, characterized in that, include: Acquisition unit, used to acquire volume information of noise signals in the vehicle; In response to the compression of the seat occupancy sensor in the vehicle, the circuit containing the seat occupancy sensor is connected; after the circuit is connected, the result collected by the seat occupancy sensor is obtained; the image acquisition device in the vehicle is controlled to acquire an image of the driver and passenger, and the image information of the image is obtained. The image information is then detected to obtain the detection results; The detection results are matched with the seats in the vehicle to obtain matching results; Based on the matching results, the results collected by the seat occupancy sensor are supplemented to determine the number of passengers in the vehicle and the location information of the passengers in the vehicle. An activation unit is configured to detect the status data of the auxiliary voice system in the vehicle in response to the volume information exceeding a first volume threshold and the quantity being greater than or equal to a quantity threshold; and to activate the sound acquisition device in response to the detected status data indicating that the auxiliary voice system is not activated, wherein the sound acquisition device corresponds to the location information. The processing unit is configured to perform echo cancellation and noise reduction processing on the original speech information acquired by the sound acquisition device to obtain a processing result; determine the speech volume of the processing result based on the average amplitude level of the processing result within a target time range; and, in response to the speech volume being less than a second volume threshold, perform enhancement processing on the processing result to obtain target speech information, wherein the volume information of the target speech information is higher than the volume information of the original speech information, and / or the clarity of the target speech information is higher than the clarity of the original speech information. The playback unit is used to control the sound playback device to play the target voice information.
4. A processor, characterized in that, The processor is used to run a program, wherein the program, when run by the processor, performs the method for processing voice information in a vehicle as described in any one of claims 1 to 2.
5. A vehicle, characterized in that, The vehicle is used to perform the method for processing voice information in a vehicle as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Personnel dialogue enhancer and method for vehicle, vehicle and storage medium
CN114664317A
Vehicle in cabin sound processing system
US20160029111A1