Vehicle cabin sound control method, device, equipment and storage medium

By acquiring information about the location of human voices in the cockpit and their sources, and combining this with image data to divide the cockpit into areas for voice enhancement and cancellation, differentiated sound control is achieved. This solves the problem of cumbersome cockpit sound adjustment in existing technologies and improves the convenience and comfort of the vehicle cockpit.

CN122493868APending Publication Date: 2026-07-31DONGFENG MOTOR CO LTD DONGFENG NISSAN PASSENGER VEHICLE CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGFENG MOTOR CO LTD DONGFENG NISSAN PASSENGER VEHICLE CO
Filing Date
2026-04-28
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing vehicle cabin sound control methods cannot adjust the sound of different cabin zones according to the auditory needs of users in the cabin, resulting in cumbersome interaction steps and affecting the entertainment and auditory experience of non-speaking occupants.

Method used

By acquiring cockpit dialogue voices and voice source location information, and combining cockpit image data to divide the cockpit into voice enhancement zones and voice cancellation zones, differentiated sound control is achieved. This includes reducing the volume of entertainment audio and generating enhanced dialogue sounds in the voice enhancement zone, and generating dialogue voice reflection waves in the voice cancellation zone to cancel out the sound.

Benefits of technology

Clear communication between passengers in the conversation zone can be ensured without the need for manual volume adjustment, avoiding auditory interference for passengers in the non-conversation zone and improving the convenience and comfort of using cockpit audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493868A_ABST
    Figure CN122493868A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of cockpit control technology and discloses a method, device, equipment, and storage medium for controlling vehicle cockpit sound. The method includes: acquiring cockpit conversational voices and corresponding voice source location information; dividing the cockpit location into voice enhancement zones and voice cancellation zones based on the voice source location information and cockpit image data; and controlling the vehicle cockpit sound based on the voice enhancement zones and the voice cancellation zones. Through this method, the auditory needs of occupants requiring conversation in certain cockpit zones are met, while preventing auditory interference from conversational voices in occupants in areas without such needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cockpit control technology, and more particularly to vehicle cockpit sound control methods, devices, equipment, and storage media. Background Technology

[0002] Currently, in-vehicle audio systems generally have basic functions such as volume adjustment, sound field balancing, and sound effects processing, and can achieve global audio control through physical buttons, touch screens, or voice commands. In real driving scenarios, there are often needs for multi-person conversations, which require control of the cabin entertainment volume.

[0003] Current vehicle cabin sound control methods primarily rely on physical buttons or central control screens to manually adjust the volume of entertainment audio. When music is playing while driving, if someone in the cabin starts chatting, to prevent the music from drowning out the conversation, the driver or front passenger will use buttons to control the entertainment domain controller to uniformly adjust the output loudness of all speakers in the vehicle, lowering the volume to ensure the conversation is heard clearly. This global manual adjustment method is cumbersome and requires the driver's attention. Furthermore, uniformly lowering the volume throughout the vehicle can deprive non-conversational occupants of a normal entertainment listening experience.

[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this invention is to provide a vehicle cabin sound control method, device, equipment, and storage medium, aiming to solve the technical problem that existing technologies cannot specifically adjust the sound of different cabin zones according to the auditory needs of users in the cabin.

[0006] To achieve the above objectives, the present invention proposes a vehicle cabin sound control method, the vehicle cabin sound control method comprising: Acquire cockpit dialogue voices and corresponding voice source location information; Based on the human voice source location information and cockpit image data, the cockpit location partition is divided into a human voice enhancement partition and a human voice cancellation partition; The vehicle cabin sound is controlled based on the human voice enhancement zone and the human voice cancellation zone.

[0007] In some embodiments, the step of dividing the cabin location partition into a voice enhancement partition and a voice cancellation partition based on the voice source location information and the cabin image data includes: Based on the location information of the human voice source and the cockpit image data, the dialogue participation zone and the dialogue listening zone are determined from the cockpit location zone; The dialogue participation zone and the dialogue listening zone are defined as voice enhancement zones; Based on the voice enhancement partition, a voice cancellation partition is determined from the cabin position partition.

[0008] In some embodiments, the step of determining dialogue participation zones and dialogue listening zones from the cockpit location zones based on the human voice source location information and the cockpit image data includes: Mouth shape features, facial features, eye features, and limb features are extracted based on the cockpit image data; Based on the location information of the human voice source and the lip shape features, the dialogue participation zone is determined; The dialogue listening zone is determined based on the dialogue participation zone, the facial features, the eye features, and the body features.

[0009] In some embodiments, the vehicle cabin sound includes entertainment audio and cabin conversation voices; The step of controlling the vehicle cabin sound based on the human voice enhancement zone and the human voice cancellation zone includes: The volume of the entertainment audio is reduced in the voice enhancement zone, and a dialogue enhancement tone is generated based on the cockpit dialogue voice, so as to enhance the cockpit dialogue voice based on the dialogue enhancement tone; In the voice cancellation zone, a voice reflection wave is generated based on the cockpit voice dialogue, and the cockpit voice dialogue is cancelled based on the voice reflection wave.

[0010] In some embodiments, the step of reducing the volume of the entertainment audio sound in the voice enhancement partition includes: The entertainment audio is split into tracks to obtain entertainment vocal audio and entertainment instrument audio; The volume of the entertainment instrument sound audio is reduced in the human voice enhancement zone according to a first volume reduction value; The volume of the entertainment voice audio is reduced in the voice enhancement zone according to a second volume reduction value, wherein the second volume reduction value is greater than the first volume reduction value.

[0011] In some embodiments, the step of obtaining cockpit dialogue voices and corresponding voice source location information includes: Acquire audio from the array microphones and audio played in the cockpit; The cockpit audio played in the audio collected by the array microphones is sound-masked to obtain the cockpit dialogue voice; The location information of the corresponding human voice source is determined based on the cockpit dialogue voice.

[0012] In some embodiments, the method further includes: The call audio is split into tracks to obtain the call voice audio and the call noise audio. The volume of the call noise audio is reduced in the target call location area to output the noise-reduced call voice audio.

[0013] Furthermore, to achieve the above objectives, the present invention also proposes a vehicle cabin sound control device, the vehicle cabin sound control device comprising: The data acquisition module is used to acquire cockpit dialogue voices and corresponding voice source location information; The data processing module is used to divide the cockpit location partition into a voice enhancement partition and a voice cancellation partition based on the voice source location information and cockpit image data; A sound control module is used to control the vehicle cabin sound based on the human voice enhancement zone and the human voice cancellation zone.

[0014] Furthermore, to achieve the above objectives, the present invention also proposes a vehicle cabin sound control device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle cabin sound control method described above.

[0015] In addition, to achieve the above objectives, the present invention also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the vehicle cockpit voice control method described above.

[0016] This invention acquires information about the location of human voices and sound sources in cockpit conversations, enabling accurate perception of the occurrence of conversations and the specific location of sound sources within the cockpit. By combining cockpit image data to divide the cockpit into voice enhancement and voice cancellation zones, it achieves accurate matching between the cockpit space and the occupants' conversational behavior and intentions. Based on these voice enhancement and voice cancellation zones, it performs differentiated control of the vehicle cockpit sound, ensuring that areas with occupants' conversational needs are met while avoiding interference with the entertainment audio experience of occupants not participating in conversations. It eliminates the need for complex manual volume adjustments, simplifies the cockpit sound control process, and improves the convenience and comfort of using vehicle cockpit audio. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the vehicle cabin sound control method of the present invention. Figure 2 This is a flowchart illustrating a second embodiment of the vehicle cabin sound control method of the present invention. Figure 3 This is a simplified flowchart of the vehicle cabin sound control method provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the module structure of the vehicle cabin sound control device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the vehicle cabin sound control method in this embodiment of the invention.

[0020] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of this application embodiment is: to obtain the cockpit dialogue voice and the corresponding voice source location information; to divide the cockpit location into a voice enhancement zone and a voice cancellation zone based on the voice source location information and cockpit image data; and to control the vehicle cockpit sound based on the voice enhancement zone and the voice cancellation zone.

[0024] In this embodiment, for ease of description, the following description will focus on the vehicle cabin voice control device as the implementing entity.

[0025] In the existing technology, manually adjusting the music volume requires multiple interactive steps, such as finding the volume adjustment button, lowering the volume until the conversation can be heard clearly, and adjusting the volume back after the conversation ends. This requires the operator's attention. Furthermore, since the volume range is the volume of the entire vehicle, adjusting the volume will force people in the cabin who are not participating in the conversation to lose their entertainment sound experience. It is impossible to adjust the sound of different zones according to the auditory needs of users in the cabin.

[0026] This application provides a solution that, by acquiring information about the location of human voices and their sources in cockpit conversations, can accurately perceive the occurrence of conversations within the cockpit and the specific location of the sound sources. By combining cockpit image data to divide the cockpit into voice enhancement zones and voice cancellation zones, it can achieve accurate matching between the cockpit space and the occupants' conversational behavior and intentions. Based on the voice enhancement zones and voice cancellation zones, differentiated control of the vehicle cockpit sound is performed, eliminating the need for complex manual volume adjustments. This ensures that occupants in cockpit zones with conversational needs have their auditory communication needs met, while also preventing the entertainment auditory experience of occupants in cockpit zones without conversational needs from being interfered with by human voices, thus improving the convenience and comfort of using vehicle cockpit audio.

[0027] As can be seen from the above embodiments, this application achieves the effect of ensuring that passengers with dialogue needs in the cabin area meet their auditory communication needs, while avoiding the auditory experience of passengers without dialogue needs in the cabin area being interfered with by the voices of conversational people.

[0028] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a vehicle cockpit voice control device. The following description uses a vehicle cockpit voice control device as an example to illustrate this embodiment and the subsequent embodiments.

[0029] Based on this, embodiments of this application provide a vehicle cabin sound control method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle cabin sound control method of this application.

[0030] In this embodiment, the vehicle cabin sound control method includes steps S10 to S30: Step S10: Obtain the cockpit dialogue voice and the corresponding voice source location information.

[0031] It should be noted that the cockpit dialogue voice is the voice signal generated by the communication between the occupants inside the vehicle cockpit. This voice signal is a representation of the existence of effective dialogue behavior in the cockpit. This signal is acquired in real time by the audio acquisition equipment in the cockpit, and is obtained after the cockpit audio playback shielding and environmental noise removal. It does not contain cockpit audio playback and environmental noise components, and is the basic audio data for intelligent cockpit sound control.

[0032] Additionally, the voice source location information is coordinate data used to identify the specific spatial location of the voices in the cockpit conversation within the vehicle cabin. Based on the voice source location information, the cabin location zone corresponding to the source of the voices in the cockpit conversation can be determined.

[0033] Understandably, real-time audio signals from inside the cockpit are acquired via an array microphone, and the acquired audio signals are transmitted to a digital signal processor (DSP) using a time division multiplexing (TDM) protocol. In the DSP, echo cancellation and noise reduction (ECNR) processing is performed. During the processing, a narrow beam is used to form a pickup beam, and time difference calculation is used to quickly locate the sound source. At the same time, near-end media stereo echo is eliminated and human voices are separated from environmental noise, ultimately obtaining clear cockpit dialogue and accurate information on the location of human voice sources.

[0034] In one feasible implementation, step S10 may include steps S11 to S13: Step S11: Acquire audio from the array microphones and audio played in the cockpit.

[0035] It should be noted that an array microphone is a multi-microphone combination acquisition device used inside the vehicle cabin to collect audio signals. The audio captured by the array microphone is the real-time acquisition of all mixed sound signals inside the vehicle cabin. This signal includes human voice signals from conversations between occupants, ambient noise signals from outside the cabin, and audio signals played from the cabin speakers, etc., and is the raw audio data for extracting human voice conversations within the cabin. The mixed audio signal is transmitted to a digital signal processor using a time-division multiplexing protocol, providing the basic data for subsequent audio processing.

[0036] Additionally, the in-cabin audio playback refers to entertainment audio signals played out by the vehicle's internal speakers. These audio signals are generated after being input from an Android audio source to the entertainment domain controller. The in-cabin audio playback includes various types such as in-cabin music, Bluetooth music, radio, audiobooks, CD music, and USB music, providing occupants with an entertainment and auditory experience. This audio is synchronously captured by the array microphones.

[0037] Understandably, after the vehicle starts, the array microphones remain active, capturing all sound within the cabin in real time. The electrical signals converted from sound vibrations are then subjected to spectral analysis, including Fourier transform, wavelet transform, and Mel-frequency cepstral coefficient transform, to obtain the audio captured by the array microphones. Simultaneously, the cabin playback audio, which will be transmitted to the speakers, is obtained from the audio output link of the entertainment domain controller. This synchronizes the acquisition of audio from the array microphones and the cabin playback audio, providing complete comparative and raw data support for subsequent cabin playback audio masking processing.

[0038] Step S12: Sound masking is applied to the cockpit audio collected by the array microphones to obtain the cockpit dialogue voice.

[0039] It should be noted that sound masking is a process that filters out and eliminates audio components in the audio collected by the array microphones that match the audio played in the cockpit. This can remove interference from the entertainment audio played in the cockpit itself and prevent the audio played in the cockpit from affecting the accuracy of extracting human voices in cockpit conversations.

[0040] Understandably, the audio captured by the array microphones and the audio played in the cockpit are synchronously input into a digital signal processor. An audio feature comparison algorithm identifies audio components in the array microphone audio that are identical to the cockpit audio and filters them out, thus masking the cockpit audio. Then, noise stripping is performed on the remaining audio signal after masking to eliminate environmental noise interference, ultimately resulting in cockpit dialogue voices containing only the voices of the occupants.

[0041] Step S13: Determine the location information of the corresponding human voice source based on the cockpit dialogue human voice.

[0042] Understandably, the clean cockpit voice dialogue is input into a digital signal processor, which then calls the ECNR algorithm. Multiple narrow beams from an array of microphones form a pickup beam, and combined with time difference of sound propagation calculations, the direction of the cockpit voice dialogue is quickly located. This direction is then converted into coordinate data matching the cockpit's location zones, ultimately determining the location information of the voice source corresponding to the cockpit voice dialogue.

[0043] In this embodiment, simultaneously acquiring audio from the array microphones and the cockpit audio provides accurate comparative data for audio masking processing. Performing sound masking on the cockpit audio eliminates entertainment audio interference from the array microphones, resulting in pure cockpit conversational voices. Based on this pure cockpit conversational voice, the ECNR algorithm is used to determine the voice source location information, improving the accuracy of sound source localization and providing reliable location data support for subsequent cockpit location zoning.

[0044] Step S20: Based on the human voice source location information and cockpit image data, divide the cockpit location partition into a human voice enhancement partition and a human voice cancellation partition.

[0045] It should be noted that the cockpit image data is RGB format image data of the cockpit interior captured by the Occupant Monitoring System (OMS) cameras and output after preprocessing. The cockpit image data undergoes color matrix correction, lens shading correction, gamma correction, sharpening, automatic exposure, and noise reduction processing by the camera's built-in digital signal processor before being transmitted to the entertainment domain controller. The cockpit image data carries information on the facial features, body postures, and spatial positions of each occupant within the cockpit, and is the raw visual data for extracting occupant behavior-related features.

[0046] In addition, cabin position zoning refers to the pre-divided independent areas according to the passenger space of the vehicle cabin. Each zone corresponds to an independent passenger space and is the basic spatial unit for realizing cabin sound zoning control. For example, in a two-row five-seat cabin, the cabin position zoning specifically includes the driver's area, the front passenger area, the left rear area, the middle rear area, and the right rear area.

[0047] Furthermore, the voice enhancement zone is a cockpit location zone identified as an area where participants in a conversation are engaged in dialogue. Occupants within this zone need to enhance the clarity of their voices to ensure smooth communication; therefore, voice enhancement processing is required in this cockpit space. The participants in the conversation include both the speaker and the listener.

[0048] Voice cancellation zones are areas within the cabin seating area that are designated as non-conversational zones. Passengers in these zones have no intention of participating in conversations, and it is necessary to block out conversational voices within the cabin to ensure an entertainment experience. These are cabin spaces that require voice cancellation processing.

[0049] Understandably, the process involves first determining the cockpit location zone corresponding to the voice source based on the location information of the voice source, then identifying the mouth shape, facial turning, and eye gaze of each occupant in the cockpit through cockpit image data, and combining this with information such as the occupant's body movement and facial orientation to determine the participants and non-participants in the conversation. The cockpit location zone where the participants in the conversation are located is designated as the voice enhancement zone, and the cockpit location zone where the non-participants in the conversation are located is designated as the voice cancellation zone, thus completing the classification and division of the cockpit location zones.

[0050] In one feasible implementation, step S20 may include steps S21 to S23: Step S21: Based on the human voice source location information and the cockpit image data, determine the dialogue participation zone and dialogue listening zone from the cockpit location zone.

[0051] It should be noted that the dialogue participation zone is the area in the cabin position zone that is determined to be a valid dialogue voice. The occupants in this area are the direct speakers of the dialogue and the direct participants in the dialogue scenario. The dialogue participation zone needs to perform voice enhancement processing to ensure the clarity of the dialogue.

[0052] Furthermore, the dialogue listening zone is a cockpit seating area identified as having the intention to participate in a dialogue but not emitting effective voice input. Occupants in this zone demonstrate their attention to the dialogue through facial expressions, eye contact, and other behaviors, making them indirect participants in the dialogue. The dialogue listening zone requires voice enhancement processing to ensure occupants can clearly hear the dialogue content.

[0053] Understandably, the audio signals collected by the array microphones within a preset time period are first processed using echo cancellation and noise suppression algorithms to obtain the location information of the human voice source and match it with the corresponding cabin location zone. Then, the cabin image data within the preset time period is acquired by the occupant monitoring system's cameras. Using a local AI model, the mouth shape, facial turning, and eye gaze of each occupant in the cabin are identified. Combined with the occupant's body activity information, areas where valid conversational voices are emitted are selected and marked as dialogue participation zones. At the same time, areas where occupants who have the intention to engage in conversation but have not spoken are selected and marked as dialogue listening zones. The definitions of dialogue participation zones and dialogue listening zones will remain in place until a preset time after which human voices disappear, such as 1 minute. If the microphones do not detect human voices in the cabin after the preset time, the definitions will be reset.

[0054] In one feasible implementation, step S21 may include steps S211 to S213: Step S211: Extract mouth features, facial features, eye features, and limb features based on the cockpit image data.

[0055] It should be noted that the mouth shape feature is a feature related to the shape and movement of the occupant's mouth extracted from cockpit image data. This feature reflects whether the occupant is speaking and is used to determine whether the occupant is the speaker in a conversation. By analyzing the pixel distribution, contour changes, and motion trajectory of the occupant's mouth region in the cockpit image, this feature can be accurately obtained and used to identify the occupant's participation in the conversation.

[0056] Additionally, facial features are extracted from cockpit image data, focusing on the overall shape and orientation of the occupant's face. These features primarily reflect the direction of facial rotation and posture, used to determine whether the occupant is paying attention to the conversation. By analyzing facial features, it can be determined whether the occupant's face is facing the conversational area, thereby assessing their intention to participate in the conversation.

[0057] Furthermore, eye features are extracted from cockpit image data, relating to the occupant's eye gaze and eyelid state. These features reflect the occupant's gaze direction and eye-opening / closing status, and are used to identify whether the occupant is paying attention to the conversation. By analyzing eye features, it can be determined whether the occupant's gaze is directed towards the conversation engagement area, thereby judging whether the occupant intends to listen to the conversation.

[0058] Additionally, body language features are extracted from cockpit image data, showing the occupant's body movements and range of motion. These features reflect the occupant's physical activity and orientation, helping to determine their engagement in conversation. Body language features can help determine whether an occupant is actively communicating or passively listening, improving the accuracy of conversation engagement identification.

[0059] Understandably, the preprocessed cockpit image data is input into the local AI model built into the entertainment domain controller. This local AI model then performs region-by-region visual analysis of the cockpit image data, targeting the corresponding image region for each occupant. Specifically, it extracts mouth shape and movement-related features, facial features related to overall facial shape and rotation direction, eye features related to gaze and eyelid state, and limb features related to range of motion and activity state, thus completing the extraction of all target features.

[0060] Step S212: Determine the dialogue participation zone based on the human voice source location information and the lip shape features.

[0061] Understandably, the process begins by matching the location information of the human voice source with the corresponding cockpit location partitions to obtain preliminary candidate vocalization partitions. Then, the extracted lip-sync features are used to verify the vocalization status of occupants within these candidate partitions, confirming that the lip-sync features of the occupants in that partition meet the feature criteria for vocalization status. The cockpit location partition where the occupant whose lip-sync features match the vocalization status is located is determined as the dialogue participation partition.

[0062] Step S213: Determine the dialogue listening zone based on the dialogue participation zone, the facial features, the eye features, and the body features.

[0063] Understandably, the established dialogue participation zones are used as a reference benchmark to analyze the characteristics of occupants in the remaining cabin zones. This involves detecting whether the facial features of occupants in each zone are facing the dialogue participation zone, whether their eye features show a gaze directed towards the dialogue participation zone, and whether their body language demonstrates high levels of attention and activity. Cabin zones that meet the criteria of facial orientation towards the dialogue participation zone, focused gaze on the dialogue participation zone, and body language consistent with a listening state are designated as dialogue listening zones.

[0064] In practice, the dialogue participation partition is removed from the cockpit position partition to obtain the auditory intent to be identified partition. The facial features, eye features, body features of each auditory intent to be identified partition, as well as the dialogue participation partition, are input into the pre-trained local large model to obtain the facial orientation, eye opening and closing status, and body activity level of each auditory intent to be identified partition. The cockpit position partition of the occupant whose facial orientation is not facing any dialogue participation partition, whose eye feature is closed, and whose body activity level is lower than the preset activity level value is marked as the dialogue shielding partition. The dialogue shielding partition is removed from the auditory intent to be identified partition to obtain the dialogue listening partition.

[0065] In this embodiment, by extracting mouth shape features, facial features, eye features, and body features from cockpit image data, comprehensive visual information related to passenger dialogue behavior can be obtained. Combining the location information of human voice sources with mouth shape features to determine dialogue participation zones can effectively improve the accuracy of dialogue voice area recognition. Based on dialogue participation zones combined with multiple types of visual features to determine dialogue listening zones, indirect participation areas that are of interest in the dialogue can be accurately screened out.

[0066] Step S22: Determine the dialogue participation partition and the dialogue listening partition as a voice enhancement partition.

[0067] Understandably, passengers in the dialogue participation zone are the speakers who directly engage in the dialogue, while passengers in the dialogue listening zone are the listeners who are paying attention to the dialogue. Passengers in both zones need to clearly obtain the dialogue audio information. Therefore, both the dialogue participation zone and the dialogue listening zone need to be designated as voice enhancement zones to ensure that all passengers who participate in or pay attention to the dialogue can have a clear auditory experience of the dialogue.

[0068] Step S23: Determine the voice cancellation zone from the cabin position zone based on the voice enhancement zone.

[0069] Understandably, by comparing all cabin position zones with voice enhancement zones one by one, eliminating all areas that have been designated as voice enhancement zones, the remaining unmarked cabin position zones are the areas where occupants who do not intend to participate in dialogue are located, and these remaining areas are designated as voice cancellation zones.

[0070] Step S30: Control the vehicle cabin sound based on the human voice enhancement zone and the human voice cancellation zone.

[0071] It should be noted that the vehicle cabin sound is the overall sound formed by the mixture of entertainment audio played in the vehicle cabin and the human voices of the occupants. Entertainment audio includes various types such as cabin music, Bluetooth music, radio, audiobooks, CD music, and USB music. This sound is played out through the speakers in the cabin and directly affects the auditory experience of the occupants.

[0072] Another approach is audio track splitting, which utilizes deep learning models to process mixed audio signals. This algorithm learns the spectral characteristics and temporal variation patterns of different sounds to separate human voices and various instrument sounds in entertainment audio into independent audio tracks, enabling differentiated control of regional sound. Track splitting is performed by a Neural Processing Unit (NPU) that executes the audio separation operation. This operation uses a deep learning model to learn the spectral characteristics and temporal variation patterns of a large amount of labeled speech data, splitting mixed multi-source call audio into independent single-source tracks.

[0073] Furthermore, the voice enhancement tone is an audio signal generated by adjusting the loudness of the speaking voices. This signal can improve the loudness and resolution of the speaking voices within the corresponding zone. The voice cancellation tone is a sound wave signal with the opposite phase to the acquired speaking voices. When this signal is superimposed on the speaking voices, it can reduce or cancel out the speaking voices, ensuring the entertainment listening experience for those who are not involved in the conversation.

[0074] Understandably, for the voice enhancement zone, the intelligent dialogue mode is activated, employing an audio track-splitting algorithm to separate entertainment audio into voice and instrument audio. Instrument audio undergoes slight attenuation, while voice audio undergoes significant attenuation, and voice enhancement sounds are played simultaneously to improve dialogue clarity. For the voice cancellation zone, the intelligent sound shield mode is activated, maintaining consistency between the entertainment audio track processing status and the function's off state, while voice cancellation sounds are played to mask dialogue voices. This differentiated zone processing achieves accurate control of the vehicle cabin sound.

[0075] In addition, when controlling the vehicle cabin sound based on the human voice enhancement zone and human voice cancellation zone, in addition to attenuating the instrument audio and human voice audio to reduce the entertainment sound played by the speakers, the entertainment sound can also be muted through the track mute control to further optimize the clarity of dialogue listening.

[0076] Specifically, when controlling the sound in the vehicle cabin, various extended functions can be realized, such as karaoke without a microphone, AI virtual panoramic sound effects, custom a cappella mode, custom pure music mode, and a cappella playback mode. All functions are based on audio track separation and independent track control.

[0077] When the microphone-free karaoke mode is enabled, the system uses a sound track-splitting algorithm to separate the vocals from the cabin music and plays songs with the vocals removed. The array microphones collect the cabin vocals, which are then processed using the ECNR algorithm. Based on the processed vocals, a vocal enhancement tone is generated and reverb is applied, creating a microphone-free karaoke effect. Simultaneously, in-vehicle audio editing software can be used to edit the microphone-free karaoke recordings using a single-track editing method, optimizing both vocals and audio. When the AI ​​virtual panoramic sound effect function is enabled, the sound track-splitting algorithm separates the cabin music vocals, and a local AI model individually adjusts different instrument tracks, allowing music without panoramic sound effects to play with a virtual panoramic sound effect. The effect improves with the increase in the number of tracks split by the NPU. When the custom a cappella mode is enabled, only the separated vocal audio is played. When the custom pure music mode is enabled, only the audio of the specified song or pure music audio excluding vocals is played. When the a cappella playback mode is enabled, the user's own a cappella vocals are played separately during the playback of microphone-free karaoke songs, making it easier to identify singing flaws.

[0078] In one feasible implementation, the method may further include steps A10 to A20: Step A10: Separate the call audio into tracks to obtain the call voice audio and the call noise audio.

[0079] It should be noted that the audio signal of the call is the audio signal transmitted in the cross-cabin communication scenario within the vehicle cabin. The cross-cabin communication scenario includes WeChat voice calls and Bluetooth calls, etc. This audio signal is the core voice carrier for passengers to conduct remote calls in the cabin. During the transmission and acquisition process, this signal will be mixed with non-human voice interference components such as environmental noise, line noise, and equipment background noise.

[0080] Additionally, the human voice audio is the pure human voice audio signal extracted from the call audio, representing the voice of the remote caller. This signal is the clear voice component that passengers need to hear, and it does not contain noise or interference. Call noise audio is the non-human voice interference audio signal extracted from the call audio. This signal contains environmental noise, line noise, equipment background noise, and other interference components that can obscure the human voice, reducing call clarity and requiring targeted volume reduction.

[0081] Understandably, the audio of a cross-cabin communication call is transmitted to a neural network processor, which then calls a trained audio track splitting algorithm. The algorithm identifies the human voice features and noise features in the call audio, and splits the mixed call audio into independent human voice audio and noise audio, thus completing the track splitting process.

[0082] Step A20: Reduce the volume of the call noise audio in the target call location area to output the noise-reduced call voice audio.

[0083] It should be noted that the target call location zone is a pre-defined dedicated location zone in the vehicle's cabin for playing call audio. This zone can be the driver's area, a space for the driver to make private calls and clearly listen to the content of the call. It uses a sound zone rendering engine to achieve independent audio control.

[0084] Additionally, the volume of the call noise audio is a parameter characterizing the loudness of the call noise audio. This parameter can be independently adjusted via the track gain control module of the digital signal processor. The adjustment range only covers the target call location zone and will not affect the audio playback status of other cockpit location zones. The noise-reduced call voice audio is a pure call voice signal after eliminating call noise audio interference. This signal retains the complete call voice components without noise masking, ensuring that the driver can clearly hear the content of long-distance calls.

[0085] Understandably, the sound partition rendering engine locks the target call location partition, calls the track gain control module of the digital signal processor, performs volume reduction processing on the call noise audio in the target call location partition, weakens the loudness of the noise signal, and keeps the volume and playback status of the call voice audio unchanged. Only the noise-reduced call voice audio without noise interference is output on the target call location partition side, thus completing the noise reduction processing of the call audio. Other partitions do not output call audio sound-related content.

[0086] In this embodiment, by performing track-splitting processing on the call audio to obtain the call voice audio and the call noise audio, it is possible to achieve independent and accurate control of the call voice and noise. The volume of the call noise audio can be reduced in a targeted manner at the target call location, which can eliminate noise interference without affecting the call voice and output clear noise-reduced call voice. This meets the usage needs of cross-cockpit communication noise reduction and privacy dialogue, and effectively improves the auditory experience and call privacy in cockpit call scenarios.

[0087] This embodiment acquires information about the location of human voices and their sources in the cabin dialogue, enabling accurate perception of the occurrence of dialogue scenarios and the specific location of sound sources within the cabin. By combining cabin image data to divide the cabin into voice enhancement and voice cancellation zones, it achieves accurate matching between the cabin space and the occupants' dialogue behavior and intentions. Based on the voice enhancement and voice cancellation zones, differentiated control is performed on the vehicle cabin sound, eliminating the need for complex manual volume adjustments. This ensures that occupants in cabin zones with dialogue needs have their auditory communication requirements met, while preventing the auditory experience of occupants in cabin zones without dialogue needs from being interfered with by the human voices in the dialogue, thus improving the convenience and comfort of using vehicle cabin audio.

[0088] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 The vehicle cabin audio includes entertainment audio and cabin conversation voices. Step S30 may also include steps S31-S32: Step S31: Reduce the volume of the entertainment audio in the voice enhancement zone, and generate a dialogue enhancement tone based on the cockpit dialogue voice, so as to enhance the cockpit dialogue voice based on the dialogue enhancement tone.

[0089] It should be noted that the voice enhancement zone prioritizes ensuring clear hearing of human voices during cabin conversations, while also appropriately preserving the auditory experience of entertainment audio. Entertainment audio refers to sounds generated by entertainment sources played within the vehicle cabin, including cabin music, Bluetooth music, radio, audiobooks, CD music, and USB music. This audio constitutes the entertainment listening content for cabin occupants, and is input from the Android audio source to the entertainment domain controller before entering the audio processing chain. The volume of the entertainment audio is a parameter characterizing its loudness, which can be independently adjusted by zone through audio track separation and track gain control.

[0090] Additionally, the dialogue enhancement tone is an enhanced speech signal generated by adjusting the loudness of the cockpit dialogue voice. This signal can improve the resolution of the cockpit dialogue voice and effectively counteract the masking effect of entertainment audio on the cockpit dialogue voice. The cockpit dialogue voice is the original speech data used to generate the dialogue enhancement tone.

[0091] Understandably, when the intelligent dialogue mode is activated in the voice enhancement zone, the entertainment audio is first separated into voice and instrument audio using an audio track splitting algorithm. The instrument audio undergoes slight attenuation, while the voice audio undergoes significant attenuation, thus reducing the overall volume of the entertainment audio. Then, based on the processed cockpit dialogue, its loudness parameter is increased to generate a dialogue enhancement tone. This enhanced tone is then mixed with the attenuated entertainment audio and played back, thus amplifying the cockpit dialogue and enabling occupants in the voice enhancement zone to clearly hear the conversation.

[0092] In one feasible implementation, the step of reducing the volume of the entertainment audio sound in the voice enhancement zone in step S31 may include steps S311 to S313: Step S311: Divide the entertainment audio into tracks to obtain entertainment vocal audio and entertainment instrument audio.

[0093] It should be noted that entertainment vocal audio is the vocal signal extracted from entertainment audio. This signal is easily masked by cockpit conversations, requiring targeted volume attenuation processing. Entertainment instrumental audio is the instrumental performance audio signal extracted from entertainment audio, such as drum sounds, piano sounds, bass sounds, guitar sounds, etc., which are the basic components of the melody and rhythm of entertainment audio.

[0094] Understandably, the entertainment audio is transmitted to a neural network processor, which then calls a trained audio track splitting algorithm to perform multi-source feature recognition and separation processing on the entertainment audio, distinguishing between human voice features and instrument sound features, and splitting the mixed entertainment audio into independent entertainment human voice audio and entertainment instrument sound audio, thus completing the track splitting process for the entertainment audio.

[0095] Step S312: Reduce the volume of the entertainment instrument sound in the human voice enhancement zone according to the first volume reduction value.

[0096] It should be noted that the first volume reduction value is a preset small volume attenuation parameter for the sound and audio of entertainment instruments. This parameter is a fixed decibel value and is used to perform a small attenuation process on the sound and audio of entertainment instruments. For example, the first volume reduction value can be set to 5dB.

[0097] Additionally, the volume of the audio frequencies of entertainment instruments is a parameter that characterizes the loudness of the audio frequencies of entertainment instruments. This parameter can be independently adjusted through the track gain control module of the digital signal processor. The adjustment range only covers the vocal enhancement zone and does not change the playback status of other cabin position zones.

[0098] Understandably, in the audio processing flow of the human voice enhancement zone, the track gain control module of the digital signal processor is called to read the preset first volume reduction value. This attenuation value is applied to the gain control of the entertainment instrument audio. The volume of the entertainment instrument audio is slightly reduced according to the first volume reduction value. The reduced entertainment instrument audio will not cover the cockpit conversation voices, and will also preserve the basic entertainment listening experience for the occupants in the human voice enhancement zone.

[0099] Step S313: Reduce the volume of the entertainment voice audio in the voice enhancement zone according to the second volume reduction value, wherein the second volume reduction value is greater than the first volume reduction value.

[0100] It should be noted that the second volume reduction value is a preset, significant volume attenuation parameter for entertainment vocal audio. This parameter is a fixed decibel value, greater than the first volume reduction value, and is used to perform substantial attenuation processing on the entertainment vocal audio. For example, the second volume reduction value can be set to 20dB. The volume of the entertainment vocal audio is a parameter characterizing its loudness. This parameter can be independently adjusted through the track gain control module of the digital signal processor, and is independent of and does not interfere with the volume adjustment of the entertainment instrumental audio.

[0101] Understandably, in the audio processing flow of the voice enhancement zone, the track gain control module of the digital signal processor is called to read the preset second volume reduction value. The second volume reduction value is greater than the first volume reduction value. This attenuation value is applied to the gain control of the entertainment voice audio. The volume of the entertainment voice audio is significantly reduced according to the second volume reduction value. The significantly attenuated entertainment voice audio will not cause auditory confusion with the cockpit conversation voice, ensuring that the occupants in the voice enhancement zone can clearly distinguish the cockpit conversation voice.

[0102] In this embodiment, by performing track-splitting processing on the entertainment audio, entertainment human voice audio and entertainment instrument audio are obtained. This enables independent and accurate control of different audio sources in the entertainment audio. By slightly reducing the entertainment instrument audio based on a first volume reduction value, the entertainment instrument audio can be reduced while preserving the entertainment experience. By significantly reducing the entertainment human voice audio based on a second volume reduction value, the obscuring interference of entertainment human voice on cockpit conversations can be eliminated. This accurately matches the dialogue clarity requirements of the human voice enhancement zone, balancing the clarity of dialogue and the passenger's entertainment auditory experience without increasing hardware costs.

[0103] Step S32: Generate a voice reflection wave in the voice cancellation zone based on the voice reflection wave in the cockpit, and cancel the voice reflection wave in the cockpit based on the voice reflection wave.

[0104] It should be noted that occupants within the voice cancellation zone are non-participants in the conversation, with no intention of engaging or paying attention to it. To maintain a complete entertainment audio listening experience, the Sound Shield mode is activated. In Sound Shield mode, inverted sound waves are played to weaken or eliminate conversational voices while maintaining normal playback of entertainment audio.

[0105] Additionally, the reflected wave of the dialogue voice is a sound wave signal with a completely opposite phase to the dialogue voice in the cabin. This signal is continuously generated based on the real-time dynamic changes of the dialogue voice in the cabin, and can adapt to the dynamic fluctuations of the voice to maintain a stable cancellation effect. The cabin dialogue voice cancellation is a sound weakening effect produced by the superposition of the reflected wave of the dialogue voice and the dialogue voice in the cabin. This effect can eliminate the interference of the dialogue voice on the voice cancellation zone without changing the playback status of the entertainment audio sound.

[0106] Understandably, when the intelligent sound shield mode is activated in the voice cancellation zone, the track processing status of the entertainment audio is kept consistent with the function's off state, without adjusting the volume or track parameters of the entertainment audio. Then, based on the collected cabin conversation voices, a voice reflection wave with opposite phase is calculated and generated in real time. This reflection wave is mixed with the entertainment audio and played back. The reflection wave is superimposed on the cabin conversation voices entering that zone, thus canceling out the cabin conversation voices and shielding them from interference with the passenger's entertainment experience in that zone.

[0107] This embodiment reduces the volume of entertainment audio in the voice enhancement zone and generates dialogue enhancement sound to strengthen the cockpit dialogue voice, which can ensure the communication hearing needs of the conversation participants and listeners. It generates dialogue voice anti-wave cancellation cockpit dialogue voice in the voice cancellation zone, which can ensure the entertainment audio hearing experience of non-conversation passengers. Through zone-differentiated sound control logic, the communication hearing needs of the conversation participants are met while the entertainment hearing needs of passengers who do not intend to converse are also guaranteed.

[0108] For example, to help understand the implementation flow of the vehicle cockpit sound control method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 , Figure 3 A simplified flowchart of a vehicle cabin sound control method is provided, specifically: When in-cabin entertainment sounds such as in-cabin music, Bluetooth music, radio, audiobooks, CD music, and USB music are detected, i.e., when the user inputs an audio control command for in-cabin playback, continuous audio collection is initiated. Then, it determines whether a valid dialogue (in-cabin conversation) has been detected. During the detection of a valid dialogue, cabin sounds are masked by comparing them with the cabin sound in the mixer to prevent sounds from the cabin itself from being picked up by the microphone. If no valid dialogue is detected, continuous audio collection continues. If a valid dialogue is detected, the location of the sound source is defined by combining OMS information (cabin image data) to obtain the location information of the human voice source corresponding to the in-cabin conversation. This location information is processed using the ENCR algorithm, which uses multiple narrow beams of a matrix microphone to form a pickup beam and calculates the time difference to quickly locate the sound source, eliminating near-end media stereo echoes and separating human voices from environmental noise. Next, the system uses a large local model to identify passengers in each cabin area except the driver's seat. Based on the identification results, the corresponding passenger status is determined, including three categories: passenger present, no passenger, and undetermined. If all areas except the driver's seat are in a state of no passenger or undetermined, the area sound control function is not triggered, and the system enters a function cooldown period, such as 10 seconds. After the cooldown is complete, the system is ready and returns to the initial conditions for identifying entertainment sound playback. If passengers in a particular area are identified, the system sequentially determines whether each area is a conversational voice area or a listener voice area. If it is a conversational voice area or a listener voice area, it is marked as a conversational participant. The conversational participant definition continues until a preset time, such as 1 minute, after the voice disappears. If the microphone does not identify a valid voice after the preset time, the definition is reset, and then entertainment voice attenuation and conversational voice enhancement are performed. If it is neither a conversational voice area nor a listener voice area, it is marked as a muted voice and a conversational voice reflection wave is generated, thus completing the intelligent cabin sound control.

[0109] Furthermore, this embodiment of the invention also proposes a storage medium storing a vehicle cabin sound control program, which, when executed by a processor, implements the steps of the vehicle cabin sound control method described above.

[0110] Reference Figure 4 , Figure 4 This is a structural block diagram of the first embodiment of the vehicle cabin sound control device of the present invention.

[0111] like Figure 4 As shown, the vehicle cabin sound control device proposed in this embodiment of the invention includes: Data acquisition module 10 is used to acquire cockpit dialogue voices and corresponding voice source location information; Data processing module 20 is used to divide the cockpit location partition into a voice enhancement partition and a voice cancellation partition based on the voice source location information and cockpit image data; The sound control module 30 is used to control the vehicle cabin sound based on the human voice enhancement zone and the human voice cancellation zone.

[0112] In some embodiments, the data processing module 20 is further configured to determine a dialogue participation zone and a dialogue listening zone from the cockpit position zone based on the human voice source location information and the cockpit image data; The dialogue participation zone and the dialogue listening zone are defined as voice enhancement zones; Based on the voice enhancement partition, a voice cancellation partition is determined from the cabin position partition.

[0113] In some embodiments, the data processing module 20 is further configured to extract mouth features, facial features, eye features, and limb features based on the cockpit image data; Based on the location information of the human voice source and the lip shape features, the dialogue participation zone is determined; The dialogue listening zone is determined based on the dialogue participation zone, the facial features, the eye features, and the body features.

[0114] In some embodiments, the vehicle cabin sound includes entertainment audio and cabin conversation voices; the sound control module 30 is further configured to reduce the volume of the entertainment audio in the voice enhancement zone and generate a dialogue enhancement tone based on the cabin conversation voices, so as to enhance the cabin conversation voices based on the dialogue enhancement tone. In the voice cancellation zone, a voice reflection wave is generated based on the cockpit voice dialogue, and the cockpit voice dialogue is cancelled based on the voice reflection wave.

[0115] In some embodiments, the sound control module 30 is further configured to split the entertainment audio sound into tracks to obtain entertainment vocal audio and entertainment instrumental audio; The volume of the entertainment instrument sound audio is reduced in the human voice enhancement zone according to a first volume reduction value; The volume of the entertainment voice audio is reduced in the voice enhancement zone according to a second volume reduction value, wherein the second volume reduction value is greater than the first volume reduction value.

[0116] In some embodiments, the data acquisition module 10 is further configured to acquire audio collected by the array microphone and audio played in the cockpit; The cockpit audio played in the audio collected by the array microphones is sound-masked to obtain the cockpit dialogue voice; The location information of the corresponding human voice source is determined based on the cockpit dialogue voice.

[0117] In some embodiments, the sound control module 30 is further configured to split the call audio into tracks to obtain call voice audio and call noise audio. The volume of the call noise audio is reduced in the target call location area to output the noise-reduced call voice audio.

[0118] The vehicle cabin sound control device provided in this application, employing the vehicle cabin sound control method described in the above embodiments, can solve the technical problem that existing technologies cannot specifically adjust the sound of different cabin zones according to the auditory needs of users within the cabin. Compared with the prior art, the beneficial effects of the vehicle cabin sound control device provided in this application are the same as those of the vehicle cabin sound control method provided in the above embodiments, and other technical features in the vehicle cabin sound control device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0119] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.

[0120] In this embodiment, the data acquisition module 10 acquires the voices of passengers in the cabin and the corresponding voice source location information; the data processing module 20 divides the cabin location into a voice enhancement zone and a voice cancellation zone based on the voice source location information and cabin image data; and the sound control module 30 controls the vehicle cabin sound based on the voice enhancement zone and the voice cancellation zone. This solution ensures that passengers in cabin areas requiring dialogue have their auditory needs met, while preventing interference from voices in passenger areas without dialogue requirements, thus improving the convenience and comfort of using vehicle cabin audio.

[0121] This application provides a vehicle cabin sound control device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the vehicle cabin sound control method in the first embodiment described above.

[0122] The following is for reference. Figure 5The diagram illustrates a structural schematic suitable for implementing a vehicle cockpit voice control device according to embodiments of this application. The vehicle cockpit voice control device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The vehicle cabin voice control device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0123] like Figure 5 As shown, the vehicle cockpit voice control device may include a processing unit 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the vehicle cockpit voice control device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the vehicle cockpit voice control equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show vehicle cockpit voice control equipment with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0124] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0125] The vehicle cabin sound control device provided in this application, employing the vehicle cabin sound control method described in the above embodiments, can solve the technical problem that existing technologies cannot specifically adjust the sound of different cabin zones according to the auditory needs of users within the cabin. Compared with the prior art, the beneficial effects of the vehicle cabin sound control device provided in this application are the same as those of the vehicle cabin sound control method provided in the above embodiments, and other technical features of this vehicle cabin sound control device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0126] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0127] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0128] Furthermore, the cockpit dialogue voices, corresponding voice source location information, and cockpit image data involved in this application were all obtained with the user's permission or consent; that is to say, when this application is applied to a specific product or technology, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations, and regulatory standards of the relevant countries and regions.

[0129] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the vehicle cockpit sound control method in the above embodiments.

[0130] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fiber, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0131] The aforementioned computer-readable storage medium may be included in the vehicle cabin voice control device; or it may exist independently and not be installed in the vehicle cabin voice control device.

[0132] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the vehicle cabin sound control device, cause the vehicle cabin sound control device to: acquire cabin conversation voices and corresponding voice source location information; divide the cabin location into voice enhancement zones and voice cancellation zones based on the voice source location information and cabin image data; and control the vehicle cabin sound based on the voice enhancement zones and the voice cancellation zones.

[0133] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0135] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0136] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described vehicle cabin sound control method. This solves the technical problem that existing technologies cannot specifically adjust the sound of different cabin zones according to the auditory needs of users within the cabin. Compared with existing technologies, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the vehicle cabin sound control method provided in the above embodiments, and will not be elaborated upon here.

[0137] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the vehicle cockpit voice control method described above.

[0138] The computer program product provided in this application can solve the technical problem that existing technologies cannot specifically adjust the sound of different cabin zones according to the auditory needs of users in the cabin. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the vehicle cabin sound control method provided in the above embodiments, and will not be repeated here.

[0139] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the content of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A vehicle cabin sound control method characterized by, The vehicle cabin sound control method includes: Acquire cockpit dialogue voices and corresponding voice source location information; Based on the human voice source location information and cockpit image data, the cockpit location partition is divided into a human voice enhancement partition and a human voice cancellation partition; The vehicle cabin sound is controlled based on the human voice enhancement zone and the human voice cancellation zone.

2. The method of claim 1, wherein, The step of dividing the cabin location into voice enhancement zones and voice cancellation zones based on the human voice source location information and the cabin image data includes: Based on the location information of the human voice source and the cockpit image data, the dialogue participation zone and the dialogue listening zone are determined from the cockpit location zone; The dialogue participation zone and the dialogue listening zone are defined as voice enhancement zones; Based on the voice enhancement partition, a voice cancellation partition is determined from the cabin position partition.

3. The method as described in claim 2, characterized in that, The step of determining the dialogue participation zone and dialogue listening zone from the cockpit location zone based on the human voice source location information and the cockpit image data includes: Mouth shape features, facial features, eye features, and limb features are extracted based on the cockpit image data; Based on the location information of the human voice source and the lip shape features, the dialogue participation zone is determined; The dialogue listening zone is determined based on the dialogue participation zone, the facial features, the eye features, and the body features.

4. The method as described in claim 1, characterized in that, The vehicle cabin audio includes entertainment audio and cabin conversation voices; The step of controlling the vehicle cabin sound based on the human voice enhancement zone and the human voice cancellation zone includes: The volume of the entertainment audio is reduced in the voice enhancement zone, and a dialogue enhancement tone is generated based on the cockpit dialogue voice, so as to enhance the cockpit dialogue voice based on the dialogue enhancement tone; In the voice cancellation zone, a voice reflection wave is generated based on the cockpit voice dialogue, and the cockpit voice dialogue is cancelled based on the voice reflection wave.

5. The method as described in claim 4, characterized in that, The step of reducing the volume of the entertainment audio sound in the voice enhancement zone includes: The entertainment audio is split into tracks to obtain entertainment vocal audio and entertainment instrument audio; The volume of the entertainment instrument sound audio is reduced in the human voice enhancement zone according to a first volume reduction value; The volume of the entertainment voice audio is reduced in the voice enhancement zone according to a second volume reduction value, wherein the second volume reduction value is greater than the first volume reduction value.

6. The method according to any one of claims 1 to 5, characterized in that, The steps for obtaining cockpit dialogue voices and corresponding voice source location information include: Acquire audio from the array microphones and audio played in the cockpit; The cockpit audio played in the audio collected by the array microphones is sound-masked to obtain the cockpit dialogue voice; The location information of the corresponding human voice source is determined based on the cockpit dialogue voice.

7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The call audio is split into tracks to obtain the call voice audio and the call noise audio. The volume of the call noise audio is reduced in the target call location area to output the noise-reduced call voice audio.

8. A vehicle cabin sound control device, characterized in that, The device includes: The data acquisition module is used to acquire cockpit dialogue voices and corresponding voice source location information; The data processing module is used to divide the cockpit location partition into a voice enhancement partition and a voice cancellation partition based on the voice source location information and cockpit image data; A sound control module is used to control the vehicle cabin sound based on the human voice enhancement zone and the human voice cancellation zone.

9. A vehicle cabin sound control device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle cockpit sound control method as claimed in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the vehicle cockpit sound control method as described in any one of claims 1 to 7.