In-vehicle sound shielding method, electronic equipment and vehicle

By combining a multimodal AI fusion architecture with computer vision and voice processing technologies, the system accurately determines the areas inside the vehicle to be shielded and controls the sound field, solving the problems of accuracy and privacy protection in in-vehicle sound shielding, reducing the risk of privacy leaks, and improving the user experience.

CN120977276APending Publication Date: 2025-11-18GREAT WALL MOTOR CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511312583.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing in-vehicle systems cannot achieve precise regional sound blocking, lack intelligent recognition of conversation content and privacy protection, leading to the risk of privacy leaks. Furthermore, traditional soundproofing methods cannot achieve precise directional audio management, which damages the user experience.

Method used

Employing a multimodal AI fusion architecture that combines computer vision and speech processing technologies, the system collects in-vehicle image and audio data to determine the directional information of the target occupant, accurately identifies the area to be shielded, and uses sound field control methods to shield audio data within the area to be shielded.

Benefits of technology

It achieves privacy protection for the conversations of target occupants, reduces the risk of privacy leaks, and improves the accuracy of in-vehicle sound management and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977276A_ABST
    Figure CN120977276A_ABST
Patent Text Reader

Abstract

The invention provides an in-vehicle sound shielding method, electronic equipment and a vehicle, and the method comprises the steps: collecting image data and audio data in the vehicle, processing the image data and the audio data, and determining directivity information used for representing a target passenger related to the current conversation content. By determining the directivity information, the in-vehicle passengers who are in conversation can be clearly known. And according to the directivity information, a to-be-shielded area in the vehicle is determined, the to-be-shielded area is the area needing sound shielding, and subsequent accurate sound shielding of the area in the vehicle is facilitated. The audio data is shielded in the to-be-shielded area through the sound field control method, so that passengers in the to-be-shielded area cannot hear conversation contents of target passengers, privacy protection of the conversation contents of the target passengers is realized, and the privacy leakage risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent cockpit of vehicle, and in particular to an in-vehicle sound shielding method, an electronic device and a vehicle. BACKGROUND

[0002] With the popularization of intelligent vehicles, the in-vehicle privacy protection problem is increasingly prominent. The traditional vehicle-mounted system cannot realize accurate regional sound shielding, lacks intelligent recognition and privacy protection mechanism for conversation content, and passengers are easy to be eavesdropped by other passengers or external devices when having private conversations in the vehicle, which exists serious privacy leakage risk. SUMMARY

[0003] Therefore, the present application aims to provide an in-vehicle sound shielding method to solve the problem that accurate regional sound shielding cannot be realized in the vehicle.

[0004] To achieve the above purpose, the first aspect of the present application provides an in-vehicle sound shielding method, comprising: collecting image data and audio data in the vehicle; processing the image data and the audio data to determine directional information for characterizing a target occupant related to the current in-vehicle conversation content; determining a to-be-shielded area in the vehicle according to the directional information; shielding the audio data in the to-be-shielded area by a sound field control method.

[0005] Optionally, the processing the image data and the audio data to determine the directional information for characterizing the target occupant related to the current in-vehicle conversation content comprises: determining the identity tag and the position information of each occupant in the vehicle based on the image data; determining the directional information based on the audio data, the identity tag and the position information of each occupant.

[0006] Optionally, the determining the directional information based on the audio data, the identity tag and the position information of each occupant comprises: determining the sound source information and the voiceprint information of the audio data, combining the identity tag and the position information to determine the sound emitting occupant information; performing format conversion on the audio data to obtain the conversation text; determining the directional information by directional analysis based on the sound emitting occupant information and the conversation text.

[0007] Optionally, the determining the to-be-shielded area in the vehicle according to the directional information comprises: determining a target privacy area in the vehicle according to the directional information. determining the to-be-masked area according to the target privacy area.

[0008] Optionally, the method for controlling a sound field to mask the audio data in the to-be-masked area comprises: generating reverse sound waves according to the audio data; emitting the reverse sound waves to the to-be-masked area.

[0009] Optionally, the method for generating the reverse sound waves according to the audio data comprises: generating initial reverse sound waves according to the audio data; determining sound wave intensities according to the audio data; adjusting intensities of the initial reverse sound waves according to the sound wave intensities to generate the reverse sound waves.

[0010] Optionally, after emitting the reverse sound waves to the to-be-masked area, the method comprises: collecting to-be-detected audio data of the to-be-masked area; determining whether the to-be-detected audio data meets preset masking conditions; in response to the to-be-detected audio data not meeting the preset masking conditions, adjusting the reverse sound waves.

[0011] Optionally, before determining the to-be-masked area in the vehicle according to the directivity information, the method comprises: determining whether a privacy protection function needs to be turned on according to the directivity information or the audio data; in response to determining that the privacy protection function needs to be turned on, determining the to-be-masked area in the vehicle according to the directivity information.

[0012] Based on the same inventive concept, a second aspect of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method of the first aspect when executing the computer program.

[0013] Based on the same inventive concept, a third aspect of the present application further provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method of the first aspect.

[0014] It can be seen from the above that the in-vehicle sound shielding method, electronic device and vehicle provided by the application, wherein the method comprises collecting image data and audio data in the vehicle, processing the image data and the audio data, and determining directional information for characterizing a target occupant related to current conversation content. By determining the directional information, the in-vehicle occupant who is talking can be clearly known. The area to be shielded in the vehicle is determined according to the directional information, that is, the area that needs to be shielded, so as to facilitate subsequent accurate in-vehicle area sound shielding. The audio data is shielded in the area to be shielded by the sound field control method, so that the occupant in the area to be shielded cannot hear the conversation content of the target occupant, realizing the private protection of the conversation content of the target occupant and reducing the risk of privacy leakage. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the application or related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 The flowchart of the in-vehicle sound shielding method of the embodiment of the application is shown in the figure. Figure 2 The in-vehicle sound shielding flowchart of the embodiment of the application is shown in the figure. Figure 3 The flowchart of the in-vehicle occupant identity confirmation and position confirmation of the embodiment of the application is shown in the figure. Figure 4 The flowchart of the voice recognition and directional analysis of the embodiment of the application is shown in the figure. Figure 5 The flowchart of the sound field control and privacy protection of the embodiment of the application is shown in the figure. Figure 6 The structural schematic diagram of the in-vehicle sound shielding device of the embodiment of the application is shown in the figure. Figure 7 The hardware structural schematic diagram of the electronic device of the embodiment of the application is shown in the figure. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of the application more clear, the application will be further described in detail below in combination with specific embodiments and with reference to the drawings.

[0018] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the present application shall be understood as the common meaning understood by a person with ordinary skills in the art to which the embodiments of the present application belong. The terms "first", "second" and similar words used in the embodiments of the present application do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar words mean that the elements or objects before the words cover the elements or objects listed after the words and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0019] As described in the background, in today's era of increasing popularity of intelligent cars, the acoustic environment and voice interaction experience of the vehicle cabin have become an important standard for measuring the intelligent level. However, the current vehicle cabin still has significant defects in privacy protection and intelligent voice control, which are manifested in the following aspects: I. Privacy leakage risk: In-vehicle sensitive conversations lack security protection and have the risk of eavesdropping The vehicle cabin is a relatively closed private space, and passengers often conduct business negotiations, personal emotional exchanges or discuss sensitive personal information. However, the acoustic system of existing vehicles lacks effective privacy protection mechanisms. The vehicle microphone array, entertainment system or networking module can become a potential security vulnerability, and if maliciously attacked or remotely controlled, the conversation content of the passengers can be easily stolen or listened to in real time by third parties.

[0020] II. Lack of intelligent recognition capability: System cannot understand conversation content and identify speaker identity Although most vehicle systems have voice command recognition functions, their intelligent level is still at a preliminary stage. The system can often only respond to preset wake-up words and simple instructions, but cannot truly understand the semantic context of natural conversations, nor can it distinguish the voiceprint identity of different passengers. For example, in a multi-person conversation scenario, the system cannot determine whether a sentence is private content, whether it should be prohibited from recording or transmitting. At the same time, due to the lack of speaker recognition technology, the system cannot achieve personalized response and permission control based on identity, thereby limiting its application in security-sensitive scenarios.

[0021] III. Insufficient sound field control technology: Traditional sound insulation methods cannot achieve precise directional audio management Traditional vehicle sound insulation mainly relies on physical material sound insulation and structural noise reduction, which can block external noise to a certain extent, but cannot actively and accurately control the sound field in the vehicle. For example, when the driver and the rear passengers are talking, the system cannot enhance the voice clarity of the specific seat area; on the contrary, when a passenger is on the phone, the system cannot limit the conversation sound to the individual area to avoid disturbing others or revealing the conversation content. This lack of directional and regional sound field control capability makes the in-vehicle audio environment still in a "one-size-fits-all" state, which is difficult to meet the needs of sound privacy and audio independence in different scenarios.

[0022] IV. Fragmented interactive experience: manual operation and physical sound insulation methods destroy user experience To solve the privacy and interference problems, some vehicles still rely on traditional methods such as manual volume adjustment, physical partitions, or headphones, which seriously destroy the seamless and natural interactive experience of the intelligent cabin. Passengers need to frequently manually operate to switch sound field modes or isolate sound, which not only distracts attention (especially for drivers, which poses a safety risk), but also reduces the overall technology and comfort of the cabin. Users expect the system to automatically recognize scene requirements, such as automatically enabling acoustic privacy zones when detecting private conversations or actively suppressing background conversations when important instructions are recognized, thereby achieving truly intelligent and invisible acoustic management, rather than relying on manual intervention.

[0023] Therefore, the present application proposes an in-vehicle sound shielding method, which adopts a multi-modal AI fusion architecture, combines computer vision and speech processing technology, and realizes intelligent recognition of in-vehicle personnel, semantic understanding of conversation content, and privacy protection based on sound field control.

[0024] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0025] The present application proposes an in-vehicle sound shielding method, referring to Figure 1 , comprising the following steps: Step 102, collecting image data and audio data in the vehicle.

[0026] Specifically, the in-vehicle sound shielding method of the present embodiment can be applied to any controller with image processing function and audio processing function at the vehicle end. The image data can be processed by the vision processing module in the controller, and the audio data can be processed by the speech processing module in the controller, realizing the fusion processing of multi-modal data by the controller.

[0027] The image data can be collected by a camera arranged in the vehicle, such as a driver monitoring camera, a passenger monitoring camera, and an in-vehicle panoramic camera. The audio data can be collected by a microphone array arranged in the vehicle. The microphone array is an audio collection system composed of multiple microphones arranged in a specific geometry, which can realize sound source positioning and spatial filtering.

[0028] The image data can be used to identify the number, identity, and seating position of the passengers in the vehicle. The audio data can be used to determine the content of the conversation of the passengers and the orientation of the passengers who are talking.

[0029] In step 104, the image data and the audio data are processed to determine the directional information of the target passenger related to the current conversation content in the vehicle.

[0030] Specifically, the image data can be processed by a visual detection model built in a visual processing module to perform face detection and feature extraction on the passengers in the vehicle. The audio data can be processed by a speech processing module to analyze the speech information in the audio data and locate the sound source in the audio data. Then, the target passenger related to the current conversation content is determined as the directional information based on the face detection result, the speech information, and the sound source location.

[0031] In determining the target passenger, the location and the content of the current speaker can be used to determine the target passenger. For example, if the current speaker is the driver and the content of the conversation includes "please close the window on the co-driver's side", the target passenger is the co-driver. For another example, if the current speaker is the passenger in the back right seat and the content of the conversation includes "stop at the intersection ahead", the target passenger is the driver. The directional information includes the target passenger corresponding to the current conversation content, and the target passenger includes the current speaker and the target passenger corresponding to the current speaker. If the target passenger corresponding to the current speaker cannot be determined, the subsequent conversation privacy protection operation does not need to be performed.

[0032] In step 106, the target shielding area in the vehicle is determined based on the directional information.

[0033] Specifically, the target shielding area is an area that needs to be shielded from the current conversation content, so that the passengers in the target shielding area cannot clearly hear the specific content included in the audio data, and the conversation content between the current talking passengers is not leaked.

[0034] After the directivity information is determined, the position of the current speaker and the position of the passenger to which the current speaker points can be determined. The space composed of the position of the current speaker and the position of the passenger to which the current speaker points is the target privacy area, and the conversation content in the target privacy area needs to be protected from being leaked. Then, the area in the vehicle space other than the target privacy area is the to-be-screened area.

[0035] Step 108: shielding the audio data in the to-be-screened area by using a sound field control method.

[0036] Specifically, the sound field control method is a technology for realizing the enhancement or inhibition of sound in a specific area by precisely controlling the propagation direction, intensity, and phase of sound waves. After the to-be-screened area is determined, the audio data can be shielded in the to-be-screened area by using the sound field control method, so that the passengers in the to-be-screened area cannot clearly hear the conversation content between the current talking passengers. At the same time, the normal conversation between the current talking passengers will not be affected.

[0037] It should be noted that the method of the embodiment is transparent to the passengers in the vehicle and does not require manual operation of any device by the current talking passengers. When the vehicle controller monitors the conversation of the passengers and needs to perform privacy protection, the sound field privacy protection function is automatically activated. The activation process of the sound field privacy protection function is transparent to the passengers in the vehicle and does not affect the normal communication in the vehicle. After the conversation ends, the controller returns to the standby state and closes the sound field privacy protection function, waiting for the next trigger.

[0038] Based on the above steps 102 to 108, the vehicle sound shielding method provided by the embodiment includes collecting image data and audio data in the vehicle, processing the image data and the audio data, and determining directivity information for representing target passengers related to the current conversation content. By determining the directivity information, the passengers currently talking in the vehicle can be clearly known. According to the directivity information, a to-be-screened area in the vehicle is determined, which is an area that needs to be screened, facilitating subsequent accurate sound screening in the vehicle area. The audio data in the to-be-screened area is shielded by using a sound field control method, so that the passengers in the to-be-screened area cannot hear the conversation content of the target passengers, realizing the privacy protection of the conversation content of the target passengers and reducing the risk of privacy leakage.

[0039] The determination of the directivity information involves a multi-modal AI fusion technology. By fusion processing of the image data and the audio data, the directivity information can be accurately determined.

[0040] In some embodiments, the processing of the image data and the audio data to determine the directivity information for representing the target passengers related to the current conversation content in the vehicle includes: determine an identity label and position information of each occupant in the vehicle based on the image data; determine the directivity information based on the audio data, the identity label and the position information of each occupant.

[0041] Specifically, the identity label of the occupant is a label capable of representing the identity type of the occupant, and the identity label can include boy, girl, adult male, adult female, old male, and old female, etc. The position information can be specifically a position coordinate of the occupant in the vehicle. The trained visual detection model is used to detect the image data, so as to determine the identity label and the position information of each occupant. Exemplarily, the visual detection model can include a YOLOv5 target detection model, a ResNet-50 network, and a Support Vector Machine (SVM) classifier. YOLOv5 is a one-stage target detection model, which is known for its fast speed, high accuracy, and easy deployment. Its core idea is “You Only Look Once”, that is, only one forward propagation on the image can predict the bounding box and class of all targets. The YOLOv5 model can reach and exceed 95% accuracy after being trained on a large enough and high-quality face dataset. At the same time, YOLOv5 provides models of multiple sizes, which can be selected according to the balance requirements of speed and accuracy.

[0042] ResNet-50 is a classic deep convolutional neural network, and its core innovation is the residual block (Residual Block) and the skip connection (Skip Connection), which solves the gradient disappearance and degradation problem of deep network difficult to train. The deep structure enables it to learn very abstract and discriminative face features. Generally, the last fully connected layer is removed, and a global average pooling layer is used, so that a face image of any size can be fixedly output as a 512-dimensional feature vector.

[0043] SVM is a very powerful supervised learning model, mainly used for classification and regression. In classification, the goal is to find an optimal hyperplane that maximizes the margin between different class data points. The incremental algorithm can incorporate new data by adjusting only the support vectors without retraining the entire model, which is very suitable for the situation that the system continuously encounters new users in actual deployment. Once the training is completed, the decision function of SVM only depends on a few “support vectors”, and the inference speed is very fast.

[0044] Further, in determining the identity label and position information, all face bounding boxes in the positioning graph are located using YOLOv5, and a cropped face image is converted into a fixed-length, highly-discriminative feature vector (512-dimensional) using ResNet-50. An SVM classifier is used to determine the specific type of personnel (e.g., boy, girl, adult male, adult female, elderly male, and elderly female) based on the feature vector, and the identity label of the occupant is obtained. Based on perspective transformation, the 2D pixel coordinates of the face in the image are converted into 3D relative position coordinates in the vehicle space, and the position information of the occupant is obtained.

[0045] For the YOLOv5 target detection model, a large number of images containing faces are collected as a training set, and the bounding boxes of the faces are labeled using a labeling tool. The YOLOv5 target detection model is trained using the training set, and the model parameters are fixed after training. During model use, image data is input into the trained YOLOv5 model, and the bounding box coordinates (x1, y1, x2, y2), confidence score, and class label of each detected face are output by the YOLOv5 model. Non-maximum suppression is applied to remove duplicate and overlapping boxes, and the final face detection result is obtained.

[0046] For the ResNet-50 network, a large-scale face dataset (such as MS-Celeb-1M, VGGFace2) is used for training. The training target usually uses a loss function for metric learning, such as TripletLoss or ArcFace Loss. The core goal of these loss functions is to make the features of different faces of the same person as close as possible in the vector space, and the features of different people as far apart as possible. After training, the model that can output a 512-dimensional vector is saved. During model use, the single face image detected and cropped by YOLOv5 is scaled to the input size required by ResNet-50 (such as 224x224). The ResNet-50 network is input for forward propagation. The 512-dimensional output layer activation value is obtained, which is the feature encoding of the face.

[0047] For SVM classifier, collect face feature vectors (from ResNet-50) and corresponding labels (e.g. Zhang San - adult male, Li Si - boy) of all known person types. Use these (512-dimension vector, label) data to train a multi-class SVM model. Use specific incremental learning algorithm to iteratively update the impact of new samples to the existing support vector set and decision function. When a new class needs to be added, usually need to use "one-to-one" or "one-to-many" strategy to expand the multi-class SVM, and the incremental learning algorithm also needs to be adjusted accordingly to support the introduction of new classes. When using the model, input the 512-dimension feature vector extracted by ResNet-50 into the trained SVM model, and the SVM outputs the probability of the feature vector belonging to each identity label or directly outputs the identity label.

[0048] Using the combination of YOLOv5+ResNet-50+SVM, high-precision identity label determination can be achieved. Each model in the combination can be optimized and upgraded independently to further improve the accuracy of target detection results in image data.

[0049] Perspective transformation is a transformation that projects an image from one perspective to another. It can map the "tilted" plane in the image (such as the floor inside the car) to a "frontal" view, thus restoring its geometric relationship in the real world. Monocular camera cannot directly obtain depth information, but if the physical structure of the space inside the car is known (a known rectangular region, such as the floor between the front seats), it can be used to estimate the location using perspective transformation. First, you need to get the camera's intrinsic matrix (focal length, principal point coordinates) and distortion coefficients. In the image, manually label the image coordinates of the four corners of a rectangular region of known physical size (for example, a 2m x 1m rectangular region on the floor inside the car). Define the coordinates of the four corners of this rectangular region in the real world (for example: (0, 0), (2, 0), (2, 1), (0, 1)), units are meters. Use the cv2.getPerspectiveTransform(src_pts, dst_pts) function of OpenCV to calculate the perspective transformation matrix (H) and the inverse perspective transformation matrix (H_inv) according to the coordinates of the four corners. Assuming that the "position" of the person is the point of his standing or sitting feet or hips in the image, use the inverse perspective transformation matrix H_inv to transform this point (x, y_bottom) in the image to the real world coordinate system: (x_world, y_world) = H_inv × [x, y_bottom, 1]. The resulting (x_world, y_world) is the relative position coordinate of the occupant relative to the origin of the previously defined rectangular region.

[0050] Further, the directionality information is determined based on the audio data, the identity label of each occupant, and the position information, including: The sound source and voiceprint of the audio data are analyzed, and the identity label and the position information are combined to determine the sound source occupant information; The audio data is converted to obtain the conversation text; Based on the sound source occupant information and the conversation text, the directionality information is determined through directionality analysis.

[0051] Specifically, based on the audio data, the Beamforming technology is used for sound source analysis to calculate the direction of the target sound source in the space. The direction angle of the sound source reaching the microphone array is accurately estimated. Based on the direction of the located target sound source, a beam is formed to enhance its voice while suppressing noise and interference in other directions, achieving voice enhancement or separation. For the enhanced single-person voice segment, a Deep Speaker model is used to determine the specific sound source occupant. At the same time, for the enhanced single-person voice segment, a Transformer model is used to convert it to text content to determine the specific speaking content.

[0052] Beamforming is not a single "model", but a collection of signal processing algorithms. The core idea is to use the different positions of multiple microphones in space to receive sound signals with slight time differences or phase differences. By aligning these signals through algorithms, a "beam" pointing in a specific direction can be artificially formed, thereby enhancing the signal in that direction. Beamforming technology not only provides directional information, but its core output "aligned and enhanced signal" provides cleaner and higher signal-to-noise ratio input for subsequent sound source recognition and speech recognition, greatly improving the performance of subsequent modules.

[0053] Deep Speaker is a speaker recognition system based on deep neural networks. The core idea is to map variable-length speech segments into a fixed-length, discriminative vector (called "embedding vector" or "voiceprint vector"). This vector can be regarded as a "digital fingerprint" of the speaker's voice. The enhanced speech input by Beamforming is input into Deep Speaker, and the continuous speech is cut into short-time frames (such as 20-40ms per frame). The short-time frame is input into the network structure (usually a recurrent neural network or a convolutional neural network), and the voiceprint embedding vector is output through the network structure.

[0054] The final speaking passenger information is determined based on the calculation of the voiceprint embedding vector, the direction, identity label and position information of the target sound source, that is, which passenger in the vehicle is speaking. In specific implementation, the approximate position of the speaking passenger can be determined through the direction of the target sound source (such as the right rear seat), the voiceprint analysis result is obtained through the calculation of the voiceprint embedding vector, and the gender and age range of the speaking passenger can be determined through the voiceprint analysis result, such as determining that the speaking passenger is a boy. At this time, the preliminary determined speaking passenger information is "boy in the right rear seat". However, since the passengers may appear head swinging during conversation, the judgment may be wrong only through the sound source. For example, the actual speaking passenger is a boy in the middle of the rear seat, but the head of the boy in the middle of the rear seat swings to the right during conversation, resulting in that the determined direction of the target sound source is the right rear seat, but the actual speaking passenger is not the passenger in the right rear seat. At this time, the identity label and position information determined before are combined to accurately determine the actual speaking passenger. Through the identity label and position information, it is determined that the passenger in the right rear seat is an adult male (not a boy), and the passenger in the middle of the rear seat is a boy. Since the conclusion "the speaking passenger is a boy" determined through the voiceprint analysis result is correct, the finally determined speaking passenger is the boy in the middle of the rear seat, not the adult male in the right rear seat. Therefore, the speaking passenger information can be accurately determined by combining the sound source, voiceprint, identity label and position information.

[0055] The Transformer is a neural network architecture based entirely on the self-attention mechanism. The speech enhanced by Beamforming is converted into a speech feature sequence, which is encoded into a high-level acoustic representation sequence consistent with the context information by the encoder in the Transformer. The decoder in the Transformer outputs the conversation text.

[0056] The Beamforming technology provides high-quality input for the subsequent Deep Speaker network and Transformer network, ensuring the availability and accuracy of the embodiment method in a real noisy environment. The accuracy of determining the speaking passenger information and the conversation text is improved.

[0057] The speaking passenger information is determined, that is, who is the speaker. The conversation text is determined, that is, what is the content of the speech. Then, the directional information can be determined according to the intention and demand of the speaking passenger contained in the speech content. The directional information includes the target passenger related to the current conversation content, and the target passenger is two or more people in conversation, including the speaking passenger and the pointing passenger. The pointing passenger is the other passenger in conversation with the speaking passenger.

[0058] In this embodiment, the BERT model is used for directionality analysis. The conversation text is input into the BERT model, and the BERT model can analyze the directional words or directional intent in the conversation text. For example, the conversation text includes "open the window on the co-driver side", it can be determined that the pointing passenger corresponding to the speaking passenger is the passenger in the co-driver position. The target passenger included in the directionality information is the speaking passenger and the co-driver position passenger.

[0059] The BERT model makes accurate judgments on the directionality information by understanding natural language instructions, linking context, and integrating various sensor information, which provides a data basis for subsequent accurate determination of the to-be-screened area. In some embodiments, the to-be-screened area in the vehicle is determined according to the directionality information, including: determining a target privacy area in the vehicle according to the directionality information; determining the to-be-screened area according to the target privacy area.

[0060] Specifically, the target passenger is included in the directionality information, that is, the current conversation is a conversation between target passengers. Then, the position of the target passenger and the area between the positions of the target passengers can be determined as the target privacy area. In the target privacy area, the conversation can proceed normally. The area in the vehicle cabin other than the target privacy area is the to-be-screened area. For example, if the target passengers include the co-driver position passenger and the rear right passenger, the to-be-screened area includes the co-driver position, the rear right position, and the area between the co-driver position and the rear right position.

[0061] In the to-be-screened area, it is necessary to ensure that the audio data is screened to prevent the conversation content between the target passengers from being leaked into the to-be-screened area. In addition, if it is determined through image data and audio data that the current speaking passenger is making a call through a mobile device, rather than talking to the passengers in the vehicle, then the area other than the area where the speaking passenger is located can be regarded as the to-be-screened area. In this way, it can be avoided that the call sound disturbs other passengers in the vehicle or the call content is leaked. Through the accurate determination of the to-be-screened area in this embodiment, it can be ensured that the sound propagating in the to-be-screened area is accurately screened through the sound field control method, and the sound screening effect is improved.

[0062] In some embodiments, the audio data is screened in the to-be-screened area through the sound field control method, including: generating a reverse sound wave according to the audio data; and emitting the reverse sound wave to the to-be-screened area.

[0063] Specifically, the reverse sound wave is generated in real time in the shielding area to cancel the sound wave of the original audio data, thereby achieving a significant shielding effect in the shielding area. The shielding effect has spatial limitations, and the reverse sound wave can only cancel the sound wave of the audio data in the shielding area. Therefore, the spatial position of the shielding area needs to be accurately determined. For example, the position coordinates of the shielding area can be estimated by beamforming technology to achieve accurate positioning of the shielding area.

[0064] Then, the frequency spectrum corresponding to the audio data can be analyzed to determine the amplitude and phase of each frequency component, generate a reverse sound wave with opposite phase, and emit the reverse sound wave in the shielding area. The reverse sound wave cancels the sound wave of the audio data, achieving a noise reduction effect in the shielding area. The occupants in the shielding area cannot hear the conversation between the target occupants, achieving the purpose of preventing privacy leakage.

[0065] In some embodiments, the generating the reverse sound wave according to the audio data comprises: generating an initial reverse sound wave according to the audio data; determining the sound wave intensity according to the audio data; adjusting the intensity of the initial reverse sound wave according to the sound wave intensity to generate the reverse sound wave.

[0066] Specifically, the Fast Fourier Transform algorithm can be used to analyze the sound wave frequency of the audio data. The Fast Fourier Transform algorithm is an efficient algorithm for converting time domain signals to frequency domain signals. It decomposes complex waveforms into a series of simple sine waves. In specific implementation, continuous audio data is cut into short periods for processing. Fast Fourier Transform calculation is performed on each frame of signal to obtain the frequency spectrum of the frame. The frequency spectrum shows the amplitude and phase information of each frequency component in the range of 20Hz~20kHz.

[0067] When generating the initial reverse sound wave, the frequency of the initial reverse sound wave is the same as that of the audio data sound wave, but the phase of each frequency component is offset, and the phase of the initial reverse sound wave = the phase of the audio data sound wave + 180°. In theory, when the phase difference is equal to 180°, the audio data sound wave can be well canceled to achieve the sound shielding effect. However, when the actual phase difference is close to 10°, the sound shielding effect will be greatly weakened. Therefore, in order to ensure the effectiveness of sound cancellation, the phase difference between the phase of the initial reverse sound wave and the phase of the audio data sound wave is determined to be within the range of 180±5°, ensuring the effectiveness of the cancellation mechanism.

[0068] Since the content of the conversation between the target passengers is changing, the volume of the conversation is also changing, and the initial reverse sound wave needs to be dynamically adjusted to maintain the best shielding effect at all times. In this embodiment, the reverse sound wave intensity is adjusted in real time by an adaptive algorithm. Specifically, the sound wave intensity of the audio data is determined in real time, and the adaptive algorithm is used to adjust the intensity of the initial reverse sound wave according to the sound wave intensity of the audio data to generate the final reverse sound wave. The adaptive algorithm takes the energy of the residual sound (i.e., the sound after the sound wave cancellation) as the cost function. The goal of the adaptive algorithm is to minimize the energy of the residual sound. The adaptive algorithm can generate the amplitude and phase of the initial reverse sound wave according to the real-time and slight changes of the residual sound. If the cancellation effect is good and the residual sound is small, the algorithm maintains the current parameters. If the cancellation effect becomes worse and the residual sound becomes larger, the adaptive algorithm can quickly adjust the parameters until the optimal point that minimizes the residual sound is found again.

[0069] The reverse sound wave generation method in this embodiment can generate a reverse sound wave with good sound cancellation effect, and the adaptive algorithm can adjust the reverse sound wave in real time to maintain the sound cancellation effect at an optimal level at all times, continuously protecting the conversation content between the target passengers from being leaked.

[0070] In some embodiments, after emitting the reverse sound wave to the area to be shielded, the method comprises: collecting audio data to be detected in the area to be shielded; determining whether the audio data to be detected meets a preset shielding condition; in response to the audio data to be detected not meeting the preset shielding condition, adjusting the reverse sound wave.

[0071] Specifically, in order to further ensure the sound shielding effect, after emitting the reverse sound wave to the area to be shielded, audio data to be detected in the area to be shielded is collected. The audio data to be detected is the audio data after shielding, and by analyzing the audio data after shielding, the current sound shielding effect can be determined. If the audio data to be detected is relatively loud and exceeds a preset sound threshold, it is determined that the preset shielding condition is not met, and the reverse sound wave parameters need to be optimized and adjusted until the audio data to be detected meets the preset shielding condition. If the audio data to be detected is relatively small, it is determined that the preset shielding condition is met, and the current reverse sound wave parameters can be maintained. Through the method of this embodiment, the sound shielding effect can be detected in real time, ensuring that a good sound shielding effect is maintained during the sound shielding process, and continuously protecting the conversation content of the target passengers from being leaked.

[0072] In some embodiments, before determining the area to be shielded in the vehicle according to the directivity information, the method comprises: determining whether the privacy protection function needs to be turned on according to the directivity information or the audio data; In response to determining that the privacy protection function needs to be turned on, a region in the vehicle to be shielded is determined according to the directivity information.

[0073] Specifically, when audio data is detected in the vehicle, if it is determined through judgment that the privacy protection function does not need to be turned on, subsequent sound shielding operation is not required. For example, if the audio data is noise, noise or music sound, etc., it is determined that the privacy protection function does not need to be turned on.

[0074] After the audio data is converted into conversation text, if it is determined through semantic analysis that the conversation is not a conversation between specific passengers, but a conversation content issued by a speaking passenger to all other passengers in the vehicle, it indicates that the current conversation content does not involve privacy, and the privacy protection function can not be turned on, and the audio data does not need to be shielded. Alternatively, if the target passenger cannot be determined according to the directivity information, the privacy protection function can also not be turned on, and the audio data does not need to be shielded. If it is determined that the privacy protection function needs to be turned on, a region in the vehicle to be shielded can be determined according to the directivity information.

[0075] In addition, the vehicle can also provide a service for the user to actively turn on the privacy protection function. After the vehicle is powered on, the user is prompted through the human-computer interaction interface whether to turn on the privacy protection function for this trip. If the user's confirmation feedback is received, the privacy protection function is turned on, and if the user's denial feedback is received, the privacy protection function is not turned on. The purpose of flexible turning on of the privacy protection function is achieved.

[0076] The method of the embodiment can accurately determine whether the privacy protection function needs to be turned on, thereby avoiding the misoperation of turning on the sound shielding when it is not required, and affecting the normal communication between the passengers in the vehicle.

[0077] It should be noted that the method of the embodiments of the present application can be executed by a single device, such as a computer or a server, etc. The method of the embodiments can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present application, and the multiple devices can interact with each other to complete the method.

[0078] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order described above and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0079] It should be noted that the embodiments of the present application can also be further described in the following manner: Figure 2 A flowchart of in-vehicle sound shielding is shown. As shown in Figure 2 , after the vehicle is powered on, the controller enters a standby state, continuously monitors the in-vehicle environment, and collects in-vehicle data, including image data collection and audio data collection. When voice activity is detected, occupant identification and voice analysis are triggered. It is determined through image data whether there is an occupant in the vehicle. If it is determined that there is an occupant, it is determined through an AI model whether the privacy protection function needs to be turned on in the vehicle. The AI model can be an image data processing model and an audio data processing model. The image data processing model includes a visual detection model, and the audio data processing model includes a sound analysis algorithm, a semantic understanding and analysis model, etc. When it is determined that the privacy protection needs to be turned on, if it is identified that it is a private conversation, a sound shielding operation is immediately performed, and a reverse sound wave is generated for the audio data. A privacy area is established inside the vehicle, and the reverse sound wave is emitted in the privacy area to cancel the audio data sound wave in the privacy area, thereby realizing the privacy protection function for the privacy area and preventing the conversation content in the privacy area from being leaked. If it is determined that the privacy protection does not need to be turned on, the current process is ended, and the controller returns to the standby state, waiting for the next in-vehicle data collection to trigger. The period of in-vehicle data collection can be flexibly set according to actual conditions.

[0080] Figure 3 A flowchart of in-vehicle occupant identity confirmation and position confirmation is shown. As shown in Figure 3 , the image data is input into the visual detection model and preprocessed. Through preprocessing, the image data can be improved, unwanted distortions can be suppressed, or certain image features that are more important for subsequent processing can be enhanced. Preprocessing can include image enhancement, image denoising, geometric transformation, color processing, etc. Face detection and feature extraction operations are performed based on the preprocessed image data. It is determined whether a face is detected, and if no face is detected, the current process is ended. If a face is detected, the corresponding occupant identity label is determined. When determining the occupant identity label, the age, gender, seating position of each occupant, conversation content between occupants, and body interaction between occupants, etc. can be used to determine the identity of each occupant, such as leader, family member, driver, and friend, etc. If the identity label of each occupant cannot be determined, the current process is ended. After the identity label of each occupant is determined, the identity label is recorded. At the same time, the seating position of each occupant can also be determined according to the image data, such as the driver position, the front passenger position, the rear seat position, or the third row position, etc. After that, the seating position of each occupant is recorded. If the seating position of each occupant cannot be determined, the current process is ended.

[0081] Figure 4A flowchart of voice recognition and directivity analysis is shown. As shown in Figure 4 audio data is collected by a microphone array. The audio data is pre-processed and noise-reduced, the purpose of pre-processing and noise-reduction is to improve signal quality, suppress noise and irrelevant interference, retain or enhance the target component of interest, so as to provide cleaner and more reliable input for subsequent models or algorithms. Specifically, pre-processing includes format standardization, silence cutting and speech activity detection, pre-emphasis, framing and windowing. Format standardization is to ensure that all audio data has a uniform format and specification, reducing unnecessary variance. Silence cutting and speech activity detection are to remove the long silent sections at the beginning and end of the audio, and only keep the valid speech part, which can significantly reduce unnecessary computational load. Pre-emphasis is to enhance the high frequency part of the signal and balance the speech spectrum to make it more flat. Framing is to cut the continuous audio signal into short frames. Windowing is to reduce the truncation effect of each frame signal at the beginning and end. Noise reduction is a technique specifically used to suppress background noise and improve signal-to-noise ratio. Traditional noise reduction methods include spectral subtraction and Wiener filtering, and modern noise reduction methods based on deep learning include spectral mapping, mask learning and end-to-end noise reduction.

[0082] Then, it is detected whether there is valid speech in the audio data, and if there is valid speech, the valid speech is separated and recognized. If there is no valid speech, the current process is ended. Speaker separation is to separate mixed speech, and speaker recognition is to determine the speaking crew corresponding to the separated speech. When performing speaker separation, first determine which time points have speech and which are silent or noise, and then separate a series of speech segments and non-speech segments. In a series of speech segments, find the exact time points of the speaking crew identity change. Extract features that can represent the speaking crew identity from each short speech segment to obtain a feature vector. Group the feature vectors of all speech segments so that the segments of the same speaking crew are grouped into the same group and the segments of different speaking crews are grouped into different groups, and finally output the separated speaking crew speech streams.

[0083] In the speaker recognition process, features are extracted from the speech of each speaker, the average of these features is calculated to obtain a voiceprint embedding vector representing the identity of each speaker, and the voiceprint embedding vector is stored in a database. The voiceprint embedding vector of the voice to be verified is extracted, and the similarity between the voiceprint embedding vectors in the database is calculated. If the similarity exceeds a threshold, the speaking crew identity can be determined.

[0084] Then, the directivity analysis is used to determine the speaking crew direction according to the audio content of the speaking crew (speaking crew), and the target crew is determined. The privacy protection decision is determined according to the target crew, which can include the target privacy area where the target crew is located, and then the current process is ended.

[0085] Figure 5 A flowchart of sound field control and privacy protection is shown. As shown, it is determined whether the privacy protection function needs to be turned on. If not, the normal sound field is maintained, and the current flow is ended. The privacy protection function can be turned on by the user. For example, after the user gets on the vehicle, the vehicle end asks the user through the human-computer interaction interface whether the privacy protection function needs to be turned on during the current trip. If the user chooses to turn on, the subsequent sound shielding operation can be triggered and executed. If the user chooses not to turn on, even if audio data is collected, the sound shielding operation does not need to be executed. It can also be determined whether the privacy protection function needs to be turned on according to the text content and directivity information of the audio data. If a specific target passenger can be determined through the text content and directivity information of the audio data, the privacy protection function can be turned on. If the specific target passenger cannot be determined through the text content and directivity information of the audio data, or the target passenger includes all passengers, the privacy protection function does not need to be turned on. Figure 5

[0086] If the privacy protection function needs to be turned on, the reverse sound wave parameters are calculated according to the audio data, the reverse sound wave is generated in the to-be-shielded area according to the reverse sound wave parameters, the reverse sound wave is superimposed with the sound wave of the audio data, and the effect of sound shielding is achieved. Then, the audio data in the to-be-shielded area is re-collected, and it is determined whether the privacy protection effect is good. If the privacy protection effect is good, the current flow is ended. If the privacy protection effect is not good, the reverse sound wave parameters are re-calculated to re-generate the reverse sound wave.

[0087] In summary, the in-vehicle sound shielding method of the embodiment includes five stages, namely, a data collection stage, an intelligent identification stage, a decision analysis stage, a sound field control stage, and an effect monitoring stage. In the data collection stage, image data is collected by an in-vehicle camera, and audio data is collected by a microphone array. In the intelligent identification stage, the visual and voice data are processed in parallel by an AI model to identify the identity tag, position information, and conversation content of each member. In the decision analysis stage, the identification results are comprehensively analyzed to determine whether the privacy protection function needs to be turned on. In the sound field control stage, the reverse sound wave is calculated and generated to achieve accurate shielding of the to-be-shielded area. In the effect monitoring stage, the shielding effect of the to-be-shielded area is monitored in real time, and the reverse sound wave parameters are adjusted in a timely manner as necessary to continuously maintain a good sound shielding effect.

[0088] Based on the same inventive concept, the application also provides an in-vehicle sound shielding device corresponding to any of the above-mentioned embodiment methods.

[0089] Referring to Figure 6 , the in-vehicle sound shielding device includes: a collection module 202 configured to collect image data and audio data in the vehicle; ​The processing module 204 is configured to process the image data and the audio data, and determine directional information for characterizing a target occupant related to current in-vehicle conversation content. The determining module 206 is configured to determine a to-be-masked area in the vehicle according to the directional information. The masking module 208 is configured to mask the audio data in the to-be-masked area by a sound field control method.

[0090] In some embodiments, the processing module 204 is further configured to determine an identity tag and position information of each occupant in the vehicle based on the image data, and determine the directional information based on the audio data, the identity tag and the position information of each occupant.

[0091] In some embodiments, the processing module 204 is further configured to determine sounder occupant information by performing sound source analysis and voiceprint analysis on the audio data, in combination with the identity tag and the position information. The audio data is format-converted to obtain conversation text. The directional information is determined by directional analysis based on the sounder occupant information and the conversation text.

[0092] In some embodiments, the determining module 206 is further configured to determine a target privacy area in the vehicle according to the directional information, and determine the to-be-masked area according to the target privacy area.

[0093] In some embodiments, the masking module 208 is configured to generate reverse sound waves according to the audio data, and emit the reverse sound waves to the to-be-masked area.

[0094] In some embodiments, the masking module 208 is further configured to generate initial reverse sound waves according to the audio data, determine sound wave intensity according to the audio data, and adjust the intensity of the initial reverse sound waves to generate the reverse sound waves according to the sound wave intensity.

[0095] In some embodiments, after emitting the reverse sound waves to the to-be-masked area, the method further comprises a detecting module configured to collect to-be-detected audio data of the to-be-masked area, determine whether the to-be-detected audio data meets a preset masking condition, and adjust the reverse sound waves in response to the to-be-detected audio data not meeting the preset masking condition.

[0096] In some embodiments, before determining the to-be-masked area in the vehicle according to the directional information, the determining module 206 is further configured to determine whether a privacy protection function needs to be turned on according to the directional information or the audio data, and determine the to-be-masked area in the vehicle according to the directional information in response to determining that the privacy protection function needs to be turned on.

[0097] For the convenience of description, the above apparatus is described in various modules in terms of functions. Of course, in the implementation of the present application, the functions of the modules can be implemented in one or more software and / or hardware.

[0098] The apparatus of the above embodiments is used to implement the corresponding in-vehicle sound shielding method of any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.

[0099] Based on the same inventive concept, the present application also provides an electronic device corresponding to the method of any of the above embodiments, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the in-vehicle sound shielding method of any of the above embodiments.

[0100] Figure 7 A more specific hardware structure of an electronic device provided by the present embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.

[0101] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the present embodiment.

[0102] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the present embodiment are implemented by software or firmware, the related program codes are stored in the memory 1020 and executed by the processor 1010.

[0103] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0104] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as a USB, a network cable, etc.) or a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.).

[0105] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.

[0106] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0107] The electronic device of the above embodiments is used to implement the corresponding in-vehicle sound shielding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0108] Based on the same inventive concept, corresponding to any of the above embodiment methods, the present application also provides a non-transitory computer readable storage medium storing computer instructions for causing the computer to execute the in-vehicle sound shielding method according to any of the above embodiments.

[0109] The computer readable medium of the embodiments can include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0110] The storage medium of the above embodiments stores computer instructions for causing the computer to perform the in-vehicle sound shielding method according to any one of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.

[0111] It can be understood that before using the technical solutions of various embodiments in the present disclosure, the type, use range, use scenario, etc. of the personal information involved will be informed to the user in an appropriate manner, and the authorization of the user will be obtained.

[0112] For example, in response to receiving the user's active request, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the technical solutions of the present disclosure according to the prompt information.

[0113] As an optional but not limited implementation manner, in response to accepting the user's active request, the way of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide personal information by the electronic device.

[0114] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0115] Those of ordinary skill in the art will realize that the foregoing discussion of any of the embodiments has been presented for the purpose of illustration and description and is not intended to be exhaustive or to limit the application to the precise forms described, and that various adaptations and modifications are possible within the scope and spirit of the application. For example, while the embodiments discussed above have been described in the context of a memory device, the embodiments discussed above can be used in other memory architectures, such as dynamic RAM (DRAM).

[0116] In addition, to simplify the description and discussion, and so as not to make the embodiments of the application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. Further, devices can be shown in block diagram form in order to avoid making the embodiments of the application difficult to understand, and this also takes into account the fact that the details regarding implementation of these block diagram devices are highly dependent on the platform in which the embodiments of the application are to be implemented (i.e., these details should be well within the understanding of one of ordinary skill in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the application, it should be apparent to one of ordinary skill in the art that the embodiments of the application can be practiced without or with variations of these specific details. Thus, the description should not be considered to be limiting in nature.

[0117] While the application has been described in connection with specific embodiments thereof, it will be understood that many modifications, variations and alternatives will be apparent to those skilled in the art as a result of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0118] The embodiments of the application are intended to cover all such modifications, variations and alternatives as come within the scope of the appended claims. Thus, any and all such modifications, variations and alternatives are intended to be included within the scope of the application.

Claims

1. A method for sound shielding inside a vehicle, characterized in that, include: Collect image and audio data from inside the vehicle; The image data and audio data are processed to determine directional information that characterizes the target occupant related to the current in-vehicle conversation. The area inside the vehicle to be shielded is determined based on the directional information; The audio data is shielded within the area to be shielded using a sound field control method.

2. The method according to claim 1, characterized in that, The process of processing the image data and the audio data to determine directional information characterizing the target occupant related to the current in-vehicle conversation includes: Based on the image data, the identification tag and location information of each passenger in the vehicle are determined; The directional information is determined based on the audio data, the identity tag of each occupant, and the location information.

3. The method according to claim 2, characterized in that, The determination of the directional information based on the audio data, each occupant's identification tag, and location information includes: By performing sound source analysis and voiceprint analysis on the audio data, and combining the identity tag and the location information, the information of the passenger who made the sound can be determined. The audio data is converted to a different format to obtain the conversation text. Based on the voice occupant information and the conversation text, the directional information is determined through directional analysis.

4. The method according to claim 1, characterized in that, Determining the area to be shielded inside the vehicle based on the directional information includes: The target privacy area inside the vehicle is determined based on the directional information; The area to be blocked is determined based on the target privacy area.

5. The method according to claim 1, characterized in that, The method of shielding the audio data in the area to be shielded by sound field control includes: Generate an inverse sound wave based on the audio data; The reverse sound wave is emitted toward the area to be shielded.

6. The method according to claim 5, characterized in that, The step of generating a reverse sound wave based on the audio data includes: An initial reverse sound wave is generated based on the audio data; The sound wave intensity is determined based on the audio data; The intensity of the initial reverse sound wave is adjusted according to the intensity of the sound wave to generate the reverse sound wave.

7. The method according to claim 5, characterized in that, After transmitting the reverse acoustic wave toward the area to be shielded, the process includes: Collect the audio data to be detected from the area to be shielded; Determine whether the audio data to be detected meets the preset shielding conditions; In response to the fact that the audio data to be detected does not meet the preset shielding conditions, the reverse sound wave is adjusted.

8. The method according to claim 1, characterized in that, Before determining the area to be shielded inside the vehicle based on the directional information, the process includes: Based on the directional information or the audio data, determine whether the privacy protection function needs to be enabled; In response to the determination that privacy protection needs to be enabled, the area inside the vehicle to be shielded is determined based on the directional information.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

10. A vehicle, characterized in that, The vehicle includes the electronic equipment as described in claim 9.

Citation Information

Patent Citations

  • An in-vehicle privacy call method and system

    CN109862472A

  • In-vehicle voice privacy protection method, system and device and storage medium

    CN118197295A

  • Vehicle control method, device, equipment and medium

    CN119811396A

  • Sound processing method and device applied in vehicle cabin, equipment and medium

    CN120110590A

  • Voice masking method, device and system and vehicle

    CN120513476A