A method and system for controlling car audio
By extracting and fusion of voice data and body infrared image data of users in the car, generating control instructions for car audio, and correcting them according to real-time volume, the problem of traditional car audio systems distracting drivers' attention and inability to accurately obtain sound outside the car is solved, achieving higher audio control accuracy and driving safety.
Patent Information
- Application Number
- CN202411254091.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-09-06
AI Technical Summary
Traditional car audio systems mainly rely on the driver's manual control to distract the driver's attention, and when the volume is too high, it is impossible to accurately obtain the sound information outside the car, resulting in safety hazards.
By obtaining the voice data and body infrared image data of the current user in the car, extracting the voice data features and body infrared image data features, and performing feature fusion features to generate fusion features for identification, and obtaining the control instructions for car audio. At the same time, the correction coefficient of the car audio is obtained based on the real-time volume outside and inside the car, and the control instructions are corrected.
On the premise of ensuring safe driving, improve the accuracy of users controlling car audio, ensure that the driver can accurately obtain important sound information from outside the car, and improve driving safety.
Smart Images

Figure CN119132300B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and particularly to a method and system for controlling automotive audio. Background Art
[0002] With the continuous development of China's economy and the continuous improvement of people's living standards, people's demand for cars has changed from a simple means of transportation to a means of transportation that meets personalized and comfortable needs. Against this background, automotive audio systems have emerged and gradually become an important part of cars. However, traditional automotive audio systems mainly rely on manual control by the driver, which will distract the driver's attention and pose unnecessary risks. In addition, users need to obtain sound information from outside the vehicle during driving, and this sound information is of great significance for improving driving safety and assisting in driving decisions. However, too high a volume of the vehicle's audio will cause users to be unable to accurately obtain sound information from outside the vehicle, creating a safety hazard.
[0003] The current technology has defects and urgently needs improvement. Summary of the Invention
[0004] One object of this application is to provide a method for controlling automotive audio, at least to solve the problem of how to improve the accuracy of users' control of automotive audio on the premise of ensuring safe driving.
[0005] To achieve the above object, some embodiments of this application provide the following aspects:
[0006] In a first aspect, some embodiments of this application further provide a method for controlling automotive audio, the method comprising:
[0007] Obtain the voice data and body infrared image data of the user currently in the vehicle;
[0008] Based on the voice data, obtain voice data features, and based on the body infrared image data, obtain body infrared image data features;
[0009] Fuse the voice data features and the body infrared image data features according to a confidence ratio to obtain fused features; wherein, the confidence ratio is determined according to the volume score of the voice data and the quality score of the body infrared image data;
[0010] Identify the fused features to obtain a control instruction for automotive audio;
[0011] Obtain the real-time volume outside the vehicle and the real-time volume inside the vehicle;
[0012] Based on the real-time sound volume outside the vehicle and the real-time volume inside the vehicle, obtain a correction coefficient for automotive audio;
[0013] Modify the control instruction of the vehicle audio based on the correction coefficient of the vehicle audio to obtain the corrected control instruction of the vehicle audio, and send the corrected control instruction of the vehicle audio to the vehicle end for control.
[0014] In a second aspect, some embodiments of the present application further provide a control system for vehicle audio, the system includes:
[0015] A voice data and body infrared image data module, which is configured to obtain the voice data and body infrared image data of the user in the vehicle currently;
[0016] A feature module, which is configured to obtain voice data features based on the voice data, and obtain body infrared image data features based on the body infrared image data;
[0017] A feature fusion module, which is configured to fuse the voice data features and the body infrared image data features according to the confidence ratio to obtain fused features; wherein, the confidence ratio is determined according to the volume score of the voice data and the quality score of the body infrared image data;
[0018] A fused feature recognition module, which is configured to recognize the fused features to obtain the control instruction of the vehicle audio;
[0019] A module for obtaining the real-time sound volume inside and outside the vehicle, which is configured to obtain the real-time volume outside the vehicle and the real-time volume inside the vehicle;
[0020] A module for obtaining the correction coefficient of vehicle audio, which is configured to obtain the correction coefficient of vehicle audio based on the real-time sound volume outside the vehicle and the real-time volume inside the vehicle;
[0021] A vehicle audio control module, which is configured to modify the control instruction of the vehicle audio based on the correction coefficient of the vehicle audio to obtain the corrected control instruction of the vehicle audio, and send the corrected control instruction of the vehicle audio to the vehicle end for control.
[0022] In a third aspect, some embodiments of the present application further provide an electronic device, the electronic device includes: one or more processors; and a memory storing computer program instructions, the computer program instructions when executed cause the processor to execute the steps of the method as described above.
[0023] In a fourth aspect, some embodiments of the present application further provide a computer-readable medium, on which computer program instructions are stored, the computer program instructions can be executed by a processor to implement the method as described above.
[0024] Fifth aspect, some embodiments of the present application further provide a computer program product, including computer programs / instructions, which when executed by a processor implement the steps of the method as described above.
[0025] Compared with the related art, in the solution provided by the embodiment of the present application, voice data and limb infrared image data of the current user in the vehicle are acquired; based on the voice data, voice data features are acquired, and based on the limb infrared image data, limb infrared image data features are acquired; the voice data features and the limb infrared image data features are feature - fused according to a confidence ratio to obtain fused features; wherein, the confidence ratio is determined according to the volume score of the voice data and the quality score of the limb infrared image data, the fused features are identified to obtain a control instruction for the car audio; the real - time volume outside the vehicle and the real - time volume inside the vehicle are acquired; based on the real - time volume outside the vehicle and the real - time volume inside the vehicle, a correction coefficient for the car audio is acquired; the control instruction for the car audio is corrected based on the correction coefficient for the car audio to obtain a corrected control instruction for the car audio, and the corrected control instruction for the car audio is sent to the vehicle terminal for control. Through the above configuration method, the present application first acquires the voice data and limb infrared image data of the current user in the vehicle, then based on the voice data, voice data features are acquired, and based on the limb infrared image data, limb infrared image data features are acquired, and the voice data features and the limb infrared image data features are feature - fused according to a confidence ratio. Since the confidence ratio is determined according to the volume score of the voice data and the quality score of the limb infrared image data, when the quality score of the limb infrared image data is low, through feature fusion with confidence, the proportion of the quality of the limb infrared image data in the fused features can be reduced to try to ensure the true intention of the driver. Since the in - vehicle audio sensor (is set on the side close to the driver's seat), when the volume score of the user's voice data is low, through feature fusion with confidence, the proportion of the volume of the user's voice data in the fused features can be reduced to try to ensure the true intention of the driver, so as to improve the user's driving experience. Since the fused features (multi - modal features) are composed of the voice data features and the limb infrared image data features, compared with single - modal features, they can better express the driver's thoughts, improve the control accuracy and intelligence of the car audio. The fused features are identified to obtain a control instruction for the car audio. Since when the user is driving a vehicle, sound information needs to be acquired from outside the vehicle, and this sound information is crucial for driving safety. Therefore, the present application calculates a correction coefficient for the car audio according to the acquired real - time volume outside the vehicle and the real - time volume inside the vehicle, and corrects the control instruction for the car audio based on the correction coefficient for the car audio to obtain a corrected control instruction for the car audio, and sends the corrected control instruction for the car audio to the vehicle terminal for control to ensure that the user can acquire sound information from outside the vehicle during the process of driving the vehicle. Therefore, the present application can improve the accuracy of the user's control of the car audio on the premise of ensuring safe driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.
[0027] Figure 1 is a schematic diagram of the main steps of a control method for automotive audio according to an embodiment of the present application;
[0028] Figure 2 is a schematic diagram of the main component structure of a control system for automotive audio according to an embodiment of the application;
[0029] Figure 3 is an exemplary structural diagram of an electronic device according to the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0031] Refer to the attached Figure 1 , Figure 1 is a schematic diagram of the main steps of a method according to an embodiment of the present application. As Figure 1 shown, the control method for automotive audio in the embodiments of the present application mainly includes the following steps S101 - step S107.
[0032] Step S101: Obtain the voice data and body infrared image data of the user in the current vehicle.
[0033] Step S102: Based on the voice data, obtain voice data features, and based on the body infrared image data, obtain body infrared image data features.
[0034] Step S103: Fuse the voice data features and the body infrared image data features according to the confidence ratio to obtain fused features. Among them, the confidence ratio is determined according to the volume score of the voice data and the quality score of the body infrared image data.
[0035] Step S104: Identify the fused features to obtain a control instruction for automotive audio.
[0036] Step S105: Obtain the real-time volume outside the vehicle and the real-time volume inside the vehicle.
[0037] Step S106: Based on the real-time sound volume outside the vehicle and the real-time volume inside the vehicle, obtain a correction coefficient for the vehicle audio.
[0038] Step S107: Correct the control instruction of the vehicle audio based on the correction coefficient of the vehicle audio to obtain a corrected control instruction of the vehicle audio, and send the corrected control instruction of the vehicle audio to the vehicle end for control.
[0039] Among them, the audio sensor for obtaining the voice data of the current user inside the vehicle is arranged on the driver's side.
[0040] In this embodiment, the voice data of the current user inside the vehicle can be collected by an audio sensor inside the vehicle (arranged on the side close to the driver's seat), and the body infrared image data of the current user inside the vehicle can be collected by a body infrared sensor inside the vehicle. Then, feature extraction is performed on the voice data to obtain voice data features, feature extraction is performed on the body infrared image data to obtain body infrared image data features, and the voice data features and the body infrared image data features are feature-fused according to the confidence ratio to obtain fused features. Among them, the confidence ratio can be obtained through formula (1).
[0041]
[0042] Among them, is the confidence ratio, β is the quality score of the body infrared image data, and those skilled in the art can customize the quality score of the body infrared image data (between 10 and 100) according to the actual usage situation or obtain the quality score of the body infrared image data by inputting the body infrared image data into a reliable scoring model. The volume score of the user's voice data can be set as follows: 0 - 30 decibels is 10 points, 40 - 80 decibels is 70 points, and above 80 decibels is 100 points. When the quality of the limb infrared image data is low, feature fusion can be performed through confidence to reduce the proportion of low-quality limb infrared image data in the fused features, so as to ensure the accuracy of the final detection result. Since the in-vehicle audio sensor is set on the side close to the driver's seat, when other passengers farther away from the audio sensor and the driver closest to the audio sensor emit sounds of the same decibel at the same time, the audio sensor will detect that the volume of the driver is louder than that of other passengers. When the volume score of the user's voice data is low, feature fusion can be performed through confidence to reduce the proportion of the volume of the user's voice data in the fused features, so as to try to ensure the true intention of the driver and ensure the accuracy of the final detection result. The fused features are input into a preset deep learning model to obtain control instructions for the car audio. The control instructions for the car audio include switching audio, increasing volume, and decreasing volume. However, since the user really needs to obtain sound information from outside the vehicle during the driving process, this sound information is crucial for driving safety. Such as traffic signal sounds: such as siren sounds, ambulance sounds, etc., these sounds can remind the driver to pay attention to special vehicles or emergencies and avoid them in time. Vehicle sounds: the horn sounds, braking sounds, etc. of other vehicles, these sounds can convey the driving intentions or warning information of other vehicles and help the driver judge the dynamics of surrounding vehicles. Therefore, in this application, an audio sensor is set outside the vehicle to obtain the real-time volume outside the vehicle, and an in-vehicle audio sensor is used to obtain the real-time volume inside the vehicle. Then, according to the real-time sound volume outside the vehicle and the real-time volume inside the vehicle, a correction coefficient for the car audio can be obtained. The correction coefficient for the car audio can be obtained through Formula 2:
[0043]
[0044] Among them, X is the correction coefficient of the car audio, f(W) is the real-time volume outside the vehicle, f(N) is the real-time volume inside the vehicle. Both f(W) and f(N) are the average volume within a period of time, which can better reflect the overall size of the volume and thus ensure the correction accuracy. a represents the start time of collecting audio data, b represents the end time of collecting audio data, and K is the sound insulation coefficient of the car. Since the models of vehicles are different, their sound insulation coefficients are also different. Those skilled in the art can customize the sound insulation coefficient of the car according to the actual use situation. Generally, the sound insulation coefficient of a vehicle can be set to 1.1. The correction coefficient of the car audio only corrects the volume adjustment in the control instructions of the car audio. Specifically, when X is less than 0.8, the increase volume instruction in the control instructions of the car audio does not take effect, and other instructions take effect, so as to control the car volume within a reasonable range, so as to ensure that the driver can accurately obtain information from outside the vehicle and thus ensure driving safety. Finally, the corrected control instructions for the car audio are sent to the car terminal for control.
[0045] In one embodiment, step S102 may further include steps S1021 to S1022:
[0046] Step S1021: Send the voice data to the trained audio extraction model to obtain the voice data features.
[0047] Step S1022: Send the limb infrared image data to the trained human body infrared data extraction model to obtain the limb infrared image data features.
[0048] In one embodiment, the trained human body infrared data extraction model is obtained according to the following steps S10221 to S10245:
[0049] Step S10221: Obtain historical limb infrared image data; wherein, the historical limb infrared image data includes historical single-person limb infrared image data and historical multi-person limb infrared image data;
[0050] Step S10222: Perform data annotation on the historical single-person limb infrared image data to obtain the annotated historical single-person limb infrared image data;
[0051] Step S10223: Perform data annotation on the limb infrared image data of the driver in the historical multi-person limb infrared image data to obtain the annotated historical multi-person limb infrared image data;
[0052] Step S10224: Based on the annotated historical single-person limb infrared image data and the annotated historical multi-person limb infrared image data, obtain a historical sample data set;
[0053] Step S10225: Based on the historical sample data set, train the initial human body infrared data extraction model to obtain the trained human body infrared data extraction model.
[0054] In this embodiment, historical limb infrared image data is obtained. Among them, the historical limb infrared image data includes historical single-person limb infrared image data and historical multi-person limb infrared image data. The historical single-person limb infrared image data is data-labeled to obtain the labeled historical single-person limb infrared image data. For example, the left tilt of the user's limb is labeled as reducing the volume, and the right tilt of the user's limb is labeled as increasing the volume. The limb infrared image data of the driver in the historical multi-person limb infrared image data is data-labeled to obtain the labeled historical multi-person limb infrared image data. According to the labeled historical single-person limb infrared image data and the labeled historical multi-person limb infrared image data, a historical sample dataset is obtained. Then, based on the historical sample dataset, an initial human infrared data extraction model is trained to obtain the trained human infrared data extraction model.
[0055] In one embodiment, the human infrared data extraction model can adopt a Convolutional Neural Network (CNN for short). It is a deep learning model that performs well in image processing and computer vision tasks. By mimicking the working mode of the human visual system, CNN can automatically learn the features in images.
[0056] In one embodiment, after step S10221, step S102216 may further be included:
[0057] Step S102216: Perform data augmentation processing on the historical limb infrared image data.
[0058] In this embodiment, it can be achieved through Rotation: Rotate the image by a certain angle, such as 90 degrees, 180 degrees, or 270 degrees. Flipping: Flip the image horizontally or vertically.
[0059] Scaling: Changing the size of the image. Cropping: Randomly selecting a part of the image. Translation: Moving the image in the x or y direction. Color Variation: Adjusting the brightness, contrast, saturation, and hue of the image. Noise Injection: Adding noise to the image, such as Gaussian noise, salt-and-pepper noise, etc. Filtering: Applying different filters, such as Gaussian blur, sharpening, edge detection, etc. Occlusion: Adding occluders to the image to simulate occlusion situations. Affine Transformation: Performing a combined transformation of tilting, scaling, and rotating on the image. Perform data augmentation on the historical limb infrared image data to increase the richness of the dataset and thereby improve the accuracy of the model.
[0060] In one embodiment, after step S101, step S108 is further included:
[0061] Step S108: Perform first preprocessing on the limb infrared image data, and use the preprocessed limb infrared image data as the input data of the trained human infrared data extraction model.
[0062] In this embodiment, the first preprocessing may include image cropping processing, image flipping processing, and noise removal processing. Use the preprocessed limb infrared image data as the input data of the trained human infrared data extraction model to improve the recognition accuracy of the final model.
[0063] In one embodiment, after step S101, step S109 is further included:
[0064] Step S109: Perform second preprocessing on the speech data, and use the preprocessed speech data as the input data of the trained audio extraction model.
[0065] In this embodiment, the second preprocessing may include denoising processing. Use the preprocessed speech data as the input data of the trained audio extraction model to improve the recognition accuracy of the final model.
[0066] It should also be noted that the training of the audio feature extraction model is an iterative process, which generally follows the following steps: Data preprocessing: First, preprocess the original audio data, which includes steps such as denoising, normalization, and segmentation, to facilitate the model's better learning and feature extraction. Data augmentation: To improve the generalization ability of the model, the preprocessed data can be augmented through various techniques such as time stretching, pitch shifting, adding noise, etc. Feature extraction: Input the augmented preprocessed data into the audio feature extraction model, and the model will attempt to extract features that are helpful for subsequent tasks (such as speech recognition, etc.). Loss calculation: According to the extracted features and the corresponding labels (in this scenario, the labels for the audio data of the main speaker), calculate the value of the loss function to evaluate the performance of the model. Backpropagation and parameter update: Use the loss value for backpropagation, calculate the gradient of each parameter, and update the parameters of the model according to the gradient to reduce the loss. Convergence judgment: After each iteration, check whether the model has converged. This is usually judged by monitoring whether the value of the loss function no longer significantly decreases within a certain threshold. Model saving: If the model has converged, save the trained audio feature extraction model; if it has not converged, continue with the next round of iterative training. End of training: When the model converges or reaches the preset number of iterations, end the training process.
[0067] It should also be noted that the audio feature extraction model can adopt a Recurrent Neural Network (RNN), which is a neural network suitable for processing sequential data. It has a memory function and can capture the temporal dynamic characteristics of the input data. RNN is very useful in processing data with sequential characteristics such as text, speech, and time series.
[0068] Based on the above steps S101 - S107, obtain the voice data and body infrared image data of the user in the current vehicle; based on the voice data, obtain voice data features, and based on the body infrared image data, obtain body infrared image data features; fuse the voice data features and the body infrared image data features according to the confidence ratio to obtain fused features; wherein, the confidence ratio is determined according to the volume score of the voice data and the quality score of the body infrared image data, identify the fused features to obtain a control instruction for the car audio; obtain the real - time volume outside the vehicle and the real - time volume inside the vehicle; based on the real - time volume outside the vehicle and the real - time volume inside the vehicle, obtain a correction coefficient for the car audio; correct the control instruction for the car audio based on the correction coefficient for the car audio to obtain a corrected control instruction for the car audio, and send the corrected control instruction for the car audio to the vehicle end for control. Through the above configuration method, the present application first obtains the voice data and body infrared image data of the user in the current vehicle, then based on the voice data, obtains voice data features, based on the body infrared image data, obtains body infrared image data features, and fuses the voice data features and the body infrared image data features according to the confidence ratio. Since the confidence ratio is determined according to the volume score of the voice data and the quality score of the body infrared image data, when the quality score of the body infrared image data is low, through feature fusion with confidence, the proportion of the quality of the body infrared image data in the fused features can be reduced to try to ensure the true intention of the driver. Since the audio sensor in the vehicle (is set on one side close to the driver's seat), when the volume score of the user's voice data is low, through feature fusion with confidence, the proportion of the volume of the user's voice data in the fused features can be reduced to try to ensure the true intention of the driver, so as to improve the user's driving experience. Since the fused features (multi - modal features) are composed of the voice data features and the body infrared image data features, compared with single - modal features, they can better express the driver's thoughts, improve the control accuracy and intelligence of the car audio. Identify the fused features to obtain a control instruction for the car audio. Since when the user is driving a vehicle, they need to obtain sound information from outside the vehicle, and this sound information is crucial for driving safety. Therefore, the present application calculates a correction coefficient for the car audio based on the obtained real - time volume outside the vehicle and the real - time volume inside the vehicle, corrects the control instruction for the car audio based on the correction coefficient for the car audio to obtain a corrected control instruction for the car audio, and sends the corrected control instruction for the car audio to the vehicle end for control to ensure that the user can obtain sound information from outside the vehicle during the process of driving the vehicle. Therefore, the present application can improve the accuracy of the user's control of the car audio on the premise of ensuring safe driving.
[0069] The step divisions of the above various methods are only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationships are included, they are within the protection scope of this patent. Adding insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core designs of the algorithm and process, are within the protection scope of this patent.
[0070] Refer to the appendix Figure 2 , Figure 2 which is a schematic diagram of the main composition structure of the control system of an automotive audio according to an embodiment of the application. As Figure 2 shown, the control system of the automotive audio in the embodiment of this application mainly includes a voice data and body infrared image data module 11, a feature module 12, a feature fusion module 13, a fusion feature recognition module 14, an in-vehicle and out-of-vehicle real-time sound volume acquisition module 15, a correction coefficient acquisition module 16 for automotive audio, and an automotive audio control module 17. In some embodiments, one or more of the voice data and body infrared image data module 11, the feature module 12, the feature fusion module 13, the fusion feature recognition module 14, the in-vehicle and out-of-vehicle real-time sound volume acquisition module 15, the correction coefficient acquisition module 16 for automotive audio, and the automotive audio control module 17 can be combined into one module. In some embodiments, the voice data and body infrared image data module 11 is configured to acquire the voice data and body infrared image data of the user in the current vehicle; the feature module 12 is configured to acquire voice data features based on the voice data and body infrared image data features based on the body infrared image data; the feature fusion module 13 is configured to perform feature fusion on the voice data features and the body infrared image data features according to the confidence ratio to obtain fusion features; the fusion feature recognition module 14 is configured to identify the fusion features to obtain a control instruction for the automotive audio; the in-vehicle and out-of-vehicle real-time sound volume acquisition module 15 is configured to acquire the real-time volume outside the vehicle and the real-time volume inside the vehicle; the correction coefficient acquisition module 16 for automotive audio is configured to acquire a correction coefficient for the automotive audio based on the real-time sound volume outside the vehicle and the real-time volume inside the vehicle; the automotive audio control module 17 is configured to correct the control instruction for the automotive audio based on the correction coefficient for the automotive audio to obtain a corrected control instruction for the automotive audio, and send the corrected control instruction for the automotive audio to the vehicle end for control.
[0071] It is worth mentioning that all the modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0072] In addition, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.
[0073] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processor to perform the steps of the method provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. As Figure 3 shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be mounted on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as, as a server array, a set of blade servers, or a multi-processor system). Among them, the components, their connections and relationships, and their functions shown herein are only examples and are not intended to limit the implementation of this application described and / or claimed herein.
[0074] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected through a bus or other means, Figure 3 taking connection through the bus as an example.
[0075] The input device 1103 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device. Examples of input devices include touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 1104 can include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors), etc. The display device can include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device can be a touchscreen.
[0076] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and input from the user can be received in any form (including acoustic input, voice input, or haptic input).
[0077] In the embodiments of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by a processor, the steps of the method provided in any one or more of the above embodiments are implemented. The computer-readable medium can be included in the electronic device described in the above embodiments; or it can exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.
[0078] The memory 1102 can be used as a non-transitory computer-readable storage medium for storing non-transitory software programs, non-transitory computer-executable programs, and modules. By running the non-transitory software programs, instructions, and modules stored in the memory 1102, the processor 1101 executes various functional applications and data processing of the server to implement the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.
[0079] The memory 1102 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include memories remotely disposed relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0080] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0081] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0082] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0083] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of this application can be implemented by hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0084] The computer program product provided by the embodiments of this application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they entirely or partially generate the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available media can be magnetic media (for example, floppy disks, hard disks, magnetic tapes), optical media (for example, DVDs), or semiconductor media (for example, solid state disks (SSDs)), etc.
[0085] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0086] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference numerals in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices recited in the apparatus claims may also be implemented by one element or device through software or hardware. The terms "first", "second", etc. are only used for descriptive distinction and do not represent any specific order, nor can they be construed as indicating or implying relative importance.
[0087] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily mention changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A method for controlling car audio, characterized in that: The method comprises: Acquire the voice data and body infrared image data of the user currently in the car; Based on the voice data, voice data features are acquired, and based on the limb infrared image data, limb infrared image data features are acquired; The voice data features and the limb infrared image data features are feature-fused according to a confidence ratio to obtain a fused feature; wherein the confidence ratio is determined according to a volume score of the voice data and a quality score of the limb infrared image data; Identifying the fusion features to obtain a control instruction for the car audio; Get the real-time volume outside the car and the real-time volume inside the car; Based on the real-time sound volume outside the vehicle and the real-time sound volume inside the vehicle, obtaining a correction coefficient for the vehicle audio; Correcting the control instruction of the car audio based on the correction coefficient of the car audio to obtain a corrected control instruction of the car audio, and sending the corrected control instruction of the car audio to the car end for control; The audio sensor for acquiring the voice data of the user currently in the vehicle is arranged on the driver's side.
2. The method for controlling the car audio according to claim 1, characterized in that: The method of acquiring voice data features based on the voice data and acquiring limb infrared image data features based on the limb infrared image data includes: Sending the voice data to the trained audio extraction model to obtain the voice data features; The limb infrared image data is sent to a trained human infrared data extraction model to obtain the limb infrared image data features.
3. The method for controlling the car audio according to claim 2, characterized in that: The method further comprises obtaining the trained human infrared data extraction model according to the following steps: Acquire historical limb infrared image data; wherein the historical limb infrared image data includes historical single-person limb infrared image data and historical multi-person limb infrared image data; Performing data labeling on the historical infrared image data of a single person's limbs to obtain the labeled historical infrared image data of a single person's limbs; Performing data labeling on the limb infrared image data of the main driver in the historical limb infrared image data of multiple people to obtain the labeled historical limb infrared image data of multiple people; Based on the annotated historical infrared image data of a single person's limbs and the annotated historical infrared image data of multiple people's limbs, a historical sample data set is acquired; Based on the historical sample data set, the initial human infrared data extraction model is trained to obtain the trained human infrared data extraction model.
4. The method for controlling the car audio according to claim 3, characterized in that: After acquiring the historical limb infrared image data, the method further includes: Data enhancement processing is performed on the historical limb infrared image data.
5. The method for controlling the car audio according to claim 2, characterized in that: After acquiring the voice data and body infrared image data of the user currently in the car, the method further includes: The limb infrared image data is subjected to a first preprocessing, and the preprocessed limb infrared image data is used as input data of the trained human infrared data extraction model.
6. The method for controlling the car audio system according to claim 2, characterized in that: After acquiring the voice data and body infrared image data of the user currently in the car, the method further includes: The speech data is subjected to a second preprocessing, and the preprocessed speech data is used as input data of the trained audio extraction model.
7. A car audio control system, characterized in that: The system comprises: A voice data and body infrared image data module, which is configured to obtain voice data and body infrared image data of a user currently in the car; A feature module, configured to obtain voice data features based on the voice data, and to obtain limb infrared image data features based on the limb infrared image data; A feature fusion module, configured to fuse the voice data features and the limb infrared image data features according to a confidence ratio to obtain a fused feature; wherein the confidence ratio is determined according to a volume score of the voice data and a quality score of the limb infrared image data; A fusion feature recognition module, which is configured to recognize the fusion feature and obtain a control instruction for the car audio; A real-time sound volume acquisition module inside and outside the vehicle, which is configured to acquire the real-time volume outside the vehicle and the real-time volume inside the vehicle; A correction coefficient acquisition module for car audio, configured to acquire a correction coefficient for car audio based on the real-time sound volume outside the car and the real-time sound volume inside the car; The car audio control module is configured to correct the car audio control instruction based on the car audio correction coefficient, obtain the corrected car audio control instruction, and send the corrected car audio control instruction to the car end for control.
8. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as claimed in any one of claims 1 to 6.
9. A computer readable medium having a computer program / instructions stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
In-vehicle dangerous scene recognition method for online car hailing
CN111091044A
Whole vehicle sound source volume self-adaptive adjustment method and device
CN112672255A