Sound effect generation methods, equipment, vehicles, storage media and products
By collecting in-vehicle scene information and using visual language models and AIGC models to generate sound effects that are adapted to the in-vehicle scene, the problem of not being able to dynamically adjust in existing technologies is solved, thereby enhancing the user's immersion and driving experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-07-03
- Publication Date
- 2026-07-31
AI Technical Summary
Existing in-vehicle audio solutions cannot fully cover complex and diverse driving environments, nor can they dynamically adjust according to scene details, resulting in weak scene adaptability and insufficient user immersion and driving experience.
By collecting in-vehicle scene information, using a visual language model to parse scene semantic tags, constructing scene prompt words based on scene semantic tags and preset prompt word templates, and using an AIGC model to generate sound effect data that matches the in-vehicle scene information, the audio is then played.
It achieves dynamic matching of sound effects to in-vehicle scenarios, enhancing the user's immersion and driving experience. It can generate appropriate sound effects based on scene details, covering complex and diverse driving scenarios.
Smart Images

Figure CN122493828A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle control technology, and in particular to a sound effect generation method, device, vehicle, storage medium and product. Background Technology
[0002] With the development of smart cockpit technology, users' demands for immersive and personalized experiences in the in-vehicle environment are increasing.
[0003] Currently, most in-vehicle scenario-based audio solutions rely on predefined scene libraries and rule matching to play sound effects adapted to the current driving scenario. Specifically, sensors identify the current driving environment information, and based on preset rules, find a fixed sound effect file (such as wind noise, rain sound, etc.) that matches the current driving environment information from the scene library and play that fixed sound effect file. This results in a limited number of pre-stored sound effects that cannot fully cover the complex and diverse driving environment, and cannot be dynamically adjusted according to scene details, resulting in weak scene adaptability.
[0004] Therefore, how to generate sound effects that are compatible with the current in-vehicle scenario to enhance user immersion and driving experience is a problem that urgently needs to be solved. Summary of the Invention
[0005] The main purpose of this application is to provide a sound effect generation method, device, vehicle, storage medium and product, which aims to generate sound effects adapted to the current in-vehicle scenario in order to enhance user immersion and driving experience.
[0006] To achieve the above objectives, this application provides a sound effect generation method, the sound effect generation method comprising: In response to sound effect generation commands, collect in-vehicle scene information of the target vehicle; The in-vehicle scene information is analyzed using a visual language model to obtain scene semantic tags; Based on the scene semantic tags and the preset prompt word template, scene prompt words are constructed; The AIGC model generates sound effect data that matches the in-vehicle scene information based on the scene prompts, and plays audio based on the sound effect data.
[0007] In one embodiment, the step of collecting the in-vehicle scene information of the target vehicle includes: Collect first video data in front of the target vehicle, second video data inside the target vehicle, driving status data of the target vehicle, environmental data of the driving environment in which the target vehicle is located, and personalized configuration information of the target vehicle; The first video data, the second video data, the driving status data, the environmental data, and the personalized configuration information are used as the in-vehicle scene information of the target vehicle.
[0008] In one embodiment, the step of parsing the in-vehicle scene information using a visual language model to obtain scene semantic labels includes: Visual features are extracted from the first and second video data using a visual language model to obtain basic scene labels, driver status labels, and obstacle labels. The visual language model is used to extract text features from the driving status data, the environmental data, and the personalized configuration information to obtain dynamic parameter labels. The basic scene label, the driver status label, the obstacle label, and the dynamic parameter label are used as scene semantic labels.
[0009] In one embodiment, the basic scene labels include road condition type, scenery type, and weather type; the dynamic parameter labels include sound effect style, vehicle speed, weather intensity, traffic element density, and steering status; the driver status labels include fatigue level; and the obstacle labels include obstacle type, relative position, relative distance, and relative speed.
[0010] In one embodiment, the prompt word template includes an ambient sound effect template and a safety sound effect template, and the scene prompt words include a first prompt word and a second prompt word. The step of constructing scene prompt words based on the scene semantic tags and the preset prompt word template includes: Based on the basic scene tags, the dynamic parameter tags, and the ambient sound effect template, a first prompt word is generated; A second prompt word is generated based on the driver status label, the obstacle label, and the safety sound effect template.
[0011] In one embodiment, the step of generating a second prompt word based on the driver status label, the obstacle label, and the safety sound effect template includes: Based on the driver status label, adjust the audio parameters and safety constraint parameters in the safety sound effect template. When the fatigue level in the driver status label exceeds a preset threshold, increase the reference playback frequency and reference volume in the audio parameters, and increase the playback priority indicator in the safety constraint parameters. Based on the obstacle labels and the adjusted safety sound effect template, a second prompt word is generated.
[0012] In one embodiment, the step of generating sound effect data adapted to the in-vehicle scene information based on the scene prompts using an AIGC model includes: Ambient sound effect data is generated based on the first prompt word using the AIGC model; The AIGC model generates secure sound effect data based on the second prompt word. Audio is played based on the ambient sound effect data and the safety sound effect data.
[0013] In one embodiment, the step of generating secure sound effect data based on the second prompt word using the AIGC model includes: The AIGC model determines the spatial orientation of the safety sound data based on the relative position in the second prompt word, wherein the spatial orientation is used to indicate that the safety sound data is played through the channel on the target vehicle corresponding to the relative position; The AIGC model determines the playback frequency of the safety sound effect data based on the relative distance in the second prompt word and the baseline playback frequency, wherein the playback frequency is negatively correlated with the relative distance; The AIGC model determines the playback volume of the safety sound effect data based on the relative speed in the second prompt word and the reference volume, wherein the playback volume is positively correlated with the relative speed.
[0014] In one embodiment, the step of playing audio based on the ambient sound effect data and the security sound effect data includes: The ambient sound data and the safety sound data are mixed to obtain composite sound data; Detect whether the target vehicle has other audio effect data that is currently to be played; If so, then the audio effect data to be played is determined from the composite audio effect data and the other audio effect data based on a preset priority, and the audio is played based on the audio effect data to be played. If not, then play the audio based on the composite sound effect data.
[0015] Furthermore, to achieve the above objectives, this application also provides a sound effect generation device, the sound effect generation device comprising: The acquisition module is used to acquire in-vehicle scene information of the target vehicle in response to the sound effect generation command; The parsing module is used to parse the in-vehicle scene information using a visual language model to obtain scene semantic tags; The construction module is used to construct scene prompt words based on the scene semantic tags and preset prompt word templates; The sound effect generation module is used to generate sound effect data that matches the vehicle scene information based on the scene prompt words using the AIGC model, and to play audio based on the sound effect data.
[0016] In addition, to achieve the above objectives, this application also proposes an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound effect generation method described above.
[0017] In addition, to achieve the above objectives, this application also proposes a vehicle that includes the electronic equipment described above.
[0018] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a program for implementing the sound effect generation method is stored, and the program for implementing the sound effect generation method is executed by a processor to implement the steps of the sound effect generation method as described above.
[0019] In addition, to achieve the above objectives, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound effect generation method described above.
[0020] This application provides a sound effect generation method. In response to a sound effect generation command, this application collects in-vehicle scene information of a target vehicle; parses the in-vehicle scene information using a visual language model to obtain scene semantic tags; constructs scene prompts based on the scene semantic tags and preset prompt word templates; finally, generates sound effect data adapted to the in-vehicle scene information based on the scene prompts using an AIGC (Artificial Intelligence Generated Content) model, and plays the corresponding audio based on the sound effect data.
[0021] In summary, this application actively collects in-vehicle scene information from the target vehicle and then uses a visual language model to parse the collected information to obtain scene semantic tags. This allows for the accurate mining of various features and details of the in-vehicle scene, eliminating reliance on predefined simple scenes. Scene prompts, constructed based on scene semantic tags and preset prompt templates, can transform scene features into precise instructions recognizable by the AIGC model, making sound effect generation more aligned with the actual in-vehicle scene. Finally, the AIGC model generates and plays adapted sound effect data based on the scene prompts. Thus, compared to the traditional method of calling pre-stored fixed sound effect files, this application not only covers a wide variety of complex in-vehicle driving scenarios but also generates adapted sound effects based on the scene details reflected in the scene semantic tags. This achieves dynamic matching of sound effects to the in-vehicle scene, ensuring the generated sound effects are highly consistent with the current in-vehicle scene. This fundamentally solves the problem of limited pre-stored sound effects and the inability to dynamically adjust them, allowing users to experience audio synchronized with the actual scene while driving, thereby enhancing user immersion and the overall driving experience. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating the first embodiment of the sound effect generation method of this application; Figure 2 This is a schematic diagram of the system architecture involved in an embodiment of the sound effect generation method of this application; Figure 3 This is a schematic diagram of the sound effect generation process according to an embodiment of the sound effect generation method of this application; Figure 4 This is a schematic diagram of the module structure of the sound effect generation device of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the sound effect generation method in the embodiments of this application.
[0025] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0026] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0027] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0028] Currently, most in-vehicle scenario-based audio solutions rely on predefined scene libraries and rule matching to play sound effects adapted to the current driving scenario. Specifically, sensors identify the current driving environment information, and based on preset rules, find a fixed sound effect file (such as wind noise, rain sound, etc.) that matches the current driving environment information from the scene library and play that fixed sound effect file. This results in a limited number of pre-stored sound effects that cannot fully cover the complex and diverse driving environment, and cannot be dynamically adjusted according to scene details, resulting in weak scene adaptability.
[0029] Therefore, how to generate sound effects that are compatible with the current in-vehicle scenario to enhance user immersion and driving experience is a problem that urgently needs to be solved.
[0030] The main solution of this application is as follows: in response to the sound effect generation command, the in-vehicle scene information of the target vehicle is collected; the in-vehicle scene information is parsed through a visual language model to obtain scene semantic tags; scene prompt words are constructed based on the scene semantic tags and preset prompt word templates; sound effect data adapted to the in-vehicle scene information is generated based on the scene prompt words through an AIGC model, and audio is played based on the sound effect data.
[0031] This application actively collects in-vehicle scene information from target vehicles and then uses a visual language model to parse the collected information to obtain scene semantic tags. This allows for the precise extraction of various features and details of the in-vehicle scene, eliminating reliance on predefined, simple scenes. Scene prompts, constructed based on scene semantic tags and preset prompt templates, can transform scene features into precise instructions recognizable by the AIGC model, making sound effect generation more aligned with the actual in-vehicle scene. Finally, the AIGC model generates and plays adapted sound effect data based on the scene prompts. Thus, compared to the traditional method of calling pre-stored fixed sound effect files, this application not only covers a wide variety of complex in-vehicle driving scenarios but also generates adapted sound effects based on the scene details reflected in the scene semantic tags. This achieves dynamic matching of sound effects to the in-vehicle scene, ensuring the generated sound effects are highly consistent with the current in-vehicle scene. This fundamentally solves the problem of limited pre-stored sound effects and the inability to dynamically adjust them, allowing users to experience audio synchronized with the actual scene while driving, thereby enhancing user immersion and the overall driving experience.
[0032] It should be noted that the execution subject of the sound effect generation method in various embodiments of this application can be a vehicle control system or an electronic device capable of performing the above functions, etc., and this embodiment does not specifically limit it in this way. The following uses a vehicle control system as the execution subject as an example to describe this embodiment and the following embodiments.
[0033] Based on this, this application proposes a sound effect generation method according to the first embodiment, please refer to... Figure 1 The sound effect generation method includes steps S10 to S40: Step S10: In response to the sound effect generation command, collect the in-vehicle scene information of the target vehicle.
[0034] It should be noted that the sound effect generation command is a start signal used to trigger the sound effect generation process, which can be triggered manually or automatically by the system. In-vehicle scene information refers to a multimodal data set that characterizes the current environment and state of the target vehicle, including but not limited to image / video data in front of the vehicle, image / video data inside the vehicle (the driver's area), the vehicle's own driving status data (such as speed, acceleration, and steering angle), external environmental data (such as rainfall and lighting), and user-personalized configuration information (such as preferred music styles).
[0035] In this embodiment, step S10 may include: Step S101: Collect first video data in front of the target vehicle, second video data inside the target vehicle, driving status data of the target vehicle, environmental data of the driving environment in which the target vehicle is located, and personalized configuration information of the target vehicle. Step S102: The first video data, the second video data, the driving status data, the environmental data, and the personalized configuration information are used as the in-vehicle scene information of the target vehicle.
[0036] It should be noted that images or video streams captured by onboard cameras deployed in front of the target vehicle (such as the forward-facing main camera) are referred to as first video data for distinction. First video data is used to perceive information such as the road environment, traffic signs, scenic elements, and dynamic obstacles in front of the vehicle. Images or video streams captured by cameras deployed inside the vehicle (such as driver monitoring cameras located behind the steering wheel or on the A-pillar) are referred to as second video data for distinction. Second video data is used to capture the driver's facial features, head posture, and eye state to identify the driver's fatigue level and attention level. Driving status data refers to vehicle motion parameters obtained through the vehicle bus (such as the CAN bus), including but not limited to vehicle speed, acceleration, and steering angle. Environmental data refers to external environmental information collected by environmental sensors (such as rain sensors and light sensors), including rainfall level, light intensity, and temperature. Personalized configuration information refers to pre-stored user preference settings, such as user-preferred music styles, sound effect modes, and volume preferences.
[0037] Upon receiving the sound effect generation command, the system immediately activates its perception unit. It captures real-time video streams from the front-view camera and driver's face via the in-vehicle camera to identify their state. Simultaneously, it obtains driving status and environmental data from speed and rain sensors, and reads the user's personalized configuration information from the cockpit domain controller. All collected data is packaged into structured in-vehicle scene information, serving as input for the next stage of semantic analysis. After collection, the system integrates these five types of data according to a preset data structure to form structured in-vehicle scene information, which then serves as input data for the subsequent VLM (Vision-Language Model) semantic analysis module. This collection process ensures the comprehensiveness, real-time nature, and multimodal characteristics of the input data, laying a data foundation for accurate analysis of driving scenarios and driver states.
[0038] In one feasible implementation, the target vehicle is an intelligent car equipped with a multimodal perception system. When the user activates the "intelligent scene sound effect" function via the central control screen, the system generates a sound effect generation command. Subsequently, the forward-facing main camera captures a sequence of images of the area in front of the vehicle at a frame rate of at least 30fps, including images of the highway surface, guardrails on both sides, and a clear sky. Simultaneously, the in-vehicle driver monitoring camera captures a sequence of facial images of the driver for subsequent fatigue analysis. The system reads data from the vehicle speed sensor via the CAN bus to obtain the current vehicle speed as 110km / h, and reads data from the steering angle sensor to obtain the steering state as straight ahead. The rain sensor detects that there is no rain, and the light sensor detects that the light intensity is 8000 lux. At the same time, the system reads the user's preset personalized configuration information from the cockpit domain controller, such as "preference style: dynamic" and "volume preference: medium-high". All the above data is packaged in real time to form structured in-vehicle scene information, which is temporarily stored in memory for subsequent processing.
[0039] In another feasible implementation, when the system detects that a vehicle has entered a tunnel and the light intensity has dropped sharply, it automatically triggers a sound effect generation command to re-collect various data in the current scene, including the forward video under low light conditions in the tunnel, the driver's facial expression under changing light conditions, vehicle speed and environmental data, etc., to ensure the timeliness of sound effect generation and scene adaptability.
[0040] Step S20: The in-vehicle scene information is parsed using a visual language model to obtain scene semantic tags.
[0041] It should be noted that a visual language model (VLM) refers to a deep learning model capable of simultaneously processing visual and textual data and learning the semantic relationships between them, such as quantized Qwen-VL and LLaVA-1.5. VLMs are deployed on automotive SoCs (System-on-Chips), such as the Qualcomm 8295 and NVIDIA Orin. In this embodiment, the VLM is used for multimodal understanding of in-vehicle scene information, outputting structured scene semantic labels. These scene semantic labels are abstract descriptions of the current driving scene, vehicle state, and driver state.
[0042] The collected multi-source data is input into VLM, which utilizes its cross-modal understanding capabilities to extract external scene semantics from the front of the vehicle image, driver state semantics from the in-vehicle image, and dynamic parameter semantics from the numerical data. The three are then fused to generate a set of structured labels.
[0043] In this embodiment, step S20 may include: Step S201: Visual features are extracted from the first video data and the second video data using a visual language model to obtain basic scene labels, driver status labels, and obstacle labels. Step S202: Extract text features from the driving status data, the environmental data, and the personalized configuration information using the visual language model to obtain dynamic parameter labels; Step S203: The basic scene label, the driver status label, the obstacle label, and the dynamic parameter label are used as scene semantic labels.
[0044] The basic scene tags include road condition type, scenery type, and weather type; the dynamic parameter tags include sound effect style, vehicle speed, weather intensity, traffic element density, and steering status; the driver status tags include fatigue level; and the obstacle tags include obstacle type, relative position, relative distance, and relative speed.
[0045] It should be noted that visual feature extraction refers to the VLM encoding the input video frames using its visual encoder (such as ViT or CLIP visual branch) to extract abstract feature vectors that represent the image content. Text feature extraction refers to the VLM semantically transforming the input numerical data (such as vehicle speed and rainfall level) using its text encoder, encoding it into feature vectors in the text space. Basic scene labels refer to semantic information describing the external driving environment, including road condition type, landscape type, and weather type. Driver state labels refer to semantic information describing the driver's current physiological and behavioral state, including fatigue level and attention level. Obstacle labels refer to semantic information describing potential hazards ahead of the vehicle, including obstacle type, relative position, relative distance, and relative speed. Dynamic parameter labels refer to quantitative parameter labels describing the vehicle's motion state and environmental intensity, including sound effect style, vehicle speed, weather intensity, traffic element density, and steering state. It should be understood that when the obstacle label indicates no obstacles, the system only generates the first prompt word and corresponding ambient sound effect data, and does not generate the second prompt word or safety sound effect data to avoid unnecessary sound effect interference.
[0046] Specifically, fatigue levels can be categorized into normal, mild, moderate, and severe based on the driver's eye condition and head posture; road condition types include highways, urban roads, mountain roads, construction zones, and intersections; scenery types include forests, seasides, snow-capped mountains, urban complexes, and farmland; weather types include sunny, rainy, snowy, foggy, and windy; obstacle types include pedestrians, non-motorized vehicles, large vehicles, stationary obstacles, and no obstacles; relative position indicates the direction of the obstacle relative to the vehicle (left / center / right); relative distance indicates the straight-line distance between the obstacle and the vehicle; relative speed indicates the approaching speed of the obstacle relative to the vehicle (positive values indicate moving away, negative values indicate moving closer); sound style refers to auditory preference parameters extracted from the user's personalized configuration information, such as "dynamic," "soothing," "natural," "technological," and "immersive"; vehicle speed is expressed in km / h; weather intensity is expressed in levels (e.g., rainfall levels 1-5); traffic element density indicates the number of pedestrians and vehicles per unit of field of vision; and turning status includes going straight, turning left, and turning right.
[0047] The first video data (image of the front of the vehicle) and the second video data (image of the driver inside the vehicle) are input into the visual encoder of the VLM. The visual encoder extracts spatiotemporal features from the images through multi-layer convolution and self-attention mechanisms, identifying external scene elements (road, scenery, weather, obstacles) and internal driver states (facial expressions, head posture, eye opening / closing), resulting in basic scene labels, driver state labels, and obstacle labels. Simultaneously, the system converts driving status data, environmental data, and personalized configuration information into text descriptions (e.g., "vehicle speed 110 km / h", "rainfall level 3"), which are input into the VLM's text encoder to extract semantic features representing dynamic parameters, resulting in dynamic parameter labels. Visual and text features are fused within the VLM through a cross-modal alignment mechanism to ensure consistency of labels across different modalities in the semantic space. Finally, the four types of labels are integrated into structured scene semantic labels for use in subsequent prompt word generation steps.
[0048] In one feasible implementation, after analyzing the image in front of the vehicle, the VLM outputs basic scene labels: road condition type = highway, scenery type = highway guardrail, weather type = sunny; after analyzing driving status data and environmental data, it outputs dynamic parameter labels: vehicle speed = 110km / h, weather intensity = none, traffic element density = low, turning status = straight; after analyzing the image of the driver inside the vehicle, it outputs driver status label: fatigue level = normal; after detecting obstacles in front of the vehicle, it outputs obstacle label: none. These labels collectively describe the scenario of "normal driving on a highway in sunny weather".
[0049] Step S30: Construct scene prompts based on the scene semantic tags and preset prompt word templates.
[0050] It should be noted that the preset prompt word template refers to a pre-designed text framework used to guide the AIGC model in generating specific sound effects, which includes fixed descriptive fields and fillable placeholders. Scene prompt words refer to structured text instructions used to drive the AIGC model to generate sound effects, obtained by dynamically filling the template according to the semantic tags of the current scene.
[0051] The information in the scene semantic tags is organized according to a preset template format to generate a set of structured prompt words, namely scene prompt words.
[0052] Step S40: Generate sound effect data adapted to the vehicle scene information based on the scene prompt words using the AIGC model, and play audio based on the sound effect data.
[0053] It should be noted that AIGC models refer to audio generation models based on generative artificial intelligence technology, such as lightweight versions of AudioLDM-S and MusicGen-small, which can generate audio content matching text prompts in real time. Sound effect data refers to the digital audio signals output by the AIGC model, including ambient sound effects used to create immersive experiences (such as wind and rain sounds) and safety sound effects used for proactive safety warnings (such as directional warning sounds).
[0054] Scene prompts are input into the AIGC model, which uses a streaming generation algorithm to synthesize sound effect data in real time that highly matches the current driving scenario and driver state. For safety sound effects, the model generates directional audio signals based on the spatial orientation information in the prompts and dynamically adjusts the frequency, volume, and urgency of the sound effects according to obstacle distance, speed, and driver fatigue level. The generated sound effect data is sent to the audio control module, which intelligently mixes the newly generated sound effects with the in-vehicle prompts according to preset mixing strategies and priority rules (with priority dynamically adjusted considering driver fatigue), ensuring that critical safety prompts are not interfered with and can effectively attract the attention of fatigued drivers. Finally, the processed sound effects are played through the in-vehicle multi-channel audio system, achieving an immersive interactive experience of "what you see is what you hear" and proactive safety warnings.
[0055] In summary, this application embodiment collects video from the front of the vehicle, video from inside the vehicle, driving status, environmental data, and personalized configuration information. It then utilizes a visual language model to simultaneously extract basic scene labels, driver status labels, obstacle labels, and dynamic parameter labels, thereby incorporating the external environment, driver status, potential hazards, and user preferences into a unified semantic parsing framework. Thus, this application embodiment breaks through the traditional static mode of "limited pre-stored sound effects + rule matching," enabling real-time perception of complex and ever-changing driving scenarios, driver fatigue states, and personalized preferences. This lays a precise semantic foundation for subsequently generating highly adaptable and dynamically adjustable sound effects, effectively solving the technical problems of pre-stored sound effects being unable to cover diverse scenarios and unable to dynamically adjust according to details, thereby improving the scene adaptability and user experience of in-vehicle sound effects.
[0056] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, the prompt word template includes an ambient sound effect template and a safety sound effect template, the scene prompt word includes a first prompt word and a second prompt word, and step S30 may include: Step S301: Generate a first prompt word based on the basic scene label, the dynamic parameter label, the driver status label, and the ambient sound effect template.
[0057] It should be noted that the prompt word templates include ambient sound effect templates and safety sound effect templates. Ambient sound effect templates refer to pre-designed text frames used to generate background environmental sound effects. They contain fixed descriptive fields and fillable placeholders, such as "Style: [Scene Style], Core Elements: [Key Scene Elements], Dynamic Parameters: [Vehicle Speed / Weather Intensity], Audio Parameters: [Duration / Volume / Reverb], Safety Constraints: [Do Not Interfere with Driving Prompt Sounds]". The scene style is an automatically derived style description based on basic scene tags and sound effect styles, matching the current driving environment and user personalization. Examples include "Open and Dynamic" (user preference: dynamic + highway in sunny weather), "Quiet and Natural" (user preference: soothing + mountain forest in rainy weather), and "Fresh and Open" (user preference: fresh + coastal highway).
[0058] The system first extracts road condition type, scenery type, and weather type from basic scene labels, and vehicle speed, weather intensity, traffic element density, and steering status from dynamic parameter labels. Then, according to preset mapping rules (such as "highway + sunny day → open and dynamic style" and "rainy day → quiet and natural style"), these labels are converted into corresponding field values in the template. Finally, all fields are concatenated into a complete natural language prompt word (hereinafter referred to as the first prompt word for distinction). This first prompt word serves as the input to the AIGC model, determining the style, content, and dynamic characteristics of the generated ambient sound effect. Its generation process is unaffected by the driver's state and focuses on creating an immersive background sound effect that matches the external environment.
[0059] For example, the system obtains the following basic scene labels: road conditions = highway, scenery = clear sky, weather = sunny; dynamic parameter labels: vehicle speed = 110km / h, weather intensity = none, traffic element density = low, turning = straight. The system calls the ambient sound effect template and fills the above information into the first prompt word: "Style: open and dynamic, core elements: continuous wind noise + slight tire friction sound, dynamic parameters: wind noise intensity increases linearly with vehicle speed, audio parameters: duration 3 seconds, volume 60dB, low reverberation intensity, safety constraint: automatically reduce volume by 30% when navigation prompt sound is triggered."
[0060] Step S302: Generate a second prompt word based on the driver status label, the obstacle label, and the safety sound effect template.
[0061] It should be noted that a safety sound effect template refers to a pre-designed text frame used to generate proactive safety warning sound effects. It contains placeholders for fields describing the type of safety event, spatial location, urgency, safety constraints, etc.
[0062] The system first extracts obstacle type, relative position, relative distance, and relative speed from obstacle labels, and fatigue level from driver status labels. Then, it determines the urgency level based on obstacle distance and speed (e.g., "high" urgency when distance is <5 meters and relative speed is negative), and adjusts the baseline parameters of the safety sound effect based on driver fatigue level (e.g., increasing volume and frequency when fatigued). Finally, this information is filled into the safety sound effect template to generate structured prompt words (hereinafter referred to as the second prompt word for distinction). This second prompt word serves as input to the AIGC model, determining the type, spatial orientation, dynamic change pattern, and playback priority of the generated safety sound effect, ensuring that safety warnings are personalized to the driver's current state.
[0063] In one feasible implementation, the system obtains obstacle labels: obstacle type = pedestrian, relative position = left, relative distance = 8 meters, relative speed = -2 km / h (approaching); driver status label: fatigue level = normal. The system calls a safety sound effect template and fills in the second prompt word: "Safety event type: pedestrian crossing, spatial orientation: left, urgency level: medium, audio parameters: playback frequency increases linearly with distance, base volume 60dB, safety constraint: the alarm sound is immediately paused when triggered, and its priority is higher than the navigation sound."
[0064] In another feasible embodiment, step S302 may include: Step S3021: Based on the driver status label, adjust the audio parameters and safety constraint parameters in the safety sound effect template. When the fatigue level in the driver status label exceeds a preset threshold, increase the reference playback frequency and reference volume in the audio parameters, and increase the playback priority indicator in the safety constraint parameters. Step S3022: Generate a second prompt word based on the obstacle label and the adjusted safety sound effect template.
[0065] It should be noted that audio parameters refer to the fields in the safety sound effect template used to control the physical characteristics of the generated sound effects, including the reference playback frequency (the fundamental frequency of the sound effect repetition or pulse) and the reference volume (the basic loudness of the sound effect). Safety constraint parameters refer to the fields in the safety sound effect template used to control the playback logic. Among them, the playback priority flag is used to indicate the priority order of this sound effect when it exists simultaneously with other in-vehicle sound effects (such as navigation sounds and alarm sounds).
[0066] The system first reads the fatigue level from the driver's status label and determines if it exceeds a preset threshold (e.g., above "mild"). If the fatigue level exceeds the threshold, it indicates that the driver is in a state of decreased attention and requires a stronger warning to effectively attract their attention. Therefore, the system automatically increases the baseline playback frequency (making the sound more urgent) and baseline volume (making the sound louder) in the safety sound effect template, while also increasing the playback priority indicator in the safety constraint parameters (e.g., raising it to a level close to an alarm sound). After these adjustments, the system then combines obstacle labels (including obstacle type, relative position, distance, and speed) to fill in the adjusted template, generating a complete second warning word. This second warning word drives the AIGC model to generate a safety sound effect with enhanced warning effect, ensuring that warning information can be effectively perceived when the driver is fatigued.
[0067] For example, the system detects frequent blinking and head drooping of the driver through the in-vehicle camera, and the VLM outputs a driver status label: Fatigue Level = Moderate. The preset threshold is "Mild," and the system determines that the fatigue level exceeds the limit. The system increases the baseline playback frequency in the safety sound effect template from the default 2Hz to 3Hz, the baseline volume from 60dB to 70dB, and the playback priority indicator from "Higher than ambient sound" to "Higher than navigation sound, second only to alarm sound." Subsequently, the system reads the obstacle label: Obstacle Type = Pedestrian, Relative Position = Left, Relative Distance = 5 meters, Relative Speed = -3km / h. Based on the adjusted template, a second prompt is generated: "Safety Event Type: Pedestrian crossing, Spatial Orientation: Left, Urgency Level: Moderate, Audio Parameters: Baseline playback frequency 3Hz, Baseline volume 70dB, Playback frequency increases linearly with decreasing distance, Safety Constraint: Alarm sound pauses immediately upon triggering, and priority is higher than navigation sound." This prompt will drive AIGC to generate a more urgent, louder safety sound effect with left-side spatial directionality. In another implementation, if the fatigue level is normal, no parameter adjustment is made, and a second prompt word is generated directly based on the original template and obstacle label to avoid excessive warnings interfering with normal driving.
[0068] In summary, this application generates a first prompt word by inputting scene semantic tags into ambient sound effect templates and safety sound effect templates, respectively. This first prompt word creates background sound effects that match the external environment. Simultaneously, a second prompt word is generated to dynamically create safety sound effects based on the driver's state and obstacle information. Furthermore, when the driver's fatigue level exceeds the limit, the application proactively increases the baseline playback frequency, volume, and priority of the safety sound effects. In this way, this application ensures both accurate adaptation of immersive ambient sound effects to the driving scenario and dynamic enhancement of the intensity and urgency of safety warnings based on the driver's fatigue level. This ensures that hazard warnings can still be effectively perceived even when the driver's attention is diminished, thereby improving both the immersive cabin experience and driving safety.
[0069] Based on the first and second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, step S40 may include: Step S401: Generate ambient sound effect data based on the first prompt word using the AIGC model; Step S402: Generate secure sound effect data based on the second prompt word using the AIGC model; Step S403: Play audio based on the ambient sound effect data and the safety sound effect data.
[0070] The system inputs the first prompt word into the AIGC model, which uses a streaming generation algorithm to synthesize ambient sound effect data matching the current driving scenario in real time. The generation delay of each sound effect is ≤2 seconds, and multi-channel output is supported. Simultaneously, the second prompt word is input into the same AIGC model (or another model instance is called in parallel) to generate corresponding safety sound effect data. The two generation processes can be executed independently in parallel or sequentially. Due to the lightweight design of the model (INT8 quantization, memory reuse), the entire generation process can be completed efficiently under the computing power constraints of the vehicle SoC, meeting the real-time requirements of the vehicle. The generated sound effect data represents two different audio contents: background environment and active safety warning. Subsequently, the audio control module mixes the ambient sound effect data and the safety sound effect data and plays them according to preset priority rules (such as safety sound effect having higher priority than ambient sound effect), ensuring that the driver can perceive the warning information first when a safety warning is present, while not completely losing the immersive background sound effect. The streaming generation algorithm means that the model generates audio data frame by frame or in blocks, outputting audio data as it is generated, without waiting for the complete sound effect to be generated before playback can begin. Multi-channel output means that the audio data generated by the model supports 5.1 or 7.1 channel format, which can be used with in-vehicle multi-channel audio systems to achieve spatial audio rendering.
[0071] Thus, the embodiments of this application reduce the model size while reducing the loss of computational accuracy through the "VLM+AIGC model joint quantization" technology (INT8 quantization), and combine memory reuse technology to adapt to the computing power constraints (computing power requirement ≤20TOPS) and power consumption requirements of automotive SoC.
[0072] In this embodiment, step S402 may include: Step S4021: Based on the relative position in the second prompt word, the AIGC model determines the spatial directionality of the safety sound effect data, wherein the spatial directionality is used to indicate that the safety sound effect data is played through the channel on the target vehicle corresponding to the relative position.
[0073] It should be noted that spatial directivity refers to the sense of direction presented by the safety sound data during playback. That is, through the volume distribution, time delay, or phase adjustment of different channels in the in-vehicle multi-channel audio system, the driver can perceive the direction of the source of danger through hearing. "The channel corresponding to the relative position" refers to a specific speaker among the multiple speakers arranged in the vehicle that is located or faces the relative position (such as the left front door speaker corresponding to the left side).
[0074] After receiving the second cue word, the AIGC model analyzes the "relative location" field and determines the spatial directivity parameters of the safety sound effect data based on this information. Specifically, when generating multi-channel audio signals, the model sets different gain, delay, or filtering coefficients for different channels, allowing the driver to perceive the direction from which the sound originates when played through the vehicle's audio system. For example, if the relative location is on the left, the model will enhance the output intensity of the left channel while reducing the output intensity of the right channel, thus simulating the effect of the sound source being on the left. This spatially directional warning helps drivers quickly locate hazards more effectively than traditional non-directional alarm sounds.
[0075] For example, if the second prompt word contains "relative position = left side", then the AIGC model uses a two-channel audio generation strategy when generating safety sound effect data: setting the gain of the left channel to 1.0, the gain of the right channel to 0.3, and introducing a slight left channel time advance (about 0.5ms) to simulate the spatial sense of the sound source being on the left. When the generated sound effect data is played through the in-vehicle audio system, the driver will clearly hear the warning sound coming from the left, thus quickly turning their attention to the left.
[0076] Step S4022: The AIGC model determines the playback frequency of the safety sound effect data based on the relative distance in the second prompt word and the baseline playback frequency, wherein the playback frequency is negatively correlated with the relative distance.
[0077] It should be noted that the AIGC model extracts the relative distance and baseline playback frequency from the second warning word, and then calculates the actual playback frequency according to a preset mapping function (such as a linear or exponential relationship). When the obstacle is far away, the playback frequency is low, and the sound effect is a slow "thump...thump...thump"; as the obstacle gets closer, the playback frequency gradually increases, becoming a rapid "thump thump thump thump," thus conveying a sense of urgency to the driver that the distance is gradually decreasing. This dynamic frequency change more intuitively reflects the changing trend of the degree of danger than a fixed frequency alarm sound.
[0078] For example, the second prompt includes "relative distance = 8 meters" and "baseline playback frequency = 2Hz". The AIGC model uses a linear mapping: playback frequency = baseline frequency × (threshold distance / relative distance), where the threshold distance is set to 10 meters, and the calculated playback frequency = 2 × (10 / 8) = 2.5Hz. When the obstacle distance is shortened to 3 meters, the playback frequency = 2 × (10 / 3) ≈ 6.7Hz.
[0079] Step S4023: The AIGC model determines the playback volume of the safety sound effect data based on the relative speed in the second prompt word and the reference volume, wherein the playback volume is positively correlated with the relative speed.
[0080] The AIGC model extracts relative speed and baseline volume from the second cue word, and then calculates the actual playback volume according to a preset mapping function (such as a linear or piecewise function). When an obstacle approaches rapidly (e.g., at a relative speed of -20 km / h), the playback volume increases significantly to strongly warn the driver; when the obstacle moves away or is stationary, the playback volume remains at the baseline level or is appropriately reduced to avoid excessive interference. This dynamic volume adjustment, combined with frequency adjustment, constructs a multi-dimensional expression of urgency.
[0081] For example, the second warning message includes "relative speed = -5 km / h" (approaching) and "base volume = 60 dB". The AIGC model uses a linear mapping: playback volume = base volume + |relative speed| × gain coefficient, with the gain coefficient set to 1 dB / (km / h), calculating the playback volume = 60 + 5 = 65 dB. If the relative speed increases to -20 km / h, the playback volume = 60 + 20 = 80 dB. Through this dynamic volume adjustment, the safety sound effect can emit a louder sound when danger is rapidly approaching, effectively alerting the driver.
[0082] In this embodiment, step S403 may include: Step S4031: Mix the ambient sound effect data and the safety sound effect data to obtain composite sound effect data; Step S4032: Detect whether the target vehicle has other sound effect data to be played. Step S4033: If yes, then determine the audio effect data to be played from the composite audio effect data and the other audio effect data based on the preset priority, and play the audio based on the audio effect data to be played. Step S4034: If not, play audio based on the composite sound effect data.
[0083] It should be noted that audio mixing refers to the process of superimposing multiple audio signals according to a certain proportion and rules to form a composite audio signal. Other audio effect data refers to various audio signals that may be played simultaneously in the vehicle cabin, excluding ambient sound effects and safety sound effects, including but not limited to navigation prompts, alarm sounds (such as forward collision warning and blind spot monitoring), collision warning sounds, and user-selected music.
[0084] The system first mixes ambient sound data and safety sound data to obtain composite sound data, allowing the two sound effects to play simultaneously without masking each other. Then, the system checks whether there are other sound effects to be played on the current target vehicle (such as navigation prompts, alarm sounds, user-selected music, etc.). If other sound effects exist, one or more of the highest priority sound effects are selected from the composite sound data and other sound effects according to preset priority rules for playback. If no other sound effects exist, the composite sound data is played directly.
[0085] In addition, the system supports real-time adjustment of sound effects' volume, spatial position, and other parameters based on dynamic parameters such as vehicle speed and steering status (e.g., enhancing wind noise simulation during acceleration and shifting sound effect spatial positioning during cornering) to achieve a more natural auditory experience. When the perception unit detects an emergency hazard (such as a vehicle braking suddenly or a pedestrian crossing), the system immediately pauses the composite sound effects generated by AIGC and amplifies the alarm sound to ensure that the driver receives the highest level of safety warning first.
[0086] In one feasible implementation, the system mixes the generated ambient sound effects (wind noise, tire friction noise) with the safety sound effects (left-side pedestrian warning sound) to obtain a composite sound effect. At this time, the system detects that the vehicle is playing the navigation prompt "Keep right 500 meters ahead." According to the preset priority: alarm sound > navigation sound > AIGC sound effect, the composite sound effect has a lower priority than the navigation sound, so the system pauses the composite sound effect and prioritizes playing the navigation prompt sound; after the navigation sound finishes playing, the composite sound effect resumes playback. Simultaneously, based on the current vehicle speed of 110 km / h, the system amplifies the wind noise intensity in the ambient sound effect in real time; when the vehicle turns, the spatial positioning of the safety sound effect is shifted to the channel corresponding to the turning direction to simulate real hearing.
[0087] In another feasible implementation, when the sensing unit detects a pedestrian suddenly crossing the road ahead, the system triggers an emergency switch: immediately pausing all AIGC-generated composite sound effects (including ambient and safety sound effects) and amplifying the alarm sound (such as a rapid "beep" sound) to play at maximum volume across all vehicle channels, ensuring the driver takes evasive action immediately. Once the danger has passed, the composite sound effects are gradually restored. Through the aforementioned mixing, priority control, real-time adaptation, and emergency switch mechanisms, this embodiment achieves intelligent coordination between AIGC sound effects and in-vehicle safety alerts, ensuring both an immersive experience and driving safety.
[0088] For example, such as Figure 2The diagram shows the system architecture. The vehicle control system includes a perception unit, a VLM semantic parsing module, a prompt word generation module, an AIGC audio generation module, an audio control module, and a playback module. The perception unit communicates with the VLM semantic parsing module to collect in-vehicle scene information of the target vehicle in real time. The VLM semantic parsing module communicates with the prompt word generation module to perform multimodal semantic parsing on the received in-vehicle scene information and output structured scene semantic tags. The prompt word generation module communicates with the AIGC audio generation module to generate a first prompt word for ambient sound effects and a second prompt word for safety sound effects based on the scene semantic tags and preset prompt word templates. The AIGC audio generation module communicates with the audio control module to generate ambient sound effect data based on the first prompt word and safety sound effect data based on the second prompt word. The audio control module communicates with the playback module to mix the ambient sound effect data and safety sound effect data and coordinate their playback with other in-vehicle sound effects according to a preset priority: alarm sound > safety sound effect > navigation sound > ambient sound effect > user-selected music. The playback module is used to play the processed audio data in real time through the vehicle's multi-channel audio system.
[0089] Based on the above system architecture, such as Figure 3The diagram illustrates the sound effect generation process. First, in response to the sound effect generation command, the sensing unit collects in-vehicle scene information of the target vehicle in real time at a frequency synchronized with the camera frame rate (≥30fps). This includes video data of the front of the vehicle, video data of the driver inside the vehicle, driving status data (vehicle speed, acceleration, steering angle), environmental data (rainfall, illumination), and personalized configuration information. Next, the Visual Language Model (VLM) performs real-time inference on the collected multimodal data, with an inference latency of ≤500ms per frame, outputting structured scene semantic tags, including basic scene tags (road conditions, scenery, weather), dynamic parameter tags (vehicle speed, weather intensity, traffic element density, steering state), driver status tags (fatigue level), and obstacle tags (obstacle type, relative position, relative distance, relative speed). If the tags are consistent for three consecutive frames, the scene is considered stable to avoid frequent sound effect switching. Then, the system's prompt word generation module constructs ambient prompt words and safety prompt words based on the stabilized scene semantic tags: ambient prompt words are based on basic scene tags and dynamic parameter tags, while safety prompt words are based on driver status tags and dynamic parameter tags. Obstacle labels; the AIGC audio generation module (i.e., the AIGC model) uses a streaming generation algorithm to generate ambient sound effect data that matches the scene in real time based on ambient cue words (default 2-3 seconds / segment, caching the end of the previous segment for smooth transition), and simultaneously generates safety sound effect data with spatial directionality, frequency negatively correlated with distance, and volume positively correlated with relative speed based on safety cue words; the audio control module mixes the ambient sound effect and safety sound effect to obtain a composite sound effect, and detects the presence of other in-vehicle sound effects (navigation, alarm, user-selected music, etc.), and plays them according to a preset priority (alarm sound > safety sound effect > navigation sound > ambient sound effect > user-selected music); at the same time, it adjusts the sound effect parameters in real time according to vehicle speed and steering status (e.g., increasing wind noise during acceleration, shifting spatial positioning during turning), and immediately pauses the AIGC sound effect and amplifies the alarm sound when the perception unit detects an emergency hazard; finally, the output unit plays the processed sound effect in real time through the in-vehicle multi-channel audio system and receives user feedback (e.g., volume adjustment, style switching), and sends the feedback data back to the cue word generation module to optimize the subsequent cue word generation logic and achieve personalized adaptation. The entire process achieves an end-to-end closed loop from visual perception to semantic parsing, from dynamic prompt word generation to multi-sound effect collaborative playback, taking into account both immersive experience and driving safety.
[0090] Thus, this embodiment of the application generates ambient sound effect data and safety sound effect data based on the first and second prompt words using an AIGC model. The safety sound effect data determines spatial directionality based on the relative position of obstacles, enabling the driver to perceive the direction of the hazard source through hearing. Simultaneously, the playback frequency and volume of the safety sound effect are dynamically adjusted based on relative distance and speed, achieving a progressive warning where "the closer the distance and the faster the speed, the louder and more urgent the warning." Finally, the ambient sound effect and safety sound effect are mixed and coordinated with other in-vehicle sound effects according to the priority order: alarm sound > safety sound effect > navigation sound > ambient sound effect. Therefore, this embodiment of the application, while creating an immersive background sound effect highly matched to the driving scenario, significantly improves the driver's ability to perceive potential hazards and enhances driving safety through spatialized, dynamic safety sound effects and intelligent priority control.
[0091] In the first feasible implementation, under a high-speed, clear weather scenario, the forward-facing camera captures images of the highway and clear sky ahead of the vehicle. The vehicle speed sensor detects a speed of 110 km / h, and the rain sensor detects no rain. The VLM semantic parsing module outputs the following tags: Road condition = highway, Scenery = highway guardrail + clear sky, Weather = sunny, Vehicle speed = 110 km / h, Turning = straight ahead. Simultaneously, the in-vehicle camera does not detect driver fatigue and there are no obstacles ahead. The prompt word generation module generates an ambient prompt word based on the basic scene tags and dynamic parameter tags: "Style: open and stable, Core elements: slight wind noise + tire friction sound, Dynamic parameters: Wind noise intensity remains stable at a speed of 110 km / h, Audio parameters: Duration 3 seconds, Volume 55dB, Low reverberation intensity, Safety constraint: Volume reduced by 40% when navigation prompt sound is triggered." Since there are no obstacles, no safety prompt word is generated. The AIGC audio generation module generates a 3-second matching ambient sound effect in real time and streams it to the audio control module. The audio control module detects the absence of navigation / alarm tones and directly outputs sound effects. When the navigation prompts "Keep right 5 kilometers ahead," it automatically reduces the AIGC sound effect volume to 33dB, restoring it after the navigation sound finishes playing. The in-car audio system uses multi-channel playback of wind noise and friction sounds to simulate the immersive experience of high-speed driving.
[0092] In the second feasible implementation, in a mountain forest rainy weather scenario, the forward-facing camera captures images of the mountain road in front of the vehicle, the trees on both sides, and raindrops. The vehicle speed sensor detects a vehicle speed of 60 km / h, and the rain sensor detects a rainfall level of 3 (moderate rain). The VLM semantic parsing module outputs the following tags: Road condition = mountain road, Scenery = mountain forest + trees, Weather = rain (moderate rain), Vehicle speed = 60 km / h, Turning = straight ahead. At the same time, the in-vehicle camera detects that the driver's fatigue level is normal, but the forward-facing camera detects a sudden obstacle ahead (such as a fallen tree), and outputs the obstacle tag: Obstacle type = stationary obstacle, Relative position = directly in front, Relative distance = 30 meters, Relative speed = 0. The prompt word generation module generates the ambient prompt word: "Style: Tranquil and natural; Core elements: Raindrops falling on the car window + birdsong in the forest (low volume) + slight tire splashing sound; Dynamic parameters: Raindrop sound intensity matches rainfall level 3; splashing sound is moderate at a vehicle speed of 60km / h; Audio parameters: Duration 3 seconds, volume 50dB, medium reverberation intensity; Safety constraint: Immediately pause when alarm sound is triggered." Simultaneously, based on obstacle tags, it generates the safety prompt word: "Safety event type: Obstacle ahead; Spatial orientation: Directly ahead; Urgency level: Low; Audio parameters: Baseline playback frequency 1Hz, base volume 45dB, playback frequency increases linearly with decreasing distance." The AIGC module generates ambient and safety sound effects. The audio control module plays birdsong through the side speakers, raindrops through the front speakers, and the safety sound effect slowly at a low frequency through the center channel. When the vehicle detects an obstacle ahead that is within 10 meters, the safety sound effect playback frequency automatically increases; when the alarm sound is triggered, all AIGC sound effects are immediately paused, and the alarm sound is played first.
[0093] In the third feasible implementation, in the coastal highway scenario, the forward-facing camera captures images of the coastal highway, the sea surface, and vegetation swayed by the sea breeze in front of the vehicle. The vehicle speed sensor detects a speed of 80 km / h, and the wind speed sensor detects a slight crosswind. The VLM semantic parsing module outputs the following tags: Road condition = coastal highway, Scenery = sea surface + vegetation, Weather = sunny + slight crosswind, Vehicle speed = 80 km / h, Turning = right turn. Simultaneously, the personalized configuration information reads the user's preferred sound effect style as "fresh". The prompt word generation module generates the following ambient prompt word: "Style: Fresh and open, Core elements: Wave sound + slight crosswind sound + rustling vegetation sound, Dynamic parameters: Wave sound volume remains stable with a vehicle speed of 80 km / h, when turning right, the right speaker's crosswind sound increases by 10%, Audio parameters: Duration 3 seconds, Volume 58dB, Reverberation intensity medium-high, Safety constraint: Does not interfere with lane departure warning sound." At this time, there are no obstacles, so no safety prompt word is generated. AIGC generates multi-channel sound effects; when turning right, the right speaker's crosswind sound slightly increases to simulate spatial positioning and enhance immersion. If a non-motorized vehicle is detected approaching from the right, an additional safety sound effect is generated and a directional warning sound is played through the right-side speaker to provide safety assistance.
[0094] This application also provides a sound effect generation device, please refer to... Figure 4 The sound effect generating device includes: The acquisition module 10 is used to acquire in-vehicle scene information of the target vehicle in response to the sound effect generation command; The parsing module 20 is used to parse the vehicle scene information through a visual language model to obtain scene semantic tags; Construction module 30 is used to construct scene prompt words based on the scene semantic tags and preset prompt word templates; The sound effect generation module 40 is used to generate sound effect data that is adapted to the vehicle scene information based on the scene prompt words through the AIGC model, and to play audio based on the sound effect data.
[0095] Optionally, the acquisition module 10 is further configured to: Collect first video data in front of the target vehicle, second video data inside the target vehicle, driving status data of the target vehicle, environmental data of the driving environment in which the target vehicle is located, and personalized configuration information of the target vehicle; The first video data, the second video data, the driving status data, the environmental data, and the personalized configuration information are used as the in-vehicle scene information of the target vehicle.
[0096] Optionally, the parsing module 20 is further configured to: Visual features are extracted from the first and second video data using a visual language model to obtain basic scene labels, driver status labels, and obstacle labels. The visual language model is used to extract text features from the driving status data, the environmental data, and the personalized configuration information to obtain dynamic parameter labels. The basic scene label, the driver status label, the obstacle label, and the dynamic parameter label are used as scene semantic labels.
[0097] Optionally, the basic scene tags include road condition type, scenery type, and weather type; the dynamic parameter tags include sound effect style, vehicle speed, weather intensity, traffic element density, and steering status; the driver status tags include fatigue level; and the obstacle tags include obstacle type, relative position, relative distance, and relative speed.
[0098] Optionally, the prompt word template includes an ambient sound effect template and a safety sound effect template, the scene prompt words include a first prompt word and a second prompt word, and the construction module 30 is further used for: Based on the basic scene tags, the dynamic parameter tags, and the ambient sound effect template, a first prompt word is generated; A second prompt word is generated based on the driver status label, the obstacle label, and the safety sound effect template.
[0099] Optionally, the building module 30 is further configured to: Based on the driver status label, adjust the audio parameters and safety constraint parameters in the safety sound effect template. When the fatigue level in the driver status label exceeds a preset threshold, increase the reference playback frequency and reference volume in the audio parameters, and increase the playback priority indicator in the safety constraint parameters. Based on the obstacle labels and the adjusted safety sound effect template, a second prompt word is generated.
[0100] Optionally, the sound effect generation module 40 is further configured to: Ambient sound effect data is generated based on the first prompt word using the AIGC model; The AIGC model generates secure sound effect data based on the second prompt word. Audio is played based on the ambient sound effect data and the safety sound effect data.
[0101] Optionally, the sound effect generation module 40 is further configured to: The AIGC model determines the spatial orientation of the safety sound data based on the relative position in the second prompt word, wherein the spatial orientation is used to indicate that the safety sound data is played through the channel on the target vehicle corresponding to the relative position; The AIGC model determines the playback frequency of the safety sound effect data based on the relative distance in the second prompt word and the baseline playback frequency, wherein the playback frequency is negatively correlated with the relative distance; The AIGC model determines the playback volume of the safety sound effect data based on the relative speed in the second prompt word and the reference volume, wherein the playback volume is positively correlated with the relative speed.
[0102] Optionally, the sound effect generation module 40 is further configured to: The ambient sound data and the safety sound data are mixed to obtain composite sound data; Detect whether the target vehicle has other audio effect data that is currently to be played; If so, then the audio effect data to be played is determined from the composite audio effect data and the other audio effect data based on a preset priority, and the audio is played based on the audio effect data to be played. If not, then play the audio based on the composite sound effect data.
[0103] The sound effect generation device provided in this application, employing the sound effect generation method described in the above embodiments, can solve the technical problem of how to generate sound effects adapted to the current in-vehicle scenario, thereby enhancing user immersion and driving experience. Compared with the prior art, the beneficial effects of the sound effect generation device provided in this application are the same as those of the sound effect generation method provided in the above embodiments, and other technical features in the sound effect generation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0104] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the sound effect generation method in Embodiment 1 above.
[0105] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic device in these embodiments may be an in-vehicle terminal (e.g., an in-vehicle navigation terminal). Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0106] like Figure 5 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0107] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0108] The electronic device provided in this application, employing the sound effect generation method described in the above embodiments, can solve the technical problem of how to generate sound effects adapted to the current in-vehicle scenario, thereby enhancing user immersion and driving experience. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the sound effect generation method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0109] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0110] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0111] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the sound effect generation method in the above embodiments.
[0112] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0113] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0114] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by an electronic device, the electronic device causes the electronic device to: respond to a sound effect generation instruction by acquiring in-vehicle scene information of the target vehicle; parse the in-vehicle scene information using a visual language model to obtain scene semantic tags; construct scene prompt words based on the scene semantic tags and a preset prompt word template; generate sound effect data adapted to the in-vehicle scene information using an AIGC model based on the scene prompt words, and play audio based on the sound effect data.
[0115] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0117] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0118] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described sound effect generation method. This solves the technical problem of how to generate sound effects adapted to the current in-vehicle scenario to enhance user immersion and driving experience. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the sound effect generation method provided in the above embodiments, and will not be repeated here.
[0119] This application provides a vehicle having the electronic equipment described above, the electronic equipment being used to perform the sound effect generation method in the above embodiments.
[0120] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound effect generation method described above.
[0121] The computer program product provided in this application can generate sound effects adapted to the current in-vehicle scenario, thereby enhancing user immersion and driving experience. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the sound effect generation method provided in the above embodiments, and will not be repeated here.
[0122] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for generating sound effects, characterized in that, The sound effect generation method includes: In response to sound effect generation commands, collect in-vehicle scene information of the target vehicle; The in-vehicle scene information is analyzed using a visual language model to obtain scene semantic tags; Based on the scene semantic tags and the preset prompt word template, scene prompt words are constructed; The AIGC model generates sound effect data that matches the in-vehicle scene information based on the scene prompts, and plays audio based on the sound effect data.
2. The sound effect generation method as described in claim 1, characterized in that, The steps for collecting the in-vehicle scene information of the target vehicle include: Collect first video data in front of the target vehicle, second video data inside the target vehicle, driving status data of the target vehicle, environmental data of the driving environment in which the target vehicle is located, and personalized configuration information of the target vehicle; The first video data, the second video data, the driving status data, the environmental data, and the personalized configuration information are used as the in-vehicle scene information of the target vehicle.
3. The sound effect generation method as described in claim 2, characterized in that, The step of parsing the in-vehicle scene information using a visual language model to obtain scene semantic labels includes: Visual features are extracted from the first and second video data using a visual language model to obtain basic scene labels, driver status labels, and obstacle labels. The visual language model is used to extract text features from the driving status data, the environmental data, and the personalized configuration information to obtain dynamic parameter labels. The basic scene label, the driver status label, the obstacle label, and the dynamic parameter label are used as scene semantic labels.
4. The sound effect generation method as described in claim 3, characterized in that, The basic scene tags include road condition type, scenery type, and weather type; the dynamic parameter tags include sound effect style, vehicle speed, weather intensity, traffic element density, and steering status; the driver status tags include fatigue level; and the obstacle tags include obstacle type, relative position, relative distance, and relative speed.
5. The sound effect generation method as described in claim 4, characterized in that, The prompt word template includes an ambient sound effect template and a safety sound effect template; the scene prompt words include a first prompt word and a second prompt word; the step of constructing scene prompt words based on the scene semantic tags and the preset prompt word templates includes: Based on the basic scene tags, the dynamic parameter tags, and the ambient sound effect template, a first prompt word is generated; A second prompt word is generated based on the driver status label, the obstacle label, and the safety sound effect template.
6. The sound effect generation method as described in claim 5, characterized in that, The step of generating a second prompt word based on the driver status label, the obstacle label, and the safety sound effect template includes: Based on the driver status label, adjust the audio parameters and safety constraint parameters in the safety sound effect template. When the fatigue level in the driver status label exceeds a preset threshold, increase the reference playback frequency and reference volume in the audio parameters, and increase the playback priority indicator in the safety constraint parameters. Based on the obstacle labels and the adjusted safety sound effect template, a second prompt word is generated.
7. The sound effect generation method as described in claim 6, characterized in that, The step of generating sound effect data adapted to the in-vehicle scene information based on the scene prompts using an AIGC model includes: Ambient sound effect data is generated based on the first prompt word using the AIGC model; The AIGC model generates secure sound effect data based on the second prompt word. Audio is played based on the ambient sound effect data and the safety sound effect data.
8. The sound effect generation method as described in claim 7, characterized in that, The step of generating safe sound effect data based on the second prompt word using the AIGC model includes: The AIGC model determines the spatial orientation of the safety sound data based on the relative position in the second prompt word, wherein the spatial orientation is used to indicate that the safety sound data is played through the channel on the target vehicle corresponding to the relative position; The AIGC model determines the playback frequency of the safety sound effect data based on the relative distance in the second prompt word and the baseline playback frequency, wherein the playback frequency is negatively correlated with the relative distance; The AIGC model determines the playback volume of the safety sound effect data based on the relative speed in the second prompt word and the reference volume, wherein the playback volume is positively correlated with the relative speed.
9. The sound effect generation method as described in claim 7, characterized in that, The step of playing audio based on the ambient sound effect data and the safety sound effect data includes: The ambient sound data and the safety sound data are mixed to obtain composite sound data; Detect whether the target vehicle has other audio effect data that is currently to be played; If so, then the audio effect data to be played is determined from the composite audio effect data and the other audio effect data based on a preset priority, and the audio is played based on the audio effect data to be played. If not, then play the audio based on the composite sound effect data.
10. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound effect generation method as described in any one of claims 1 to 9.
11. A vehicle, characterized in that, The vehicle includes the electronic equipment as described in claim 10.
12. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the sound effect generation method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the sound effect generation method as described in any one of claims 1 to 9.