Deep learning active sound design system and method

By combining a deep learning model with fast and slow refresh rate inputs, an active sound design system has been developed to address the issues of high computational load and insufficient responsiveness in electric vehicles. This system enables fast-response and personalized active sound design, thereby enhancing the driving experience.

CN120977281APending Publication Date: 2025-11-18HYUNDAI MOTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411206315.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2024-08-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing active sound design systems for electric vehicles suffer from high computational load and insufficient responsiveness, making it difficult to meet the challenges of personalized customization and rapid response.

Method used

A deep learning model is used to combine fast and slow refresh rate inputs, allocate processing resources, generate active sound design, use FRRI to quickly respond to motor-related inputs, SRRI to generate a sound track library, and output synthesized sound through a speaker.

Benefits of technology

It enables rapid response and personalized customization of active sound design for electric vehicles, reduces computational load, and improves the driver's driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977281A_ABST
    Figure CN120977281A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning active sound design system and method. Systems and methods for active sound design (ASD) generation are provided. The system may include one or more speakers and a computing device including a processor and a memory. The memory may be configured to store instructions that, when executed by the processor, are configured to cause the processor to receive one or more inputs for synthetic sound generation, classify the one or more inputs as a fast refresh rate input (FRRI) or a slow refresh rate input (SRRI), allocate one or more processing resources according to a refresh rate, and transmit the one or more processing resources to the processor. An ASD is generated based on the one or more inputs, and the synthesized sound is played on the one or more speakers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to systems and methods for generating Active Sound Design (ASD). Background Technology

[0002] Active Sound Design (ASD) has become a mandatory feature for many electric vehicles (EVs). Therefore, having an ASD that sets you apart in the market is crucial, to the point that competitors are hiring top talent and / or showcasing their unique technologies. North American customers tend to want vehicles that express their individuality and value personalized and customizable features.

[0003] The customer expressed interest in creating their own active sound design. However, they were unaware of the inherent skill and technical challenges of implementing an Active Sound System (ASD) on their own. Furthermore, the customer desired ASD to be responsive to current driving conditions (i.e., without lag). However, increasing ASD responsiveness increases the required computational load, and this increased computational load raises costs. Summary of the Invention

[0004] According to one object of the present invention, a system for generating Active Sound Design (ASD) is provided. The system may include one or more speakers and a computing device, the computing device including a processor and a memory. The memory may be configured to store instructions, which, when executed by the processor, are configured to cause the processor to: receive one or more inputs for synthesized sound generation; classify the one or more inputs as Fast Refresh Rate Input (FRRI) or Slow Refresh Rate Input (SRRI); and allocate one or more processing resources according to the refresh rate. When executed by the processor, the instructions may be configured to cause the processor to generate ASD based on one or more inputs, wherein generation includes: modifying one or more weights of a deep learning model using SRRI and one or more cue words; processing the weights and cue words using the deep learning model to output one or more looping sound files to a track library, thereby forming one or more audio tracks; modifying one or more dynamics of a wave synthesis ASD module using FRRI; and generating synthesized sound using one or more audio tracks and the wave synthesis ASD module. The cue words may include inputs to the deep learning model. When the instructions are executed by the processor, the instructions can be configured to cause the processor to play synthesized sound on one or more speakers.

[0005] According to an exemplary implementation, FRRI may include one or more inputs selected from: throttle position; motor speed; wheel speed; brake position; vehicle gravity; and motor load.

[0006] According to an exemplary implementation, SRRI may include one or more inputs selected from the following: time of day; one or more calendar dates; location; driving mode; weather; traffic conditions; aggressiveness; complexity; musicality; one or more stored individual model weights; one or more shared model weights; and user ASD history.

[0007] According to an exemplary implementation, a deep learning model may include a diffusion model.

[0008] According to an exemplary implementation, prompt words can be defined by the way a deep learning model is built and trained.

[0009] According to an exemplary embodiment, the synthesized sound can be the sound of a synthesized power system.

[0010] According to an exemplary implementation, the system may include a vehicle.

[0011] According to an exemplary implementation, one or more speakers can be connected to the vehicle.

[0012] According to an exemplary implementation, one or more audio tracks can be manipulated by one or more ASD dynamic curves.

[0013] According to an exemplary implementation, each of one or more audio tracks may include multiple ASD dynamic curves for each FRRI.

[0014] According to an exemplary implementation, when the instructions are executed by the processor, the instructions may be configured to enable the processor to: enable a first user to share one or more tracks of a first audio track library with a second user.

[0015] According to one object of the present invention, a method for ASD generation is provided. The method may include: receiving one or more inputs for synthesized sound generation, and classifying the one or more inputs into FRRI or SRRI. Since the near-instantaneous response to these inputs is crucial for driver dynamic perception, FRRI can be assigned to the direct operation of a wave synthesis ASD module. SRRI can be assigned to the input weights of a trained deep learning model. When appropriate processing power and memory are available, the model can be configured to generate audio tracks for a track library as looping sound files. Since the generation of new audio tracks is not critical to driver dynamic perception, new audio tracks can be generated at a slower rate than the output of the wave synthesis module. The ASD module utilizing audio tracks from the track library and FRRI can be configured to output synthesized sound. The method may include playing the synthesized sound on one or more speakers.

[0016] According to an exemplary implementation, FRRI may include one or more inputs selected from: throttle position; motor speed; wheel speed; brake position; vehicle gravity; and motor load.

[0017] According to an exemplary implementation, SRRI may include one or more inputs selected from the following: time of day; one or more calendar dates; location; driving mode; weather; traffic conditions; aggressiveness; complexity; musicality; one or more stored individual model weights; one or more shared model weights; and user ASD history.

[0018] According to an exemplary implementation, a deep learning model may include a diffusion model.

[0019] According to an exemplary implementation, prompt words can be defined by the way a deep learning model is built and trained.

[0020] According to an exemplary embodiment, the synthesized sound can be the sound of a synthesized power system.

[0021] According to an exemplary implementation, one or more speakers can be connected to the vehicle.

[0022] According to an exemplary implementation, the method may include manipulating one or more audio tracks using one or more ASD dynamic curves.

[0023] According to an exemplary implementation, each of one or more audio tracks may include multiple ASD dynamic curves for each FRRI.

[0024] According to an exemplary implementation, the method may include enabling a first user to share one or more tracks from a first audio track library with a second user. Attached Figure Description

[0025] Various non-limiting and non-exhaustive embodiments of the subject matter are illustrated in conjunction with the accompanying drawings, which form part of the detailed description, and are used together with the detailed description to explain the principles of the subject matter discussed below. Unless otherwise specified, the drawings mentioned in this description should be understood as not drawn to scale, and unless otherwise stated, the same reference numerals refer to the same parts throughout the various drawings.

[0026] Figure 1 A configuration for a vehicle generated by Active Sound Design (ASD) according to an exemplary embodiment of the present invention is shown.

[0027] Figure 2 The process for integrating a deep learning system with an ASD system in a vehicle according to an exemplary embodiment of the present invention is illustrated.

[0028] Figure 3 A flowchart of a method for ASD generation according to an exemplary embodiment of the present invention is shown.

[0029] Figure 4 An active voice sharing process according to an exemplary embodiment of the present invention is illustrated.

[0030] Figure 5 An example architecture of a vehicle according to an exemplary embodiment of the present invention is shown.

[0031] Figure 6 Example elements of a computing device according to an exemplary embodiment of the present invention are shown. Detailed Implementation

[0032] The specific embodiments described below are provided by way of example only and not by way of limitation. Furthermore, they are not intended to be limited by any express or implied theory presented in the foregoing background or the specific embodiments below.

[0033] Reference will now be made in detail to various exemplary embodiments of the subject matter, examples of which are illustrated in the accompanying drawings. While various embodiments are discussed herein, it should be understood that they are not intended to be limited to these embodiments. Rather, the presented embodiments are intended to cover alternatives, modifications, and equivalents that may be included within the spirit and scope of the various embodiments defined by the appended claims. Furthermore, numerous specific details are set forth in this particular embodiment to provide a comprehensive understanding of the embodiments of the subject matter. However, the embodiments may be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the described embodiments.

[0034] Some parts of the following detailed description are presented as procedures, logic blocks, processes, and other symbolic representations relating to the manipulation of data within an electrical device. These descriptions and representations are means used by those skilled in the art of data processing to most effectively communicate the substance of their work to others skilled in the art. In this application, procedures, logic blocks, processes, etc., are considered as one or more self-consistent processes or instructions that lead to a desired result. A process is one that requires the physical manipulation of physical quantities. Typically, although not essential, these quantities may take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated in electronic systems, devices, and / or components.

[0035] However, it should be remembered that these and similar terms will be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specifically stated as will be apparent from the discussion below, it should be understood that throughout the description of the implementation, discussions using terms such as “determine,” “communicate,” “take,” “compare,” “monitor,” “calibrate,” “estimate,” “start,” “provide,” “receive,” “control,” “transmit,” “isolate,” “generate,” “align,” “synchronize,” “identify,” “maintain,” “display,” “switch,” etc., refer to the actions and processes of electronic items such as processors, sensor processing units (SPUs), processors of sensor processing units, application processors of electronic devices / systems, etc., or combinations thereof. This item manipulates and converts data represented as physical (electronic and / or magnetic) quantities in registers and memories into other data similarly represented as physical quantities within memory or registers or other such information storage, transmission, processing, or display components.

[0036] It should be understood that the term "vehicle" or "of a vehicle" or other similar terms as used herein generally include motor vehicles, such as passenger vehicles including sports utility vehicles (SUVs), buses, trucks, various commercial vehicles, vessels including various boats and ships, aircraft, etc., and includes hybrid vehicles, electric vehicles, plug-in hybrid electric vehicles, hydrogen-powered vehicles, and other alternative fuel vehicles (e.g., vehicles derived from non-petroleum fuels). As mentioned herein, a hybrid vehicle is a vehicle having two or more power sources, such as both gasoline power and electric power. In various respects, a vehicle may include an internal combustion engine system as disclosed herein.

[0037] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. These terms are intended only to distinguish one component from another, and they do not limit the nature, order, or sequence of the constituent components. It will also be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of the stated feature, value, step, operation, element, and / or component, but do not exclude the presence or inclusion of one or more other features, values, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more related enumerations. Throughout the specification, unless explicitly stated otherwise, the word “comprising” and variations such as “including” or “including” will be understood to imply the inclusion of the stated element, but do not exclude any other element. Furthermore, the terms “unit,” “device,” “component,” and “module” described in the specification refer to a unit for performing at least one function and operation, and this unit may be implemented by hardware components or software components and combinations thereof.

[0038] While exemplary embodiments are described as utilizing multiple units to perform exemplary processes, it should be understood that exemplary processes may also be performed by one or more modules. Furthermore, it should be understood that the term controller / control unit refers to a hardware device including a memory and a processor, specifically programmed to perform the processes described herein. The memory is configured to store modules, and the processor is specifically configured to execute said modules to perform one or more processes further described below.

[0039] Furthermore, the control logic of the present invention can be implemented as a non-volatile computer-readable medium containing executable program instructions that are executed by a processor, controller, or the like. Examples of computer-readable media include, but are not limited to, ROM, RAM, optical disc (CD)-ROM, magnetic tape, floppy disk, flash drive, smart card, and optical data storage device. The computer-readable recording medium can also be distributed across a network-connected computer system, allowing the computer-readable medium to be stored and executed in a distributed manner, for example, via a telematics server or a controller area network (CAN).

[0040] Unless otherwise stated or obvious from the context, as used herein, the term "approximately" is understood to mean within the normal tolerance range in the field, such as within two standard deviations of the mean. "Approximately" can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. All numerical values ​​provided herein are modified by the term "approximately" unless clearly stated from the context.

[0041] The embodiments described herein can be discussed in the general context of processor-executable instructions residing on some form of non-volatile processor-readable medium (e.g., program modules) that are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or distributed as needed.

[0042] In the accompanying drawings, a single block may be described as performing one or more functions; however, in practice, the one or more functions performed by that block may be performed in a single component or across multiple components, and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, logic, circuits, and steps are generally described according to their functions. Whether such a function is implemented as hardware or software depends on the specific application and design constraints imposed on the system as a whole. Those skilled in the art may implement the described functions in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention. Furthermore, the exemplary device vibration sensing system and / or electronic device described herein may include components in addition to those shown, including well-known components.

[0043] Unless specifically described as implemented in a particular manner, the various techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. Any feature described as a module or component may also be implemented together in an integrated logic device or separately as a discrete but interoperable logic device. If implemented in software, the techniques may be implemented at least in part by a non-volatile processor-readable storage medium including instructions that, when executed, perform one or more of the methods described herein. The non-volatile processor-readable data storage medium may form part of a computer program product (which may include packaging material).

[0044] Non-volatile processor-readable storage media may include random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, and other known storage media. Additionally or alternatively, the technology may be implemented at least in part by a processor-readable communication medium that carries or transmits code in the form of instructions or data structures, and can be accessed, read, and / or executed by a computer or other processor.

[0045] The various implementations described herein can be executed by one or more processors, such as one or more motion processing units (MPUs), sensor processing units (SPUs), host processors or their cores, digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), field-programmable gate arrays (FPGAs), programmable logic controllers (PLCs), complex programmable logic devices (CPLDs), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein, or other equivalent integrated or discrete logic circuits. As used herein, the term "processor" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. As used in this subject matter specification, the term "processor" can refer to virtually any computing processing unit or device, including but not limited to single-core processors; single-core processors with software multithreading capabilities; multi-core processors; multi-core processors with software multithreading capabilities; multi-core processors employing hardware multithreading technology; parallel platforms; and parallel platforms with distributed shared memory. Furthermore, processors can utilize nanoscale architectures, such as, but not limited to, molecular and quantum dot-based transistors, switches, and gates, to optimize space utilization or enhance the performance of user devices. Processors can also be implemented as a combination of computing processing units.

[0046] Furthermore, in some aspects, the functionality described herein can be provided within dedicated software or hardware modules configured as described herein. Additionally, the technology can be fully implemented in one or more circuit or logic elements. A general-purpose processor can be a microprocessor, but alternatively, the processor can be any processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of an SPU / MPU and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors with an SPU core, an MPU core, or any other combination of such configurations. One or more components of the SPU or electronic device described herein can be embodied in the form of a “chip,” a “package,” or one or more integrated circuits (ICs).

[0047] According to exemplary embodiments, systems and methods for generating Active Sound Design (ASD) are provided.

[0048] Now for reference Figure 1 The illustration depicts a configuration for ASD generation of a vehicle 100 according to an exemplary embodiment of the present invention. According to the exemplary embodiment, the vehicle 100 may include an EV.

[0049] According to an exemplary embodiment, vehicle 100 may include one or more sensors configured to detect and / or record sound, such as one or more microphones 105. According to an exemplary embodiment, vehicle 100 may include one or more speakers 110 configured to play one or more sounds. According to an exemplary embodiment, vehicle 100 may include a computing device 115. The computing device 115 may include a processor 120, memory 125, and / or a user interface 130 (e.g., a graphical user interface). The computing device 115 may be configured to send and / or receive instructions / data, etc., via one or more external systems (e.g., via a cloud 135) through wired and / or wireless connections.

[0050] According to an exemplary embodiment, one or more microphones 105 and / or one or more speakers 110 may communicate electronically with one or more computing devices 115. One or more computing devices 115 may be separate from one or more microphones 105 and / or one or more speakers 110, and / or one or more computing devices 115 may be integrated into one or more microphones 105 and / or one or more speakers 110.

[0051] The memory 125 may be configured to store programming instructions that, when executed by the processor 120, may be configured to cause the processor 120 to perform one or more tasks, such as receiving one or more inputs from one or more microphones 105 and / or user interface 130, performing ASD generation using deep learning, and / or performing other suitable tasks.

[0052] Now for reference Figure 2 The process 200 for combining a deep learning system with an ASD system of a vehicle (e.g., vehicle 100) according to an exemplary embodiment of the present invention.

[0053] According to an exemplary embodiment, the system of the present invention (e.g., vehicle 100) processes critical inputs with fast response times and passive inputs or direct manual inputs from users with slow changes, respectively, to greatly reduce the computing resources required by the artificial intelligence (AI) control system.

[0054] According to an exemplary embodiment, the system inputs (e.g., system input 205) can be virtually any dynamic variable (driven or passive). According to an exemplary embodiment, system input 205 may include: slow refresh rate input 210 (e.g., time of day, one or more calendar dates (e.g., one or more nearby calendar dates), location, driving mode, weather, traffic conditions, aggressiveness, complexity, musicality, one or more stored personal and / or shared model weights and user ASD history, and other suitable slow refresh rate inputs) and fast refresh rate input 215 (e.g., motor conditions (e.g., throttle position, motor speed, motor load), wheel speed, brake position, and vehicle gravity, and other suitable fast refresh rate inputs).

[0055] ASD requires a near-instantaneous or sharp instantaneous response time to motor-related inputs in order to make it sound pleasing to the driver. These inputs are classified as Fast Refresh Rate Input (FRRI) 215. All other inputs that are not important to ASD response are classified as Slow Refresh Rate Input (SRRI) 210.

[0056] According to an exemplary implementation, SRRI 210 can be configured to change the weights 230 of the deep learning model 235, and FRRI 215 can be configured to change the dynamics of the waveform synthesized audio track (stem) of the ASD via the waveform synthesized ASD module 245.

[0057] According to an exemplary embodiment, the trained model cue words and weights 220 can have a low processing priority. Cue words 225 are the inputs 235 of the deep learning model. According to an exemplary embodiment, SRRI 210 can be used to determine the values ​​of the weights 230 for those cue words 225.

[0058] The cue word 225 can be defined by how the deep learning model 235 is constructed and trained. Examples can be strong, calm, loud, metallic, etc., and can be determined by the designer (e.g., the user) to best suit the sound experience the designer desires.

[0059] According to an exemplary embodiment, the trained deep learning model 235 may have a low processing priority. According to an exemplary embodiment, the deep learning model 235 may be an AI algorithm configured to process the weights 230 of the cue words 225 and output one or more looping sound files to the audio track library 240.

[0060] According to the exemplary embodiment, the diffusion model is the preferred deep learning model 235. However, it should be noted that other suitable deep learning models 235 can be combined while maintaining the spirit and function of the invention. Using the diffusion model, by applying denoising to already created audio tracks or pre-loaded libraries, the model's results can be more nuanced and more controlled by design. The diffusion model can also be more conducive to allowing users to share or download "tuning".

[0061] According to an exemplary embodiment, the audio track library 240 may have a low processing priority. According to an exemplary embodiment, the audio track library 240 may include a collection of sounds manipulated by one or more ASD dynamic curves. According to an exemplary embodiment, the wave synthesis ASD module 245 may have a high processing priority. According to an exemplary embodiment, the system may be configured to allocate one or more processing resources based on the refresh rate, wherein low processing priority is allocated to the SRRI and high processing priority is allocated to the FRRI.

[0062] ASD is a wave synthesis type, therefore it requires audio track library 240.

[0063] According to an exemplary embodiment, FRRI 215 can be configured to determine the dynamic gain and filter values ​​of an audio track, as in conventional ASD. According to an exemplary embodiment, each track in the track library 240 may include multiple curves for each FRRI 215.

[0064] According to an exemplary embodiment, the wave synthesis ASD module 245 can be configured to output synthesized dynamic system sound 250 and / or other suitable sound.

[0065] Now for reference Figure 3The present invention provides an exemplary description of a method 300 for ASD generation according to an exemplary embodiment of the present invention.

[0066] In 310, a user can directly input one or more feature ASD settings using a user interface (e.g., a vehicle infotainment system), and in 315, the system can read one or more SRRIs.

[0067] In step 320, the weights of the deep learning model can be set using one or more SRRIs, and in step 325, the deep learning model can be configured to create / generate one or more audio tracks to populate the audio track library. It should be noted that, according to an exemplary embodiment, the deep learning model can be configured to create audio tracks even if the user does not have ASD enabled. According to an exemplary embodiment, steps 320 and / or 325 can have low processing priority. According to an exemplary embodiment, one or more audio tracks can include one or more looped sound files. According to an exemplary embodiment, the deep learning model can be configured to generate audio tracks for the audio track library when appropriate processing power and memory are available. Since the generation of new audio tracks is not critical to the driver's dynamic perception, new audio tracks can be generated at a slower rate than the output of the wave synthesis module.

[0068] At 305, according to the exemplary implementation, the user can enable the ASD. At 330, after the user enables the ASD, the user can drive the vehicle normally, and at 335, the system can read one or more FRRIs.

[0069] According to an exemplary embodiment, at 340, using audio tracks from a track library and FRRI, the system can construct a sound (e.g., synthesize a powertrain sound and / or other suitable sound). Then, at 345, the sound can be played through one or more speakers in the vehicle. According to an exemplary embodiment, step 340 can have a high processing priority.

[0070] According to an exemplary embodiment, the system of the present invention may include a linking system for users to share their audio track libraries and hidden cue words from deep learning diffusion models. This adds a social component to the system and allows customers to feel a sense of ownership over their own voices.

[0071] For example, such as Figure 4 As shown, an exemplary process 400 for active voice sharing according to an exemplary embodiment of the present invention is depicted.

[0072] According to an exemplary embodiment, a software system can be provided that synthetically generates a time-based FRRI series to produce a demonstration of user-generated sounds. This would allow users to showcase their powertrain sounds to other users before they decide to upload their powertrain sounds to their "Our Vehicles" section.

[0073] For example, a first user (user #1) can use a deep learning ASD system 405 to generate one or more saved audio recordings 410, and a second user (user #2) can use a deep learning ASD system 415 to generate one or more saved audio recordings 420. The system can be configured to enable users to share audio via, for example, an active audio sharing hub 425. According to an exemplary embodiment, the active audio sharing hub 425 can be an Internet-based system and / or other suitable systems.

[0074] According to an exemplary implementation, the saved and / or shared sounds may include a soundtrack library, model cue words and weights, and / or sound seeds.

[0075] According to an exemplary embodiment, the active voice sharing hub 425 can be configured to facilitate the use of a voice-sharing social system and / or enable one or more users to rate shared voices. According to an exemplary embodiment, rating shared voices can increase the visibility of the rated voices within the voice-sharing social system. This will allow users to find the best voices.

[0076] According to an exemplary embodiment, the system can be configured to enable users to save their sound libraries and inputs to machine learning diffusion models. According to an exemplary embodiment, the system can be configured to enable users to use their saved favorite dynamic system sounds as a basis to create one or more new sounds.

[0077] Now for reference Figure 5 An exemplary vehicle system architecture 500 of a vehicle according to an exemplary embodiment of the present invention is provided. The following discussion of the vehicle system architecture 500 is sufficient to understand one or more components of the vehicle 100.

[0078] like Figure 5As shown, vehicle system architecture 500 may include an engine 502, an electric motor or propulsion device, and various sensors 504 to 518 for measuring various parameters of vehicle system architecture 500. In a gas-powered or hybrid vehicle with a fuel-powered engine, for example, sensors 504 to 518 may include an engine temperature sensor 504, a battery voltage sensor 506, an engine rotations per minute (RPM) sensor 508, and / or a throttle position sensor 510. If the vehicle is an electric vehicle or a hybrid vehicle, the vehicle may include an electric motor, and accordingly, for example, sensors 504 to 518 may include sensors such as a battery monitoring system 512 (to measure battery current, voltage, and / or temperature), a motor current sensor 514 and a motor voltage sensor 516, and a motor position sensor 518 such as a resolver and an encoder.

[0079] For example, operating parameter sensors common to both types of vehicles may include: a position sensor 534 (e.g., an accelerometer, gyroscope, and / or inertial measurement unit); a speed sensor 536; and / or an odometer sensor 538. The vehicle system architecture 500 may also include a clock 542, enabling the system to determine vehicle time and / or date during operation. The clock 542 may be encoded into an onboard computing device 520 (which may be a separate device), or multiple clocks may be used.

[0080] Vehicle system architecture 500 may include various sensors that operate to collect information about the environment in which the vehicle is traveling. For example, these sensors may include: a position sensor 544 (e.g., a Global Positioning System (GPS) device); object detection sensors (e.g., one or more cameras 546); a LiDAR sensor system 548; and / or a radar and / or sonar system 550. Sensors may include an environmental sensor 552, such as a humidity sensor, precipitation sensor, light sensor, and / or an ambient temperature sensor. The object detection sensor may be configured to enable vehicle system architecture 500 to detect objects within a given distance range of the vehicle in any direction, while the environmental sensor 552 may be configured to collect data about environmental conditions within the area where the vehicle is traveling. According to an exemplary embodiment, vehicle system architecture 500 may include one or more lights 554 (e.g., headlights, floodlights, strobes, etc.).

[0081] During operation, information can be transmitted from sensors to an onboard computing device 520 (e.g., computing device 115, computing equipment 600). The onboard computing device 520 can be configured to analyze data captured by sensors and / or received from data providers, and optionally, can be configured to control the operation of the vehicle system architecture 500 based on the analysis results. For example, the onboard computing device 520 can be configured to perform the following controls: controlling braking via a brake controller 522; controlling direction via a steering controller 524; and controlling speed and acceleration via a throttle controller 526 (in a gas-powered vehicle), an electric motor speed controller 528 (e.g., a current level controller in an electric vehicle), a differential transmission controller 530 (in a vehicle with a transmission), and / or other controllers. The brake controller 522, as described herein, may include a pedal force sensor, a pedal angle sensor, and / or a simulator temperature sensor.

[0082] Geographic location information can be transmitted from location sensor 544 to onboard computing device 520. Onboard computing device 520 can then access a map of the environment corresponding to the location information to determine known fixed features of the environment, such as streets, buildings, stop signs, and / or stop / drive signals. Images captured from camera 546 and / or object detection information captured from sensors such as LiDAR 548 can be transmitted from those sensors to onboard computing device 520. The object detection information and / or captured images can be processed by onboard computing device 520 to detect objects near the vehicle. Any known or potentially known techniques for object detection based on sensor data and / or captured images can be used in the embodiments disclosed in this document.

[0083] Now for reference Figure 6 This document provides an illustration of an exemplary architecture for computing device 600. According to exemplary embodiments, one or more functions of the present invention may be implemented by a computing device such as computing device 600 or a computing device similar to computing device 600. Computing device 600 may be a quantum computer, a classical computer, and / or have one or more components configured to perform one or more quantum and / or classical computing functions. Computing device 115 and / or vehicle-mounted computing device 520 may be examples of computing device 600 and / or may include one or more components of computing device 600.

[0084] Figure 6 The hardware architecture represents an exemplary implementation of a representative computing device configured to implement at least a portion of the systems / devices (e.g., vehicle 100) and method / control logic (e.g., process 200, method 300, and process 400) as described herein.

[0085] Some or all components of computing device 600 may be implemented as hardware, software, and / or a combination of hardware and software. Hardware may include, but is not limited to, one or more electronic circuits. Electronic circuits may include, but are not limited to, passive components (e.g., resistors and capacitors) and / or active components (e.g., amplifiers and / or microprocessors). Passive and / or active components may be adapted, arranged, and / or programmed to perform one or more of the methods, processes, or functions described herein.

[0086] like Figure 6 As shown, computing device 600 may include a user interface 602 (e.g., a graphical user interface), a central processing unit (CPU) 606, a system bus 610, a memory 612, and a hardware entity 614. The memory 612 is connected to and accessible by other parts of computing device 600 via the system bus 610, and the hardware entity 614 is connected to the system bus 610. The user interface may include input and output devices configured to facilitate user-software interaction for controlling the operation of computing device 600. Input devices may include, but are not limited to, a physical and / or touch keyboard 640. Input devices may be connected via wired or wireless connections (e.g., […]). (Connection) is connected to computing device 600. Output devices may include, but are not limited to, a speaker 642, a display 644, and / or a light-emitting diode 646.

[0087] At least a portion of hardware entity 614 may be configured to perform actions involving accessing and utilizing memory 612, which may be random access memory (RAM), a disk drive and / or compact disc-only memory (CD-ROM), and other suitable memory types. Hardware entity 614 may include a disk drive unit 616, which includes a computer-readable storage medium 618 on which a set of one or more instructions 620 (e.g., programming instructions, such as, but not limited to, software code) may be stored, one or more instructions 620 configured to implement one or more of the methods, processes, or functions described herein. During execution of instructions 620 by computing device 600, instructions 620 may also reside wholly or at least partially within memory 612 and / or CPU 606.

[0088] Memory 612 and CPU 606 may also constitute a machine-readable medium. As used herein, the term "machine-readable medium" refers to a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store a set of one or more instructions 620. The term "machine-readable medium" as used herein also refers to any medium capable of storing, encoding, or carrying a set of instructions 620 for execution by computing device 600 of the instructions 620 and causing computing device 600 to perform any or more methods of the present invention. According to various embodiments, one or more computer application programs 624 may be stored in memory 612.

[0089] The foregoing description includes examples of the subject matter invention. Of course, for the purposes of describing the subject matter, it is impossible to describe every conceivable combination of components or methods; however, it should be understood that many other combinations and arrangements of the subject matter invention are possible. Accordingly, the claimed subject matter is intended to cover all such alternatives, modifications, and variations falling within the spirit and scope of the appended claims.

[0090] Specifically, and with regard to the various functions performed by the aforementioned components, devices, systems, etc., unless otherwise indicated, the terminology used to describe such components (including references to “devices”) is intended to correspond to any component that performs the specified function of the described component (e.g., functionally equivalent), even if it is not structurally equivalent to the disclosed structure, and can perform that function in the exemplary aspects of the claimed subject matter described herein.

[0091] The interactions between several components have been described in the foregoing system and components. It is understood that such a system and components may include those components or designated sub-components, designated components or portions of sub-components, and / or additional components, and, depending on the various permutations and combinations described above, sub-components may also be implemented as components communicatively connected to other components, rather than being included within a parent component (hierarchical). Furthermore, it should be noted that one or more components may be combined into a single component providing a set of functionalities, or divided into several separate sub-components. Any component described herein may also interact with one or more other components not specifically described herein.

[0092] Furthermore, while specific features of the subject invention have been disclosed with respect to only one of several embodiments, such features may be combined with one or more other features in other embodiments, such as one or more other features in other embodiments that may be desirable and advantageous for any given or particular application. Moreover, with respect to the use of the terms “comprising,” “including,” “having,” “containing,” variations thereof, and other similar words in the specific embodiments or claims, these terms are intended to be inclusive in a manner similar to the term “comprising” as an open transitional phrase, without excluding any additional or other elements.

[0093] Therefore, the embodiments and examples set forth herein are presented in order to best explain various alternative embodiments of the invention and their particular applications, thereby enabling those skilled in the art to make and utilize embodiments of the invention. However, those skilled in the art will recognize that the foregoing descriptions and examples are presented for illustrative and exemplary purposes only. The descriptions set forth are not intended to be exhaustive or to limit the embodiments of the invention to the precise forms disclosed.

Claims

1. A system for active sound design generation, comprising: One or more speakers; and A computing device includes a processor and a memory, wherein the memory is configured to store instructions, which, when executed by the processor, are configured to cause the processor to: Receive one or more inputs for synthesizing sound; Classify one or more inputs as fast refresh rate inputs or slow refresh rate inputs; One or more processing resources are allocated based on the refresh rate, where: Low processing priority is assigned to slow refresh rate inputs; High processing priority is assigned to fast refresh rate inputs; Active sound design based on one or more inputs, wherein the generation includes: Using a slow refresh rate input and one or more cue words to change one or more weights of a deep learning model, wherein the cue words are the input to the deep learning model; Using deep learning models to process weights and cue words to output one or more looping sound files to an audio track library, thereby forming one or more audio tracks; Use fast refresh rate input to change one or more dynamics of the wave synthesis active sound design module; Synthetic sound is generated using one or more audio tracks and wave synthesis active sound design modules; Synthetic sound is played on one or more speakers.

2. The system for active sound design and generation according to claim 1, wherein, The fast refresh rate input includes one or more inputs selected from the following: Throttle position; Motor speed; Wheel speed; Brake position; Vehicle gravity; and Motor load.

3. The system for active sound design and generation according to claim 1, wherein, The slow refresh rate input includes one or more inputs selected from the following: Time of day; One or more calendar dates; Location; Driving modes; weather; Traffic conditions; Aggressiveness; Complexity; Musicality; One or more stored individual model weights; One or more shared model weights; as well as User-initiated voice design history.

4. The system for active sound design and generation according to claim 1, wherein, The deep learning model includes a diffusion model.

5. The system for active sound design and generation according to claim 1, wherein, The prompt words are defined by the way the deep learning model is built and trained.

6. The system for active sound design and generation according to claim 1, wherein, The synthesized sound is the sound of a synthesized power system.

7. The system for active sound design generation according to claim 1, further comprising a vehicle; in, One or more speakers are connected to the vehicle.

8. The system for active sound design and generation according to claim 1, wherein, The one or more audio tracks are controlled by one or more active sound design dynamic curves.

9. The system for active sound design and generation according to claim 8, wherein, Each of the one or more audio tracks includes multiple active sound design dynamic curves for each fast refresh rate input.

10. The system for active sound design generation according to claim 1, wherein, When the processor executes the instructions, the instructions are further configured to enable the processor to: enable the first user to share one or more tracks of the first audio track library with the second user.

11. A method for active sound design generation, comprising: Receive one or more inputs for synthesizing sound; Classify one or more inputs as fast refresh rate inputs or slow refresh rate inputs; One or more processing resources are allocated based on the refresh rate, where: Low processing priority is assigned to slow refresh rate inputs; High processing priority is assigned to fast refresh rate inputs; Using a computing device including a processor and memory, an active sound design is generated based on one or more inputs, wherein the generation includes: Using a slow refresh rate input and one or more cue words to change one or more weights of a deep learning model, wherein the cue words are the input to the deep learning model; Using deep learning models to process weights and cue words to output one or more looping sound files to an audio track library, thereby forming one or more audio tracks; Use fast refresh rate input to change one or more dynamics of the wave synthesis active sound design module; Synthetic sound is generated using one or more audio tracks and wave synthesis active sound design modules; Synthetic sound is played on one or more speakers.

12. The method for active sound design generation according to claim 11, wherein, The fast refresh rate input includes one or more inputs selected from the following: Throttle position; Motor speed; Wheel speed; Brake position; Vehicle gravity; and Motor load.

13. The method for active sound design generation according to claim 11, wherein, The slow refresh rate input includes one or more inputs selected from the following: Time of day; One or more calendar dates; Location; Driving modes; weather; Traffic conditions; Aggressiveness; Complexity; Musicality; One or more stored individual model weights; One or more shared model weights; as well as User-initiated voice design history.

14. The method for active sound design generation according to claim 11, wherein, The deep learning model includes a diffusion model.

15. The method for active sound design generation according to claim 11, wherein, The prompt words are defined by the way the deep learning model is built and trained.

16. The method for active sound design generation according to claim 11, wherein, The synthesized sound is the sound of a synthesized power system.

17. The method for active sound design generation according to claim 11, wherein, One or more speakers are connected to the vehicle.

18. The method for active sound design generation according to claim 11, further comprising: Manipulate one or more audio tracks by designing dynamic curves using one or more active sound designs.

19. The method for active sound design generation according to claim 18, wherein, Each of the one or more audio tracks includes multiple active sound design dynamic curves for each fast refresh rate input.

20. The method for active sound design generation according to claim 11, further comprising: This allows the first user to share one or more tracks from the first audio track library with the second user.