Music motivation hierarchical design method and system for multi-working-condition active sound of automobile
By employing a layered design method based on musical motifs, the harmony and stability issues of active sound in electric vehicles under various operating conditions were resolved. This approach enabled unified and efficient sound design across multiple operating conditions, thereby enhancing the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-03-13
Smart Images

Figure CN121662001A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle acoustics technology, and in particular relates to a method and system for layering musical motifs in active sound for automobiles under multiple operating conditions. Background Technology
[0002] With the increasing popularity of electric vehicles, the focus of in-vehicle acoustic environment has gradually shifted from noise suppression to the active construction and control of sound. Active sound control technology shapes the acoustic atmosphere created during vehicle operation by introducing controllable sound sources inside the vehicle. This not only enhances the driving experience and sense of security but also, to a certain extent, meets users' comprehensive needs for comfort and emotional experience.
[0003] Existing active sound technologies for electric vehicles generally focus on sound generation or playback control. They typically drive a preset sound source to vary based on operating parameters such as vehicle speed, motor speed, or acceleration to achieve active in-vehicle sound. In practice, most solutions follow the sound logic of traditional gasoline vehicles, using the internal combustion engine's order tones or their variations as the primary sound source, generating corresponding sounds through simulation, sampling, or spectrum reconstruction. For example, CN110525364B discloses an active sound system for electric vehicles and its sound control method, which calculates a virtual internal combustion engine speed based on the vehicle's speed or motor speed, and emits corresponding internal combustion engine order tones based on this virtual speed. This type of technology is relatively straightforward in engineering implementation and can enhance the feedback during driving to some extent. However, its sound performance is limited by the characteristics of internal combustion engine sounds, resulting in a relatively monotonous overall style and failing to reflect the independent design space of electric vehicles at the sound level. In addition, existing active sound solutions mostly focus on adjusting the physical parameters of sound, such as frequency, amplitude or order changes, while paying insufficient attention to the structural characteristics of sound as an auditory object itself, and lacking a systematic consideration of the audibility, harmony and long-term listening acceptance of sound from the sound source level.
[0004] As electric vehicles gradually become the primary mode of transportation for users, the sound of traditional internal combustion engines is no longer a necessary reference point for active sound design for users lacking experience in driving internal combustion engines. Music, with its inherent harmony, scene adaptability, and personalized expressive characteristics, is suitable for use as a sound source in the field of active sound generation for electric vehicles. In musicology, variations are based on a core musical motif (i.e., the smallest musical fragment), which develops into a complete piece of music. Its logic of "unified core motif + derived variation" is highly consistent with the essential characteristic of active sound in electric vehicles that dynamically changes with operating conditions while maintaining overall consistency. In the field of musicology, composers create musical motifs through two paths: one is intuitive creation based on artistic inspiration, capturing the core of the motif through immediate artistic perception; the other is a scene-oriented rational design logic, first anchoring the scene style and emotional needs, and then determining the core sound elements such as tonality, timbre, and rhythmic framework. Different keys carry different emotional meanings due to differences in pitch structure (e.g., C major is bright and open, suitable for uplifting emotions, while D major is bright and passionate, fitting for a striving scene). Different timbres adapt to different texture needs due to differences in medium properties (e.g., strings are warm and suitable for lyrical scenes, while brass is powerful and suitable for solemn contexts). Ultimately, these differences are refined into a core musical motif that can be developed and varied.
[0005] To translate this logical alignment into a practically applicable in-vehicle sound source, the key lies in developing musical motifs tailored to the diverse driving conditions of automobiles. However, existing music-based active sound designs for electric vehicles suffer from the following shortcomings: First, driving conditions can be categorized into acceleration, deceleration, and constant speed. Drivers and passengers exhibit significant differences in their perception of acceleration, power, and emotional needs across these three conditions, requiring differentiated sound expression while maintaining overall harmony to create a unified brand image and prevent driver and passenger distraction. Existing technologies struggle to achieve a balance between these two aspects. Second, the sound source parameters for in-vehicle active sound encompass multiple dimensions, including note combinations, tonality, pitch, time signature, tempo, duration, and timbre. Simultaneous design of these multiple parameters can easily lead to a "parameter explosion" problem, resulting in low design efficiency.
[0006] In view of this, the present invention provides a layered design scheme for the music motif of active sound under multiple operating conditions in automobiles, and uses the music motif as the basic sound source of active sound, thereby solving the above-mentioned problems of the prior art while meeting consumers' strong demand for personalized and diversified in-vehicle sound quality. Summary of the Invention
[0007] This invention provides a method and system for layered design of musical motifs for active sound in automobiles under multiple operating conditions, the specific technical contents of which are as follows: A method for hierarchical design of music motifs for active sound in automobiles under multiple operating conditions, constructing music motifs adapted to vehicle driving conditions and using them as the basic sound source for in-vehicle active sound, the construction of the music motifs includes the following steps: S1. Perform note design to determine the note sequence of the musical motif. A set of candidate note sequences is generated based on a preset note sequence length, and corresponding note design sound samples are generated for each candidate note sequence under at least one set of preset music parameter constraints. Each note design sound sample is evaluated and screened to determine the target note sequence used for the music motif. S2. Perform parameter coordination to determine the musical parameters of the musical motif, excluding the note sequence. While keeping the target note sequence unchanged, candidate parameter configurations are constructed for each music parameter and corresponding parameter coordination sound samples are generated. The parameter coordination sound samples are evaluated and screened to determine the target music parameter configuration used for the music motif. S3. Perform timbre matching to determine the timbre of the musical motif for different operating conditions. Collect sound samples from candidate instruments or synthesizers and extract the amplitude envelope ADSR curve of each sound sample. Based on the auditory requirements of different working conditions and the duration of the Attack, Decay, Sustain, and Release stages in the ADSR curve, select target timbres from candidate instruments or synthesizers for acceleration, deceleration, and constant speed working conditions. S4, Perform sound source generation Based on the established note sequence, musical parameters, and timbre, corresponding musical motifs are generated to serve as the basic sound sources for in-vehicle active sound under acceleration, deceleration, and constant speed conditions, respectively.
[0008] Further, in step S1: for the preset note sequence length, each note position is used as a test factor, and the seven basic notes of the modern music system, do, re, mi, fa, sol, la, si, are used as the level set of this factor. Several candidate note sequence sets are generated by Latin hypercube sampling. The preset music parameters include: setting the mode to C major, setting the scale degree to the fifth degree, setting the time signature to 3 / 4 time, setting the tempo to 60 beats / minute, and setting the note value to 3 quarter notes.
[0009] Furthermore, in step S1: the evaluation and screening adopts a listening and review test method to organize a review panel to subjectively score each note design sound sample, and selects the highest-scoring note design sound samples based on the average score; in order to reduce the impact of playback loudness differences on the listening experience, each note design sound sample needs to be uniformly normalized in loudness or peak value before the listening and review test so that its average output level at the playback device is consistent.
[0010] Furthermore, the selected note design sound samples are further discriminated by combining objective psychoacoustic evaluation indicators, including at least loudness, sharpness, and roughness, to determine the target note sequence of the musical motif; the loudness is used to characterize the strength of the sound in the subjective perception of the human ear, the sharpness is used to describe the concentration of sound energy in the high-frequency region, and the roughness is used to characterize the unsmoothness and harshness caused by rapid amplitude modulation in the sound.
[0011] Furthermore, in step S2: First, the coordination design of modes and pitches is carried out: based on the determined note sequence, it is mapped to different candidate modes. For each candidate mode, each note in the note sequence needs to be converted into a pitch representation under the corresponding mode, and a corresponding sound sample is generated. Through subjective scoring and objective psychoacoustic evaluation indicators in the listening test, the mode and pitch combination with excellent performance is selected and used as the basic pitch framework of the musical motif. Then, the time signature, tempo, and note value are coordinated and designed separately. While keeping the established note sequence, mode, and pitch unchanged, the appropriate settings for each musical parameter are selected individually through the method of controlling variables and a combination of subjective and objective evaluation. Finally, based on the determined note sequence, mode, pitch, time signature, tempo, and note value, corresponding parameter coordination sound samples are generated. By subjectively evaluating and objectively analyzing these sound samples again, it is verified whether each musical parameter still maintains a good synergistic effect in the combined state. If the combined effect is found to be unsatisfactory, fine-tuning can be performed within the determined range of musical parameters until a stable musical motif parameter configuration with overall auditory performance is obtained.
[0012] Furthermore, in step S3, in the amplitude envelope ADSR curve, the Attack phase is used to describe the rising process of sound from the beginning to the peak, the Decay phase is used to describe the falling process after the peak to the stable segment, the Sustain phase is used to describe the continuous process of maintaining stability, and the Release phase is used to describe the release process of sound decaying from the sustain segment to the end. By comparing the corresponding ADSR curves of each candidate instrument or synthesizer, the instrument or synthesizer sound with the shortest duration of the Attack phase is identified as the acceleration tone that emphasizes response speed and dynamics, the instrument or synthesizer sound with the longest duration of the Release phase is identified as the deceleration tone that emphasizes mellowness and soothingness, and the instrument or synthesizer sound with the longest duration of the Sustain phase is identified as the constant speed tone that emphasizes smoothness and continuity.
[0013] Preferably, in step S3, in order to ensure the consistency of the active sound under multiple working conditions in terms of style, the principle of selecting the same type of instrument is adopted when determining the timbre corresponding to each working condition. The three timbres used for acceleration, deceleration and constant speed working conditions come from the same instrument category, so as to avoid the timbre difference being too large when switching working conditions, which would cause auditory fragmentation. The instrument category includes percussion instruments, bowed string instruments, wind instruments and plucked instruments.
[0014] Further, in step S3, the step of obtaining the amplitude envelope ADSR curve includes: Acquire single-shot samples of candidate musical instruments or synthesizers and discretize them to obtain time-domain signals; The time-domain signal is normalized and filtered to obtain the filtered signal. ; Constructing analytic signals And take its modulus value to obtain the envelope signal. ; The envelope signal is smoothed, and the smoothed signal is used as the amplitude envelope ADSR curve. in, Represents the Hilbert transform. j The imaginary unit, It is a complex modulus.
[0015] Furthermore, dynamic thresholds are used to determine the boundaries of each stage in the ADSR curve: The start of the Attack phase Defined as the first sampling point in the ADSR curve that meets the preset zero-crossing / oscillation start-up conditions, and the end point of the Attack phase. Defined as the first peak sampling point in the ADSR curve; the Decay phase begins from the end of the Attack phase. To the point ,point Defined as the first sampling point in the ADSR curve when it falls back to the first set threshold level from the first peak; the Sustain phase starts from point To the point ,point Defined as the sampling point where the amplitude of the ADSR curve first falls below the second set threshold; the Release phase starts from point... To the point ,point Defined as the sampling point in the ADSR curve where the amplitude first falls below the termination threshold; The duration of each stage is calculated as follows:
[0016]
[0017] in, , , and These represent the durations of the Attack, Decay, Sustain, and Release phases, respectively. The sampling rate.
[0018] A hierarchical design system for music motifs based on the above method includes a subjective and objective evaluation and screening unit, a note sequence determination unit, a music parameter determination unit, a timbre matching unit, and a sound source synthesis and output unit, wherein: The subjective and objective evaluation and screening unit includes: a subjective scoring acquisition module, used to organize subjective listening evaluation of sound samples and obtain subjective scores; an objective index module, used to calculate objective indicators for sound samples, including at least loudness, sharpness and roughness; and a fusion screening decision module, used to jointly judge the subjective scores and objective indicators to determine the optimal sound sample. The note sequence determination unit includes: a candidate note sequence generation module, used to generate a set of candidate note sequences under a preset note sequence length constraint, and to use experimental design sampling to filter the note combination space to obtain a comprehensive candidate set; a candidate sample rendering and export module, used to render each candidate note sequence into a playable candidate sound source sample file under the condition of fixing at least one set of music parameters other than the note sequence; and a target note sequence determination module, used to filter each candidate sound source sample file through the subjective and objective evaluation and filtering unit, and to take the note sequence corresponding to the optimal candidate sound source sample as the target note sequence. The music parameter determination unit includes a parameter coordination module, which is used to construct candidate parameter configurations and generate corresponding samples for the mode, pitch, time signature, tempo, and time value parameters respectively, while keeping the target note sequence unchanged; and a target music parameter determination module, which is used to filter each of the samples through the subjective and objective evaluation and screening unit, and take the music parameters corresponding to the optimal sample as the target music parameters. The timbre matching unit includes: a timbre acquisition and timbre library management module, used to acquire sound samples of candidate instruments or synthesizers through an audio input / output interface and store them as a timbre sample library; an ADSR curve generation module, used to perform preprocessing and envelope extraction on the timbre samples and generate ADSR curves and their duration parameters for each stage; and a working condition timbre matching module, used to establish working condition adaptation rules based on the duration parameters of each stage of the ADSR curve, and to select target timbres for corresponding acceleration, deceleration, and constant speed working conditions from the timbre sample library respectively. The sound source synthesis and output unit includes: a multi-condition sound source synthesis module, used to synthesize the target note sequence and the target music parameters with the target timbre of each condition to form a multi-condition sound source; and a multi-condition sound source output module, used to perform loudness and / or peak normalization on the output sound source and output it through the audio input / output interface as a sound source for use by the vehicle active sound system.
[0019] Compared with the prior art, the present invention has at least the following beneficial effects: This invention systematically designs the active sound of electric vehicles from the sound source level, breaking through the existing technology's active sound construction method which mainly relies on single internal combustion engine sound materials or simple parameter mapping. It achieves a unified design of active sound under multiple operating conditions. Compared to existing technologies that rely on internal combustion engine order tones or isolated sound samples for sound control, this invention uses structured musical motifs as the basic sound source. This ensures that the active sound maintains a consistent overall auditory style across different driving conditions such as acceleration, deceleration, and constant speed, effectively avoiding abrupt or disjointed sound changes during condition transitions. This significantly improves the continuity and stability of the in-vehicle active sound. Furthermore, because it does not rely on traditional internal combustion engine sound characteristics, this invention provides electric vehicles with a more diverse and customizable selection of active sound sources, which is beneficial for enhancing product differentiation and user experience.
[0020] Furthermore, this invention introduces a progressive, layered design architecture of note design, parameter coordination, and timbre matching during the sound source design process. This transforms the active sound construction process from empirical selection into a calculable and reproducible technical flow, while avoiding the "parameter explosion" problem caused by multi-parameter synchronous design, significantly improving design efficiency. By systematically screening musical motif note combinations and then coordinating the configuration of mode, pitch, rhythmic structure, and time parameters, the generated active sound exhibits better harmony and acceptability under long-term playback conditions. Simultaneously, through a timbre matching method based on amplitude envelope characteristics, different timbres are correlated with vehicle driving conditions, ensuring that the active sound meets the directional requirements of the driving conditions while avoiding auditory discomfort caused by an overly dispersed timbre style.
[0021] Furthermore, this invention combines subjective auditory evaluation with objective psychoacoustic indicators to comprehensively optimize candidate sound sources, thereby effectively controlling key indicators such as loudness, sharpness, and roughness of active sounds, thus reducing the stimulation to the human ear and the risk of auditory fatigue during long-term use.
[0022] The layered design of the music motifs in this invention (finished sound sources for three operating conditions) can be directly adapted to the automotive application requirements of embedded devices. The embedded devices only need to retrieve the corresponding finished sound source from the storage module based on the vehicle's operating condition signals (such as vehicle speed and acceleration) for playback. This mode, which completes the adaptation optimization in advance during the design phase, eliminates the need for any complex calculations by the embedded devices, requiring only simple signal matching and retrieval. This reduces the software design complexity of the embedded devices and improves the stability of automotive applications.
[0023] In summary, this invention has significant technical advantages over existing technologies in terms of adaptability to multiple working conditions, sound consistency, auditory comfort, sound source designability, and design efficiency, and has good engineering application value. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0025] Figure 1 This is a schematic flowchart of the music motif layered design method provided in an embodiment of the present invention; Figure 2 This is an ADSR envelope curve diagram of the pipa provided in an embodiment of the present invention; Figure 3 This is the ADSR envelope curve diagram of Ruan provided in the embodiment of the present invention; Figure 4 This is an ADSR envelope curve diagram of the guzheng provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the hierarchical design system framework for music motifs provided in an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of the present invention.
[0027] Example 1 This embodiment provides a method for layered design of musical motifs for active sound in automobiles under multiple operating conditions.
[0028] Musical motifs can be characterized by parameters such as notes, modes, pitches, time signatures, tempos, durations, and timbre. To achieve harmonious unity of musical motifs under three typical driving conditions—acceleration, deceleration, and constant speed—this embodiment adopts a layered design strategy and combines subjective and objective sound quality evaluation methods to design musical motifs as active sound sources within the vehicle.
[0029] like Figure 1 As shown, the hierarchical design method for musical motifs described in this embodiment mainly includes the following steps: I. Note Design The design of musical motifs aims to construct the smallest sound segments that can reflect the basic characteristics of active sound. Its core lies in generating comparable candidate note sequences from a finite set of notes, and then filtering the candidate sequences under the condition of fixing other musical parameters, so as to determine the candidate note combinations used for subsequent multi-condition active sound sources.
[0030] First, the note set and note sequence length are determined. The note set uses the seven basic notes of the modern music system, represented by the letters C, D, E, F, G, A, and B, corresponding to do, re, mi, fa, sol, la, and si. To cover a sufficiently rich combination space within an acceptable experimental scale, while ensuring a controllable screening workload, this embodiment fixes the number of notes in the musical motif to six, so that each candidate musical motif consists of six notes in a specific order. This setting makes the note combination space large enough to reflect the auditory differences between different sequences, while also facilitating sampling screening through experimental design methods.
[0031] After determining the 6-note sequence, the candidate scheme generation stage begins. Since randomly combining the 6 note positions on the 7-note matrix would generate a massive number of candidate schemes (117,649), directly exhaustively searching would be too costly. Therefore, this embodiment employs a Latin hypercube experimental design method to screen the note combination samples, obtaining a finite candidate set that is evenly distributed and has good coverage within the combination space. Each note position is used as an experimental factor, and the seven note letters are used as the level set of that factor. Candidate sequences are generated through Latin hypercube sampling. In this embodiment, 30 candidate note sequences were ultimately determined, as shown in Table 1.
[0032] To ensure comparability between different candidate note combinations, after the candidate sequences are generated, other musical motif parameters besides the note sequences need to be standardized and fixed. In this embodiment, for the aforementioned 30 candidate note sequences, the key is uniformly set to C major, the pitch degree is uniformly set to the 5th degree, the time signature is uniformly set to 3 / 4 time, the tempo is uniformly set to 60 beats per minute (bpm), and the note duration is uniformly set to three quarter notes. Subsequently, a corresponding sound sample file is generated for each candidate note sequence.
[0033] In the sound sample generation stage, to ensure reproducibility, any common digital audio generation method can be used to convert the above music parameters into playable sound sources. For example, a unified timbre can be driven by a MIDI event sequence (e.g., uniformly using piano or simple sine / square wave synthesizer timbre to avoid timbre differences from interfering with each other), and exported as a WAV file with a uniform sampling rate (e.g., 44.1kHz or 48kHz) and a uniform bit depth (e.g., 16bit or 24bit). In specific implementation, the duration of each quarter note is converted to 1 second at 60 bpm, so the total duration of each candidate musical motif consisting of 3 quarter notes is 3 seconds. Then, the 6 notes are mapped to this 3-second timeline according to a predetermined rhythm (e.g., the 6 notes can be evenly distributed within 3 seconds, with each note lasting 0.5 seconds, or a unified rhythm template including legato / staccato can be used, but this must be consistent across all candidate schemes), thereby forming 30 candidate sound source files that can be directly compared. To reduce the impact of loudness differences on subjective listening experience, all samples can be normalized to a uniform loudness or peak value after generation, ensuring consistent average output levels on playback devices and avoiding the bias of "louder = better sound".
[0034] After the sound samples are generated, the subjective evaluation stage begins. The subjective evaluation employs a listening test method, organizing a judging panel to score 30 candidate samples. During implementation, judging members wear high-performance, high-fidelity headphones (e.g., Sennheiser HD650) and use a high-performance computer equipped with a subjective evaluation module (e.g., LMS Jury Test) for unified playback and scoring recording. The distribution of judging panel members in this embodiment is shown in Table 2. To reduce sequence effects and fatigue effects, the sample playback order can be randomized or balanced using a Latin square, and each judge is required to complete the listening test in a quiet environment. The scoring method uses a rating scale, with 11 levels from 0 to 10, and each level is associated with evaluation terms from "very bad" to "excellent" (as shown in Table 3) to ensure judges provide subjective scores under a unified semantic scale. After listening to all candidate samples according to this process, the average subjective score for each scheme is calculated. In this embodiment, the subjective scoring results for the 30 candidate samples are shown in Table 1, ultimately obtaining the three schemes with the highest subjective scores. Then, the optimal solution is determined by combining objective psychoacoustic evaluation indicators such as loudness, sharpness, and roughness.
[0035] Table 1. Candidate sequences of note combinations and corresponding subjective scoring results
[0036] Table 2 Composition of the Jury
[0037] Table 3 Rating Scale
[0038] In the objective evaluation stage, in order to make the screening results more stable and reliable, psychoacoustic objective indicators are introduced into the secondary judgment of subjective high-scoring schemes. The indicators used include at least loudness, sharpness and roughness.
[0039] Loudness is used to characterize the intensity of sound as perceived subjectively by the human ear, and can be calculated based on critical band theory. First, the input time-domain sound signal is decomposed into multiple critical bands (commonly set to 24 Bark critical bands) through a set of bandpass filters, and the specific loudness of each critical band is calculated. Then, the specific loudness of each critical band is integrated to obtain the overall loudness value, which is calculated as follows:
[0040] in, N The total loudness of a sound is expressed in one sone. Indicates the first z The specific loudness at each critical band is related to the sound pressure level within that band and the perceived weight of human hearing. According to the Zwicker loudness model, specific loudness is typically expressed as a nonlinear function of the sound pressure level exceeding the human hearing threshold, for example:
[0041] in, Indicates the first z The sound pressure level of the human ear corresponding to each critical band; It is a non-linear exponent, usually taken as 0.23~0.3. Indicates the first z The sound pressure level of each critical band is calculated as follows:
[0042] in, For reference sound pressure (20 μ Pa), Indicates the first z The root mean square value of the critical band signal, i.e.
[0043] That is, the sound sample The first bandpass filter obtained z Time-domain signals within the critical band:
[0044] in, Indicates the first z A bandpass filter operator for each critical band.
[0045] Sharpness describes the concentration of sound energy in the high-frequency region. Higher sharpness indicates a higher proportion of high-frequency components in the sound, resulting in greater stimulation to the human ear. Sharpness can be calculated based on loudness distribution, and its expression is:
[0046] in, S Sharpness is expressed in acum; z This indicates the location parameter of the critical zone, which usually corresponds to the critical zone number or the center frequency of the critical zone. This represents a weighting function related to the location of the critical band, used to enhance the weight of the high-frequency critical band in sharpness calculations. It typically employs a piecewise linear function.
[0047] in, This indicates the starting critical zone position for weighted calculation (typically a value of 15). k This represents the high-frequency enhancement factor (usually taken as 0.2). In the low-frequency and mid-frequency regions ( In the Bark region, the human ear's sensitivity to the sharpness of sound changes little, therefore the weight remains 1; while in the high-frequency region ( (Bark) As the frequency increases, the human ear becomes more sensitive to sharpness, requiring a linearly increasing weighting function to amplify the influence of this critical band.
[0048] Roughness is used to characterize the unevenness and harshness caused by rapid amplitude modulation in sound. It is closely related to the modulation frequency and modulation depth in the sound signal and can be obtained through modulation analysis of the signal. Its typical calculation form is:
[0049] in, R This represents surface roughness, measured in asper. Indicates the first i The modulation frequency corresponding to each modulation component Indicates the first i The modulation depth of each modulation component This represents the weighting coefficient related to the modulation frequency.
[0050] Specifically: Hilbert transform is used to construct sound samples. Analyzed signal , Represents the Hilbert transform. j The imaginary unit is used; then the amplitude envelope signal is obtained. The frequency spectrum is then obtained by performing a Fast Fourier Transform on the signal. Within the modulation frequency range (e.g., 15~300Hz), several significant modulation peaks in the spectrum are identified and denoted as M modulation components. i Each modulation component has a modulation frequency and modulation amplitude modulation depth Defined as the normalization of the modulation component amplitude relative to the envelope mean, i.e.
[0051] The human ear is sensitive to different modulation frequencies, typically being most sensitive around 70Hz. To reflect this characteristic, in this embodiment, the weighting coefficient... Adopt the following form:
[0052] in, This is the modulation frequency that the human ear is most sensitive to (70Hz). This is the frequency spread factor, typically taken as 30Hz.
[0053] After calculating the three objective indicators mentioned above, the subjective average score corresponding to each candidate scheme is jointly analyzed with the loudness, sharpness, and roughness data. If the subjective score is significantly better than other schemes, it can be directly selected as the preferred scheme; if the subjective scores of multiple schemes are similar, the scheme with lower roughness and lower sharpness is given priority, while also considering the scheme with loudness within a reasonable range, in order to avoid excessively strong or weak auditory perception.
[0054] Table 4. Subjective and objective evaluation results of the three schemes
[0055] Table 4 summarizes the subjective and objective evaluation results of the three candidate schemes with the highest subjective scores in this embodiment. It is clear that Scheme 3 has the best subjective sound quality. Furthermore, compared to other note combination schemes, Scheme 3 exhibits lower coarsness and sharpness, indicating that the musical motif designed in this scheme is less stimulating to the human ear and more easily accepted. Therefore, this embodiment ultimately determines the note combination corresponding to Scheme 3 as the note combination for the musical motif.
[0056] II. Parameter Coordination After determining the musical motif's note combination, to ensure good auditory consistency and acceptability under various driving conditions of electric vehicles, it is necessary to further coordinate and design other structural parameters of the musical motif. By configuring musical parameters such as mode, pitch, time signature, tempo, and time value, the musical motif can form a stable, controllable, and suitable basic sound source for in-vehicle active sound applications while maintaining its note structure.
[0057] Because users from different regions and cultural backgrounds have different perceptions of pitch relationships, melody direction, and rhythm, this embodiment introduces auditory preferences tailored to the target user group as a design constraint during the parameter coordination stage to avoid feelings of unfamiliarity or discomfort from musical motifs during long-term use in the vehicle. Taking electric vehicle applications in the Chinese market as an example, the parameter coordination process prioritizes modes and pitch structures that conform to the auditory habits of Chinese users, thus providing a clear direction for subsequent parameter combination selection.
[0058] Building upon this foundation, the coordination design step of mode and pitch class proceeds. Mode defines the interval relationships between notes and is a crucial factor influencing the overall harmony of a musical motif. In this embodiment, based on a predetermined note combination, it is mapped to different candidate modes for comparison, such as natural major, natural minor, and the traditional Chinese pentatonic scale. For each candidate mode, each note in the note combination is converted into a pitch class representation under the corresponding mode, generating a corresponding sound sample. When generating samples, the note order remains unchanged; only the mode and pitch class system they belong to are altered, ensuring that differences between samples originate solely from mode and pitch class settings. Through subjective auditory evaluation and objective acoustic index analysis of the samples, mode and pitch class combinations that perform better in terms of overall harmony, stability, and acceptability are selected and used as the basic pitch framework for the musical motif. The subjective and objective evaluation methods are the same as described above and will not be repeated here.
[0059] After determining the key and pitch, the next step is to coordinate the design of the time signature. The time signature describes the rhythmic organization of a musical motif, directly affecting the rhythmic and stable feel of the sound over time. In this embodiment, several commonly used time signatures are selected as candidates, such as 2 / 4, 3 / 4, and 4 / 4 time signatures. While maintaining the note combinations, key, and pitch, musical motif samples are constructed for each time signature. For each time signature setting, the distribution of notes within the beat must be uniformly planned to ensure a reasonable correspondence between the note starting point and the beat structure, avoiding a blurred or unstable rhythmic center. By comparing the auditory effects of different time signature samples under continuous playback conditions, the time signature setting with clear rhythm, high stability, and suitability for long-term playback in a vehicle is selected.
[0060] After the time signature is determined, the coordination and design of tempo parameters proceeds. Tempo, usually expressed in beats per minute (bpm), is a crucial parameter affecting the dynamic perception of musical motifs. In this embodiment, several representative tempo values are selected as candidates, such as 60 bpm, 90 bpm, and 120 bpm. With other parameters fixed, musical motif samples are generated at the corresponding tempos. By changing the tempo parameters, the changes in tension, relaxation, or stability of the same musical motif at different time scales can be observed. Through subjective evaluation and objective indicator analysis, tempo parameters that neither produce a sense of haste or unease due to excessive speed nor sound sluggishness or lack of presence due to excessive slowness are selected for subsequent multi-condition active sound applications.
[0061] After determining the tempo parameters, the next step is to coordinate and design the note values. Note values describe the duration of a single note on the time axis and their combinations, and are a crucial factor influencing the rhythmic subtlety of a musical motif. In this embodiment, based on the determined time signature and tempo parameters, various candidate note value combinations are constructed, such as a single eighth note, a dotted quarter note, a half note, or a composite note structure composed of multiple notes of different lengths. For each note value combination, it is essential to ensure that the musical motif ends naturally within a complete cycle, and that there are no obvious time breaks or rhythmic abrupt changes between notes. By comparing the auditory performance of different note value combinations under continuous playback conditions, a note value configuration scheme with clear rhythmic layers and good overall coherence is selected.
[0062] After each parameter has been individually screened, a comprehensive coordination and verification between them is required. Specifically, the selected key, pitch, time signature, tempo, and time value parameters are combined to form a complete musical motif parameter configuration scheme, and a corresponding sound sample is generated. This sample is then subject to further subjective evaluation and objective index analysis to verify whether each parameter maintains a good synergistic effect in the combined state, avoiding situations where a single parameter performs well but the overall listening experience deteriorates after combination. If the combination effect is found to be unsatisfactory, minor adjustments can be made within the established parameter range until a musical motif parameter configuration with stable overall auditory performance is obtained.
[0063] Table 5 Musical motifs after parameter coordination
[0064] Table 6. Subjective and objective evaluation results of musical motifs before and after parameter coordination
[0065] In this embodiment, the musical motif after parameter coordination and the corresponding subjective and objective evaluation results are shown in Tables 5 and 6, respectively. It can be seen that the subjective and objective sound quality of the musical motif is significantly improved after parameter coordination.
[0066] III. Timbre Matching To meet the active sound requirements of different driving conditions, it is necessary to further combine the sound-producing characteristics of traditional Chinese musical instruments and conduct timbre matching of musical motifs. Traditional Chinese musical instruments can be divided into four categories: percussion instruments, bowed string instruments, wind instruments, and plucked string instruments. Comparing the sound-producing characteristics of the four types of instruments, it can be seen that the sound of percussion instruments is too abrupt, the sound of wind instruments and bowed string instruments is too slow, and plucked string instruments are more suitable for the active sound requirements of different driving conditions.
[0067] In this embodiment, the amplitude envelope curve (ADSR) is used as a parameterized object to characterize the dynamic characteristics of an instrument's timbre. The ADSR consists of four phases: the Attack phase describes the rise of sound from its initial peak; the Decay phase describes the fall from the peak to the stable phase; the Sustain phase describes the sustained stability; and the Release phase describes the decay of sound from the sustain phase to its end. Since differences in structure and materials among different instruments result in varying ADSR phase durations, quantifying the duration of each ADSR phase allows for comparison of differences in auditory dimensions such as intensity, soothingness, and stability, enabling the matching of appropriate timbres for different driving conditions.
[0068] 1. Acquisition and preprocessing of musical instrument sound samples Sound samples from musical instruments are collected in an anechoic chamber using professional recording equipment at a sampling rate of 48kHz. After collection, the recorded samples need to be preprocessed, including amplitude normalization and bandpass filtering.
[0069] The purpose of amplitude normalization is to eliminate inconsistencies in amplitude scale caused by differences in recording levels or playing dynamics. Normalization can be achieved by the following formula:
[0070] in, This represents the discrete-time sampling points of the sound sample signal. n amplitude, This represents the maximum amplitude of the sample over the entire time range. This is the normalized signal obtained.
[0071] After normalization, to ensure the envelope extraction covers the instrument's fundamental frequency and its main harmonic components, while suppressing irrelevant low-frequency drift and ultra-high-frequency noise, a Butterworth bandpass filter is applied to the normalized signal. The filter's cutoff frequency range is set to 50Hz to 8kHz. Since different instruments have different sound frequencies, the cutoff frequency range can be flexibly adjusted according to the characteristics of each instrument to still cover its fundamental frequency and main harmonic range. The Butterworth filter is defined by its Z-transform, as shown below:
[0072] in, and Represents the coefficients of the numerator and denominator polynomials of the filter; P The filter order is set to 4 in this embodiment to ensure that the filter roll-off and phase characteristics meet the requirements for extraction.
[0073] 2. Envelope Extraction Based on Hilbert Transform To accurately characterize the change of amplitude over time, an analytic signal is constructed using the analytic signal method. It is composed of the filtered signal Its Hilbert transform constitutes, i.e. Instantaneous envelope Defined as the magnitude of the analytic signal, i.e. The instantaneous phase can be obtained from the amplitude of the analytical signal.
[0074] To avoid Gibbs oscillations in the envelope and improve boundary detection stability, the envelope is... Applying a moving average filter with a window length of 10ms yields a smooth envelope. The number of sampling points corresponding to the window length is denoted as N w Its value is determined by Calculations show that The sampling rate.
[0075] 3. Calculate the duration of each stage of ADSR. This embodiment uses dynamic thresholds to determine the boundaries of each stage.
[0076] The start of the Attack phase Defined as the first sampling point that meets the preset zero-crossing / oscillation start-up conditions, for example The first point is used to characterize the moment when vocalization begins; the end point of the Attack phase. Defined as the envelope reaching its global maximum value. The location of this point corresponds to the peak time. The end of the Decay phase. Defined as the first sampling point where the envelope falls back from the peak to a certain threshold level, for example... The first point; the end of the Sustain phase. Defined as the sampling point where the envelope enters the decay phase from the hold phase and first falls below another threshold level, for example... The first point; the end point of the Release phase is defined as the first sampling point when the envelope decays to near the background noise or below the termination threshold, for example The threshold coefficient introduced in the above boundary detection (such as...) and The settings can be flexibly adjusted according to the shape of the instrument envelope, but should remain consistent in the same round of instrument comparison to ensure comparability.
[0077] After the boundaries are determined, the duration of each of the four stages is calculated:
[0078]
[0079] in , , , , All are time indices based on the number of sampling points. To reduce the impact of random noise or differences in single performances, multiple samples can be collected for the same instrument under the same performance method. The duration of each sample can be calculated, and the mean or median can be used as the representative ADSR parameter for the instrument.
[0080] 4. Working condition tone matching After obtaining the ADSR stage duration parameters of the candidate instruments, the vehicle enters the condition timbre matching stage. In this embodiment, the auditory requirements of the vehicle under various conditions are abstracted into three categories: acceleration conditions emphasize response speed and power, deceleration conditions emphasize smoothness and comfort, and constant speed conditions emphasize stability and continuity.
[0081] Instruments with shorter attack durations have a faster onset and a more impactful sound, making them more suitable for acceleration-related tones; instruments with longer release durations have a smoother finish and a softer sound, making them more suitable for deceleration-related tones; and instruments with longer sustain durations have a more stable sustain and a more even sound, making them more suitable for steady-state tones. By comparing the durations of each candidate instrument, they can be ranked and selected according to their suitability for different tones.
[0082] To ensure stylistic consistency of the active sound across multiple operating conditions, this invention employs a principle of selectively choosing instruments of the same category when selecting candidate timbres. Specifically, the three timbres used for acceleration, deceleration, and constant speed are preferably from the same instrument category to avoid significant timbre differences during condition switching, which could lead to auditory disjointedness. In one embodiment targeting the Chinese market, candidate instruments can be limited to traditional Chinese folk instruments, further focusing on plucked string instruments as a single category. Within this category, three representative instruments are selected according to the aforementioned ADSR matching rules, used for acceleration, deceleration, and constant speed conditions, respectively. Based on this process, the pipa, ruan, and guzheng are determined as the musical motif timbres for acceleration, deceleration, and constant speed conditions, respectively. The ADSR envelope curves of these three instruments are shown below. Figures 2-4 As shown in the figure (for easier observation, the ADSR curves in the figure have been shifted upwards along the y-axis by a certain distance). The figure reveals that the ADSR characteristics of these three instruments exhibit common features of plucked string instruments, namely a fast attack phase, relatively short decay and sustain phases. Specifically, the pipa has a relatively shorter attack phase, the ruan's decay / release characteristics are more suitable for a slower, more decelerated listening experience, and the guzheng's sustain phase is relatively longer and more conducive to creating a uniform and stable listening experience.
[0083] IV. Sound Source Generation After timbre matching is completed, the timbre samples of the selected instruments are synthesized with the previously determined musical motif note combinations and their coordinated mode, pitch, time signature, tempo, and time parameters to form three sets of musical motif sound sources for acceleration, deceleration, and constant speed conditions, respectively. During synthesis, the note structure of the musical motifs is kept consistent; only timbre samples or synthesizer timbre mappings are replaced, and the output sound sources undergo uniform loudness / peak value normalization to avoid inconsistent in-vehicle experience caused by differences in the background loudness of different timbres. Subsequently, the in-vehicle active sound system calls the corresponding musical motif sound source based on the operating condition recognition results, and uses cross-fade-in / fade-out or parameter smooth transition methods to connect adjacent operating condition timbres during operating condition switching to ensure auditory continuity.
[0084] The above is the main content of the hierarchical design method for musical motifs described in this embodiment. Its core innovation lies in: (1) Establish the design logic with music motif as the core, ensure the harmony of the sound of the three working conditions and the consistency of the brand by unifying the core music motif, and achieve the recognition of the working conditions by differentiating the timbre. (2) Construct a three-layer progressive hierarchical design architecture to solve the parameter explosion problem: First layer (basic layer): Prioritize the determination of note combination parameters, which serve as the core components of musical motifs and provide a constraining basis for subsequent parameter design; The second layer (middle layer): Design non-core parameters such as tonality, pitch, time signature, tempo, and duration, and form a unified musical framework based on the combination of notes in the basic layer; The third layer (adaptation layer): The timbre parameters are designed based on the ADSR curve. The Attack, Decay, Sustain, and Release durations of the ADSR curve are directly related to the sound characteristics of the instrument and are precisely bound to the requirements of automotive operating conditions. That is, the acceleration condition is matched with a short Attack time (to enhance the sense of power), the constant speed condition is matched with a long Sustain time (to ensure stability), and the deceleration condition is matched with a long Release time (to improve the smoothness). (3) By combining a unified core motivation, hierarchical parameter design, and ADSR-operating condition mapping, this solution can simultaneously achieve the three major goals of harmony, differentiation, and high design efficiency. Moreover, the results of the hierarchical design (finished audio sources for three types of operating conditions) can be directly adapted to the automotive application requirements of embedded devices: the audio sources have been classified and produced according to acceleration / deceleration / uniform speed operating conditions during the design phase. The embedded device only needs to retrieve the corresponding finished audio source from the storage module for playback based on the vehicle operating condition signal (such as vehicle speed and acceleration). This mode of completing the adaptability optimization in advance during the design phase does not require the embedded device to perform any complex calculations. It only requires simple signal matching and retrieval, which reduces the software design complexity of the embedded device and improves the stability of automotive applications, adapting to the reliability requirements of harsh automotive environments.
[0085] Example 2 Based on the above method, this embodiment provides a hierarchical design system for music motifs in automotive multi-condition active sound. The system's core consists of a processing unit and a storage unit, along with audio acquisition and monitoring / playback hardware and external interfaces. This completes a closed-loop process from candidate note sequence generation, subjective and objective evaluation and selection, parameter coordination, timbre ADSR extraction and condition matching, to multi-condition sound source synthesis and export, supporting in-vehicle access. The processing unit can be a PC workstation or a server CPU / GPU. The storage unit stores the note set, candidate sequences, scoring records, psychoacoustic features, ADSR parameter tables, selection results from each round, and the final multi-condition sound source. The audio acquisition unit acquires sound samples from candidate instruments / synthesizers and writes them into a timbre library with a uniform sampling rate and bit depth. The monitoring / playback unit organizes subjective listening tests to ensure consistency between the playback link and loudness calibration, thereby making different candidate samples comparable. The external interface exports the final output acceleration, deceleration, and constant speed condition sound sources and their metadata to the in-vehicle active sound system or simulation platform, and provides switching transition strategy parameters to ensure auditory continuity.
[0086] like Figure 5 The diagram shown is a framework schematic of the music motif layered design system described in this embodiment.
[0087] The system first establishes a set of target driving conditions and defines a set of musical motif parameters. This parameter set includes at least note sequences, as well as key, pitch, time signature, tempo, duration, and timbre, for subsequent hierarchical design and constraint control. To reduce the combination space and increase candidate coverage, the system generates a finite candidate set during the note design phase using a candidate note sequence generation module. For example, each note position is considered an experimental factor, and the mother set of notes is considered a factor level set. Multiple sets of candidate note sequences are output using an experimental design sampling method. Subsequently, the candidate sample rendering and export module combines each set of candidate note sequences with the fixed structural parameters of the current round to generate a uniformly structured event sequence and render it as a playable audio sample. Simultaneously, the samples undergo uniform sampling rate, duration, and amplitude normalization to avoid the bias of "higher volume leading to higher subjective scores," ensuring that subjective evaluation focuses solely on differences in melodic structure and auditory characteristics.
[0088] During the evaluation and decision-making stage, the system uses the subjective scoring acquisition module to complete the randomization or balance control of the sample playback order, the configuration of the scoring scale, and the summary of scoring records, forming the subjective scoring statistics for each candidate musical motif. Based on this, the objective index module performs secondary discrimination on the high-scoring subjective candidates, calculating at least indicators such as loudness, sharpness, and roughness, and integrates them with the subjective scores to make a decision, thereby outputting the target note sequence.
[0089] After the note sequence is determined, the system enters the parameter coordination stage. While keeping the target note sequence unchanged, candidate configurations are constructed in layers for structural parameters such as mode / pitch, time signature, tempo and duration, and corresponding samples are generated. The optimal configuration is then selected through subjective evaluation and objective index analysis. The overall listening experience after the combination of multiple parameters is comprehensively verified and adjusted within the necessary range. Finally, a stable target parameter configuration for music motifs that is suitable for long-term playback in the vehicle is obtained, providing a unified melody skeleton and structural rules for subsequent working condition sound source generation.
[0090] The timbre acquisition and timbre library management module acquires sound samples from candidate instruments or synthesizers and writes them into the timbre library. Subsequently, the ADSR curve generation module performs signal processing on each timbre sample to generate an ADSR curve and four-stage duration parameters. In specific implementation, the module first performs amplitude normalization and bandpass filtering on the time-domain signal to highlight the effective sound frequency band and suppress low-frequency drift and high-frequency noise. Then, it constructs an analytical signal through Hilbert transform and obtains the instantaneous envelope by taking the modulus. The envelope is then subjected to moving average or equivalent smoothing to improve boundary detection stability. Subsequently, the peak value is determined on the smoothed envelope, and a dynamic thresholding strategy is used to obtain the boundary indices of the Attack, Decay, Sustain, and Release stages. Then, the duration parameters of each stage are obtained by combining the sampling rate. If necessary, the mean or median of multiple samples of the same timbre can be taken as representative ADSR parameters to reduce the influence of performance differences and random noise. After obtaining the ADSR parameters of each timbre, the working condition timbre matching module establishes mapping rules based on the design goal of "accelerating to emphasize responsiveness, decelerating to emphasize smoothness of fallback, and constant speed to emphasize smoothness of continuity". It prioritizes matching timbres with shorter attack duration to acceleration conditions, timbres with longer release duration to deceleration conditions, and timbres with longer sustain duration to constant speed conditions. It also introduces the constraint of preferential selection of similar instruments to ensure that the timbres of the three working conditions come from the same instrument category or the same timbre family as much as possible, so as to reduce the timbre discontinuity and auditory abruptness when switching working conditions.
[0091] Once the target note sequence, target music parameter configuration, and target timbres for the three operating conditions are determined, the system uses a multi-condition sound source synthesis and output unit to generate the final sound source. This involves mapping the same musical motif and melody skeleton onto the accelerated, decelerated, and constant-speed target timbres for synthesis and rendering, maintaining a consistent note structure while only changing the timbre and envelope dynamics. The output sound source is then normalized to a uniform loudness or peak value, forming a set of multi-condition sound sources that can be directly deployed. The external interface module exports the multi-condition sound sources and their metadata (such as operating condition labels, suggested playback levels, and transition parameters) to the in-vehicle active sound system or simulation platform, enabling it to call the corresponding sound source based on the operating condition recognition results.
[0092] The above system can assist in executing the sound source design method described in Embodiment 1, and has the relevant functional modules and beneficial effects. For technical details not described in detail in this embodiment, please refer to the sound source design method provided in Embodiment 1 of this invention.
[0093] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for layered design of musical motifs for active sound in automobiles under multiple operating conditions, characterized in that, Constructing a musical motif adapted to the vehicle's driving conditions and using it as the basic sound source for in-vehicle active sound, the construction of the musical motif includes: S1. Perform note design to determine the note sequence of the musical motif. A set of candidate note sequences is generated based on a preset note sequence length, and corresponding note design sound samples are generated for each candidate note sequence under at least one set of preset music parameter constraints. Each note design sound sample is evaluated and screened to determine the target note sequence used for the music motif. S2. Perform parameter coordination to determine the musical parameters of the musical motif, excluding the note sequence. While keeping the target note sequence unchanged, candidate parameter configurations are constructed for each music parameter and corresponding parameter coordination sound samples are generated. The parameter coordination sound samples are evaluated and screened to determine the target music parameter configuration used for the music motif. S3. Perform timbre matching to determine the timbre of the musical motif for different operating conditions. Collect sound samples from candidate instruments or synthesizers and extract the amplitude envelope ADSR curve of each sound sample. Based on the auditory requirements of different working conditions and the duration of the Attack, Decay, Sustain, and Release stages in the ADSR curve, select target timbres from candidate instruments or synthesizers for acceleration, deceleration, and constant speed working conditions. S4, Perform sound source generation Based on the established note sequence, musical parameters, and timbre, corresponding musical motifs are generated to serve as the basic sound sources for in-vehicle active sound under acceleration, deceleration, and constant speed conditions, respectively.
2. The hierarchical design method for musical motifs as described in claim 1, characterized in that, In step S1: For the preset note sequence length, each note position is used as a test factor, and the seven basic notes of the modern music system, do, re, mi, fa, sol, la, si, are used as the level set of this factor. Several candidate note sequence sets are generated by Latin hypercube sampling. The preset music parameters include: setting the mode to C major, setting the scale degree to the fifth degree, setting the time signature to 3 / 4 time, setting the tempo to 60 beats / minute, and setting the note value to 3 quarter notes.
3. The hierarchical design method for musical motifs as described in claim 1, characterized in that, In step S1: The evaluation and screening adopts the listening test method to organize the review panel to subjectively score each note design sound sample, and selects the highest-scoring note design sound samples according to the average score; in order to reduce the impact of the difference in playback loudness on the listening experience, the loudness or peak value of each note design sound sample needs to be uniformly normalized before the listening test so that the average output level at the playback device is consistent.
4. The hierarchical design method for musical motifs as described in claim 3, characterized in that, By combining objective psychoacoustic evaluation indicators, including at least loudness, sharpness, and roughness, a secondary discrimination is performed on several selected note design sound samples to determine the target note sequence of the musical motif; the loudness is used to characterize the strength of the sound in the subjective perception of the human ear, the sharpness is used to describe the concentration of sound energy in the high-frequency region, and the roughness is used to characterize the unevenness and harshness of the sound caused by rapid amplitude modulation.
5. The hierarchical design method for musical motifs as described in claim 4, characterized in that, In step S2: First, the coordination design of modes and pitches is carried out: based on the determined note sequence, it is mapped to different candidate modes. For each candidate mode, each note in the note sequence needs to be converted into a pitch representation under the corresponding mode, and a corresponding sound sample is generated. Through subjective scoring and objective psychoacoustic evaluation indicators in the listening test, the mode and pitch combination with excellent performance is selected and used as the basic pitch framework of the musical motif. Then, the time signature, tempo, and note value are coordinated and designed separately. While keeping the established note sequence, mode, and pitch unchanged, the appropriate settings for each musical parameter are selected individually through the method of controlling variables and a combination of subjective and objective evaluation. Finally, based on the determined note sequence, mode, pitch, time signature, tempo, and note value, corresponding parameter coordination sound samples are generated. By subjectively evaluating and objectively analyzing these sound samples again, it is verified whether each musical parameter still maintains a good synergistic effect in the combined state. If the combined effect is found to be unsatisfactory, fine-tuning can be performed within the determined range of musical parameters until a stable musical motif parameter configuration with overall auditory performance is obtained.
6. The hierarchical design method for musical motifs as described in claim 1, characterized in that, In step S3, in the amplitude envelope ADSR curve, the Attack phase is used to describe the rising process of sound from the beginning to the peak, the Decay phase is used to describe the falling process after the peak to the stable segment, the Sustain phase is used to describe the continuous process of maintaining stability, and the Release phase is used to describe the release process of sound decaying from the sustain segment to the end. By comparing the corresponding ADSR curves of each candidate instrument or synthesizer, the instrument or synthesizer sound with the shortest duration of the Attack phase is identified as the acceleration tone that emphasizes response speed and dynamics, the instrument or synthesizer sound with the longest duration of the Release phase is identified as the deceleration tone that emphasizes mellowness and soothingness, and the instrument or synthesizer sound with the longest duration of the Sustain phase is identified as the constant speed tone that emphasizes smoothness and continuity.
7. The hierarchical design method for musical motifs as described in claim 1, characterized in that, In step S3, to ensure the consistency of the active sound under multiple working conditions in terms of style, the principle of selecting the same type of instrument is adopted when determining the corresponding timbre for each working condition. The three timbres used for acceleration, deceleration and constant speed working conditions come from the same instrument category, so as to avoid the auditory fragmentation caused by the large difference in timbre when switching working conditions. The instrument category includes percussion instruments, bowed string instruments, wind instruments and plucked instruments.
8. The hierarchical design method for musical motifs as described in claim 1, characterized in that, In step S3, the step of obtaining the amplitude envelope ADSR curve includes: Acquire single-shot samples of candidate musical instruments or synthesizers and discretize them to obtain time-domain signals; The time-domain signal is normalized and filtered to obtain the filtered signal. ; Constructing analytic signals And take its modulus value to obtain the envelope signal. ; The envelope signal is smoothed, and the smoothed signal is used as the amplitude envelope ADSR curve. in, Represents the Hilbert transform. j The imaginary unit, It is a complex modulus.
9. The hierarchical design method for musical motifs as described in claim 8, characterized in that, Dynamic thresholds are used to determine the boundaries of the Attack, Decay, Sustain, and Release stages in the ADSR curve: The start of the Attack phase Defined as the first sampling point in the ADSR curve that meets the preset zero-crossing / oscillation start-up conditions, and the end point of the Attack phase. Defined as the first peak sampling point in the ADSR curve; the Decay phase begins from the end of the Attack phase. To the point ,point Defined as the first sampling point in the ADSR curve when it falls back to the first set threshold level from the first peak; the Sustain phase starts from point To the point ,point Defined as the sampling point in the ADSR curve where the amplitude first falls below the second set threshold; The Release phase starts from point To the point ,point Defined as the sampling point in the ADSR curve where the amplitude first falls below the termination threshold; The duration of each stage is calculated as follows: in, , , and These represent the durations of the Attack, Decay, Sustain, and Release phases, respectively. The sampling rate.
10. A hierarchical design system for musical motifs based on the method of any one of claims 1 to 9, characterized in that, It includes a subjective and objective evaluation and screening unit, a note sequence determination unit, a music parameter determination unit, a timbre matching unit, and a sound source synthesis and output unit, among which: The subjective and objective evaluation and screening unit includes: a subjective scoring acquisition module, used to organize subjective listening evaluation of sound samples and obtain subjective scores; an objective index module, used to calculate objective indicators for sound samples, including at least loudness, sharpness and roughness; and a fusion screening decision module, used to jointly judge the subjective scores and objective indicators to determine the optimal sound sample. The note sequence determination unit includes: a candidate note sequence generation module, used to generate a set of candidate note sequences under a preset note sequence length constraint, and to use experimental design sampling to filter the note combination space to obtain a comprehensive candidate set; a candidate sample rendering and export module, used to render each candidate note sequence into a playable candidate sound source sample file under the condition of fixing at least one set of music parameters other than the note sequence; and a target note sequence determination module, used to filter each candidate sound source sample file through the subjective and objective evaluation and filtering unit, and to take the note sequence corresponding to the optimal candidate sound source sample as the target note sequence. The music parameter determination unit includes a parameter coordination module, which is used to construct candidate parameter configurations and generate corresponding samples for the mode, pitch, time signature, tempo, and time value parameters respectively, while keeping the target note sequence unchanged; and a target music parameter determination module, which is used to filter each of the samples through the subjective and objective evaluation and screening unit, and take the music parameters corresponding to the optimal sample as the target music parameters. The timbre matching unit includes: a timbre acquisition and timbre library management module, used to acquire sound samples of candidate instruments or synthesizers through an audio input / output interface and store them as a timbre sample library; an ADSR curve generation module, used to perform preprocessing and envelope extraction on the timbre samples and generate ADSR curves and their duration parameters for each stage; and a working condition timbre matching module, used to establish working condition adaptation rules based on the duration parameters of each stage of the ADSR curve, and to select target timbres for corresponding acceleration, deceleration, and constant speed working conditions from the timbre sample library respectively. The sound source synthesis and output unit includes: a multi-condition sound source synthesis module, used to synthesize the target note sequence and the target music parameters with the target timbre of each condition to form a multi-condition sound source; and a multi-condition sound source output module, used to perform loudness and / or peak normalization on the output sound source and output it through the audio input / output interface as a sound source for use by the vehicle active sound system.
Citation Information
Patent Citations
An active sound generation system for electric vehicles and its sound control method
CN110525364B