A sound quality evaluation method and system for actively synthesized sound of an electric vehicle

The synthetic quality index (SQI), constructed through time-frequency analysis of sound pressure levels and indicators of standard deviation and mutation rate, solves the problem of acoustic mutations in the active synthetic sound of electric vehicles at the splicing of sound sources, and achieves efficient and scientific sound quality assessment, which is suitable for a variety of sound source scenarios.

CN120636472BActive Publication Date: 2025-10-17JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511126910.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-17
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively evaluate the acoustic mutation characteristics of electric vehicle active synthetic sounds at the sound source splicing points, which affects auditory comfort and system professionalism. In addition, the subjective evaluation cost is high and the objective evaluation indicators are insufficient.

Method used

A method based on time-frequency analysis of sound pressure level is adopted. Through short-time Fourier transform and A-weighting processing, the standard deviation and mutation rate indicators are calculated, and a synthesis quality index (SQI) is constructed to quantitatively evaluate the quality of synthesized sound.

Benefits of technology

It achieves accurate quality evaluation of the active synthetic sound of electric vehicles, reduces human resources and time costs, improves the scientific nature and adaptability of the evaluation, and is applicable to a variety of sound source scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636472B_ABST
    Figure CN120636472B_ABST
Patent Text Reader

Abstract

The application discloses a kind of active synthesis sound quality evaluation method and system of electric vehicle, for the objective quantitative evaluation of the sound quality of electric vehicle active sound piece synthesis after sound.The method is based on the sound pressure level distribution characteristics analysis at the sound synthesis position, combined with the time-frequency characteristics of active synthesis sound, the sound pressure level distribution near the synthesis position is divided into four types, and the corresponding standard deviation index, mutation rate index and energy correction coefficient are respectively constructed, and then a comprehensive evaluation index is proposed.The index can reflect the objective quality of synthesis sound in terms of mutation perception, smooth transition, etc., with good subjective-objective consistency.The system is based on short-time Fourier transform and A-weighting processing to realize the structured analysis of sound signals, and is suitable for the synthesis sound quality evaluation of engine sound, music sound, animal sound and other active sound sources.Compared with the traditional subjective evaluation method, the application can significantly improve the efficiency and accuracy of sound quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of sound synthesis / processing, and particularly relates to a sound quality evaluation method and system for actively synthesized sound of an electric vehicle. BACKGROUND

[0002] With the rapid development of the automobile industry towards lightweight, electrification, and intelligentization, multi-cylinder engines in traditional fuel vehicles are gradually replaced by small displacement engines or driving motors. Electric vehicles not only meet the demand for energy efficiency and environmental friendliness, but also significantly weaken the sound generated by their power systems, resulting in a lack of the dynamic and brand recognition brought by traditional vehicle engines. Especially in high-end vehicles, sound is not only a feedback of driving state, but also an important embodiment of brand individualization design. At the same time, the sound quality of the vehicle interior space has gradually become an important factor affecting the driving comfort and brand perception experience of consumers.

[0003] To make up for the lack of auditory perception in electric vehicles, automobile manufacturers have introduced active sound systems to synthesize specific sound signals inside or outside the vehicle. Such active sound systems can simulate the sound of traditional internal combustion engines, and can also generate music segments or bionic sounds with emotional colors and brand tones, thereby creating personalized and immersive acoustic experiences without relying on mechanical vibration sources. The diversity and flexibility of actively synthesized sound provide a new path for improving the overall acoustic quality, enhancing user experience, and meeting regulatory requirements.

[0004] In the development process of active sound systems, the quality evaluation of synthesized sound is a key task. Especially in the process of splicing segments of electric vehicle active sound, due to the significant differences in time and frequency dimensions of different sound source segments, non-natural transient impact signals often occur at the splicing position. Such acoustic "breakpoints" are easily perceived as abnormal sounds or mutations, which not only destroy the coherence of the sound, but also seriously affect the auditory comfort of users and the professional expression of the system. Therefore, how to accurately capture these acoustic mutation features at the synthesis position and reflect their impact on sound quality in a quantitative way has become a technical problem to be solved.

[0005] Currently, subjective evaluation and objective evaluation based on psychoacoustic parameters are the main evaluation methods. Subjective evaluation requires the organization of a review team with specific characteristics, such as recruiting personnel who meet the gender and age layer coverage requirements when customizing synthesized sound for young female users, which poses great challenges in terms of human resources, time cost, and implementation difficulty. Although the objective evaluation method reduces human factors to some extent, the current evaluation dimensions are still mainly single or limited indicators such as loudness and sharpness, which are difficult to fully reflect the impact of complex structures such as transient impact and energy distribution mutation in synthesized sound on auditory perception.

[0006] In addition, considering that the synthesized sound signal often contains complex spectral components such as music, simulated engine sound or animal sound, the auditory perception effects of different types of sound sources are different when splicing, and a systematic sound quality evaluation method and system applicable to various sound source scenes are urgently needed. SUMMARY

[0007] To solve the above technical problems, the application provides a sound quality evaluation method and system for active synthesized sound of electric vehicles, which can take into account the auditory psychological model and physical signal characteristics, has good subjective and objective consistency, supports fast, automatic and repeatable sound quality analysis, and meets the needs of new generation electric vehicle active sound design and development.

[0008] Specifically, the technical solution provided by the application is as follows:

[0009] A sound quality evaluation method for active synthesized sound of electric vehicles, comprising the steps of:

[0010] obtaining the time domain synthesis position of the synthesized sound, and two segments of sound sources before and after each synthesis position;

[0011] respectively performing time-frequency analysis on the sound pressure level of the synthesized sound and each sound source to obtain the corresponding time-frequency information;

[0012] determining the time-frequency distribution type near the synthesis position according to the time-frequency characteristics of the synthesized sound at the synthesis position;

[0013] respectively calculating the standard deviation index and the mutation rate index corresponding to each time-frequency distribution type;

[0014] constructing a sound pressure level distribution feature statistical index based on the standard deviation index and the mutation rate index;

[0015] correcting the sound pressure level distribution feature statistical index to obtain a synthesis quality index for evaluating the quality of the synthesized sound.

[0016] Further, the types include a first type, a second type, a third type and a fourth type; the first type means that the frequency at the synthesis position is not the fundamental frequency or harmonic frequency of any of the two segments of sound sources before and after; the second type means that the frequency at the synthesis position is only the fundamental frequency or harmonic frequency of the front segment of sound sources; the third type means that the frequency at the synthesis position is only the fundamental frequency or harmonic frequency of the rear segment of sound sources; and the fourth type means that the frequency at the synthesis position is both the fundamental frequency or harmonic frequency of the front segment of sound sources and the fundamental frequency or harmonic frequency of the rear segment of sound sources.

[0017] Further, for the synthesis area of the first type and the fourth type, the standard deviation index is calculated; and for the synthesis area of the second type and the third type, the mutation rate index is calculated.

[0018] Furthermore, the calculation step of the standard deviation index includes:

[0019] Determine the perceptual interval of the composite area: ,

[0020] in, is the time corresponding to the synthetic position, is the set interval width;

[0021] Get the matrix containing time-frequency structure information within the perception interval :

[0022] ,

[0023] Pair Matrix Perform A weighting processing to obtain the matrix :

[0024] ,

[0025] Among them, M and N represent matrices and The number of rows and columns, The perception interval The matrix obtained by Fourier transforming the synthetic sound signal in The element in the mth row and nth column of , where m is the number of the discrete frequency and n is the number of the discrete time; is a matrix The element in row m and column n of , is the weighted gain coefficient of A;

[0026] Matrix-based , calculate the amplitude A ij : ,

[0027] in, is the frequency value at the corresponding synthesis position, is the perception interval The discrete moments contained within Representation matrix No. Rank Column elements;

[0028] Amplitude A ij Normalize and calculate its standard deviation: , ,

[0029] in, After normalization Aij , A min and A max are respectively A ij the minimum and maximum values in j ; is the average value in j , the standard deviation, the smaller the standard deviation, the better the quality of the synthesized sound.

[0030] Preferably, the Fourier transform employs a short-time discrete Fourier transform based on Kaiser window function; the A-weighting gain coefficient wherein is the corresponding sound pressure level correction value at different frequencies according to the IEC61672-1 international standard, used to simulate the sensitivity of human ears to different frequencies.

[0031] Further, the calculation step of the mutation rate index comprises:

[0032] calculating the average amplitude in the time period before and after the perception interval , denoted as and :

[0033] , ,

[0034] For the second type, the mutation rate index is: ,

[0035] For the third type, the mutation rate index is: ,

[0036] and the smaller the value, the better the quality of the synthesized sound.

[0037] Further, the sound pressure level distribution characteristic statistical index is composed of the standard deviation index of the first type and the fourth type synthesis area and the mutation rate index of the second type and the third type synthesis area: , is a sound pressure level distribution characteristic statistical index representing frequency, and are respectively the normalized standard deviation index of the first type and the fourth type synthesis area.

[0038] Further, the calculation formula of the synthesis quality index SQI is: The smaller the value of the SQI is, the better the synthesized sound quality is; wherein, is the energy correction factor corresponding to the frequency f, and , is the energy of the corresponding frequency within the perceptual interval , , is the total energy of each frequency within the perceptual interval , .

[0039] A sound quality evaluation system based on the above method, comprising a synthesized region type determination module, a standard deviation index calculation module, a mutation rate index calculation module and a synthesized quality index calculation module; the synthesized region type determination module is used to determine the type of the synthesized region according to the time-frequency characteristics of the synthesized sound at the synthesized position; the standard deviation index calculation module is used to calculate the standard deviation index of the first type and the fourth type of synthesized region, the mutation rate index calculation module is used to calculate the mutation rate index of the second type and the third type of synthesized region, and the synthesized quality index calculation module is used to calculate the synthesized quality index used for evaluating the synthesized sound quality according to the output values of the standard deviation index calculation module and the mutation rate index calculation module.

[0040] Further, the system further comprises an audio input module for inputting the synthesized sound and the sound source, and a sound playing module, and further comprises a display module for displaying the time-frequency analysis diagram and the sound quality evaluation index.

[0041] Compared with the prior art which mainly relies on subjective evaluation or traditional psychoacoustic parameters to evaluate the quality of the active sound of the electric vehicle, the present application proposes four typical types based on the sound pressure level distribution mode by deeply analyzing the sound pressure level variation characteristics at the synthesized position of the sound segment, and combines statistical quantitative indexes such as standard deviation, mutation rate and energy correction factor to construct a systematic, quantitative and highly targeted sound quality objective evaluation method. The method can effectively capture the mutation characteristics caused by the sound source switching in the synthesized sound, thereby realizing accurate measurement of the synthesized sound quality, and overcoming the defects of the traditional evaluation method, such as heavy reliance on manpower, strong subjectivity of evaluation and lack of targeted evaluation index.

[0042] ​​By introducing short-time Fourier analysis, A-weighted perceptual correction, etc., the evaluation system constructed by the application is not only suitable for the synthesis of music segments, but also can be extended to various active sound scenes such as engine sound and animal sound, and has good universality and adaptability. Experiments prove that the objective indicators obtained by the method are highly consistent with subjective perception, significantly improving the efficiency and scientificity of sound quality evaluation, providing strong technical support for the rapid development, personalized customization and quality control of active sound of electric vehicles, and laying a solid foundation for the construction of automobile brand sound identification system. BRIEF DESCRIPTION OF DRAWINGS

[0043] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, together with the embodiments of the application, to explain the application, and do not constitute a limitation on the application.

[0044] Figure 1 is a time-frequency graph of two music sound sources provided by an embodiment of the application, the subgraph (a) is a time-frequency graph of the first music, and the subgraph (b) is a time-frequency graph of the second music;

[0045] Figure 2 is a time-frequency graph of a synthesized sound obtained by synthesizing two sound sources at different positions, the subgraph (a) is a time-frequency graph of the synthesized sound obtained by synthesizing the first sound source at a position of 10% of the time length of the first sound source and the second sound source, and the subgraph (b) is a time-frequency graph of the synthesized sound obtained by synthesizing the first sound source at a position of 100% of the time length of the first sound source and the second sound source;

[0046] Figure 3 is a time-frequency graph of a synthesized sound containing four types of sound pressure level distribution characteristics provided by an embodiment of the application;

[0047] Figure 4 is a comparison diagram of subjective and objective evaluation results of different synthesized sound qualities provided by an embodiment of the application;

[0048] Figure 5 is a flowchart of a sound quality evaluation method provided by an embodiment of the application. DETAILED DESCRIPTION

[0049] To make the purpose, technical scheme and advantages of the embodiments of the application more clear, the technical scheme in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0050] Embodiment one

[0051] The embodiment provides a sound quality evaluation method of active synthesized sound of an electric vehicle.

[0052] The existing active sound sources are mainly engine sound, animal sound and music sound. Two active music segments are designed in the embodiment, each of which contains six notes, and the note time values are 0.5s, 0.5s, 3s, 0.5s, 0.5s and 3s. The total time length is 8s. The sound characteristics near the synthesis position of the music segments are analyzed in depth.

[0053] As shown in Figure 1 , two subgraphs in the figure are time-frequency graphs of the two music segments. Subgraph (a) is a time-frequency graph of the first music segment, and subgraph (b) is a time-frequency graph of the second music segment. It can be observed from the time-frequency graphs of the two music segments that there are multiple highlighted areas parallel to the time axis at some frequencies. The reason for this is that according to the sound production principle of musical instruments, when a musical instrument produces the sound of a corresponding note, it will simultaneously produce sounds corresponding to the fundamental frequency and harmonic frequency of the note. Therefore, these frequency values correspond to the fundamental frequency and harmonic frequency of each note of the music source respectively. The harmonic frequency is an integer or fractional multiple of the fundamental frequency. In the figure, the harmonic component is represented by Hn , wherein n represents the order of the harmonic frequency component.

[0054] Since the driving state of the vehicle is variable, the two music sources can be synthesized at any synthesis position. That is, the previous music source may not be able to sound completely when synthesized, and it is necessary to switch to the next music source. For example, the previous source is synthesized with the next source at the 10% position (0.8s) of the source time length.

[0055] In the embodiment, the two music sources are directly synthesized at multiple synthesis positions (10% to 100% positions, with 10% positions as intervals) by a direct synthesis method, and the quality of the synthesized sound is measured by subjective evaluation. The subjective evaluation adopts a rating method, and the rating scale is shown in Table 1. The average subjective score of the synthesized sound at each synthesis position is 5.6, indicating that the subjective feeling is difficult.

[0056] Table 1 Subjective evaluation scale

[0057]

[0058] In order to explore the reasons for the insufficient sound quality of the synthesized sound, time-frequency analysis is performed on the directly synthesized sound, as shown in Figure 2As shown in the figure, the two subgraphs are the time-frequency graphs of the synthesized sound obtained by directly synthesizing two pieces of sound sources at different positions. Subgraph (a) is the time-frequency graph of the synthesized sound after synthesizing the first piece of sound source at 10% of its duration with the second piece of sound source, and subgraph (b) is the time-frequency graph of the synthesized sound after synthesizing the first piece of sound source at 100% of its duration with the second piece of sound source. It can be seen that there is a highlighted band perpendicular to the time axis at the synthesis position. The reason is that the time-domain waveforms of the two pieces of music sound sources are different, so the sound pressure level of the sound source at the synthesis position is prone to sudden changes. Based on the principle of uncertainty, such a time extremely short and energy highly concentrated sudden signal will exhibit the characteristics of wideband energy distribution in the frequency domain, that is, in the time-frequency graph, it is embodied as a highlighted band perpendicular to the time axis. The obvious mutation of the sound pressure level at the synthesis position will be felt as a sudden "click" sound in the actual evaluation, thereby causing the sound quality to decrease and the subjective experience to be poor.

[0059] Through the above analysis of the sound synthesis problem, it can be found that the quality of the synthesized sound of the sound source segment is related to the distribution characteristics of the sound pressure level at the synthesis position. For example, Figure 3 As shown, taking the case of synthesizing at the 10% position of the sound source as an example, the distribution characteristics of the sound pressure level near the synthesis position are locally enlarged and displayed. As mentioned earlier, in Figure 1 In the time-frequency graph of the music sound source, there are highlighted areas at the fundamental frequency and harmonic frequency, indicating that the sound pressure level at the corresponding frequency is high. Therefore, according to the distribution of the sound pressure level on both sides of the synthesis position, the synthesis region can be divided into four categories: Type 1 is the frequency at the synthesis position that is not the fundamental frequency and harmonic frequency of either of the two pieces of sound source. The sound pressure level of this region has the distribution characteristics of high sound pressure level at the synthesis position and low sound pressure level on both sides; Type 2 is the frequency at the synthesis position corresponding to the fundamental frequency and harmonic frequency of the previous sound source, but not the fundamental frequency and harmonic frequency of the latter sound source. The sound pressure level of this region has the distribution characteristics that the sound pressure level near the synthesis position decreases from large to small; Type 3 is the frequency at the synthesis position corresponding to the fundamental frequency and harmonic frequency of the latter sound source, but not the fundamental frequency and harmonic frequency of the previous sound source. The sound pressure level of this region has the distribution characteristics that the sound pressure level near the synthesis position increases from small to large; Type 4 is the frequency at the synthesis position that is consistent with the fundamental frequency and harmonic frequency of the two pieces of sound source. The sound pressure level of this region has the distribution characteristics of relatively high near the synthesis position.

[0060] According to the distribution characteristics of the sound pressure level of the above four types, further select the corresponding statistical indicators for each type to represent the change of the sound pressure level, thereby proposing objective indicators for evaluating the quality of the synthesized sound. Among them, Type 1 can be regarded as pulse fluctuation, and Type 4 can be regarded as sawtooth fluctuation. The changes of the sound pressure level of these two types can be measured by the standard deviation index σ . Types 2 and 3 can be regarded as step fluctuation, and the change of the sound pressure level can be represented by the mutation rate β .

[0061] Specifically, the process of establishing the standard deviation, mutation rate index, and objective evaluation index of synthetic sound quality is as follows:

[0062] Step 1: Determine the temporal awareness range around the composition location Tspan , thereby determining the calculation range of the above statistical indicators.

[0063] (1)

[0064] in is the time corresponding to the synthetic position. Combined with the human ear perception characteristics, determine is 200 milliseconds.

[0065] Step 2: Get Tspan A matrix containing time-frequency structure information in the time range .

[0066] (2)

[0067] (3)

[0068] in, Yes Tspan The signal within the time range is obtained by short-time Fourier transform based on Kaiser window function The matrix m Rank n column matrix elements, and , m is the number of discrete frequency, n is the number of discrete time. Further, considering the auditory characteristics of the human ear, The matrix is ​​processed by A weighting to obtain Matrix, where Matrix elements . is the gain coefficient of A weighting, Different frequencies are obtained according to the IEC61672-1 international standard The corresponding sound pressure level correction value is used to simulate the sensitivity of the human ear to different frequencies.

[0069] Step 3: Calculate the standard deviation indicators of type 1 and type 4. First, by Matrix gets amplitude A ij .

[0070] (4)

[0071] in are the frequency values ​​corresponding to type 1 and type 4.

[0072] The amplitudes are then normalized and the standard deviation is calculated as follows:

[0073] (5)

[0074] (6)

[0075] where, , , .

[0076] To ensure the value range of all statistical indicators remain consistent, multiply the above standard deviation by 2 and map it to the [0, 1] interval. The statistical indicator expressions for representing type 1 and type 4 sound pressure level changes are as follows:

[0077] (7)

[0078] Obviously, and the smaller the value, the better the quality of the synthesized sound.

[0079] Step 4: Calculate the mutation rate indicators representing type 2 and type 3. Calculate the average amplitude in the time period before and after the synthesis position , denoted as and .

[0080] (8)

[0081] (9)

[0082] where is the frequency value corresponding to type 2 and type 3.

[0083] The calculation method of the mutation rate indicator β is as follows:

[0084] (10)

[0085] (11)

[0086] Obviously, and the smaller the value, the better the quality of the synthesized sound.

[0087] Step 5: Determine the correction factor . Considering that the greater the energy of the sound signal, the more obvious the human perception. Therefore, introduce a correction coefficient based on the proportion of sound energy to correct the above four statistical indicators.

[0088] (12)

[0089] where, is corresponding to energy of the frequency, . is the total energy in the frequency range, .

[0090] Step 6: Calculate the synthesis quality index SQI (Synthesis Quality Index). Considering the above four characterization indexes, the expression of SQI is as follows:

[0091] (13)

[0092] wherein, is energy correction factor corresponding to the frequency. is a statistical index representing the sound pressure level distribution characteristics at the frequency, . The smaller the value of SQI, the better the quality of the synthesized sound.

[0093] The cross-fade method can effectively improve the sound pressure level mutation problem at the synthesis position and improve the quality of the synthesized sound. To verify the effectiveness of the objective evaluation index of the synthesized sound quality, the synthesis of the sound source is carried out at the aforementioned multiple synthesis positions by the cross-fade method with 30 different parameter settings, and the quality of the synthesized sound is evaluated by subjective and objective methods. The evaluation results are shown in Figure 4 It can be seen that the subjective and objective indexes have good consistency, and the Pearson correlation coefficient is -0.91 (the absolute value of the correlation coefficient is greater than 0.9), which confirms the effectiveness of the objective evaluation index of the synthesis sound quality of the active sound of the electric vehicle.

[0094] The above is the main content of the sound quality evaluation method provided by the present application, and the overall technical process is shown in Figure 5 . It should be noted that:

[0095] (1) Although the active sound synthesis sound quality objective evaluation index is proposed in this embodiment taking the electric vehicle as an example, it is also applicable to other active sound application situations and scenarios, such as the objective evaluation of the synthesis sound quality of the electric vehicle outside pedestrian warning sound, the active sound of the self-driving car and other transportation tools.

[0096] (2) The music sound includes single tone and chord sound, wherein the single tone is a single pitch tone (including overtone); the chord sound is a harmony formed by two or more different pitch tones sounding at the same time. This embodiment takes single tone as an example to propose the corresponding index, but for the synthesis of chord sound, the sound pressure level distribution characteristics near the synthesis position are still the same as single tone, so the index proposed in this embodiment is also applicable to the evaluation of music chord sound.​​

[0097] (3) The engine sound has similar characteristics as the music sound, the fundamental frequency of the music sound can correspond to the main order of the engine sound, and the harmonic frequency of the music sound can correspond to the harmonic order of the engine sound, so the indexes proposed in the embodiment are also applicable to the evaluation of the sound quality of the engine sound synthesis.

[0098] (4) The animal sound is generated in a similar manner as human sound, mainly relying on the compression of air flow in the lungs (or air sac) to drive the vibration of the vocal cords (or membrane structure), and the sound is formed after being modulated by the sound channel and radiated through the opening. The animal sound also contains fundamental frequency and harmonic frequency components, which are essentially the same as the fundamental frequency and harmonic frequency components of the music sound, so the indexes proposed in the embodiment are also applicable to the evaluation of the sound quality of the animal sound synthesis.

[0099] (5) The embodiment designs two music source segments to study the problem of proposing objective evaluation indexes of the sound quality of the active sound synthesis of the electric vehicle, but the notes and time values in the source segments can be adjusted and designed according to the actual needs of the active sound, and the indexes proposed in the embodiment are also applicable to other music sources.

[0100] Embodiment Two

[0101] Based on the above method, the embodiment provides a sound quality evaluation system for the active sound synthesis of the electric vehicle, which mainly includes a synthesis region type determination module, a standard deviation index calculation module, a mutation rate index calculation module, and a synthesis quality index calculation module.

[0102] The synthesis region type determination module is used to determine the type of the synthesis region according to the time-frequency characteristics of the synthesis sound at the synthesis position. The standard deviation index calculation module is used to calculate the standard deviation index of the first type and the fourth type of synthesis region, and the mutation rate index calculation module is used to calculate the mutation rate index of the second type and the third type of synthesis region. The synthesis quality index calculation module is used to calculate the synthesis quality index for evaluating the quality of the synthesis sound according to the output values of the standard deviation index calculation module and the mutation rate index calculation module.

[0103] In some embodiments, the system includes an audio input module for inputting the synthesis sound and the source, and a sound playing module, and further includes a display module for displaying the time-frequency analysis diagram and the detection results.

[0104] The above system can execute the sound quality evaluation method described in embodiment one, has the corresponding functional modules and beneficial effects of the method, and the technical details not described in detail in the embodiment can be referred to the sound quality evaluation method provided in embodiment one of the application.

[0105] Those skilled in the art can clearly understand the technical solutions of the embodiments from the above description of the embodiments, and the embodiments can be implemented by means of software plus a general hardware platform, or by hardware. Based on such an understanding, the above technical solutions, essentially or in terms of related art, can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, or an optical disk, and includes a plurality of instructions to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0106] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of the different aspects of the present application as described above; for the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A sound quality evaluation method for active synthetic sound of electric vehicles, characterized in that: Including steps: Obtain the time domain synthesis position of the synthesized sound, and the two sound sources before and after each synthesis position; Perform time-frequency analysis of the sound pressure level of the synthesized sound and each sound source to obtain the corresponding time-frequency information; Determine the time-frequency distribution type near the synthesis position according to the acquired time-frequency information; the types include a first type, a second type, a third type, and a fourth type; the first type means that the frequency at the synthesis position is not the fundamental frequency and harmonic frequency of any of the two preceding and following sound sources; the second type means that the frequency at the synthesis position is only the fundamental frequency or harmonic frequency of the preceding sound source; the third type means that the frequency at the synthesis position is only the fundamental frequency or harmonic frequency of the following sound source; the fourth type means that the frequency at the synthesis position is both the fundamental frequency or harmonic frequency of the preceding sound source and the fundamental frequency or harmonic frequency of the following sound source; Calculate the standard deviation index and mutation rate index of the synthetic region corresponding to each time-frequency distribution type respectively: for the synthetic regions of the first and fourth types, calculate their standard deviation index; for the synthetic regions of the second and third types, calculate their mutation rate index; Construct the statistical index of sound pressure level distribution characteristics based on the standard deviation index and mutation rate index; The synthetic quality index for evaluating the synthetic sound quality is obtained based on the correction of statistical indicators of sound pressure level distribution characteristics.

2. The sound quality evaluation method according to claim 1, wherein: The calculation steps of the standard deviation indicator include: Determine the perceptual interval of the composite area: , in, is the time corresponding to the synthetic position, is the set interval width; Get the matrix containing time-frequency structure information within the perception interval : , Pair Matrix Perform A weighting processing to obtain the matrix : , Among them, M and N represent matrices and The number of rows and columns, The perception interval The matrix obtained by Fourier transforming the synthetic sound signal in The element in the mth row and nth column of , where m is the number of the discrete frequency and n is the number of the discrete time; is a matrix The element in row m and column n of , is the weighted gain coefficient of A; Matrix-based , calculate the amplitude A ij : , in, is the frequency value at the corresponding synthesis position, is the perception interval The discrete moments contained within Representation matrix No. Rank Column elements; Amplitude A ij Normalize and calculate its standard deviation: , , in, After normalization A ij , A min and A max They are A ij exist j The minimum and maximum values ​​in the row; for exist j The average value of the row, is the standard deviation. The smaller the standard deviation, the better the quality of the synthesized sound.

3. The sound quality evaluation method according to claim 2, wherein: The Fourier transform adopts the short-time discrete Fourier transform based on the Kaiser window function; the A weighted gain coefficient ,in Different frequencies are obtained according to the IEC61672-1 international standard The corresponding sound pressure level correction value is used to simulate the sensitivity of the human ear to different frequencies.

4. The sound quality evaluation method according to claim 3, wherein: The calculation steps of the mutation rate index include: Calculating the perception interval Inside The average amplitude of the time period before and after is recorded as and : , , For the second type, the mutation rate index is: , For the third type, the mutation rate index is: , and The smaller the value, the better the quality of the synthesized sound.

5. The sound quality evaluation method according to claim 4, wherein: The statistical index of the sound pressure level distribution characteristics is composed of the standard deviation index of the first type and the fourth type synthesis area and the mutation rate index of the second type and the third type synthesis area: , It is a representation Statistical indicators of the sound pressure level distribution characteristics at the frequency, and They are the normalized standard deviation indicators of the first and fourth type synthetic areas.

6. The sound quality evaluation method according to claim 5, wherein: The calculation formula of the composite quality index SQI is: , the smaller the SQI value is, the better the quality of the synthesized sound is; yes The energy correction factor corresponding to the frequency, and , is the perception interval Internal correspondence The energy of the frequency, , is the perception interval The total energy of each frequency within .

7. A sound quality evaluation system based on the method according to any one of claims 1 to 6, characterized in that: It includes a synthesis area type determination module, a standard deviation index calculation module, a mutation rate index calculation module and a synthesis quality index calculation module; the synthesis area type determination module is used to determine the type of the synthesis area according to the time-frequency characteristics of the synthesized sound at the synthesis position; The standard deviation index calculation module is used to calculate the standard deviation index of the first type and the fourth type of synthesis areas, the mutation rate index calculation module is used to calculate the mutation rate index of the second type and the third type of synthesis areas, and the synthesis quality index calculation module is used to calculate the synthesis quality index for evaluating the quality of the synthesized sound based on the output values ​​of the standard deviation index calculation module and the mutation rate index calculation module.

8. The sound quality evaluation system according to claim 7, wherein: The system further comprises an audio input module for inputting synthesized sound and sound source, a sound playing module, and a display module for displaying a time-frequency analysis graph and sound quality evaluation indicators.

Citation Information

Patent Citations

  • Electric vehicle active sound production system design method

    CN110481470A

  • Synthesized speech evaluation system and synthesized speech evaluation method

    JP2010060846A