Audio signal generation method and device, related equipment and medium

By using semantic information and acoustic parameter models to generate target acoustic parameter groups in new energy vehicles, the problem of personalized configuration of the simulated sound wave system of new energy vehicles is solved, personalized generation of audio signals is achieved, and user experience is improved.

CN120708635APending Publication Date: 2025-09-26XIAOMI EV TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410355205.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The simulated sound wave system of new energy vehicles cannot personalize the acoustic parameters, has poor user adjustment flexibility, cannot meet actual needs, and affects the driving experience.

Method used

By determining the semantic information and inputting the preset acoustic parameter model, a target acoustic parameter group is generated, and then an audio signal that meets the user's needs is generated. The semantic information and acoustic parameter model are used to realize personalized configuration of acoustic parameters.

Benefits of technology

The universality and reliability of audio signal generation are improved, and the user has low difficulty in learning and using it. It can generate audio signals that meet user needs and enhance the driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708635A_ABST
    Figure CN120708635A_ABST
Patent Text Reader

Abstract

The invention relates to an audio signal generation method and device, related equipment and a medium. The audio signal generation method comprises the following steps: determining semantic information of an audio signal to be generated; inputting the semantic information into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model; and generating an audio signal conforming to the semantic information according to the target acoustic parameter group. The purpose of individually configuring the acoustic parameters is achieved through the semantic information and the acoustic parameter model, so that the audio signals meeting the user requirements can be generated, and the user experience is improved. Besides, the target acoustic parameter group is obtained by utilizing the semantic information and the acoustic parameter model, on one hand, the learning and use difficulty of the user is relatively low, and the universality of the generated audio signal is improved, and on the other hand, the acoustic parameter group can be adjusted through intuitive semantics to generate the audio signal conforming to the semantics, so that the user experience is improved. And the reliability of generating the audio signal conforming to the user demand semantics is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio technology, and in particular to an audio signal generation method, apparatus, related equipment, and medium. Background Art

[0002] Compared to traditional internal combustion engine vehicles, new energy vehicles utilize a quieter motor system, resulting in exceptionally quiet driving. However, the lack of acoustic feedback regarding driving status can make new energy vehicles less appealing to some drivers who value driving pleasure, social interaction, and information. To address this lack, new energy vehicles often incorporate a sound simulation feature, using virtual synthesis to compensate for the lack of engine sound.

[0003] However, the current development status of the simulated sound field of new energy vehicles is that the OEMs provide one or more pre-designed sound styles during the development phase, allowing users to freely choose to switch or adjust the overall output volume of the simulated sound system while driving. The adjustment flexibility is poor, and it is impossible to personalize the acoustic parameters to meet the actual needs of users. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides an audio signal generation method, apparatus, related equipment and medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating an audio signal, comprising:

[0006] determining semantic information of the audio signal to be generated;

[0007] Inputting the semantic information into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model, wherein the target acoustic parameter group includes a parameter value of each acoustic parameter to be adjusted;

[0008] An audio signal conforming to the semantic information is generated according to the target acoustic parameter group.

[0009] According to a second aspect of the embodiments of the present disclosure, there is provided an audio signal generating apparatus, including:

[0010] A first determining module is configured to determine semantic information of the audio signal to be generated;

[0011] a second determining module configured to input the semantic information into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model, wherein the target acoustic parameter group includes a parameter value of each acoustic parameter to be adjusted;

[0012] A generating module is configured to generate an audio signal that conforms to the semantic information according to the target acoustic parameter group.

[0013] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0014] processor;

[0015] a memory for storing processor-executable instructions;

[0016] The processor is configured to execute the executable instructions to implement the audio signal generation method described in the first aspect of the embodiment of the present disclosure.

[0017] According to a fourth aspect of an embodiment of the present disclosure, an audio system is provided, comprising the electronic device as described in the third aspect of the embodiment of the present disclosure and an acoustic parameter model, wherein the acoustic parameter model is used to input semantic information and output a target acoustic parameter group.

[0018] According to a fifth aspect of the embodiments of the present disclosure, a vehicle is provided, comprising the audio system according to the fourth aspect of the embodiments of the present disclosure.

[0019] According to a sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the audio signal generation method provided in the first aspect of the embodiment of the present disclosure are implemented.

[0020] According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the audio signal generation method provided in the first aspect of the embodiment of the present disclosure.

[0021] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0022] The above technical solution utilizes semantic information and a preset acoustic parameter model to obtain a target acoustic parameter group, and then generates an audio signal based on the semantic information based on the target acoustic parameter group. This personalized configuration of acoustic parameters is achieved through the semantic information and acoustic parameter model, thereby enabling the generation of audio signals that meet user needs and enhance the user experience. Furthermore, utilizing semantic information and an acoustic parameter model to obtain a target acoustic parameter group reduces user learning and usage difficulty, improving the universality of generated audio signals. Furthermore, intuitive semantics can be used to adjust the acoustic parameter group to generate an audio signal that conforms to the semantics, improving the reliability of generating audio signals that meet the semantics of user needs.

[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0025] Figure 1 The figure is a flowchart of a method for generating an audio signal according to an exemplary embodiment.

[0026] Figure 2 The figure is a schematic diagram showing a method for generating simulated sound waves according to an exemplary embodiment.

[0027] Figure 3 The figure is a flowchart showing a method for training an acoustic parameter model according to an exemplary embodiment.

[0028] Figure 4 The figure is a block diagram of an audio signal generating apparatus according to an exemplary embodiment.

[0029] Figure 5 It is a block diagram of an electronic device according to an exemplary embodiment.

[0030] Figure 6 is a block diagram of a vehicle according to an exemplary embodiment. DETAILED DESCRIPTION

[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0032] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0033] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0034] In the audio field, similar to analog sound, car audio systems have developed to the point where they offer a wide range of parameters that users can freely adjust and save as "custom" presets. However, due to the large number of adjustable parameters and the higher algorithmic complexity of analog sound systems, it is difficult to simply extract these parameters and provide them as user-friendly adjustment options.

[0035] In the related art, a user preference model can be constructed based on the user's rating results for the preset sound effects, and the preset sound effects can be modulated with the calculation results of the model until the user is satisfied with the sound effects. In the process of exploring user preferences, users are required to audition a large amount of audio that they do not prefer, which reduces the user experience; and this method requires users to divide their energy and judge while listening, which is not suitable for in-vehicle driving scenarios that require users to concentrate highly. In addition, in the related art, it is also possible to obtain a simulated sound experience that is the same as the target expectation by modifying parameters in the actual vehicle control based on the expected performance of the simulated sound waves in each pre-stored mode. This method cannot achieve user-defined expected sound waves, and is still limited to the car manufacturer designing the expected sound waves and then the actual car adjusting the parameters to approximate the expected sound waves. Therefore, it is impossible to obtain an audio signal that conforms to the semantic information based on the semantic information to meet the actual needs of the user.

[0036] In light of this, the present disclosure provides an audio signal generation method, apparatus, related devices, and media. These methods utilize semantic information and a preset acoustic parameter model to derive a target acoustic parameter set, and then generate an audio signal based on this target acoustic parameter set to meet user needs. Furthermore, utilizing the acoustic parameter model to derive the target acoustic parameter set reduces user learning and usage complexity, improving the universality of audio signal generation.

[0037] Figure 1 FIG. 1 is a flow chart of a method for generating an audio signal according to an exemplary embodiment. Figure 1 As shown, the method may include the following steps.

[0038] In step S11 , semantic information of the audio signal to be generated is determined.

[0039] Optionally, the audio signal may include an in-vehicle entertainment audio signal, such as an audio signal of in-vehicle multimedia playback, and may also include an analog sound wave signal.

[0040] In the present disclosure, the semantic information of the audio signal to be generated can be determined according to the user's selection. In one embodiment, the user can customize the semantic information according to needs and input it into the device that executes the audio signal generation method. In another embodiment, different semantic information is pre-set and output, and the user can select the semantic information of the audio signal to be generated from the pre-set semantic information according to needs. For example, multiple semantic information are pre-set, such as comfortable-intense, exquisite-rough, and the user can select the semantic information of the audio signal to be generated from these pre-set semantic information.

[0041] For example, the semantic information may be a semantic text, such as comfortable-intense, refined-rough, etc., or a numerical value used to represent the semantic text, such as the numerical value 1 used to represent comfortable, the numerical value 5 used to represent intense, etc.

[0042] In step S12, the semantic information is input into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model. The target acoustic parameter group includes a parameter value of each acoustic parameter to be adjusted.

[0043] The acoustic parameters to be adjusted can be all the parameters required in the audio signal generation process. However, considering the large number of parameters required in the audio signal generation process, which results in a high algorithmic complexity of the acoustic parameter model, the present disclosure also allows for extracting some parameters from all the parameters required in the audio signal generation process as the acoustic parameters to be adjusted, thereby reducing the algorithmic complexity of the acoustic parameter model. The specific implementation method for extracting the acoustic parameters to be adjusted and the training method for the acoustic parameter model will be described in detail below and will not be described here.

[0044] In step S13 , an audio signal conforming to the semantic information is generated according to the target acoustic parameter group.

[0045] In the present disclosure, after obtaining the target acoustic parameter set output by the acoustic parameter model, the target acoustic parameter set can be input into the audio signal generation algorithm to obtain an audio signal that conforms to the semantic information determined in step S11. For example, if the audio signal can be a simulated sound signal, after obtaining the target acoustic parameter set, the target acoustic parameter set can be input into the simulated sound model to obtain a simulated sound signal that conforms to the semantic information.

[0046] For example, Figure 2 FIG. 1 is a schematic diagram showing a method for generating simulated sound waves according to an exemplary embodiment. Figure 2 As shown, the simulated sound wave system may include an originally designed simulated sound wave, an acoustic parameter model, and a modulated simulated sound wave. The originally designed simulated sound wave may include a simulated sound wave model and an originally designed acoustic parameter group for the simulated sound wave. The simulated sound wave model is the sound calculation algorithm for the simulated sound wave and cannot be changed. The simulated sound wave model may use the originally designed acoustic parameter group for the simulated sound wave to generate an originally designed simulated sound wave signal. The acoustic parameter model may generate a target acoustic parameter group based on the semantic information input by the user, and input the target acoustic parameter group into the originally designed acoustic parameter group for the simulated sound wave to obtain a new simulated sound wave acoustic parameter group, so that the simulated sound wave model generates a new simulated sound wave signal based on the new simulated sound wave acoustic parameter group, and the generated new simulated sound wave signal conforms to the semantic information input by the user.

[0047] The original acoustic parameter set for simulated sound waves is the set of acoustic parameters for simulated sound waves that comes with the OEM's mass-produced vehicles. Changing these parameters will directly change the acoustic output of the simulated sound waves, that is, directly alter the output simulated sound signal. Typically, due to the large number of original acoustic parameter sets for simulated sound waves and the high complexity of understanding them, users cannot change these acoustic parameters, or can only modify one parameter, the output volume. However, in the present disclosure, the target acoustic parameter set can be directly derived through a semantic information set and an acoustic parameter model, thereby enabling the modification of some parameters of the original acoustic parameter set for simulated sound waves, except for the output volume, with a lower level of user learning and usage difficulty.

[0048] The above technical solution utilizes semantic information and a preset acoustic parameter model to obtain a target acoustic parameter group, and then generates an audio signal based on the semantic information based on the target acoustic parameter group. This personalized configuration of acoustic parameters is achieved through the semantic information and acoustic parameter model, thereby enabling the generation of audio signals that meet user needs and enhance the user experience. Furthermore, utilizing semantic information and an acoustic parameter model to obtain a target acoustic parameter group reduces user learning and usage difficulty, improving the universality of generated audio signals. Furthermore, intuitive semantics can be used to adjust the acoustic parameter group to generate an audio signal that conforms to the semantics, improving the reliability of generating audio signals that meet the semantics of user needs.

[0049] In addition, when the audio signal is an analog sound signal, it can achieve the purpose of user-defined modulation of the analog sound, and can provide drivers who have driving pleasure needs, driving social needs, and driving information needs with analog sound that conforms to their selected semantics, thereby improving the driving experience.

[0050] The following describes the training method of the acoustic parameter model.

[0051] Figure 3 FIG. 1 is a flow chart showing an acoustic parameter model training method according to an exemplary embodiment. Figure 3 As shown, the training method may include the following steps.

[0052] In step S31 , a training data set is obtained. The training data set includes multiple groups of training samples. The training samples include sample semantic information and a sample acoustic parameter group corresponding to an audio signal that conforms to the sample semantic information.

[0053] In the present disclosure, the sample acoustic parameter group corresponding to the audio signal refers to the sample acoustic parameters used when generating the audio signal, wherein the sample acoustic parameters include the sample parameter values ​​of each acoustic parameter to be adjusted.

[0054] In one embodiment, step S31 may further include:

[0055] Extracting acoustic parameters to be adjusted from an audio signal generation algorithm;

[0056] According to the acoustic parameters to be adjusted, a plurality of sample acoustic parameter groups are generated, and sample semantic information corresponding to the sample audio signal corresponding to each sample acoustic parameter group is obtained to obtain a training data set.

[0057] For example, an audio signal generation system can be constructed to extract the acoustic parameters to be adjusted from the audio signal generation algorithms involved in the audio signal generation system. The audio signal generation algorithms include: a first type of algorithm for generating a single-track audio signal; a second type of algorithm for generating a master-track audio signal from the single-track audio signal; and a third type of algorithm for processing the master-track audio signal to obtain an audio signal to be played.

[0058] Taking the audio signal as an analog sound signal as an example, the constructed audio signal generation system can include a sound source module, a mixing module, and a post-processing module. Among them, the sound source module is responsible for generating a real-time analog sound signal associated with the driving status, and each independently operated sound source module can generate a single-track audio signal. The mixing module is designed to mix and superimpose the single-track audio signals output by multiple sound source modules to generate a total track audio signal. The post-processing module contains a variety of optional sound effect algorithms, which can be used to modulate the total track audio signal into an audio signal to be played. It should be understood that if the number of sound source modules is two, two single-track audio signals can be obtained, and the mixing module mixes and superimposes the two single-track audio signals to generate a total track audio signal.

[0059] As for the sound source module, it includes a first type of algorithm. It is assumed that the audio signal generation system includes two sound source modules, and one sound source module is used to generate a frequency-shifted single-track audio signal according to a frequency shifting algorithm, and the other audio module is used to generate a harmonically processed single-track audio signal according to a harmonic algorithm. The frequency shifting algorithm is used to obtain a frequency shift phase difference based on the frequency of the initial audio signal, the distance between the initial audio signal and the listening point, the frequency shifting rate, and the intensity change rate of the audio signal, and phase-shift the initial audio signal according to the frequency shifting phase difference to obtain a frequency-shifted single-track audio signal. The harmonic algorithm is used to obtain a harmonically processed single-track audio signal based on the amplitude of each harmonic, the frequency conversion rate, the motor speed, and the order of each harmonic.

[0060] In one embodiment, the frequency shift algorithm is shown in formulas (1) and (2), and the harmonic algorithm is shown in formula (3):

[0061]

[0062]

[0063] in, represents the frequency shift phase difference, f0 represents the frequency of the initial audio signal, R0 represents the distance between the initial audio signal and the listening point, and R0 = 0, v r Characterizes the frequency shift rate, Gain r Characterizes the rate of change of the intensity of the audio signal, t characterizes the time interval between two adjacent samples, c characterizes the speed of sound, sig m1 Characterize the single-track audio signal after frequency shift, sig preMake represents the initial audio signal, and τ represents the current time point.

[0064]

[0065] Among them, Sig m2 Represents a single-track audio signal after harmonic processing, n represents the number of harmonics, Amp i represents the harmonic amplitude of the i-th harmonic, k represents the frequency conversion rate, rpm represents the motor speed, Ord i represents the harmonic order of the i-th harmonic, and t represents the time interval between two adjacent samplings.

[0066] For example, an audio signal generation system includes two sound source modules, one of which is used to obtain a frequency-shifted single-track audio signal according to a frequency shift algorithm, and the other is used to obtain a harmonically processed single-track audio signal according to a harmonic algorithm.

[0067] The second type of algorithm is used to obtain a total track audio signal based on the frequency-shifted single-track audio signal, the harmonically processed single-track audio signal, and the loudness of the corresponding single-track audio signals;

[0068] The second type of algorithm is shown in formula (4):

[0069]

[0070] Among them, Sig M Represents the total track audio signal, m represents the number of single track audio signals, Gain j Represents the loudness of the j-th single-track audio signal, Sig mj Represents the j-th single-track audio signal.

[0071] Following the above example, the audio signal generation system includes two sound source modules, namely, m=2, Sig M =Gain1×Sig m1 +Gain2×Sig m2 , where Gain1 represents the loudness of the single-track audio signal corresponding to the single-track audio signal after frequency shifting, and Gain2 represents the loudness of the single-track audio signal corresponding to the single-track audio signal after harmonic processing.

[0072] The third type of algorithm includes a delay algorithm and a frequency domain modulation algorithm; the delay algorithm is used to obtain a delayed total track audio signal based on the total track audio signal and the delay amount; the frequency domain modulation algorithm is used to obtain an audio signal to be played based on a filter coefficient, the delayed total track audio signal at the current moment, and the delayed total track audio signal at the previous moment.

[0073] In this disclosure, the delay algorithm is shown in formula (5):

[0074] Sig M (τ)=Sig M (τ-t0) (5)

[0075] Among them, τ represents the current moment, Sig M (τ) represents the total track audio signal at the current moment obtained through the delay algorithm, and t0 represents the delay amount.

[0076] The frequency domain modulation algorithm is shown in formula (6):

[0077] F(τ)=α*Sig M (τ)+(1-α)*Sig M (τ-1) (6)

[0078] Among them, F(τ) represents the audio signal to be played at the current time τ, α represents the coefficient of the filter, Sig M (τ-1) represents the total track audio signal at the previous moment (τ-1) obtained through the delay algorithm.

[0079] Accordingly, extracting the acoustic parameters to be adjusted from the audio signal generation algorithm may include:

[0080] Extracting acoustic parameters to be adjusted from at least one of the first, second, and third algorithms;

[0081] Among them, in the first type of algorithm, at least one of the frequency shift rate, the intensity change rate of the audio signal, the harmonic amplitude, the frequency conversion rate and the harmonic order is used as the acoustic parameter to be adjusted; in the second type of algorithm, the loudness of the single-track audio signal is used as the acoustic parameter to be adjusted; in the third type of algorithm, the delay amount and / or the filter coefficient is used as the acoustic parameter to be adjusted.

[0082] In a first embodiment, the acoustic parameters to be adjusted are extracted from the first type of algorithm. In the above formula (1), the frequency shift rate affects the speed at which the frequency of the frequency-shifted single-track audio signal changes with changes in the vehicle motor speed, and the audio signal intensity change rate affects the speed at which the loudness of the frequency-shifted single-track audio signal changes with changes in the vehicle speed. Therefore, in the frequency shift algorithm in the first type of algorithm, the frequency shift rate and / or the audio signal intensity change rate can be used as the acoustic parameters to be adjusted.

[0083] In formula (3), the frequency conversion rate affects the speed at which the frequency of the single-track audio signal after harmonic processing changes with the change of the vehicle motor speed, the harmonic amplitude affects the speed at which the loudness of the single-track audio signal after harmonic processing changes with the change of the vehicle speed, and the harmonic order affects the frequency distribution of the single-track audio signal after harmonic processing. Therefore, in the harmonic algorithm in the first category of algorithms, at least one of the frequency conversion rate, the harmonic amplitude and the harmonic order can be used as the acoustic parameter to be adjusted.

[0084] In a second embodiment, the acoustic parameter to be adjusted is extracted from the second type of algorithm. In formula (4), the loudness of a single-track audio signal affects the weight of each single-track audio signal when it is superimposed. Therefore, in the second type of algorithm, the loudness of a single-track audio signal can be used as the acoustic parameter to be adjusted.

[0085] In a third embodiment, the acoustic parameter to be adjusted is extracted from the third type of algorithm. In formula (5), the delay amount affects the delay length in the delay algorithm. Therefore, in the delay algorithm in the third type of algorithm, the delay amount can be used as the acoustic parameter to be adjusted. In formula (6), the filter coefficient affects the frequency domain response of the frequency domain modulation algorithm. Therefore, in the frequency domain modulation algorithm in the third type of algorithm, the filter coefficient can be used as the acoustic parameter to be adjusted.

[0086] In a fourth embodiment, the acoustic parameters to be adjusted can be extracted from any two or three of the first, second, and third algorithms. The specific method is as shown in the above embodiment and will not be repeated here.

[0087] After extracting the acoustic parameters to be adjusted in the above manner, multiple sample acoustic parameter groups are generated according to the acoustic parameters to be adjusted, and sample semantic information corresponding to the sample audio signals corresponding to each of the sample acoustic parameter groups is obtained to obtain a training data set.

[0088] In the present disclosure, a plurality of sample acoustic parameter groups are generated according to the acoustic parameters to be adjusted, and sample semantic information corresponding to the sample audio signal corresponding to each sample acoustic parameter group is obtained to obtain a training data set, which may include:

[0089] First, the high-frequency semantic tags and the semantic dimensions corresponding to each high-frequency semantic tag are determined.

[0090] In one embodiment, if the audio signal is a simulated sound signal, the semantics of the driver's historical needs can be counted, and then high-frequency semantic tags with higher frequency of occurrence can be determined based on the semantics of the historical needs. For example, multiple semantic tags are determined based on the semantics of the historical needs, and the frequency of occurrence of each semantic tag is counted to see whether it is greater than a preset frequency. If it is, the semantic tag is determined to be a high-frequency semantic tag. Afterwards, the corresponding semantic dimension is counted in each high-frequency semantic tag. The semantic dimension corresponding to the high-frequency semantic tag refers to the degree of the semantics corresponding to the high-frequency semantic tag. For example, the high-frequency semantic tag includes semantic tag A, and the semantic dimensions corresponding to semantic tag A are: very comfortable - relatively comfortable - medium - relatively intense - very intense.

[0091] In another embodiment, the specific implementation method for determining high-frequency semantic tags and the semantic dimensions corresponding to each high-frequency semantic tag can be: for each acoustic parameter to be adjusted, within the value range of the acoustic parameter, determine multiple first parameter values ​​of the acoustic parameter according to the first division granularity; based on the multiple first parameter values ​​of each acoustic parameter to be adjusted, determine multiple groups of first sample acoustic parameter groups, and obtain the first sample audio signals corresponding to the first sample acoustic parameter groups, the first sample acoustic parameter group including any first parameter value of each acoustic parameter to be adjusted; count the first semantic feedback of the first group of users on the first sample audio signal, and determine the high-frequency semantic tags and the semantic dimensions corresponding to each high-frequency semantic tag based on the first semantic feedback.

[0092] For example, assuming the acoustic parameter to be adjusted is the frequency shift rate, and the frequency shift rate range is [1, 20], if the partition interval of the first partition granularity representation is 5, then the first parameter value of the frequency shift rate can be 5, 10, 15, and 20, respectively. Alternatively, it can be any four values ​​with a partition interval of 5, which is not specifically limited in this disclosure. Similarly, the first parameter value of each acoustic parameter to be adjusted can be obtained. Then, for each acoustic parameter to be adjusted, a parameter value is selected from its first parameter values, and the parameter values ​​of each selected acoustic parameter to be adjusted are combined into a first sample acoustic parameter group, thereby obtaining multiple groups of first sample acoustic parameter groups. Then, for each group of first sample acoustic parameter groups, the first sample acoustic parameter group is input into an audio signal generation algorithm to obtain a first sample audio signal corresponding to the first sample acoustic parameter group. Thereafter, a first group of users is summoned to listen to the first sample audio signal, and the semantic vocabulary of the first group of users' subjective evaluations is summarized. The first group of users can be a team of experts familiar with subjective psychoacoustic evaluations, and the number of users included in the first group can be 15. In the semantic vocabulary of the subjective evaluation of the first group of users, the frequently appearing words are collected in the form of adjective pairs and divided into semantic dimensions. For example, the high-frequency semantic tags include semantic tag A, and the semantic dimensions corresponding to the semantic tag A obtained by division are: very comfortable - relatively comfortable - medium - relatively intense - very intense. It should be understood that in the present disclosure, semantic dimensions can also be represented by numbers, for example, the number 1 represents the semantic dimension of very comfortable, the number 2 represents the semantic dimension of relatively comfortable, the number 3 represents the semantic dimension of medium, the number 4 represents the semantic dimension of relatively intense, and the number 5 represents the semantic dimension of very intense.

[0093] Furthermore, after summarizing the semantic vocabulary of the subjective evaluations of the first group of users, the first sample acoustic parameter group corresponding to the first sample audio signal with poor subjective listening experience can be determined as a prohibited acoustic parameter group. This ensures that the prohibited acoustic parameter group is not included in the training dataset during generation, thereby improving the efficiency of acoustic parameter model training.

[0094] Next, for each acoustic parameter to be adjusted, multiple second parameter values ​​of the acoustic parameter are determined according to the second partitioning granularity within the value range of the acoustic parameter. The first partitioning granularity is larger than the second partitioning granularity, i.e., the partitioning interval represented by the first partitioning granularity is larger than the partitioning interval represented by the second partitioning granularity. For example, if the partitioning interval represented by the first partitioning granularity is 5, the partitioning interval represented by the second partitioning granularity can be any value among 1, 2, 3, and 4.

[0095] Afterwards, multiple groups of second sample acoustic parameter groups are determined according to multiple second parameter values ​​of each acoustic parameter to be adjusted, and second sample audio signals corresponding to the second sample acoustic parameter groups are obtained. The second sample acoustic parameter groups include any second parameter value of each acoustic parameter to be adjusted.

[0096] The specific method of obtaining the second sample audio signal is similar to the specific method of obtaining the first sample audio signal, and is not repeated here.

[0097] Finally, second semantic feedback of the second group of users on the second sample audio signal based on the semantic dimension corresponding to each high-frequency semantic tag is obtained, and sample semantic information of each second sample audio signal is counted based on the second semantic feedback to obtain a training data set.

[0098] For example, the second group of users can be a public review team that is familiar with car driving and has requirements for vehicle acoustic performance, and the number of users in the second group can be 40. When the second group of users listens to the second sample audio signal, they select the semantic dimension that is closest to the second sample audio signal in the semantic dimension corresponding to each high-frequency semantic label. That is, for each second sample audio signal, each user will select the semantic dimension that is closest to the second sample audio signal in the semantic dimension corresponding to each high-frequency semantic label, or select the number corresponding to the semantic dimension as the sample semantic information that the second sample audio signal conforms to. In this way, based on the second semantic feedback of the second group of users, the sample semantic information of each second sample audio signal can be counted to obtain a training data set.

[0099] It is worth noting that for each second audio sample signal, the second group of users selects the semantic dimension that is closest to the second audio sample signal being auditioned from each high-frequency semantic label. For each high-frequency semantic label, the semantic dimension selected by the largest number of users is determined as the sample semantic information for the second audio sample signal under that high-frequency semantic label. For example, if the number of users in the second group is 40, and if, for second audio sample signal 1, 30 users select a medium semantic dimension under semantic label A, then the sample semantic information for second audio sample signal 1 is determined to be a medium semantic dimension under semantic label A.

[0100] At this point, the training data set can be obtained according to the above method.

[0101] In addition, in the present disclosure, in order to facilitate the user to select accurate semantic information, the method may further include: outputting high-frequency semantic tags and semantic dimensions corresponding to each high-frequency semantic tag, so that the user can select a target semantic dimension in each high-frequency semantic tag;

[0102] Accordingly, step S11 of determining the semantic information of the audio signal to be generated may include: determining the semantic information of the audio signal to be generated according to a target semantic dimension selected by a user.

[0103] For example, the target semantic dimension selected by the user or a number corresponding to the target semantic dimension may be determined as the semantic information of the audio signal to be generated.

[0104] In this way, predetermined high-frequency semantic tags and the semantic dimensions corresponding to each high-frequency semantic tag can be output. In this way, the user can select a suitable target semantic dimension from the predetermined semantic dimensions as the semantic information of the audio signal to be generated, thereby improving the reliability of the generated audio signal.

[0105] return Figure 3 , the training method may further include step S32.

[0106] In step S32, the preset initial model is trained according to the training data set to obtain an acoustic parameter model.

[0107] In one embodiment, the sample semantic information may be used as an input parameter of a preset initial model, and the sample acoustic parameter group may be used as an output parameter of the preset initial model. The initial model may be trained to obtain an acoustic parameter model.

[0108] For example, the sample semantic information is used as the input parameter of the preset initial model, and the sample acoustic parameter group is used as the output parameter of the preset initial model. When the number of training times, training time or model error meets the preset conditions, the training is terminated to obtain the acoustic parameter model.

[0109] In another embodiment, the initial model is a generative adversarial network, which includes a generative network and a discriminative network. Accordingly, step S32 may further include:

[0110] First, the discriminant network is trained according to the training data set until the discriminant network meets the first constraint.

[0111] In this embodiment, the discriminant network can be a semantic prediction model, and the discriminant network can be trained by using the sample acoustic parameter group as the model input parameter and the sample semantic information as the model output parameter, and training the discriminant network until the discriminant network meets the first constraint condition.

[0112] For example, the parameter value of each acoustic parameter to be adjusted in the sample acoustic parameter group can be normalized and used as a model input parameter. The acoustic parameter to be adjusted is used as an input dimension. Assuming that the extracted acoustic parameters to be adjusted are 16, the input layer of the discriminant network has at least 16 nodes, that is, there are 16 parameter values ​​of the acoustic parameters to be adjusted in the sample acoustic parameter group. Assuming that the number of high-frequency semantic labels is 5, the output layer of the discriminant network has at least 5 nodes, each node corresponds to a high-frequency semantic label, that is, each node is used to output the estimated semantic dimension under the high-frequency semantic label. In addition, the discriminant network also includes two hidden layers, wherein the nodes of the hidden layer can be 20*10. In this way, the sample acoustic parameter group is used as the model input parameter and the sample semantic information is used as the model output parameter, and the discriminant model is trained until the discriminant network meets the first constraint condition.

[0113] The first constraint condition may include but is not limited to the number of training times reaching a preset number or the duration of this round of training being greater than a preset duration.

[0114] Next, the generative network is trained based on the sample semantic information and the discriminative network that satisfies the first constraint until the generative network satisfies the second constraint.

[0115] In this embodiment, the generative network may also include an input layer, a hidden layer, and an output layer. The input layer of the generative network is used to input sample semantic information, and the output layer is used to output an estimated acoustic parameter group. Therefore, the input layer of the generative network has at least 5 nodes, each node corresponds to a high-frequency semantic label, that is, each node is used to input the semantic dimension under the high-frequency semantic label. The output layer of the generative network has at least 16 nodes, each node is used to output a parameter value of an acoustic parameter to be adjusted. In addition, the hidden layer of the generative network is two layers, and the nodes of the hidden layer can be 20*10.

[0116] For example, the specific implementation method of training the generative network according to the sample semantic information and the discriminant network that satisfies the first constraint condition until the generative network satisfies the second constraint condition can be: inputting the sample semantic information into the generative network to obtain the first acoustic parameter group corresponding to the sample semantic information output by the generative network, and inputting the first acoustic parameter group into the discriminant network that satisfies the first constraint condition to obtain the first semantic information output by the discriminant network; training the generative network according to the sample semantic information and the first semantic information until the generative network satisfies the second constraint condition.

[0117] For example, the sample semantic information is used as the true value, and the first semantic information is used as the semantic information estimated by the discriminant network. If the semantic information estimated by the discriminant network is consistent with the sample semantic information, then the first acoustic parameter group generated by the characterization generation network is accurate, that is, the first acoustic parameter group is the sample acoustic parameter group corresponding to the audio signal of the sample semantic information in the training data set. Therefore, in this embodiment, the parameters of the generation network can be adjusted according to the error between the sample semantic information and the first semantic information to realize the training of the generation network, so as to make the first acoustic parameter group output by the generation network as much as possible the sample acoustic parameter group corresponding to the audio signal of the sample semantic information in the training data set.

[0118] The second constraint condition may include but is not limited to the number of training times reaching a preset number or the duration of this round of training being greater than a preset duration.

[0119] Finally, if the generative network that satisfies the second constraint condition does not satisfy the preset training termination condition, the step of training the discriminant network according to the training data set is returned to execute until the discriminant network satisfies the first constraint condition, and until the generative network that satisfies the second constraint condition satisfies the preset training termination condition; if the generative network that satisfies the second constraint condition satisfies the preset training termination condition, the generative network is determined as an acoustic parameter model.

[0120] The training termination condition may be that the accuracy of the generated network exceeds a preset threshold. For example, when a generated network that satisfies the second constraint is obtained, the accuracy of the generated network is determined. If the accuracy reaches the preset threshold, the generated network is determined as the acoustic parameter model. Otherwise, the training process returns to the step of training the discriminant network based on the training dataset until the discriminant network satisfies the first constraint, and until the accuracy of the generated network that satisfies the second constraint reaches the preset threshold.

[0121] For example, the loss function of the generative adversarial network can be:

[0122]

[0123] Among them, G(z) is the generative network, D(x) is the discriminative network, E[] is the loss function, Z is the sample semantic information, x is the first acoustic parameter group, xp data (x) is the probability distribution of the first acoustic parameter group, zp z (z) is the probability distribution of the sample semantic information.

[0124] By adopting the above technical solution and utilizing a generative adversarial network to generate an acoustic parameter model, the generated acoustic parameter model has good stability, strong anti-interference ability, and high output accuracy.

[0125] At this point, the acoustic parameter model can be obtained in the above manner. Afterwards, the acoustic parameter model can be solidified and embedded in an audio system, for example, an in-vehicle simulated sound system. The user can select the semantic information he wants to modulate from the semantic dimension under each high-frequency semantic label in the human-computer interaction interface. For example, if a "very intense" simulated sound wave under semantic label A is required, the value "5" can be selected under semantic label A. Afterwards, the device that executes the audio signal generation method can input the semantic information selected by the user into the acoustic parameter model to obtain the target acoustic parameter group. Finally, the target acoustic parameter group is input into the simulated sound wave algorithm model to change and modulate the output simulated sound wave in real time to a sound wave that conforms to the semantic information, thereby achieving the purpose of customizing the simulated sound wave based on semantic information.

[0126] It should be understood that the human-computer interaction interface described herein can be a user-operable display screen, or can also be a device capable of detecting user voice or user actions. For example, the human-computer interaction interface is a device for detecting user voice, which can collect user voice to obtain semantic information of the user input from the user voice. This disclosure does not specifically limit this.

[0127] Based on the same inventive concept, the present disclosure also provides an audio signal generating device. Figure 4 FIG. 1 is a block diagram of an audio signal generating device according to an exemplary embodiment. Figure 4 As shown, the audio signal generating device 400 may include:

[0128] A first determining module 401 is configured to determine semantic information of the audio signal to be generated;

[0129] A second determining module 402 is configured to input the semantic information into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model, wherein the target acoustic parameter group includes a parameter value of each acoustic parameter to be adjusted;

[0130] The generating module 403 is configured to generate an audio signal that conforms to the semantic information according to the target acoustic parameter group.

[0131] Optionally, the audio signal generating device 400 may include:

[0132] A first acquisition module is configured to acquire a training data set, wherein the training data set includes multiple groups of training samples, and the training samples include sample semantic information and a sample acoustic parameter group corresponding to the audio signal that conforms to the sample semantic information;

[0133] The training module is configured to train a preset initial model according to the training data set to obtain the acoustic parameter model.

[0134] Optionally, the initial model is a generative adversarial network, which includes a generative network and a discriminative network; and the training module includes:

[0135] a first training submodule, configured to train the discriminant network according to the training data set until the discriminant network satisfies a first constraint;

[0136] a second training submodule, configured to train the generative network according to the sample semantic information and the discriminant network satisfying the first constraint condition, until the generative network satisfies the second constraint condition;

[0137] an execution submodule, configured to return to the step of training the discriminant network according to the training data set until the discriminant network satisfies the first constraint, and until the generation network that satisfies the second constraint satisfies the preset training termination condition, if the generation network that satisfies the second constraint does not satisfy the preset training termination condition;

[0138] The second determining submodule is configured to determine the generating network that satisfies the second constraint condition as the acoustic parameter model if the generating network satisfies a preset training termination condition.

[0139] Optionally, the first training submodule is configured to: use the sample acoustic parameter group as a model input parameter and the sample semantic information as a model output parameter to train the discriminant network until the discriminant network satisfies a first constraint condition;

[0140] The second training submodule is configured to: input the sample semantic information into the generative network to obtain a first acoustic parameter group corresponding to the sample semantic information output by the generative network, and input the first acoustic parameter group into a discriminative network that satisfies a first constraint condition to obtain first semantic information output by the discriminative network;

[0141] The generation network is trained according to the sample semantic information and the first semantic information until the generation network satisfies a second constraint condition.

[0142] Optionally, the training module may further include:

[0143] The third training submodule is configured to: use the sample semantic information as the input parameter of the preset initial model, use the sample acoustic parameter group as the output parameter of the preset initial model, train the initial model, and obtain the acoustic parameter model.

[0144] Optionally, the first acquisition module may include:

[0145] an extraction submodule, configured to extract acoustic parameters to be adjusted from an audio signal generation algorithm;

[0146] The generating submodule is configured to generate a plurality of sample acoustic parameter groups according to the acoustic parameters to be adjusted, and obtain sample semantic information corresponding to the sample audio signal corresponding to each of the sample acoustic parameter groups to obtain a training data set.

[0147] Optionally, the audio signal generation algorithm includes: a first type of algorithm for generating a single-track audio signal, a second type of algorithm for generating a full-track audio signal based on the single-track audio signal, and a third type of algorithm for processing the full-track audio signal to obtain an audio signal to be played;

[0148] The first type of algorithm includes a frequency shift algorithm and a harmonic algorithm; the frequency shift algorithm is used to obtain a frequency shift phase difference based on the frequency of the initial audio signal, the distance between the initial audio signal and the listening point, the frequency shift rate, and the intensity change rate of the audio signal, and phase-shift the initial audio signal according to the frequency shift phase difference to obtain a frequency-shifted single-track audio signal; the harmonic algorithm is used to obtain a harmonically processed single-track audio signal based on the amplitude of each harmonic, the frequency conversion rate, the motor speed, and the order of each harmonic;

[0149] The second type of algorithm is used to obtain a total track audio signal based on the frequency-shifted single-track audio signal, the harmonically processed single-track audio signal, and the loudness of the corresponding single-track audio signals;

[0150] The third type of algorithm includes a delay algorithm and a frequency domain modulation algorithm; the delay algorithm is used to obtain a delayed total track audio signal based on the total track audio signal and the delay amount; the frequency domain modulation algorithm is used to obtain an audio signal to be played based on a filter coefficient, the delayed total track audio signal at the current moment, and the delayed total track audio signal at the previous moment.

[0151] Optionally, the extraction submodule is configured to:

[0152] Extracting acoustic parameters to be adjusted from at least one of the first type of algorithms, the second type of algorithms, and the third type of algorithms;

[0153] Among them, in the first type of algorithm, at least one of the frequency shift rate, the intensity change rate of the audio signal, the harmonic amplitude, the frequency conversion rate and the harmonic order is used as the acoustic parameter to be adjusted; in the second type of algorithm, the loudness of the single-track audio signal is used as the acoustic parameter to be adjusted; in the third type of algorithm, the delay amount and / or the filter coefficient is used as the acoustic parameter to be adjusted.

[0154] Optionally, the generating submodule is configured to:

[0155] Determine high-frequency semantic tags and the semantic dimensions corresponding to each high-frequency semantic tag;

[0156] For each acoustic parameter to be adjusted, determining a plurality of second parameter values ​​of the acoustic parameter according to a second division granularity within a value range of the acoustic parameter;

[0157] Determining a plurality of second sample acoustic parameter groups according to the plurality of second parameter values ​​of each acoustic parameter to be adjusted, and obtaining second sample audio signals corresponding to the second sample acoustic parameter groups, wherein the second sample acoustic parameter groups include any second parameter value of each acoustic parameter to be adjusted;

[0158] Obtain second semantic feedback from the second group of users on the second sample audio signal based on the semantic dimension corresponding to each high-frequency semantic tag, and count sample semantic information of each second sample audio signal based on the second semantic feedback to obtain a training data set.

[0159] Optionally, the generating submodule is further configured to:

[0160] For each acoustic parameter to be adjusted, within a value range of the acoustic parameter, determining a plurality of first parameter values ​​of the acoustic parameter according to a first division granularity, where the first division granularity is greater than the second division granularity;

[0161] Determining a plurality of first sample acoustic parameter groups according to a plurality of first parameter values ​​of each acoustic parameter to be adjusted, and obtaining first sample audio signals corresponding to the first sample acoustic parameter groups, wherein the first sample acoustic parameter groups include any first parameter value of each acoustic parameter to be adjusted;

[0162] First semantic feedback of the first group of users on the first sample audio signal is counted, and high-frequency semantic tags and semantic dimensions corresponding to each high-frequency semantic tag are determined according to the first semantic feedback.

[0163] Optionally, the audio signal generating device 400 may further include:

[0164] an output module configured to output high-frequency semantic tags and semantic dimensions corresponding to each high-frequency semantic tag, so that a user can select a target semantic dimension from each high-frequency semantic tag;

[0165] The first determination module 401 is configured to determine semantic information of the audio signal to be generated according to the target semantic dimension selected by the user.

[0166] Optionally, the audio signal includes an analog sound wave signal.

[0167] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0168] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the audio signal generating method provided by the present disclosure are implemented.

[0169] The present disclosure also provides an electronic device, comprising: a processor;

[0170] a memory for storing processor-executable instructions;

[0171] The processor is configured to execute the executable instructions to implement the steps of the audio signal generation method provided by the present disclosure.

[0172] Figure 5 8 is a block diagram of an electronic device according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0173] Reference Figure 5 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0174] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0175] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0176] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0177] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0178] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0179] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0180] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0181] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0182] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0183] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0184] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for executing the above-mentioned audio signal generation method when executed by the programmable device.

[0185] The present disclosure also provides an audio system, comprising the electronic device provided by the present disclosure and an acoustic parameter model, wherein the acoustic parameter model is used to input semantic information and output a target acoustic parameter group.

[0186] The present disclosure also provides a vehicle including the audio system provided by the present disclosure.

[0187] Figure 6 6 is a block diagram illustrating a vehicle according to an exemplary embodiment. For example, vehicle 600 may be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or another type of vehicle. Vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0188] Reference Figure 6 The vehicle 600 may include various subsystems, such as an infotainment system 610, a perception system 620, a decision control system 630, a drive system 640, and a computing platform 650. The vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. For example, the vehicle 600 may also include an audio system, which may include Figure 5 The electronic device and acoustic parameter model shown are used to input semantic information and output a target acoustic parameter set. Furthermore, the audio system can be a simulated sound system. Furthermore, each subsystem and each component of vehicle 600 can be interconnected via wired or wireless means.

[0189] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, a navigation system, and the like.

[0190] The perception system 620 may include several sensors for sensing information about the environment surrounding the vehicle 600. For example, the perception system 620 may include a global positioning system (which may be a GPS system, a BeiDou system, or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter-wave radar, an ultrasonic radar, and a camera.

[0191] The decision control system 630 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.

[0192] The drive system 640 may include components that provide power to the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be an internal combustion engine, an electric motor, an air compression engine, or a combination thereof. The engine is capable of converting energy provided by the energy source into mechanical energy.

[0193] Some or all functions of the vehicle 600 are controlled by a computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652. The processor 651 may execute instructions 653 stored in the memory 652.

[0194] The processor 651 can be any conventional processor, such as a commercially available CPU. The processor can also include a graphics processor (GPU), a field programmable gate array (FPGA), a system on chip (SOC), an application specific integrated circuit (ASIC), or a combination thereof.

[0195] The memory 652 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0196] In addition to instructions 653 , memory 652 may also store data, such as road maps, route information, and vehicle location, direction, speed, etc. The data stored in memory 652 may be used by computing platform 650 .

[0197] In the embodiment of the present disclosure, the processor 651 may execute the instruction 653 to complete all or part of the steps of the above-mentioned audio signal generating method.

[0198] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.

[0199] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art after reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific functions of the described components, even if structurally not equivalent to the disclosed structures. In addition, although specific features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and beneficial for any given or specific application. In addition, with respect to the terms "including," "having," "having," "having," or variations thereof used in the specific embodiments or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0200] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

[0201] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0202] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.

[0203] Although terms such as "first", "second" and "third" may be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer or section from another component, part, region, layer or section. Therefore, without departing from the teachings of each example, the first component, part, region, layer or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" can explicitly or implicitly include at least one such feature. In the description herein, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically defined.

Claims

1. A method for generating an audio signal, characterized in that: include: determining semantic information of the audio signal to be generated; Inputting the semantic information into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model, wherein the target acoustic parameter group includes a parameter value of each acoustic parameter to be adjusted; An audio signal conforming to the semantic information is generated according to the target acoustic parameter group.

2. The method according to claim 1, characterized in that The acoustic parameter model is trained in the following way: Acquire a training data set, the training data set including multiple groups of training samples, the training samples including sample semantic information and a sample acoustic parameter group corresponding to an audio signal that conforms to the sample semantic information; The preset initial model is trained according to the training data set to obtain the acoustic parameter model.

3. The method according to claim 2, characterized in that The initial model is a generative adversarial network, which includes a generative network and a discriminative network; The step of training a preset initial model based on the training data set to obtain the acoustic parameter model includes: Training the discriminant network according to the training data set until the discriminant network satisfies a first constraint; Training the generative network according to the sample semantic information and the discriminant network that satisfies the first constraint condition until the generative network satisfies the second constraint condition; If the generated network that satisfies the second constraint condition does not satisfy the preset training termination condition, returning to the step of training the discriminant network according to the training data set until the discriminant network satisfies the first constraint condition, and until the generated network that satisfies the second constraint condition satisfies the preset training termination condition; If the generated network that meets the second constraint condition meets a preset training termination condition, the generated network is determined as the acoustic parameter model.

4. The method according to claim 3, characterized in that The step of training the discriminant network according to the training data set until the discriminant network satisfies a first constraint condition comprises: Using the sample acoustic parameter group as a model input parameter and the sample semantic information as a model output parameter, the discriminant network is trained until the discriminant network satisfies a first constraint condition; The step of training the generative network according to the sample semantic information and the discriminant network satisfying the first constraint condition until the generative network satisfies the second constraint condition comprises: Inputting the sample semantic information into the generative network to obtain a first acoustic parameter group corresponding to the sample semantic information output by the generative network, and inputting the first acoustic parameter group into a discriminative network that satisfies a first constraint condition to obtain first semantic information output by the discriminative network; The generation network is trained according to the sample semantic information and the first semantic information until the generation network satisfies a second constraint condition.

5. The method according to claim 2, characterized in that The step of training a preset initial model based on the training data set to obtain the acoustic parameter model includes: The sample semantic information is used as an input parameter of a preset initial model, the sample acoustic parameter group is used as an output parameter of the preset initial model, and the initial model is trained to obtain the acoustic parameter model.

6. The method according to claim 2, characterized in that The obtaining of the training data set includes: Extracting acoustic parameters to be adjusted from an audio signal generation algorithm; A plurality of sample acoustic parameter groups are generated according to the acoustic parameters to be adjusted, and sample semantic information corresponding to the sample audio signals corresponding to each of the sample acoustic parameter groups is obtained to obtain a training data set.

7. The method according to claim 6, characterized in that The audio signal generation algorithm includes: a first type of algorithm for generating a single-track audio signal, a second type of algorithm for generating a full-track audio signal based on the single-track audio signal, and a third type of algorithm for processing the full-track audio signal to obtain an audio signal to be played; The first type of algorithm includes a frequency shift algorithm and a harmonic algorithm; the frequency shift algorithm is used to obtain a frequency shift phase difference based on the frequency of the initial audio signal, the distance between the initial audio signal and the listening point, the frequency shift rate, and the intensity change rate of the audio signal, and phase-shift the initial audio signal according to the frequency shift phase difference to obtain a frequency-shifted single-track audio signal; the harmonic algorithm is used to obtain a harmonically processed single-track audio signal based on the amplitude of each harmonic, the frequency conversion rate, the motor speed, and the order of each harmonic; The second type of algorithm is used to obtain a total track audio signal based on the frequency-shifted single-track audio signal, the harmonically processed single-track audio signal, and the loudness of the corresponding single-track audio signals; The third type of algorithm includes a delay algorithm and a frequency domain modulation algorithm; the delay algorithm is used to obtain a delayed total track audio signal based on the total track audio signal and the delay amount; the frequency domain modulation algorithm is used to obtain an audio signal to be played based on a filter coefficient, the delayed total track audio signal at the current moment, and the delayed total track audio signal at the previous moment.

8. The method according to claim 7, characterized in that The step of extracting the acoustic parameters to be adjusted from the audio signal generation algorithm includes: Extracting acoustic parameters to be adjusted from at least one of the first type of algorithms, the second type of algorithms, and the third type of algorithms; Among them, in the first type of algorithm, at least one of the frequency shift rate, the intensity change rate of the audio signal, the harmonic amplitude, the frequency conversion rate and the harmonic order is used as the acoustic parameter to be adjusted; in the second type of algorithm, the loudness of the single-track audio signal is used as the acoustic parameter to be adjusted; in the third type of algorithm, the delay amount and / or the filter coefficient is used as the acoustic parameter to be adjusted.

9. The method according to claim 6, characterized in that The step of generating a plurality of sample acoustic parameter groups according to the acoustic parameters to be adjusted, and obtaining sample semantic information corresponding to the sample audio signals corresponding to each of the sample acoustic parameter groups to obtain a training data set includes: Determine high-frequency semantic tags and the semantic dimensions corresponding to each high-frequency semantic tag; For each acoustic parameter to be adjusted, determining a plurality of second parameter values ​​of the acoustic parameter according to a second division granularity within a value range of the acoustic parameter; Determining a plurality of second sample acoustic parameter groups according to the plurality of second parameter values ​​of each acoustic parameter to be adjusted, and obtaining second sample audio signals corresponding to the second sample acoustic parameter groups, wherein the second sample acoustic parameter groups include any second parameter value of each acoustic parameter to be adjusted; Obtain second semantic feedback from the second group of users on the second sample audio signal based on the semantic dimension corresponding to each high-frequency semantic tag, and count sample semantic information of each second sample audio signal based on the second semantic feedback to obtain a training data set.

10. The method according to claim 9, characterized in that The determining of the high-frequency semantic tags and the semantic dimension corresponding to each high-frequency semantic tag includes: For each acoustic parameter to be adjusted, within a value range of the acoustic parameter, determining a plurality of first parameter values ​​of the acoustic parameter according to a first division granularity, where the first division granularity is greater than the second division granularity; Determining a plurality of first sample acoustic parameter groups according to a plurality of first parameter values ​​of each acoustic parameter to be adjusted, and obtaining first sample audio signals corresponding to the first sample acoustic parameter groups, wherein the first sample acoustic parameter groups include any first parameter value of each acoustic parameter to be adjusted; First semantic feedback of the first group of users on the first sample audio signal is counted, and high-frequency semantic tags and semantic dimensions corresponding to each high-frequency semantic tag are determined according to the first semantic feedback.

11. The method according to any one of claims 1 to 10, characterized in that The method further comprises: Outputting high-frequency semantic tags and semantic dimensions corresponding to each high-frequency semantic tag, so that a user can select a target semantic dimension from each high-frequency semantic tag; The determining of semantic information of the audio signal to be generated includes: The semantic information of the audio signal to be generated is determined according to the target semantic dimension selected by the user.

12. The method according to any one of claims 1 to 10, characterized in that The audio signal includes an analog sound wave signal.

13. An audio signal generating device, characterized in that: include: A first determining module is configured to determine semantic information of the audio signal to be generated; a second determining module configured to input the semantic information into a preset acoustic parameter model to obtain a target acoustic parameter group output by the acoustic parameter model, wherein the target acoustic parameter group includes a parameter value of each acoustic parameter to be adjusted; A generating module is configured to generate an audio signal that conforms to the semantic information according to the target acoustic parameter group.

14. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions to implement the audio signal generation method according to any one of claims 1 to 12.

15. An audio system, characterized in that The device comprises the electronic device as claimed in claim 14 and an acoustic parameter model, wherein the acoustic parameter model is used to input semantic information and output a target acoustic parameter group.

16. A vehicle, characterized in that: Comprising the audio system of claim 15.

17. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

18. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 12 when the computer program is executed by a processor.