Audio data mixing method and system based on prediction algorithm
Through the audio data mixing method based on prediction algorithm, combined with audio characteristics and user attention, the problem of poor audio data mixing effect and efficiency in the prior art is solved, and more efficient and accurate audio data mixing is achieved, improving the user experience.
Patent Information
- Application Number
- CN202510067105.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing audio data mixing technology lacks the use of prediction algorithms to combine audio data characteristics and user feedback, resulting in poor mixing effect and efficiency and poor user experience.
Using an audio data mixing method based on a prediction algorithm, a mixing strategy is determined based on this information by obtaining the audio data to be mixed, and a mixing strategy is determined based on this information to achieve more efficient and accurate mixing of audio data.
Improve the effect and efficiency of audio mixing, give users a better sound experience, and achieve a more accurate mixing strategy by comprehensively considering audio characteristics and user attention.
Smart Images

Figure CN119937974A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to an audio data mixing method and system based on a prediction algorithm. Background Art
[0002] The accurate mixing and display of audio data can effectively improve the accuracy of information dissemination and the user's sound environment experience. Therefore, how to effectively and accurately mix different audio data is an important technical issue. However, most of the existing audio data mixing technologies still use artificially set rules to directly mix audio data from different sources, and do not use prediction algorithms to combine the characteristics of audio data and user feedback to achieve more accurate mixing. Therefore, the effect and efficiency of audio mixing are poor, and the user experience is also relatively general. It can be seen that the existing technology has defects that need to be solved urgently. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide an audio data mixing method and system based on a prediction algorithm, which can achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect and provide users with a better sound experience.
[0004] In order to solve the above technical problems, the first aspect of the present invention discloses an audio data mixing method based on a prediction algorithm, the method comprising: Acquire multiple audio data to be mixed; Extracting audio feature parameters corresponding to each of the audio data based on a feature extraction model; Determining user attention corresponding to each of the audio data based on historical user feedback records; According to the audio feature parameters and the user attention, based on a prediction algorithm, a mixing strategy corresponding to the multiple audio data is determined; the mixing strategy is used to indicate an operation of mixing the multiple audio data.
[0005] As an optional implementation, in the first aspect of the present invention, the audio feature parameters include one or more of audio duration, audio volume change, audio instrument type and audio text content.
[0006] As an optional implementation, in the first aspect of the present invention, extracting audio feature parameters corresponding to each audio data based on a feature extraction model includes: For each of the audio data, the audio data is input into a trained classifier algorithm to determine an audio type parameter corresponding to the audio data; the audio type parameter includes one or more of a vocal type, a single instrument type, a mixed instrument type, and a music genre type; the classifier algorithm is trained by a training data set including a plurality of training audio data and corresponding audio type parameter annotations; Based on the audio type parameter, determining a target feature extraction model from a plurality of candidate feature extraction models; The audio data is input into the target feature extraction model to obtain audio feature parameters corresponding to the audio data.
[0007] As an optional implementation, in the first aspect of the present invention, determining a target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter comprises: For each candidate feature extraction model, calculating the similarity between the audio type annotation in the training data set corresponding to the candidate feature extraction model and the audio type parameter; The candidate feature extraction model with the highest similarity is determined as the target feature extraction model; wherein the candidate feature extraction model is trained by the following steps: Training a preset fully convolutional neural network based on a training data set including a plurality of training audio data and corresponding audio feature parameter annotations until convergence; The parameters of the feature extraction layer in the fully convolutional neural network are fixed to obtain a candidate feature extraction model.
[0008] As an optional implementation, in the first aspect of the present invention, determining the user attention corresponding to each audio data based on the historical user feedback record includes: Acquire historical user feedback records including feedback records of multiple users on different types of audio data; For each of the audio data, according to the audio feature parameter corresponding to the audio data, a plurality of similar audio data are screened out from all the evaluated audio data in the historical user feedback record; the feature similarity between the audio feature of the similar audio data and the audio feature parameter is greater than a first similarity threshold; The weighted sum average of the user feedback parameters corresponding to each of the similar audio data is calculated to obtain the user attention corresponding to the audio data.
[0009] As an optional implementation, in the first aspect of the present invention, the user feedback parameter is the user's attention duration to the similar audio data, and the attention duration is calculated by the following steps: When playing mixed audio including the similar audio data and other audio data to the user, obtaining a user face image; Based on the face attention recognition model, continuously identifying the attention parameter corresponding to the user's face image; First, continuously reduce the proportion of the similar audio data in the mixed audio until the attention parameter is lower than a preset first parameter threshold, and calculate the reduction duration; Then continue to increase the proportion of the similar audio data in the mixed audio until the attention parameter is higher than a preset second parameter threshold, and calculate the increase duration; the second parameter threshold is higher than the first parameter threshold; The ratio of the reduced duration to the increased duration is calculated to obtain the attention duration corresponding to the similar audio data.
[0010] As an optional embodiment, in the first aspect of the present invention, when calculating the user attention, the weighted calculation weight corresponding to the user feedback parameter corresponding to each of the similar audio data is the product of a first weight and a second weight; the first weight is inversely proportional to the audio duration and / or average audio volume of the similar audio data; the second weight is directly proportional to the feature similarity corresponding to the similar audio data.
[0011] As an optional implementation, in the first aspect of the present invention, determining the mixing strategy corresponding to the multiple audio data based on the prediction algorithm according to the audio feature parameters and the user attention includes: The objective function is set to a user satisfaction level corresponding to the hybrid strategy that is greater than a satisfaction threshold; the user satisfaction level is obtained by inputting the hybrid parameters corresponding to all the audio data in the hybrid strategy and the user attention into a trained user satisfaction prediction model; the hybrid parameters include a hybrid volume weight, an audio mixing mode, and a hybrid audio part; Setting restrictions includes: The predicted audio average volume corresponding to the mixed volume weight of any audio data in the mixed strategy is greater than the volume threshold; the predicted audio average volume is obtained by inputting any audio data and the corresponding mixed volume weight into a volume prediction model; The parameter similarity between the mixing parameters corresponding to any two of the audio data in the mixing strategy is greater than a second similarity threshold; Based on a dynamic programming algorithm, the multiple audio data are iteratively calculated according to the objective function and the limiting conditions to obtain a mixing strategy corresponding to the multiple audio data.
[0012] A second aspect of an embodiment of the present invention discloses an audio data mixing system based on a prediction algorithm, the system comprising: An acquisition module, used for acquiring a plurality of audio data to be mixed; An extraction module, used for extracting audio feature parameters corresponding to each of the audio data based on a feature extraction model; A determination module, used to determine the user attention corresponding to each of the audio data based on historical user feedback records; A prediction module is used to determine a mixing strategy corresponding to the multiple audio data based on a prediction algorithm according to the audio feature parameters and the user attention; the mixing strategy is used to indicate an operation of mixing the multiple audio data.
[0013] As an optional implementation, in the second aspect of the present invention, the audio feature parameters include one or more of audio duration, audio volume change, audio instrument type and audio text content.
[0014] As an optional implementation, in the second aspect of the present invention, the specific manner in which the extraction module extracts the audio feature parameters corresponding to each of the audio data based on the feature extraction model includes: For each of the audio data, the audio data is input into a trained classifier algorithm to determine an audio type parameter corresponding to the audio data; the audio type parameter includes one or more of a vocal type, a single instrument type, a mixed instrument type, and a music genre type; the classifier algorithm is trained by a training data set including a plurality of training audio data and corresponding audio type parameter annotations; Based on the audio type parameter, determining a target feature extraction model from a plurality of candidate feature extraction models; The audio data is input into the target feature extraction model to obtain audio feature parameters corresponding to the audio data.
[0015] As an optional implementation, in the second aspect of the present invention, the specific manner in which the extraction module determines the target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter includes: For each candidate feature extraction model, calculating the similarity between the audio type annotation in the training data set corresponding to the candidate feature extraction model and the audio type parameter; The candidate feature extraction model with the highest similarity is determined as the target feature extraction model; wherein the candidate feature extraction model is trained by the following steps: Training a preset fully convolutional neural network based on a training data set including a plurality of training audio data and corresponding audio feature parameter annotations until convergence; The parameters of the feature extraction layer in the fully convolutional neural network are fixed to obtain a candidate feature extraction model.
[0016] As an optional implementation, in the second aspect of the present invention, the determination module determines the specific manner of user attention corresponding to each audio data based on historical user feedback records, including: Acquire historical user feedback records including feedback records of multiple users on different types of audio data; For each of the audio data, according to the audio feature parameter corresponding to the audio data, a plurality of similar audio data are screened out from all the evaluated audio data in the historical user feedback record; the feature similarity between the audio feature of the similar audio data and the audio feature parameter is greater than a first similarity threshold; The weighted sum average of the user feedback parameters corresponding to each of the similar audio data is calculated to obtain the user attention corresponding to the audio data.
[0017] As an optional implementation, in the second aspect of the present invention, the user feedback parameter is the user's attention duration to the similar audio data, and the attention duration is calculated by the following steps: When playing mixed audio including the similar audio data and other audio data to the user, obtaining a user face image; Based on the face attention recognition model, continuously identifying the attention parameter corresponding to the user's face image; First, continuously reduce the proportion of the similar audio data in the mixed audio until the attention parameter is lower than a preset first parameter threshold, and calculate the reduction duration; Then continue to increase the proportion of the similar audio data in the mixed audio until the attention parameter is higher than a preset second parameter threshold, and calculate the increase duration; the second parameter threshold is higher than the first parameter threshold; The ratio of the reduced duration to the increased duration is calculated to obtain the attention duration corresponding to the similar audio data.
[0018] As an optional embodiment, in the second aspect of the present invention, when calculating the user attention, the weighted calculation weight corresponding to the user feedback parameter corresponding to each of the similar audio data is the product of a first weight and a second weight; the first weight is inversely proportional to the audio duration and / or average audio volume of the similar audio data; the second weight is directly proportional to the feature similarity corresponding to the similar audio data.
[0019] As an optional implementation, in the second aspect of the present invention, the prediction module determines the specific manner of the mixing strategy corresponding to the multiple audio data based on the prediction algorithm according to the audio feature parameters and the user attention, including: The objective function is set to a user satisfaction level corresponding to the hybrid strategy that is greater than a satisfaction threshold; the user satisfaction level is obtained by inputting the hybrid parameters corresponding to all the audio data in the hybrid strategy and the user attention into a trained user satisfaction prediction model; the hybrid parameters include a hybrid volume weight, an audio mixing mode, and a hybrid audio part; Setting restrictions includes: The predicted audio average volume corresponding to the mixed volume weight of any audio data in the mixed strategy is greater than the volume threshold; the predicted audio average volume is obtained by inputting any audio data and the corresponding mixed volume weight into a volume prediction model; The parameter similarity between the mixing parameters corresponding to any two of the audio data in the mixing strategy is greater than a second similarity threshold; Based on a dynamic programming algorithm, the multiple audio data are iteratively calculated according to the objective function and the limiting conditions to obtain a mixing strategy corresponding to the multiple audio data.
[0020] The third aspect of the present invention discloses another audio data mixing system based on a prediction algorithm, the system comprising: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute part or all of the steps in the audio data mixing method based on the prediction algorithm disclosed in the first aspect of the present invention.
[0021] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute part or all of the steps in the audio data mixing method based on the prediction algorithm disclosed in the first aspect of the present invention.
[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention can extract audio feature parameters corresponding to each audio data based on a feature extraction model, and then determine the user attention corresponding to each audio data based on historical user feedback records, so as to comprehensively and accurately predict the mixing strategies corresponding to multiple audio data, thereby being able to achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 It is a flowchart of an audio data mixing method based on a prediction algorithm disclosed in an embodiment of the present invention.
[0025] Figure 2 It is a structural schematic diagram of an audio data mixing system based on a prediction algorithm disclosed in an embodiment of the present invention.
[0026] Figure 3 It is a structural schematic diagram of another audio data mixing system based on a prediction algorithm disclosed in an embodiment of the present invention.
[0027] Figure 4 It is a code diagram of an audio data mixing dynamic programming algorithm disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or equipment.
[0030] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0031] The present invention discloses an audio data mixing method and system based on a prediction algorithm, which can extract audio feature parameters corresponding to each audio data based on a feature extraction model, and then determine the user attention corresponding to each audio data based on historical user feedback records, so as to comprehensively and accurately predict the mixing strategies corresponding to multiple audio data, thereby being able to achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience. The following are detailed descriptions.
[0032] Embodiment 1 See also Figure 1 , Figure 1 is a flow chart of an audio data mixing method based on a prediction algorithm disclosed in an embodiment of the present invention. Figure 1 The described audio data mixing method based on the prediction algorithm can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 1 As shown, the audio data mixing method based on the prediction algorithm may include the following operations: 101. Obtain multiple audio data to be mixed.
[0033] 102. Extract audio feature parameters corresponding to each audio data based on the feature extraction model. 103. Based on historical user feedback records, determine the user attention corresponding to each audio data. 104. Determine a mixing strategy corresponding to the multiple audio data based on the prediction algorithm according to the audio feature parameters and the user's attention.
[0034] Optionally, the mixing strategy is used to indicate an operation of mixing multiple audio data.
[0035] It can be seen that the above-mentioned embodiments of the invention can extract the audio feature parameters corresponding to each audio data based on the feature extraction model, and then determine the user attention corresponding to each audio data based on the historical user feedback records, so as to comprehensively and accurately predict the mixing strategies corresponding to multiple audio data, thereby being able to achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience.
[0036] As an optional embodiment, in the above steps, the audio feature parameters include one or more of audio duration, audio volume change, audio instrument type and audio text content.
[0037] It can be seen that through the above-mentioned optional embodiments, the content of the audio feature parameters is clarified, which can more accurately and comprehensively characterize the characteristics of the audio data, so as to facilitate subsequent audio mixing predictions, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effects and provide users with a better sound experience.
[0038] As an optional embodiment, in the above step, extracting audio feature parameters corresponding to each audio data based on the feature extraction model includes: For each audio data, the audio data is input into a trained classifier algorithm to determine an audio type parameter corresponding to the audio data; optionally, the audio type parameter includes one or more of a vocal type, a single instrument type, a mixed instrument type, and a music genre type; the classifier algorithm is trained by a training data set including a plurality of training audio data and corresponding audio type parameter annotations; Determine a target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter; The audio data is input into a target feature extraction model to obtain audio feature parameters corresponding to the audio data.
[0039] It can be seen that through the above-mentioned optional embodiments, it is possible to determine the audio type parameters of the audio data based on the trained classifier algorithm and then screen out the target feature extraction model to extract the audio feature parameters corresponding to the audio data, so as to facilitate subsequent audio mixing predictions, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, thereby giving users a better sound experience.
[0040] As an optional embodiment, in the above step, determining a target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter includes: For each candidate feature extraction model, calculate the similarity between the audio type annotation and the audio type parameter in the training data set corresponding to the candidate feature extraction model; The candidate feature extraction model with the highest similarity is determined as the target feature extraction model; wherein the candidate feature extraction model is trained by the following steps: Training a preset fully convolutional neural network based on a training data set including a plurality of training audio data and corresponding audio feature parameter annotations until convergence; The parameters of the feature extraction layer in the fully convolutional neural network are fixed to obtain a candidate feature extraction model.
[0041] It can be seen that through the above-mentioned optional embodiments, a suitable feature extraction model can be screened out based on the similarity between the audio type annotations and the audio type parameters in the training data set corresponding to the candidate feature extraction model, so as to facilitate subsequent audio mixing predictions, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effects, thereby giving users a better sound experience.
[0042] As an optional embodiment, in the above step, determining the user attention corresponding to each audio data based on the historical user feedback record includes: Acquire historical user feedback records including feedback records of multiple users on different types of audio data; For each audio data, according to the audio feature parameters corresponding to the audio data, multiple similar audio data are screened out from all evaluated audio data in the historical user feedback records; optionally, the feature similarity between the audio features of the similar audio data and the audio feature parameters is greater than a first similarity threshold; The weighted sum average of the user feedback parameters corresponding to each similar audio data is calculated to obtain the user attention corresponding to the audio data.
[0043] It can be seen that through the above-mentioned optional embodiments, the user attention corresponding to each audio data can be determined based on the user feedback parameters corresponding to the audio data similar to each audio data in the evaluated audio data in the historical user feedback records, so as to facilitate the subsequent audio mixing prediction, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience.
[0044] As an optional embodiment, in the above steps, the user feedback parameter is the user's attention duration to the similar audio data, and the attention duration is calculated by the following steps: When playing mixed audio including similar audio data and other audio data to the user, obtaining a user face image; Based on the face attention recognition model, continuously identify the attention parameters corresponding to the user's face image; First, continuously reduce the proportion of similar audio data in the mixed audio until the attention parameter is lower than the preset first parameter threshold, and calculate the reduction duration; Then continue to increase the proportion of similar audio data in the mixed audio until the attention parameter is higher than a preset second parameter threshold, and calculate the increase duration; optionally, the second parameter threshold is higher than the first parameter threshold; The ratio of the reduced duration to the increased duration is calculated to obtain the attention duration corresponding to the similar audio data.
[0045] It can be seen that through the above optional embodiments, the content and calculation steps of the user feedback parameters are clarified, which can accurately characterize the user's attention duration on the audio data, so as to facilitate subsequent audio mixing prediction, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, thereby giving the user a better sound experience.
[0046] As an optional embodiment, in the above steps, when calculating user attention, the weighted calculation weight corresponding to the user feedback parameter corresponding to each similar audio data is the product of the first weight and the second weight; the first weight is inversely proportional to the audio duration and / or the average audio volume of the similar audio data; the second weight is directly proportional to the feature similarity corresponding to the similar audio data.
[0047] It can be seen that through the above-mentioned optional embodiments, the weight rules corresponding to different similar audio data in the user attention calculation are defined, so that the calculated user attention accurately represents the user's attention duration on the audio data, so as to facilitate subsequent audio mixing prediction, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, thereby giving the user a better sound experience.
[0048] As an optional embodiment, in the above steps, determining the mixing strategy corresponding to the multiple audio data based on the prediction algorithm according to the audio feature parameters and the user's attention includes: The objective function is set to that the user satisfaction corresponding to the hybrid strategy is greater than the satisfaction threshold; optionally, the user satisfaction is obtained by inputting the hybrid parameters corresponding to all audio data in the hybrid strategy and the user attention into the trained user satisfaction prediction model; the hybrid parameters include the hybrid volume weight, the audio mixing mode and the hybrid audio part; Setting restrictions includes: The predicted audio average volume corresponding to the mixed volume weight of any audio data in the mixed strategy is greater than the volume threshold; optionally, the predicted audio average volume is obtained by inputting any audio data and the corresponding mixed volume weight into a volume prediction model; The parameter similarity between the mixing parameters corresponding to any two audio data in the mixing strategy is greater than a second similarity threshold; Based on the dynamic programming algorithm, multiple audio data are iteratively calculated according to the objective function and the limiting conditions to obtain the hybrid strategy corresponding to the multiple audio data.
[0049] In a specific embodiment, Figure 4As shown, the scheme of dynamic programming model and prediction model is implemented based on Python code. It can be seen that through the above optional embodiments, it is possible to perform iterative calculations through the dynamic programming algorithm based on the preset objective function and limiting conditions to obtain the mixing strategy corresponding to the multiple audio data, so as to achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience.
[0050] Embodiment 2 See also Figure 2 , Figure 2 is a schematic diagram of the structure of an audio data mixing system based on a prediction algorithm disclosed in an embodiment of the present invention. Figure 2 The described audio data mixing system based on the prediction algorithm can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 2 As shown, the audio data mixing system based on the prediction algorithm may include: The acquisition module 201 is used to acquire a plurality of audio data to be mixed.
[0051] The extraction module 202 is used to extract the audio feature parameters corresponding to each audio data based on the feature extraction model. The determination module 203 is used to determine the user attention corresponding to each audio data based on the historical user feedback records. The prediction module 204 is used to determine the mixing strategy corresponding to the multiple audio data based on the prediction algorithm according to the audio feature parameters and the user's attention.
[0052] Optionally, the mixing strategy is used to indicate an operation of mixing multiple audio data.
[0053] It can be seen that the above-mentioned embodiments of the invention can extract the audio feature parameters corresponding to each audio data based on the feature extraction model, and then determine the user attention corresponding to each audio data based on the historical user feedback records, so as to comprehensively and accurately predict the mixing strategies corresponding to multiple audio data, thereby being able to achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience.
[0054] As an optional embodiment, the audio feature parameters include one or more of audio duration, audio volume change, audio instrument type and audio text content.
[0055] It can be seen that through the above-mentioned optional embodiments, the content of the audio feature parameters is clarified, which can more accurately and comprehensively characterize the characteristics of the audio data, so as to facilitate subsequent audio mixing predictions, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effects and provide users with a better sound experience.
[0056] As an optional embodiment, the specific method of extracting the audio feature parameters corresponding to each audio data based on the feature extraction model by the extraction module includes: For each audio data, the audio data is input into a trained classifier algorithm to determine an audio type parameter corresponding to the audio data; optionally, the audio type parameter includes one or more of a vocal type, a single instrument type, a mixed instrument type, and a music genre type; the classifier algorithm is trained by a training data set including a plurality of training audio data and corresponding audio type parameter annotations; Determine a target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter; The audio data is input into a target feature extraction model to obtain audio feature parameters corresponding to the audio data.
[0057] It can be seen that through the above-mentioned optional embodiments, it is possible to determine the audio type parameters of the audio data based on the trained classifier algorithm and then screen out the target feature extraction model to extract the audio feature parameters corresponding to the audio data, so as to facilitate subsequent audio mixing predictions, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, thereby giving users a better sound experience.
[0058] As an optional embodiment, the specific manner in which the extraction module determines the target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter includes: For each candidate feature extraction model, calculate the similarity between the audio type annotation and the audio type parameter in the training data set corresponding to the candidate feature extraction model; The candidate feature extraction model with the highest similarity is determined as the target feature extraction model; wherein the candidate feature extraction model is trained by the following steps: Training a preset fully convolutional neural network based on a training data set including a plurality of training audio data and corresponding audio feature parameter annotations until convergence; The parameters of the feature extraction layer in the fully convolutional neural network are fixed to obtain a candidate feature extraction model.
[0059] It can be seen that through the above-mentioned optional embodiments, a suitable feature extraction model can be screened out based on the similarity between the audio type annotations and the audio type parameters in the training data set corresponding to the candidate feature extraction model, so as to facilitate subsequent audio mixing predictions, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effects, thereby giving users a better sound experience.
[0060] As an optional embodiment, the determination module determines the specific manner of user attention corresponding to each audio data based on historical user feedback records, including: Acquire historical user feedback records including feedback records of multiple users on different types of audio data; For each audio data, according to the audio feature parameters corresponding to the audio data, multiple similar audio data are screened out from all evaluated audio data in the historical user feedback records; optionally, the feature similarity between the audio features of the similar audio data and the audio feature parameters is greater than a first similarity threshold; The weighted sum average of the user feedback parameters corresponding to each similar audio data is calculated to obtain the user attention corresponding to the audio data.
[0061] It can be seen that through the above-mentioned optional embodiments, the user attention corresponding to each audio data can be determined based on the user feedback parameters corresponding to the audio data similar to each audio data in the evaluated audio data in the historical user feedback records, so as to facilitate the subsequent audio mixing prediction, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, and thus provide users with a better sound experience.
[0062] As an optional embodiment, the user feedback parameter is the user's attention duration to similar audio data, and the attention duration is calculated by the following steps: When playing mixed audio including similar audio data and other audio data to the user, obtaining a user face image; Based on the face attention recognition model, continuously identify the attention parameters corresponding to the user's face image; First, continuously reduce the proportion of similar audio data in the mixed audio until the attention parameter is lower than the preset first parameter threshold, and calculate the reduction duration; Then continue to increase the proportion of similar audio data in the mixed audio until the attention parameter is higher than a preset second parameter threshold, and calculate the increase duration; optionally, the second parameter threshold is higher than the first parameter threshold; The ratio of the reduced duration to the increased duration is calculated to obtain the attention duration corresponding to the similar audio data.
[0063] It can be seen that through the above optional embodiments, the content and calculation steps of the user feedback parameters are clarified, which can accurately characterize the user's attention duration on the audio data, so as to facilitate subsequent audio mixing prediction, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, thereby giving the user a better sound experience.
[0064] As an optional embodiment, when calculating user attention, the weighted calculation weight corresponding to the user feedback parameter corresponding to each similar audio data is the product of the first weight and the second weight; the first weight is inversely proportional to the audio duration and / or the average audio volume of the similar audio data; the second weight is directly proportional to the feature similarity corresponding to the similar audio data.
[0065] It can be seen that through the above-mentioned optional embodiments, the weight rules corresponding to different similar audio data in the user attention calculation are defined, so that the calculated user attention accurately represents the user's attention duration on the audio data, so as to facilitate subsequent audio mixing prediction, and assist in achieving more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effect, thereby giving the user a better sound experience.
[0066] As an optional embodiment, the prediction module determines the specific manner of the mixing strategy corresponding to the multiple audio data based on the prediction algorithm according to the audio feature parameters and the user's attention, including: The objective function is set to that the user satisfaction corresponding to the hybrid strategy is greater than the satisfaction threshold; optionally, the user satisfaction is obtained by inputting the hybrid parameters corresponding to all audio data in the hybrid strategy and the user attention into the trained user satisfaction prediction model; the hybrid parameters include the hybrid volume weight, the audio mixing mode and the hybrid audio part; Setting restrictions includes: The predicted audio average volume corresponding to the mixed volume weight of any audio data in the mixed strategy is greater than the volume threshold; optionally, the predicted audio average volume is obtained by inputting any audio data and the corresponding mixed volume weight into a volume prediction model; The parameter similarity between the mixing parameters corresponding to any two audio data in the mixing strategy is greater than a second similarity threshold; Based on the dynamic programming algorithm, multiple audio data are iteratively calculated according to the objective function and the limiting conditions to obtain the hybrid strategy corresponding to the multiple audio data.
[0067] It can be seen that through the above-mentioned optional embodiments, it is possible to perform iterative calculations based on preset objective functions and limiting conditions through a dynamic programming algorithm to obtain mixing strategies corresponding to multiple audio data, so as to achieve more efficient and accurate mixing of audio data based on audio features and user attention, so as to effectively improve the sound effects and provide users with a better sound experience.
[0068] Embodiment 3 See also Figure 3 , Figure 3 This is another audio data mixing system based on a prediction algorithm disclosed in an embodiment of the present invention. Figure 3 The described audio data mixing system based on the prediction algorithm is applied in a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 3 As shown, the audio data mixing system based on the prediction algorithm may include: A memory 301 storing executable program codes; a processor 302 coupled to the memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the audio data mixing method based on the prediction algorithm described in the first embodiment.
[0069] Embodiment 4 An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the audio data mixing method based on a prediction algorithm described in the first embodiment.
[0070] Embodiment 5 An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the audio data mixing method based on the prediction algorithm described in the first embodiment.
[0071] The above describes specific embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0072] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0073] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0074] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may be in the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of this specification may be in the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0075] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0076] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0078] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0079] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0080] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0081] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0082] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0083] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0084] Finally, it should be noted that the audio data mixing method and system based on the prediction algorithm disclosed in the embodiment of the present invention only discloses the preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An audio data mixing method based on a prediction algorithm, characterized in that: The method comprises: Acquire multiple audio data to be mixed; Extracting audio feature parameters corresponding to each of the audio data based on a feature extraction model; Determining user attention corresponding to each of the audio data based on historical user feedback records; According to the audio feature parameters and the user attention, based on a prediction algorithm, a mixing strategy corresponding to the multiple audio data is determined; the mixing strategy is used to indicate an operation of mixing the multiple audio data.
2. The audio data mixing method based on the prediction algorithm according to claim 1, characterized in that: The audio feature parameters include one or more of audio duration, audio volume change, audio instrument type, and audio text content.
3. The audio data mixing method based on the prediction algorithm according to claim 1, characterized in that: The extracting audio feature parameters corresponding to each of the audio data based on the feature extraction model includes: For each of the audio data, the audio data is input into a trained classifier algorithm to determine an audio type parameter corresponding to the audio data; the audio type parameter includes one or more of a vocal type, a single instrument type, a mixed instrument type, and a music genre type; the classifier algorithm is trained by a training data set including a plurality of training audio data and corresponding audio type parameter annotations; Based on the audio type parameter, determining a target feature extraction model from a plurality of candidate feature extraction models; The audio data is input into the target feature extraction model to obtain audio feature parameters corresponding to the audio data.
4. The audio data mixing method based on the prediction algorithm according to claim 3 is characterized in that: The step of determining a target feature extraction model from a plurality of candidate feature extraction models based on the audio type parameter comprises: For each candidate feature extraction model, calculating the similarity between the audio type annotation in the training data set corresponding to the candidate feature extraction model and the audio type parameter; The candidate feature extraction model with the highest similarity is determined as the target feature extraction model; wherein the candidate feature extraction model is trained by the following steps: Training a preset fully convolutional neural network based on a training data set including a plurality of training audio data and corresponding audio feature parameter annotations until convergence; The parameters of the feature extraction layer in the fully convolutional neural network are fixed to obtain a candidate feature extraction model.
5. The audio data mixing method based on the prediction algorithm according to claim 1, characterized in that: The determining the user attention corresponding to each audio data based on the historical user feedback record includes: Acquire historical user feedback records including feedback records of multiple users on different types of audio data; For each of the audio data, according to the audio feature parameter corresponding to the audio data, a plurality of similar audio data are screened out from all the evaluated audio data in the historical user feedback record; the feature similarity between the audio feature of the similar audio data and the audio feature parameter is greater than a first similarity threshold; The weighted sum average of the user feedback parameters corresponding to each of the similar audio data is calculated to obtain the user attention corresponding to the audio data.
6. The audio data mixing method based on the prediction algorithm according to claim 5, characterized in that: The user feedback parameter is the user's attention duration to the similar audio data, and the attention duration is calculated by the following steps: When playing mixed audio including the similar audio data and other audio data to the user, obtaining a user face image; Based on the face attention recognition model, continuously identifying the attention parameter corresponding to the user's face image; First, continuously reduce the proportion of the similar audio data in the mixed audio until the attention parameter is lower than a preset first parameter threshold, and calculate the reduction duration; Then, the proportion of the similar audio data in the mixed audio is continuously increased until the attention parameter is higher than a preset second parameter threshold, and the increase duration is calculated; The second parameter threshold is higher than the first parameter threshold; The ratio of the reduced duration to the increased duration is calculated to obtain the attention duration corresponding to the similar audio data.
7. The audio data mixing method based on prediction algorithm according to claim 5, characterized in that: When calculating the user attention, the weighted calculation weight corresponding to the user feedback parameter corresponding to each similar audio data is the product of the first weight and the second weight; the first weight is inversely proportional to the audio duration and / or the average audio volume of the similar audio data; the second weight is directly proportional to the feature similarity corresponding to the similar audio data.
8. The audio data mixing method based on prediction algorithm according to claim 1, characterized in that: The step of determining the mixing strategy corresponding to the plurality of audio data based on the prediction algorithm according to the audio feature parameters and the user attention includes: The objective function is set to a user satisfaction level corresponding to the hybrid strategy that is greater than a satisfaction threshold; the user satisfaction level is obtained by inputting the hybrid parameters corresponding to all the audio data in the hybrid strategy and the user attention into a trained user satisfaction prediction model; the hybrid parameters include a hybrid volume weight, an audio mixing mode, and a hybrid audio part; Setting restrictions includes: The predicted audio average volume corresponding to the mixed volume weight of any audio data in the mixed strategy is greater than the volume threshold; the predicted audio average volume is obtained by inputting any audio data and the corresponding mixed volume weight into a volume prediction model; The parameter similarity between the mixing parameters corresponding to any two of the audio data in the mixing strategy is greater than a second similarity threshold; Based on a dynamic programming algorithm, the multiple audio data are iteratively calculated according to the objective function and the limiting conditions to obtain a mixing strategy corresponding to the multiple audio data.
9. An audio data mixing system based on a prediction algorithm, characterized in that: The system comprises: An acquisition module, used for acquiring a plurality of audio data to be mixed; An extraction module, used for extracting audio feature parameters corresponding to each of the audio data based on a feature extraction model; A determination module, used to determine the user attention corresponding to each of the audio data based on historical user feedback records; A prediction module is used to determine a mixing strategy corresponding to the multiple audio data based on a prediction algorithm according to the audio feature parameters and the user attention; the mixing strategy is used to indicate an operation of mixing the multiple audio data.
10. An audio data mixing system based on a prediction algorithm, characterized in that: The system comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the audio data mixing method based on the prediction algorithm as described in any one of claims 1-8.
Citation Information
Patent Citations
Audio processing method and device
CN107293308A
Audio and video playing method, computer device and computer readable storage medium
CN110213663A
Audio processing method and device, and storage medium
CN114067827A
Music multi-modal data-based user long and short term preference recommendation prediction method
CN114254205A
Deep learning based method and system for processing sound quality characteristics
US20210264938A1
Cited By
Mental health guidance data processing method and system based on AI algorithm
CN120600237A