Audio processing method and device, model training method and device, storage medium and equipment
By weighting the word vectors of audio descriptions and reference descriptions, the problem of imprecise audio descriptions is solved, the accuracy of similarity calculation is improved, and the accuracy of audio retrieval and user experience in accessible services are enhanced.
Patent Information
- Application Number
- CN202510796770.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-03
AI Technical Summary
The audio descriptions generated by existing audio understanding models are not precise enough, resulting in low accuracy of the output results of technical applications and poor user experience.
By weighting the word vectors of the audio description and the reference description, distinguishing the weights of the word vectors, and calculating the similarity between the audio description and the reference description, the accuracy of the similarity calculation is improved.
Improved the accuracy of audio description, improved the accuracy of audio retrieval and user experience in accessibility services.
Smart Images

Figure CN120748437A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of audio processing technology, and in particular to an audio processing method, a model training method and apparatus, a storage medium, and a device. Background Art
[0002] Related technologies can use audio understanding models to understand the content of audio and generate audio descriptions to describe the audio content. The audio descriptions generated by the audio understanding models can be applied in technical applications involving audio processing. For example, in some examples, audio descriptions can be used in audio retrieval to locate audio clips. In other examples, audio descriptions can also be used in accessibility services to deliver audio content in the form of audio descriptions to the service recipients.
[0003] However, if the audio description generated by the audio understanding model is not precise enough, the output results of the technical application will be less accurate and the user experience will be poor. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides an audio processing method, a model training method and apparatus, a storage medium and a device.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio processing method, comprising: obtaining a first vector of an audio description of audio, and a second vector of a reference description of the audio, the first vector and the second vector containing word vectors; performing weighted processing on the first vector to obtain a third vector, and performing weighted processing on the second vector to obtain a fourth vector, the weighted processing being used to distinguish the weights of the word vectors contained in the first vector and the second vector; determining the similarity between the audio description and the reference description based on the third vector and the fourth vector; and measuring the audio description according to the similarity.
[0006] By using weighted processing to differentiate the weights of the word vectors contained in the first vector, a third vector is obtained; by differentiating the weights of the word vectors contained in the second vector, a fourth vector is obtained. This can quantify the importance of the word vectors in the first and second vectors, so that when determining the similarity between the audio description and the reference description based on the third and fourth vectors, the differentiated expression of the sound in the audio description and the reference description can be captured based on the importance of the word vectors, thereby improving the accuracy of the similarity calculation. Furthermore, by measuring the audio description based on the similarity between the audio description and the reference description, the accuracy of the measurement can be improved. Therefore, using accurate audio descriptions in the processing of technical applications can improve the accuracy of the results of technical application processing. For example, in the technical application of audio retrieval, the accuracy of audio segment positioning can be improved based on accurate audio descriptions. In barrier-free services, the audio descriptions delivered to the service recipients are accurate, which can improve the user experience.
[0007] In some exemplary embodiments of the present disclosure, the weighted processing of the first vector to obtain the third vector, and the weighted processing of the second vector to obtain the fourth vector, include: determining a first weight of the word vector contained in the first vector based on the frequency of occurrence of the vocabulary in the audio description and the inverse document frequency of the vocabulary, the vocabulary and the word vector corresponding to each other; weighting the word vector contained in the first vector based on the first weight to obtain the third vector; determining a second weight of the word vector contained in the second vector based on the frequency of occurrence of the vocabulary in the reference description and the inverse document frequency of the vocabulary; and weighting the word vector contained in the second vector based on the second weight to obtain the fourth vector.
[0008] Based on the frequency of occurrence and inverse document frequency of the words in the audio description, a first weight is determined for the word vectors included in the first vector, and the word vectors included in the first vector are weighted by the first weight to obtain a third vector. This allows the importance of the word vectors included in the third vector to be distinguished. Similarly, based on the frequency of occurrence and inverse document frequency of the words in the reference description, a second weight is determined for the word vectors included in the second vector, and the word vectors included in the second vector are weighted by the second weight to obtain a fourth vector. This allows the importance of the word vectors included in the fourth vector to be distinguished. Thus, determining the similarity between the audio description and the reference description based on the third and fourth vectors can improve the accuracy of the similarity calculation.
[0009] In some exemplary embodiments of the present disclosure, the first weight of the word vector contained in the first vector is determined based on the frequency of occurrence of the vocabulary in the audio description and the inverse document frequency of the vocabulary, including: weighting the inverse document frequency of the vocabulary based on the frequency of occurrence of the vocabulary in the audio description in the audio description to obtain the first weight of the word vector corresponding to the vocabulary in the first vector.
[0010] Based on the frequency of occurrence of the vocabulary in the audio description, the inverse document frequency of the vocabulary is weighted to obtain a first weight of the word vector corresponding to the vocabulary in the first vector. The first weight can accurately measure the importance of the word vector in the first vector.
[0011] In some exemplary embodiments of the present disclosure, the weighted processing of the first vector to obtain the third vector, and the weighted processing of the second vector to obtain the fourth vector, include: weighting the word vectors related to the sound features contained in the first vector to obtain the third vector; weighting the word vectors related to the sound features contained in the second vector to obtain the fourth vector.
[0012] The word vectors associated with the sound features can be used to distinguish different sounds. By weighting the word vectors associated with the sound features contained in the first and second vectors, the importance of the word vectors associated with the sound features in the first and second vectors can be highlighted, thereby improving the accuracy of the similarity calculation.
[0013] In some exemplary embodiments of the present disclosure, measuring the audio description according to the similarity includes: measuring the audio description based on the similarity between the audio description and the reference description, and the recall rate of the target model used to generate the audio description.
[0014] Measuring audio descriptions from two dimensions: the similarity between the audio description and the reference description, and the recall rate of the model. This can prevent the results from being dominated by a single dimension and improve the accuracy of the measurement results.
[0015] In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of the first similarity among multiple second similarities; in the ranking, the similarity difference between adjacent rankings is greater than a preset threshold; wherein, the first similarity refers to the similarity between the audio description and a reference description of the audio; the second similarity refers to the similarity between the audio description and other reference descriptions other than the reference description.
[0016] In some exemplary embodiments of the present disclosure, the method for obtaining the ranking includes: for the first similarity and any two similarities among the multiple second similarities, based on the difference between one of the similarities and the other similarity being greater than the preset threshold, determining that the ranking of one of the similarities is higher than the ranking of the other similarity; or based on the one of the similarities being greater than the other similarity and the difference between the one of the similarities and the other similarity being less than or equal to the preset threshold, determining that the ranking of one of the similarities is equal to the ranking of the other similarity.
[0017] In some exemplary embodiments of the present disclosure, the method for obtaining the ranking includes: clustering the first similarity and the multiple second similarities according to the distance between the similarities to obtain one or more clusters; sorting the one or more clusters based on the mean of the elements in the clusters; wherein the ranking of the elements in the same cluster is equal to the ranking of the cluster.
[0018] By clustering the first similarity and multiple second similarities, and then sorting the clusters according to the mean of the elements in the clusters, the sorting of the clusters is used as the sorting of the elements in the clusters, and ensuring that the difference in the mean of the elements between two adjacent clusters is greater than a preset threshold, the impact of disturbances in the similarity calculation on the sorting results can be reduced, thereby improving the accuracy of the sorting.
[0019] In some exemplary embodiments of the present disclosure, the sorting of the one or more clusters based on the means of the elements in the clusters includes: merging clusters whose means differ by less than a preset threshold based on the means of the elements in the clusters.
[0020] Based on the mean of the elements in the clusters, clusters whose mean differences are less than a preset threshold are merged, which can improve the distinction between different clusters and the accuracy of cluster sorting.
[0021] According to a second aspect of an embodiment of the present disclosure, an audio processing method is provided, comprising obtaining audio and processing the audio based on a target model to obtain an audio description of the audio; wherein the target model is determined based on the measurement results of any embodiment of the above-mentioned first aspect.
[0022] Because the target model is determined based on the measurement results of the first aspect above, generating an audio description based on the target model can ensure the accuracy of the audio description. When the audio description is used as the basis for generating the final output result, the accuracy of the final output result can be improved. For example, in the technical application of audio retrieval, accurate audio description can improve the accuracy of audio segment positioning. In barrier-free services, it can ensure that the audio description delivered to the service recipient is accurate, improving the user experience.
[0023] According to a third aspect of an embodiment of the present disclosure, a model training method is provided, including: processing audio based on a target model to generate an audio description of the audio; encoding the audio description and a reference description of the audio to obtain a first vector of the audio description and a second vector of the reference description, the first vector and the second vector containing word vectors; performing weighted processing on the first vector to obtain a third vector, and performing weighted processing on the second vector to obtain a fourth vector, the weighted processing being used to distinguish the weights of the word vectors contained in the first vector and the weights of the word vectors contained in the second vector; determining the similarity between the audio description and the reference description based on the third vector and the fourth vector; and training the target model based on the similarity.
[0024] By using a weighted process to differentiate the weights of the word vectors contained in the first vector, a third vector is obtained; and by using a weighted process to differentiate the weights of the word vectors contained in the second vector, a fourth vector is obtained. This quantifies the importance of the word vectors in the first and second vectors. Thus, when determining the similarity between the audio description and the reference description based on the third and fourth vectors, the importance of the word vectors can be used to capture the differential representation of the sound in the audio description and the reference description, thereby improving the accuracy of the similarity calculation between the audio description and the reference description. Furthermore, the similarity between the audio description and the reference description can be used to measure the accuracy of the audio description generated by the target model. The higher the accuracy of the audio description generated by the target model, the stronger the target model's ability to generate audio descriptions. By feeding the similarity between the audio description and the reference description into the training of the target model, the target model's ability to generate audio descriptions can be improved. Furthermore, when the target model is used to perform tasks related to audio processing, the accuracy of the task execution is improved. For example, in the technical application of audio retrieval, accurate audio descriptions can improve the accuracy of audio segment location. In accessibility services, accurate audio descriptions delivered to service recipients can improve the user experience.
[0025] In some exemplary embodiments of the present disclosure, the training of the target model based on the similarity includes: measuring the audio description based on the similarity and the recall rate of the target model; and training the target model based on the measurement result of the audio description.
[0026] Measuring the audio descriptions generated by the target model based on two dimensions: the similarity between the audio description and the reference description, and the recall rate of the target model. This prevents a single dimension from dominating the measurement results and improves their accuracy. Furthermore, training the target model based on these measurement results can improve its ability to generate audio descriptions.
[0027] In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of the first similarity among multiple second similarities; in the ranking, the similarity difference between adjacent rankings is greater than a preset threshold; wherein, the first similarity refers to the similarity between the audio description and a reference description of the audio; the second similarity refers to the similarity between the audio description and other reference descriptions other than the reference description.
[0028] In the sorting of the first similarity and the plurality of second similarities, the difference in similarity between adjacent sorts is greater than a preset threshold, thereby reducing the impact of disturbances in the similarity calculation on the sorting result and improving the differentiation of similarities between the audio description and different reference descriptions. By improving the differentiation of similarities between the audio description and different reference descriptions.
[0029] According to a fourth aspect of an embodiment of the present disclosure, there is provided an audio processing device, including:
[0030] An acquisition module is configured to obtain a first vector of an audio description of an audio and a second vector of a reference description of the audio, wherein the first vector and the second vector include word vectors.
[0031] A weighting module is used to perform weighted processing on the first vector to obtain a third vector, and to perform weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is used to distinguish the weights of the word vectors included in the first vector and the second vector.
[0032] The first determining module is configured to determine a similarity between the audio description and the reference description based on the third vector and the fourth vector.
[0033] The second determining module is configured to measure the audio description according to the similarity.
[0034] In some exemplary embodiments of the present disclosure, the weighting module includes:
[0035] A first determining unit is configured to determine a first weight of a word vector included in the first vector based on an occurrence frequency of a word in the audio description and an inverse document frequency of the word, wherein the word corresponds to the word vector.
[0036] The first processing unit is configured to weight the word vectors included in the first vector based on the first weight to obtain a third vector.
[0037] The first determining unit is further configured to determine a second weight of a word vector included in the second vector based on an occurrence frequency of a word in the reference description and an inverse document frequency of the word.
[0038] The first processing unit is further configured to weight the word vectors included in the second vector based on the second weight to obtain a fourth vector.
[0039] In some exemplary embodiments of the present disclosure, the first determination unit is configured to weight the inverse document frequency of the vocabulary in the audio description based on the frequency of occurrence of the vocabulary in the audio description, to obtain a first weight of the word vector corresponding to the vocabulary in the first vector.
[0040] In some exemplary embodiments of the present disclosure, the weighting module includes:
[0041] The second processing unit is configured to weight the word vectors related to the sound features contained in the first vector to obtain a third vector.
[0042] The third processing unit is configured to weight the word vectors related to the sound features contained in the second vector to obtain a fourth vector.
[0043] In some exemplary embodiments of the present disclosure, the second determination module is configured to measure the audio description based on a similarity between the audio description and the reference description, and a recall rate of a target model used to generate the audio description.
[0044] In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of the first similarity among multiple second similarities; in the ranking, the similarity difference between adjacent rankings is greater than a preset threshold; wherein, the first similarity refers to the similarity between the audio description and a reference description of the audio; the second similarity refers to the similarity between the audio description and other reference descriptions other than the reference description.
[0045] In some exemplary embodiments of the present disclosure, the device further includes a third determination module, configured to determine, for the first similarity and any two similarities among the multiple second similarities, that the ranking of one of the similarities is higher than the ranking of the other similarity based on the difference between one of the similarities and the other similarity being greater than the preset threshold; or to determine that the ranking of one of the similarities is equal to the ranking of the other similarity based on the one of the similarities being greater than the other similarity and the difference between the one of the similarities and the other similarity being less than or equal to the preset threshold.
[0046] In some exemplary embodiments of the present disclosure, the apparatus further includes:
[0047] The clustering module is configured to cluster the first similarity and the plurality of second similarities according to distances between the similarities to obtain one or more clusters.
[0048] A sorting module is configured to sort the one or more clusters based on the means of the elements in the clusters.
[0049] The order of elements in the same cluster is equal to the order of the cluster.
[0050] In some exemplary embodiments of the present disclosure, the sorting module is further configured to: merge, based on the means of the elements in the clusters, the clusters whose means differ from each other by less than the preset threshold.
[0051] According to a fifth aspect of the embodiments of the present disclosure, there is provided an audio processing device, including:
[0052] Acquisition module, used to obtain audio.
[0053] The processing module is used to process the audio based on the target model to obtain an audio description of the audio.
[0054] The target model is determined based on the measurement results of the first aspect mentioned above.
[0055] According to a sixth aspect of an embodiment of the present disclosure, there is provided a model training device, comprising:
[0056] The generation module is used to process the audio based on the target model to generate an audio description of the audio.
[0057] An encoding module is used to encode the audio description and the reference description of the audio to obtain a first vector of the audio description and a second vector of the reference description, wherein the first vector and the second vector include word vectors.
[0058] A weighting module is used to perform weighted processing on the first vector to obtain a third vector, and to perform weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is used to distinguish the weights of the word vectors included in the first vector and the second vector.
[0059] A determination module is configured to determine a similarity between the audio description and the reference description based on the third vector and the fourth vector.
[0060] A training module is used to train the target model based on the similarity.
[0061] In some exemplary embodiments of the present disclosure, the training module includes:
[0062] A determining unit is configured to measure the audio description based on the similarity and a recall rate of the target model.
[0063] A training unit is used to train the target model based on the measurement result of the audio description.
[0064] In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of the first similarity among multiple second similarities; in the ranking, the similarity difference between adjacent rankings is greater than a preset threshold; wherein, the first similarity refers to the similarity between the audio description and a reference description of the audio; the second similarity refers to the similarity between the audio description and other reference descriptions other than the reference description.
[0065] According to the seventh aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: implement the method of any one of the embodiments in the first to third aspects above.
[0066] According to the eighth aspect of the embodiments of the present disclosure, a non-temporary computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to execute the method of any one of the embodiments in the first to third aspects above.
[0067] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0069] Figure 1 The present invention is a flowchart illustrating an audio processing method according to some embodiments of the present disclosure.
[0070] Figure 2 is a schematic diagram illustrating a vector encoding method according to some embodiments of the present disclosure.
[0071] Figure 3 is a schematic diagram illustrating another vector encoding method according to some embodiments of the present disclosure.
[0072] Figure 4 is a schematic diagram illustrating a weighting method according to some embodiments of the present disclosure.
[0073] Figure 5 is a schematic diagram illustrating another weighting method according to some embodiments of the present disclosure.
[0074] Figure 6 is a schematic diagram illustrating yet another weighting method according to some embodiments of the present disclosure.
[0075] Figure 7 is a schematic diagram illustrating a sorting method according to some embodiments of the present disclosure.
[0076] Figure 8 is a schematic diagram illustrating another sorting method according to some embodiments of the present disclosure.
[0077] Figure 9 is a schematic diagram illustrating a clustering method according to some embodiments of the present disclosure.
[0078] Figure 10 is a flowchart illustrating another audio processing method according to some embodiments of the present disclosure.
[0079] Figure 11 The present invention is a flowchart illustrating an audio retrieval method according to some embodiments of the present disclosure.
[0080] Figure 12 This is a flowchart of a model training method according to some embodiments of the present disclosure.
[0081] Figure 13 is a schematic diagram illustrating a model training process according to some embodiments of the present disclosure.
[0082] Figure 14 is a block diagram of an audio processing device according to some embodiments of the present disclosure.
[0083] Figure 15 is a block diagram of another audio processing device according to some embodiments of the present disclosure.
[0084] Figure 16 It is a block diagram of a model training device according to some embodiments of the present disclosure.
[0085] Figure 17 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0086] Some embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, for the sake of clarity and brevity, descriptions of features known in the art may be omitted.
[0087] In some exemplary embodiments, the audio can be processed by an audio understanding model to generate an audio description for describing the audio content. The accuracy of the audio description of the audio content (hereinafter referred to as the audio description) can be determined by the similarity between the audio description and a reference description of the audio (which can be exemplarily understood as a standard description of the audio content, which can be obtained by manual annotation or other methods). However, the related art has the following defects when evaluating the similarity between the audio description and the reference description:
[0088] 1. Insufficient Semantic Capture
[0089] Related technologies are generally based on N-gram technology (an algorithm based on statistical language models) to match vocabulary between audio descriptions and reference descriptions. However, even if this matching method expresses the same semantics, as long as there are slight changes in vocabulary (such as synonyms or sentence structure adjustments), it will be mistakenly judged as words with different semantics. METEOR (an indicator used to evaluate the semantic similarity and alignment between machine-generated text and reference text) has improved the recognition of simple vocabulary changes (such as stems and basic synonyms), but it is still unable to recognize more complex equivalent semantics (especially context-related semantics).
[0090] 2. Cross-domain migration failure
[0091] Metrics used to evaluate image descriptions (i.e., descriptions of image content) are difficult to effectively transfer to audio descriptions. The objects, attributes, and spatial relationships described in images correspond to more abstract, temporally specific events and complex acoustic relationships in audio. Metrics used to evaluate image descriptions cannot account for these differences.
[0092] 3. The accuracy of similarity assessment is low
[0093] Based on a pre-trained language model, a method for determining the similarity between an audio description and a reference description assigns the same weight to common words and key words (which can be exemplarily understood as words related to sound features, such as happy, sad, etc., but not limited to the words listed here) when calculating the similarity. As a result, key words are ignored in the similarity calculation, resulting in low accuracy of the similarity calculation.
[0094] 4. Lack of recognition / inadequate differentiation
[0095] Related technologies generally focus on comparing audio descriptions of an audio file with reference descriptions of the same audio file, but lack the ability to compare the audio description of an audio file with reference descriptions of other audio files. This results in an inability to distinguish between different audio files. This can lead to ambiguous audio descriptions, making them incapable of meeting the requirements of technical applications such as audio retrieval.
[0096] In response to related technologies, an embodiment of the present disclosure provides an audio processing solution. The technical solution of the embodiment of the present disclosure is described below in conjunction with exemplary embodiments.
[0097] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0098] Figure 1 This is a flow chart of an audio processing method according to some embodiments of the present disclosure, which is used in electronic devices, such as desktop computers, portable computers, mobile phones, distributed processing nodes, servers, etc., but is not limited to the devices listed here. Figure 1 As shown, in some embodiments of the present disclosure, the audio processing method may include the following.
[0099] In step S101 , a first vector of an audio description of audio and a second vector of a reference description of audio are obtained.
[0100] In some exemplary embodiments, the audio description referred to in the embodiments of the present disclosure can be illustratively understood as one or more paragraphs of text used to describe the audio content. For example, in some examples, the audio description can be, for example, "impact sound." Of course, this is merely an example and not the only limitation on audio description.
[0101] In some exemplary embodiments, the audio description referred to in the embodiments of the present disclosure can be generated, for example, by an audio understanding model (hereinafter referred to as a target model). The target model is trained to analyze and process audio and generate an audio description that describes the audio content. The accuracy of the audio description generated by the target model is related to the capabilities of the target model.
[0102] In some cases, the audio description generated by the target model is ambiguous and less accurate. For example, if the sound in the audio is a car crash, the audio description generated by the target model may be "crash." In other cases, the audio description generated by the target model can accurately describe the audio. For example, in the example above, if the sound in the audio is a car crash, the audio description generated by the target model is also "car crash."
[0103] In some exemplary embodiments, the reference description of the audio can be illustratively understood as a standard description of the audio. The reference description can also be illustratively understood as a textual description. For example, if the sound in the audio is a car crash, the reference description can be "car crash sound." Of course, this is merely an example of a reference description and not a limitation.
[0104] In some exemplary embodiments, the reference description of the audio may be obtained through manual annotation, but is not limited to manual annotation.
[0105] In some exemplary embodiments, the audio description generated by the target model and the reference description of the audio can be vector-encoded using a preset language model to obtain a first vector corresponding to the audio description and a second vector of the reference description, wherein the first vector and the second vector contain word vectors.
[0106] It should be noted that, in general, word vectors in the same vector have the same weight.
[0107] for example, Figure 2 is a schematic diagram illustrating a vector encoding method according to some embodiments of the present disclosure. Figure 3 FIG. 1 is a schematic diagram of another vector encoding method according to some embodiments of the present disclosure. Figure 2, assuming that the audio description generated by the target model is "there are two kinds of roaring sounds". In some exemplary vector encoding methods, the audio description can be encoded into a first vector by Sentence-BERT (an improved BERT model). The first vector includes three word vectors, namely, a word vector A11 corresponding to the word "there are", a word vector A12 corresponding to the word "two kinds", and a word vector A13 corresponding to the word "roaring sound". Among them, in the first vector, the weights of the word vector A11, the word vector A12, and the word vector A13 are the same. The weights of the word vector A11, the word vector A12, and the word vector A13 are the same, indicating that the word vector A11, the word vector A12, and the word vector A13 are equally important for the description of the sound.
[0108] See also Figure 3 , assuming that the reference description of the audio is "there are vehicle roars and airplane roars". In some exemplary vector encoding methods, the reference description can be encoded into a second vector by the Sentence-BERT model. The second vector includes six word vectors, namely, a word vector A21 corresponding to the word "there are", a word vector A22 corresponding to the word "vehicle", a word vector A23 corresponding to the word "roar", a word vector A24 corresponding to the word "and", a word vector A25 corresponding to the word "airplane", and a word vector A26 corresponding to the word "roar". Among them, in the first vector, the weights of word vector A21, word vector A22, word vector A23, word vector A24, word vector A25, and word vector A26 are the same.
[0109] In step S103 , weighted processing is performed on the first vector to obtain a third vector, and weighted processing is performed on the second vector to obtain a fourth vector.
[0110] The weighting of the first vector and the second vector can be understood as weighting the word vectors in the first vector and the second vector. For the first vector or the second vector, the weighting process can distinguish the weights of the word vectors contained in the vector. In some embodiments, the weight of the word vector is related to the importance of the word vector. The higher the importance of the word vector, the greater the weight, and the lower the importance of the word vector, the smaller the weight.
[0111] In the embodiment of the present disclosure, the importance of the word vector related to the sound feature is higher than the weight of the word vector that is not related to the sound feature. For example, if the audio description is "there are two kinds of roaring sounds", the weight of the word vector corresponding to "roaring sound" is higher than the weight of the word vector corresponding to "two kinds" and "there are". Among them, the weight of the word vector of "two kinds" and the weight of the word vector of "there are" can be the same or different. For another example, if the reference description is "there are roaring sounds of vehicles and roaring sounds of airplanes", the weight of the word vector corresponding to "roaring sound", the weight of the word vector corresponding to "vehicles", and the weight of the word vector corresponding to "airplanes" are higher than the weight of the word vector corresponding to "there are", and the weight of the word vector corresponding to "and". Among them, the weight of the word vector corresponding to "roaring sound", the weight of the word vector corresponding to "vehicles", and the weight of the word vector corresponding to "airplanes" can be the same or different.
[0112] In some exemplary embodiments, a dictionary may be established that stores vocabulary related to sound features. When weighting the first vector and the second vector, vocabulary related to the sound features contained in the audio description and the reference description may be determined based on the dictionary. Then, based on the correspondence between these vocabulary words and the word vectors in the first vector and the second vector, high weights may be assigned to the word vectors corresponding to these vocabulary words, while low weights may be assigned to the word vectors of other vocabulary words, thereby achieving the purpose of distinguishing the weights of different word vectors.
[0113] for example, Figure 4 FIG. 1 is a schematic diagram illustrating a weighting method according to some embodiments of the present disclosure. Figure 4 As shown in the figure, assume the audio description is "There are two kinds of roaring sounds." According to the dictionary, "roaring sound" is a word related to sound features. The word vector corresponding to "roaring sound" in vector a is A13, the word vector corresponding to "there is" in vector a is A11, and the word vector corresponding to "two kinds" in vector a is A12. Therefore, weight Q1 can be applied to A13, and weight Q2 can be applied to A11 and A12, where Q1 is greater than Q2.
[0114] Of course, the above is merely an example of the weighting method involved in the embodiment of the present disclosure, and is not the only limitation.
[0115] By differentiating the weights of word vectors, word vectors related to sound features are given larger weights, which can improve sensitivity to key features of sound events.
[0116] In step S105 , the similarity between the audio description and the reference description is determined based on the third vector and the fourth vector.
[0117] In some embodiments, the similarity between the third vector and the fourth vector can be calculated to determine the similarity between the third vector and the fourth vector as the similarity between the audio description and the reference description. For example, in some examples, the cosine similarity between the third vector and the fourth vector can be calculated based on the cosine similarity calculation method, and the cosine similarity between the third vector and the fourth vector can be determined as the similarity between the audio description and the reference description. The expression is as follows:
[0118]
[0119] Among them, SentBERT represents the cosine similarity between the third vector and the fourth vector, E Ci_weighted represents the third vector, E Ri_weighted Represents the fourth vector.
[0120] Of course, this is only an exemplary similarity calculation method, not the only method. In fact, any similarity calculation method in the related art can be applied in the embodiments of the present disclosure.
[0121] Since the word vectors in the third vector and the fourth vector have been distinguished in importance through weights, determining the similarity between the audio description and the reference description based on the third vector and the fourth vector can improve the accuracy of the similarity calculation.
[0122] In step S107 , the audio description is measured according to the similarity between the audio description and the reference description.
[0123] In some embodiments, when the similarity between the audio description and the reference description is greater than a preset similarity, the audio description is accurate; when the similarity between the audio description and the reference description is less than or equal to the preset similarity, the audio description is inaccurate.
[0124] The above-mentioned embodiment of the present disclosure distinguishes the weights of the word vectors contained in the first vector by weighted processing to obtain a third vector; distinguishes the weights of the word vectors contained in the second vector to obtain a fourth vector, which can quantify the importance of the word vectors in the first vector and the second vector, so that when determining the similarity between the audio description and the reference description based on the third vector and the fourth vector, the differentiated expression of the sound in the audio description and the reference description can be captured according to the importance of the word vector, thereby improving the accuracy of the similarity calculation. Further, the audio description is measured according to the similarity between the audio description and the reference description, which can improve the accuracy of the measurement, so that the use of accurate audio descriptions in the processing of technical applications can improve the accuracy of the results of technical application processing. For example, in the technical application of audio retrieval, the accuracy of audio segment positioning can be improved based on accurate audio descriptions. In barrier-free services, the audio description delivered to the service recipient is accurate, which can improve the user experience.
[0125] Figure 5 FIG. 1 is a schematic diagram illustrating another weighting method according to some embodiments of the present disclosure. Figure 5 As shown, in some exemplary implementations, the weighting method provided by the embodiments of the present disclosure may include the following steps.
[0126] In step S501 , a first weight of a word vector included in a first vector is determined based on the frequency of occurrence of a word in the audio description and the inverse document frequency of the word.
[0127] The word vectors in the first vector correspond one-to-one to the vocabulary in the audio description.
[0128] The TF-IDF (term frequency-inverse document frequency) technology provided by the related art shows that the higher the frequency of occurrence of a word in a text, the lower the inverse document frequency in other texts, and the stronger the ability of the word to distinguish categories. Similarly, in an embodiment of the present disclosure, the higher the frequency of occurrence of a word in the audio description, the lower the inverse document frequency in other texts, and the stronger the ability of the word to distinguish different sounds, and accordingly, the higher the importance of the word in the audio description. Therefore, in some embodiments, the first weight of the word vector included in the first vector can be determined based on the frequency of occurrence of the words in the audio description in the audio description and the inverse document frequency of the words in the audio description.
[0129] For example, in some embodiments, the audio description and the inverse document frequency of each word in the audio description can be input into a preset weight recognition model. The weight recognition model is used to encode the audio description into a first vector and count the frequency of occurrence of the words contained in the audio description in the audio description. Based on the frequency of occurrence of the words contained in the audio description in the audio description and the inverse document frequency of the words, the first weight of the word vector contained in the first vector is determined. Among them, the weight recognition model can be obtained by training based on the model training method provided by the relevant technology, and the training sample can include the text for describing the audio and the inverse document frequency of the words contained in the text. The reference data for training can include the vector corresponding to the text and the weight of the word vector contained in the vector.
[0130] For another example, in some other implementations, the inverse document frequency of a word in the audio description may be weighted based on the frequency of occurrence of the word in the audio description to obtain a first weight of the word vector corresponding to the word in the first vector.
[0131] For example, assuming there are d words in the audio description, the frequency of occurrence of the w-th word in the audio description is TF(w,d), and the inverse document frequency of the w-th word is IDF(w). Then the first weight of the word vector of the w-th word in the first vector can be expressed as:
[0132] TFIDF(w)=TF(w,d)×IDF(w) (2)
[0133] Based on the frequency of occurrence of the vocabulary in the audio description, the inverse document frequency of the vocabulary is weighted to obtain a first weight of the word vector corresponding to the vocabulary in the first vector. The first weight can accurately measure the importance of the word vector in the first vector.
[0134] In step S503 , the word vectors in the first vector are weighted based on the first weights of the word vectors included in the first vector to obtain a third vector.
[0135] Continuing with the above example, assuming that the word vector corresponding to the w-th word in the first vector is Embedding(w), then the weighting of Embedding(w) can be expressed as:
[0136]
[0137] Among them, E W Represents the word vector obtained after weighting.
[0138] It should be noted that although the above method for weighting the first vector is described using the first vector as an example, in other feasible implementations, the above method can also be used to weight the second vector to obtain a fourth vector. That is, in some implementations, the second weights of the word vectors included in the second vector can be determined based on the frequency of occurrence of the words in the reference description and the inverse document frequency of the words in the reference description; based on the second weights of the word vectors included in the second vector, the word vectors included in the second vector are weighted to obtain the fourth vector.
[0139] In the above embodiment of the present disclosure, based on the frequency of occurrence and inverse document frequency of the vocabulary in the audio description, a first weight of the word vector contained in the first vector is determined, and the word vector contained in the first vector is weighted by the first weight to obtain a third vector, which can distinguish the importance of the word vector contained in the third vector. Similarly, based on the frequency of occurrence and inverse document frequency of the vocabulary in the reference description, a second weight of the word vector contained in the second vector is determined, and the word vector contained in the second vector is weighted by the second weight to obtain a fourth vector, which can distinguish the importance of the word vector contained in the fourth vector. Therefore, based on the third vector and the fourth vector, the similarity between the audio description and the reference description is determined, which can improve the accuracy of the similarity calculation.
[0140] Figure 6 FIG. 1 is a schematic diagram of another weighting method according to some embodiments of the present disclosure. Figure 6As shown, in some exemplary implementations, the weighting method provided by the embodiments of the present disclosure may include the following steps.
[0141] In step S601, the word vectors related to the sound feature contained in the first vector are weighted to obtain a third vector.
[0142] In step S603, the word vectors related to the sound feature contained in the second vector are weighted to obtain a fourth vector.
[0143] In some exemplary embodiments, the first vector can be input into a preset model, and the word vector related to the sound feature contained in the first vector can be output through the preset model. The preset model is trained to identify the word vector related to the sound feature contained in the vector. The training data of the preset model includes a vector, which includes word vectors related to the sound feature. The word vectors related to the sound feature contained in the vector are annotated with preset labels. The training method of the preset model can be found in the relevant technology, and will not be repeated in the embodiments of the present disclosure.
[0144] It should be noted that, in some implementations, the word vectors related to the sound features contained in the second vector can also be obtained by the above method.
[0145] In some embodiments, the weights used to weight the word vectors related to the sound features contained in the first vector and the weights used to weight the word vectors related to the sound features contained in the second vector may be the same or different. The weights may be illustratively understood as set values. Alternatively, in other embodiments, they may be randomly generated values.
[0146] In the embodiment of the present disclosure, the method of weighting the word vectors related to the sound features can be referred to the aforementioned embodiment and will not be repeated here.
[0147] In the above-mentioned embodiment of the present disclosure, by weighting the word vectors related to the sound features contained in the first vector and the second vector, the importance of the word vectors related to the sound features in the first vector and the second vector can be highlighted, thereby improving the accuracy of the similarity calculation.
[0148] For example, in some embodiments of the present disclosure, measuring the audio description according to the similarity between the audio description and the reference description may include: measuring the audio description based on the similarity between the audio description and the reference description, and the recall rate of the target model used to generate the audio description.
[0149] For example, assuming that the similarity between the audio description and the reference description is SentBERT and the recall rate of the target model is Recall, the accuracy of the audio description can be evaluated by the following expression.
[0150]
[0151] Measuring audio descriptions from two dimensions: the similarity between the audio description and the reference description, and the recall rate of the model. This can prevent the results from being dominated by a single dimension and improve the accuracy of the measurement results.
[0152] In some implementations of the disclosed embodiments, the recall rate of the target model is related to the ranking of a first similarity among multiple second similarities, and the difference in similarity between adjacent rankings of the first similarity and the multiple second similarities is greater than a preset threshold. The first similarity refers to the similarity between the audio description and a reference description of the audio; the second similarity refers to the similarity between the audio description and a reference description other than the reference description.
[0153] For example, suppose the vector of audio description generated by the target model is C i , C i Used to describe audio i. There are j preset reference descriptions (j is an integer greater than 1), where the i-th reference description is the reference description corresponding to audio i. The following confusion matrix can be calculated:
[0154] M[i,h]=cos(C i ,R h ) (5)
[0155] Among them, R h is the vector describing the hth reference, and the value of h ranges from 1 to j. i and R j It can be exemplarily understood as a vector obtained by weighting using the weighting method in the above embodiment, or it can also be a vector obtained without being weighted using the weighting method in the above embodiment.
[0156] Assuming that the following first relationship exists when sorting similarity, it is determined that M[i,h] is ranked higher than M[i,a], the value range of a is 1 to j, and a is not equal to h.
[0157] M[i,h]>M[i,a]+τ (6)
[0158] Among them, τ is the preset threshold.
[0159] Assuming the following second relationship holds when sorting by similarity, it is determined that the sorting of M[i,h] is equal to M[i,a].
[0160] M[i,a]<M[i,h]<M[i,a]+τ (7)
[0161] For example, Figure 7 is a schematic diagram of a sorting method according to some embodiments of the present disclosure. Figure 7 In , reference description 1 and audio description 1 are descriptions of audio 1, reference description 2 is a description of audio 2, and reference description 3 is a description of audio 2. Audio 1, audio 2, and audio 3 are different audios. The similarity between the vector of reference description 1 and the vector of audio description 1 is S1, the similarity between the vector of reference description 2 and the vector of audio description 1 is S2, and the similarity between the vector of reference description 3 and the vector of audio description 1 is S3. Among them, S1 is greater than S2+τ; S2 is greater than S3 and less than S3+τ, then the ranking of S1 is 1, the ranking of S2 is 2, and the ranking of S3 is 2. The recall rate of the target model can be expressed as:
[0162]
[0163] Here, rank(S1) in expression (8) represents the ranking of S1, and "3" is the number of reference descriptions.
[0164] In the sorting of the first similarity and multiple second similarities, the similarity difference between adjacent sortings is greater than a preset threshold, which can reduce the impact of disturbances in the similarity calculation on the sorting results and improve the differentiation of the similarities between the audio description and different reference descriptions.
[0165] For example, Figure 8 FIG. 1 is a schematic diagram illustrating another sorting method according to some embodiments of the present disclosure. Figure 8 As shown, in some exemplary implementations, the sorting method provided by the embodiments of the present disclosure may include the following steps.
[0166] In step S801 , the first similarity and the plurality of second similarities are clustered according to the distance between the similarities to obtain one or more clusters.
[0167] The first similarity refers to the similarity between the audio description and the reference description of the audio; the second similarity refers to the similarity between the audio description and other reference descriptions other than the reference description.
[0168] In some implementations, the absolute value of the difference between any two similarities may be calculated, and similarities whose absolute value of the difference between the similarities is less than a preset value may be grouped into one category.
[0169] For example, Figure 9 is a schematic diagram of a clustering method according to some embodiments of the present disclosure, such as Figure 9As shown in the figure, assuming that the first similarity a1 has a similarity value of 70%, and there are three second similarities, namely b1, b2, and b3. Among them, the similarity value of b1 is 70.1%, the similarity value of b2 is 71%, and the similarity value of b3 is 73%. When the preset value is 0.2%, a1 and b1 can be aggregated into one cluster, b2 can be aggregated into one cluster alone, and b3 can be aggregated into one cluster alone, resulting in three clusters.
[0170] In other embodiments, the Euclidean distance between similarities can be calculated to cluster similarities whose Euclidean distance is less than a preset distance into one category. The method for calculating the Euclidean distance between similarities can be found in the Euclidean distance calculation method provided in the related art and will not be described in detail here.
[0171] It should be noted that the above two clustering methods are merely examples and are not the only limitations on the clustering methods involved in the embodiments of the present disclosure.
[0172] In step S803 , one or more clusters are sorted based on the mean values of the elements in the clusters.
[0173] The order of elements in the same cluster is equal to the order of the cluster.
[0174] by Figure 9 Take the clustering result shown in as an example. Assume that the cluster where a1 and b1 are located is cluster 1, the cluster where b2 is located is cluster 2, and the cluster where b3 is located is cluster 3. By calculating the mean of a1 and b1, we can get the mean of the elements in cluster 1 to be 70.05%. By calculating the mean of b2 (in Figure 9 In the example shown, the similarity value of b2 itself) is obtained that the mean of the elements in cluster 2 is 71%. By calculating the mean of b3 (in Figure 9 In the example shown, the mean of the elements in cluster 3 is 73%, which is the similarity value of b3 itself. Assuming the preset threshold for sorting is 0.5%, the ranking results based on the mean of the elements in each cluster, in descending order of mean, are: Cluster 3, Cluster 2, Cluster 1.
[0175] Among them, in cluster 3, the order of elements a1 and b1 is the same, which is the order of cluster 1 in the three clusters.
[0176] It should be noted that if the preset threshold is 1% in the above example, and the difference (0.95%) between the mean of the elements in cluster 1 (70.05%) and the mean of the elements in cluster 2 (71%) is less than the preset threshold (1%), then cluster 1 and cluster 2 need to be merged. In other words, in some embodiments, clusters whose mean differences are less than the preset threshold can be merged based on the means of the elements in the clusters.
[0177] By clustering the first similarity and multiple second similarities, and then sorting the clusters according to the mean of the elements in the clusters, the sorting of the clusters is used as the sorting of the elements in the clusters, and it is ensured that between two adjacent clusters, the mean of the elements in the higher-ranked cluster is greater than the mean of the elements in the lower-ranked cluster plus a preset threshold. This can reduce the impact of disturbances in the similarity calculation on the sorting results and improve the accuracy of the sorting.
[0178] Figure 10 FIG. 1 is a flow chart of another audio processing method according to some embodiments of the present disclosure. Figure 10 As shown, in some implementations, the audio processing method provided by the embodiments of the present disclosure may include the following steps.
[0179] In step S1001, audio is obtained.
[0180] In step S1003, the audio is processed based on the target model to obtain an audio description of the audio.
[0181] Among them, the target model can be based on the above Figures 1-9 The measurement result in any one of the embodiments is determined.
[0182] Take its application in audio retrieval as an example. Figure 11 FIG. 1 is a flow chart of an audio retrieval method according to some embodiments of the present disclosure. Figure 11 As shown, in some embodiments, the above Figures 1-9 The method provided in any embodiment evaluates the audio description generated by the target model. If the evaluation passes, the target model can be used in audio retrieval. First, the user enters the search criteria into the audio retrieval application. The audio retrieval application sends the audio clip to the target model, which generates an audio description of the audio clip. The target model then matches the audio description of the audio clip with the search criteria to obtain audio clips that meet the search criteria. The audio clips are then fed back to the user.
[0183] It should be understood that the above is only an exemplary application of the embodiment of the present disclosure, and not the only application.
[0184] Since the target model is based on the above Figures 1-9 The audio description is generated based on the target model, which is determined by the measurement results of any embodiment, and the accuracy of the audio description can be ensured. When the audio description is used as the basis for generating the final output result, the accuracy of the final output result can be improved. For example, in the technical application of audio retrieval, the accuracy of audio segment positioning can be improved based on accurate audio description. In barrier-free services, it can ensure that the audio description delivered to the service recipient is accurate, thereby improving the user experience.
[0185] Figure 12 This is a flow chart of a model training method according to some embodiments of the present disclosure. Figure 12 As shown, in some embodiments, the model training method provided by the embodiments of the present disclosure includes the following steps.
[0186] In step S120 , the audio is processed based on the target model to generate an audio description of the audio.
[0187] In step S122, the audio description of the audio and the reference description of the audio are encoded to obtain a first vector of the audio description and a second vector of the reference description.
[0188] The first vector and the second vector contain word vectors.
[0189] In step S124 , weighted processing is performed on the first vector to obtain a third vector, and weighted processing is performed on the second vector to obtain a fourth vector.
[0190] The weighting process is used to differentiate the weights of the word vectors included in the first vector and the weights of the word vectors included in the second vector.
[0191] In step S126 , the similarity between the audio description and the reference description is determined based on the third vector and the fourth vector.
[0192] In step S128, the target model is trained based on the similarity between the audio description and the reference description.
[0193] The execution method and beneficial effects of the above steps S120 to S126 can be found in the above Figures 1-9 Any embodiment in will not be described in detail here.
[0194] The model training involved in the embodiments of the present disclosure can be performed in the model training stage or in the intensive training stage, and the embodiments of the present disclosure do not make specific limitations.
[0195] For example, Figure 13 FIG. 1 is a schematic diagram illustrating a model training process according to some embodiments of the present disclosure. Figure 13As shown, after the similarity calculation is completed for the audio description generated by the target model through steps S120 to S126, the similarity calculation result, the audio description, and the audio described by the audio description can be used as input to the target model for model training. The target model training method can be found in related art and will not be further described here.
[0196] In the above embodiment, weighted processing is used to differentiate the weights of the word vectors contained in the first vector to obtain a third vector; and weighted processing is used to differentiate the weights of the word vectors contained in the second vector to obtain a fourth vector. This can quantify the importance of the word vectors in the first and second vectors. Therefore, when determining the similarity between the audio description and the reference description based on the third and fourth vectors, the differential expression of the sound in the audio description and the reference description can be captured based on the importance of the word vectors, thereby improving the accuracy of the similarity calculation between the audio description and the reference description. In addition, the similarity between the audio description and the reference description can be used to measure the accuracy of the audio description generated by the target model. The higher the accuracy of the audio description generated by the target model, the stronger the target model's ability to generate audio descriptions. By feeding the similarity between the audio description and the reference description into the training of the target model, the target model's ability to generate audio descriptions can be improved. Furthermore, when the target model is used to perform tasks related to audio processing, the accuracy of task execution is improved. For example, in the technical application of audio retrieval, accurate audio descriptions can improve the accuracy of audio segment location. In accessibility services, accurate audio descriptions delivered to service recipients can improve the user experience.
[0197] In some exemplary embodiments of the present disclosure, training a target model based on the similarity between the audio description and the reference description may include: measuring the audio description based on the similarity between the audio description and the reference description, and the recall rate of the target model; and training the target model based on the measurement results.
[0198] Measuring the audio descriptions generated by the target model based on two dimensions: the similarity between the audio description and the reference description, and the recall rate of the target model. This prevents a single dimension from dominating the measurement results and improves their accuracy. Furthermore, training the target model based on these measurement results can improve its ability to generate audio descriptions.
[0199] In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of a first similarity among multiple second similarities. In the ranking of the first similarity and the multiple second similarities, the difference in similarity between adjacent rankings is greater than a preset threshold; wherein the first similarity refers to the similarity between the audio description and a reference description of the audio; and the second similarity refers to the similarity between the audio description and a reference description other than the reference description.
[0200] In the sorting of the first similarity and the plurality of second similarities, the difference in similarity between adjacent sorts is greater than a preset threshold, thereby reducing the impact of disturbances in the similarity calculation on the sorting result and improving the differentiation of similarities between the audio description and different reference descriptions. By improving the differentiation of similarities between the audio description and different reference descriptions.
[0201] Figure 14 is a block diagram of an audio processing device according to some embodiments of the present disclosure.
[0202] Reference Figure 14 , the audio processing device 140 includes:
[0203] The acquisition module 141 is configured to obtain a first vector of an audio description of an audio and a second vector of a reference description of the audio, wherein the first vector and the second vector include word vectors.
[0204] The weighting module 142 is configured to perform weighted processing on the first vector to obtain a third vector, and to perform weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is configured to differentiate the weights of the word vectors included in the first vector and the second vector.
[0205] The first determining module 143 is configured to determine a similarity between the audio description and the reference description based on the third vector and the fourth vector.
[0206] The second determining module 144 is configured to measure the audio description according to the similarity.
[0207] In some exemplary embodiments of the present disclosure, the weighting module 142 includes:
[0208] A first determining unit is configured to determine a first weight of a word vector included in the first vector based on an occurrence frequency of a word in the audio description and an inverse document frequency of the word, wherein the word corresponds to the word vector.
[0209] The first processing unit is configured to weight the word vectors included in the first vector based on the first weight to obtain a third vector.
[0210] The first determining unit is further configured to determine a second weight of a word vector included in the second vector based on an occurrence frequency of a word in the reference description and an inverse document frequency of the word.
[0211] The first processing unit is further configured to weight the word vectors included in the second vector based on the second weight to obtain a fourth vector.
[0212] In some exemplary embodiments of the present disclosure, the first determining unit is configured to:
[0213] Based on the frequency of occurrence of the vocabulary in the audio description, the inverse document frequency of the vocabulary is weighted to obtain a first weight of the word vector corresponding to the vocabulary in the first vector.
[0214] In some exemplary embodiments of the present disclosure, the weighting module 142 includes:
[0215] The second processing unit is configured to weight the word vectors related to the sound features contained in the first vector to obtain a third vector.
[0216] The third processing unit is configured to weight the word vectors related to the sound features contained in the second vector to obtain a fourth vector.
[0217] In some exemplary embodiments of the present disclosure, the second determination module 144 is configured to measure the audio description based on the similarity between the audio description and the reference description, and the recall rate of the target model used to generate the audio description. In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of the first similarity among multiple second similarities; in the ranking, the difference in similarity between adjacent rankings is greater than a preset threshold; wherein the first similarity refers to the similarity between the audio description and the reference description of the audio, and the second similarity refers to the similarity between the audio description and reference descriptions other than the reference description.
[0218] In some exemplary embodiments of the present disclosure, the device further includes a third determination module, configured to determine, for the first similarity and any two similarities among the multiple second similarities, that the ranking of one of the similarities is higher than the ranking of the other similarity based on the difference between one of the similarities and the other similarity being greater than the preset threshold; or to determine that the ranking of one of the similarities is equal to the ranking of the other similarity based on the one of the similarities being greater than the other similarity and the difference between the one of the similarities and the other similarity being less than or equal to the preset threshold.
[0219] In some exemplary embodiments of the present disclosure, the apparatus further includes:
[0220] The clustering module is configured to cluster the first similarity and the plurality of second similarities according to distances between the similarities to obtain one or more clusters.
[0221] A sorting module is configured to sort the one or more clusters based on the means of the elements in the clusters.
[0222] The order of elements in the same cluster is equal to the order of the cluster.
[0223] In some exemplary embodiments of the present disclosure, the sorting module is further configured to: merge, based on the means of the elements in the clusters, the clusters whose means differ from each other by less than the preset threshold.
[0224] Regarding the above Figure 14 The specific manner in which each module performs operations in the device in the embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.
[0225] Figure 15 FIG. 1 is a block diagram of another audio processing device according to some embodiments of the present disclosure. Figure 15 , the audio processing device 150 includes:
[0226] The acquisition module 151 is used to obtain audio.
[0227] The processing module 152 is configured to process the audio based on the target model to obtain an audio description of the audio.
[0228] Wherein, the target model is based on the above Figure 1-9 The measurement result of any embodiment is determined.
[0229] Regarding the above Figure 15 The specific manner in which each module performs operations in the device in the embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.
[0230] Figure 16 It is a block diagram of a model training device according to some embodiments of the present disclosure.
[0231] Reference Figure 16 , the model training device 160 includes:
[0232] The generating module 161 is configured to process the audio based on the target model to generate an audio description of the audio.
[0233] The encoding module 162 is configured to encode the audio description and the reference description of the audio to obtain a first vector of the audio description and a second vector of the reference description, wherein the first vector and the second vector include word vectors.
[0234] The weighting module 163 is configured to perform weighted processing on the first vector to obtain a third vector, and to perform weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is configured to differentiate the weights of the word vectors included in the first vector and the second vector.
[0235] The determination module 164 is configured to determine a similarity between the audio description and the reference description based on the third vector and the fourth vector.
[0236] The training module 165 is configured to train the target model based on the similarity.
[0237] In some exemplary embodiments of the present disclosure, the training module 165 includes:
[0238] A determining unit is configured to measure the audio description based on the similarity and a recall rate of the target model.
[0239] A training unit is used to train the target model based on the measurement result of the audio description.
[0240] In some exemplary embodiments of the present disclosure, the recall rate of the target model is related to the ranking of the first similarity among multiple second similarities; in the ranking, the similarity difference between adjacent rankings is greater than a preset threshold; wherein, the first similarity refers to the similarity between the audio description and a reference description of the audio; the second similarity refers to the similarity between the audio description and other reference descriptions other than the reference description.
[0241] Regarding the above Figure 16 The specific manner in which each module performs operations in the device in the embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.
[0242] In some embodiments, an electronic device is further provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: implement the method of any of the above method embodiments.
[0243] Figure 17 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 17 As shown, the electronic device 170 is implemented as a general-purpose computing device. Components of the electronic device 170 may include, but are not limited to, the at least one processing unit 171, the at least one storage unit 172, and a bus 173 connecting different system components (including the storage unit 172 and the processing unit 171).
[0244] The storage unit stores program code, which can be executed by the processing unit 171, so that the processing unit 171 performs the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above. For example, the processing unit 171 can perform any of the above method embodiments.
[0245] The storage unit 172 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1721 and / or a cache memory unit 1722 , and may further include a read-only memory unit (ROM) 1723 .
[0246] The storage unit 172 may also include a program / utility 1724 having a set (at least one) of program modules 1725, such program modules 1725 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0247] Bus 173 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0248] The electronic device 170 can also communicate with one or more external devices 700 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 170, and / or any device that enables the electronic device 170 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 175. Furthermore, the electronic device 170 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 176. As shown, the network adapter 176 communicates with other modules of the electronic device 170 via a bus 173. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the electronic device 170, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0249] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0250] The present disclosure also provides a non-transitory computer-readable storage medium. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the above method embodiments.
[0251] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.
[0252] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art after reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific functions of the described components, even if structurally not equivalent to the disclosed structures. In addition, although specific features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and beneficial for any given or specific application. In addition, with respect to the terms "including," "having," "having," "having," or variations thereof used in the specific embodiments or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0253] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0254] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An audio processing method, characterized in that: include: Obtaining a first vector of an audio description of audio, and a second vector of a reference description of the audio, the first vector and the second vector comprising word vectors; Performing weighted processing on the first vector to obtain a third vector, and performing weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is used to distinguish weights of word vectors included in the first vector and the second vector; determining a similarity between the audio description and the reference description based on the third vector and the fourth vector; The audio description is measured according to the similarity.
2. The method according to claim 1, characterized in that The step of performing weighted processing on the first vector to obtain a third vector and performing weighted processing on the second vector to obtain a fourth vector includes: determining a first weight of a word vector included in the first vector based on a frequency of occurrence of a word in the audio description and an inverse document frequency of the word, the word corresponding to the word vector; weighting the word vectors included in the first vector based on the first weight to obtain a third vector; determining a second weight of a word vector included in the second vector based on an occurrence frequency of a word in the reference description and an inverse document frequency of the word; Based on the second weight, the word vectors included in the second vector are weighted to obtain a fourth vector.
3. The method according to claim 2, characterized in that The determining, based on the frequency of occurrence of words in the audio description and the inverse document frequency of the words, a first weight of the word vector included in the first vector includes: Based on the frequency of occurrence of the vocabulary in the audio description, the inverse document frequency of the vocabulary is weighted to obtain a first weight of the word vector corresponding to the vocabulary in the first vector.
4. The method according to claim 1, wherein The step of performing weighted processing on the first vector to obtain a third vector and performing weighted processing on the second vector to obtain a fourth vector includes: weighting the word vectors related to the sound feature contained in the first vector to obtain a third vector; The word vectors related to the sound feature contained in the second vector are weighted to obtain a fourth vector.
5. The method according to any one of claims 1 to 4, characterized in that The measuring the audio description according to the similarity includes: The audio description is measured based on a similarity between the audio description and the reference description and a recall of a target model used to generate the audio description.
6. The method according to claim 5, characterized in that The recall rate of the target model is related to the ranking of the first similarity among the plurality of second similarities; In the sorting, the similarity difference between adjacent sorts is greater than a preset threshold; wherein the first similarity refers to the similarity between the audio description and a reference description of the audio; The second similarity refers to a similarity between the audio description and other reference descriptions other than the reference description.
7. The method according to claim 6, characterized in that The method for obtaining the ranking includes: For the first similarity and any two similarities among the plurality of second similarities, determining that the ranking of one of the similarities is higher than the ranking of the other similarity based on a difference between the one similarity and the other similarity being greater than the preset threshold; or Based on the one similarity being greater than the other similarity and the difference between the one similarity and the other similarity being less than or equal to the preset threshold, it is determined that the ranking of the one similarity is equal to the ranking of the other similarity.
8. The method according to claim 6, characterized in that The method for obtaining the ranking includes: Clustering the first similarity and the plurality of second similarities according to distances between the similarities to obtain one or more clusters; sorting the one or more clusters based on means of elements in the clusters; The order of elements in the same cluster is equal to the order of the cluster.
9. The method according to claim 8, characterized in that The sorting of the one or more clusters based on the means of the elements in the clusters comprises: Based on the means of the elements in the clusters, the clusters whose means have a difference less than the preset threshold are merged.
10. An audio processing method, characterized in that: include: Get audio; Processing the audio based on the target model to obtain an audio description of the audio; The target model is determined based on the measurement result of any one of claims 1-9.
11. A model training method, characterized in that: include: Processing the audio based on the target model to generate an audio description of the audio; Encoding the audio description and the reference description of the audio to obtain a first vector of the audio description and a second vector of the reference description, wherein the first vector and the second vector include word vectors; Performing weighted processing on the first vector to obtain a third vector, and performing weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is used to distinguish weights of word vectors included in the first vector and the second vector; determining a similarity between the audio description and the reference description based on the third vector and the fourth vector; Based on the similarity, the target model is trained.
12. The method according to claim 11, characterized in that The training of the target model based on the similarity includes: Measuring the audio description based on the similarity and the recall rate of the target model; The target model is trained based on the measurement results of the audio description.
13. The method according to claim 12, characterized in that The recall rate of the target model is related to the ranking of the first similarity among the plurality of second similarities; In the sorting, the similarity difference between adjacent sorts is greater than a preset threshold; wherein the first similarity refers to the similarity between the audio description and a reference description of the audio; The second similarity refers to a similarity between the audio description and other reference descriptions other than the reference description.
14. An audio processing device, characterized in that: include: an acquisition module, configured to obtain a first vector of an audio description of an audio, and a second vector of a reference description of the audio, wherein the first vector and the second vector comprise word vectors; a weighting module, configured to perform weighted processing on the first vector to obtain a third vector, and to perform weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is used to differentiate weights of word vectors included in the first vector and the second vector; a first determining module, configured to determine a similarity between the audio description and the reference description based on the third vector and the fourth vector; The second determining module is configured to measure the audio description according to the similarity.
15. The device according to claim 14, characterized in that The weighting module includes: a first determining unit, configured to determine a first weight of a word vector included in the first vector based on an occurrence frequency of a word in the audio description and an inverse document frequency of the word, the word corresponding to the word vector; a first processing unit, configured to weight the word vectors included in the first vector based on the first weight to obtain a third vector; The first determining unit is further configured to determine a second weight of a word vector included in the second vector based on an occurrence frequency of a word in the reference description and an inverse document frequency of the word; The first processing unit is further configured to weight the word vectors included in the second vector based on the second weight to obtain a fourth vector.
16. An audio processing device, characterized in that: include: Acquisition module, used to obtain audio; a processing module, configured to process the audio based on a target model to obtain an audio description of the audio; The target model is determined based on the measurement result according to any one of claims 1 to 9.
17. A model training device, characterized in that: include: A generation module, configured to process the audio based on the target model to generate an audio description of the audio; an encoding module, configured to encode the audio description and the reference description of the audio to obtain a first vector of the audio description and a second vector of the reference description, wherein the first vector and the second vector include word vectors; a weighting module, configured to perform weighted processing on the first vector to obtain a third vector, and to perform weighted processing on the second vector to obtain a fourth vector, wherein the weighted processing is used to differentiate weights of word vectors included in the first vector and the second vector; a determination module, configured to determine a similarity between the audio description and the reference description based on the third vector and the fourth vector; A training module is used to train the target model based on the similarity.
18. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the method according to any one of claims 1 to 13.
19. A non-transitory computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Text matching method and device, electronic equipment and storage medium
CN111898380A
Likaraoke detection method and device, electronic equipment and computer readable storage medium
CN113470693A
Word vector representation method based on self-attention mechanism and self-attention model
CN113761934A
Method and device for determining text similarity, equipment and storage medium
CN115374775A
Automatic music description generation method
CN116259289A