Methods, apparatus, electronic devices, and storage media for generating speech feature descriptions
The method addresses the challenge of describing voice data features by recognizing attribute labels and generating natural language descriptions, enhancing accuracy and conciseness through techniques like pre-trained models and large-scale models.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-25
AI Technical Summary
Existing technologies struggle to accurately and efficiently describe the features of voice data, such as speaker information, voice quality, and emotional information, leading to redundant and non-focused descriptions.
A method and apparatus for generating voice feature descriptions by acquiring target voice data, recognizing it to obtain multiple attribute labels, determining a target label, and generating natural language information to describe the voice features, using techniques like pre-trained models and large-scale models to improve accuracy and focus.
The method enhances the accuracy and conciseness of voice feature descriptions by filtering out redundant information and emphasizing key features, improving the naturalness and focus of the descriptions.
Smart Images

Figure 2026053484000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, particularly to the fields of artificial intelligence such as voice technology and large-scale models. Specifically, it relates to a method, apparatus, electronic device, and storage medium for generating voice feature descriptions.
Background Art
[0002] With the development of artificial intelligence, voice technology is being increasingly widely applied in life. For example, in the scenario of an online meeting, the collected voice can be converted into text in real time and displayed to the user. Also, for example, in the field of search, the voice input by the user can be recognized to determine the user's search intention.
Summary of the Invention
[0003] This application provides a method, apparatus, electronic device, and storage medium for generating voice feature descriptions. The specific technical solutions are as follows.
Means for Solving the Problems
[0004] According to one aspect of this application, a step of obtaining target voice data, a step of recognizing the target voice data to obtain a plurality of voice attribute labels of the target voice data, a step of determining a target voice attribute label from the plurality of voice attribute labels, a step of generating natural language description information corresponding to the target voice data based on the target voice attribute label, where the natural language description information is used to describe the voice features of the target voice data, and a method for generating a voice feature description is provided.
[0005] According to another aspect of this application, a first acquisition module configured to obtain target voice data, A recognition module configured to recognize the target audio data and obtain multiple audio attribute labels of the target audio data, A first determination module configured to determine a target voice attribute label from the plurality of voice attribute labels, An apparatus for generating a speech feature description is provided, comprising a first generation module configured to generate natural language description information corresponding to the target speech data based on the target speech attribute labels, wherein the natural language description information is used to describe the speech features of the target speech data.
[0006] According to another aspect of this application, At least one processor, Includes memory that is communicably connected to at least one processor, The provided electronic device stores instructions that can be executed by the at least one processor in the memory, and when the instructions are executed by the at least one processor, it causes the at least one processor to execute the method for generating the speech feature description described in the above embodiment.
[0007] According to another aspect of this application, there is a non-temporary computer-readable storage medium in which computer instructions are stored, the computer instructions being used to cause a computer to perform the method for generating the speech feature description described in the above embodiment.
[0008] According to another aspect of this application, a computer program is provided, wherein when the computer program is executed by a processor, the steps of a method for generating a speech feature description as described in the above embodiment are realized.
[0009] Please understand that this section is not intended to identify the main or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will be better understood from the following description. [Brief explanation of the drawing]
[0010] The drawings are provided for the purpose of better understanding this technical proposal and do not limit this application. [Figure 1] This is a schematic flowchart of a method for generating an audio feature description according to one embodiment of this application. [Figure 2] This is a schematic flowchart of a method for generating an audio feature description according to another embodiment of this application. [Figure 3] This is a schematic flowchart of a method for generating an audio feature description according to another embodiment of this application. [Figure 4] This is a schematic flowchart of a method for generating an audio feature description according to another embodiment of this application. [Figure 5] This is a schematic diagram of a device for generating speech feature descriptions according to one embodiment of this application. [Figure 6] This is a block diagram of an electronic device that implements a method for generating speech feature descriptions according to an embodiment of this application. [Modes for carrying out the invention]
[0011] The following description, with reference to the drawings, illustrates exemplary embodiments of this application and includes various details of these embodiments to aid understanding. These are merely illustrative examples. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and configurations are omitted in the following description.
[0012] In addition, the acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant regulations of laws and regulations in various countries and do not violate public order and good customs.
[0013] A method, apparatus, electronic device, and storage medium for generating an audio feature description according to an embodiment of this application will be described below with reference to the drawings.
[0014] In some scenarios, for audio data, it may be necessary to extract speaker information, voice quality information, emotional information, etc., and describe the audio features of the audio data based on these information. Therefore, an embodiment of this application proposes a method for generating an audio feature description.
[0015] FIG. 1 is a schematic flowchart of a method for generating an audio feature description according to an embodiment of this application.
[0016] The method for generating an audio feature description according to an embodiment of this application can be executed by an apparatus for generating an audio feature description according to an embodiment of this application, and the apparatus can be set in an electronic device.
[0017] Here, the electronic device may be any device with a computing function, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal may be a hardware device with various operating systems, touch screens, and / or displays, such as an in-vehicle device, a mobile phone, a tablet, a personal digital assistant, a wearable device, etc.
[0018] As shown in FIG. 1, the method for generating the audio feature description includes the following steps 101 to 104.
[0019] In step 101, target audio data is acquired.
[0020] In this application, the target voice data may be any voice data. For example, the target voice data may be voice data acquired in real time by an audio collection device, or may be audio data extracted from voice and video data, or may be voice data obtained by preprocessing the original voice data, or may be voice data obtained by other means, but is not limited thereto.
[0021] As an example, by performing voice segment detection on the original voice data, silent segments can be removed to obtain the target voice data. Thereby, by removing silent segments from the voice data through voice segment detection, the influence on the calculation of the overall volume, pitch, speech rate, etc. due to silence can be avoided.
[0022] In step 102, the target voice data is recognized to obtain a plurality of voice attribute labels of the target voice data.
[0023] Here, the voice attribute labels can be used to identify the age, gender, accent, speech rate, pitch, volume, vocal undulation, emotion, etc. of the speaker.
[0024] For example, the voice attribute labels can include, but are not limited to, age labels, gender labels, accent labels, speech rate labels, pitch labels, volume labels, emotion labels, etc.
[0025] For example, the age labels can include child, teenager, young adult, middle-aged person, elderly person, etc., and the speech rate labels can include slow speech rate, normal speech rate, fast speech rate, etc.
[0026] For example, the plurality of voice attribute labels of the target voice data can include young adult, high volume, fast speech rate, intense emotion, etc.
[0027] In this application, for different speech attributes, a recognition policy that matches the speech attribute based on target speech data can be adopted to determine the speech attribute label.
[0028] For example, a pre-trained speech representation learning model can be used to recognize target speech data and determine speech attribute labels to identify the speaker's age, gender, accent, and other characteristics.
[0029] For example, a speech rate value can be calculated based on the text obtained by speech recognition of target speech data, and the corresponding speech label can be determined based on the calculated speech rate value.
[0030] For example, the speech rate value can be determined based on the ratio of the total number of phonemes to the speech duration, and the speech rate label for the target speech data can be determined based on the speech rate value and the speech rate range corresponding to each speech rate label. For example, the speech rate label corresponding to a speech rate value of less than 40% is "slow speech rate," and the speech rate label corresponding to a speech rate value of 40% or more and 60% or less is "normal speech rate."
[0031] For example, the volume value of the target audio data can be calculated, and based on the volume value, volume labels such as low volume, normal volume, and high volume can be determined.
[0032] For example, the pitch value of target voice data can be calculated, and based on this pitch value and pitch data of a different gender, pitch labels such as low pitch, normal pitch, and high pitch can be determined.
[0033] For example, the degree of fluctuation in voice can be determined based on the variance of pitch values.
[0034] For example, a pre-trained emotion recognition model can be used to recognize target audio data and determine the emotion label corresponding to that data.
[0035] In step 103, the target voice attribute label is determined from multiple voice attribute labels.
[0036] To improve the focus of information, this application allows for screening multiple voice attribute labels and determining a target voice attribute label from among them.
[0037] The target voice attribute labels may be one or more, and are not limited thereto in this application.
[0038] One possible implementation is to determine a target voice attribute label from multiple voice attribute labels based on the attribute information of the voice attribute labels. The attribute information of the voice attribute labels may be data that describes the characteristics of the voice attribute labels.
[0039] For example, the attribute information of a speech attribute label may include, but is not limited to, the sampling probability corresponding to the speech attribute label, the quantization value corresponding to the speech attribute label, and whether or not the speech attribute label is a neutral attribute label.
[0040] The sampling probability corresponding to an audio attribute label can refer to the probability that an audio recognition label is selected.
[0041] Regarding the quantization values corresponding to speech attribute labels, for example, the quantization value corresponding to the speech rate label refers to the speech rate value, and the quantization value corresponding to the volume label refers to the volume value, and so on.
[0042] Neutral attribute labels refer to unbiased attribute labels. For example, "normal speech rate," "normal volume," and "normal pitch" all belong to the category of neutral attribute labels.
[0043] For example, among multiple audio attribute labels, an audio attribute label other than the neutral attribute label can be determined as the target attribute label.
[0044] Another possible implementation involves obtaining the need to generate a speech feature description corresponding to the target speech data, and then determining the target speech attribute label from multiple speech attribute labels based on that need. The need to generate the speech feature description may be entered by the user or determined according to the requirements of the application scenario, but is not limited to these.
[0045] For example, if the need to generate speech feature descriptions is to describe a speaker's accent, speech rate, and emotion, then from a set of multiple speech attribute labels, the speech attribute labels associated with accent, speech rate, and emotion can be determined as the target speech attribute labels.
[0046] In step 104, natural language description information corresponding to the target audio data is generated based on the target audio attribute labels.
[0047] Natural language description information can be used to describe the speech features of target speech data, such as volume features, speech rate features, timbre features, and emotion features.
[0048] For example, natural language description information may be in text format or audio format, but is not limited to these.
[0049] For example, natural language description information can be generated using a large-scale model based on target speech attribute labels.
[0050] As an example, prompt information can be generated based on target speech attribute labels and feature description requirements, and a large-scale model can be used to process this prompt information and generate natural language description information. The prompt information is used to instruct the large-scale model to generate a natural language description for describing the speech features. Feature description requirements may include, but are not limited to, character count requirements and output format requirements for the natural language requirement information.
[0051] For example, if the target speech attribute labels include high volume and fast speaking speed, the generated natural language description might be something like "speaker has a high volume and speaks quickly."
[0052] For example, after generating natural language description information, it can be displayed or pushed to the corresponding related object.
[0053] The method for generating speech feature descriptions according to the embodiments of this application is applicable to a variety of scenarios.
[0054] For example, it can be applied to customer service scenarios, where the natural language description information of the generated speech can be used to accurately assess the user's emotional state and provide personalized services.
[0055] For example, it can be applied to smart education scenarios, where teachers can use natural language description information of students' speech to understand changes in students' emotional fluctuations and speech rate, and provide personalized instructional feedback to teachers.
[0056] In the embodiments of this application, target audio data is recognized, multiple audio attribute labels are obtained, a target audio attribute label is determined from the multiple audio attribute labels, the audio attribute labels are screened, and natural language description information for describing the audio features of the target audio data is generated based on the screened target audio attribute label, thereby reducing redundant descriptions and improving the accuracy and conciseness of the descriptions.
[0057] Figure 2 is a schematic flowchart of a method for generating an audio feature description according to another embodiment of this application.
[0058] The method for generating the speech feature description includes the following steps 201 to 205.
[0059] In step 201, the target audio data is acquired.
[0060] In step 202, the target audio data is recognized, and multiple audio attribute labels of the target audio data are obtained.
[0061] In this application, steps 201 to 202 can refer to any one embodiment in each of the embodiments of this application, so they will not be described again here.
[0062] In step 203, the key voice attribute label is determined from multiple voice attribute labels.
[0063] Here, key speech attribute labels may be understood as important attribute labels for describing and distinguishing speech features. Alternatively, key speech attribute labels may be understood as attribute labels for identifying key speech features.
[0064] For example, key speech attribute labels may include, but are not limited to, fast speech rate, high volume, high pitch, happy, angry, etc.
[0065] One possible implementation is that if the quantization value corresponding to a second speech attribute label among multiple speech attribute labels is greater than the corresponding first threshold, the second speech attribute label can be determined as the key speech attribute label.
[0066] For example, if the second audio attribute label is "high volume" and the corresponding quantization value is 100 dB, which is greater than a predetermined value of 90 dB, then "high volume" can be determined as the key audio attribute label.
[0067] Please note that if the second audio attribute label is different, the corresponding first threshold may be different, and this is not limited to that case.
[0068] Therefore, the method for screening key speech attribute labels based on quantized values corresponding to speech attribute labels and corresponding first thresholds is simple, and the naturalness and focus of the natural language description information can be improved by generating natural language description information based on the screened speech attribute labels. Furthermore, different screening needs can be met by adjusting the thresholds.
[0069] Another possible implementation involves determining whether a voice attribute label belongs to a highlighting attribute label, and if a third voice attribute label among several belongs to a highlighting attribute label, this third voice attribute label can be determined as the key voice attribute label.
[0070] Here, the emphasis attribute label can refer to the audio attribute label that needs to be emphasized. For example, the emphasis attribute label may be pre-set.
[0071] For example, emphasis attribute labels could include each emotion label, such as low volume, high volume, fast speech rate, low pitch, high pitch, etc.
[0072] Therefore, by determining the speech attribute labels belonging to the emphasized attribute labels as key speech attribute labels, it is possible to screen for speech attribute labels that need attention, filter out less prominent features, and generate natural language description information based on the screened speech attribute labels. This preserves key speech features in the target speech data, improving the naturalness and focus of the information in the natural language description information.
[0073] Key voice attribute features can be screened using one or two of the above forms, but are not limited to these.
[0074] In step 204, the target voice attribute label is determined based on the key voice attribute label.
[0075] For example, the key voice attribute label can be directly determined as the target voice attribute label.
[0076] For example, target voice attribute labels can be determined based on the historical frequency of key voice attribute labels. For instance, the first predetermined number of voice attribute labels with the highest historical frequency among the key voice attribute labels can be determined as the target voice attribute labels. This improves the accuracy of the target voice attribute labels.
[0077] In step 205, natural language description information corresponding to the target audio data is generated based on the target audio attribute labels.
[0078] In this application, step 205 can refer to any one embodiment in each of the embodiments of this application, and therefore will not be described again here.
[0079] In the embodiments of this application, a key speech attribute label is determined from a plurality of speech attribute labels, a target speech attribute label is determined based on the key speech attribute label, and natural language description information is generated based on the target speech attribute label. This ensures that the key speech features are included in the natural language description information, thereby improving the focus and accuracy of the description information.
[0080] Figure 3 is a schematic flowchart of a method for generating an audio feature description according to another embodiment of this application.
[0081] As shown in Figure 3, the method for generating the speech feature description includes the following steps 301 to 305.
[0082] In step 301, the target audio data is acquired.
[0083] In step 302, the target audio data is recognized, and multiple audio attribute labels of the target audio data are obtained.
[0084] In this application, steps 301 to 302 can refer to any one embodiment in each of the embodiments of this application, so they will not be described again here.
[0085] In step 303, the sampling probability corresponding to the speech attribute label is obtained.
[0086] In this application, each audio attribute label has a corresponding sampling probability, and the sampling probabilities of different audio attribute labels may be the same or different, but are not limited to this.
[0087] For example, the sampling probability corresponding to a speech attribute label can be determined based on the frequency of occurrence of the speech attribute label in the historical data. For instance, a higher frequency corresponds to a higher sampling probability. The frequency of occurrence in the historical data can refer to the frequency of occurrence of the speech attribute label in the natural language description of previously generated speech.
[0088] The sampling probability corresponding to each audio attribute label may be constant or updated, but this is not a limitation.
[0089] Since multiple speech attribute labels may include neutral attribute labels such as normal speech rate and normal volume, to reduce the probability of a neutral attribute label being selected, a neutral attribute label can be determined from multiple speech attribute labels as an example. If the first speech attribute label among the multiple speech attribute labels is a neutral attribute label, the initial sampling probability of the first speech attribute label can be updated based on the quantization value corresponding to the first speech attribute label to obtain the sampling probability.
[0090] The updated sampling probability is smaller than the initial sampling probability. The quantized value corresponding to the first speech attribute label can refer to the quantized value of the speech attribute to which the first speech attribute label belongs.
[0091] For example, the sampling probability p of the first speech attribute can be calculated by the following equation (1).
number
[0092] Assuming the speech rate value is in the range of [40% to 60%], the speech rate label is "Normal speech rate". For example, if the speech rate value of the target audio data is 45%, the corresponding speech rate label is "Normal speech rate", and "Normal speech rate" is a neutral attribute label, with a quantization value of 45% corresponding to "Normal speech rate". Then, using equation (1) above, the sampling probability corresponding to "Normal speech rate" can be calculated, and the initial sampling probability corresponding to "Normal speech rate" can be updated and replaced with the calculated sampling probability.
[0093] Therefore, for neutral attribute labels among multiple speech attribute labels, a sampling probability lower than the initial sampling probability can be obtained by updating them based on the quantization value corresponding to the neutral attribute label. This reduces the probability of neutral attribute labels being selected, filters out inconspicuous speech features, increases attention to prominent speech features, and improves the naturalness and information focus of the natural language description.
[0094] In step 304, multiple voice attribute labels are sampled based on sampling probabilities to obtain the target voice attribute label.
[0095] In this application, a target voice attribute label can be obtained by sampling multiple voice attribute labels based on the sampling probability and number of samples of multiple voice attribute labels of the target voice data.
[0096] For example, if the target audio data has seven audio attribute labels and the sampling rate is two, two audio attribute labels can be selected based on the sampling probabilities of these seven labels to generate a natural language description.
[0097] The number of samples can be set or determined according to actual needs, and is not limited to this.
[0098] In step 305, natural language description information corresponding to the target audio data is generated based on the target audio attribute labels.
[0099] In this application, step 305 can refer to any one embodiment in each of the embodiments of this application, and therefore will not be described again here.
[0100] In the embodiments of this application, by sampling multiple speech attribute labels based on the sampling probabilities of multiple speech attribute labels and obtaining a target attribute label, the target speech attribute label can be identified more quickly, thereby improving the efficiency of generating natural language description information.
[0101] Figure 4 is a schematic flowchart of a method for generating an audio feature description according to another embodiment of this application. As shown in Figure 4, the method for generating the speech feature description includes the following steps 401 to 405.
[0102] In step 401, the target audio data is acquired. In step 402, the target audio data is recognized, and multiple audio attribute labels of the target audio data are obtained. In step 403, the target voice attribute label is determined from multiple voice attribute labels.
[0103] Steps 401 to 403 in this application can refer to any one embodiment in each of the embodiments of this application, and therefore will not be described again here.
[0104] In step 404, the target description template is determined based on the target voice attribute labels.
[0105] Here, a descriptive template can refer to a template for describing speech features. Exemplarily, a descriptive template may be, but is not limited to, complete natural language description information or a natural language description with slot markers.
[0106] For example, one description template is "The speaker's speaking speed is [speaking speed] and their emotion is [emotion]", and this description template contains two slots: [speaking speed] and [emotion].
[0107] One possible implementation is to set up at least one template category, each template category having a corresponding descriptive template. Based on the target audio attribute label, a target template category can be determined from at least one template category, and then a target descriptive template can be determined from the descriptive templates corresponding to the target template category.
[0108] For example, by comparing the target voice attribute label with the descriptive information of each template category, it is possible to determine the target template category that matches the target voice attribute label.
[0109] For example, by combining various voice attributes, multiple voice attribute combinations can be obtained, and each voice attribute combination can be designated as a single template category. Based on the target voice attribute label, a target voice attribute combination can be determined, and the template category corresponding to that target voice attribute combination can be designated as the target template category.
[0110] For example, speech attribute combinations may include speech rate-emotion, speech rate-emotion-age, pitch-speech rate-emotion, etc. If the target speech attribute labels include fast speech rate and intense emotion, the corresponding speech attribute combination is speech rate-emotion, and the category corresponding to speech rate-emotion becomes the target template category.
[0111] For example, one or more templates can be randomly selected from the description templates corresponding to the target template category to be designated as the target description templates, or the target description templates can be determined by sampling description templates based on the sampling probability of the description templates corresponding to the target template category, or a predetermined number of templates with the highest historical usage rate can be selected as the target description templates based on the historical usage rate of each description template corresponding to the target template category.
[0112] Therefore, by selecting a target template category based on the target speech attribute label and determining the target description template from the description templates corresponding to the target template category, the accuracy of the selected description template can be improved, thereby improving the accuracy of natural language description information.
[0113] Selectively, for each combination of speech attributes, a large-scale model can be used to generate descriptive templates belonging to the corresponding template category based on the speech attribute combination. This improves the diversity of descriptive templates.
[0114] In step 405, natural language description information is generated based on the target description template.
[0115] This application enables the generation of natural language descriptions based on target description templates and target speech attribute labels.
[0116] One possible implementation is to use target speech attribute labels to fill in slots in the target description template and obtain natural language description information.
[0117] Another possible implementation involves pre-constructing a thesaurus for each speech attribute label, determining the target synonym for the target speech attribute label from the thesaurus corresponding to the target speech attribute label, and then filling in the slots of the target description template based on the target synonym to obtain natural language description information.
[0118] For example, if the speech attribute label is "high pitch," the corresponding thesaurus entry would be "shrill, sharp, high, unpleasant." If the target speech attribute label includes "high pitch," synonyms such as "shrill and sharp" can be selected from the thesaurus to generate a natural language description.
[0119] For example, if the speech attribute label is "fast speech rate," the corresponding thesaurus entries might be "like a machine gun, machine gun-like, uninterrupted." If the target speech attribute label includes "fast speech rate," synonyms such as "machine gun-like" can be selected from the thesaurus to generate a natural language description.
[0120] For example, target synonyms for generating natural language description information can be randomly selected from a thesaurus corresponding to a target phonetic attribute label. Alternatively, target synonyms can be obtained by sampling based on the sampling probability of each synonym in the thesaurus. The sampling probability of synonyms may be constant or updated based on the historical frequency of use of the synonyms, but is not limited to this.
[0121] Therefore, by selecting target synonyms from a thesaurus corresponding to the target phonetic attribute label and generating natural language description information, the repeated use of the same word can be reduced, making the description more diverse and clearer, and improving the diversity and flexibility of the description.
[0122] For example, by using a large-scale model to generate synonyms for each speech attribute label and building a thesaurus based on the generated synonyms, the thesaurus's coverage can be expanded to meet various descriptive needs.
[0123] For example, by constructing corresponding thesauruses for different speech speed labels such as slow, fast, and normal speech speed, the detail and diversity of the description can be improved.
[0124] Due to the high co-occurrence frequency of certain attribute combinations, the description template may be called repeatedly. Based on this, for example, historical speech feature description data can be obtained, and the co-occurrence frequency of each speech attribute combination can be determined based on the historical speech feature description data. For speech attribute combinations whose co-occurrence frequency is greater than a second threshold, a description template for the template category corresponding to that speech attribute combination can be generated based on the co-occurrence frequency.
[0125] Historical speech feature description data can refer to natural language description information previously generated to describe the speech features of speech data.
[0126] For example, by determining the number of templates based on co-occurrence frequency and generating a corresponding template description template based on the given combination of speech attributes, repeated template calls for high-frequency attribute combinations can be avoided.
[0127] For example, if the co-occurrence frequency of speech rate and emotion is greater than a second threshold, the number of templates in the template category corresponding to that combination of speech attributes will be N = log(co-occurrence frequency) × cardinality, and N difference description templates can be generated using a large-scale model.
[0128] Therefore, for audio attribute combinations with high co-occurrence frequency, a description template for the template category corresponding to that audio attribute combination can be generated based on the co-occurrence frequency. This reduces the probability of repeatedly calling the description template for high-frequency attribute combinations and enriches the diversity of descriptions.
[0129] In the embodiments of this application, the accuracy of selecting a target description template can be improved by determining the target description template based on the target speech attribute label, and the generation efficiency is high because natural language description information is generated based on the target description template.
[0130] To realize the above embodiment, the embodiment of this application further provides an apparatus for generating speech feature descriptions. Figure 5 is a schematic diagram of an apparatus for generating speech feature descriptions according to one embodiment of this application.
[0131] As shown in Figure 5, the device 500 that generates speech feature descriptions is A first acquisition module 510 configured to acquire target audio data, A recognition module 520 is configured to recognize the target audio data and obtain multiple audio attribute labels of the target audio data, A first determination module 530 is configured to determine a target voice attribute label from the plurality of voice attribute labels, A first generation module 540 configured to generate natural language description information corresponding to the target speech data based on the target speech attribute labels, the natural language description information includes the first generation module 540 used to describe the speech features of the target speech data.
[0132] Selectively, the first decision module 530, A key voice attribute label is determined from the aforementioned plurality of voice attribute labels. The system is configured to determine the target voice attribute label based on the key voice attribute label.
[0133] Selectively, the first decision module 530, In response to the fact that the quantization value corresponding to the second voice attribute label among the plurality of voice attribute labels is greater than the corresponding first threshold, the second voice attribute label is determined to be the key voice attribute label, and The system is configured to perform at least one of the following: determining the third voice attribute label as a key voice attribute label in response to the fact that the third voice attribute label among the plurality of voice attribute labels belongs to the emphasis attribute label.
[0134] Selectively, the first decision module 530, The sampling probability corresponding to the aforementioned audio attribute label is obtained, The system is configured to sample the plurality of voice attribute labels based on the sampling probability and obtain the target voice attribute label.
[0135] Selectively, the first decision module 530, In response to the determination that a first voice attribute label among the plurality of voice attribute labels is a neutral attribute label, the system is configured to update the initial sampling probability of the first voice attribute label based on the quantization value corresponding to the first voice attribute label, and to obtain the sampling probability. The aforementioned sampling probability is smaller than the aforementioned initial sampling probability.
[0136] Selectively, the first generation module 540 is, Based on the aforementioned target voice attribute labels, a target description template is determined. It is configured to generate natural language description information based on the aforementioned target description template.
[0137] Selectively, the first generation module 540 is, The target synonyms for the target voice attribute labels are determined from the thesaurus corresponding to the target voice attribute labels. Based on the aforementioned target synonyms, the system is configured to fill in the slots in the target description template and obtain natural language description information.
[0138] Selectively, the first generation module 540 is, Based on the aforementioned target voice attribute labels, determine the target template category from at least one template category. The system is configured to determine the target description template from the description templates corresponding to the target template category.
[0139] The device that selectively generates the speech feature description is A second acquisition module configured to acquire historical speech feature description data, A second determination module configured to determine the co-occurrence frequency of each voice attribute combination based on the aforementioned historical voice feature description data, The system may further include a second generation module configured to generate a descriptive template for a template category corresponding to any one combination of voice attributes, based on the co-occurrence frequency, in response to the co-occurrence frequency of any one combination of voice attributes being greater than a second threshold.
[0140] The interpretation and explanation of the aforementioned embodiment of the method for generating speech feature descriptions are also applicable to the apparatus for generating speech feature descriptions in that embodiment, and therefore will not be explained again here.
[0141] In the embodiments of this application, target audio data is recognized to obtain multiple audio attribute labels, a target audio attribute label is determined from the multiple audio attribute labels, the audio attribute labels are screened, and natural language description information for describing the audio features of the target audio data is generated based on the screened target audio attribute label. This reduces redundant descriptions and improves the accuracy and conciseness of the descriptions.
[0142] According to embodiments of this application, the application further provides an electronic device, a readable storage medium, and a computer program.
[0143] Figure 6 is a schematic block diagram of an exemplary electronic device 600 that can be used to carry out embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other appropriate computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are illustrative and are not intended to limit the realization of the present application described and / or required herein.
[0144] As shown in Figure 6, device 600 includes a computing unit 601, which can perform various appropriate operations and processes based on computer programs stored in ROM (Read-Only Memory) 602 or computer programs loaded from storage unit 608 into RAM (Random Access Memory) 603. RAM 603 can also store various programs and data necessary for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are connected to each other via bus 604. An I / O (Input / Output) interface 605 is also connected to bus 604.
[0145] Multiple components in device 600 are connected to the I / O interface 605, which includes input units 606 such as a keyboard and mouse, output units 606 such as various displays and speakers, storage units 608 such as magnetic disks and optical disks, and communication units 609 such as a network card, modem, and wireless communication transceiver. The communication units 609 allow device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.
[0146] The computing unit 601 may be a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various dedicated AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as a method for generating a speech feature description. For example, in some embodiments, a method for generating a speech feature description can be implemented as a computer software program tangibly contained in a machine-readable medium such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed into device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the method for generating the speech feature description can be performed. Selectively, in another embodiment, the computing unit 601 may be configured to perform a method for generating a speech feature description by any other suitable method (for example, by firmware).
[0147] Various embodiments of the systems and technologies described herein can be implemented as digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments are implemented by one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor may be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0148] Program code for carrying out the method of this application can be written using any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, and when the program code is executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagrams will be performed. The program code may be fully executed on a machine, partially executed on a machine, partially executed on a machine and partially executed on a remote machine as a standalone software package, or fully executed on a remote machine or server.
[0149] In the context of this application, a machine-readable medium may be a tangible medium that contains or stores programs used by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections based on one or more lines, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0150] To provide user interaction, the systems and techniques described herein can be implemented on a computer, which has a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball), and the user can provide input to the computer through the keyboard and pointing device. Other types of devices can also provide user interaction, for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from the user may be received in any form (including acoustic input, voice input, or haptic input).
[0151] The systems and technologies described herein can be implemented in a computing system including backend components (e.g., as a data server), a computing system including middleware components (e.g., an application server), a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser, through which the user can interact with embodiments of the systems and technologies described herein), or in a computing system including any combination of such backend components, middleware components, and frontend components. The components of the system can be interconnected via digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.
[0152] A computer system can include a client and a server. The client and server are generally geographically separated and typically interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers that have a client-server relationship with each other. The server may be a cloud server, also called a cloud computing server or cloud host, a hosting product within a cloud computing service ecosystem, addressing the management difficulties and limited business scalability of traditional physical hosts and VPS (Virtual Private Server) services. The server may also be a server for distributed systems or a server incorporating blockchain technology.
[0153] According to embodiments of this application, the application further provides a computer program that performs a method for generating an audio feature description according to the above embodiments of this application when the computer program is executed by a processor.
[0154] The steps can be rearranged, added, or deleted using the various forms of flows shown above. For example, each step described in this application may be performed in parallel, sequentially, or in a different order, as long as the desired results of the invention disclosed herein are achieved, and are not limited herein.
[0155] The specific embodiments described above do not limit the scope of protection of this application. Those skilled in the art can make various modifications, combinations, subcombinations, and substitutions depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection.
Claims
1. Steps to acquire target audio data, The steps include: recognizing the target audio data and obtaining multiple audio attribute labels for the target audio data; The steps include determining a target voice attribute label from the aforementioned plurality of voice attribute labels, A step of generating natural language description information corresponding to the target audio data based on the target audio attribute labels, wherein the natural language description information is used to describe the audio features of the target audio data. A method for generating a speech feature description that includes [specific elements].
2. The step of determining the target voice attribute label from the aforementioned plurality of voice attribute labels is: The steps include determining a key voice attribute label from the aforementioned plurality of voice attribute labels, The steps include determining the target voice attribute label based on the key voice attribute label, A method for generating the speech feature description described in claim 1, including the method described in claim 1.
3. The step of determining a key voice attribute label from the aforementioned plurality of voice attribute labels is: The steps include determining the second voice attribute label as a key voice attribute label in response to the fact that the quantization value corresponding to the second voice attribute label among the plurality of voice attribute labels is greater than the corresponding first threshold, and In response to the fact that a third voice attribute label among the plurality of voice attribute labels belongs to the emphasis attribute label, the third voice attribute label is determined to be the key voice attribute label. A method for generating the speech feature description according to claim 2, comprising at least one of the following.
4. The step of determining the target voice attribute label from the aforementioned plurality of voice attribute labels is: The steps include obtaining the sampling probability corresponding to the aforementioned audio attribute label, The steps include sampling the plurality of voice attribute labels based on the sampling probability to obtain the target voice attribute label, A method for generating the speech feature description described in claim 1, including the method described in claim 1.
5. The step of obtaining the sampling probability corresponding to the aforementioned audio attribute label is: In response to the determination that a first voice attribute label among the plurality of voice attribute labels is a neutral attribute label, the initial sampling probability of the first voice attribute label is updated based on the quantization value corresponding to the first voice attribute label, and the sampling probability is obtained. The method for generating an audio feature description according to claim 4, wherein the sampling probability is smaller than the initial sampling probability.
6. The step of generating natural language description information corresponding to the target speech based on the target speech attribute label is: The steps include determining a target description template based on the aforementioned target voice attribute labels, The steps include generating the natural language description information based on the target description template, A method for generating the speech feature description described in claim 1, including the method described in claim 1.
7. The step of generating the natural language description information based on the target description template is: The steps include determining the target synonym of the target voice attribute label from a thesaurus corresponding to the target voice attribute label, The steps include: filling in the slots of the target description template based on the target synonyms to obtain the natural language description information; A method for generating the speech feature description described in claim 6, including the method described in claim 6.
8. The step of determining a target description template based on the aforementioned target voice attribute label is: The steps include determining a target template category from at least one template category based on the target voice attribute label, The steps include determining the target description template from the description templates corresponding to the target template category, A method for generating the speech feature description described in claim 6, including the method described in claim 6.
9. Steps to obtain historical speech feature description data, The steps include determining the co-occurrence frequency of each combination of speech attributes based on the aforementioned historical speech feature description data, A step of generating a descriptive template for a template category corresponding to any one combination of voice attributes based on the co-occurrence frequency, in response to the co-occurrence frequency of any one combination of voice attributes being greater than a second threshold, A method for generating an audio feature description according to claim 8, further comprising:
10. A first acquisition module configured to acquire target audio data, A recognition module configured to recognize the target audio data and obtain multiple audio attribute labels of the target audio data, A first determination module configured to determine a target voice attribute label from the plurality of voice attribute labels, A first generation module configured to generate natural language description information corresponding to the target audio data based on the target audio attribute labels, wherein the natural language description information is used to describe the audio features of the target audio data. A device that generates speech feature descriptions, including [specific features].
11. At least one processor, A memory that is communicably connected to at least one of the processors, Includes, An electronic device wherein the memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor causes the at least one processor to execute the method for generating a speech feature description according to any one of claims 1 to 9.
12. A non-temporary computer-readable storage medium in which computer instructions are stored, The computer instruction is a non-temporary computer-readable storage medium that causes a computer to execute the method for generating a speech feature description according to any one of claims 1 to 9.
13. It is a computer program, A computer program wherein, when executed by a processor, the steps of a method for generating a speech feature description according to any one of claims 1 to 9 are realized.