Voice feature description generation method and device, electronic equipment and storage medium
By identifying and filtering speech attribute tags to generate natural language descriptions, the redundancy problem of speech feature descriptions is solved, and a more accurate and concise speech feature description is achieved.
Patent Information
- Application Number
- CN202510850301.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have difficulty in efficiently generating natural language descriptions of speech features, resulting in redundant and inaccurate descriptions.
By acquiring target speech data, identifying multiple speech attribute tags, filtering out key or target speech attribute tags, generating natural language description information, and using a large model for feature description.
It improves the accuracy and conciseness of speech feature description, reduces redundancy, and enhances information focus and naturalness.
Smart Images

Figure CN120708612A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, in particular to artificial intelligence fields such as speech technology and large models, and specifically to a method, device, electronic device and storage medium for generating speech feature descriptions. Background Art
[0002] With the development of artificial intelligence, voice technology is increasingly being used in everyday life. For example, in online meetings, collected speech can be converted into text in real time and displayed to users. In the search field, voice input can be recognized to determine the user's search intent. Summary of the Invention
[0003] This application provides a method, device, electronic device, and storage medium for generating speech feature descriptions. The specific solution is as follows:
[0004] According to one aspect of the present application, a method for generating a speech feature description is provided, comprising:
[0005] Obtain target voice data;
[0006] Recognizing the target voice data to obtain multiple voice attribute labels of the target voice data;
[0007] Determining a target speech attribute label from the multiple speech attribute labels;
[0008] According to the target speech attribute label, natural language description information corresponding to the target speech data is generated; wherein the natural language description information is used to describe the speech characteristics of the target speech data
[0009] According to another aspect of the present application, there is provided a device for generating a speech feature description, comprising:
[0010] A first acquisition module, configured to acquire target voice data;
[0011] A recognition module, configured to recognize the target speech data to obtain a plurality of speech attribute labels of the target speech data;
[0012] A first determining module, configured to determine a target speech attribute label from the plurality of speech attribute labels;
[0013] The first generating module is used to generate natural language description information corresponding to the target speech data according to the target speech attribute label; wherein the natural language description information is used to describe the speech features of the target speech data.
[0014] According to another aspect of the present application, an electronic device is provided, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiment.
[0018] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to the above embodiment.
[0019] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.
[0022] Figure 1 A flowchart of a method for generating a speech feature description according to an embodiment of the present application is provided;
[0023] Figure 2 A flowchart of a method for generating a speech feature description provided in another embodiment of the present application;
[0024] Figure 3 A flowchart of a method for generating a speech feature description provided in another embodiment of the present application;
[0025] Figure 4 A flowchart of a method for generating a speech feature description provided in another embodiment of the present application;
[0026] Figure 5 A schematic diagram of the structure of a device for generating speech feature descriptions provided in one embodiment of the present application;
[0027] Figure 6 This is a block diagram of an electronic device used to implement the method for generating a speech feature description according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.
[0030] The following describes the method, device, electronic device, and storage medium for generating speech feature descriptions according to embodiments of the present application with reference to the accompanying drawings.
[0031] In some scenarios, it may be necessary to extract speaker information, voice information, emotional information, etc. from speech data and describe the speech features of the speech data based on this information. Based on this, the present invention proposes a method for generating speech feature description.
[0032] Figure 1 A flowchart of a method for generating a speech feature description according to an embodiment of the present application is provided.
[0033] The method for generating a speech feature description according to an embodiment of the present application may be performed by a device for generating a speech feature description according to an embodiment of the present application, and the device may be configured in an electronic device.
[0034] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.
[0035] like Figure 1 As shown, the method for generating the speech feature description includes:
[0036] Step 101: Acquire target voice data.
[0037] In the present application, the target voice data can be any voice data, for example, the target voice data can be voice data collected in real time by an audio collection device, or it can be audio data extracted from audio and video data, etc., or it can be voice data obtained by preprocessing the original voice data, or voice data obtained by other means, without limitation.
[0038] As an example, voice activity detection can be performed on the original speech data to remove silent segments and obtain the target speech data. Thus, by removing silent segments from the speech data through voice activity detection, the influence of silence on the calculation of global volume, pitch, speaking rate, etc. can be avoided.
[0039] Step 102: Recognize the target speech data to obtain multiple speech attribute labels of the target speech data.
[0040] Among them, speech attribute tags can be used to identify the speaker's age, gender, accent, speaking speed, pitch, volume, voice fluctuation, emotion, etc.
[0041] For example, the speech attribute tags may include but are not limited to age tags, gender tags, accent tags, speaking speed tags, pitch tags, volume tags, emotion tags, etc.
[0042] For example, age labels may include children, teenagers, youth, middle-aged, elderly, etc., and speech speed labels may include slow speech speed, normal speech speed, fast speech speed, etc.
[0043] For example, the multiple voice attribute labels of the target voice data may include youth, high volume, fast speaking speed, intense emotion, etc.
[0044] In this application, for different speech attributes, a recognition strategy matching the speech attribute can be adopted according to the target speech data to determine the speech attribute label.
[0045] As an example, a pre-trained speech representation learning model can be used to recognize target speech data to determine speech attribute labels for identifying the speaker's age, gender, accent, etc.
[0046] As an example, a speech rate value may be calculated based on text obtained by performing speech recognition on target speech data, and a corresponding speech tag may be determined according to the calculated speech rate value.
[0047] For example, the speaking rate value can be determined based on the ratio between the total number of phonemes and the duration of speech, and the speaking rate label of the target speech data can be determined based on the speaking rate value and the speaking rate range corresponding to each speaking rate label. For example, the speaking rate label corresponding to a speaking rate value less than 40% is slow speaking rate, the speaking rate label corresponding to a speaking rate value less than 40% is slow speaking rate, the speaking rate label corresponding to a speaking rate value greater than or equal to 40% and less than or equal to 60% is normal speaking rate, etc.
[0048] As an example, the volume value of the target voice data may be calculated, and a volume label, such as low volume, normal volume, high volume, etc., may be determined based on the volume value.
[0049] As an example, the pitch value of the target speech data may be calculated, and a pitch label, such as low pitch, normal pitch, high pitch, etc., may be determined based on the pitch value and pitch data of different genders.
[0050] As an example, the degree of sound fluctuation may be determined based on the pitch value variance.
[0051] As an example, a pre-trained emotion recognition model may be used to recognize target speech data to determine an emotion label corresponding to the target speech data.
[0052] Step 103: Determine a target speech attribute label from the multiple speech attribute labels.
[0053] In order to improve information focus, in this application, multiple voice attribute tags can be screened and a target voice attribute tag can be determined from the multiple voice attribute tags.
[0054] The target speech attribute label may be one or more, which is not limited in this application.
[0055] As a possible implementation, a target speech attribute tag may be determined from multiple speech attribute tags based on attribute information of the speech attribute tags, wherein the attribute information of the speech attribute tag may be data for describing characteristics of the speech attribute tag.
[0056] Exemplarily, the attribute information of the speech attribute label may include, but is not limited to, a sampling probability corresponding to the speech attribute label, a quantization value corresponding to the speech attribute label, whether the speech attribute label is a neutral attribute label, and the like.
[0057] The sampling probability corresponding to the speech attribute label may refer to the probability of the speech recognition label being selected.
[0058] For the quantized value corresponding to the speech attribute label, for example, the quantized value corresponding to the speech rate label refers to the speech rate value, and the quantized value corresponding to the volume label refers to the volume value, etc.
[0059] Among them, neutral attribute labels refer to non-biased attribute labels, such as normal speaking speed, normal volume, normal pitch, etc., which are all neutral attribute labels.
[0060] Exemplarily, other speech attribute tags except the neutral attribute tag among the multiple speech attribute tags may be determined as target attribute tags.
[0061] As another possible implementation, a requirement for generating a speech feature description corresponding to the target speech data can be obtained. Based on this requirement, a target speech attribute label can be determined from multiple speech attribute labels. The requirement for generating a speech feature description can be user-input or determined based on application scenario requirements, and this is not limited to this.
[0062] For example, if the generation requirement of speech feature description is to describe the speaker's accent, speaking speed and emotion, then the speech attribute label associated with accent, speaking speed and emotion can be determined from multiple speech attribute labels as the target speech attribute label.
[0063] Step 104: Generate natural language description information corresponding to the target speech data according to the target speech attribute label.
[0064] The natural language description information may be used to describe the speech characteristics of the target speech data, such as volume characteristics, speech speed characteristics, timbre characteristics, and emotion characteristics.
[0065] For example, the natural language description information may be in text form or voice form, which is not limited.
[0066] For example, a large model can be used to generate natural language description information based on the target speech attribute label.
[0067] As an example, prompt information can be generated based on the target speech attribute labels and feature description requirements, and the prompt information can be processed using a large model to generate natural language description information. The prompt information instructs the large model to generate a natural language description for describing speech features. Feature description requirements may include, but are not limited to, word count requirements and output format requirements for the natural language description information.
[0068] For example, the target speech attribute labels include high volume and fast speaking speed, and the generated natural language description information may be "the speaker's volume is high and the speaking speed is fast."
[0069] For example, after the natural language description information is generated, it can be displayed or pushed to the corresponding associated objects.
[0070] The method for generating speech feature description in the embodiment of the present application can be applied in various scenarios.
[0071] As an example, it can be applied in customer service scenarios. The natural language description information of the generated voice can be used to accurately assess the user's emotional state and provide personalized services.
[0072] As an example, it can be applied in intelligent education scenarios. For example, teachers can understand students' emotional fluctuations, changes in speech speed, etc. through the natural language description information of students' voices, and provide teachers with personalized teaching feedback.
[0073] In an embodiment of the present application, by identifying the target voice data, obtaining multiple voice recognition tags, and determining the target voice attribute tag from the multiple voice attribute tags, the voice attribute tags are screened, and based on the screened target voice attribute tag, natural language description information is generated to describe the voice characteristics of the target voice data, which can reduce redundant descriptions and improve the accuracy and conciseness of the description.
[0074] Figure 2 A flowchart of a method for generating speech feature description provided in another embodiment of the present application.
[0075] like Figure 2 As shown, the method for generating the speech feature description includes:
[0076] Step 201: Acquire target voice data.
[0077] Step 202: Recognize the target speech data to obtain multiple speech attribute labels of the target speech data.
[0078] In this application, steps 201-202 can be implemented in any of the embodiments of this application, so they will not be repeated here.
[0079] Step 203: Determine a key speech attribute tag from the multiple speech attribute tags.
[0080] The key speech attribute tags may be understood as important attribute tags used to describe and distinguish speech features, or the key speech attribute tags may be understood as attribute tags used to identify key speech features.
[0081] For example, key speech attribute tags may include but are not limited to fast speaking speed, high volume, high pitch, happy, angry, etc.
[0082] As a possible implementation manner, if the quantization value corresponding to the second speech attribute in the plurality of speech attribute labels is greater than the corresponding first threshold, the second speech attribute label may be determined to be a key speech attribute label.
[0083] For example, the second voice attribute label is high volume, and the corresponding quantization value is 100dB, which is greater than the preset value of 90dB. It can be determined that high volume is the key voice attribute label.
[0084] It is understandable that different second voice attribute labels may correspond to different first thresholds, and this is not limited to this.
[0085] Thus, based on the magnitude of the quantized value corresponding to the voice tag attribute and the corresponding first threshold, the key voice attribute tags are screened. The method is simple, and natural language description information is generated based on the screened voice attribute tags, which can improve the naturalness and information focus of the natural language description information. Moreover, the threshold can be adjusted to meet different screening requirements.
[0086] As another possible implementation, by determining whether the speech attribute tag is an enhanced attribute tag, if the third speech attribute tag among the multiple speech attribute tags is an enhanced attribute tag, the third speech attribute tag may be determined as a key speech attribute tag.
[0087] The enhanced attribute tag may refer to a speech attribute tag that needs to be emphasized. For example, the enhanced attribute tag may be pre-set.
[0088] For example, the enhanced attribute tags may include various emotion tags, low volume, high volume, fast speaking speed, low pitch, high pitch, etc.
[0089] Therefore, by determining the speech attribute tags that belong to the enhanced attribute tags as key speech attribute tags, the speech attribute tags that need to be paid attention to can be screened out, non-significant features can be filtered out, and natural language description information can be generated based on the screened speech attribute tags, so that the key speech features in the target speech data can be retained, and the naturalness and information focus of the natural language description information can be improved.
[0090] It should be noted that any one or both of the above methods can be used to screen key speech attribute features, and there is no limitation to this.
[0091] Step 204: Determine the target speech attribute label according to the key speech attribute label.
[0092] Exemplarily, the key speech attribute label may be directly determined as the target speech attribute label.
[0093] For example, the target speech attribute labels can be determined based on the historical occurrence rates of the key speech attribute labels. For example, the first preset number of speech attribute labels with the highest historical occurrence rates among the key speech attribute labels can be determined as the target speech attribute labels. This can improve the accuracy of the target speech attribute labels.
[0094] Step 205: Generate natural language description information corresponding to the target speech data according to the target speech attribute label.
[0095] In this application, step 205 can be implemented by referring to any one of the embodiments of this application, so it will not be repeated here.
[0096] In an embodiment of the present application, by determining a key speech attribute tag from multiple speech attribute tags, determining a target speech attribute tag based on the key speech attribute tag, and generating natural language description information based on the target speech attribute tag, the natural language description information can include key speech features, thereby improving the information focus and accuracy of the description.
[0097] Figure 3 A flowchart of a method for generating speech feature description provided in another embodiment of the present application.
[0098] like Figure 3 As shown, the method for generating the speech feature description includes:
[0099] Step 301: Acquire target voice data.
[0100] Step 302: Recognize the target speech data to obtain multiple speech attribute labels of the target speech data.
[0101] In this application, steps 301-302 can be implemented in any of the embodiments of this application, so they will not be repeated here.
[0102] Step 303: Obtain the sampling probability corresponding to the speech attribute label.
[0103] In the present application, each speech attribute label has a corresponding sampling probability, and the sampling probabilities of different speech attribute labels may be the same or different, and there is no limitation on this.
[0104] For example, the sampling probability corresponding to the speech attribute tag can be determined based on the frequency of occurrence of the speech attribute tag in historical data. For example, the higher the frequency of occurrence, the higher the sampling probability. The frequency of occurrence in historical data can refer to the frequency of occurrence of the speech attribute tag in previously generated natural language description information for speech.
[0105] It should be noted that the sampling probability corresponding to each speech attribute label can be fixed or updated, and there is no limitation on this.
[0106] Since multiple speech attribute labels may include neutral attribute labels, such as normal speaking speed and normal volume, in order to reduce the probability of neutral attribute labels being selected, based on this, illustratively, a neutral attribute label of multiple speech attribute labels is determined. If a first speech attribute label of the multiple speech attribute labels is a neutral attribute label, the initial sampling probability of the first speech attribute label can be updated according to the quantized value corresponding to the first speech attribute label to obtain a sampling probability.
[0107] Among them, the updated sampling probability is less than the initial sampling probability.
[0108] The quantized value corresponding to the first speech attribute label may refer to the quantized value of the speech attribute to which the first speech attribute label belongs.
[0109] Exemplarily, the sampling probability p of the first speech attribute can be calculated using the following formula (1):
[0110]
[0111] Here, k represents the adjustment coefficient, and x represents the quantized value corresponding to the first speech attribute.
[0112] Assume that the speaking rate value is between [40%, 60%], and the speaking rate label is normal. For example, if the speaking rate value of the target speech data is 45%, then the corresponding speaking rate label is normal. Normal speaking rate belongs to a neutral attribute label, and the quantization value corresponding to normal speaking rate is 45%. Then, the sampling probability corresponding to normal speaking rate can be calculated based on the above formula (1), and the initial sampling probability corresponding to normal speaking rate is updated to the calculated sampling probability to replace the initial sampling probability.
[0113] Therefore, for the neutral attribute label among multiple speech attribute labels, the sampling probability can be updated to be smaller than the initial sampling probability according to the quantized value corresponding to the neutral attribute label, thereby reducing the probability of the neutral attribute label being selected, thereby filtering out insignificant speech features and increasing the attention to significant speech features, thereby improving the naturalness and information focus of the natural language description.
[0114] Step 304: Sample multiple speech attribute labels according to the sampling probability to obtain a target speech attribute label.
[0115] In the present application, multiple speech attribute labels may be sampled according to the sampling probabilities and sampling quantities of the multiple speech attribute labels of the target speech data to obtain the target speech attribute label.
[0116] For example, the target speech data has 7 speech attribute labels and the number of samples is 2. Then, based on the sampling probabilities of the 7 speech attribute labels, 2 speech attribute labels can be selected to generate a natural language description.
[0117] It should be noted that the number of samples can be set or determined according to actual needs and is not limited thereto.
[0118] Step 305: Generate natural language description information corresponding to the target speech data according to the target speech attribute label.
[0119] In this application, step 305 can be implemented in any of the embodiments of this application, so it will not be described in detail here.
[0120] In an embodiment of the present application, multiple speech attribute labels are sampled based on the sampling probabilities of the multiple speech attribute labels to obtain a target attribute label, so that the target speech attribute label can be located more quickly, thereby improving the efficiency of generating natural language description information.
[0121] Figure 4 A flowchart of a method for generating speech feature description provided in another embodiment of the present application.
[0122] like Figure 4 As shown, the method for generating the speech feature description includes:
[0123] Step 401: Acquire target voice data.
[0124] Step 402: Recognize the target speech data to obtain multiple speech attribute labels of the target speech data.
[0125] Step 403: Determine a target speech attribute label from the multiple speech attribute labels.
[0126] In this application, steps 401-403 can be implemented by referring to any one of the embodiments of this application, so they will not be repeated here.
[0127] Step 404: Determine a target description template based on the target speech attribute label.
[0128] The description template may refer to a template for describing speech features. For example, the description template may be complete natural language description information or a natural language description with slot marks, which is not limited thereto.
[0129] For example, a description template is "the speaker's speaking speed is [speaking speed], and the emotion is [emotion]". This template includes two slots: [speaking speed] and [emotion].
[0130] As a possible implementation method, at least one template category is set, each template category has a corresponding description template, and the target template category can be determined from at least one template category based on the target speech attribute label, and then the target description template can be determined from the description templates corresponding to the target template category.
[0131] Exemplarily, the target speech attribute label may be matched with the description information of each template category to determine the target template category that matches the target speech attribute label.
[0132] Exemplarily, the various speech attributes can be combined to obtain multiple speech attribute combinations, each speech attribute combination can be used as a template category, and the target speech attribute combination can be determined based on the target speech attribute label, and the template category corresponding to the target speech attribute combination can be used as the target template category.
[0133] For example, speech attribute combinations may include speaking speed-emotion, speaking speed-emotion-age, pitch-speaking speed-emotion, etc., and the target speech attribute labels include fast speaking speed and intense emotion. Then the corresponding speech attribute combination is speaking speed-emotion, and the category corresponding to speaking speed-emotion is the target template category.
[0134] Exemplarily, one or several templates can be randomly selected from the description templates corresponding to the target template category as the target description templates, or the description templates can be sampled according to the sampling probability of the description templates corresponding to the target template category to determine the target description template, or a preset number of templates with the highest historical usage rates can be selected as the target description templates based on the historical usage rates of the description templates corresponding to the target template category.
[0135] Therefore, by selecting the target template category according to the target speech attribute label, and then determining the target description template from the description templates corresponding to the target template category, the accuracy of the selected description template can be improved, and thus the accuracy of the natural language description information can be improved.
[0136] Optionally, for each voice attribute combination, a large model can be used to generate a description template under a template category corresponding to the voice attribute combination, thereby increasing the diversity of the description templates.
[0137] Step 405: Generate natural language description information according to the target description template.
[0138] In this application, a natural language description can be generated based on the target description template and the target speech attribute label.
[0139] As a possible implementation method, the target speech attribute labels can be used to fill in the slots of the target description template to obtain natural language description information.
[0140] As another possible implementation method, a synonym library for each voice attribute tag can be pre-built, and the target synonym of the target voice attribute tag can be determined from the synonym library corresponding to the target voice attribute tag. According to the target synonym, the slots of the target description template can be filled to obtain natural language description information.
[0141] For example, the speech attribute label is high pitch, and the corresponding synonym library is {sharp, high-pitched, harsh}. If the target speech attribute label includes high pitch, you can select a synonym word from the synonym library, such as sharp, to generate a natural language description.
[0142] For example, the speech attribute label is fast speaking speed, and the corresponding synonym library is {like a machine gun, like a machine gun, and endlessly talking}, etc. If the target speech attribute label includes fast speaking speed, you can select a synonym word from the synonym library, such as like a machine gun, to generate a natural language description.
[0143] For example, a synonym can be randomly selected from the synonym library corresponding to the target speech attribute tag as the target synonym for generating the natural language description information, or the target synonym can be obtained by sampling based on the sampling probability of each synonym in the synonym library. The sampling probability of the synonym can be fixed or updated based on the historical usage frequency of the synonym, which is not limited to this.
[0144] Therefore, target synonyms are selected from the synonym library corresponding to the target speech attribute label to generate natural language description information, thereby reducing the repeated use of the same words, making the description more diverse and vivid, and improving the diversity and flexibility of the description.
[0145] For example, for each speech attribute tag, a large model can be used to generate synonyms corresponding to the speech attribute tag, and a synonym library can be constructed based on the generated synonyms, thereby expanding the coverage of the synonym library and meeting different description requirements.
[0146] For example, for different speaking speed labels, such as slow speaking speed, fast speaking speed, normal speaking speed, etc., corresponding synonym libraries can be constructed to improve the delicacy and diversity of the description.
[0147] Since some attribute combinations occur frequently, repeated calls to description templates may occur. Based on this, for example, historical speech feature description data can be obtained, and the co-occurrence frequency of each speech attribute combination can be determined based on the historical speech feature description data. For speech attribute combinations with a co-occurrence frequency greater than a second threshold, a description template of the template category corresponding to the speech attribute combination can be generated based on the co-occurrence frequency.
[0148] The historical speech feature description data may refer to previously generated natural language description information used to describe speech features of speech data.
[0149] For example, the number of templates can be determined based on the co-occurrence frequency, and then a corresponding number of difference description templates can be generated based on the speech attribute combination, thereby avoiding repeated template calls for high-frequency attribute combinations.
[0150] For example, the co-occurrence frequency of speech rate and emotion is the second threshold, then the number of templates of the template category corresponding to this voice attribute combination is N = log (co-occurrence frequency) × cardinality, and a large model can be used to generate N differentiated description templates.
[0151] Therefore, for speech attribute combinations with relatively high co-occurrence frequencies, description templates of the template categories corresponding to the speech attribute combinations can be generated based on the co-occurrence frequencies, thereby reducing the probability of repeatedly calling description templates for high-frequency attribute combinations and enriching the diversity of descriptions.
[0152] In an embodiment of the present application, by determining a target description template based on a target speech attribute label, the accuracy of selecting a target description template can be improved, and natural language description information can be generated based on the target description template with high generation efficiency.
[0153] In order to implement the above embodiment, the embodiment of the present application also proposes a device for generating a speech feature description. Figure 5 A schematic diagram of the structure of a device for generating speech feature descriptions provided in one embodiment of the present application.
[0154] like Figure 5 As shown, the speech feature description generating device 500 includes:
[0155] A first acquisition module 510 is used to acquire target voice data;
[0156] The recognition module 520 is configured to recognize the target speech data to obtain a plurality of speech attribute labels of the target speech data;
[0157] A first determining module 530 is configured to determine a target speech attribute label from the plurality of speech attribute labels;
[0158] The first generating module 540 is configured to generate natural language description information corresponding to the target speech data according to the target speech attribute label; wherein the natural language description information is used to describe speech features of the target speech data.
[0159] Optionally, the first determining module 530 is configured to:
[0160] Determining a key speech attribute tag from the multiple speech attribute tags;
[0161] The target speech attribute label is determined according to the key speech attribute label.
[0162] Optionally, the first determining module 530 is configured to perform at least one of the following:
[0163] In response to a quantization value corresponding to a second speech attribute in the plurality of speech attribute labels being greater than a corresponding first threshold, determining the second speech attribute label as the key speech attribute label;
[0164] In response to a third voice attribute tag among the plurality of voice attribute tags being an enhanced attribute tag, the third voice attribute tag is determined to be the key voice attribute tag.
[0165] Optionally, the first determining module 530 is configured to:
[0166] Obtaining the sampling probability corresponding to the speech attribute label;
[0167] The plurality of speech attribute labels are sampled according to the sampling probability to obtain the target speech attribute label.
[0168] Optionally, the first determining module 530 is configured to:
[0169] In response to determining that a first speech attribute label of the plurality of speech attribute labels is a neutral attribute label, updating an initial sampling probability of the first speech attribute label according to a quantization value corresponding to the first speech attribute label to obtain the sampling probability;
[0170] The sampling probability is smaller than the initial sampling probability.
[0171] Optionally, the first generating module 540 is configured to:
[0172] Determining a target description template according to the target speech attribute label;
[0173] The natural language description information is generated according to the target description template.
[0174] Optionally, the first generating module 540 is configured to:
[0175] Determining a target synonym for the target speech attribute tag from a synonym library corresponding to the target speech attribute tag;
[0176] The target description template is slot-filled according to the target synonyms to obtain the natural language description information.
[0177] Optionally, the first generating module 540 is configured to:
[0178] Determining a target template category from at least one template category according to the target speech attribute label;
[0179] The target description template is determined from the description templates corresponding to the target template category.
[0180] Optionally, the device may further include:
[0181] The second acquisition module is used to obtain historical speech feature description data;
[0182] a second determining module, configured to determine the co-occurrence frequency of each speech attribute combination based on the historical speech feature description data;
[0183] The second generating module is configured to generate, in response to a co-occurrence frequency of any speech attribute combination being greater than a second threshold, a description template of a template category corresponding to the any speech attribute combination according to the co-occurrence frequency.
[0184] It should be noted that the explanation of the above-mentioned embodiment of the method for generating speech feature description is also applicable to the device for generating speech feature description of this embodiment, so it will not be repeated here.
[0185] In an embodiment of the present application, by identifying the target voice data, obtaining multiple voice recognition tags, and determining the target voice attribute tag from the multiple voice attribute tags, the voice attribute tags are screened, and based on the screened target voice attribute tag, natural language description information is generated to describe the voice characteristics of the target voice data, which can reduce redundant descriptions and improve the accuracy and conciseness of the description.
[0186] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0187] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0188] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.
[0189] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0190] The computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the method for generating a speech feature description. For example, in some embodiments, the method for generating a speech feature description can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for generating a speech feature description described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the method for generating a speech feature description in any other appropriate manner (eg, by means of firmware).
[0191] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0192] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0193] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0194] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0195] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0196] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0197] According to an embodiment of the present application, the present application also provides a computer program product, which, when an instruction processor in the computer program product executes, executes the method for generating speech feature description proposed in the above embodiment of the present application.
[0198] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.
[0199] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A method for generating a speech feature description, comprising: Obtain target voice data; Recognizing the target voice data to obtain multiple voice attribute labels of the target voice data; Determining a target speech attribute label from the multiple speech attribute labels; According to the target speech attribute label, natural language description information corresponding to the target speech data is generated; wherein the natural language description information is used to describe the speech features of the target speech data.
2. The method according to claim 1, wherein Determining a target speech attribute label from the multiple speech attribute labels includes: Determining a key speech attribute tag from the multiple speech attribute tags; The target speech attribute label is determined according to the key speech attribute label.
3. The method according to claim 2, wherein: Determining a key speech attribute tag from the multiple speech attribute tags includes at least one of the following: In response to a quantization value corresponding to a second speech attribute in the plurality of speech attribute labels being greater than a corresponding first threshold, determining the second speech attribute label as the key speech attribute label; In response to a third voice attribute tag among the plurality of voice attribute tags being an enhanced attribute tag, the third voice attribute tag is determined to be the key voice attribute tag.
4. The method according to claim 1, wherein Determining a target speech attribute label from the multiple speech attribute labels includes: Obtaining the sampling probability corresponding to the speech attribute label; The plurality of speech attribute labels are sampled according to the sampling probability to obtain the target speech attribute label.
5. The method according to claim 4, wherein: The obtaining of the sampling probability corresponding to the speech attribute includes: In response to determining that a first speech attribute label of the plurality of speech attribute labels is a neutral attribute label, updating an initial sampling probability of the first speech attribute label according to a quantization value corresponding to the first speech attribute label to obtain the sampling probability; The sampling probability is smaller than the initial sampling probability.
6. The method of claim 1, wherein: Generating natural language description information corresponding to the target speech according to the target speech attribute label includes: Determining a target description template according to the target speech attribute label; The natural language description information is generated according to the target description template.
7. The method according to claim 6, wherein: Generating the natural language description information according to the target description template includes: Determining a target synonym for the target speech attribute tag from a synonym library corresponding to the target speech attribute tag; The target description template is slot-filled according to the target synonyms to obtain the natural language description information.
8. The method of claim 6, wherein: The step of determining a target description template according to the target speech attribute tag includes: Determining a target template category from at least one template category according to the target speech attribute label; The target description template is determined from the description templates corresponding to the target template category.
9. The method of claim 8, further comprising: Obtain historical speech feature description data; determining the co-occurrence frequency of each speech attribute combination based on the historical speech feature description data; In response to the co-occurrence frequency of any speech attribute combination being greater than a second threshold, a description template of a template category corresponding to the any speech attribute combination is generated according to the co-occurrence frequency.
10. A device for generating a speech feature description, comprising: A first acquisition module, configured to acquire target voice data; A recognition module, configured to recognize the target speech data to obtain a plurality of speech attribute labels of the target speech data; A first determining module, configured to determine a target speech attribute label from the plurality of speech attribute labels; The first generating module is used to generate natural language description information corresponding to the target speech data according to the target speech attribute label; wherein the natural language description information is used to describe the speech features of the target speech data.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.