Text generation method and device, computer equipment and storage medium

The audio representation and large language model are extracted through the audio Transformer model to generate text embedding and feature alignment, which solves the problem of insufficient multimodal information fusion in the prior art, and achieves the generation of more accurate and diverse audio description text.

CN120067388APending Publication Date: 2025-05-30PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146101.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing automatic audio description system has shortcomings in multimodal information fusion and cannot comprehensively consider multiple information sources, resulting in the generated description text being incomplete and accurate enough.

Method used

By obtaining the target audio, the audio representation is extracted using the audio Transformer model; the prompt text is obtained, and the word segmentation is used to generate text embeddings; the audio representation is downsampled and aligned with the text embedding; the aligned features are decoded using the large language model to generate description text.

Benefits of technology

It realizes the generation of more diverse, accurate and realistic audio description text, improving the performance and user experience of the automatic audio description system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067388A_ABST
    Figure CN120067388A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and the field of financial science and technology and medical health, and discloses a text generation method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a target audio, and extracting the audio representation of the target audio through an audio Transform model; obtaining a prompt text of the target audio, and performing word segmentation processing on the prompt text by using a large language model to generate text embedding; down-sampling the audio representation, and aligning the down-sampling audio representation with the text embedding; decoding the aligned audio representation and text embedding by using the large language model to generate a description text of the target audio; therefore, more diversified, accurate and real audio description texts can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, as well as the fields of fintech and healthcare, and particularly relates to a text generation method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] Currently, with the rapid development of artificial intelligence technology, the automatic audio description technology has gradually become a research hotspot. Automatic audio description aims to generate natural, accurate, and expressive text descriptions for the input audio signals. This technology has broad application prospects in multiple fields, such as voice assistants, virtual anchors, voice customer service, voice translation, and speech synthesis in the fintech field, and application scenarios such as intelligent medical customer service in the healthcare field. In these application scenarios, the system needs to be able to accurately understand the content of the audio information in order to finally generate coherent and fluent descriptive texts to meet the users' needs for information acquisition and interaction. However, in the healthcare field, for example, one of the main problems faced by current automatic audio description systems is the lack of multi-modal information fusion, unable to comprehensively consider multiple information sources, and the generated descriptive texts may not be comprehensive and accurate enough. For example, when describing a patient's condition, it may only be based on superficial audio information and not combined with other hint information such as imaging examination results, resulting in incomplete and inaccurate descriptions. That is, in the prior art, the text descriptions generated by current automatic audio description systems are often not diverse enough and are also difficult to fully capture the fine-grained content of the audio, resulting in the lack of depth and details in the generated text descriptions and unable to meet the users' needs for high-quality audio descriptions.

[0003] Based on this, how to provide a text generation method, apparatus, computer device, and computer-readable storage medium that can generate more diverse, accurate, and realistic audio description texts is an urgent problem to be solved by those skilled in the art currently. Summary of the Invention

[0004] In view of the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a text generation method, apparatus, computer device, and computer-readable storage medium, aiming to solve the problem of how to generate more diverse, accurate, and realistic audio description texts.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] In the first aspect, the present invention provides a text generation method, which includes:

[0007] Obtain a target audio, and use an audio Transformer model to extract the audio representation of the target audio;

[0008] Obtain the prompt text of the target audio, and use a large language model to perform word segmentation on the prompt text to generate a text embedding;

[0009] Downsample the audio representation and align it with the text embedding;

[0010] Use the large language model to decode the aligned audio representation and the text embedding to generate a description text of the target audio.

[0011] In a second aspect, the present invention provides a text generation device, which includes:

[0012] An extraction module, configured to obtain a target audio and extract an audio representation of the target audio by using an audio Transformer model;

[0013] A processing module, configured to obtain the prompt text of the target audio, and use a large language model to perform word segmentation on the prompt text to generate a text embedding;

[0014] A downsampling module, configured to downsample the audio representation and align it with the text embedding;

[0015] A decoding module, configured to use the large language model to decode the aligned audio representation and the text embedding to generate a description text of the target audio.

[0016] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned text generation method is implemented.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above-mentioned text generation method is implemented.

[0018] Compared with the prior art, the present invention provides a text generation method, device, computer device, and computer-readable storage medium. By obtaining a target audio, extracting an audio representation of the target audio by using an audio Transformer model; obtaining the prompt text of the target audio, and using a large language model to perform word segmentation on the prompt text to generate a text embedding; downsampling the audio representation and aligning it with the text embedding; using the large language model to decode the aligned audio representation and the text embedding to generate a description text of the target audio; thus, the present invention can generate more diverse, accurate, and real audio description texts. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 Schematic diagram of the application environment of a text generation method provided by an embodiment of the present invention.

[0021] Figure 2 Schematic flow diagram of a text generation method provided by an embodiment of the present invention.

[0022] Figure 3 Schematic diagram of the program modules of a text generation device provided by an embodiment of the present invention.

[0023] Figure 4 Schematic diagram of the structure of a computer device provided by an embodiment of the present invention.

[0024] Figure 5 Another schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0026] It should be understood that when used in the specification of the present invention and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0027] It should also be understood that the term " / and" as used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0028] As used in the specification of the present invention and the appended claims, the term "if" may be construed contextually as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be construed contextually to mean "once determined" or "in response to determining" or "once [described condition or event] is detected" or "in response to detecting [described condition or event]".

[0029] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0030] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present invention means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0031] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or posterior. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0032] In order to illustrate the technical solution of the present invention, specific embodiments will be used for illustration below.

[0033] A text generation method provided by an embodiment of the present invention can be applied, for example, in Figure 1In the application environment shown, the client communicates with the server through a network. Among them, the client includes but is not limited to computer devices such as a personal digital assistant (PDA), a tablet computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, and a personal digital assistant (PDA). The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0034] Please refer to Figure 2 , an embodiment of the present invention provides a text generation method, where the method includes the following steps:

[0035] S100. Obtain a target audio, and use an audio Transformer model to extract the audio representation of the target audio;

[0036] S200. Obtain the prompt text of the target audio, and use a large language model to perform word segmentation on the prompt text to generate a text embedding;

[0037] S300. Downsample the audio representation and align it with the text embedding;

[0038] S400. Use the large language model to decode the aligned audio representation and the text embedding to generate the description text of the target audio.

[0039] Specifically, when implemented, the text generation method of this embodiment achieves the technical effect of generating more diverse, accurate, and realistic audio description texts through a series of carefully designed steps.

[0040] First, in step S100, an audio Transformer model is used to extract the audio representation of the target audio. This model can capture fine-grained features in the audio, such as pitch, rhythm, timbre, etc., providing a rich audio information basis for subsequent text generation.

[0041] Next, in step S200, the large language model performs word segmentation on the prompt text and generates a text embedding. This step can not only understand the semantics of the prompt text but also convert the text information into an embedding form that can be processed by subsequent models, enabling the fusion of audio and text information in the same feature space.

[0042] Then, in step S300, the audio representation is downsampled and aligned with the text embedding. This crucial step ensures the synchronization of audio and text features in terms of time and space, enabling the subsequent model to comprehensively consider the context information of audio and text during the decoding process. This alignment operation helps the subsequent model better understand the relationship between the audio content and the text prompt, thereby generating descriptive text that more conforms to the actual audio content.

[0043] Finally, in step S400, a large language model is used to decode the aligned audio representation and text embedding to generate the descriptive text of the target audio. The powerful text generation ability of the large language model can generate natural, fluent, and diverse text descriptions based on the fused features. Since the large language model has learned a large number of language patterns and expressions during the training process, it can generate descriptive texts of various styles and details, not only improving the accuracy of the description but also enhancing the realism and expressiveness of the description.

[0044] In summary, through the fine-grained extraction of audio features, the semantic understanding of text embeddings, the context fusion of feature alignment, and the efficient decoding of the large language model in this embodiment, the method of this embodiment can generate more diverse, accurate, and realistic audio descriptive texts, effectively improving the performance and user experience of the automatic audio description system.

[0045] It can be understood that the text generation method provided by the embodiments of the present invention can be applied to text generation scenarios related to the fintech field. The following are some specific examples:

[0046] 1. Automatic generation of financial reports

[0047] Application scenario:

[0048] Financial institutions need to regularly generate various reports, such as quarterly financial reports, market analysis reports, risk assessment reports, etc. These reports usually contain a large amount of data and complex analysis content. Manual writing is time-consuming, laborious, and prone to errors. The text generation method of the present invention can automatically generate these reports, improving the writing efficiency and accuracy.

[0049] Specific implementation:

[0050] Obtain the target audio: Record or collect audio about financial data and analysis, such as analysts' explanations, recordings of market research meetings, etc.

[0051] Extract the audio representation: Use the audio Transformer model to extract the feature representation of the audio.

[0052] Obtain the prompt text: Provide relevant prompt text, such as the report title, key data indicators, analysis dimensions, etc.

[0053] Generate text embeddings: Use a large language model to tokenize the prompt text and generate text embeddings.

[0054] Feature alignment: Align the audio representation and the text embeddings in terms of features.

[0055] Generate descriptive text: Generate the text content of the financial report through the decoding of the large language model.

[0056] Example:

[0057] Target audio: An analyst's explanation of the quarterly market performance of a certain industry.

[0058] Prompt text: Industry name, quarter, key data indicators (such as market size, growth rate, performance of major enterprises, etc.).

[0059] Generated descriptive text:

[0060] In the third quarter of 2024, the market size of the fintech industry reached 500 billion yuan, a year-on-year increase of 20%. Among them, the payment field had the most significant growth, with a year-on-year increase of 25%, mainly due to the popularization of mobile payment and the application of innovative payment technologies. In terms of risk management, the overall risk level of the industry remained stable, and the non-performing loan ratio remained below 2%. In the future, with the further improvement of regulatory policies and continuous technological innovation, the fintech industry is expected to continue to grow rapidly.

[0061] 2. Generation of financial product recommendation text

[0062] Application scenarios:

[0063] Financial institutions need to recommend suitable financial products to customers, such as wealth management products, insurance products, funds, etc. Traditional recommendation methods rely on manual writing and explanation, which are inefficient and prone to information inconsistency. The text generation method of the present invention can automatically generate recommendation texts for financial products, improving the accuracy and consistency of recommendations.

[0064] Specific implementation:

[0065] Obtain the target audio: Record or collect the explanatory audio of financial products, such as the explanations of salespersons, the audio parts of product introduction videos, etc.

[0066] Extract the audio representation: Use an audio Transformer model to extract the feature representation of the audio.

[0067] Obtain the prompt text: Provide relevant prompt text, such as product name, product features, target population, etc.

[0068] Generate text embeddings: Use a large language model to tokenize the prompt text and generate text embeddings.

[0069] Feature alignment: Align the features of the audio representation and the text embedding.

[0070] Generate descriptive text: Generate recommended text for financial products through the decoding of a large language model.

[0071] Example:

[0072] Target audio: The explanation of the features and returns of a certain financial product by a salesperson.

[0073] Prompt text: Product name, product type, expected return, risk level, applicable population, etc.

[0074] Generated descriptive text:

[0075] We recommend that you pay attention to our "Stable Profit Financial Plan". This is a fixed-income product with medium to low risk, and the expected annualized return rate is 4.5%. It is suitable for investors with a lower risk appetite. The product term is 1 year, and the principal and interest will be repaid in a lump sum at maturity. The safety of the funds is guaranteed. Our professional team will provide you with full-service to ensure the maximization of your investment returns.

[0076] Through the above application, the text generation method of the present invention can significantly improve the efficiency and quality of information processing in the field of fintech, enhance the user experience, and has important practical application value.

[0077] It can be understood that the text generation method provided by the embodiments of the present invention can also be applied to text generation scenarios related to the field of medical and health. The following are some specific examples:

[0078] 1. Automatic generation of medical reports

[0079] Application scenario:

[0080] Hospitals and medical institutions need to generate a large number of medical reports, such as diagnostic reports, inspection reports, medical records, etc. These reports usually contain complex medical terms and detailed inspection results. Manual writing is time-consuming and laborious and prone to errors. The text generation method of the present invention can automatically generate these reports, improve the writing efficiency and accuracy, and at the same time reduce the workload of medical staff.

[0081] Specific implementation:

[0082] Obtain the target audio: Record or collect audio such as doctors' diagnostic explanations and oral reports of inspection results.

[0083] Extract the audio representation: Use the audio Transformer model to extract the feature representation of the audio.

[0084] Obtain the prompt text: Provide relevant prompt text, such as patient basic information, inspection items, preliminary diagnosis, etc.

[0085] Generate text embeddings: Use a large language model to tokenize the prompt text and generate text embeddings.

[0086] Feature alignment: Align the audio representation and the text embeddings.

[0087] Generate descriptive text: Generate the text content of the medical report through decoding by a large language model.

[0088] Example:

[0089] Target audio: A doctor's explanation of the electrocardiogram examination results of a certain patient.

[0090] Prompt text: Patient's name, gender, age, examination item (electrocardiogram).

[0091] Generated descriptive text:

[0092] Patient: Zhang San, male, 45 years old. The electrocardiogram examination results show normal sinus rhythm, heart rate of 78 beats per minute, PR interval of 0.16 seconds, QRS complex of 0.08 seconds, and QT interval of 0.36 seconds. No obvious ST segment elevation or depression is seen, and the T wave morphology is normal. Preliminary diagnosis: Normal electrocardiogram.

[0093] 2. Generation of health consultation replies

[0094] Application scenarios:

[0095] Online health consultation platforms need to answer users' health questions promptly and accurately. Traditional answering methods rely on manual writing, which is inefficient and prone to information inconsistency. The text generation method of the present invention can automatically generate reply texts for health consultations, improve the efficiency and quality of answering, and enhance the user experience.

[0096] Specific implementation:

[0097] Obtain the target audio: Record or collect the audio of the user's consultation question, such as the health question input by the user through voice.

[0098] Extract the audio representation: Use an audio Transformer model to extract the feature representation of the audio.

[0099] Obtain the prompt text: Provide relevant prompt text, such as keywords of the user's question, historical consultation records, etc.

[0100] Generate text embeddings: Use a large language model to tokenize the prompt text and generate text embeddings.

[0101] Feature alignment: Align the audio representation and the text embeddings.

[0102] Generate descriptive text: Generate a response text for health consultation through the decoding of a large language model.

[0103] Example:

[0104] Target audio: The user consults about the dietary precautions for hypertension.

[0105] Prompt text: Hypertension, diet, precautions.

[0106] Generated descriptive text:

[0107] Patients with hypertension should pay attention to the following points in their diet: 1) Low-salt diet, with a daily salt intake of no more than 5 grams; 2) Eat more fresh vegetables and fruits to supplement minerals such as potassium and magnesium; 3) Limit the intake of high-fat and high-cholesterol foods, such as animal offal, fried foods, etc.; 4) Appropriate intake of high-quality protein, such as fish, lean meat, soy products, etc.; 5) Control the total calorie intake and keep the weight within the normal range; 6) Avoid alcohol and smoking, as these habits will exacerbate the condition of hypertension. It is recommended that you regularly monitor your blood pressure, take medicine according to doctor's advice, and maintain good living habits.

[0108] Through the above application, the text generation method of the present invention can significantly improve the efficiency and quality of information processing in the field of medical health, enhance the user experience, and has important practical application value.

[0109] Furthermore, in one embodiment, the text generation method, wherein the step S100, obtaining the target audio, and using the audio Transformer model to extract the audio representation of the target audio, specifically includes the steps:

[0110] Obtain the audio to be processed, perform noise reduction and normalization processing on the audio to be processed to obtain the target audio;

[0111] Perform frame splitting and feature extraction processing on the target audio to obtain the Mel spectrogram features of the target audio;

[0112] Use the Mel spectrogram features as input, and generate the audio representation of the target audio through the pre-trained audio Transformer model.

[0113] In specific implementation, this embodiment realizes the generation of high-quality audio representation through a series of fine audio processing and feature extraction operations, providing a solid foundation for subsequent text generation. First, by obtaining the audio to be processed and performing noise reduction and normalization processing, the noise interference in the audio can be effectively removed, and the audio signal is adjusted to a uniform amplitude range, thereby improving the audio quality and ensuring the accuracy of subsequent processing. Then, the target audio is framed and feature extracted to obtain Mel spectrum features. This process converts the continuous audio signal into a spectrum feature suitable for model processing. The Mel spectrum feature can better simulate the human ear's perception characteristics of sounds of different frequencies, so that the model can more effectively capture the key information of the audio. Finally, the Mel spectrum feature is input into the pre-trained audio Transformer model to generate an audio representation of the target audio. With its powerful feature extraction capability, the audio Transformer model can further mine the deep semantic information in the audio, and the generated audio representation not only contains the surface features of the audio, but also contains rich information such as the semantics and emotions of the audio. These technical effects work together to make the generated audio representation more accurate, rich and representative, providing high-quality input for subsequent text generation tasks, and helping to generate more accurate and diverse audio description texts.

[0114] The specific implementation process of the steps in this embodiment is roughly as follows:

[0115] 1) Get the audio to be processed and preprocess it

[0116] 11) Audio Collection

[0117] Get the audio to be processed from the specified audio source. The audio source can be a locally stored audio file, a real-time audio stream (such as microphone input), or a network audio resource. For example, if the application scenario is a voice assistant, the audio to be processed may come from the real-time recording of the microphone of the user's device.

[0118] 12) Noise reduction

[0119] Apply a noise reduction algorithm to the acquired audio to be processed to remove background noise and interference signals. Common noise reduction methods include spectral subtraction, Wiener filtering, and noise reduction models based on deep learning. For example, a deep learning noise reduction model can be used to identify and remove noise components through a trained neural network, retaining the main information in the audio, such as human voice or instrument sound.

[0120] 13) Normalization

[0121] Normalize the denoised audio to adjust the amplitude range of the audio signal to meet the requirements of subsequent processing. Usually, the amplitude of the audio signal is normalized to the range of [-1, 1]. Normalization can be achieved by calculating the maximum absolute value of the audio signal and then dividing each sample value by this maximum value to ensure the amplitude consistency of the audio signal.

[0122] 2) Perform frame splitting and feature extraction on the target audio

[0123] 21) Audio frame splitting

[0124] Split the normalized target audio into a series of short-time frames. The purpose of frame splitting is to convert the continuous audio signal into small segments suitable for processing, facilitating the extraction of local features. Usually, the frame length is selected to be 20 - 40 milliseconds, and the frame shift is 10 - 20 milliseconds. For example, with parameter settings of a frame length of 25 milliseconds and a frame shift of 10 milliseconds, the audio is frame-split by means of a sliding window, and a window function (such as a Hamming window) can be used to reduce the discontinuity at the frame edges.

[0125] 22) Mel-spectrum feature extraction

[0126] Calculate the Mel-spectrum features for each audio frame. First, perform a fast Fourier transform (FFT) on each frame to obtain the spectrogram. Then, filter the spectrogram through a Mel filter bank to obtain the Mel spectrum. The Mel spectrum can better simulate the perception characteristics of the human ear for sounds of different frequencies. For example, for an audio with a sampling rate of 16 kHz, 128-dimensional Mel-spectrum features can be calculated. These Mel-spectrum features will be used as the input to the audio Transformer model.

[0127] 3) Generate audio representations using a pre-trained audio Transformer model

[0128] 31) Model input preparation

[0129] Use the extracted sequence of Mel-spectrum features as the input and prepare to input it into the pre-trained audio Transformer model, ensuring that the format and dimension of the Mel-spectrum features meet the input requirements of the model.

[0130] 32) Forward propagation to generate audio representations

[0131] The Mel spectrogram features are input into a pre-trained audio Transformer model for forward propagation. The audio Transformer model processes the Mel spectrogram features deeply through a multi-layer Transformer architecture, utilizing the self-attention mechanism and the feed-forward neural network. The self-attention mechanism can capture the long-range dependencies between audio frames and extract the global features of the audio; the feed-forward neural network performs non-linear transformations on the features to further extract the high-level semantic information of the audio. After being processed by multiple layers of Transformer, the model outputs the audio representation of the target audio. These audio representations can capture the rich semantic and emotional information of the audio and provide high-quality feature inputs for subsequent text generation tasks.

[0132] Through the above process, this embodiment can achieve the conversion from the original audio signal to a high-quality audio representation, laying a solid foundation for generating accurate and diverse audio description texts.

[0133] Further, in one embodiment, for the text generation method, in step S200, obtaining the prompt text of the target audio and performing word segmentation processing on the prompt text by using a large language model to generate text embeddings specifically includes the steps of:

[0134] Obtaining the prompt text of the target audio, organizing the content of the prompt text, and generating a target text;

[0135] Performing word segmentation processing on the target text by using the pre-trained large language model to generate the text embeddings.

[0136] In specific implementation, through content organization and word segmentation processing, this embodiment realizes the conversion from the original prompt text to high-quality text embeddings, providing accurate and semantic-rich text features for subsequent text generation tasks. First, obtaining the prompt text of the target audio and organizing the content can remove the noise and non-standard content in the text, generate a standard and clear target text, ensure the text quality, and provide a clean data basis for subsequent processing. Then, performing word segmentation processing on the target text by using the pre-trained large language model to generate text embeddings. The large language model has strong language understanding ability, can segment the text into meaningful tokens, and generate high-dimensional text embeddings for each token. These embeddings not only contain the semantic information of the tokens but also the semantic relationships and context information between the tokens. In this way, the generated text embeddings can accurately express the semantics of the prompt text, provide rich semantic features for the alignment of audio and text and subsequent decoding and generation, contribute to generating more accurate and diverse audio description texts, and improve the performance and output quality of the entire text generation system.

[0137] Among them, the specific implementation process of the steps in this embodiment is roughly as follows:

[0138] 1) Obtain the prompt text of the target audio

[0139] 11) Source of the prompt text

[0140] Determine the source of the prompt text. The prompt text can be directly input by the user. For example, in the application scenario of audio description generation, the user may input a text describing the content of the target audio, such as "This is a piece of lively jazz music". It can also be extracted from the metadata associated with the target audio, such as information like the tags and introductions of the audio file. If the target audio is the background music of a certain video, relevant prompt text can be obtained from the subtitles or descriptions of the video.

[0141] 2) Organize the content of the prompt text

[0142] 21) Text cleaning

[0143] Clean the obtained prompt text to remove irrelevant characters (such as extra spaces, punctuation marks, special characters, etc.). For example, replace multiple consecutive spaces in the text with a single space, and remove the spaces at the beginning and end of the text.

[0144] Check whether there are spelling mistakes or grammar mistakes in the text and correct them. Spelling check tools or grammar check tools can be used to automatically detect and correct these mistakes.

[0145] 22) Text normalization

[0146] Convert the prompt text into a unified format. For example, convert all text to lowercase to reduce lexical variations.

[0147] Perform sentence splitting on the text, splitting the long text into multiple sentences for subsequent word segmentation. For example, split "This is a piece of lively jazz music. The music rhythm is lively." into two sentences: "This is a piece of lively jazz music." and "The music rhythm is lively."

[0148] 23) Target text generation

[0149] Merge the cleaned and normalized text into the final target text. Ensure that the target text is coherent, clear, and accurate, and can accurately describe the content of the target audio. For example, the final target text may be: "This is a piece of lively jazz music. The music rhythm is lively and full of vitality."

[0150] 3) Perform word segmentation processing through a pre-trained large language model

[0151] 31) Select a suitable large language model

[0152] Select an appropriate large language model for word segmentation according to the language and application scenario of the prompt text.

[0153] 32) Perform word segmentation

[0154] Input the target text into the selected large language model for word segmentation. The model will segment the text into individual tokens according to its internal vocabulary and word segmentation algorithm. During the word segmentation process, the model will also consider the context information of the text to ensure the accuracy of word segmentation. For example, when processing the phrase "start over", the model will judge whether to segment it into two tokens "start" and "over" or as a whole token "start over" according to the context, because different segmentation methods may affect subsequent semantic understanding.

[0155] 4) Generate text embeddings

[0156] 41) Obtain token embeddings

[0157] While performing word segmentation, the large language model will generate corresponding embedding vectors for each token. These embedding vectors are points in a high-dimensional space and can represent the semantic information of the tokens. For example, the embedding vector of the token "lively" will contain semantic features related to emotions such as happiness and pleasure. Each token in the model's vocabulary has a unique embedding vector, and these vectors are learned during the model training process and can capture the usage patterns and semantic relationships of the tokens in a large amount of text data.

[0158] 42) Integrate text embeddings

[0159] Integrate the embedding vectors of all tokens into an embedding representation of the entire target text. A common method is to perform average pooling on all token embedding vectors, that is, add the corresponding dimensions of all token embedding vectors and then divide by the number of tokens to obtain a fixed-length text embedding vector. This text embedding vector can comprehensively represent the semantic information of the entire target text. For example, for the target text "This is a lively jazz piece", after average pooling of its token embedding vectors, the obtained text embedding vector will contain comprehensive information about semantic features such as "lively" and "jazz", providing a semantic basis for subsequent fusion and decoding with the audio representation.

[0160] Through the above process, this embodiment can achieve the conversion from the original prompt text to high-quality text embeddings, providing rich semantic features for generating accurate and diverse audio description texts, and improving the performance and output quality of the entire text generation system.

[0161] Furthermore, in one embodiment, the text generation method, wherein the step S300 of downsampling the audio representation and aligning it with the text embedding specifically includes the steps:

[0162] Analyze the dimensions and time steps of the audio representation and the text embedding to generate an analysis result;

[0163] Determine a downsampling strategy based on the analysis result, and use a pre-constructed linear layer to perform a downsampling operation on the audio representation using the downsampling strategy to align the dimensions and time steps of the audio representation with the text embedding.

[0164] In specific implementation, this embodiment realizes the alignment of the audio representation and the text embedding in terms of dimensions and time steps through precise analysis and downsampling operations, providing high-quality feature inputs for subsequent joint decoding. First, by analyzing the dimensions and time steps of the audio representation and the text embedding, the differences between the two can be accurately understood, generating a detailed analysis result. This step is a prerequisite for achieving alignment because it provides the necessary information to formulate an appropriate downsampling strategy. Then, determine the downsampling strategy based on the analysis result, and use a pre-constructed linear layer to perform a downsampling operation on the audio representation. The linear layer can effectively adjust the dimensions of the audio representation to match those of the text embedding, and at the same time adjust the time steps through appropriate downsampling methods (such as interval sampling or time window aggregation) to ensure synchronization in the time series. This alignment operation not only improves the efficiency of feature fusion but also enhances the subsequent model's comprehensive understanding ability of audio and text information, thereby generating more accurate and diverse audio description texts in the subsequent decoding process, significantly improving the overall performance and output quality of the text generation system.

[0165] Among them, the specific implementation process of the steps in this embodiment is roughly as follows:

[0166] 1) Analyze the dimensions and time steps of the audio representation and the text embedding

[0167] 11) Obtain the feature dimensions and time steps

[0168] Obtain the dimensions and time steps of the audio representation. The audio representation is usually a two-dimensional array, where each row represents a time step and each column represents a feature dimension. For example, the audio representation may have 100 time steps, and each time step has 2048 feature dimensions.

[0169] Obtain the dimensions and sequence length of the text embedding. The text embedding is also a two-dimensional array, where each row represents a token and each column represents a feature dimension. For example, the text embedding may have 50 tokens, and each token has 768 feature dimensions.

[0170] 12) Generate an analysis result

[0171] Compare the dimensions and time steps of the audio representation and the text embedding to generate an analysis result. The analysis result includes:

[0172] Dimension difference: The difference between the feature dimension (2048) of the audio representation and the feature dimension (768) of the text embedding.

[0173] Time step difference: The difference between the time step (100) of the audio representation and the sequence length (50) of the text embedding.

[0174] For example, the analysis result may show that the dimension of the audio representation is 2.66 times that of the text embedding, and the time step is 2 times that of the text embedding.

[0175] 2) Determine the downsampling strategy

[0176] 21) Determine the dimension downsampling strategy

[0177] Based on the dimension difference, determine the dimension downsampling strategy. If the dimension of the audio representation is significantly higher than that of the text embedding, a linear transformation can be used to reduce the dimension of the audio representation to the same dimension as the text embedding. For example, use a linear layer to convert the 2048-dimensional audio feature vector into a 768-dimensional vector.

[0178] Select appropriate linear layer parameters (weight matrix and bias vector) to ensure that the converted features can retain important audio information.

[0179] 22) Determine the time step downsampling strategy

[0180] Based on the time step difference, determine the time step downsampling strategy. If the time step of the audio representation is significantly more than the sequence length of the text embedding, methods such as interval sampling or time window aggregation can be used to reduce the time step.

[0181] For example, use interval sampling to select a feature vector every 2 time steps; or use time window aggregation to take the average of the feature vectors of every 2 time steps to generate a new time step feature sequence.

[0182] 3) Perform downsampling operations using a pre-constructed linear layer

[0183] 31) Dimension downsampling operation

[0184] Input the audio representation into the pre-constructed linear layer for dimension downsampling. The weight matrix and bias vector of the linear layer have been learned during the pre-training stage and can convert the high-dimensional audio feature vector into a low-dimensional feature vector.

[0185] 32) Time step downsampling operation

[0186] Downsample the audio representation according to the determined time-step downsampling strategy. If interval sampling is adopted, select the feature vectors at the set intervals; if time-window aggregation is adopted, calculate the average value of the feature vectors within each time window.

[0187] Through the above process, this embodiment can achieve the precise alignment of the audio representation and the text embedding in terms of dimension and time step, providing high-quality feature input for subsequent joint decoding, and significantly improving the overall performance and output quality of the text generation system.

[0188] Further, in one embodiment, the text generation method, wherein, after determining the downsampling strategy according to the analysis result and using the pre-constructed linear layer to perform the downsampling operation on the audio representation using the downsampling strategy to align the dimension and time step of the audio representation with the text embedding, it specifically further includes the steps:

[0189] Verify the alignment result of the dimension and time step of the audio representation and the text embedding;

[0190] When the alignment result does not meet the preset requirements, adjust the downsampling strategy, and re-perform the downsampling operation on the audio representation according to the adjusted downsampling strategy to align the dimension and time step of the audio representation with the text embedding.

[0191] In specific implementation, this embodiment ensures the precise alignment of the audio representation and the text embedding in terms of dimension and time step through a verification and adjustment mechanism, thereby improving the reliability and accuracy of the text generation system. Specifically, first verify whether the dimension and time step of the downsampled audio representation and the text embedding are exactly matched. This step can timely detect possible minor errors in the alignment process. If the alignment result does not meet the preset requirements, for example, there are minor differences in dimension or time step, the system will automatically adjust the downsampling strategy and re-perform the downsampling operation until the audio representation and the text embedding are completely aligned. This iterative optimization process not only improves the accuracy of feature alignment but also enhances the system's adaptability to different audio and text data, ensuring that the generated descriptive text is more accurate, diverse, and real, and significantly improving the overall performance and output quality of the text generation system.

[0192] Further, in one embodiment, the text generation method, wherein, the step S400, using the large language model to decode the aligned audio representation and the text embedding to generate the descriptive text of the target audio, specifically includes the steps:

[0193] Perform feature fusion on the aligned audio representation and the text embedding to generate a fused feature representation;

[0194] Taking the fused feature representation as an input, multiple description texts of the target audio are generated through decoding by the large language model.

[0195] Further, for the text generation method, after taking the fused feature representation as an input and generating multiple description texts of the target audio through decoding by the large language model, the method specifically further includes the steps of:

[0196] Calculating the similarity between each description text and the target audio by using a pre-trained similarity model to obtain a similarity list;

[0197] Based on the similarity list, selecting the description text with the highest similarity to the target audio for output or display.

[0198] In specific implementation, this embodiment realizes the generation of high-quality, diverse and most suitable description texts for the target audio from the aligned audio representation and text embedding through feature fusion, decoding generation, similarity calculation and best text selection, significantly improving the accuracy and practicality of the text generation system. First, feature fusion is performed on the aligned audio representation and text embedding to generate a fused feature representation. This step integrates the semantic information of audio and text, providing rich context information for subsequent decoding. Then, the fused feature representation is input into the large language model for decoding to generate multiple description texts. This step utilizes the powerful text generation ability of the large language model to generate description texts in various styles and details, increasing the diversity of descriptions. Then, the similarity between each description text and the target audio is calculated by using a pre-trained similarity model to obtain a similarity list. This step ensures a high degree of correlation between the generated text and the audio content by quantifying the matching degree between the description text and the target audio. Finally, based on the similarity list, the description text with the highest similarity to the target audio is selected for output or display. This step ensures that the finally output text best matches the content of the target audio, improving the accuracy and practicality of the generated text. Such a technical effect enables the text generation system to generate more natural, fluent and expressive audio description texts, meeting the needs of different users and significantly enhancing the practicality and user experience of the system.

[0199] Among them, the specific implementation process of the steps of this embodiment is roughly as follows:

[0200] 1) Perform feature fusion on the aligned audio representation and text embedding

[0201] 11) Feature concatenation

[0202] Concatenate the aligned audio representation and the text embedding in the feature dimension. For example, if the dimension of the audio representation is da and the dimension of the text embedding is dt, then the dimension of the concatenated feature will be da + dt.

[0203] 12) Feature weighted summation

[0204] In addition to concatenation, weighted summation can also be performed on the audio representation and the text embedding. Different weights are assigned according to the importance of the audio and the text in the target task, and then weighted summation is carried out. For example, if the audio information is more critical in the description generation, a higher weight can be given to the audio representation.

[0205] 13) Multi-layer perceptron (MLP) fusion

[0206] Use a multi-layer perceptron (MLP) to learn the non-linear relationship between the audio representation and the text embedding, and output a fused feature representation. The MLP can contain several hidden layers, and each hidden layer is followed by an activation function (such as ReLU) to increase the expressive power of the model.

[0207] For example, a simple MLP structure can be: input layer (concatenation of audio representation and text embedding) -> hidden layer 1 (1024 neurons, ReLU activation) -> hidden layer 2 (512 neurons, ReLU activation) -> output layer (fused feature representation).

[0208] 2) Generate multiple descriptive texts through the decoding of the large language model

[0209] 21) Initialize the decoder state

[0210] According to the architecture of the large language model, initialize the internal state of the decoder. For a Transformer-based model, this usually includes initializing the query, key, and value vectors in the self-attention mechanism. These vectors can be randomly initialized or preset according to the fused feature representation.

[0211] At the same time, initialize the hidden state of the decoder, which can be a zero vector or an initial hidden state obtained through a linear transformation from the fused feature representation.

[0212] 22) Gradually decode to generate text

[0213] Use the decoder of the large language model to gradually generate descriptive texts. In each step of decoding, the decoder will predict the next token according to the current input (an element in the sequence of fused feature representations) and the previous decoding state (including the generated text and the hidden state).

[0214] For example, in the first-step decoding, the input is the first element of the fused feature representation sequence and the initialized decoder state. The model outputs the probability distribution of the first token, and then selects the token with the highest probability as the output of the current step.

[0215] The selected token is added to the generated text sequence, and the embedding of this token is used as one of the inputs for the next-step decoding. Meanwhile, the hidden state of the decoder is updated to prepare for the next-step decoding. This process is repeated until the generated text sequence reaches the preset maximum length or encounters an end marker (such as <end>)。

[0216] To generate multiple descriptive texts, a beam search strategy can be adopted, considering multiple candidate tokens and selecting the sequence with the highest comprehensive score. For example, setting the beam width to 5, the model will generate 5 candidate text sequences and select the sequence with the highest score as the final output.

[0217] 3) Calculate the similarity using a pre-trained similarity model

[0218] 31) Calculate the similarity

[0219] Calculate the similarity between each generated descriptive text and the target audio using a pre-trained similarity model. The similarity model can be based on cosine similarity, Jaccard similarity, or a more complex deep learning model.

[0220] For example, input the feature representations of the generated descriptive text and the target audio into the similarity model to calculate the similarity value between them. The similarity model can be a pre-trained neural network that can accurately evaluate the matching degree between the two by learning a large amount of audio-text pair data.

[0221] 32) Generate a similarity list

[0222] Store the calculated similarity values in a list, with each generated descriptive text corresponding to a similarity value. For example, if 5 descriptive texts are generated, the similarity list may be [0.85, 0.78, 0.90, 0.82, 0.88], where each value represents the similarity between the corresponding descriptive text and the target audio.

[0223] 4) Select the best descriptive text based on the similarity list

[0224] 41) Select the best descriptive text

[0225] Based on the similarity list, select the descriptive text with the highest similarity to the target audio. For example, from the similarity list [0.85, 0.78, 0.90, 0.82, 0.88], select the descriptive text with the highest similarity (similarity is 0.90).

[0226] Output or display the selected descriptive text to ensure that the finally output text best matches the content of the target audio, improving the accuracy and practicality of the generated text.

[0227] Through the above process, this embodiment can achieve the generation of high-quality, diverse, and most suitable descriptive texts for the target audio from the aligned audio representation and text embedding, significantly improving the accuracy and practicality of the text generation system.

[0228] As can be seen from the above method embodiments, the text generation method provided by the present invention includes: obtaining a target audio, and using an audio Transformer model to extract an audio representation of the target audio; obtaining a prompt text of the target audio, and using a large language model to perform word segmentation processing on the prompt text to generate a text embedding; performing downsampling on the audio representation and aligning it with the text embedding; and using the large language model to decode the aligned audio representation and the text embedding to generate a description text of the target audio. In this way, more diverse, accurate, and realistic audio description texts can be generated through the method of the present invention.

[0229] It should be understood that although the present application provides method operation steps as described in the embodiments or flowcharts, based on routine or non-creative labor, there may be more or fewer operation steps, and these operation steps are not necessarily executed in the order of the embodiments or flowcharts. The step order listed in the embodiments or flowcharts is only one way among many step execution orders and does not represent the only execution order. It should be noted that there is not necessarily a certain sequence between the above steps. Those of ordinary skill in the art can understand from the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel, or the execution can be exchanged, etc. Moreover, at least a part of the steps in the embodiments or flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately, alternately, or synchronously with at least a part of other steps or sub-steps or stages of other steps.

[0230] Based on the above method embodiments, please refer to Figure 3 , another embodiment of the present invention further provides a text generation device, where the device includes:

[0231] An extraction module 11, configured to obtain a target audio and use an audio Transformer model to extract an audio representation of the target audio;

[0232] A processing module 12, configured to obtain a prompt text of the target audio and use a large language model to perform word segmentation processing on the prompt text to generate a text embedding;

[0233] A downsampling module 13, configured to perform downsampling on the audio representation and align it with the text embedding;

[0234] A decoding module 14, configured to use the large language model to decode the aligned audio representation and the text embedding to generate a description text of the target audio.

[0235] Further, in one embodiment, the text generation device, wherein the extraction module 11 is specifically configured to:

[0236] Obtain the audio to be processed, perform noise reduction and normalization processing on the audio to be processed to obtain the target audio;

[0237] Perform frame splitting and feature extraction processing on the target audio to obtain the Mel spectrogram features of the target audio;

[0238] Use the Mel spectrogram features as input, and generate the audio representation of the target audio through the pre-trained audio Transformer model.

[0239] Further, in one embodiment, the text generation device, wherein the processing module 12 is specifically configured to:

[0240] Obtain the prompt text of the target audio, organize the content of the prompt text to generate the target text;

[0241] Perform word segmentation processing on the target text through the pre-trained large language model to generate the text embedding.

[0242] Further, in one embodiment, the text generation device, wherein the downsampling module 13 is specifically configured to:

[0243] Analyze the dimensions and time steps of the audio representation and the text embedding to generate an analysis result;

[0244] Determine a downsampling strategy according to the analysis result, and use the pre-constructed linear layer to perform a downsampling operation on the audio representation using the downsampling strategy to align the dimensions and time steps of the audio representation and the text embedding.

[0245] Further, in one embodiment, after the text generation device determines the downsampling strategy according to the analysis result, and uses the pre-constructed linear layer to perform a downsampling operation on the audio representation using the downsampling strategy to align the dimensions and time steps of the audio representation and the text embedding, it further includes:

[0246] Verify the alignment result of the dimensions and time steps of the audio representation and the text embedding;

[0247] When the alignment result does not meet the preset requirements, adjust the downsampling strategy, and re-perform the downsampling operation on the audio representation according to the adjusted downsampling strategy to align the dimensions and time steps of the audio representation and the text embedding.

[0248] Further, in one embodiment, the text generation device, wherein the decoding module 14 is specifically configured to:

[0249] Perform feature fusion on the aligned audio representation and the text embedding to generate a fused feature representation;

[0250] Use the fused feature representation as input and generate multiple description texts of the target audio through decoding by the large language model.

[0251] Further, in the text generation device, after using the fused feature representation as input and generating multiple description texts of the target audio through decoding by the large language model, it further includes:

[0252] Calculate the similarity between each description text and the target audio using a pre-trained similarity model to obtain a similarity list;

[0253] Based on the similarity list, select the description text with the highest similarity to the target audio for output or display.

[0254] It should be noted that in the device embodiment of the present invention, for the information interaction, execution process, etc. between the above modules, since they are based on the same concept as the method embodiment of the present invention, their specific functions and the technical effects brought are specifically described in the foregoing method embodiment part, and will not be elaborated here.

[0255] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of the text generation method in any one of the above method embodiments.

[0256] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which can be a client, and its internal structure diagram can be as Figure 5 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of the text generation method in any of the above method embodiments.

[0257] Those skilled in the art can understand that Figure 4 and Figure 5 the structural schematic diagram shown in the figure is only a schematic diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more components than those shown in the figure, or combine some components, or have different component arrangements.

[0258] Among them, the so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0259] Among them, the memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or will be output.

[0260] Based on the above method embodiments, another embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the text generation method in any of the above method embodiments. The computer-readable storage medium may be non-volatile or volatile.

[0261] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, and the technical effects brought by the functions / steps, reference can be made to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0262] Those of ordinary skill in the art can understand that all or part of the processes in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. The memory components or memories of the operating environment disclosed herein are intended to include one or more of these and / or any other suitable types of memories.

[0263] Those skilled in the art can clearly understand that, for the convenience and brevity of description, in the device embodiments of the present invention, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0264] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0265] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0266] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0267] It should be noted that in the embodiments of this application, if there are software tools or components that do not belong to our company, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.< / end>

Claims

1. A text generation method, characterized in that: include: Obtaining target audio, and extracting an audio representation of the target audio using an audio Transformer model; Obtaining a prompt text of the target audio, and performing word segmentation processing on the prompt text using a large language model to generate a text embedding; downsampling the audio representation and aligning it with the text embedding; The aligned audio representation and the text embedding are decoded using the large language model to generate a description text of the target audio.

2. The text generation method according to claim 1, characterized in that: The step of obtaining the target audio and extracting the audio representation of the target audio using the audio Transformer model includes: Acquire the audio to be processed, perform noise reduction and normalization on the audio to be processed, and obtain the target audio; Performing frame segmentation and feature extraction processing on the target audio to obtain Mel frequency spectrum features of the target audio; The mel-spectrogram feature is used as input to generate the audio representation of the target audio through the pre-trained audio Transformer model.

3. The text generation method according to claim 1, characterized in that: The obtaining of the prompt text of the target audio, and performing word segmentation processing on the prompt text using a large language model to generate text embedding, includes: Acquire the prompt text of the target audio, organize the content of the prompt text, and generate the target text; The target text is segmented by the pre-trained large language model to generate the text embedding.

4. The text generation method according to claim 1, characterized in that: The downsampling the audio representation and aligning it with the text embedding comprises: Analyze the dimensions and time steps of the audio representation and the text embedding to generate analysis results; A downsampling strategy is determined according to the analysis result, and a pre-built linear layer is used to downsample the audio representation using the downsampling strategy to align the audio representation with the dimension and time step of the text embedding.

5. The text generation method according to claim 4, characterized in that: After determining the downsampling strategy according to the analysis result, and using the pre-constructed linear layer to perform a downsampling operation on the audio representation using the downsampling strategy to align the audio representation with the dimension and time step of the text embedding, the method further includes: Verifying the alignment of the audio representation with the dimensions and time steps of the text embedding; When the alignment result does not meet the preset requirements, the downsampling strategy is adjusted, and the audio representation is re-downsampled according to the adjusted downsampling strategy to align the audio representation with the dimension and time step of the text embedding.

6. The text generation method according to claim 1, characterized in that: The step of decoding the aligned audio representation and the text embedding using the large language model to generate a description text of the target audio includes: Performing feature fusion on the aligned audio representation and the text embedding to generate a fused feature representation; The fused feature representation is taken as input, and a plurality of description texts of the target audio are generated after decoding by the large language model.

7. The text generation method according to claim 6, characterized in that: After taking the fused feature representation as input and decoding the large language model to generate a plurality of description texts of the target audio, the method further includes: Calculate the similarity between each of the description texts and the target audio using a pre-trained similarity model to obtain a similarity list; Based on the similarity list, the description text having the highest similarity to the target audio is selected for output or display.

8. A text generation device, characterized in that: include: An extraction module, used to obtain target audio and extract an audio representation of the target audio using an audio Transformer model; A processing module, used for obtaining a prompt text of the target audio, and performing word segmentation processing on the prompt text using a large language model to generate a text embedding; a downsampling module for downsampling the audio representation and aligning it with the text embedding; A decoding module is used to decode the aligned audio representation and the text embedding using the large language model to generate a description text of the target audio.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the text generation method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text generation method according to any one of claims 1 to 7 is implemented.