Audio analysis method and system, and device, storage medium and computer program product
By acquiring client interest features and using neural network technology to analyze audio, the problem of high efficiency and low cost in extracting key information from audio and video streams has been solved, realizing automated and real-time audio information extraction and summary generation.
Patent Information
- Application Number
- PCT/CN2024/089191
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
In fields such as cybersecurity and surveillance, existing technologies suffer from low processing efficiency and high costs when extracting key information from large amounts of audio and video streams.
By acquiring the client's interest features, it is determined whether the target audio matches the client's interest features, and the matching parts are converted into content summary text. Neural network technology is used for audio analysis, including the application of speech recognition, classification language model and summary language model.
It enables automated extraction of audio information of interest to the client, reduces manual intervention, improves the efficiency and accuracy of information extraction, and provides real-time and effective information.
Smart Images

Figure CN2024089191_30102025_PF_FP_ABST
Abstract
Description
Audio analysis methods and systems, devices, storage media and computer program products Technical Field
[0001] The embodiments of the present invention relate to the field of image processing technology, and more particularly to an audio analysis method and system, device, storage medium and computer program product. Background Technology
[0002] With the continuous development of technology, various intelligent network devices are being used more and more widely in daily production and life, greatly improving the efficiency of production and life. At the same time, it has brought another troublesome problem: the amount of information is too large and the processing is relatively cumbersome. For example, obtaining the required information from a large amount of media streams will require some human intervention. The greater the proportion of human intervention, the more information is processed, the greater the cost, the lower the processing efficiency, and the higher the cost.
[0003] In certain application areas, such as network security and surveillance, there is a need for analyzing large amounts of audio and video streams. This requires extracting key information from the raw data, such as identifying a person, finding sensitive information, and obtaining useful information. However, when useful data is mixed in with a large amount of background data, it can lead to high processing costs. Technical issues
[0004] The problem solved by the embodiments of the present invention is to provide an audio analysis method and system, device, storage medium and computer program product, which is beneficial to providing greater convenience and value to the client. Technical solutions
[0005] To address the aforementioned problems, this invention provides an audio analysis method, comprising: acquiring the client's interest features; acquiring target audio; determining whether the target audio matches the client's interest features; and if they match, acquiring the content summary text of the target audio.
[0006] Optionally, acquiring the target audio includes: acquiring real-time audio as the target audio; or, acquiring the audio of real-time video as the target audio.
[0007] Optionally, the client's interest features are obtained, including: providing multiple interest terms; obtaining one or more interest terms selected by the client as interest features; determining whether the target audio matches the client's interest features, including: determining whether the target audio matches any one or more interest terms in the interest features; if they match, obtaining the content summary text of the target audio, including: if they match, obtaining one or more matching interest terms as target terms; obtaining the content summary text of the part of the target audio corresponding to the target terms.
[0008] Optionally, determining whether the target audio matches any one or more interest terms in the interest features includes: converting the target audio into corresponding target text; determining whether the target text matches any one or more interest terms in the interest features; and obtaining the content summary text of the part corresponding to the target term in the target audio, including: generating the content summary text of the part corresponding to the target term in the target audio based on the target text.
[0009] Optionally, neural network technology can be used to convert the target audio into the corresponding target text.
[0010] Optionally, neural network technology is used to convert the target audio into the corresponding target text, including: establishing a speech recognition model using neural network technology; and using the speech recognition model to convert the target audio into the corresponding target text.
[0011] Optionally, determining whether the target text matches any one or more interest terms in the interest features includes: performing content analysis on the target text to obtain the target features of the target text; determining whether the target features include any one or more interest terms in the interest features; if the target features include any one or more interest terms in the interest features, then the text is determined to match; otherwise, the text is determined to not match.
[0012] Optionally, content analysis is performed on the target text to obtain target features, including: segmenting the target text to obtain multiple paragraphs; obtaining the category label corresponding to each paragraph; and obtaining multiple category labels as target features of the target text.
[0013] Optionally, determining whether the target feature includes any one or more interest terms among the interest features includes: determining whether multiple category labels include any one or more interest terms among the client's interest features; generating a content summary text of the part corresponding to the target term in the target audio based on the target text includes: obtaining the category label corresponding to the target term as the target label; generating a content summary text of the paragraph corresponding to the target label.
[0014] Optionally, neural network technology can be used to obtain the classification labels corresponding to each paragraph.
[0015] Optionally, neural network technology is used to obtain the classification labels corresponding to each paragraph, including: establishing a classification language model using neural network technology, the classification language model including text content and classification labels with a mapping relationship; and using the classification language model to obtain the classification labels corresponding to the text content of each paragraph.
[0016] Optionally, a classification language model is established using neural network technology, including: providing a text dataset, which includes corresponding training text content and training classification labels; training the text dataset using neural network technology to obtain a classification language model; and using the classification language model to obtain the classification labels corresponding to the text content of each paragraph, including: analyzing the paragraphs using the classification language model to obtain training text content that matches the text content of the paragraphs; and obtaining the training classification labels corresponding to the training text content as the classification labels corresponding to the paragraphs.
[0017] Optionally, neural network technology can be used to generate a content summary text of the paragraph corresponding to the target label.
[0018] Optionally, neural network technology is used to generate content summary text of the paragraphs corresponding to the target tags, including: establishing a summary language model using neural network technology, the summary language model being used to generate summaries of text content; and using the summary language model to analyze the text content of the paragraphs corresponding to the target tags to obtain the corresponding content summary text.
[0019] Optionally, after obtaining the content summary text of the target audio, the method further includes: transmitting the content summary text to the client.
[0020] Optionally, an audio analysis method can be deployed locally; or an audio analysis method can be deployed in the cloud.
[0021] Optionally, in the audio analysis method using local deployment, a hardware accelerator can be set up for audio analysis.
[0022] Optionally, the target audio can be deleted if it does not match the client's interests.
[0023] Accordingly, this embodiment of the invention also provides an audio analysis system, including: an interest feature acquisition module for acquiring the interest features of a client; a target audio acquisition module for acquiring target audio; a judgment module for judging whether the target audio matches the interest features of the client; and a content summary text acquisition module for acquiring the content summary text of the target audio if they match.
[0024] Accordingly, embodiments of the present invention also provide a device including at least one memory and at least one processor, wherein the memory stores one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the audio analysis method provided in the embodiments of the present invention.
[0025] Accordingly, embodiments of the present invention also provide a storage medium storing one or more computer instructions, which are used to implement the audio analysis method provided in the embodiments of the present invention.
[0026] Accordingly, embodiments of the present invention also provide a computer program product, including computer programs / instructions, which, when executed by a processor, implement the audio analysis method provided in embodiments of the present invention. Beneficial effects
[0027] Compared with the prior art, the technical solution of the present invention has the following advantages: In the audio analysis method provided by the present invention, the client's interest features are obtained, the target audio is obtained, and it is determined whether the target audio matches the client's interest features. If they match, the content summary text of the target audio is obtained. In the present invention, for the target audio, the part that matches the client's interest features is generated into content summary text. Thus, the present invention can automatically extract the part of the target audio that the client is interested in and provide the client with effective information in the target audio in the form of text. The client does not need to make additional text records of the target audio, which is beneficial to providing the client with greater convenience and better value. Attached Figure Description
[0028] Figure 1 is a flowchart of an embodiment of the audio analysis method of the present invention.
[0029] Figure 2 is a functional block diagram of an embodiment of the audio analysis system of the present invention.
[0030] Figure 3 is a hardware structure diagram of an embodiment of the device provided by the present invention. Embodiments of the present invention
[0031] As the background technology shows, there is currently a demand for the analysis of large amounts of audio and video streams in certain application areas, such as network security and monitoring. This requires extracting key information of interest from the raw information, such as identifying a person, finding sensitive information, and obtaining useful information.
[0032] To address the aforementioned technical problems, embodiments of the present invention provide an audio analysis method. Referring to Figure 1, a flowchart of an embodiment of the audio analysis method of the present invention is shown.
[0033] In this embodiment, the audio analysis method includes the following basic steps:
[0034] Step S1: Obtain the client's interest features;
[0035] Step S2: Obtain the target audio;
[0036] Step S3: Determine whether the target audio matches the client's interest characteristics;
[0037] Step S4: If they match, obtain the content summary text of the target audio.
[0038] In this embodiment of the invention, for the target audio, a content summary text is generated from the part that matches the client's interest characteristics. This embodiment of the invention can automatically extract the part of the target audio that the client is interested in and provide the client with effective information in the target audio in the form of text. The client does not need to make additional text records of the target audio, which is beneficial to providing the client with greater convenience and better value.
[0039] To make the above-mentioned objectives, features and advantages of the embodiments of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below.
[0040] Referring to Figure 1, perform step S1: Obtain the client's interest features.
[0041] Client-side interest characteristics represent the audio content that a client is interested in. Obtaining these characteristics prepares us for subsequent verification of whether target audio matches the client's interests.
[0042] In this embodiment, obtaining the client's interest features includes: performing step S11: providing multiple interest terms.
[0043] It provides multiple interest terms, offering clients a variety of interest references so they can choose content that interests them. Moreover, providing interest terms for clients to choose from is clear and convenient, and it also helps to clearly define the client's interests in text, making it easier and more accurate to subsequently determine whether the target audio matches the client's interests.
[0044] As an example, multiple interest-based entries are provided, including cooking, news, conferences, etc.
[0045] Execution step S12: Obtain one or more interest terms selected by the client as interest features.
[0046] After the client makes a selection, the one or more interest terms selected by the client represent the client's interests and serve as the client's interest features.
[0047] Step S2: Obtain the target audio.
[0048] The target audio is the audio listened to by the client, which is the object of analysis in this embodiment to determine whether it matches the client's interest characteristics.
[0049] In this embodiment, obtaining the target audio includes: performing step S21: obtaining real-time audio as the target audio; or, obtaining the audio of real-time video as the target audio.
[0050] By acquiring real-time audio as the target audio, the audio being listened to by the client can be captured in real time without the need for additional recording steps on the client side. This allows for real-time extraction of information from the target audio, which in turn improves the timeliness of the extracted information and provides the client with effective information in real time.
[0051] Alternatively, by acquiring the audio of real-time video as the target audio, the audio of the video being watched by the client can be captured in real time without the need for additional recording steps such as recording or screen recording on the client side. This allows for real-time extraction of information from the target audio, which is beneficial for achieving real-time information extraction and thus improving the timeliness of the information extracted from the target audio, providing the client with effective information in real time.
[0052] For example, if the client is listening to an audio program or watching a cooking video, then in this embodiment, the real-time cooking audio is obtained as the target audio.
[0053] As an example, in this embodiment, real-time audio is acquired as the target audio.
[0054] Step S3: Determine whether the target audio matches the client's interest characteristics.
[0055] Determine whether the target audio matches the client's interest characteristics to identify whether the target audio aligns with the client's interests, which will serve as the basis for deciding whether to generate a content summary text.
[0056] Specifically, in this embodiment, determining whether the target audio matches the client's interest features includes: determining whether the target audio matches any one or more interest terms in the interest features.
[0057] Determining whether the target audio matches any one or more interest terms in the interest features means that the target audio matches any one interest term in the interest features, or the target audio matches any multiple interest terms in the interest features, all indicate that the target audio matches the client's interest features.
[0058] In this embodiment, determining whether the target audio matches any one or more interest terms in the interest features includes: performing step S31: converting the target audio into the corresponding target text.
[0059] The target audio is converted into corresponding target text, which then serves as the object of subsequent analysis. First, the target audio is converted into text-based target text, and then the target text is analyzed to further analyze the target audio. This text analysis approach is simpler and easier to implement. Furthermore, the target text is used to prepare for generating subsequent content summary text.
[0060] In this embodiment, neural network technology is used to convert the target audio into the corresponding target text.
[0061] Neural network technology possesses powerful learning capabilities, exhibiting excellent self-learning and adaptive abilities. Through training, it can automatically extract features from input data and learn complex relationships between data, giving it a unique advantage in handling complex, uncertain, and nonlinear problems. Furthermore, neural network technology has strong generalization capabilities; a trained neural network can process untrained related data and obtain appropriate results, demonstrating good fault tolerance and robustness. It can maintain stable performance under different environments and conditions, resulting in a high accuracy rate in converting target audio into corresponding target text.
[0062] In this embodiment, neural network technology is used to convert the target audio into the corresponding target text, including: establishing a speech recognition model using neural network technology.
[0063] By using neural network technology to train relevant data in a targeted manner and establishing a speech recognition model, a more accurate audio analysis method suitable for this embodiment can be developed.
[0064] In this embodiment, a speech recognition model is used to convert the target audio into the corresponding target text.
[0065] Speech recognition models built using targeted training can improve the accuracy of converting target audio into corresponding target text.
[0066] In other embodiments, the target audio can be converted into the corresponding target text using an existing speech recognition model without additional data training.
[0067] Step S32: Determine whether the target text matches any one or more interest terms in the interest features.
[0068] Determining whether the target text matches any one or more interest terms in the interest features means that the target text matches any one interest term in the interest features, or the target text matches any multiple interest terms in the interest features, all indicate that the target text matches the client's interest features.
[0069] In this embodiment, determining whether the target text matches any one or more interest terms in the interest features includes: performing content analysis on the target text to obtain the target features of the target text.
[0070] Content analysis of the target text is performed to obtain target features. Target features represent the content characteristics of the target text. By using target features to determine whether the content of the target text matches the client's interest characteristics, the judgment can be made more accurately. Furthermore, extracting target features for judgment simplifies the judgment process and helps to improve judgment efficiency.
[0071] In this embodiment, content analysis is performed on the target text to obtain the target features of the target text, including: segmenting the target text to obtain multiple paragraphs.
[0072] Segmenting the target text into multiple paragraphs allows for separate analysis of each paragraph. This reduces the computational cost of each analysis, improves efficiency, and facilitates the selection of paragraphs that match the client's interests from the complex target text. Consequently, it reduces the computational cost of extracting the content summary text and improves the efficiency of audio analysis.
[0073] In this embodiment, the target text is segmented based on punctuation marks to obtain multiple paragraphs, or the target text is segmented based on pause time to obtain multiple paragraphs.
[0074] In this embodiment, the category tags corresponding to each paragraph are obtained.
[0075] The system obtains category tags for each paragraph to represent its content features. Using category tags to represent corresponding paragraphs allows for better categorization of multiple paragraphs and provides a concise representation of each paragraph. By comparing category tags and interest terms, the system can determine whether the target text matches the client's interest features, which helps save computing power and improve analysis efficiency.
[0076] In this embodiment, neural network technology is used to obtain the category label corresponding to each paragraph.
[0077] Neural network technology possesses powerful learning capabilities, exhibiting excellent self-learning and adaptive abilities. Through training, it can automatically extract features from input data and learn complex relationships between data, giving it a unique advantage in handling complex, uncertain, and nonlinear problems. Furthermore, neural network technology has strong generalization capabilities; a trained neural network can process related data that has not been trained and obtain appropriate results for that data. This gives neural networks good fault tolerance and robustness, enabling them to maintain stable performance under different environments and conditions, thus achieving high accuracy in obtaining classification labels for each paragraph.
[0078] In this embodiment, neural network technology is used to obtain the classification labels corresponding to each paragraph, including: using neural network technology to establish a classification language model, the classification language model including text content and classification labels with a mapping relationship.
[0079] By using neural network technology to train relevant data in a targeted manner and establish a classification language model, it is beneficial to make the audio analysis method applicable to this embodiment more accurate. The classification language model includes text content and classification labels with a mapping relationship, so that text information can be input into the classification language model and corresponding classification labels can be output accordingly.
[0080] As an example, in this embodiment, neural network technology is used to build large-scale language models (LLMs) as classification language models.
[0081] In other embodiments, neural network technology can also be used to build small language models (SLMs) as classification language models.
[0082] In this embodiment, a classification language model is established using neural network technology, including: providing a text dataset, which includes corresponding training text content and training classification labels.
[0083] The text dataset includes corresponding training text content and training classification labels, which are used for targeted data training to obtain a targeted classification language model.
[0084] In this embodiment, neural network technology is used to train the text dataset to obtain a classification language model.
[0085] By using neural network technology to train the text dataset specifically for this embodiment, a classification language model suitable for this embodiment can be obtained, thereby improving the accuracy of the audio analysis method.
[0086] In this embodiment, a classification language model is used to obtain the classification labels corresponding to the text content of each paragraph.
[0087] We use a classification language model to obtain the classification labels corresponding to the text content of each paragraph, in order to prepare for the subsequent comparison of the classification labels of each paragraph with interest terms.
[0088] In other embodiments, instead of performing additional data training, existing language models can be used to obtain the classification labels corresponding to the text content of each paragraph.
[0089] In this embodiment, the classification language model is used to obtain the classification label corresponding to the text content of each paragraph, including: using the classification language model to analyze the paragraph and obtain training text content that matches the text content of the paragraph.
[0090] The corresponding classification labels are obtained by matching the training text content with the text content of paragraphs.
[0091] Accordingly, in this embodiment, the training classification labels corresponding to the training text content are obtained as the classification labels corresponding to the paragraphs.
[0092] As an example, in this embodiment, the target text is segmented to obtain three paragraphs, and the category tags for each paragraph are cooking, news and meetings.
[0093] In this embodiment, multiple classification labels are obtained as target features of the target text.
[0094] Multiple category labels serve as target features of the target text, representing its content characteristics.
[0095] In this embodiment, it is determined whether the target feature includes any one or more interest terms in the interest features.
[0096] The algorithm for determining whether a target text matches an interest feature is transformed into determining whether the target feature includes one or more interest terms. By comparing the target text with the interest terms, the algorithm for determining the target text is simplified and the efficiency of the determination is improved.
[0097] Accordingly, in this embodiment, if the target feature includes any one or more interest terms in the interest features, then it is determined to be consistent; otherwise, it is determined to be inconsistent.
[0098] In this embodiment, determining whether the target feature includes any one or more interest terms among the interest features includes: determining whether multiple classification labels include any one or more interest terms among the client's interest features.
[0099] The system compares multiple category labels with interest terms in the client's interest features. If the multiple category labels include one or more interest terms in the interest features, it means that the target feature includes any one or more interest terms in the interest features. Accordingly, it means that the target text matches the client's interest features.
[0100] Execute step S4: If they match, obtain the content summary text of the target audio.
[0101] Obtain the content summary text of the target audio, which is used to provide the client with text content of interest.
[0102] In this embodiment, for the target audio, the portion that matches the client's interest features is generated as a content summary text. Thus, this embodiment of the invention can automatically extract the portion of the target audio that the client is interested in and provide the client with effective information in the target audio in the form of text. This eliminates the need for the client to make additional text recordings of the target audio, thereby providing greater convenience and value to the client.
[0103] In this embodiment, if they match, the content summary text of the target audio is obtained, including: step S41: if they match, one or more matching interest terms are obtained as target terms.
[0104] One or more matching terms of interest are selected as target terms and used as the basis for generating subsequent content summary text.
[0105] Specifically, in this embodiment, multiple category tags are compared with interest terms in the client's interest features to determine whether the multiple category tags match the interest terms. If there is a category tag that matches the interest terms in the interest features, the corresponding interest term is selected as the target term.
[0106] As an example, in this embodiment, the client's interest features include multiple interest terms such as cooking, news, and travel. The target text has multiple paragraphs with corresponding category tags such as cooking, cooking, conferences, and news. The category tags cooking and news match the interest terms cooking and news, thus the matching interest terms cooking and news are obtained as the target terms.
[0107] Execute step S42 to obtain the content summary text of the part corresponding to the target term in the target audio.
[0108] By obtaining the content summary text of the corresponding part of the target term in the target audio, the content summary text of the part of the target audio that the customer is interested in can be output in a targeted manner, which helps to save computing power in generating content summary text and improve the generation rate of content summary text.
[0109] In this embodiment, obtaining the content summary text of the part corresponding to the target term in the target audio includes: generating the content summary text of the part corresponding to the target term in the target audio based on the target text.
[0110] After converting the target audio into target text, the corresponding parts of the target terms are extracted and simplified to obtain the content summary text.
[0111] In this embodiment, based on the target text, a content summary text corresponding to the target term in the target audio is generated, including: obtaining the category tag corresponding to the target term as the target tag.
[0112] Obtain the category tags corresponding to the target terms as target tags, which will serve as the basis for generating content summary text for paragraphs that match the client's interests.
[0113] As an example, in this embodiment, the client's interest features include multiple interest terms such as cooking, news, and travel. The target text has multiple paragraphs with category tags such as cooking, cooking, conferences, and news. The category tags cooking and news match the interest terms cooking and news. The matching interest terms cooking and news are obtained as the target terms, and the category tags cooking and news corresponding to the target terms are used as the target tags.
[0114] In this embodiment, a summary text of the paragraph corresponding to the target tag is generated.
[0115] The paragraphs corresponding to the target tags represent paragraphs that match the client's interests, i.e., paragraphs that the client is interested in. This generates a summary text of the content of the paragraphs corresponding to the target tags, providing the client with content in a text format that is of interest to them.
[0116] As an example, in this embodiment, the client's interest features include multiple interest terms such as cooking, news, and travel. The target text has multiple paragraphs with corresponding category tags such as cooking, cooking, conferences, and news. The category tags cooking and news match the interest terms cooking and news. The matching interest terms cooking and news are obtained as target terms. In the target text, the content summary text of the paragraphs corresponding to the target terms cooking and news is generated.
[0117] As an example, a summary text of cooking steps is generated for the paragraph corresponding to the target term "cooking," a summary text of news summaries is generated for the paragraph corresponding to the target term "news," and a summary text of meeting minutes is generated for the paragraph corresponding to the target term "meeting."
[0118] In this embodiment, neural network technology is used to generate the content summary text of the paragraph corresponding to the target tag.
[0119] Neural network technology possesses powerful learning capabilities, exhibiting excellent self-learning and adaptive abilities. Through training, it can automatically extract features from input data and learn complex relationships between data, giving it a unique advantage in handling complex, uncertain, and nonlinear problems. Furthermore, neural network technology has strong generalization capabilities; a trained neural network can process related data that has not been trained and obtain appropriate results for that data. This gives neural networks good fault tolerance and robustness, enabling them to maintain stable performance under different environments and conditions, thus resulting in high accuracy in generating content summary text for paragraphs corresponding to target tags.
[0120] In this embodiment, neural network technology is used to generate content summary text of the paragraph corresponding to the target tag, including: establishing a summary language model using neural network technology, which is used to generate a summary of the text content.
[0121] By using neural network technology to train relevant data in a targeted manner and establish a summary language model, it is beneficial to more accurately apply the audio analysis method in this embodiment. The classification language model is used to generate summaries of text content, so that the text information of the paragraph can be input into the summary language model and the corresponding content summary text can be output accordingly.
[0122] As an example, in this embodiment, neural network technology is used to build large-scale language models (LLMs) as the summary language model.
[0123] In other embodiments, neural network technology can also be used to build small language models (SLMs) as summary language models.
[0124] As an example, in this embodiment, the same large-scale language model is used as both the classification language model and the summarization language model.
[0125] In other embodiments, different large-scale language models may be used as the classification language model and the summarization language model, respectively.
[0126] In this embodiment, a summary language model is used to analyze the text content of the paragraph corresponding to the target tag to obtain the corresponding content summary text.
[0127] We use a summary language model to analyze the text content of paragraphs corresponding to target tags and provide customers with summary content of interest.
[0128] In other embodiments, without additional data training, an existing language model can be used to generate the content summary text of the paragraph corresponding to the target label.
[0129] In this embodiment, after obtaining the content summary text of the target audio, the method further includes: performing step S5: transmitting the content summary text to the client.
[0130] The content summary text is transmitted to the client, providing the client with valuable content information.
[0131] In this embodiment, if the target audio does not match the client's interest characteristics, the target audio is deleted.
[0132] If the target audio does not match the client's interest characteristics, indicating that the client is not interested in the target audio, then the target audio is deleted, which helps save storage space and ensures the analysis efficiency of the audio analysis method in this embodiment.
[0133] Specifically, in this embodiment, if the target audio does not match the client's interest characteristics, the target text corresponding to the target audio is deleted.
[0134] In other embodiments, the target audio may still be retained if it does not match the client's interest characteristics.
[0135] In this embodiment, a local deployment method is used for audio analysis.
[0136] Using a local deployment method for audio analysis helps protect client privacy.
[0137] In this embodiment, the audio analysis method using a local deployment approach employs a hardware accelerator.
[0138] Hardware accelerators are used to ensure the analysis speed of the audio analysis method, thereby ensuring the real-time provision of content text summaries to the client by the audio analysis method in this embodiment.
[0139] Specifically, in this embodiment, a hardware accelerator is set up to perform neural network technology. When designing the neural network, it is necessary to adapt it to the hardware. Through hardware and software co-design, it is beneficial to ensure the accuracy and speed of the neural network technology, thereby ensuring the ability to process in real time.
[0140] In other embodiments, the audio analysis method can also be performed using a cloud-based deployment approach.
[0141] Accordingly, the present invention also provides an audio analysis system. Figure 2 is a functional block diagram of an embodiment of the audio analysis system of the present invention.
[0142] In this embodiment, the audio analysis system 50 includes: an interest feature acquisition module 501, used to acquire the interest features of the client; a target audio acquisition module 502, used to acquire the target audio; a judgment module 503, used to judge whether the target audio matches the interest features of the client; and a content summary text acquisition module 504, used to acquire the content summary text of the target audio if they match.
[0143] The interest feature acquisition module 501 is used to acquire the interest features of the client.
[0144] Client-side interest characteristics represent the audio content that clients are interested in. Obtaining these characteristics prepares us for subsequent verification of whether target audio matches the client's interests.
[0145] In this embodiment, obtaining the client's interest features includes providing multiple interest terms.
[0146] It provides multiple interest terms, offering clients a variety of interest references so they can choose content that interests them. Moreover, providing interest terms for clients to choose from is clear and convenient, and it also helps to clearly define the client's interests in text, making it easier and more accurate to subsequently determine whether the target audio matches the client's interests.
[0147] As an example, multiple interest-based terms are provided, including cooking, news, conferences, etc.
[0148] In this embodiment, one or more interest terms selected by the client are obtained as interest features.
[0149] After the client makes a selection, the one or more interest terms selected by the client represent the client's interests and serve as the client's interest features.
[0150] The target audio acquisition module 502 is used to acquire the target audio.
[0151] The target audio is the audio listened to by the client, which is the object of analysis in this embodiment to determine whether it matches the client's interest characteristics.
[0152] In this embodiment, obtaining the target audio includes: obtaining real-time audio as the target audio; or, obtaining the audio of real-time video as the target audio.
[0153] By acquiring real-time audio as the target audio, the audio being listened to by the client can be captured in real time without the need for additional recording steps on the client side. This allows for real-time extraction of information from the target audio, which in turn improves the timeliness of the extracted information and provides the client with effective information in real time.
[0154] Alternatively, by acquiring the audio of real-time video as the target audio, the audio of the video being watched by the client can be captured in real time without the need for additional recording steps such as recording or screen recording on the client side. This allows for real-time extraction of information from the target audio, which is beneficial for achieving real-time information extraction and thus improving the timeliness of the information extracted from the target audio, providing the client with effective information in real time.
[0155] For example, if the client is listening to an audio program or watching a cooking video, then in this embodiment, the real-time cooking audio is obtained as the target audio.
[0156] As an example, in this embodiment, real-time audio is acquired as the target audio.
[0157] The judgment module 503 is used to determine whether the target audio matches the client's interest characteristics.
[0158] Determine whether the target audio matches the client's interest characteristics to identify whether the target audio aligns with the client's interests, which will serve as the basis for deciding whether to generate a content summary text.
[0159] Specifically, in this embodiment, determining whether the target audio matches the client's interest features includes: determining whether the target audio matches any one or more interest terms in the interest features.
[0160] Determining whether the target audio matches any one or more interest terms in the interest features means that the target audio matches any one interest term in the interest features, or the target audio matches any multiple interest terms in the interest features, all indicate that the target audio matches the client's interest features.
[0161] In this embodiment, determining whether the target audio matches any one or more interest terms in the interest features includes converting the target audio into the corresponding target text.
[0162] The target audio is converted into corresponding target text, which then serves as the object of subsequent analysis. First, the target audio is converted into text-based target text, and then the target text is analyzed to further analyze the target audio. This text analysis approach is simpler and easier to implement. Furthermore, the target text is used to prepare for generating subsequent content summary text.
[0163] In this embodiment, neural network technology is used to convert the target audio into the corresponding target text.
[0164] Neural network technology possesses powerful learning capabilities, exhibiting excellent self-learning and adaptive abilities. Through training, it can automatically extract features from input data and learn complex relationships between data, giving it a unique advantage in handling complex, uncertain, and nonlinear problems. Furthermore, neural network technology has strong generalization capabilities; a trained neural network can process untrained related data and obtain appropriate results, demonstrating good fault tolerance and robustness. It can maintain stable performance under different environments and conditions, resulting in a high accuracy rate in converting target audio into corresponding target text.
[0165] In this embodiment, neural network technology is used to convert the target audio into the corresponding target text, including: establishing a speech recognition model using neural network technology.
[0166] By using neural network technology to train relevant data in a targeted manner and establishing a speech recognition model, a more accurate audio analysis method suitable for this embodiment can be developed.
[0167] In this embodiment, a speech recognition model is used to convert the target audio into the corresponding target text.
[0168] Speech recognition models built using targeted training can improve the accuracy of converting target audio into corresponding target text.
[0169] In other embodiments, the target audio can be converted into the corresponding target text using an existing speech recognition model without additional data training.
[0170] In this embodiment, it is determined whether the target text matches any one or more interest terms in the interest features.
[0171] Determining whether the target text matches any one or more interest terms in the interest features means that the target text matches any one interest term in the interest features, or the target text matches any multiple interest terms in the interest features, all indicate that the target text matches the client's interest features.
[0172] In this embodiment, determining whether the target text matches any one or more interest terms in the interest features includes: performing content analysis on the target text to obtain the target features of the target text.
[0173] Content analysis of the target text is performed to obtain target features. Target features represent the content characteristics of the target text. By using target features to determine whether the content of the target text matches the client's interest characteristics, the judgment can be made more accurately. Furthermore, extracting target features for judgment simplifies the judgment process and helps to improve judgment efficiency.
[0174] In this embodiment, content analysis is performed on the target text to obtain the target features of the target text, including: segmenting the target text to obtain multiple paragraphs.
[0175] Segmenting the target text into multiple paragraphs allows for separate analysis of each paragraph. This reduces the computational cost of each analysis, improves efficiency, and facilitates the selection of paragraphs that match the client's interests from the complex target text. Consequently, it reduces the computational cost of extracting the content summary text and improves the efficiency of audio analysis.
[0176] In this embodiment, the target text is segmented based on punctuation marks to obtain multiple paragraphs, or the target text is segmented based on pause time to obtain multiple paragraphs.
[0177] In this embodiment, the category tags corresponding to each paragraph are obtained.
[0178] The system obtains category tags for each paragraph to represent its content features. Using category tags to represent corresponding paragraphs allows for better categorization of multiple paragraphs and provides a concise representation of each paragraph. By comparing category tags and interest terms, the system can determine whether the target text matches the client's interest features, which helps save computing power and improve analysis efficiency.
[0179] In this embodiment, neural network technology is used to obtain the category label corresponding to each paragraph.
[0180] Neural network technology possesses powerful learning capabilities, exhibiting excellent self-learning and adaptive abilities. Through training, it can automatically extract features from input data and learn complex relationships between data, giving it a unique advantage in handling complex, uncertain, and nonlinear problems. Furthermore, neural network technology has strong generalization capabilities; a trained neural network can process related data that has not been trained and obtain appropriate results for that data. This gives neural networks good fault tolerance and robustness, enabling them to maintain stable performance under different environments and conditions, thus achieving high accuracy in obtaining classification labels for each paragraph.
[0181] In this embodiment, neural network technology is used to obtain the classification labels corresponding to each paragraph, including: using neural network technology to establish a classification language model, the classification language model including text content and classification labels with a mapping relationship.
[0182] By using neural network technology to train relevant data in a targeted manner and establish a classification language model, it is beneficial to make the audio analysis method applicable to this embodiment more accurate. The classification language model includes text content and classification labels with a mapping relationship, so that text information can be input into the classification language model and corresponding classification labels can be output accordingly.
[0183] As an example, in this embodiment, neural network technology is used to build large-scale language models (LLMs) as classification language models.
[0184] In other embodiments, neural network technology can also be used to build small language models (SLMs) as classification language models.
[0185] In this embodiment, a classification language model is established using neural network technology, including: providing a text dataset, which includes corresponding training text content and training classification labels.
[0186] The text dataset includes corresponding training text content and training classification labels, which are used for targeted data training to obtain a targeted classification language model.
[0187] In this embodiment, neural network technology is used to train the text dataset to obtain a classification language model.
[0188] By using neural network technology to train the text dataset specifically for this embodiment, a classification language model suitable for this embodiment can be obtained, thereby improving the accuracy of the audio analysis method.
[0189] In this embodiment, a classification language model is used to obtain the classification labels corresponding to the text content of each paragraph.
[0190] We use a classification language model to obtain the classification labels corresponding to the text content of each paragraph, in order to prepare for the subsequent comparison of the classification labels of each paragraph with interest terms.
[0191] In other embodiments, instead of performing additional data training, existing language models can be used to obtain the classification labels corresponding to the text content of each paragraph.
[0192] In this embodiment, the classification language model is used to obtain the classification label corresponding to the text content of each paragraph, including: using the classification language model to analyze the paragraph and obtain training text content that matches the text content of the paragraph.
[0193] The corresponding classification labels are obtained by matching the training text content with the text content of paragraphs.
[0194] Accordingly, in this embodiment, the training classification labels corresponding to the training text content are obtained as the classification labels corresponding to the paragraphs.
[0195] As an example, in this embodiment, the target text is segmented to obtain three paragraphs, and the category tags for each paragraph are cooking, news and meetings.
[0196] In this embodiment, multiple classification labels are obtained as target features of the target text.
[0197] Multiple category labels serve as target features of the target text, representing its content characteristics.
[0198] In this embodiment, it is determined whether the target feature includes any one or more interest terms in the interest features.
[0199] The algorithm for determining whether a target text matches an interest feature is transformed into determining whether the target feature includes one or more interest terms. By comparing the target text with the interest terms, the algorithm for determining the target text is simplified and the efficiency of the determination is improved.
[0200] Accordingly, in this embodiment, if the target feature includes any one or more interest terms in the interest features, then it is determined to be consistent; otherwise, it is determined to be inconsistent.
[0201] In this embodiment, determining whether the target feature includes any one or more interest terms among the interest features includes: determining whether multiple classification labels include any one or more interest terms among the client's interest features.
[0202] The system compares multiple category labels with interest terms in the client's interest features. If the multiple category labels include one or more interest terms in the interest features, it means that the target feature includes any one or more interest terms in the interest features. Accordingly, it means that the target text matches the client's interest features.
[0203] The content summary text acquisition module 504 is used to acquire the content summary text of the target audio if a match is found.
[0204] Obtain the content summary text of the target audio, which is used to provide the client with text content of interest.
[0205] In this embodiment, for the target audio, the portion that matches the client's interest features is generated as a content summary text. Thus, this embodiment of the invention can automatically extract the portion of the target audio that the client is interested in and provide the client with effective information in the target audio in the form of text. This eliminates the need for the client to make additional text recordings of the target audio, thereby providing greater convenience and value to the client.
[0206] In this embodiment, if they match, the content summary text of the target audio is obtained, including: if they match, one or more matching interest terms are obtained as target terms.
[0207] One or more matching terms of interest are selected as target terms and used as the basis for generating subsequent content summary text.
[0208] Specifically, in this embodiment, multiple category tags are compared with interest terms in the client's interest features to determine whether the multiple category tags match the interest terms. If there is a category tag that matches the interest terms in the interest features, the corresponding interest term is selected as the target term.
[0209] As an example, in this embodiment, the client's interest features include multiple interest terms such as cooking, news, and travel. The target text has multiple paragraphs with corresponding category tags such as cooking, cooking, conferences, and news. The category tags cooking and news match the interest terms cooking and news, thus the matching interest terms cooking and news are obtained as the target terms.
[0210] In this embodiment, the content summary text of the part corresponding to the target term in the target audio is obtained.
[0211] By obtaining the content summary text of the corresponding part of the target term in the target audio, the content summary text of the part of the target audio that the customer is interested in can be output in a targeted manner, which helps to save computing power in generating content summary text and improve the generation rate of content summary text.
[0212] In this embodiment, obtaining the content summary text of the part corresponding to the target term in the target audio includes: generating the content summary text of the part corresponding to the target term in the target audio based on the target text.
[0213] After converting the target audio into target text, the corresponding parts of the target terms are extracted and simplified to obtain the content summary text.
[0214] In this embodiment, based on the target text, a content summary text corresponding to the target term in the target audio is generated, including: obtaining the category tag corresponding to the target term as the target tag.
[0215] Obtain the category tags corresponding to the target terms as target tags, which will serve as the basis for generating content summary text for paragraphs that match the client's interests.
[0216] As an example, in this embodiment, the client's interest features include multiple interest terms such as cooking, news, and travel. The target text has multiple paragraphs with category tags such as cooking, cooking, conferences, and news. The category tags cooking and news match the interest terms cooking and news. The matching interest terms cooking and news are obtained as the target terms, and the category tags cooking and news corresponding to the target terms are used as the target tags.
[0217] In this embodiment, a summary text of the paragraph corresponding to the target tag is generated.
[0218] The paragraphs corresponding to the target tags represent paragraphs that match the client's interests, i.e., paragraphs that the client is interested in. This generates a summary text of the content of the paragraphs corresponding to the target tags, providing the client with content in a text format that is of interest to them.
[0219] As an example, in this embodiment, the client's interest features include multiple interest terms such as cooking, news, and travel. The target text has multiple paragraphs with corresponding category tags such as cooking, cooking, conferences, and news. The category tags cooking and news match the interest terms cooking and news. The matching interest terms cooking and news are obtained as target terms. In the target text, the content summary text of the paragraphs corresponding to the target terms cooking and news is generated.
[0220] As an example, a summary text of cooking steps is generated for the paragraph corresponding to the target term "cooking," a summary text of news summaries is generated for the paragraph corresponding to the target term "news," and a summary text of meeting minutes is generated for the paragraph corresponding to the target term "meeting."
[0221] In this embodiment, neural network technology is used to generate the content summary text of the paragraph corresponding to the target tag.
[0222] Neural network technology possesses powerful learning capabilities, exhibiting excellent self-learning and adaptive abilities. Through training, it can automatically extract features from input data and learn complex relationships between data, giving it a unique advantage in handling complex, uncertain, and nonlinear problems. Furthermore, neural network technology has strong generalization capabilities; a trained neural network can process related data that has not been trained and obtain appropriate results for that data. This gives neural networks good fault tolerance and robustness, enabling them to maintain stable performance under different environments and conditions, thus resulting in high accuracy in generating content summary text for paragraphs corresponding to target tags.
[0223] In this embodiment, neural network technology is used to generate content summary text of the paragraph corresponding to the target tag, including: establishing a summary language model using neural network technology, which is used to generate a summary of the text content.
[0224] By using neural network technology to train relevant data in a targeted manner and establish a summary language model, it is beneficial to more accurately apply the audio analysis method in this embodiment. The classification language model is used to generate summaries of text content, so that the text information of the paragraph can be input into the summary language model and the corresponding content summary text can be output accordingly.
[0225] As an example, in this embodiment, neural network technology is used to build large-scale language models (LLMs) as the summary language model.
[0226] In other embodiments, neural network technology can also be used to build small language models (SLMs) as summary language models.
[0227] As an example, in this embodiment, the same large-scale language model is used as both the classification language model and the summarization language model.
[0228] In other embodiments, different large-scale language models may be used as the classification language model and the summarization language model, respectively.
[0229] In this embodiment, a summary language model is used to analyze the text content of the paragraph corresponding to the target tag to obtain the corresponding content summary text.
[0230] We use a summary language model to analyze the text content of paragraphs corresponding to target tags and provide customers with summary content of interest.
[0231] In other embodiments, without additional data training, an existing language model can be used to generate the content summary text of the paragraph corresponding to the target label.
[0232] In this embodiment, after obtaining the content summary text of the target audio, the method further includes: performing step S5: transmitting the content summary text to the client.
[0233] The content summary text is transmitted to the client, providing the client with valuable content information.
[0234] In this embodiment, if the target audio does not match the client's interest characteristics, the target audio is deleted.
[0235] If the target audio does not match the client's interest characteristics, indicating that the client is not interested in the target audio, then the target audio is deleted, which helps save storage space and ensures the analysis efficiency of the audio analysis method in this embodiment.
[0236] Specifically, in this embodiment, if the target audio does not match the client's interest characteristics, the target text corresponding to the target audio is deleted.
[0237] In other embodiments, the target audio may still be retained if it does not match the client's interest characteristics.
[0238] In this embodiment, a local deployment method is used for audio analysis.
[0239] Using a local deployment method for audio analysis helps protect client privacy.
[0240] In this embodiment, the audio analysis method using a local deployment approach employs a hardware accelerator.
[0241] Hardware accelerators are used to ensure the analysis speed of the audio analysis method, thereby ensuring the real-time provision of content text summaries to the client by the audio analysis method in this embodiment.
[0242] Specifically, in this embodiment, a hardware accelerator is set up to perform neural network technology. When designing the neural network, it is necessary to adapt it to the hardware. Through hardware and software co-design, it is beneficial to ensure the accuracy and speed of the neural network technology, thereby ensuring the ability to process in real time.
[0243] In other embodiments, the audio analysis method can also be performed using a cloud-based deployment approach.
[0244] This invention also provides a device that can implement the audio analysis method provided in this invention by loading a program. An optional hardware structure of the terminal device provided in this invention, as shown in FIG3, includes: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.
[0245] In this embodiment, the number of processor 01, communication interface 02, memory 03, and communication bus 04 is at least one, and the processor 01, communication interface 02, and memory 03 communicate with each other through communication bus 04. Communication interface 02 can be an interface of a communication module for network communication, such as the interface of a GSM module. Processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. Memory 03 may include high-speed RAM and may also include non-volatile memory (NVM), such as at least one disk storage device. Memory 03 stores one or more computer instructions, which are executed by processor 01 to implement the audio analysis method provided in this embodiment of the present invention.
[0246] It should be noted that the aforementioned terminal device may also include other devices (not shown) that may not be essential to understanding the content disclosed in the embodiments of the present invention; given that these other devices may not be essential for understanding the content disclosed in the embodiments of the present invention, the embodiments of the present invention will not describe them one by one.
[0247] This invention also provides a storage medium storing one or more computer instructions for implementing the audio analysis method provided in this invention.
[0248] In this embodiment of the invention, for the target audio, a content summary text is generated from the part that matches the client's interest characteristics. This embodiment of the invention can automatically extract the part of the target audio that the client is interested in and provide the client with effective information in the target audio in the form of text. The client does not need to make additional text records of the target audio, which is beneficial to providing the client with greater convenience and better value.
[0249] The embodiments of the present invention described above are combinations of elements and features of the present invention. Unless otherwise stated, elements or features may be considered optional. Individual elements or features may be practiced without combination with other elements or features. Furthermore, embodiments of the present invention may be constructed by combining some elements and / or features. The order of operations described in the embodiments of the present invention may be rearranged. Some constructions of any embodiment may be included in another embodiment and may be replaced by corresponding constructions of another embodiment. It will be apparent to those skilled in the art that claims in the appended claims that are not expressly referenced in each other may be combined to form embodiments of the present invention, or may be included as new claims in amendments made after the filing of this application.
[0250] Embodiments of the present invention can be implemented by various means, such as hardware, firmware, software, or combinations thereof. In a hardware configuration, the method according to an exemplary embodiment of the present invention can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc. In a firmware or software configuration, embodiments of the present invention can be implemented in the form of modules, processes, functions, etc. Software code can be stored in memory units and executed by a processor. The memory units are located inside or outside the processor and can send data to and receive data from the processor via various known means.
[0251] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.
[0252] Accordingly, embodiments of the present invention also provide a computer program product, including computer programs / instructions, which, when executed by a processor, implement the audio analysis method provided in embodiments of the present invention.
[0253] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art may make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. An audio analysis method, characterized in that, include: Obtain the client's interest characteristics; Obtain the target audio; Determine whether the target audio matches the client's interest characteristics; If they match, the content summary text of the target audio is obtained.
2. The audio analysis method as described in claim 1, characterized in that, Acquiring the target audio includes: acquiring real-time audio as the target audio; or, The audio from the real-time video is obtained as the target audio.
3. The audio analysis method as described in claim 1, characterized in that, Obtain the client's interest characteristics, including: providing multiple interest terms; The interest features are obtained by acquiring one or more interest terms selected by the client. Determining whether the target audio matches the client's interest features includes: determining whether the target audio matches any one or more interest terms among the interest features; If they match, the content summary text of the target audio is obtained, including: if they match, one or more matching interest terms are obtained as target terms; Obtain the content summary text of the portion corresponding to the target term in the target audio.
4. The audio analysis method as described in claim 3, characterized in that, Determining whether the target audio matches any one or more interest terms in the interest features includes: converting the target audio into the corresponding target text; Determine whether the target text matches any one or more interest terms in the interest features; Obtaining the content summary text of the part corresponding to the target term in the target audio includes: generating the content summary text of the part corresponding to the target term in the target audio based on the target text.
5. The audio analysis method as described in claim 4, characterized in that, The target audio is converted into the corresponding target text using neural network technology.
6. The audio analysis method as described in claim 5, characterized in that, Converting the target audio into corresponding target text using neural network technology includes: establishing a speech recognition model using neural network technology; The speech recognition model is used to convert the target audio into the corresponding target text.
7. The audio analysis method as described in claim 4, characterized in that, Determining whether the target text matches any one or more interest terms in the interest features includes: performing content analysis on the target text to obtain the target features of the target text; Determine whether the target feature includes any one or more interest terms among the interest features; If the target feature includes any one or more interest terms from the interest features, then the match is determined; Otherwise, the judgment is inconsistent.
8. The audio analysis method as described in claim 7, characterized in that, The target text is subjected to content analysis to obtain target features, including: segmenting the target text to obtain multiple paragraphs; Obtain the category label corresponding to each paragraph; Multiple classification labels are obtained as target features of the target text.
9. The audio analysis method as described in claim 8, characterized in that, Determining whether the target feature includes any one or more interest terms among the interest features includes: determining whether any one or more interest terms among the client's interest features are included in the plurality of classification labels; Based on the target text, generate a content summary text of the part corresponding to the target term in the target audio, including: obtaining the category tag corresponding to the target term as the target tag; Generate a summary text of the paragraph corresponding to the target tag.
10. The audio analysis method as described in claim 8, characterized in that, The classification label corresponding to each paragraph is obtained by using neural network technology.
11. The audio analysis method as described in claim 10, characterized in that, The classification label for each paragraph is obtained using neural network technology, including: establishing a classification language model using neural network technology, wherein the classification language model includes text content and classification labels with a mapping relationship; The classification language model is used to obtain the classification labels corresponding to the text content of each paragraph.
12. The audio analysis method as described in claim 11, characterized in that, A classification language model is established using neural network technology, including: providing a text dataset, wherein the text dataset includes corresponding training text content and training classification labels; The text dataset is trained using the neural network technology to obtain the classification language model; Obtaining classification labels corresponding to the text content of each paragraph using the classification language model includes: analyzing the paragraph using the classification language model to obtain training text content that matches the text content of the paragraph; The training classification labels corresponding to the training text content are obtained as the classification labels corresponding to the paragraph.
13. The audio analysis method as described in claim 9, characterized in that, The content summary text of the paragraph corresponding to the target tag is generated using neural network technology.
14. The audio analysis method as described in claim 13, characterized in that, Generating a content summary text of the paragraph corresponding to the target tag using neural network technology includes: establishing a summary language model using neural network technology, wherein the summary language model is used to generate a summary of the text content; The text content of the paragraph corresponding to the target tag is analyzed using the summarization language model to obtain the corresponding content summary text.
15. The audio analysis method as described in claim 1, characterized in that, After obtaining the content summary text of the target audio, the method further includes: transmitting the content summary text to the client.
16. The audio analysis method as described in claim 1, characterized in that, The audio analysis method is performed using a local deployment approach; or, The audio analysis method is implemented using a cloud-based deployment approach.
17. The audio analysis method as described in claim 16, characterized in that, In the audio analysis method using local deployment, a hardware accelerator is set up to perform the audio analysis.
18. The audio analysis method as described in claim 1, characterized in that, If the target audio does not match the client's interest characteristics, then the target audio is deleted.
19. An audio analysis system, characterized in that, include: The interest feature acquisition module is used to acquire the client's interest features; The target audio acquisition module is used to acquire the target audio. The judgment module is used to determine whether the target audio matches the interest characteristics of the client; The content summary text acquisition module is used to acquire the content summary text of the target audio if a match is found.
20. A device, characterized in that, It includes at least one memory and at least one processor, the memory storing one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the audio analysis method as described in any one of claims 1-18.
21. A storage medium, characterized in that, The storage medium stores one or more computer instructions, which are used to implement the audio analysis method as described in any one of claims 1-18.
22. A computer program product, comprising computer programs / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the audio analysis method according to any one of claims 1-18.
Citation Information
Patent Citations
Information push method and device based on multi-screen interaction scene
CN103517100A
Audio abstract text creation method based on speech recognition and creation device thereof
CN108305622A
Video abstract generation method and device
CN115190357A
Personal digest distribution apparatus, distribution method thereof, program thereof, and personal digest distribution system
JP2004260297A