Method, apparatus, electronic device, and readable storage medium for audio information extraction

By dividing the role and time of the audio text and extracting audio information in combination with the determination conditions, the problem of low audio information extraction efficiency and accuracy in the prior art is solved, and more efficient and accurate information acquisition is achieved.

CN114255751BActive Publication Date: 2025-07-04阳光保险集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111499605.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-07-04
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

The prior art cannot efficiently and accurately obtain target information when extracting audio information.

Method used

By dividing roles in the audio text, a dialogue set of each dialogue character is obtained, and a time division is made for each dialogue, data judgment is made based on the judgment conditions, and the target judgment content is obtained.

Benefits of technology

Improve the accuracy and efficiency of audio information extraction and reduce the error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255751B_ABST
    Figure CN114255751B_ABST
Patent Text Reader

Abstract

This application belongs to the technical field of data processing, and discloses a method, device, electronic device and readable storage medium for audio information extraction. The method includes: performing text conversion on a target audio to obtain an audio text; performing role division on the audio text to respectively obtain a dialogue set for each dialogue role, where the dialogue set contains at least one dialogue; respectively performing time division on each dialogue in the dialogue set of each dialogue role to obtain the dialogue time information of each dialogue role; and performing data determination on the audio text according to the dialogue sets and dialogue time information of each dialogue role and a determination condition to obtain target determination content. In this way, when performing audio information extraction, the efficiency and accuracy of audio information extraction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and more particularly, to a method, apparatus, electronic device, and readable storage medium for audio information extraction. Background Art

[0002] With the rapid development of the Internet, there are more and more application scenarios for audio information extraction. For example, after a phone call or voice communication with a customer, it is usually necessary to analyze the generated audio and extract content based on the audio information to better serve the customer.

[0003] However, in the prior art, when extracting audio information, it is usually impossible to efficiently and accurately obtain the information to be extracted from the audio.

[0004] Therefore, when extracting audio information, how to improve the efficiency and accuracy of audio information extraction is a technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method, apparatus, electronic device, and readable storage medium for audio information extraction, so as to improve the efficiency and accuracy of audio information extraction when performing audio information extraction.

[0006] On the one hand, a method for audio information extraction is provided, including:

[0007] Perform text conversion on the target audio to obtain an audio text;

[0008] Perform role division on the audio text to obtain a dialogue set for each dialogue role, where the dialogue set contains at least one dialogue;

[0009] Perform time division on each dialogue in the dialogue set of each dialogue role to obtain the dialogue time information of each dialogue role;

[0010] Perform data determination on the audio text according to the dialogue set and dialogue time information of each dialogue role, and a determination condition, to obtain target determination content.

[0011] In the above implementation process, performing role division on the audio text, and performing division on the dialogue time information of each dialogue role, and searching for target determination content according to the dialogue set and dialogue time information of each dialogue role, and a determination condition, can reduce the error rate, and can efficiently and accurately obtain the target determination content.

[0012] In one implementation, performing role division on the audio text to obtain a dialogue set for each dialogue role includes:

[0013] Divide the audio text according to the specified characters to obtain at least one conversation;

[0014] Use a pre-trained role division model to respectively determine the speaker of each conversation and the role probabilities of each conversation role;

[0015] For each conversation, perform the following steps:

[0016] Determine the maximum value among the role probabilities corresponding to each role of a conversation;

[0017] Add a conversation to the conversation set of the conversation role corresponding to the maximum value.

[0018] In the above implementation process, by using a pre-trained role division model to divide the roles of the audio text, in subsequent steps, audio information can be extracted based on the conversation roles, improving the accuracy of audio information extraction.

[0019] In one implementation, using a pre-trained role division model to respectively determine the speaker of each conversation and the role probabilities of each conversation role includes:

[0020] Divide at least one conversation according to the specified number of conversations to obtain at least one conversation group;

[0021] For each conversation group, perform the following steps:

[0022] Input each conversation in a conversation group into the pre-trained role division model respectively to obtain the initial role probabilities of each conversation for each conversation role respectively;

[0023] Perform weighted summation of the initial role probabilities of each conversation of the same conversation role respectively;

[0024] Determine the weighted summation result of each conversation of the same conversation role as the role probability of the corresponding conversation role in a conversation group for the corresponding conversation.

[0025] In the above implementation process, by using a pre-trained role division model to perform weighted summation of the initial role probabilities of each conversation role of each conversation respectively, in this way, the role probabilities of each conversation for the corresponding conversation role can be respectively determined, improving the accuracy of role division.

[0026] In one implementation, perform time division on each conversation in the conversation set of each conversation role to obtain the conversation time information of each conversation role, including:

[0027] Perform time division on the conversations in the conversation sets of each conversation role respectively to obtain the conversation time intervals of each conversation respectively;

[0028] According to the dialogue time intervals corresponding to each dialogue and the number of characters included in each dialogue, determine the character time intervals of each character respectively;

[0029] Obtain dialogue time information based on the dialogue time intervals of each dialogue and the character time intervals of each character.

[0030] In the above implementation process, divide the time of each dialogue to obtain the dialogue time interval of each dialogue respectively. In this way, the dialogue time information can be determined according to the dialogue time interval and the character time interval of each dialogue, improving the accuracy of the character time interval.

[0031] In one implementation manner, divide the time of each dialogue in the dialogue set of each dialogue role to obtain the dialogue time interval of each dialogue respectively, including:

[0032] If there is an overlap in the dialogue time intervals of the dialogues corresponding to at least two dialogue roles, divide the dialogue according to the specified character to obtain multiple sub-dialogues;

[0033] Obtain the dialogue time interval corresponding to each sub-dialogue respectively;

[0034] Update each dialogue time interval according to the dialogue time interval corresponding to each sub-dialogue.

[0035] In the above implementation process, divide the dialogue with an overlap in the dialogue time intervals of the dialogues corresponding to at least two dialogue roles, and obtain the dialogue time interval corresponding to each sub-dialogue. In this way, each dialogue time interval can be updated according to the dialogue time interval corresponding to each sub-dialogue, improving the accuracy of each dialogue time interval.

[0036] In one implementation manner, according to the dialogue time intervals corresponding to each dialogue and the number of characters included in each dialogue, determine the character time intervals of each character respectively, including:

[0037] For each dialogue in the audio text, perform the following steps respectively:

[0038] According to the dialogue time interval of a dialogue and the number of characters included in a dialogue, determine the average character time length of each character in a dialogue;

[0039] Divide the dialogue time interval of a dialogue according to the average character time length to obtain the character time intervals of each character respectively.

[0040] In the above implementation process, according to the obtained conversation time intervals of each conversation and the number of characters in each conversation, the character time interval of each character is determined respectively. In this way, the character time interval of each character can be determined respectively, improving the accuracy of the character time interval.

[0041] In one implementation, the determination conditions include a first determination condition and a second determination condition. The first determination condition includes a first determination element and a first conversation role. The second determination condition includes a second determination element, a second conversation role, and a set determination duration. According to the conversation sets and conversation time information of each conversation role, and the determination conditions, data determination is performed on the audio text to obtain target determination content, including:

[0042] The following steps are executed in a loop:

[0043] From the conversation set of the first conversation role, a first determination conversation set is filtered out. Among them, the first determination conversation set includes a first initial conversation and the conversations after the first initial conversation, where the initial value of the first initial conversation is the first conversation of the first conversation role.

[0044] In chronological order, according to the first determination condition, each conversation in the first determination conversation set is determined respectively.

[0045] If it is determined that the first determination content corresponding to the first determination condition is obtained, the conversation containing the first determination content is determined as the second initial conversation.

[0046] According to the set determination duration, a second determination conversation set of the second conversation role is filtered out from the conversation set of the second conversation role.

[0047] According to the second determination element, the conversations in the second determination conversation set are determined.

[0048] If it is determined that the second determination content corresponding to the second determination condition is obtained, the determination element determination process is stopped; otherwise, the first conversation of the first conversation role after the second initial conversation is determined as the new first initial conversation, where the target determination content includes the first determination content and the second determination content.

[0049] In the above implementation process, based on the conversation sets, conversation time information of each conversation role, and the determination conditions, querying the target determination content can reduce the error rate and obtain the target determination content more accurately.

[0050] On the one hand, a device for extracting audio information is provided, including:

[0051] A conversion unit for performing text conversion on the target audio to obtain an audio text.

[0052] A role division unit, which is used to divide the audio text into roles, and respectively obtain the dialogue sets of each dialogue role, where each dialogue set contains at least one dialogue;

[0053] A time division unit, which is used to respectively divide each dialogue in the dialogue set of each dialogue role to obtain the dialogue time information of each dialogue role;

[0054] A determination unit, which is used to perform data determination on the audio text according to the dialogue sets and dialogue time information of each dialogue role, as well as the determination conditions, to obtain the target determination content.

[0055] In one implementation, the role division unit is specifically used for:

[0056] Divide the audio text according to the specified characters to obtain at least one dialogue;

[0057] Use a pre-trained role division model to respectively determine the speaker of each dialogue, and respectively obtain the role probabilities of each dialogue role;

[0058] For each dialogue, perform the following steps respectively:

[0059] Determine the maximum value among the role probabilities corresponding to a dialogue;

[0060] Add a dialogue to the dialogue set of the dialogue role corresponding to the maximum value.

[0061] In one implementation, the role division unit is specifically used for:

[0062] Divide at least one dialogue according to the specified number of dialogues to obtain at least one dialogue group;

[0063] For each dialogue group, perform the following steps respectively:

[0064] Respectively input each dialogue in a dialogue group into a pre-trained role division model, and respectively obtain the initial role probabilities of each dialogue for each dialogue role;

[0065] Respectively perform weighted summation on the initial role probabilities of each dialogue of the same dialogue role;

[0066] Respectively determine the weighted summation result of each dialogue of the same dialogue role as the role probability of the corresponding dialogue role of the corresponding dialogue in a dialogue group.

[0067] In one implementation, the time division unit is specifically used for:

[0068] Perform time division on each dialogue in the dialogue sets of each dialogue role to respectively obtain the dialogue time interval of each dialogue;

[0069] According to the dialogue time intervals corresponding to each dialogue and the number of characters included in each dialogue, determine the character time intervals for each character respectively;

[0070] Obtain the dialogue time information according to the dialogue time intervals of each dialogue and the character time intervals of each character.

[0071] In one implementation, the time division unit is specifically configured to:

[0072] If there are at least two dialogue roles with overlapping dialogue time intervals, divide the dialogue according to the specified character to obtain multiple sub-dialogues;

[0073] Obtain the dialogue time intervals corresponding to each sub-dialogue respectively;

[0074] Update the dialogue time intervals of each dialogue according to the dialogue time intervals corresponding to each sub-dialogue.

[0075] In one implementation, the time division unit is specifically configured to:

[0076] For each dialogue in the audio text, perform the following steps respectively:

[0077] According to the dialogue time interval of a dialogue and the number of characters included in a dialogue, determine the average character time length of each character in a dialogue;

[0078] Divide the dialogue time interval of a dialogue according to the average character time length to obtain the character time intervals of each character respectively.

[0079] In one implementation, the determination conditions include a first determination condition and a second determination condition. The first determination condition includes a first determination element and a first dialogue role. The second determination condition includes a second determination element, a second dialogue role, and a set determination duration. The determination unit is specifically configured to:

[0080] Loop to perform the following steps:

[0081] From the dialogue set of the first dialogue role, filter out the first determination dialogue set, where the first determination dialogue set includes the first initial dialogue and the dialogues after the first initial dialogue, and the initial value of the first initial dialogue is the first dialogue of the first dialogue role;

[0082] According to the first determination condition, determine each dialogue in the first determination dialogue set in chronological order;

[0083] If it is determined to obtain the first determination content corresponding to the first determination condition, determine the dialogue including the first determination content as the second initial dialogue;

[0084] According to the set determination duration, from the conversation set of the second conversation role, screen out the second determination conversation set of the second conversation role;

[0085] According to the second determination element, determine the conversations in the second determination conversation set;

[0086] If it is determined that the second determination content corresponding to the second determination condition is obtained, stop the determination element determination process; otherwise, determine the conversation of the first conversation role after the second initial conversation as the new first initial conversation, where the target determination content includes the first determination content and the second determination content.

[0087] On the one hand, an electronic device is provided, including a processor and a memory. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the method provided in any of the above optional implementation manners of audio information extraction are run.

[0088] On the one hand, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in any of the above optional implementation manners of audio information extraction are run.

[0089] On the one hand, a computer program product is provided. When the computer program product runs on a computer, the computer is caused to execute the steps of the method provided in any of the above optional implementation manners of audio information extraction.

[0090] In the method, device, electronic device, and readable storage medium for audio information extraction provided by the embodiments of the present application, text conversion is performed on the target audio to obtain an audio text; role division is performed on the audio text to respectively obtain the conversation sets of each conversation role, and each conversation set contains at least one conversation; time division is respectively performed on each conversation in the conversation sets of each conversation role to obtain the conversation time information of each conversation role; according to the conversation sets and conversation time information of each conversation role, and the determination condition, data determination is performed on the audio text to obtain the target determination content. In this way, the accuracy of audio information extraction can be improved when performing audio information extraction.

[0091] Other features and advantages of the present application will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0093] Figure 1 It is a schematic diagram of the architecture of an audio information extraction system provided by an embodiment of the present application;

[0094] Figure 2 It is an implementation flowchart of a method for audio information extraction provided by an embodiment of the present application;

[0095] Figure 3 It is an example diagram of a method for determining the dialogue time interval of interrupting conversations provided by an embodiment of the present application;

[0096] Figure 4 It is an example diagram of a method for determining the character time interval of each character in a dialogue provided by an embodiment of the present application;

[0097] Figure 5 It is a detailed implementation flowchart of a method for audio information extraction provided by an embodiment of the present application;

[0098] Figure 6 It is a structural block diagram of a device for audio information extraction provided by an embodiment of the present application;

[0099] Figure 7 It is a schematic diagram of the structure of an electronic device in an embodiment of the present implementation manner. Specific implementation manner

[0100] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0101] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0102] First, some terms involved in the embodiments of the present application are described to facilitate the understanding of those skilled in the art.

[0103] Terminal device: It can be a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system device, a personal navigation device, a personal digital assistant, an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the terminal device can support any type of user interface (such as a wearable device), etc.

[0104] Server: It can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.

[0105] Automatic Speech Recognition (ASR): It is a technology that converts human speech into text.

[0106] Natural Language Processing (NLP): A technology for human-machine interaction communication using the natural language used by humans for communication.

[0107] Sampling rate: It is the number of samples extracted from a continuous signal and composed into a discrete signal per unit time.

[0108] Bit width: It refers to the amount of data that can be transferred by the memory or video memory at one time.

[0109] In order to improve the accuracy of audio information extraction when performing audio information extraction, the embodiments of the present application provide a method, an apparatus, an electronic device, and a readable storage medium for audio information extraction.

[0110] Refer to Figure 1As shown in the figure, it is a schematic architecture diagram of an audio information extraction system provided by an embodiment of the present application. The audio information extraction system includes a user terminal and a server.

[0111] User terminal: used to provide audio for the server.

[0112] Server: used to perform text conversion on the target audio to obtain the audio text, perform role division on the audio text to respectively obtain the dialogue sets of each dialogue role in the audio text, and respectively perform time division on each dialogue in the dialogue sets of each dialogue role to obtain the dialogue time information of each dialogue role, and based on the dialogue sets and dialogue time information of each dialogue role, and the determination condition, perform data determination on the audio text to obtain the target determination content.

[0113] In one implementation, the server receives the audio sent by the user terminal and performs text conversion on the audio. Here, the audio is the conversation between the customer and the agent. The server divides the audio text according to punctuation marks, including but not limited to commas, periods, question marks, and exclamation marks, to obtain multiple conversations. The server inputs the multiple conversations into a binary classification model to respectively obtain the initial role probabilities of each conversation being the customer and the agent, and weights and sums the respective initial role probabilities of each conversation with the same conversation role, to obtain the weighted sum results of each conversation with the same conversation role, and determines the role probability of the conversation role of each conversation as the weighted sum result. The server compares the role probabilities of each conversation, determines the maximum value among the role probabilities corresponding to each conversation, and respectively adds each conversation to the conversation set of the conversation role corresponding to the maximum value. The server divides each conversation according to commas to respectively obtain the conversation time interval of each conversation, and determines the character time interval of each character according to the number of characters in each conversation. The server uses keyword search technology to start determining from the first conversation of the audio text based on the conversation sets of each conversation role, conversation time information, the first determination condition, and the second determination condition. The first determination condition includes a first determination element and a first conversation role. Here, the first determination condition can be "asking about the occupation + agent", the first determination element is "asking about the occupation", and the first conversation role is "agent". The second determination condition includes a second determination element, a second conversation role, and a set determination duration. The second determination condition can be "answering about the occupation + customer + within 15 seconds", the second determination element is "answering about the occupation", the second conversation role is "customer", and the set determination duration is "within 15 seconds". The server filters out the first determination conversation set from the conversation set of the first conversation role, and sequentially determines each conversation in the first determination conversation set according to the first determination condition. The server obtains the first determination content corresponding to the first determination condition, and determines the conversation containing the first determination content as the second initial conversation. And according to the set determination duration, filters out the second determination conversation set of the second conversation role from the conversation set of the second conversation role, and determines the conversations in the second determination conversation set according to the second determination condition. If the second determination content corresponding to the second determination condition is obtained, the server stops the determination element determination process.

[0114] Otherwise, determine the first conversation of the first conversation role after the second initial conversation of the server as the new first initial conversation, and restart the loop determination until the target determination content is found or after the entire audio text is determined, stop the determination element determination process, where the target determination content includes the first determination content and the second determination content.

[0115] Optionally, the method of using a pre-trained model can also be adopted to determine the audio text, or other methods can be used, which are not limited herein.

[0116] In this way, when extracting audio information, the dialogue roles of the audio text can be divided, and the dialogue time information of each dialogue role can be divided. Based on the dialogue sets and dialogue time information of each dialogue role, the determination conditions and the target determination content can be determined, which can reduce the error rate and more accurately obtain the target determination content.

[0117] In the embodiments of the present application, only the server is taken as an example of the execution entity for illustration. In actual applications, the execution entity can also be other electronic devices such as terminal devices, which are not limited herein.

[0118] Refer to Figure 2 As shown, it is a flowchart of the implementation of a method for extracting audio information provided by the embodiments of the present application. The specific implementation process of this method is as follows:

[0119] Step 200: Perform text conversion on the target audio to obtain the audio text.

[0120] Specifically, after receiving the target audio, the server sets the encoding format, sampling rate, and bit width of the target audio according to the requirements of the audio conversion technology, and uses the audio conversion technology to perform text conversion on the target audio to obtain the audio text of the target audio.

[0121] Optionally, the ASR technology can be used to perform text conversion on the target audio.

[0122] In actual applications, other technologies can also be used to perform text conversion on the target audio according to the actual application situation, which are not limited herein.

[0123] In this way, the audio conversion technology can be used to perform text conversion on the target audio to obtain the audio text.

[0124] Step 201: Divide the roles of the audio text to obtain the dialogue sets of each dialogue role respectively.

[0125] Specifically, to execute step 201, the following steps can be executed:

[0126] S2011: Divide the audio text according to the specified character to obtain at least one dialogue.

[0127] Specifically, the server divides the audio text according to the specified character in the audio text to obtain one or more dialogues.

[0128] Among them, the specified character can be a comma, a period, a question mark, etc. In actual applications, the specified character can also be other characters, which are not limited herein.

[0129] S2012: Use a pre-trained role classification model to respectively determine the speakers of each conversation, and obtain the role probabilities for each conversation role respectively.

[0130] Specifically, to execute step S2012, the following steps can be adopted:

[0131] Step 1: Divide at least one conversation according to the specified number of conversations to obtain at least one conversation group.

[0132] Among them, the specified number of conversations can be one conversation or multiple conversations, and there is no restriction here.

[0133] Among them, after executing Step 1, the following steps can also be respectively executed for each conversation group:

[0134] Step A: Input each conversation in a conversation group into the pre-trained role classification model respectively to obtain the initial role probabilities of each conversation for each conversation role respectively.

[0135] Specifically, the server inputs each conversation in a conversation group into the pre-trained role classification model respectively to obtain the initial role probabilities of each conversation for each conversation role respectively.

[0136] Optionally, the pre-trained role classification model can be a binary classification model. In practical applications, the pre-trained role classification model can also be other models, and there is no restriction here.

[0137] In one implementation, the server inputs a conversation between a consultant and a customer into the binary classification model, and respectively obtains the output results of the binary classification model in the above conversation. The output results indicate that the initial probability of the conversation role being a consultant is 80%, and the initial probability of the conversation role being a customer is 20%.

[0138] Furthermore, when inputting a conversation into the pre-trained role classification model, the output result of the pre-trained role classification model can also be to perform conversation role annotation after each conversation, for example, 0 or 1.

[0139] Among them, 0 or 1 represents the conversation role. For example, 0 represents the customer, and 1 represents the consultant.

[0140] In practical applications, other forms of annotation can also be used, and there is no restriction here.

[0141] In this way, the initial role probabilities of each conversation for each conversation role can be respectively obtained according to the pre-trained role classification model.

[0142] Step B: Weighted sum the initial role probabilities of each conversation of the same conversation role respectively.

[0143] In one implementation, the server inputs four conversations between the consultant and the customer into a binary classification model. Among them, the server determines that the first two conversations are of the same conversation role, and the last two conversations are of another conversation role. The server respectively obtains the output results of the binary classification model in the above conversations. The output results show that the initial role probability of the first conversation role being the consultant is 80%, the initial role probability of the first conversation role being the customer is 20%, the initial role probability of the second conversation role being the consultant is 70%, the initial role probability of the second conversation role being the customer is 30%, the initial role probability of the third conversation role being the consultant is 10%, the initial role probability of the third conversation role being the customer is 90%, the initial role probability of the fourth conversation role being the consultant is 20%, and the initial role probability of the fourth conversation role being the customer is 80%. Among them, the weights are all 1. Then, the initial role probabilities of the first two conversations are weighted and summed respectively. The role probability of the resulting conversation role being the consultant is 75%, and the role probability of the resulting conversation role being the customer is 25%. The initial role probabilities of the last two sentences are weighted and summed respectively. Among them, the weights are all 1. The role probability of the resulting conversation role being the consultant is 15%, and the role probability of the resulting conversation role being the customer is 85%.

[0144] Optionally, the weight of each conversation can be 1 or other values, which is not limited here.

[0145] Step C: Respectively determine the role probabilities of the corresponding conversations of a conversation group for the weighted sum results of each conversation of the same conversation role.

[0146] Specifically, the server respectively determines the role probabilities of the corresponding conversations of the above conversation group for the weighted sum results of each conversation of the same conversation role.

[0147] In one implementation, the weighted sum results of the four conversations between the consultant and the customer are that the role probability of the conversation role of the first two conversations being the consultant is 75%, and the role probability of being the customer is 25%. The role probability of the conversation role of the last two conversations being the consultant is 15%, and the role probability of being the customer is 85%.

[0148] Furthermore, the server can also respectively determine the role probabilities of the corresponding conversation roles of the corresponding conversations for the average values of the formaldehyde sum results of each conversation of the same conversation role.

[0149] Among them, after executing step S2012, the server can also respectively execute the following steps for each conversation:

[0150] Step 1: Determine the maximum value among the probabilities of each role corresponding to a conversation.

[0151] Specifically, the server compares the probabilities of each role in a conversation, obtains the maximum value among the probabilities of each role, and determines the conversation role corresponding to the maximum value as the conversation role of this conversation.

[0152] In one implementation, if the role probability of the conversation role being a consultant in a conversation is 80% and the role probability of being a customer is 20%, then the role probabilities of the two conversation roles are compared, and the consultant with a role probability of 80% is determined as the conversation role of this conversation.

[0153] Optionally, other methods can also be used to determine the maximum value, which is not limited here.

[0154] In this way, the maximum value among the probabilities of each role corresponding to a conversation can be determined, and the conversation role corresponding to the role probability of the maximum value is determined as the conversation role of this conversation.

[0155] Step 2: Add a conversation to the conversation set of the conversation role corresponding to the maximum value.

[0156] Specifically, the server adds a conversation to the conversation set of the conversation role corresponding to the maximum value among the probabilities of each role in this conversation.

[0157] Among them, the conversation set contains at least one conversation.

[0158] Step 202: Perform time division on each conversation in the conversation set of each conversation role respectively to obtain the conversation time information of each conversation role.

[0159] Specifically, to execute Step 202, the following steps can be performed:

[0160] S2021: Perform time division on each conversation in the conversation set of each conversation role respectively to obtain the conversation time interval of each conversation.

[0161] Specifically, to execute Step S2021, the following steps can be adopted:

[0162] Step 1: If there is an overlap in the conversation time intervals corresponding to at least two conversation roles, divide the conversation according to the specified character to obtain multiple sub-conversations.

[0163] Specifically, if there is an overlapping part in the conversation intervals corresponding to at least two conversation roles in the audio text, the server divides the conversation according to the specified character to obtain multiple sub-conversations.

[0164] Step 2: Obtain the conversation time interval corresponding to each sub-conversation respectively.

[0165] Step 3: Update the conversation time intervals of each conversation according to the conversation time intervals corresponding to each sub-conversation.

[0166] Specifically, the server updates the conversation time intervals corresponding to each conversation according to the obtained conversation time intervals corresponding to multiple sub-conversations and according to the conversation time intervals of each sub-conversation.

[0167] In one implementation, the server sorts multiple sub-conversations according to the start time of each obtained sub-conversation.

[0168] Refer to Figure 3 As shown, it is an example diagram of a method for determining the conversation time interval of an interrupting conversation in an embodiment of the present application. The example diagram includes two conversations, which are respectively the conversations between a consultant and a customer. One conversation of the consultant has a conversation time interval of [10, 30], and one conversation of the customer has a conversation time interval of [11, 20]. The server uses four punctuation marks, namely commas, periods, question marks, and exclamation marks, as delimiters to divide the above conversation of the consultant into two sub-conversations, and regards the two sub-conversations as two new conversations. The conversation time intervals of the above two new conversations are [10, 15] and [15, 30] respectively. Then, the conversation time intervals corresponding to each of the conversations of the consultant and the customer are updated as follows: The first conversation is the first half of the consultant's conversation, with a conversation time interval of [10, 15], the second conversation is the customer's conversation, with a conversation time interval of [16, 25], and the third conversation is the second half of the consultant's conversation, with a conversation time interval of [26, 41], where the time unit is seconds.

[0169] In this way, the interrupting conversation can be re-divided into conversations, and the conversation time intervals of the divided conversations can be updated according to the start time of the divided conversations.

[0170] Optionally, sorting can also be performed according to the end time of each sub-conversation and the average conversation time of each sub-conversation.

[0171] In practical applications, sorting of each sub-conversation can also be performed according to the actual application situation, which is not limited herein.

[0172] Further, after updating the conversation time intervals of the sub-conversations, the other conversations after this conversation in the audio text are updated in chronological order of the conversation time intervals.

[0173] In this way, multiple sub-conversations can be sorted according to the conversation time intervals corresponding to the multiple sub-conversations, and the conversation time intervals corresponding to the multiple sub-conversations can be updated.

[0174] S2022: Determine the character time interval for each character respectively according to the conversation time interval corresponding to each conversation and the number of characters included in each conversation.

[0175] Specifically, the server determines the average character duration corresponding to each character in each conversation according to the conversation interval corresponding to each conversation and the number of characters included in each conversation, and determines the character time interval for each character according to the start time of the first character in each conversation and the average character duration corresponding to each character.

[0176] Among them, the number of characters included in each conversation may not include the number of non-text characters. For example, non-text characters may be punctuation marks.

[0177] In one implementation, the server determines the average character duration corresponding to each character according to the ratio between the time interval of a conversation and the number of characters included in the conversation. According to the start time of the first character in the conversation, the character time interval of the first character is from the start time to the start time plus the average character duration, and the character time interval of the second character is from the end time of the first character to the end time of the first character plus the average character duration, and so on. The character time interval for each character in a conversation can be determined respectively.

[0178] Among them, the average character duration of non-text characters in each conversation can be considered as 0 or can be considered as a negligible duration.

[0179] For example, if the character time interval of a character before a punctuation mark is [1000, 1600], then the character time interval of this character is changed to [1000, 1599], and the character time interval of this punctuation mark is [1599, 1600], and the time unit is microseconds.

[0180] Optionally, the time interval of a conversation can be represented by D, the number of characters included in a conversation can be represented by L, the average character duration corresponding to a character can be represented by G, and the start time of the first character can be represented by S. In practical applications, it can also be represented in other ways, which is not limited here.

[0181] Among them, the start time and end time in the conversation time interval and the character time interval are preferably selected as integers.

[0182] Optionally, the time unit can be seconds, milliseconds, microseconds, etc.

[0183] Refer to Figure 4As shown, it is an example diagram of a method for determining the character time interval of each character of a conversation according to an embodiment of the present application. The example diagram includes a conversation of a consultant, and the content of the conversation is "while talking", and the conversation time interval is [2300, 4700]. The server determines the character time interval of each character according to the conversation time interval corresponding to the conversation and the number of characters contained in the above conversation. The character time interval of each character is: the character time interval of the character "在" is [2300, 2700], the character time interval of the character "说" is [2700, 3100], the character time interval of the character "话" is [31000, 3500], the character time interval of the character "的" is [3500, 3900], the character time interval of the character "时" is [3900, 4300], and the character time interval of the character "候" is [4300, 4700], wherein the time unit is microseconds.

[0184] In this way, the character time interval of each character can be determined according to the dialogue time interval corresponding to the dialogue and the number of characters included in the dialogue.

[0185] S2023: Obtaining conversation time information according to the conversation time interval of each conversation and the character time interval of each character.

[0186] Specifically, the server obtains the conversation time information according to the conversation time interval of each conversation and the character time interval of each character.

[0187] Furthermore, spelling correction can be performed on the converted audio text.

[0188] Step 203: Perform data judgment on the audio text according to the dialogue set and dialogue time information of each dialogue role, as well as the judgment condition, to obtain the target judgment content.

[0189] Specifically, executing step 203, the following steps may be performed:

[0190] S2031: Filter out a first determination dialogue set from the dialogue set of the first dialogue role.

[0191] The first determined dialogue set includes the first initial dialogue and dialogues after the first initial dialogue.

[0192] The initial value of the first initial dialogue is the first dialogue of the first dialogue role.

[0193] In this way, the first determination dialogue set can be screened out from the dialogue set of the first dialogue role.

[0194] S2032: Determine each conversation in the first determination conversation set in chronological order according to the first determination condition.

[0195] Specifically, the server determines each conversation in the first determination conversation set according to the first determination condition in chronological order.

[0196] Optionally, the first determination condition can be "asking about occupation + consultant".

[0197] Among them, the first determination condition can include a first determination element and a first conversation role.

[0198] Optionally, the first determination condition can also include a time limit.

[0199] In practical applications, the first determination condition can be set according to the actual application situation, and no limitation is imposed here.

[0200] Optionally, the server can use any one or any combination of methods such as keyword search technology, regular expression matching, and NLP model prediction to determine the target determination content.

[0201] In one implementation, the server uses the method of a pre-trained NLP model to determine each conversation in the consultant conversation set according to "asking about occupation + consultant". For the conversation where the first determination content is determined, output 1, and for the conversation where the first determination content is not determined, output 0.

[0202] Among them, other content can also be output, and no limitation is imposed here.

[0203] In practical applications, other technologies can also be used for determination, and no limitation is imposed here.

[0204] In one implementation, the first determination condition can be "consultant + at all times + asking about occupation". The server determines each conversation with the role of consultant after the first initial conversation in sequence according to the first determination condition.

[0205] Among them, the determination condition includes a first determination condition and a second determination condition.

[0206] In one implementation, the first determination condition can be "consultant + at all times + asking about occupation". The server determines each conversation with the role of consultant in sequence according to the first determination condition. The first determination content can be "Consultant, [16, 23], 'Mr. Wang, may I ask if you are an explorer?'", and the conversation containing the first determination content is determined as the second initial conversation, where the time unit is seconds.

[0207] In this way, each conversation in the first determination conversation set can be determined according to the first determination condition in chronological order.

[0208] S2033: If it is determined that the first determination content corresponding to the first determination condition is obtained, the conversation including the first determination content is determined as the second initial conversation.

[0209] Specifically, after the server determines the first determination content corresponding to the first determination condition, the conversation including the first determination content is determined as the second initial conversation.

[0210] S2034: According to the set determination duration, filter out the second determination conversation set of the second conversation role from the conversation set of the second conversation role.

[0211] Specifically, within the set determination duration after the server determines the second initial conversation, the server filters out the second determination conversation set of the second conversation role from the conversation set of the second conversation role.

[0212] Among them, the set determination duration can be set to 10 seconds.

[0213] In practical applications, the set determination duration can be set according to the actual application situation, and there is no limitation here.

[0214] In this way, within the set determination duration, the conversations of the conversation role after the second initial conversation can be filtered out as the target role.

[0215] S2035: Determine the conversations in the second determination conversation set according to the second determination condition.

[0216] Specifically, the server determines each conversation in the second determination conversation set according to the second determination condition.

[0217] Among them, the second determination condition can include the conversation role, time interval, and keywords.

[0218] Optionally, the second determination condition can be "customer + set determination duration + answer occupation".

[0219] In practical applications, the second determination condition can be set according to the actual application situation, and there is no limitation here.

[0220] In one implementation, the server uses the method of a pre-trained NLP model to determine each conversation in the customer conversation set according to "ask occupation + consultant", outputs 1 for the conversation where the second determination content is determined, and outputs 0 for the conversation where the first determination content is not determined.

[0221] Among them, other content can also be output, and there is no limitation here.

[0222] In this way, the conversations of the filtered target role can be determined according to the second determination condition.

[0223] S2036: If it is determined that the second determination content corresponding to the second determination condition is obtained, stop the determination element determination process; otherwise, determine the dialogue of the first first dialogue role after the second initial dialogue as the new first initial dialogue, and execute step S2031.

[0224] Among them, the target determination content includes the first determination content and the second determination content.

[0225] Specifically, if the second determination content corresponding to the second determination condition is obtained, the server stops the determination element determination process; otherwise, the server determines the dialogue of the first first dialogue role after the second initial dialogue as the new first initial dialogue.

[0226] In this way, after determining the first determination content corresponding to the first determination condition, if the second determination content corresponding to the second determination condition is not determined within the set time, this step is looped until the target determination content is determined or the audio text is determined.

[0227] Furthermore, after the server determines the target determination content, it converts the target determination content into data formats such as JavaScript Object Notation (JSON) and Extensible Markup Language (XML) for information recording, records it in a file, database, or transmits it to other hosts.

[0228] Among them, the information recording may include information such as dialogue roles, time intervals, and target determination content.

[0229] In practical applications, other formats can also be used for recording, which is not limited here.

[0230] In this way, data determination can be performed on the audio text according to the first determination condition and the second determination condition to obtain the target determination content.

[0231] Refer to Figure 5 As shown, it is a detailed implementation flowchart of a method for extracting audio information provided by an embodiment of the present application, including:

[0232] Step 500: Perform text conversion on the target audio to obtain the audio text.

[0233] Step 501: Divide the audio text according to the specified characters to obtain at least one dialogue.

[0234] Step 502: Use the pre-trained role division model to respectively determine the speaker of each dialogue and the role probability of each dialogue role.

[0235] Step 503: Determine the maximum value among the probabilities of each role corresponding to each conversation.

[0236] Step 504: Add each conversation to the conversation set of the conversation role corresponding to the maximum value.

[0237] Step 505: Perform time division on the conversations in the conversation set of each conversation role to obtain the conversation time interval of each conversation respectively.

[0238] Step 506: Determine the character time interval of each character respectively according to the conversation time interval corresponding to each conversation and the number of characters included in each conversation.

[0239] Step 507: Obtain the conversation time information according to the conversation time interval of each conversation and the character time interval of each character.

[0240] Step 508: In chronological order, judge the first initial conversation in the first judgment conversation set one by one according to the first judgment condition.

[0241] Step 509: Judge whether the first judgment content corresponding to the first judgment condition is judged. If so, execute Step 510; otherwise, execute Step 516.

[0242] Step 510: Determine the conversation containing the first judgment content as the second initial conversation.

[0243] Step 511: Screen out the second judgment conversation set of the second conversation role from the conversation set of the second conversation role according to the set judgment duration.

[0244] Step 512: Judge the conversations in the second judgment conversation set according to the second judgment element.

[0245] Step 513: Judge whether the second judgment content corresponding to the second judgment condition is judged. If so, execute Step 516; otherwise, execute Step 514.

[0246] Step 514: Judge whether the second initial conversation is the last conversation of the audio text. If so, execute Step 516; otherwise, execute Step 515.

[0247] Step 515: Determine the conversation of the first first conversation role after the second initial conversation as the new first initial conversation, and execute Step 508.

[0248] Step 516: Stop the judgment element judgment process.

[0249] Specifically, when executing Step 500 - Step 516, for the specific steps, refer to the above Step 200 - Step 203, which will not be elaborated here.

[0250] In the embodiments of the present application, the audio text is divided into dialogue roles, and a dialogue set for each dialogue role is generated respectively. Then, according to the determination conditions, the target determination content is queried in the dialogue set of the corresponding dialogue role, which improves the efficiency of audio information extraction. The dialogue time information of each dialogue role is divided, and based on each dialogue role, dialogue time information, and determination elements, the target determination content is determined, which reduces the error rate of determination and improves the accuracy of determining the target determination content.

[0251] Based on the same inventive concept, an audio information extraction device is also provided in the embodiments of the present application. Since the principle of solving problems by the above device and equipment is similar to that of an audio information extraction method, therefore, the implementation of the above device can refer to the implementation of the method, and the repeated parts will not be described again.

[0252] As Figure 6 shown, it is a schematic structural diagram of an audio information extraction device provided by the embodiments of the present application, including:

[0253] A conversion unit 601, configured to perform text conversion on the target audio to obtain an audio text;

[0254] A role division unit 602, configured to divide the audio text into roles, and respectively obtain a dialogue set for each dialogue role, where the dialogue set includes at least one dialogue;

[0255] A time division unit 603, configured to respectively perform time division on each dialogue in the dialogue set of each dialogue role to obtain the dialogue time information of each dialogue role;

[0256] A determination unit 604, configured to perform data determination on the audio text according to the dialogue set and dialogue time information of each dialogue role, and the determination conditions, to obtain the target determination content.

[0257] In one implementation manner, the role division unit 602 is specifically configured to:

[0258] Divide the audio text according to the specified character to obtain at least one dialogue;

[0259] Adopt a pre-trained role division model to respectively determine the speaker of each dialogue, and respectively obtain the role probabilities of each dialogue role;

[0260] For each dialogue, respectively perform the following steps:

[0261] Determine the maximum value among the role probabilities corresponding to a dialogue;

[0262] Add a dialogue to the dialogue set of the dialogue role corresponding to the maximum value.

[0263] In one implementation, the role division unit 602 is specifically configured to:

[0264] Divide at least one conversation into at least one conversation group according to the specified number of conversations;

[0265] For each conversation group, perform the following steps:

[0266] Input each conversation in a conversation group into a pre-trained role division model respectively, and obtain the initial role probabilities of each conversation for each conversation role respectively;

[0267] Perform weighted summation on the initial role probabilities of each conversation of the same conversation role respectively;

[0268] Determine the role probability of the corresponding conversation role of the corresponding conversation in a conversation group respectively by the weighted summation result of each conversation of the same conversation role.

[0269] In one implementation, the time division unit 603 is specifically configured to:

[0270] Perform time division on each conversation in the conversation set of each conversation role, and obtain the conversation time interval of each conversation respectively;

[0271] Determine the character time interval of each character respectively according to the conversation time interval corresponding to each conversation and the number of characters included in each conversation;

[0272] Obtain the conversation time information according to the conversation time interval of each conversation and the character time interval of each character.

[0273] In one implementation, the time division unit 603 is specifically configured to:

[0274] If there is an overlap in the conversation time intervals corresponding to at least two conversation roles, divide the conversation according to the specified character to obtain multiple sub-conversations;

[0275] Obtain the conversation time interval corresponding to each sub-conversation respectively;

[0276] Update the conversation time intervals of each conversation according to the conversation time interval corresponding to each sub-conversation.

[0277] In one implementation, the time division unit 603 is specifically configured to:

[0278] For each conversation in the audio text respectively, perform the following steps:

[0279] Determine the average character time length of each character in a conversation according to the conversation time interval of a conversation and the number of characters included in a conversation;

[0280] Divide the conversation time interval of a conversation according to the average character time length, and obtain the character time interval of each character respectively.

[0281] In one implementation, the determination conditions include a first determination condition and a second determination condition. The first determination condition includes a first determination element and a first conversation role. The second determination condition includes a second determination element, a second conversation role, and a set determination duration. The determination unit 604 is specifically configured to:

[0282] Loop to execute the following steps:

[0283] From the conversation set of the first conversation role, screen out the first determination conversation set, where the first determination conversation set includes the first initial conversation and the conversations after the first initial conversation, and the initial value of the first initial conversation is the first conversation of the first conversation role;

[0284] According to the time sequence and the first determination condition, determine each conversation in the first determination conversation set respectively;

[0285] If it is determined that the first determination content corresponding to the first determination condition is obtained, then determine the conversation including the first determination content as the second initial conversation;

[0286] According to the set determination duration, screen out the second determination conversation set of the second conversation role from the conversation set of the second conversation role;

[0287] Determine the conversations in the second determination conversation set according to the second determination element;

[0288] If it is determined that the second determination content corresponding to the second determination condition is obtained, then stop the determination element determination process, otherwise, determine the first conversation of the first conversation role after the second initial conversation as the new first initial conversation, where the target determination content includes the first determination content and the second determination content.

[0289] Figure 7 Shows a schematic structural diagram of an electronic device 7000. Refer to Figure 7 As shown, the electronic device 7000 includes: a processor 7010 and a memory 7020. Optionally, it may further include a power supply 7030, a display unit 7040, and an input unit 7050.

[0290] The processor 7010 is the control center of the electronic device 7000, connects each component through various interfaces and lines, and executes various functions of the electronic device 7000 by running or executing software programs and / or data stored in the memory 7020, so as to perform overall monitoring of the electronic device 7000.

[0291] In the embodiments of the present application, when the processor 7010 calls the computer program stored in the memory 7020, it executes a method for extracting audio information provided by the embodiment shown in Figure 2 In the embodiments provided in

[0292] Optionally, the processor 7010 may include one or more processing units; preferably, the processor 7010 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, applications, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 7010 either. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be implemented separately on independent chips.

[0293] The memory 7020 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, various applications, etc.; the data storage area may store data created according to the use of the electronic device 7000, etc. In addition, the memory 7020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices, etc.

[0294] The electronic device 7000 further includes a power supply 7030 (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 7010 through a power management system, so as to manage functions such as charging, discharging, and power consumption through the power management system.

[0295] The display unit 7040 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 7000. In the embodiments of the present invention, it is mainly used to display the display interfaces of various applications in the electronic device 7000 and objects such as text and pictures displayed in the display interfaces. The display unit 7040 may include a display panel 7041. The display panel 7041 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0296] The input unit 7050 can be used to receive information such as numbers or characters input by the user. The input unit 7050 may include a touch panel 7051 and other input devices 7052. Among them, the touch panel 7051, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 7051).

[0297] The electronic device 7000 may further include one or more sensors, such as a pressure sensor, a gravity acceleration sensor, a proximity light sensor, etc. Of course, according to the needs in specific applications, the above-mentioned electronic device 7000 may further include other components such as a camera. Since these components are not the key components used in the embodiments of the present application, therefore, in Figure 7 it is not shown and will not be described in detail.

[0298] Those skilled in the art can understand that Figure 7 this is merely an example of an electronic device and does not constitute a limitation on the electronic device. It may include more or fewer components than those shown in the figure, or combine certain components, or different components.

[0299] In the embodiments of the present application, a readable storage medium stores a computer program. When the computer program is executed by a processor, the communication device can perform the various steps in the above embodiments.

[0300] For the convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.

Claims

1. A method for extracting audio information, characterized in that, including: performing text conversion on the target audio to obtain audio text; performing role division on the audio text to respectively obtain a dialogue set for each dialogue role, where the dialogue set contains at least one dialogue; respectively performing time division on each dialogue in the dialogue set of each dialogue role to obtain dialogue time information for each dialogue role; the step of respectively performing time division on each dialogue in the dialogue set of each dialogue role to obtain dialogue time information for each dialogue role includes: performing time division on each dialogue in the dialogue sets of each dialogue role to respectively obtain a dialogue time interval for each dialogue; respectively determining a character time interval for each character according to the dialogue time interval corresponding to each dialogue and the number of characters included in each dialogue; obtaining the dialogue time information according to the dialogue time intervals of each dialogue and the character time intervals of each character; performing data determination on the audio text according to the dialogue sets and dialogue time information of each dialogue role and a determination condition to obtain target determination content; the determination condition includes a first determination condition and a second determination condition. The first determination condition includes a first determination element and a first dialogue role. The second determination condition includes a second determination element, a second dialogue role, and a set determination duration. The step of performing data determination on the audio text according to the dialogue sets and dialogue time information of each dialogue role and the determination condition to obtain target determination content includes: repeatedly executing the following steps: screening out a first determination dialogue set from the dialogue set of the first dialogue role, where the first determination dialogue set includes a first initial dialogue and the dialogues after the first initial dialogue, and the initial value of the first initial dialogue is the first dialogue of the first dialogue role; respectively determining, according to the first determination condition in chronological order, each dialogue in the first determination dialogue set; if it is determined that first determination content corresponding to the first determination condition is obtained, then determining the dialogue containing the first determination content as a second initial dialogue; screening out a second determination dialogue set of the second dialogue role from the dialogue set of the second dialogue role according to the set determination duration; determining the dialogues in the second determination dialogue set according to the second determination element; if it is determined that second determination content corresponding to the second determination condition is obtained, then stopping the determination element determination process; otherwise, determining the first dialogue of the first dialogue role after the second initial dialogue as a new first initial dialogue, where the target determination content includes the first determination content and the second determination content.

2. The method according to claim 1, characterized in that, the step of performing role division on the audio text to respectively obtain a dialogue set for each dialogue role includes: dividing the audio text according to a specified character to obtain at least one dialogue; using a pre-trained role division model to respectively determine the speaker of each dialogue as the role probability of each dialogue role; respectively performing the following steps for each dialogue: determining the maximum value among the role probabilities corresponding to a dialogue; Add the said one conversation to the conversation set of the conversation role corresponding to the maximum value.

3. The method according to claim 2, wherein Using the pre-trained role division model, respectively determine the speaker of each conversation and the role probability of each conversation role, including: Divide the said at least one conversation according to the specified number of conversations to obtain at least one conversation group; For each conversation group respectively, perform the following steps: Input each conversation in the said one conversation group into the pre-trained role division model respectively, and respectively obtain the initial role probability of each conversation for each conversation role; Perform weighted summation on the initial role probabilities of each conversation of the same conversation role respectively; Determine the weighted summation result of each conversation of the same conversation role as the role probability of the corresponding conversation role of the corresponding conversation in the said one conversation group respectively.

4. The method according to claim 1, characterized in that, The time division of each conversation in the conversation set of each conversation role, respectively obtaining the conversation time interval of each conversation, including: If it is determined that there is an overlap in the conversation time intervals of different conversation roles, then divide the overlapping conversations according to the specified characters to respectively obtain multiple sub-conversations of the overlapping conversations; Obtain the conversation time interval corresponding to each sub-conversation respectively; Update the conversation time intervals of each conversation based on the conversation time intervals corresponding to the multiple sub-conversations.

5. The method according to claim 1, characterized in that, According to the conversation time intervals corresponding to each conversation and the number of characters included in each conversation, respectively determine the character time interval of each character, including: For each conversation in the audio text respectively, perform the following steps: According to the conversation time interval of a conversation and the number of characters included in the said one conversation, determine the average character time length of each character in the said one conversation; Divide the conversation time interval of the said one conversation according to the average character time length to respectively obtain the character time interval of each character.

6. An apparatus for extracting audio information, characterized in that, Including: A conversion unit for performing text conversion on the target audio to obtain an audio text; A role division unit for performing role division on the audio text to respectively obtain the conversation set of each conversation role, and at least one conversation is included in the conversation set; A time division unit for respectively performing time division on each conversation in the conversation set of each conversation role to obtain the conversation time information of each conversation role; The said respectively performing time division on each conversation in the conversation set of each conversation role to obtain the conversation time information of each conversation role includes: performing time division on each conversation in the conversation set of each conversation role to respectively obtain the conversation time interval of each conversation; according to the conversation time intervals corresponding to each conversation and the number of characters included in each conversation, respectively determine the character time interval of each character; according to the conversation time intervals of each conversation and the character time intervals of each character, obtain the said conversation time information; A determination unit, configured to perform data determination on the audio text according to the dialogue sets and dialogue time information of each dialogue role, and a determination condition, to obtain target determination content; the determination condition includes a first determination condition and a second determination condition, the first determination condition includes a first determination element and a first dialogue role, the second determination condition includes a second determination element, a second dialogue role, and a set determination duration, and performing data determination on the audio text according to the dialogue sets and dialogue time information of each dialogue role, and the determination condition, to obtain target determination content, includes: circularly executing the following steps: screening out a first determination dialogue set from the dialogue set of the first dialogue role, where the first determination dialogue set includes a first initial dialogue and the dialogues after the first initial dialogue, where the initial value of the first initial dialogue is the first dialogue of the first dialogue role; sequentially determining each dialogue in the first determination dialogue set according to the first determination condition in chronological order; if it is determined that the first determination content corresponding to the first determination condition is obtained, determining the dialogue including the first determination content as the second initial dialogue; screening out a second determination dialogue set of the second dialogue role from the dialogue set of the second dialogue role according to the set determination duration; determining the dialogues in the second determination dialogue set according to the second determination element; if it is determined that the second determination content corresponding to the second determination condition is obtained, stopping the determination element determination process, otherwise, determining the first dialogue of the first dialogue role after the second initial dialogue as the new first initial dialogue, where the target determination content includes the first determination content and the second determination content.

7. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-5 is run.

8. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-5 is run.

Citation Information

Patent Citations

  • Video search method and video search device

    CN106021496A

  • Interactive system, method, terminal and medium for audio character segmentation and word recognition

    CN108597521A