Data processing methods, devices, equipment, and storage media based on large models

CN119851661BActive Publication Date: 2026-08-14BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]在双录离线质检的场景中,当前解决方案主要是在音视频录制的过程中记录打点信息,以此来区分音视频中的各个阶段,换言之,以此来确定出音视频的具体流程,但是,现有很多需要质检的音视频在录制的过程中并无打点信息,使得无法获知整个音视频中的具体流程,导致无法有效进行质检

Benefits of technology

[0020]这样,本公开方案能够根据目标对话音频得到多个目标音频片段,以得到目标文本内容,进而对该目标文本内容进行意图提取,得到与问题相关的问题文本和/或与回答相关的答复文本,再根据目标对话音频中的时间信息,对所提取到的文本(与问题相关的问题文本,或与回答相关的答复文本)进一步处理,以得到带有时间信息的目标对话文本,如此,自动且精确地从目标对话音频中提取出对话文本,同时,实现了对目标对话音频的音频分段,为后续提升对音频数据的质检效率提供了有力支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851661B_ABST
    Figure CN119851661B_ABST
Patent Text Reader

Abstract

This disclosure provides a data processing method, apparatus, device, and storage medium based on a large model, relating to the field of data processing technology, particularly to the fields of artificial intelligence, big data, and large models. The specific implementation scheme is as follows: Speech activity detection is performed on target dialogue audio to obtain multiple target audio segments containing speech activity; based on the multiple target audio segments, target text content is obtained; using a large model, intent extraction is performed on the target text content, and at least one of the following is extracted: question text related to the question, and answer text related to the answer; based on the time information in the target dialogue audio and the extracted text, a target dialogue text with time information associated with the target dialogue audio is obtained, wherein the target dialogue text includes at least one of the following: a question-answer pair with time information, and question text with time information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of artificial intelligence, big data, and large models. Background Technology

[0002] Offline quality inspection is a common and widely used scenario for dual-recording quality inspection. The existing backlog of data requiring quality inspection according to dual-recording regulations is enormous. Using offline quality inspection can significantly reduce the workload of manual review, lower labor costs for quality inspection, and, to some extent, also address the risks and oversights associated with manual quality inspection.

[0003] In the scenario of offline quality inspection of dual recording, the current solution mainly records the point information during the audio and video recording process to distinguish the various stages in the audio and video. In other words, it determines the specific process of the audio and video. However, many audio and video recordings that need quality inspection do not have point information during the recording process, making it impossible to know the specific process of the entire audio and video and thus making it impossible to perform quality inspection effectively. Summary of the Invention

[0004] This disclosure provides a data processing method, apparatus, device, and storage medium based on a large model.

[0005] According to one aspect of this disclosure, a data processing method based on a large model is provided, comprising:

[0006] Speech activity detection is performed on the target dialogue audio to obtain multiple target audio segments with speech activity;

[0007] Based on the multiple target audio segments, the target text content is obtained;

[0008] Using a large model, intent extraction is performed on the target text content, and at least one of the following is extracted: question text related to the question, and response text related to the answer;

[0009] Based on the time information in the target dialogue audio and the extracted text, a target dialogue text with time information associated with the target dialogue audio is obtained, wherein the target dialogue text includes at least one of the following: a question-and-answer pair with time information, or a question text with time information.

[0010] According to another aspect of this disclosure, a data processing apparatus based on a large model is provided, comprising:

[0011] The audio segmentation unit is used to detect speech activity in the target dialogue audio and obtain multiple target audio segments with speech activity.

[0012] An audio conversion unit is used to obtain target text content based on the plurality of target audio segments;

[0013] The text processing unit is configured to use a large model to extract intent from the target text content and extract at least one of the following: question text related to the question, and response text related to the answer; based on the time information in the target dialogue audio and the extracted text, obtain target dialogue text with time information associated with the target dialogue audio, wherein the target dialogue text includes at least one of the following: question-answer pairs with time information, and question text with time information.

[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0015] At least one processor; and

[0016] The memory is communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0019] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0020] In this way, the disclosed solution can obtain multiple target audio segments from the target dialogue audio to obtain target text content, and then extract intent from the target text content to obtain question text related to the question and / or response text related to the answer. Then, based on the time information in the target dialogue audio, the extracted text (question text related to the question or response text related to the answer) is further processed to obtain target dialogue text with time information. Thus, the dialogue text is automatically and accurately extracted from the target dialogue audio. At the same time, audio segmentation of the target dialogue audio is realized, which provides strong support for improving the efficiency of subsequent quality inspection of audio data.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0023] Figure 1 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 1 ;

[0024] Figure 2 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 2 ;

[0025] Figure 3 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 3 ;

[0026] Figure 4 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 4 ;

[0027] Figure 5 This is a schematic diagram showing the result of voice activity detection of a target dialogue audio according to an embodiment of this application;

[0028] Figure 6 This is a flowchart illustrating a data processing method based on a large model according to an embodiment of this application in a specific example. Figure 1 ;

[0029] Figure 7(a) is a schematic diagram of the processing flow of a video-to-audio module according to an embodiment of the present application;

[0030] Figure 7(b) is a schematic diagram of the training and inference process of a general text classification model according to an embodiment of this application;

[0031] Figure 7(c) is a schematic diagram of speech template matching according to an embodiment of this application;

[0032] Figure 8 This is a flowchart illustrating a data processing method based on a large model according to an embodiment of this application in a specific example. Figure 2 ;

[0033] Figure 9 This is a schematic diagram of the structure of a data processing device based on a large model according to an embodiment of this application;

[0034] Figure 10 This is a block diagram of an electronic device used to implement the data processing method based on a large model according to the embodiments of this disclosure. Detailed Implementation

[0035] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0036] In this document, the term "and / or" merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document indicates any combination of at least two of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this document refer to and distinguish between multiple similar technical terms, not to restrict the order or to limit there to only two. For example, "first feature" and "second feature" refer to two categories / two features; the first feature can be one or more, and the second feature can also be one or more.

[0037] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can still be practiced even without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0038] Many audio and video recordings that require quality inspection lack data entry information (i.e., timestamps, event descriptions, and other markers added at key points in the audio and video recording). This makes it impossible to know the specific process of each stage in the entire audio and video recording process, and consequently, it is impossible to improve the efficiency of quality inspection capabilities through AI.

[0039] Based on this, the present disclosure provides a data processing method that can automatically and accurately extract dialogue text from audio or video data without marker information, thereby determining the current stage and enabling segmented markers. This provides strong support for improving the efficiency of quality inspection of audio or video data.

[0040] Specifically, Figure 1 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0041] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes:

[0042] Step S101: Perform speech activity detection on the target dialogue audio to obtain multiple target audio segments with speech activity.

[0043] For example, in one instance, after performing speech activity detection on the target dialogue audio, the target dialogue audio can be segmented based on the detection results to obtain multiple target audio segments with relatively "active" speech activity.

[0044] Step S102: Based on the multiple target audio segments, obtain the target text content.

[0045] Step S103: Using a large model, extract intent from the target text content and extract at least one of the following: question text related to the question, and response text related to the answer.

[0046] Step S104: Based on the time information in the target dialogue audio and the extracted text, obtain the target dialogue text with time information associated with the target dialogue audio.

[0047] Here, the target dialogue text includes at least one of the following: a question-and-answer pair with time information, or a question text with time information.

[0048] For example, in one example, a question-and-answer pair with time information could specifically include a question text with time information and an answer text with time information that is related to the question text. For instance, a question-and-answer pair with time information might specifically be:

[0049] "0:00:00~0:00:16.235714: In order to regulate the sales behavior of sales personnel and to better protect your legitimate rights and interests, in accordance with relevant regulations, we will record key aspects of my sales process using audio and video recording. Do you agree?"

[0050] "0:00:16.235714~0:00:17.219000: Agree, agree."

[0051] In this way, the disclosed solution can obtain multiple target audio segments from the target dialogue audio to obtain target text content, and then extract intent from the target text content to obtain question text related to the question and / or response text related to the answer. Then, based on the time information in the target dialogue audio, the extracted text (question text related to the question or response text related to the answer) is further processed to obtain target dialogue text with time information. Thus, the dialogue text is automatically and accurately extracted from the target dialogue audio. At the same time, audio segmentation of the target dialogue audio is realized, which provides strong support for improving the efficiency of subsequent quality inspection of audio data.

[0052] It should be noted that there are relatively obvious pauses in actual dialogue scenarios. Therefore, the proposed solution performs speech activity detection on the target dialogue audio to obtain the audio activity portion. For example, the Voice Activity Detection (VAD) algorithm can be used to detect speech activity in the target dialogue audio. This provides strong support for the subsequent accurate acquisition of the target text content, while also improving the efficiency of text generation and further enhancing the accuracy of audio segmentation.

[0053] Here, in one example, the target dialogue audio may specifically be audio data recorded using an audio acquisition device, such as audio data obtained after recording using an audio acquisition device in a dual-recording scenario.

[0054] Alternatively, in another example, the target dialogue audio can be audio data obtained by extracting audio from video captured by a video capture device, such as audio data obtained by extracting audio from video captured by a video capture device in a dual-recording scenario. For example, in one example, the FastForward moving picture experts group (FFmpeg) tool is used to extract audio from the recorded video to obtain the target dialogue audio.

[0055] Furthermore, in one example, after extracting audio data from the video, the audio data can be converted into a format, such as converting the extracted audio data into a Waveform Audio File Format (WAV), thereby obtaining the target dialogue audio that meets the preset audio format requirements. This facilitates the subsequent rapid transcription of the target dialogue audio into text.

[0056] It is understandable that this public solution does not impose specific restrictions on the source or extraction method of the target dialogue audio.

[0057] It should be noted that the "dual recording" scenario referred to in this public solution can specifically refer to the scenario of recording the voice or video of both parties in a conversation. For example, in the scenario of a salesperson and a customer conducting business or transactions, the conversation content and operation process between the salesperson and the customer can be recorded by a data acquisition device (such as a video acquisition device or an audio acquisition device).

[0058] In a specific example, intent extraction can be performed as follows: specifically, the intent extraction of the target text content described above (e.g., step S103) can specifically include: using a large model (e.g., a Large Language Model (LLM)) to extract intent from the target text content. This allows for efficient and accurate extraction of question and answer texts from the target text content, significantly improving data processing efficiency and ensuring the accuracy of data extraction.

[0059] Figure 2 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 2 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.

[0060] Furthermore, the method includes at least a portion of the following: For example... Figure 2 As shown, it includes:

[0061] Step S201: Perform speech activity detection on the target dialogue audio to obtain multiple target audio segments with speech activity.

[0062] It should be noted that the relevant content regarding the target dialogue audio and the target audio segment can be found in the example above, and will not be repeated here.

[0063] Step S202: Based on the multiple target audio segments, obtain the text segment of each target audio segment.

[0064] For example, in one instance, a speech recognition model (such as a convolution-augmented transformer, or Conformer) can be used to perform speech recognition on each target audio segment, obtaining text segments for each target audio segment. Furthermore, to facilitate the rapid acquisition of target dialogue text with time information later, the text segments of each target audio segment can also include time information.

[0065] Step S203: Denoise the text segments of each target audio segment to obtain the denoised text segments of each target audio segment.

[0066] Step S204: Concatenate the text segments of each denoised target audio segment to obtain the target text content.

[0067] Step S205: Using a large model, extract intent from the target text content and extract at least one of the following: question text related to the question, and response text related to the answer.

[0068] Step S206: Based on the time information in the target dialogue audio and the extracted text, obtain the target dialogue text with time information associated with the target dialogue audio.

[0069] Here, the target dialogue text includes at least one of the following: a question-and-answer pair with time information, or a question text with time information.

[0070] In this way, the disclosed solution can perform noise reduction processing on the text segments of each target audio segment after obtaining the text segments of each target audio segment. This effectively improves the text quality of the obtained target text content, provides data support for subsequent intent extraction, and also provides strong support for improving the efficiency of audio data quality inspection.

[0071] It should be noted that the text segment of the target audio segment mentioned above can be regarded as a dialogue scenario, such as a dialogue scenario between a salesperson and a customer. The dialogue content in this scenario may contain audio data that is not related to business content. Therefore, in order to improve the efficiency of subsequent quality inspection, noise reduction processing can be performed after obtaining the text segment of the target audio segment to remove irrelevant content.

[0072] Specifically, in a specific example, denoising can be performed in the following manner; specifically, the above-described denoising process on the text segments of each target audio segment to obtain the denoised text segments of each target audio segment (e.g., step S203) can specifically include:

[0073] Step S203-1: Using the first model, process the text segment of the target audio segment to determine whether there are any statements in the text segment that are unrelated to the target business scenario.

[0074] Here, the target business scenario refers to the business scenario to which the target dialogue audio is intended.

[0075] Step S203-2: If it is determined that there are statements that are irrelevant to the target business scenario, delete the statements that are irrelevant to the target business scenario from the text segment of the target audio segment to obtain the text segment of the denoised target audio segment.

[0076] Here, it can be understood that if it is determined that there are no statements unrelated to the target business scenario, in other words, all statements in the text segment of the target audio segment are related to the target business scenario, then there is no need to perform deletion processing.

[0077] In this way, the disclosed solution can use the first model to identify statements that are irrelevant to the target business scenario from the text segments of each target audio segment, and then delete such statements. Thus, the text segments of each target audio segment are denoised, which effectively improves the text quality of the obtained target text content and lays the foundation for improving the quality inspection efficiency in the future, such as improving the quality inspection efficiency of dual recording data.

[0078] Furthermore, in one example, the above-described method of using the first model to process the text segment of the target audio segment to determine whether there are statements in the text segment that are unrelated to the target business scenario (e.g., step S203-1) can specifically include:

[0079] Step S203-1-1: Using the first model, classify the sentences in the text segment of the target audio segment to obtain the classification result of the text segment of the target audio segment.

[0080] Here, the classification result indicates that the statements in the text fragment belong to one of the following: noise, question, or answer; noise indicates that it is irrelevant to the target business scenario; question indicates a question related to the target business scenario; and answer indicates a response related to the target business scenario.

[0081] Step S203-1-2: Based on the classification results of the text segments for the target audio segment, determine whether there is noise in the text segments.

[0082] In other words, the first model can be a pre-trained classification model, such as a multi-classification model. Furthermore, in this example, the classification ability of the first model is used to determine the category of the sentences in the text segment of the target audio segment. If it is determined that there are sentences in the category of noise, the sentences in the category of noise are deleted. In this way, the foundation is laid for improving the text quality of the obtained target text content, and at the same time, it also lays the foundation for improving the quality inspection efficiency, such as improving the quality inspection efficiency of dual-recording data.

[0083] Figure 3 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 3This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 and Figure 2 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.

[0084] Furthermore, the method includes at least a portion of the following: For example... Figure 3 As shown, it includes:

[0085] Step S301: Perform speech activity detection on the target dialogue audio to obtain multiple target audio segments with speech activity.

[0086] It should be noted that the relevant content regarding the target dialogue audio and the target audio segment can be found in the example above, and will not be repeated here.

[0087] Step S302: Based on the multiple target audio segments, obtain the text segment of each target audio segment.

[0088] It should be noted that the relevant content of the text segment about the target audio segment can be referred to the above example, and will not be repeated here.

[0089] Step S303: Denoise the text segments of each target audio segment to obtain the denoised text segments of each target audio segment.

[0090] Step S304: Concatenate the text segments of each denoised target audio segment to obtain the target text content.

[0091] Step S305: Using a large model, perform intent extraction on the target text content and extract at least one of the following: question text related to the question, and response text related to the answer.

[0092] Step S306: Based on the time information in the target dialogue audio and the extracted text, obtain the target dialogue text with time information associated with the target dialogue audio.

[0093] Here, the target dialogue text includes at least one of the following: a question-and-answer pair with time information, or a question text with time information.

[0094] Step S307: Determine the question-and-answer segments involved in the text segments of each target audio segment after noise reduction.

[0095] Here, step S307 can be executed after step S303, or it can be executed after the target dialogue text is obtained. The present disclosure does not specify the execution order of step S307.

[0096] For example, in one scenario, the text segment of the target audio segment can be vectorized first. For instance, the text segment of the target audio segment can be input into a vectorization service engine to obtain the vectorized text content. Secondly, in the vector index library, the K (integers greater than or equal to 1) recall results with the highest semantic similarity to the vectorized text content can be retrieved. Finally, based on the K recall results, the discourse template corresponding to the text segment of the denoised target audio segment can be determined. Then, based on the discourse template corresponding to the text segment of the denoised target audio segment, the question-and-answer process involved in the text segment of the denoised target audio segment can be determined.

[0097] Here, the vector index library is constructed based on data obtained by vectorizing multiple pre-set script templates.

[0098] Step S308: Merge the content of the associated question-and-answer segments in the target dialogue text to obtain the target dialogue text marked with question-and-answer segments.

[0099] For example, content belonging to the same question-and-answer segment in the target dialogue text can be merged to obtain target dialogue text labeled with question-and-answer segments. Alternatively, content from question-and-answer segments with hierarchical dependencies in the target dialogue text can be merged to obtain target dialogue text labeled with question-and-answer segments. The above are merely illustrative examples; in practical applications, merging can be performed based on scenario requirements, and this disclosure does not impose specific limitations on this.

[0100] In this way, the disclosed solution can automatically and efficiently obtain the target dialogue text marked with question-and-answer segments and time information based on the target dialogue audio. This effectively solves the problem that existing audio and video that need quality inspection cannot determine the specific process due to the lack of marker information, and thus cannot effectively carry out artificial intelligence (AI) quality inspection. At the same time, it also effectively avoids the problems of high cost and oversight of human quality inspection caused by the inability to conduct AI quality inspection, thereby laying the foundation for improving quality inspection efficiency, such as improving the quality inspection efficiency of dual-recorded data.

[0101] Figure 4 This is an illustrative flowchart of a data processing method based on a large model according to an embodiment of this application. Figure 4 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1 to 3 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.

[0102] Furthermore, the method includes at least a portion of the following: For example... Figure 4 As shown, it includes:

[0103] Step S401: Perform speech activity detection on the target dialogue audio to obtain multiple initial audio segments with speech activity.

[0104] For example, in one instance, the Voice Activity Detection (VAD) algorithm is used to detect the target dialogue audio, resulting in... Figure 5 The detection results shown can include multiple speech activity intervals, such as speech activity interval 1, speech activity interval 2, ..., speech activity interval 7. Based on these speech activity intervals, the target dialogue audio can be segmented to obtain initial audio segments corresponding to each speech activity interval. These initial audio segments are then saved in WAV format. For example, the following seven initial audio segments can be obtained:

[0105] Initial audio segment 1: 001_0_17219.wav;

[0106] Initial audio segment 2: 002_18770_39100.wav;

[0107] Initial audio segment 3: 003_39410_45959.wav;

[0108] Initial audio segment 4: 004_48619_50519.wav;

[0109] Initial audio segment 5: 005_56849_57639.wav;

[0110] Initial audio segment 6: 006_58719_59299.wav;

[0111] Initial audio clip 7: 007_60219_75700.wav.

[0112] It should be noted that the naming format of the initial audio segment is idx_start_end.wav, where idx represents the sequential number of the initial audio segment, start represents the start time of the initial audio segment in the target dialogue audio, and end represents the end time of the initial audio segment in the target dialogue audio.

[0113] Step S402: Determine the initial audio segments whose frequencies fall within the target frequency range from the plurality of initial audio segments, and use the initial audio segments that fall within the target frequency range as the target audio segments.

[0114] It is understandable that the target dialogue audio may contain non-human voice noise data (such as keyboard typing sounds, traffic noise, etc.). Therefore, in order to further improve the accuracy of the obtained target text content and avoid noise interference, the present invention can also perform non-human voice filtering on multiple initial audio segments, for example, by using the target frequency range to perform non-human voice filtering on multiple initial audio segments.

[0115] Furthermore, in one example, the target frequency range mentioned above can be specifically the frequency range of human voices. Furthermore, the frequency values ​​of the initial audio segment can be specifically the average or median of the frequency values ​​corresponding to each time frame in the initial audio segment.

[0116] For example, taking 7 initial audio segments as an example, firstly, the average frequency value of the initial audio segment is obtained by taking the frequency value corresponding to each time frame in the initial audio segment. In this way, the average frequency value of each initial audio segment is obtained. Secondly, from the 7 initial audio segments, the initial audio segments whose average frequency value falls within the target frequency range are determined to be: initial audio segment 1, initial audio segment 2, initial audio segment 3, and initial audio segment 7. Finally, initial audio segment 1 and initial audio segment 2, which fall within the target frequency range, are taken as target audio segment 2, and initial audio segment 3 and initial audio segment 7, which also fall within the target frequency range, are taken as target audio segments.

[0117] Alternatively, in another example, it can be determined whether the audio intensity falling within the target frequency range in the initial audio segment is greater than or equal to a preset threshold. If so, it indicates that the initial audio segment contains human voices, and in this case, the initial audio segment can be used as the target audio segment; otherwise, that is, if the audio intensity falling within the target frequency range in the initial audio segment is less than the preset threshold, it indicates that the initial audio segment only contains non-human voice noise, and in this case, the initial audio segment needs to be removed. In this way, non-human voice filtering of multiple initial audio segments can be achieved.

[0118] Furthermore, in one example, the target frequency range can be determined before identifying initial audio segments whose frequencies fall within the target frequency range from the plurality of initial audio segments (e.g., before step S402), specifically including:

[0119] Determine the spectral information corresponding to the time frames contained in each initial audio segment;

[0120] Based on the spectral information corresponding to the time frames contained in each initial audio segment, the frequency values ​​corresponding to the time frames contained in each initial audio segment are obtained.

[0121] The target frequency range is determined based on the frequency values ​​corresponding to the time frames contained in each initial audio segment.

[0122] Thus, this disclosed solution provides a refined method for obtaining the target audio range. This method is simple, efficient, and can quickly obtain the target audio range, thereby providing strong support for obtaining highly accurate target audio segments. At the same time, it also lays the foundation for improving the text quality of the obtained target text content, and further lays the foundation for improving quality inspection efficiency, such as improving the quality inspection efficiency of dual-recording data.

[0123] Step S403: Based on the multiple target audio segments, obtain the target text content.

[0124] Step S404: Using a large model, perform intent extraction on the target text content and extract at least one of the following: question text related to the question, and response text related to the answer.

[0125] Step S405: Based on the time information in the target dialogue audio and the extracted text, obtain the target dialogue text with time information associated with the target dialogue audio.

[0126] Here, the target dialogue text includes at least one of the following: a question-and-answer pair with time information, or a question text with time information.

[0127] Thus, this disclosed solution provides a refined method for extracting multiple target audio segments from target dialogue audio. This method can easily and efficiently extract high-quality target audio segments, thereby providing strong support for the subsequent accurate extraction of key text information and the resulting high-quality dialogue text. Furthermore, for audio or video data without punctuation information, this disclosed solution can still automatically and accurately extract the dialogue text and achieve segmented punctuation, thus providing strong support for improving the efficiency of subsequent quality inspection of audio or video data.

[0128] The following description, with specific examples and accompanying figures, further illustrates the present disclosure. This disclosure provides a video (or audio) segmentation method based on a large model. For example, in this example, the segmentation can be functionally divided into the following modules: video-to-audio conversion module, speech recognition module, dialogue denoising module, intent extraction module, time backtracking module, speech matching module, and segment merging module. Both the intent extraction module and the speech matching module can be implemented using large model technology. Figure 6 As shown, the modules are interconnected to form a complete end-to-end video segmentation service, which can greatly improve the accuracy and efficiency of video review in dual-recording offline quality inspection scenarios.

[0129] Specifically, such as Figure 6As shown, the specific solutions disclosed herein include:

[0130] (I) Video to Audio Module

[0131] The video-to-audio module is mainly used to extract audio from the video to be inspected in order to obtain the dialogue audio of the video to be inspected (that is, the target dialogue audio mentioned above). For example, the Fast Forward Image Experts Group (FFmpeg) tool can be used to extract audio from the video to be inspected in order to obtain the target dialogue audio.

[0132] It should be noted that in real-world scenarios, there may be situations where the image stream in the video to be inspected is out of sync with the audio stream in the dialogue audio. In such cases, the time information of the image stream in the video to be inspected and the time information of the audio stream in the target dialogue audio can be obtained. Then, based on the time interval between the time information of the image stream and the time information of the audio stream, the image stream in the video and the audio stream in the dialogue audio can be aligned.

[0133] Here, after obtaining the target dialogue audio, the audio format can be converted to the desired format, such as 16kHz sampling rate, mono, or WAV. This facilitates subsequent transcription of the dialogue audio.

[0134] (II) Speech Recognition Module

[0135] The video-to-audio module is mainly used to process the target dialogue audio to obtain multiple text segments with timestamps (i.e., text segments corresponding to the target audio segments mentioned above). The speech recognition module can specifically include three sub-modules: audio segmentation, non-human voice filtering, and audio transcription. This allows for the separation of human voice pauses in the target dialogue audio, providing strong support for subsequent dialogue scene construction. Specifically, as shown in Figure 7(a), it includes:

[0136] (1) Audio cut-off module

[0137] In actual dual-recording scenarios, since there are usually noticeable pauses in the text-to-speech (TTS) broadcast (such as the conversation between the account manager and the customer), the target conversation audio can be segmented to obtain multiple audio segments (that is, the multiple initial audio segments mentioned above) based on this audio characteristic.

[0138] For example, using the Voice Activity Detection (VAD) algorithm, voice activity can be detected in the target dialogue audio to obtain results such as... Figure 5The detection results shown are used to segment the target dialogue audio. The initial audio segments obtained after segmentation are saved in WAV format to obtain multiple initial audio segments in WAV format (for example, the audio segments shown below). At the same time, the timestamp of each audio segment in the target dialogue audio is recorded.

[0139] The initial audio segments obtained after segmentation are as follows:

[0140] 001_0_17219.wav;

[0141] 002_18770_39100.wav;

[0142] 003_39410_45959.wav;

[0143] 004_48619_50519.wav;

[0144] 005_56849_57639.wav;

[0145] 006_58719_59299.wav;

[0146] 007_60219_75700.wav;

[0147] 008_76270_77450.wav;

[0148] 009_88150_93420.wav;

[0149] 010_94550_95830.wav;

[0150] 011_97380_127310.wav;

[0151] 012_128540_131400.wav;

[0152] 013_132040_132810.wav;

[0153] 014_132810_166999.wav;

[0154] 015_166999_169309.wav;

[0155] 016_170069_171979.wav.

[0156] It should be noted that the naming format of the above initial audio segment is idx_start_end.wav, where idx represents the sequential number of the initial audio segment, start represents the start time of the initial audio segment in the target dialogue audio, and end represents the end time of the initial audio segment in the target dialogue audio.

[0157] (2) Non-human voice filtering submodule

[0158] In actual dual-recording scenarios, the audio and video recording process may take place in an open environment, so there will be noise such as non-human voices (e.g., keyboard typing, traffic noise). In other words, the initial audio segment obtained after the above segmentation may contain non-human voices. Based on this, a non-human voice filtering submodule can be used to remove non-human voices from the initial audio segment obtained above, so as to achieve non-human voice filtering of the initial audio segment.

[0159] Specifically, firstly, the frequency of human voice is calculated using the Fast Fourier Transform. For example, the specific steps are as follows:

[0160] Step a: Divide the initial audio segment into multiple time frames;

[0161] Step b: Calculate the spectral information of each time frame using Fast Fourier Transform (ratio).

[0162] (e.g., amplitude spectrum and power spectrum, etc.);

[0163] Step c: For a given time frame, determine the spectral information of each element in that time frame.

[0164] The initial frequency value corresponding to the frequency point;

[0165] Step d: From the multiple initial frequency values ​​corresponding to the time frame, determine the location of the human...

[0166] At least one target frequency value within the sound frequency range;

[0167] At this point, repeat steps c and d above to obtain the target frequency corresponding to each time frame.

[0168] Rate value;

[0169] Step e: Take the average of all target frequency values ​​as the human voice frequency of the initial audio segment.

[0170] In this way, the frequency range of human voices can be obtained (corresponding to the target frequency range mentioned above). Then, based on this frequency range, the overall intensity of human voices in the initial audio signal can be evaluated. For example, it can be determined whether the overall intensity of human voices falling within the frequency range in the initial audio signal is greater than a preset threshold. If so, the initial audio signal is considered to contain human voices; otherwise, it can be considered that the initial audio information does not contain human voices. This processing method can effectively determine whether there are human voices in an audio segment and eliminate interference from non-human voice noise, thus effectively improving transcription accuracy and speech recognition performance.

[0171] For example, by filtering out non-human voices from each initial audio segment obtained after audio segmentation, the following target audio segment is obtained:

[0172] 001_0_17219.wav;

[0173] 002_18770_39100.wav;

[0174] 003_39410_45959.wav;

[0175] 007_60219_75700.wav;

[0176] 009_88150_93420.wav;

[0177] 011_97380_127310.wav;

[0178] 012_128540_131400.wav;

[0179] 014_132810_166999.wav;

[0180] 015_166999_169309.wav;

[0181] 016_170069_171979.wav.

[0182] (3) Audio transcription submodule

[0183] The audio transcription submodule is primarily used to transcribe non-human voice filtered target audio segments using an Automatic Speech Recognition (ASR) model, obtaining text segments for each target audio segment. Furthermore, it can assign corresponding timestamps to the text segments of each target audio segment based on the timestamps (such as start and end times) recorded in the audio segmentation submodule.

[0184] For example, the target audio segments obtained above can be transcribed using the Conformer model to obtain text segments with timestamps. For instance, transing the target audio segments above yields the following text segments:

[0185] 0:00:00~0:00:17.219000: To standardize the sales behavior of sales personnel and to better protect...

[0186] To protect your legal rights, in accordance with relevant regulations, we will record key aspects of my sales process.

[0187] Please record this. Do you agree? Yes, yes.

[0188] 0:00:18.770000~0:00:39.100000: Please note that this audio and video recording process is for...

[0189] It will be crucial for you to protect your rights in the future. Please carefully read the details of the documents you sign and answer the relevant questions truthfully.

[0190] Question: If I make any promises to you that are inconsistent with the contents of the written documents, I suggest you consult with me in writing.

[0191] This form of confirmation is to better protect your legitimate rights and interests.

[0192] 0:00:39.410000~0:00:45.959000: Please send the tester for route B, gmh test, to test the route.

[0193] Show your qualification certificate.

[0194] 0:01:00.219000~0:01:15.700000: Hello Mr. / Ms. ××, before confirming your purchase, in order to protect your legal rights, we need to collect your facial and identity information for personal verification. Please hold your ID card with the back facing you.

[0195] The camera completes the collection of identity information.

[0196] 0:01:28.150000~0:01:33.420000: Please hold your ID card and face the camera to complete the identity verification.

[0197] Information collection.

[0198] 0:01:37.380000 0:02:07.310000: You are about to purchase ×× Test B-channel ×× product.

[0199] Product number 01462, high risk, high income, say goodbye to working and take a step further, purchase amount is 20.

[0200] This product, priced at 34,567 yuan, is a wealth management plan from a subsidiary of ×× Wealth Management, not from ×× Savings or...

[0201] The wealth management product and its operating entity are Earth's First Wealth Management Group and XX Financial, which is not the product's distributor.

[0202] Are you aware that we bear the investment management responsibility behind these products?

[0203] 0:02:08.540000~0:02:11.400000: Understood.

[0204] 0:02:12.810000~0:02:46.999000: You need to purchase a wealth management product called "Three Books" from this bank.

[0205] This product is a high-risk, high-return wealth management product, offering a step up from your current job. Our bank's risk rating for this product is R-5, which represents the highest risk, with very large fluctuations in returns and a very high risk of principal loss. Corresponding wealth management products include private equity funds. Your risk rating in our bank is A-2, which is considered conservative and below R-5. Investment expectations...

[0206] Investment period and risk tolerance, investment objectives, expected investment goals. Would you like to make your own investment decision?

[0207] I believe this product is perfectly suited to my needs.

[0208] 0:02:46.999000~0:02:49.309000: Do you want to confirm?

[0209] 0:02:50.069000~0:02:51.979000: Confirmed.

[0210] (III) Dialogue Denoising Module

[0211] In practical dual-recording scenarios, the text segments of the target audio clips can be viewed as multi-turn dialogues between the account manager and the customer. Theoretically, each round of dialogue follows the prescribed scripts for dual-recording, which are related to business content. However, because the recording process is in an open environment, some dialogue unrelated to business content may be recorded during the normal question-and-answer session. Therefore, this proposed solution utilizes a dialogue denoising module to filter the text segments of each target audio clip obtained from the audio transcription, thereby achieving the goal of dialogue denoising.

[0212] Specifically, the dialogue denoising module is mainly used to process the text segments of the target audio segment through a classification model to determine whether there are statements in the text segment that are unrelated to the business content.

[0213] Here, taking the Universal Text Classification (UTC) model as an example, as shown in Figure 7(b), the specific steps for denoising using this model include:

[0214] (1) Offline: Model Training

[0215] Step a: Use the document annotation for collaborative optimization (doccano) tool to annotate each text data in the text dataset of the dialogue scenario so that each text data has a corresponding label, thus obtaining a labeled text dataset.

[0216] Here, the labels include questions, answers, and noise. The "question" label indicates that the current text data belongs to the text of the account manager broadcasting the product or the text of the question raised to the customer. The "answer" label indicates that the current text data is the text of the customer's reply. The "noise" label indicates that the current text data is irrelevant to the current business scenario.

[0217] Step b: Divide the labeled text dataset into training set, validation set and test set according to a preset ratio (e.g., 8:1:1).

[0218] Step c: Train the UTC model using the training set, and evaluate the trained model using the validation set and test set. Then, fine-tune the parameters of the UTC model based on the evaluation results until a UTC model whose evaluation results meet the preset requirements is obtained.

[0219] Step d: Utilize containerization technology to deploy the trained UTC model on a server and combine it with a backend service framework to provide inference services.

[0220] In practical inference scenarios, a Fast Application Programming Interface (Fast API) backend framework developed in a specific programming language (such as Python) can also be used to provide inference services and enable GPUs to improve inference performance.

[0221] (2) Online: Model Inference

[0222] Step e: Input the text segment obtained from the audio transcription into the trained UTC model to classify the sentences in the text segment and obtain the classification result of the text segment. Then, based on the classification result of the text segment, determine whether the sentences in the text segment belong to "noise". If the sentences in the text segment belong to "noise", the sentences can be deleted from the text segment or the text segment containing the sentences can be deleted. In this way, by traversing each text segment, the text segments after dialogue denoising can be obtained.

[0223] (iv) Intent Extraction Module

[0224] In actual dual-recording scenarios, in addition to conversations unrelated to business content, there are also instances of customers answering questions prematurely (i.e., customers answering the account manager's questions before the account manager has finished speaking the script prescribed for the dual-recording scenario). This makes it impossible to find obvious pauses in the voice when segmenting the audio of the conversation, thus making it impossible to separate the audio related to the customer's response from the audio related to the account manager's questions.

[0225] To address the aforementioned issues, this proposed solution utilizes an intent extraction module, such as a large model, to extract question text related to the questions and response text related to the answers from the multi-turn dialogue content between the account manager and the customer.

[0226] Here, the task of the large model is to distinguish between text that belongs to "questions" and text that belongs to "answers". For example, the large model can roughly divide the text in the text into three categories: "question", "answer", and "please do an action". "Question" means that the text extracted is a question from the account manager; "answer" means that the text extracted is a reply from the customer; and "please do an action" means that the text extracted is a question raised by the account manager that does not require a verbal answer from the customer, but requires a response through an action.

[0227] Specifically, the steps for intention extraction include:

[0228] Step a: Merge the text segments of each denoised target audio segment according to the time sequence to obtain the merged text content (that is, the target text content mentioned above).

[0229] Step b: Concatenate the merged target text content into the initial prompt template to obtain the target prompt template. Input the target prompt template into the large model to obtain the intent extraction results. From the intent extraction results, mark the question text belonging to "question" or "please do an action" as "Q" and the reply text belonging to "answer" as "A".

[0230] Here, to guide the large model to extract intent from text content, a suitable prompt can be designed. For example, in one example, the initial prompt could be:

[0231] "prompt": "You are a professional financial analyst who will follow the requirements below from the dialogue information"

[0232] Some key information was extracted from it. The [dialogue information] section contains questions from the account manager and the customer during the dual recording process.

[0233] In answering the dialogue text, your task is to pay attention to the question-and-answer information in the dialogue, such as the account manager's inquiry.

[0234] Questions, customer responses, etc.; able to understand the meaning behind the dialogue, including the intentions represented by the customer's answers.

[0235] Wait. \n[Dialogue Information]\n{content}\n[Example Answer]\n{template}\n\n[Answer Restrictions]\n1. Involving

[0236] The information provided must be true and can be found in the [dialogue information]. Fabrication is not allowed.\n2. Do not include in the output.

[0237] HTML tags such as '”'

[0238] Furthermore, by concatenating the text content of Example 1 into the initial prompt template, the target prompt template is obtained. After inputting this target prompt template into the large model, the following extraction results are obtained:

[0239] The text content of Example 1:

[0240] 0:00:00~0:00:17.219000: To standardize the sales behavior of sales personnel and to better protect...

[0241] To protect your legal rights, in accordance with relevant regulations, we will record key aspects of my sales process.

[0242] Please record this. Do you agree? Yes, yes.

[0243] 0:00:18.770000~0:00:39.100000: Please note that this audio and video recording process is for...

[0244] It will be crucial for you to protect your rights in the future. Please carefully read the details of the documents you sign and answer the relevant questions truthfully.

[0245] Question: If I make any promises to you that are inconsistent with the contents of the written documents, I suggest you consult with me in writing.

[0246] This form of confirmation is to better protect your legitimate rights and interests.

[0247] 0:00:39.410000~0:00:45.959000: Please have the tester on route B, gmh test, demonstrate.

[0248] Qualification certificate.

[0249] 0:01:00.219000~0:01:15.700000: Hello Mr. / Ms. ××, before confirming your purchase, in order to protect your legal rights, we need to collect your facial and identity information for personal verification. Please hold your ID card with the back facing you.

[0250] The camera completes the collection of identity information.

[0251] Extraction results:

[0252] Q: Do you agree?

[0253] I suggest you confirm this with me in writing so that we can better protect your legal rights.

[0254] Please have the test driver (gmh ​​test) from route B demonstrate his / her credentials.

[0255] Please hold your ID card with the back facing the camera to complete the identity information collection.

[0256] A: Agreed, agreed.

[0257] (V) Time Rewind Module

[0258] The time-tracing module first uses the extracted statements as dividing nodes and then segments the merged target text content. The text between the marker "Q" and the previous marker "A" is used as the query stage, or the text between the marker "Q" and the previous marker "Q" is used as the query stage, and the text marked "A" is used as the intention response stage. In this way, the QA division result is obtained.

[0259] Furthermore, timestamps are used to backtrack through each stage in the QA segmentation results to obtain the timestamp for each stage. For example, if "Q" and "A" in the query stage correspond to the same time period, the timestamps can be segmented according to word length.

[0260] For example, the timestamp for the text corresponding to the question phase ([Do you agree? Yes.]) is between 0 and 9 seconds. In this case, the timestamp can be segmented according to the word length. For instance, based on the rule of one word per second, we get: "Q": [0-7 seconds Do you agree?], "A": [7-9 seconds Yes.]

[0261] Specific examples are as follows:

[0262] (1) QA division process:

[0263] 0:00:00~0:00:17.219000: To standardize the sales behavior of sales personnel and to better protect...

[0264] To protect your legal rights, in accordance with relevant regulations, we will record key aspects of my sales process.

[0265] (Q1) Do you agree? (A2) Yes, I agree.

[0266] 0:00:18.770000~0:00:39.100000: Please note that this audio and video recording process is for...

[0267] It will be crucial for you to protect your rights in the future. Please carefully read the details of the documents you sign and answer the relevant questions truthfully.

[0268] Question: If I make any promises to you that are inconsistent with the contents of the written documents (Q2), I suggest you contact me.

[0269] Please confirm in writing to better protect your legal rights.

[0270] 0:00:39.410000~0:00:45.959000:(Q3) Please have the tester gmh test the B route.

[0271] The test displays the qualification certificate.

[0272] 0:01:00.219000~0:01:15.700000: Hello Mr. / Ms. ××, before confirming your purchase, in order to protect your legal rights, we need to collect your facial and identity information for personal verification (Q4). Please hold your ID card in front of you.

[0273] Face the camera to complete the identity information collection.

[0274] (2) QA segmentation results:

[0275] Question 1: To regulate the sales behavior of sales personnel and to better protect your legitimate rights...

[0276] In accordance with relevant regulations, we will record key aspects of my sales process using audio and video recordings.

[0277] Do you agree?

[0278] Response text 2: Agreed.

[0279] Question 2: It is important to remind you that this audio and video recording process is very important for protecting your rights in the future.

[0280] The key is that you carefully read the specific contents of the document you sign and answer the relevant questions truthfully. If I make any decisions regarding this matter...

[0281] Any commitments that are inconsistent with the content of the written documents are recommended that you confirm them with me in writing for further clarification.

[0282] We will do our best to protect your legal rights.

[0283] Question text 3: Please have the tester gmh test from the ×× test route B show his / her qualification certificate.

[0284] Question Text 4: Hello Mr. / Ms. ××, in order to protect your legal rights before confirming the purchase, we need to collect your information.

[0285] Please hold your ID card with the back facing the camera to verify your identity.

[0286] collection.

[0287] Furthermore, based on the timestamps, the text content of each part in each QA segmentation result is traced back in time to obtain the dialogue text with timestamps (that is, the target dialogue text mentioned above).

[0288] Here, if multiple parts of the QA segmentation result are within the same time interval, the time interval can be divided according to the number of words in each part of the text within the same time interval. This time interval is based on the start and end times contained in the timestamp.

[0289] For example, if the text segment is "0:00:00~0:00:09: Do you agree? Agreed.", the QA segmentation result is: Q: "Do you agree?", A: "Agreed.". Since the text content of each part in the QA segmentation result is within the same time interval (i.e., "0:00:00~0:00:09"), then based on the proportion of text words in each part, this time interval is further divided. That is, the time interval for Q: "Do you agree?" is 0:00:00~0:00:07, and the time interval for A: "Agreed." is 0:00:07~0:00:09.

[0290] Continuing with the above extraction results as an example, after dividing the merged text content according to the extraction results and obtaining the following QA division results, time backtracking is performed on each part of the text content in the QA division results based on the timestamps in the speech recognition module, resulting in the following time backtracking results:

[0291] The time backtracking results are as follows:

[0292] 0:00:00~0:00:16.235714: To regulate the sales behavior of sales personnel and to better protect...

[0293] To protect your legal rights, in accordance with relevant regulations, we will record key aspects of my sales process.

[0294] Please record this section. Do you agree?

[0295] 0:00:16.235714~0:00:17.219000: Agreed.

[0296] 0:00:18.770000~0:00:39.100000: Please note that this audio and video recording process is for...

[0297] It will be crucial for you to protect your rights in the future. Please carefully read the details of the documents you sign and answer the relevant questions truthfully.

[0298] Question: If I make any promises to you that are inconsistent with the contents of the written documents, I suggest you consult with me in writing.

[0299] This form of confirmation is to better protect your legitimate rights and interests.

[0300] 0:00:39.410000~0:00:45.959000: Please have the tester from XX test route B demonstrate.

[0301] Qualification certificate.

[0302] 0:01:00.219000~0:01:15.700000: Hello Mr. / Ms. ××, before confirming your purchase, in order to protect your legal rights, we need to collect your facial and identity information for personal verification. Please hold your ID card with the back facing you.

[0303] The camera completes the collection of identity information.

[0304] At this point, a timestamped dialogue text can be obtained, which includes at least one of the following: a timestamped question-and-answer pair, or a timestamped question text.

[0305] (vi) Script Matching Module

[0306] In actual dual-recording scenarios, to ensure the accuracy and consistency of the segmentation, each segment corresponds to a specific script template. Using these script templates, semantic matching of the speech recognition results during dual-recording can be performed to map the dialogue text onto the corresponding script template, thereby determining the specific segment corresponding to the dialogue text.

[0307] To achieve this functionality, the dialogue matching module employs Ernie3.0 (Enhanced Representation Through Knowledge Integration 3.0) as its base model, combined with In-batch Negative Sampling (In-batch Negatives) techniques. Supervised training is performed on a publicly available literature search dataset to obtain a semantic indexing model. Based on this model, a vector index library is constructed. The Approximate Nearest Neighbors (ANN) engine from the hierarchical navigable small world library (hnswlib) enables rapid retrieval of corresponding dialogue templates. This fully leverages advanced natural language processing technology and an index retrieval engine, providing an efficient and reliable solution for segmentation in dual-recording scenarios.

[0308] Specifically, as shown in Figure 7(c), the specific steps of the script matching module include:

[0309] Step a: Organize the pre-set script templates and input the organized script templates into the vectorization service engine based on Ernie 3.0 to perform vectorization processing on the script templates. Then, use the vectorized script templates and the ANN engine to build a vector index library.

[0310] Step b: Integrate the dialogue text obtained from the time backtracking module, and input the integrated dialogue text into the vectorization service engine based on Ernie 3.0 to obtain vectorized dialogue text. Use the vectorized dialogue text to search in the vector index library to obtain the top K recall results with the highest semantic similarity.

[0311] Step c: Based on the first K recall results, determine the stage to which each text content in the dialogue text belongs, so as to obtain the stage information to which each text content in the dialogue text belongs.

[0312] (vii) Process Merging Module

[0313] The specific steps of the process merging module include:

[0314] Step a: Merge text content with the same stage information in the dialogue text according to the time sequence, and also merge the timestamps corresponding to the text content with the same stage information in the dialogue text.

[0315] Step b: For steps involving actions (such as identity verification, document display, qualification certificate display, signature recognition, etc.), the timestamps need to be extended based on the time of adjacent steps.

[0316] Step c: Obtain the dialogue text annotated with stage information.

[0317] Furthermore, Figure 8 A complete flowchart of the disclosed solution is provided, such as Figure 8 As shown, the overall process of this disclosed solution includes:

[0318] Step S801: Use the FFmpeg audio and video stream processing tool to extract the dialogue audio from the video to be inspected and perform format conversion to obtain the format-converted dialogue audio.

[0319] Step S802: Perform speech activity detection on the format-converted dialogue audio to obtain multiple audio segments and the timestamps corresponding to each audio segment. Transcribe each audio segment to obtain the text segment of each audio segment.

[0320] Here, after obtaining each audio segment, non-human voice filtering can be performed on each audio segment. For the specific processing procedure, please refer to the example above, which will not be repeated here.

[0321] Step S803: Denoise the text segment of the audio segment to remove text content that does not belong to the dual recording scenario, and obtain the denoised text segment.

[0322] Step S804: Merge the denoised text fragments to obtain merged text content, and extract intent from the merged text content to obtain at least one of the following: question text related to the question, and response text related to the answer.

[0323] Step S805: Based on the extracted text, perform QA segmentation on the merged text content, and perform time matching on each part of the QA segmented text according to the timestamp to obtain the dialogue text with timestamp.

[0324] Step S806: Integrate the obtained dialogue text to obtain the integrated dialogue text; use the dialogue template to perform dialogue matching on the integrated dialogue text to determine the segment information to which each text content belongs in the integrated dialogue text.

[0325] Step S807: Based on the stage information and time information of each text content in the integrated dialogue text, merge the text content with the same stage information to obtain the dialogue text marked with stage information.

[0326] In summary, the disclosed solution has the following advantages, including:

[0327] First, the quality of the dialogue content is higher. This disclosed solution performs non-human voice filtering on the multiple audio segments obtained after audio segmentation, ensuring the quality of audio transcription; moreover, this disclosed solution also performs dialogue denoising processing on the multiple text segments obtained after audio transcription. For example, it uses text classification to divide each text segment into three categories: questions, answers, and noise, and then removes the text segments that belong to noise, thereby obtaining a cleaner dialogue content.

[0328] Secondly, the distinction between question and answer content is higher. This disclosed solution utilizes a large model to extract information from text content to obtain question text related to the question and answer text related to the answer, effectively avoiding the problem of inaccurate segmentation of question text and answer text, thereby improving the distinction between question and answer content.

[0329] Third, quality inspection efficiency is improved. This disclosed solution can accurately extract business-related and timestamped dialogue text from target dialogue audio, and realize audio segmentation of target dialogue audio, thereby significantly improving the quality inspection efficiency of audio data.

[0330] This disclosure also provides a data processing apparatus based on a large model, such as... Figure 9 As shown, it includes:

[0331] The audio segmentation unit 901 is used to perform speech activity detection on the target dialogue audio to obtain multiple target audio segments with speech activity.

[0332] The audio conversion unit 902 is used to obtain target text content based on the plurality of target audio segments;

[0333] The text processing unit 903 is configured to use a large model to extract intent from the target text content and extract at least one of the following: question text related to the question, and answer text related to the answer; based on the time information in the target dialogue audio and the extracted text, obtain a target dialogue text with time information associated with the target dialogue audio, wherein the target dialogue text includes at least one of the following: question-answer pairs with time information, and question text with time information.

[0334] In a specific example of the disclosed solution, the audio conversion unit is specifically used for:

[0335] Based on the multiple target audio segments, a text segment of each target audio segment is obtained;

[0336] The text segments of each target audio segment are denoised to obtain the denoised text segments of each target audio segment.

[0337] The text segments of each denoised target audio segment are concatenated to obtain the target text content.

[0338] In a specific example of the disclosed solution, the audio conversion unit is specifically used for:

[0339] Using the first model, the text segment of the target audio segment is processed to determine whether there are statements in the text segment that are unrelated to the target business scenario; where the target business scenario refers to the business scenario that the target dialogue audio is aimed at.

[0340] If it is determined that there are statements that are irrelevant to the target business scenario, delete the statements that are irrelevant to the target business scenario from the text segment of the target audio segment to obtain the text segment of the target audio segment after noise reduction.

[0341] In a specific example of the disclosed solution, the audio conversion unit is specifically used for:

[0342] Using the first model, the sentences in the text segments of the target audio segment are classified to obtain the classification results of the text segments of the target audio segment. The classification results indicate that the sentences in the text segments belong to one of the following: noise, question, or answer; noise indicates that it is irrelevant to the target business scenario; question indicates a question related to the target business scenario; and answer indicates a response related to the target business scenario.

[0343] Based on the classification results of text segments for the target audio segment, it is determined whether there is noise in the text segment.

[0344] In a specific example of the disclosed solution, the text processing unit is further configured to:

[0345] Identify the question-and-answer segments involved in the text segments of each target audio segment after noise reduction;

[0346] The content of the associated question-and-answer segments in the target dialogue text is merged to obtain the target dialogue text marked with question-and-answer segments.

[0347] In a specific example of the disclosed solution, the audio segmentation unit is specifically used for:

[0348] Speech activity detection is performed on the target dialogue audio to obtain multiple initial audio segments containing speech activity;

[0349] From the plurality of initial audio segments, the initial audio segments whose frequencies fall within the target frequency range are determined, and the initial audio segments falling within the target frequency range are used as the target audio segments.

[0350] In a specific example of the disclosed solution, the audio segmentation unit is further configured to:

[0351] Determine the spectral information corresponding to the time frames contained in each initial audio segment;

[0352] Based on the spectral information corresponding to the time frames contained in each initial audio segment, the frequency values ​​corresponding to the time frames contained in each initial audio segment are obtained.

[0353] The target frequency range is determined based on the frequency values ​​corresponding to the time frames contained in each initial audio segment.

[0354] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.

[0355] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0356] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0357] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0358] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0359] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0360] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as large model-based data processing methods. For example, in some embodiments, the large model-based data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the large model-based data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform a large model-based data processing method by any other suitable means (e.g., by means of firmware).

[0361] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0362] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0363] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0364] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0365] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0366] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0367] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0368] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A data processing method based on a large model, comprising: Speech activity detection is performed on the target dialogue audio to obtain multiple target audio segments with speech activity; Based on the multiple target audio segments, the target text content is obtained; Using a large model, intent extraction is performed on the target text content, and at least one of the following is extracted: question text related to the question, and response text related to the answer; Based on the time information in the target dialogue audio and the extracted text, multiple target dialogue texts with time information are obtained that are associated with the target dialogue audio. Among these multiple target dialogue texts, at least one of the following types is included: question-answer pairs with time information and question texts with time information. The step of obtaining the target text content based on the plurality of target audio segments includes: Based on the multiple target audio segments, a text segment of each target audio segment is obtained; Using the first model, the sentences in the text segments of the target audio segment are classified to obtain the classification results of the text segments of the target audio segment. The classification results indicate that the sentences in the text segments belong to one of the following: noise, question, or answer; noise indicates that it is irrelevant to the target business scenario; question indicates a question related to the target business scenario; and answer indicates a response related to the target business scenario. Based on the classification results of the text segments of the target audio segments, the text segments of each target audio segment are denoised to obtain the denoised text segments of each target audio segment, so that the denoised text segments of each target audio segment are related to the target business scenario, wherein the target business scenario refers to the business scenario targeted by the target dialogue audio. The text segments of each denoised target audio segment are concatenated to obtain the target text content.

2. The method according to claim 1, wherein, Based on the classification results of text segments for the target audio segments, the text segments of each target audio segment are denoised to obtain denoised text segments of each target audio segment, including: Based on the classification results of the text segments of the target audio segment, the text segments of the target audio segment are processed to determine whether there are statements in the text segments that are irrelevant to the target business scenario; If it is determined that there are statements that are irrelevant to the target business scenario, delete the statements that are irrelevant to the target business scenario from the text segment of the target audio segment to obtain the text segment of the target audio segment after noise reduction.

3. The method according to claim 1 or 2, further comprising: Identify the question-and-answer segments involved in the text segments of each target audio segment after noise reduction; The content of the associated question-and-answer segments in the target dialogue text is merged to obtain the target dialogue text marked with question-and-answer segments.

4. The method according to claim 1 or 2, wherein, The process of detecting speech activity in the target dialogue audio to obtain multiple target audio segments containing speech activity includes: Speech activity detection is performed on the target dialogue audio to obtain multiple initial audio segments containing speech activity; From the plurality of initial audio segments, the initial audio segments whose frequencies fall within the target frequency range are determined, and the initial audio segments falling within the target frequency range are used as the target audio segments.

5. The method according to claim 4, further comprising: Determine the spectral information corresponding to the time frames contained in each initial audio segment; Based on the spectral information corresponding to the time frames contained in each initial audio segment, the frequency values ​​corresponding to the time frames contained in each initial audio segment are obtained. The target frequency range is determined based on the frequency values ​​corresponding to the time frames contained in each initial audio segment.

6. A data processing device based on a large model, comprising: The audio segmentation unit is used to detect speech activity in the target dialogue audio and obtain multiple target audio segments with speech activity. An audio conversion unit is used to obtain target text content based on the plurality of target audio segments; The text processing unit is used to extract intent from the target text content using a large model, and extract at least one of the following: question text related to the question, and response text related to the answer; based on the time information in the target dialogue audio and the extracted text, to obtain multiple target dialogue texts with time information associated with the target dialogue audio, wherein the multiple target dialogue texts include at least one of the following: question-answer pairs with time information, and question text with time information; Specifically, the audio conversion unit is used for: Based on the multiple target audio segments, a text segment of each target audio segment is obtained; Using the first model, the sentences in the text segments of the target audio segment are classified to obtain the classification results of the text segments of the target audio segment. The classification results indicate that the sentences in the text segments belong to one of the following: noise, question, or answer; noise indicates that it is irrelevant to the target business scenario; question indicates a question related to the target business scenario; and answer indicates a response related to the target business scenario. Based on the classification results of the text segments of the target audio segments, the text segments of each target audio segment are denoised to obtain the denoised text segments of each target audio segment, so that the denoised text segments of each target audio segment are related to the target business scenario, wherein the target business scenario refers to the business scenario targeted by the target dialogue audio. The text segments of each denoised target audio segment are concatenated to obtain the target text content.

7. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Learner input degree analysis method and system based on multi-modal large language model

    CN118037103A

  • Multi-modal video question answering method and device and computer equipment

    CN118194230A

  • Clinical test informed agreement signing method, signing system and electronic equipment

    CN118964543A