Index information generation method and device, electronic equipment and medium

By generating speaker-related and interference information label sequences by acquiring the spectral and prosodic features of audio and video, the problem of lack of contextual relevance in audio and video index information in existing technologies is solved, and the efficiency of acquiring key information is improved.

CN122019827APending Publication Date: 2026-05-12VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, audio and video index information generated based on rules or shallow statistical models lacks contextual relevance, resulting in low efficiency for users to obtain key information in audio and video scenarios.

Method used

By acquiring the spectral and prosodic features of audio and video, speaker-related and interference information label sequences are generated, and combined with the audio and video text, structured index information is generated.

Benefits of technology

Ensuring that the generated index information has contextual relevance improves the efficiency for users to obtain key information in audio and video scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019827A_ABST
    Figure CN122019827A_ABST
Patent Text Reader

Abstract

The invention discloses an index information generation method and device, electronic equipment and a medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring audio and video features of a first audio and video, wherein the audio and video features comprise at least one of a spectrum feature and a rhythm feature; a label sequence is determined according to the audio and video features, the label sequence comprises a label corresponding to each audio and video frame of the first audio and video, and the labels comprise at least one of a speaker related label and an interference information label; and generating structured index information of the first audio and video according to the tag sequence and the first text corresponding to the first audio and video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an index information generation method, apparatus, electronic device, and medium. Background Technology

[0002] With the increasing popularity of long-form audio content such as podcasts, interviews, and training sessions, users' need for quick access to key information in audio and video content is becoming increasingly prominent.

[0003] Among related technologies, automatic speech recognition (ASR) technology can be used to first convert audio content into text, and then generate text summaries through rule-based or shallow statistical models.

[0004] However, indexing information such as summaries and tables of contents derived from rules or shallow statistical models lacks contextual relevance and logic, making it difficult to support indexing and knowledge retrieval of long text content. Therefore, there is a problem of low efficiency for users to obtain key information in audio and video scenarios. Summary of the Invention

[0005] The purpose of this application is to provide an index information generation method, apparatus, electronic device, and medium that can solve the problem of low efficiency for users to obtain key information in audio and video scenarios.

[0006] In a first aspect, embodiments of this application provide an index information generation method, the method comprising: acquiring audio-visual features of a first audio-visual video, the audio-visual features including at least one of spectral features and prosodic features; determining a tag sequence based on the audio-visual features, the tag sequence including tags corresponding to each audio-visual frame of the first audio-visual video, the tags including at least one of speaker-related tags and interference information tags; and generating structured index information of the first audio-visual video based on the tag sequence and a first text corresponding to the first audio-visual video.

[0007] Secondly, embodiments of this application provide an index information generation system, which may include: an audio feature extraction model, a tag generation model, and a structured content generation model; the audio feature extraction model is used to obtain audio and video features of a first audio and video, the audio and video features including at least one of spectral features and prosodic features; the tag generation model is used to determine a tag sequence based on the audio and video features, the tag sequence including tags corresponding to each audio and video frame of the first audio and video, the tags including at least one of speaker-related tags and interference information tags; the structured content generation model is used to generate structured index information of the first audio and video based on the tag sequence and the first text corresponding to the first audio and video.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, audio-visual features of a first audio-visual video can be obtained, including at least one of spectral features and prosodic features. A tag sequence is determined based on these features, comprising tags corresponding to each frame of the first audio-visual video, including at least one of speaker-related tags and interference information tags. Structured index information of the first audio-visual video is generated based on the tag sequence and the corresponding first text. This approach ensures that the generated index information has contextual relevance, resulting in high-quality index information and improving the efficiency of users obtaining key information in audio-visual scenarios. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating an index information generation method provided in an embodiment of this application;

[0014] Figure 2 This is an architecture diagram of an index information generation system provided in an embodiment of this application;

[0015] Figure 3 This is a partial architecture diagram of an index information generation system provided in an embodiment of this application;

[0016] Figure 4 This is a partial architecture diagram of an index information generation system provided in an embodiment of this application;

[0017] Figure 5 This is a partial architecture diagram of an index information generation system provided in an embodiment of this application;

[0018] Figure 6 This is a partial architecture diagram of an index information generation system provided in an embodiment of this application;

[0019] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0020] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0022] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0023] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0024] The following is an explanation of the terms and concepts used in this application specification.

[0025] Automatic Speech Recognition (ASR) is a technology that converts human speech into a text sequence.

[0026] Bidirectional Encoder Representations from Transformers (BERT) is a pre-trained language model that can capture bidirectional contextual information and is widely used in natural language understanding tasks.

[0027] Bidirectional Long Short-Term Memory (Bi-LSTM) is a recurrent neural network structure that can simultaneously consider the contextual information before and after a sequence, and is often used for speech and text sequence analysis.

[0028] Convolutional Recurrent Neural Network (CRNN) is a deep learning model that combines convolutional neural networks and recurrent neural networks, and is suitable for processing multidimensional time-series data such as audio.

[0029] Speaker log error rate (DER): A metric used to evaluate the accuracy of speaker segmentation (i.e., determining "who is speaking and when") tasks; for example, DER is used to evaluate log performance, including error classification, missed detections, and false alarms.

[0030] Large Language Model (LLM): refers to a large-scale artificial intelligence model trained on massive amounts of text data, which has powerful natural language understanding and generation capabilities.

[0031] Mel-Frequency Cepstral Coefficients (MFCCs): Acoustic features extracted from speech signals that effectively reflect the characteristics of human ear perception, and are widely used in the field of speech recognition.

[0032] Natural Language Processing (NLP) is the field of artificial intelligence that studies how to enable computers to understand, interpret, and generate human language.

[0033] Noise suppression (NS): a signal processing technique used to reduce background noise in audio signals.

[0034] Voice Activity Detection (VAD): Voice activity detection is a technique used to identify segments of human speech in an audio stream.

[0035] The index information generation method, apparatus, electronic device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0036] The index information generation method, apparatus, electronic device, and medium provided in this application can be applied to scenarios where multimedia streams such as audio or video are transcribed into structured text. Specifically, the method of this application can be applied to scenarios including but not limited to the following:

[0037] (1) Podcast / Audiobook Transcription and Textification: Automatically convert long audio interviews, reviews, and stories into interactive texts with chapter indexes, core summaries, and argument relationship diagrams to improve learning and review efficiency.

[0038] (2) Online education / training course structuring: intelligent segmentation of lectures and training videos to generate course outlines, key summary and knowledge point association network to assist students in self-study and review.

[0039] (3) Intelligent organization of online meeting / interview records: Automatically distinguish multiple speakers, extract key points of meeting minutes and resolutions, construct the logical framework of discussion topics, and support the retrospective of the speaking process from the conclusion.

[0040] (4) Media content post-production and retrieval: Automatically generate timestamp indexes, content summaries and key topic tags for long videos such as news programs and documentaries, greatly simplifying material management and content retrieval processes.

[0041] (5) Customer service quality inspection and knowledge mining: Analyze customer service call recordings, automatically extract common questions, standard scripts and service points, and build a question and answer knowledge base for service quality assessment and agent training.

[0042] (6) Judicial / Government Records Assistance Generation: High-precision transcription and structuring of audio from formal occasions such as court hearings and hearings, automatically generating procedural catalogs, summaries of points of contention, and presentation of evidence chains.

[0043] In the index information generation method of this application embodiment, audio-visual features of a first audio-visual video can be obtained, including at least one of spectral features and prosodic features; a tag sequence is determined based on the audio-visual features, the tag sequence including tags corresponding to each audio-visual frame of the first audio-visual video, the tags including at least one of speaker-related tags and interference information tags; and structured index information of the first audio-visual video is generated based on the tag sequence and the first text corresponding to the first audio-visual video. Through this scheme, since the tag sequence containing speaker-related and interference information in the first audio-visual video can be determined based on the audio spectral features of the first audio-visual video, and the tag sequence is combined with the text content of the first audio-visual video to generate index information, the generated index information can be ensured to have contextual relevance. This results in high-quality index information for audio-visual videos, improving the efficiency of users obtaining key information in audio-visual scenarios.

[0044] For example, consider a podcast episode about the development of artificial intelligence. Originally, it was just two hours of continuous audio, which was unstructured content. To find the relevant content, such as "difficulties in training data for large models," one would have to listen to the entire episode. However, after using the system described in this application, the audio is first segmented into semantic units such as "background of AI development," "technical principles of large models," and "pain points in training data." Then, a text with three levels of headings is generated. Each heading contains core arguments and coherent summaries. Technical terms such as "annotated data cost" are explained in the text. Clicking on the text allows for precise navigation to the corresponding audio segment. It also supports adding personal notes. This transforms the originally messy audio into structured knowledge units that allow for quick access to core information and interactive processing. This makes it more convenient for both users to find information and for machines to process it, thereby improving the efficiency of users in obtaining key information in audio and video scenarios.

[0045] The index information generation method provided in this application is executed by an index information generation device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this method. The following will use an index information generation device as an example to illustrate the index information generation method provided in this application.

[0046] This application provides a method for generating index information. Figure 1 A flowchart illustrating the index information generation method provided in this application is shown, as follows: Figure 1 As shown, the index information generation method may include the following steps 101 to 103.

[0047] Step 101: The index information generation device acquires the audio and video features of the first audio and video.

[0048] In some embodiments of this application, the audio-visual features of the first audio-visual video may include at least one of the spectral features and prosodic features of the first audio-visual video.

[0049] In some embodiments of this application, the first audio and video may be an audio and video podcast, audio and video courseware, film and television works, telephone recordings, court hearing audio and video, etc., and this application does not limit the embodiments.

[0050] In some embodiments of this application, the index information generating device may first extract features from the first audio and video to obtain the audio and video features of the first audio and video.

[0051] In some embodiments of this application, the spectral characteristics of the first audio and video may include any of the following: log-melspectrogram, MFCC, filter bank energy, etc.

[0052] In some embodiments of this application, the prosodic features of the first audio and video may include, but are not limited to, extracting prosodic-related parameters such as fundamental frequency (F0), energy, and formants.

[0053] It can be understood that the spectral features of the first audio / video frame include the spectral features of each audio / video frame in the aforementioned audio / video frame sequence. Similarly, the prosodic features of the first audio / video frame include the prosodic features of each audio / video frame in the aforementioned audio / video frame sequence. In short, the audio / video features of the first audio / video frame include the spectral features and prosodic features of each audio / video frame in the first audio / video frame.

[0054] In some embodiments of this application, the spectral characteristics of the first audio and video can be referred to as the main channel characteristics, and the prosodic characteristics of the first audio and video can be referred to as the auxiliary channel characteristics.

[0055] In some embodiments of this application, the index information generation device may first perform front-end processing such as voice activity detection and noise suppression on the first audio and video.

[0056] Specifically, through speech activity detection, speech segments containing human speech can be identified in the first audio and video, and non-speech segments (such as silence and background noise) can be marked; then, noise reduction processing is performed on each speech segment in the first audio and video to improve the speech quality of each speech segment.

[0057] In some embodiments of this application, after the front-end processing is completed, the index information generation device can perform frame-by-frame processing on the first audio and video after the front-end processing to obtain an audio and video frame sequence, and extract the spectral features and prosodic features of each audio and video frame in the audio and video frame sequence to obtain the audio and video features of the first audio and video.

[0058] For example, the index information generation device can extract the 80-dimensional or 40-dimensional log-Mel spectrum of each audio / video frame from the audio / video frame sequence as the main channel feature (i.e., spectral feature). Of course, the spectral feature can also be a log-Mel spectrum of other dimensions.

[0059] For example, the index information generation device can perform frame processing on the first audio and video according to a frame length of 25 milliseconds and a step size of 10 milliseconds, and then extract 40-dimensional MFCC features for each audio and video frame to ensure a sensitive response to the pitch and prosody information of the speech signal.

[0060] For example, prosodic parameters such as fundamental frequency (F0), energy, and formants of each audio and video frame can be extracted from the audio and video frame sequence, and the extracted features can be used as auxiliary feature channels.

[0061] In some embodiments of this application, the front-end processing and feature extraction processing of the first audio and video can be performed by... Figure 2 The audio feature extraction model 201 in the index information generation system 200 shown is executed.

[0062] It is understandable that the accuracy of the first audio and video feature front-end processing and feature extraction is directly related to the accuracy and stability of subsequent semantic segmentation and index information generation.

[0063] Step 102: The index information generation device determines the tag sequence based on the audio and video features of the first audio and video.

[0064] The aforementioned tag sequence may include tags corresponding to each audio / video frame of the first audio / video, and the tags corresponding to each audio / video frame may include at least one of speaker-related tags and interference information tags.

[0065] In some embodiments of this application, the aforementioned speaker-related tags may include at least one of the following: speaker identity tag, speaker emotion tag, and speaking rhythm tag.

[0066] In some embodiments of this application, the speaker-related tags corresponding to each audio / video frame may include at least one of the following: speaker identity tag, speaker emotion tag, and speaking rhythm tag, which are specifically determined according to the audio / video content of each audio / video frame.

[0067] In some embodiments of this application, the aforementioned speaker identity label can be an anonymous label, such as "speaker A" or "speaker B", used to distinguish different speakers in the first audio and video.

[0068] In some embodiments of this application, the speaker emotion label described above can be a label indicating an emotional change.

[0069] For example, the speaker emotion label corresponding to the audio and video frame can indicate a change in the speaker's emotion. For instance, if the index information generating device detects that speaker B's emotion was stable in the previous speech, speaker B's emotion became significantly more agitated in the current speech.

[0070] For example, if the index information generating device detects that speaker A's tone was calm in the last time he spoke, but his tone was excited in the current speech, it indicates that speaker A's emotions are running high.

[0071] In some embodiments of this application, the aforementioned speaking rhythm label can be a label indicating changes in the speaker's speaking rhythm, i.e., a relative rhythm label.

[0072] For example, the speaking rhythm label corresponding to the audio and video frame can indicate whether a speaker's speaking rhythm has changed compared to the speaker's usual rhythm. For instance, if the speaker usually speaks slowly, but speaks much faster this time.

[0073] Of course, in actual implementation, emotion levels and rhythm levels can also be preset. The emotion level is determined by comparing the emotion tone parameter with the emotion tone parameter corresponding to the emotion level, and the matched emotion level is used as the emotion label. The same applies to rhythm labels.

[0074] In some embodiments of this application, the speaking rhythm label can reflect the speaker's behavior and may therefore also be referred to as a "behavior label" or "action label".

[0075] In some embodiments of this application, the interference information label can be a mask for interference information, which can indicate the location of the interference information in the first audio and video.

[0076] In some embodiments of this application, the interference information tag can be used to identify at least one of the following in audio and video frames:

[0077] Background noise, such as ambient noise and equipment background noise;

[0078] Non-lexical sounds, such as laughter and coughing;

[0079] Disfluency in speech, such as the use of interjections, repetitions, and meaningless filler words.

[0080] In some embodiments of this application, the background noise of the first audio / video can be determined by signal-to-noise ratio estimation and spectral silence detection based on the spectral features of the first audio / video, or by VAD based on the spectral features of the first audio / video. Non-lexical sounds in the first audio / video can be determined by the spectral features of the first audio / video. Disfluency in speech in the first audio / video can be determined by tone recognition analysis of the prosodic features of the first audio / video.

[0081] It should be noted that after step 102 is executed, each audio and video frame in the first audio and video contains a tag.

[0082] For example, each audio / video frame may carry at least one of the following: speaker identity tag, speaker emotion tag, speaking rhythm, and interference information tag.

[0083] It is understandable that the above tag sequence can be used to achieve semantic and acoustic linkage of the first audio and video audio segmentation, as well as the generation of subsequent structured index information.

[0084] In some embodiments of this application, the speaker-related tags mentioned above include speaker identity tags, and step 102 may include steps 102A and 102B as described below.

[0085] Step 102A: The index information generation device obtains the embedding vector of at least one speech segment based on the spectral features of the first audio and video by using the speaker-related tag generation model.

[0086] Each speech segment comprises at least one consecutive audio / video frame from the first audio / video frame. It is understood that each of these at least one audio / video frame includes speech content; that is, these at least one audio / video frame does not include non-speech frames.

[0087] In some embodiments of this application, the embedding vector of a speech segment can represent the speaker identity information of the speech in that speech segment. The embedding vector of a speech segment can also be called the speaker embedding vector of the speech segment.

[0088] In some embodiments of this application, the embedding vector of a speech segment can be a fixed-dimensional embedding vector. The embedding vector of the speech segment is discriminative, so that embedding vectors containing the same speaker can be clustered in the embedding space.

[0089] In some embodiments of this application, the speaker-related tag generation model may include a speaker digitization module, and the index information generation device may obtain the embedding vector of at least one speech segment based on the spectral features of the first audio and video through the speaker digitization module.

[0090] Specifically, each audio / video frame in the first audio / video file has a tag that identifies whether it is a speech frame. The speaker logging module can divide the spectral features of the first audio / video frame into at least one spectral feature segment based on the tag of each frame, with each spectral feature segment corresponding to a speech segment in the first audio / video file. For example, the speaker logging module can divide the spectral features of the first audio / video frame into at least one spectral feature segment based on the tag and the first step length. Alternatively, it can adaptively divide the spectral features of the first audio / video frame into at least one spectral feature segment based on the tag of each frame, for example, using a non-speech frame as the segmentation identifier. The first step length can be 2 seconds, 1.5 seconds, or 3 seconds, etc. Then, the speaker logging module generates an embedding vector for the speech segment corresponding to each spectral feature segment based on all spectral features in each spectral feature segment.

[0091] In some embodiments of this application, the speaker log module described above can be any model capable of extracting speaker identity representation.

[0092] For example, the speaker log module mentioned above can be an enhanced channel attention mechanism time delay neural network (ECAPA-TDNN) model or an x-vector model based on residual network (ResNet).

[0093] In some embodiments of this application, the speaker-related label generation model described above can be a three-layer stacked neural network model. This three-layer stacked neural network model can include an input layer, an intermediate layer, and an output layer.

[0094] The input layer, also known as the bottom layer, is used to receive the spectral features of the first audio and video. This layer is usually composed of convolutional layers (CNN) or fully connected layers, and is responsible for performing preliminary feature fusion and transformation on the spectral features, extracting low-level acoustic patterns from the spectral features of each audio and video frame.

[0095] The intermediate layer, also known as the context coding layer, can be used to perform contextual modeling on the spectral features of the first audio / video file to extract a deep, context-rich speech content representation for each frame of the first audio / video file. The intermediate layer can be an acoustic encoder. It can be understood that the speech content representation can be frame-level.

[0096] The output layer, also known as the higher layer, contains the aforementioned speaker log module. The input to the speaker log module can be the speech content representation output from the intermediate layer. The speaker log module can aggregate all speech content representations belonging to the same speech segment through a pooling layer (such as statistical pooling, calculating the mean and standard deviation) to form a unified segment-level feature representation; and based on this segment-level feature representation, generate a fixed-dimensional speaker embedding vector. This speaker embedding vector aims to ensure that the vectors of different speech segments from the same speaker are close in distance in the embedding space, while the vectors of different speakers are far apart.

[0097] In some embodiments of this application, speaker identification is accomplished through a dedicated SpeakerDiarization module. This module employs an advanced speaker representation extraction model (e.g., ECAPA-TDNN or ResNet-based x-vector) to generate fixed-dimensional embedding vectors for each speech segment. In unsupervised mode, the system estimates the number of speakers by performing cluster analysis on the embedding vectors (e.g., spectral clustering, such as MFCC) and assigns an anonymous speaker label (e.g., "Speaker A", "Speaker B") to each speech segment. The accuracy is evaluated by metrics such as Speaker Error Rate (DER).

[0098] In some embodiments of this application, the intermediate layer can extract the aforementioned speech content representation of each audio and video frame from the spectral features of the first audio and video by capturing the long-distance temporal dependencies between audio and video frames.

[0099] For example, assuming the first audio / video is an episode of a continuously updated video podcast, the speaker-related tag generation model can use an acoustic encoder to search for semantics of spectral features that match the semantics of the spectral features in historical podcasts based on the spectral features of each audio / video frame in the first audio / video, in order to help understand the semantics of the audio / video frame and thus obtain the speech content representation of the audio / video frame.

[0100] In some embodiments of this application, the acoustic encoder may be an encoder based on Transformer or a similar architecture.

[0101] In some embodiments of this application, adjacent speech segments may or may not overlap.

[0102] For example, taking overlap as an example, two adjacent speech segments include the same audio / video frame. It should be noted that "the speaker log module generates an embedding vector for the speech segment corresponding to each spectral feature segment based on all spectral features in each spectral feature segment," includes: the speaker log module aggregating all speech content representations corresponding to each spectral feature segment to generate an embedding vector for the speech segment corresponding to each spectral feature segment.

[0103] In some embodiments of this application, in order to accurately model the dynamic evolution of the speaker's speech features, a joint distributed coding function can be introduced into the acoustic encoder or intermediate layer. :

[0104] (Formula 1)

[0105] In Formula 1, This represents the speaker composite feature distribution at time point t, where each audio / video frame corresponds to a time point t. and These are attention-weighted static and dynamic coefficients, used to emphasize the MFCC features of the current audio / video frame and their first derivatives. The importance of This refers to the MFCC change between two adjacent audio and video frames. This joint distributed coding function, by introducing a dynamic guidance mechanism, establishes a sensitive response mapping for speech rate changes, pitch jitter, and semantic breakpoints in the first audio and video frame, effectively improving the accuracy of subsequent content unit segmentation and the discrimination performance of acoustic abrupt changes.

[0106] Among them, the above-mentioned speaker composite feature distribution is a specific implementation of multi-granularity speaker features, which are high-level speaker-related features extracted from different granularities such as frame level and segment level.

[0107] It should be noted that in Formula 1 above... This represents the deep, context-rich speech content representation of audio / video frame t. In other words, the speaker composite feature distribution at time point t is a specific manifestation of the speech content representation of the audio / video frame corresponding to time point t.

[0108] Step 102B: The index information generation device performs clustering processing on at least one embedded vector to obtain the speaker identity label corresponding to each audio and video frame.

[0109] It should be noted that "the index information generation device performs clustering processing on at least one embedding vector to obtain speaker identity tags corresponding to each audio and video frame" can be understood as: by performing clustering processing on the at least one embedding vector, the speaker identity tag corresponding to each audio and video frame corresponding to the at least one embedding vector can be obtained. For audio and video frames in the first audio and video that have no speech activity, no speaker identity tag is configured.

[0110] In some embodiments of this application, the index information generation device can perform clustering processing on the above-mentioned at least one embedding vector through a speaker-related tag generation model to obtain clustering results; and determine the speaker identity tag corresponding to each audio and video frame based on the clustering results.

[0111] In some embodiments of this application, the index information generation device can, in unsupervised mode, employ a first clustering algorithm to perform cluster analysis on at least one embedding vector to estimate the number of speakers in each embedding vector and assign an anonymous speaker identity label to each speaker. The first clustering algorithm may include, but is not limited to, spectral clustering and hierarchical clustering.

[0112] Specifically, (1) Speaker number estimation: The number of clusters in at least one of the above embedding vectors is determined by calculating the eigengap of the similarity matrix or by using the Bayesian information criterion. (2) Clustering process: A similarity matrix is ​​constructed based on cosine similarity or Euclidean distance, and spectral clustering is applied to obtain the cluster label for each cluster. It can be understood that the cluster label is an anonymous label. (3) Frame-level label mapping: The cluster label of each cluster is assigned to all audio and video frames in that cluster.

[0113] In some embodiments of this application, the speaker-related label generation model described above, especially the speaker log module, can be pre-trained in a supervised manner on a large speaker recognition dataset (such as VoxCeleb) to learn the embedding space that distinguishes different speakers. The accuracy of the speaker module is evaluated using metrics such as the speaker error rate (DER).

[0114] It should be noted that steps 102A and 102B above are illustrated using the example of generating speaker identity tags corresponding to audio and video frames through a speaker-related tag generation model. In actual implementation, the speaker-related tag generation model can also generate emotion tags and rhythm tags corresponding to audio and video frames.

[0115] In other words, the speaker-related label generation model can be used to extract multi-granularity speaker features, sentiment tone parameters, and rhythm dynamic vectors. The weights of the speaker-related label generation model can be enhanced through adversarial training to ensure strong robustness and transferability. It can be understood that multi-granularity speaker features are used to determine speaker identity labels, sentiment tone parameters are used to determine speaker emotion labels, and rhythm dynamic vectors are used to determine speech rhythm labels.

[0116] For example, consider a speaker-related label generation model that is a three-layer stacked neural network. This three-layer stacked neural network model can be a multi-task model. Specifically,

[0117] Bottom layer: Receives MFCC feature sequences and extracts basic acoustic features through multiple fully connected layers or convolutional layers.

[0118] The intermediate layer: Based on the basic acoustic features, further processing is performed through multiple parallel task-specific paths. Specifically, the speaker log branch contains a joint distributed coding function to generate a speaker composite feature distribution; the emotion recognition branch extracts emotion tone parameters; and the rhythm analysis branch extracts rhythm dynamic vectors.

[0119] Output layer: Each task-specific path corresponds to an output port to output the corresponding feature vector. The speaker log branch outputs the speaker embedding vector, the emotion recognition branch outputs the emotion vector, and the rhythm analysis branch outputs the rhythm vector.

[0120] For example, taking “Speaker A” as an example, the speaker log branch can first perform feature matching in the “speaker” tag library based on the multi-granularity speaker features (i.e., speech content representation, which is a high-level feature) corresponding to a speech segment. If “Speaker A” is matched, then “Speaker A” will be used as the speaker identity tag for speech segment 1.

[0121] For example, a three-layer stacked deep neural network model can compare the emotional tone parameters corresponding to speech segment 2 with the historical emotional tone parameters of "speaker A" (such as the emotional tone parameters corresponding to the previous speech), and determine the emotion label corresponding to speech segment 2 based on the comparison result; then, the emotion label of speech segment 2 is assigned to each audio and video frame in speech segment 2. Of course, in actual implementation, the three-layer stacked deep neural network model can also preset the emotion level, and determine the emotion level by comparing the emotional tone parameters with the emotional tone parameters corresponding to the emotion level, and use the matched emotion level as the emotion label corresponding to speech segment 2. The same applies to rhythm labels. For the method of determining speech rhythm labels, please refer to emotion labels, which will not be repeated here.

[0122] In some embodiments of this application, the aforementioned mood tone parameter can be a three-dimensional vector. ,in, These represent the parameters of pleasure, arousal, and dominance, respectively. These tone-based parameters not only enhance contextual awareness during semantic analysis but also provide a basis for potential topic distribution in directory construction and summary generation.

[0123] In some embodiments of this application, the speaker-related label generation model can first perform speaker identity label recognition, and after recognizing the speaker identity label, perform emotion recognition and rhythm recognition on a per-speaker basis to obtain emotion labels and rhythm labels.

[0124] In some embodiments of this application, if a certain speech segment does not match a speaker identity tag in the aforementioned "speaker" tag library, the speaker-related tag generation model can generate a new speaker tag for the speech segment and add the tag and the corresponding speaker composite features to the "speaker" tag library. Alternatively, the speaker-related tag generation model can also distribute the spectral features or speaker composite features of the speech segment to the interference information tag generation model described below, which will then generate a speaker identity tag for the speech segment.

[0125] Thus, the speaker-related label generation model can automatically and deeply extract high-level semantic labels such as speaker identity and emotion from the spectral features of the first audio and video, thereby providing richer and more accurate contextual description features for the subsequent generation of structured line information, which can improve the accuracy of generating structured index information.

[0126] Thus, by using a speaker-related tag generation model based on spectral features, embedding vectors for at least one audio-visual frame of a speech segment can be obtained, with each speech segment including at least one consecutive audio-visual frame in the first audio-visual frame. Furthermore, clustering of these embedding vectors yields speaker identity tags corresponding to each audio-visual frame. This improves the accuracy of generating speaker identity tags for each audio-visual frame, providing richer and more accurate contextual description features for subsequent generation of structured index information, thereby enhancing the accuracy of generating structured index information.

[0127] In some embodiments of this application, step 102 may include step 102C.

[0128] Step 102C: The index information generation device performs context-aware silence detection, tone component recognition, and noise recognition processing on the audio and video features of the first audio and video through the interference information tag generation model, and obtains the interference information tag corresponding to each audio and video frame.

[0129] In some embodiments of this application, the interference information label generation model can be a Bi-LSTM network model.

[0130] The Bi-LSTM network model can employ forward and backward propagation mechanisms to perform context-aware silence detection and tone component recognition on the audio and video features of the first audio and video frame, thereby obtaining the interference information labels corresponding to each audio and video frame.

[0131] For example, a Bi-LSTM network model can employ forward and backward propagation mechanisms to perform context-aware silence detection and tone component recognition on the prosodic features of the first audio / video frame, thereby identifying whether speech disfluency exists in each audio / video frame. In other words, interference information tags indicate speech disfluency in the audio / video frames.

[0132] For example, the Bi-LSTM network model can employ forward and backward propagation mechanisms to perform context-aware noise recognition processing on the spectral features of the first audio and video frame to determine whether background noise and non-lexical sounds exist in each audio and video frame. That is, interference information tags indicate background noise or non-lexical sounds in the audio and video frame.

[0133] In some embodiments of this application, the Bi-LSTM network model can also perform context-aware speech recognition processing on the spectral features of the first audio and video to obtain preliminary text corresponding to the first audio and video.

[0134] In some embodiments of this application, the Bi-LSTM network model can be used to analyze silence detection and tone component recognition in unstructured temporal signals, with particularly dense annotation of tone words, repeated words, and meaningless filler words to obtain an interference information label sequence for the first audio and video. The interference information label sequence can provide acoustic density constraints for the segmentation of the first audio and video.

[0135] It should be noted that the above embodiments are illustrated by using audio and video features as the input of the Bi-LSTM network model. In actual implementation, the audio and video frame sequence of the first audio and video can also be used as the input. This application embodiment does not limit this.

[0136] In some embodiments of this application, after the speaker-related label generation model outputs the speaker-related labels corresponding to the first audio and video, the speaker-related labels can be input into a Bi-LSTM network model to perform context-aware joint modeling and calibration processing to obtain calibrated speaker-related labels. That is, the calibrated speaker-related labels from the Bi-LSTM network model can then be used in the calculation of navigation information.

[0137] In some embodiments of this application, the above formula 1 can also be set in the Bi-LSTM network model to more accurately verify the speaker identity label.

[0138] Thus, because interference information label generation models, such as Bi-LSTM network models, have excellent temporal modeling capabilities, they can accurately distinguish between effective speech and meaningless filler words, silence, background noise, and other interference information in the first audio and video. Therefore, interference information labels can accurately reflect the interference information in the first audio and video, thereby providing key acoustic clues for content purification and accurate segmentation of the first audio and video through interference information labels.

[0139] It is understandable that by introducing a multi-scale convolutional neural network based on audio and video features (i.e., the three-layer stacked neural network model mentioned above) and a Bi-LSTM network model, it is possible to accurately identify the speaker's identity, tone features and emotional state in the first audio and video, and to perform semantic separation processing on the ambiguous, overlapping speech and background noise in the first audio and video, thereby significantly improving the fidelity and comprehensibility of the first audio and video in terms of information dimension.

[0140] Step 103: The index information generation device generates structured index information for the first audio and video based on the above-mentioned tag sequence and the first text corresponding to the first audio and video.

[0141] In some embodiments of this application, the structured index information may include at least one of the following: summary, topic, hierarchical directory, relational mapping network, and target text corresponding to the first audio and video.

[0142] This relational mapping network can represent the mapping relationship between viewpoints and arguments, and the mapping relationship between questions and answers.

[0143] In some embodiments of this application, the hierarchical directory and relational mapping network described above are both temporally aligned with the first audio / video and the first text of the first audio / video, and a bidirectional mapping relationship is established. Therefore, by manipulating the text elements in the hierarchical directory or relational mapping network, the playback progress of the first audio / video can be switched, and the displayed content of the first text can be adjusted.

[0144] In some embodiments of this application, the target text mentioned above may be the text corresponding to the first audio / video and hierarchical directory, such as the third text described below.

[0145] Thus, since the aforementioned structured index information includes at least one of hierarchical directory, summary, and relational mapping network, users can quickly and accurately browse and acquire knowledge of the first audio and video content in a non-linear and cross-cutting manner based on the structured index information, thereby improving the flexibility of browsing the first audio and video content.

[0146] In some embodiments of this application, the index information generation device can generate structured index information of the first audio and video based on the tag sequence corresponding to the first audio and video and the first text corresponding to the first audio and video, through a structured content generation model.

[0147] In some embodiments of this application, the first text may include N semantically complete content units, where N is a positive integer.

[0148] In some embodiments of this application, the above-mentioned N content units are content units corresponding to audio segments obtained by segmenting the first audio and video based on the prosodic features, spectral features, and semantic features of the first audio and video.

[0149] For the segmentation method of obtaining N content units from the first text segmentation, please refer to the relevant description of step 104 in the following embodiments. To avoid repetition, it will not be repeated here.

[0150] It is understandable that after receiving N content units, the structured content generation model can transform them into interactive text blocks with titles, abstracts, argument levels, and indexes through a multi-level natural language modeling process.

[0151] In some embodiments of this application, the above-described structured content generation model can employ hierarchical natural language processing techniques to perform the following processing on each content unit: extracting core arguments using a BERT variant model; generating a table of contents and chapter titles corresponding to the first text based on the core arguments; automatically generating a summary using the Text Rank algorithm in conjunction with the core arguments; and constructing a tree-like table of contents based on the chronological order of events based on the core arguments.

[0152] In some embodiments of this application, such as Figure 4 As shown, the structured content generation model 203 described above may include:

[0153] 1. Argument Extraction Unit 2031, or Argument Extractor: It adopts a multi-head attention mechanism to identify the core propositions in the content unit and distinguishes between factual statements and opinion expressions through comparative learning, thereby obtaining the proposition tensor matrix and the proposition-text correlation matrix corresponding to the first audio and video.

[0154] 2. A structured content generation unit 2032, wherein the structured content generation unit includes at least one of the following:

[0155] 2.1) The abstract generation module 321 can be used to generate coherent abstracts using a fusion of extractive and generative methods. In simple terms, it can extract key phrases (such as the word chunks described below) from the first text based on the proposition tensor matrix extracted by the argument extraction unit and the proposition-text correlation matrix. Then, it uses a pointer network mechanism through a Transformer decoder to generate a coherent abstract, where the pointer network retains the extracted key phrases. In essence, the decoder can combine the semantics of the key phrases and use a pointer network approach to further expand and connect them to obtain the abstract.

[0156] 2.2) Directory building module 322 can be used to establish a three-level hierarchical directory structure. The top level is the overall theme of audio and video, the middle level is divided into stages according to the time development axis or the middle level is divided into stages by semantic clustering of arguments, and the bottom level corresponds to the specific arguments.

[0157] For example, the first layer is the overall theme of the podcast, such as "Discussion on the Ethics of Artificial Intelligence". The second layer is the stages divided in chronological order, such as "Background Introduction", "Technological Risks", "Legal Regulation", and "Summary". The third layer is the specific arguments in each stage, such as "Algorithmic Bias" and "Data Security" in the "Technological Risks" stage.

[0158] For example, timeline clustering preserves the original chronological order of the first text, which is the basic table of contents. However, semantic clustering, which disrupts the chronological order, might group content discussing the same argument or topic at different times together to form a new thematic view. For instance, it could group all arguments mentioning "data privacy" into a single cluster, regardless of when they appeared in the first audio or video.

[0159] 2.3) Relationship mapping module 323 can be used to create a hyperlink network between text elements, including the supporting relationship between viewpoints and arguments, and the logical connection between question-and-answer pairs.

[0160] For a detailed description of the structured content generation model, please refer to the relevant descriptions in the following embodiments.

[0161] In some embodiments of this application, the first text is the ASR text of the first audio / video.

[0162] In some embodiments of this application, the first text may include N semantically complete content units, where N is a positive integer; step 103 may include steps 103A to 103C below.

[0163] Step 103A: The index information generation device performs temporal alignment and textual encoding processing on the above N content units and the above tag sequence to obtain N multidimensional feature encoding sequences that correspond one-to-one with the N content units.

[0164] In some embodiments of this application, each multidimensional feature encoding sequence includes a text encoding of a content unit and a text encoding of a tag corresponding to that content unit. In other words, the index information generation device can align and encode non-text features such as speaker identity tags, emotion tags, rhythm tags, and non-information masks corresponding to the first audio and video with the content units corresponding to the timestamps, and jointly transform them into a unified, text-based multidimensional feature encoding sequence.

[0165] Step 103B: The index information generation device performs proposition structured decoding and association mapping processing on N multidimensional feature encoding sequences through the argument extraction unit to obtain the proposition tensor matrix and the proposition-text association matrix.

[0166] In this context, each proposition tensor in the aforementioned proposition tensor matrix can be used to indicate the semantic information of at least one argument and its corresponding supporting statement in the first text; the aforementioned proposition-text association matrix can be used to indicate the association strength between each proposition tensor and the corresponding supporting statement in the first text.

[0167] In some embodiments of this application, the index information generation device can first input N multidimensional feature encoding sequences into a first BERT model. The first BERT model can match corresponding CLS tags for each multidimensional feature encoding sequence from a preset artificial CLS tag library. Then, based on the multidimensional feature encoding sequences and the corresponding CLS tags, the first BERT model can determine the hidden state sequence corresponding to each multidimensional feature encoding sequence. This hidden state sequence contains the context-aware representation of each semantic unit in the multidimensional feature encoding sequence. Then, the N hidden state sequences corresponding one-to-one with the N content units are input into the argument extractor to perform proposition structured decoding and association mapping processing to obtain the proposition tensor matrix and the proposition-text association matrix. Alternatively, the first BERT model can first extract the key features of the multidimensional feature encoding sequences, such as simplifying the multidimensional feature encoding sequences using an attention mechanism. Then, the simplified multidimensional feature encoding sequences are used to match CLS tags in the artificial CLS tag library, and based on the matched CLS tags and the simplified multidimensional feature encoding sequences, the hidden state sequence corresponding to each multidimensional feature encoding sequence is determined.

[0168] In some embodiments of this application, the argument extraction unit can be an argument extraction unit based on a multi-head attention mechanism. Each attention head of the argument extraction unit can independently focus on the contextual logical differences between factual statements and subjective opinions in each multi-dimensional feature encoding sequence. Then, the recognition results of each of the multiple attention heads on the N multi-dimensional feature encoding sequence are fused to obtain the proposition tensor matrix and the proposition-text association matrix.

[0169] It is understandable that different attention heads have different attention preferences. For example, different attention heads can be trained using different training and labeled data to make them focus on different preferences.

[0170] For example, each attention head may learn to identify which parts of a text are objective facts, which are subjective opinions, and the logical relationships between them (such as support, refutation, explanation, etc.).

[0171] Among them, multiple attention heads have different attention preferences, such as focusing on semantic coherence, logical structure, or entity association respectively.

[0172] Specifically, some heads may be more concerned with semantic coherence and consistency, for example, judging whether the arguments explain or prove a point.

[0173] Other heads may pay more attention to logical connectors or argument structures, such as recognizing prompts like "because" and "for example".

[0174] Some heads may focus more on the co-occurrence and correlation of entities or events. For example, judging whether the facts mentioned in the argument are directly related to the claims in the viewpoint.

[0175] These attention points work independently, and their outputs are fused together to form a comprehensive and robust relational judgment, which is encoded in the proposition tensor matrix of the output and the proposition-text correlation matrix.

[0176] In some embodiments of this application, the optimization objectives for each attention head can be defined in conjunction with a contrastive learning strategy:

[0177] Formula 2

[0178] in, The discriminative contrastive loss function refers to the argument of a paragraph, where i represents the index of the triple. For expressing opinions, This represents an irrelevant sentence. For supporting arguments, in contrastive learning, each training sample typically consists of an "anchor" (opinion sentence), a "positive sample" (supporting argument), and a "negative sample" (irrelevant sentence), i.e. , and A triplet is formed, i is used to traverse these sample pairs or triplets, and Q represents the number of triplets. i and Q are both positive integers. This is a preset margin threshold used to control the minimum distance difference between "positive sample pairs" and "negative sample pairs". During training, loss is only incurred when the similarity difference between positive and negative sample pairs is less than δ, thus forcing the model to learn a more discriminative representation. Loss Function The aim is to maximize the semantic distance between supporting statements and the target proposition (determined by opinion statements) and suppress similarity with non-opinion statements, thereby achieving high-precision separation of opinion extraction. It can be calculated at the paragraph-level semantic dimension, extracting argument features from the divided paragraphs. This function can minimize the feature distance of the same argument and maximize the feature distance between different arguments. Opinion statements and argument statements have the same meaning; both refer to the statement containing the argument or the statement that can be used to extract an argument.

[0179] It can be understood that Formula 2 above is a variant of Triplet Margin Loss, the purpose of which is:

[0180] Make the point sentence Supporting arguments The representation should be as close as possible, even if Make it as large as possible.

[0181] Make the point sentence Unrelated sentences The meaning is to stay as far away as possible, that is As small as possible.

[0182] By introducing δ, the model's discriminative power is enhanced to force a difference of at least δ between the two.

[0183] In this embodiment of the application, for each attention head, after training the attention head to converge the above loss function by using the corresponding training data, the attention head has the ability to identify opinion statements, supporting statements and irrelevant statements.

[0184] Step 103C: The index information generation device uses a structured content generation model to perform structured fusion and organization processing on the first text based on the proposition tensor matrix and the proposition-original text association matrix to obtain structured index information.

[0185] In some embodiments of this application, the first text is subjected to structured fusion and organization processing based on the proposition tensor matrix and the proposition-text association matrix to obtain structured index information, which may include:

[0186] Based on the semantic information indicated by the proposition tensor matrix and the correlation strength indicated by the proposition-original text correlation matrix, the first text is subjected to structured fusion and organization processing to obtain structured index information.

[0187] It should be noted that, in the embodiments of this application, "based on the proposition tensor matrix" refers to: the semantic information indicated by the proposition tensor matrix; and "based on the proposition-text association matrix" refers to the association strength indicated by the proposition-text association matrix. Furthermore, the proposition-text association matrix can also directly or indirectly indicate the position of the corresponding argument and supporting statements in the first text for each proposition.

[0188] It should be noted that each proposition tensor in the above proposition tensor matrix can represent the semantic information of an argument and the supporting statements corresponding to that argument.

[0189] Thus, since the argument extraction unit can perform structured decoding and association mapping on the multidimensional feature encoding sequences corresponding to N content units to obtain the proposition tensor matrix indicating the semantic information of the arguments and corresponding supporting statements in the first text, and the proposition-text association matrix indicating the strength of the association between the proposition and the supporting statements in the first text; therefore, the proposition tensor matrix and the proposition-text association matrix can provide a unified and rich core data foundation for the subsequent generation of various logical index information, ensuring that the internal logic of the structured index information is consistent with the argument-related logic of the first audio and video, and improving the accuracy of the structured index information.

[0190] In some embodiments of this application, the structured index information includes a summary, and step 103C may include steps 103C1 and 103C2.

[0191] Step 103C1: The index information generation device uses the summary generation module in the structured content generation unit to locate and extract key phrases corresponding to key propositions from the first text based on the proposition-original text association matrix.

[0192] Step 103C2: The index information generation device generates coherent text connecting key phrases based on the proposition tensor matrix through the generation module, so as to form a summary of the first audio and video.

[0193] In some embodiments of this application, the above-mentioned summary generation module may adopt a fusion of extraction and generation strategies. By introducing a pointer network, key phrases in the original sentences are highlighted and retained in the source paragraph of the first text, and are used as part of the input of the Transformer decoder. The Transformer decoder then generates coherent text that connects the key phrases, thereby enhancing the coherence and information retention of the summary.

[0194] Specifically, the summary generation module may include a chunk extraction unit and a Transformer encoder. The chunk extraction unit is used to locate and extract key chunks corresponding to key propositions from the first text based on the proposition-text association matrix; the Transformer encoder is used to generate a coherent summary including keyword chunks based on the semantic information indicated by the proposition tensor matrix and the key chunks, using a pointer network generation mechanism.

[0195] In some embodiments of this application, the chunk extraction unit can extract a key chunk from each supporting statement of each key argument.

[0196] It is understandable that the above summary retains both the key information of the original text and semantic coherence.

[0197] Thus, since key phrases can be located and extracted from the first text based on the proposition-original text association matrix, and a coherent text connecting these key phrases can be generated based on the semantics indicated by the proposition tensor matrix, and this coherent text can be used as a summary of the first audio-visual content, it can be ensured that the final summary retains the keyword phrases from the first audio-visual content, is logically consistent with the first audio-visual content, and has strong readability. This ensures the accuracy of the summary content and the fluency of the language expression, thereby effectively reflecting the main content of the first audio-visual content. In other words, the summary in this application is a high-quality summary text of the first audio-visual content.

[0198] In some embodiments of this application, the structured index information mentioned above includes a hierarchical directory, and step 103C may include steps 103C3 to 103C6 below.

[0199] Step 103C3: The index information generation device uses the directory construction module in the structured content generation model to perform semantic clustering on the N content units based on the proposition tensor matrix, and obtains at least two topic clusters.

[0200] Step 103C4: The index information generation device obtains a summary title for each topic cluster based on the core arguments of each topic cluster through the directory construction module.

[0201] Step 103C5: The index information generation device, through the directory construction module, organizes the arguments corresponding to the proposition tensors belonging to the same topic cluster into child nodes under that topic cluster based on the proposition-text association matrix.

[0202] Step 103C6: The index information generation device organizes the topic of the first text, the general titles of at least two topic clusters, and the child nodes under each topic cluster into a hierarchical directory containing at least three logical levels through the directory construction module.

[0203] In some embodiments of this application, taking the construction of a hierarchical directory with three logical levels as an example, the directory construction module can logically establish a three-level tree-like directory. The root node of the three-level tree-like directory is the overall semantic theme of the first audio and video (i.e., the theme of the first text), the intermediate nodes are semantically clustered according to the time order of the content units, and the leaf nodes are at least one independent argument corresponding to each semantic cluster.

[0204] In some embodiments of this application, the topic of the first text can be abstracted from the semantics of the intermediate node, or it can be the original title of the first audio or video.

[0205] In some embodiments of this application, intermediate nodes may include at least one, which may be determined based on the semantic clustering results of the content units.

[0206] In some embodiments of this application, the index information generation device performs semantic clustering on N content units based on the proposition tensor matrix to obtain at least two topic clusters, including at least one of the following:

[0207] Method 1: The index information generation device performs semantic clustering on N content units according to the proposition tensor matrix and in accordance with time constraints to obtain at least two topic clusters.

[0208] Method 2: The index information generation device performs semantic clustering on N content units based on the proposition tensor matrix and according to the cross-time constraint method to obtain at least two topic clusters.

[0209] In one approach, the index information generating device can generate a hierarchical directory based on at least two topic clusters obtained in approach 1. It can be understood that the hierarchical directory generated according to approach 1 is aligned with the first text.

[0210] In another approach, the index information generating device can generate another hierarchical directory based on at least two topic clusters obtained in approach 2.

[0211] It is understandable that the directory generated according to method 2 will disrupt the chronological order of the content in the first text. In other words, the first text can be rearranged according to the topic clusters and their child nodes; and then a hierarchical directory can be generated based on the rearranged first text, the extracted topic clusters and related child nodes.

[0212] In another approach, the index information generation device can generate a hybrid embedded hierarchical directory based on at least two topic clusters obtained in approach 1 and at least two topic clusters obtained in approach 2.

[0213] Specifically, this hybrid embedded hierarchical directory is obtained by integrating the directories obtained through method 1 and method 2. For example, under the overall theme of this hybrid embedded hierarchical directory, there are two sets of subdirectories: one set of subdirectories is generated according to the topic clusters and their child nodes in method 1, and the other set of subdirectories is generated according to the topic clusters and their child nodes in method 2. In this way, the two subdirectories in the hybrid embedded hierarchical directory can be switched through a directory switching control.

[0214] In some embodiments of this application, for method 2 above, the directory building module can perform topic clustering using the following structure optimization function:

[0215] (Formula 3)

[0216] in, This indicates the topic clustering results. This represents the span between adjacent nodes on the time axis. The KL divergence represents the distribution of themes in the first audio and video, in other words... This can represent the span difference of the j-th node (or level) on the time axis. and This is a structural balance coefficient between time and semantics; This can reflect the "degree of dissatisfaction" of the current directory structure. The optimization goal of Formula 3 is to make... Minimize the structure to obtain a directory tree that is reasonably distributed in both time and semantics. This structural optimization ensures that the constructed directory conforms to the natural rhythm of podcast topic development and avoids the repetitive and scattered arrangement of homogeneous content; j represents the j-th node, and M is the number of nodes. Adjacent nodes refer to two nodes with semantic similarity at the same level, not those that are temporally adjacent. It can be understood that nodes can represent arguments or propositions, or they can represent content units.

[0217] It should be noted that the hierarchical directory is aligned with and mapped to the first audio and video in the time domain. Thus, by inputting a directory entry (such as a summary title for a topic cluster) within the hierarchical directory, the first audio and video can be switched to the playback position corresponding to that entry, achieving dual control over both audio / video and text.

[0218] Thus, since the semantics indicated by the propositional tensor matrix can be used to semantically cluster content units to form topic clusters, and a hierarchical directory structure including the upper-level generalization of topic clusters (i.e., the overall theme generalization of the first text), topic clusters, and child nodes of topic clusters can be constructed, it can be ensured that the hierarchical directory can reflect the topic hierarchy of the first audio and video, thereby enabling the hierarchical directory to provide users with a clear index path for quickly understanding both the macro framework and micro details of long content (i.e. the content of the first audio and video).

[0219] In some embodiments of this application, the structured index information includes a relational mapping network; step 103 above may include steps 103D to 103G below.

[0220] Step 103D: The index information generation device identifies question-type blocks in the first text based on the syntactic features of the statements in the first text through the relation mapping module.

[0221] Step 103E: The index information generation device, through the relation mapping module, identifies argument-type blocks and corresponding supporting blocks in the first text based on the proposition-text association matrix and the proposition tensor matrix, and identifies the answer-type blocks in the first text corresponding to each question-type block.

[0222] Step 103F: The index information generation device establishes a supporting relationship link between each argument block and its corresponding supporting block, and establishes a question-and-answer relationship link between each question block and its corresponding answer block through the relationship mapping module. It also generates a hyperlink path index through graph embedding to obtain the relationship mapping network.

[0223] It should be noted that "generating hyperlink path indexes via graph embedding" includes generating hyperlink path indexes based on each supporting relationship link and each question-and-answer relationship link using graph embedding. In other words, the hyperlink path index includes: the hyperlink index corresponding to each supporting relationship link and the hyperlink index corresponding to each question-and-answer relationship link.

[0224] In some embodiments of this application, the relationship mapping network includes all supporting relationship links and question-and-answer relationship links.

[0225] In some embodiments of this application, a "chunk" may include phrases, sentences, keywords, etc. For example, an argument-type chunk may include at least one opinion statement, or at least one word expressing an opinion, etc.

[0226] In some embodiments of this application, the relation mapping module can automatically construct a viewpoint-argument-question triplet network based on the syntactic features of the statements in the first text, the proposition-original text association matrix, and the proposition tensor matrix. This network connects each viewpoint node to its upstream and downstream logical support elements, such as the question or argument corresponding to the viewpoint. Hyperlink path indexing can be completed through graph embedding, and the data can be transmitted to a visualization model for magazine-style layout. All text elements in this layout structure have been time-coded with the first audio and video to ensure bidirectional jump and highlight tracking capabilities when entering the next stage of the interactive presentation interface.

[0227] It should be noted that after generating the hierarchical directory, the first text can be formatted according to the hierarchical directory, and each entry in the directory can be embedded into the first text to obtain the third text. Therefore, the above-mentioned "connecting the text elements in the hierarchical directory with the timecode of the first audio and video" actually refers to aligning the hierarchical directory with the first audio and video in the time domain.

[0228] In some embodiments of this application, the aforementioned relation mapping network can be directly integrated or embedded in a third text. This allows readers to access the corresponding argument statement by clicking on a specific argument statement while reading the third text; or to access all the answer statements associated with a specific question by clicking on a specific question.

[0229] For example, a viewpoint may correspond to multiple arguments, or it may correspond to a question, meaning that the viewpoint answers a question or raises a question. Therefore, the above triple can be "viewpoint-argument-question", where the viewpoint is connected to both the arguments and the question, while there is no direct connection between the question and the arguments.

[0230] In some embodiments of this application, each block in the relation mapping network can be referred to as a node.

[0231] For example, the viewpoint-argument-question triplet network is a directed graph structure, which includes:

[0232] Use opinion blocks, argument blocks, and question blocks as nodes;

[0233] The supporting relationship links are directed edges from the viewpoint node to the argument node;

[0234] The question-and-answer relationship link is used as a directed edge from the question node to the answer node;

[0235] The third type of link is used as a directed edge from the problem node to the viewpoint node.

[0236] Thus, by initially screening question-type phrases in the first text through syntactic features, and by analyzing the correspondence between arguments and evidence, and between questions and answers, we can clearly and accurately locate "question-answer phrase pairs" and "argument-argument phrase pairs" in the first text based on these correspondences. Then, by establishing a hyperlink path index between phrase pairs through graph embedding, we can obtain a relation mapping network. Through the relation mapping network, we can achieve rapid locking, switching, and aggregation between questions and answers or arguments and evidence, thereby further improving the practicality of the index information.

[0237] It is understandable that structured content generation units can be used to construct a structured and hierarchical text representation system for the first audio and video content, creating a structured and hierarchical text format corresponding to the first audio and video content. On the one hand, based on proposition extraction, summary generation, and hierarchical directory construction mechanisms, structured text with logical coherence and knowledge density can be generated. On the other hand, by introducing a relational mapping network, semantic links can be established, enabling users to quickly grasp the core content of the first audio and video content, jump to key information, and improve information acquisition efficiency.

[0238] In the method for generating line information provided in this application embodiment, since the label sequence containing speaker-related and interference information in the first audio and video can be determined according to the audio spectrum characteristics of the first audio and video, and the label sequence is combined with the text content of the first audio and video to generate index information, it can ensure that the generated index information has contextual relevance. In this way, high-quality index information of audio and video can be obtained, improving the efficiency of users obtaining key information in audio and video scenarios.

[0239] In some embodiments of this application, before step 103 above, the index information generation method provided in the embodiments of this application may further include the following step 104.

[0240] Step 104: The index information generation device uses a multimodal segmentation model to perform multimodal fusion decision processing on the prosodic features of the first audio and video, the second text corresponding to the first audio and video, and the above-mentioned tag sequence, to obtain the content segmentation result corresponding to the first audio and video.

[0241] The above content segmentation results can be used to segment the first audio and video into N semantically complete audio and video segments, which correspond one-to-one with the above N content units, where N is a positive integer.

[0242] In some embodiments of this application, the second text may be the same as or different from the first text.

[0243] For example, the first text is ASR text, and the second text is text obtained by Bi-LSTM conversion.

[0244] For example, both the first and second texts are ASR texts.

[0245] In some embodiments of this application, the multimodal segmentation model may also be referred to as a multimodal segmentation engine.

[0246] In some embodiments of this application, the multimodal segmentation engine can divide a continuous audio-video stream (such as the first audio-video mentioned above) into multiple candidate audio segments by analyzing the semantic coherence of the first audio-video and detecting changes in acoustic features. Then, an attention mechanism can be used to calculate the semantic similarity between adjacent audio segments, and combined with a speaker switching detection algorithm, the optimal segmentation point can be determined, thus obtaining N audio segments. In other words, a segmentation method that first determines the initial segmentation point and then performs clustering according to the temporal order can be adopted.

[0247] In some embodiments of this application, the multimodal segmentation engine can take over the tag sequence corresponding to the first audio and video. The core task is to cut the continuous audio and video content into the smallest semantic units with logical consistency and structural closure (such as the above N content units). In this process, the multimodal segmentation engine can integrate the acoustic mutation information corresponding to the first audio and video with the semantic similarity matrix corresponding to the second text, and construct a cross-modal collaborative optimization mechanism through deep fusion and decision modeling to achieve accurate segmentation of the first audio and video.

[0248] In some embodiments of this application, after the multimodal segmentation engine divides the first audio and video into N audio and video segments, the index information generation device can send each audio segment into an ASR engine. This engine adopts a mainstream end-to-end model architecture, which can ensure high transcription accuracy and the ability to model long texts. The ASR engine outputs preliminary transcribed text with word-level timestamps. Subsequently, the preliminary transcribed text is optimized by a text post-processing module.

[0249] For example, the text post-processing module may perform at least one of the following on the initially transcribed text:

[0250] Punctuation recovery: Recovering periods, commas, question marks, etc. using a pre-trained language model.

[0251] Text normalization: unifies the format of numbers, dates, currencies, etc.

[0252] Spoken Language Standardization: Based on the optimized configuration, filler words such as "uh," "um," and repetitive words can be selectively removed or retained to adapt to different application scenarios. This optimized configuration can indicate which words should be retained in the text and which words should be deleted.

[0253] In some embodiments of this application, the multimodal segmentation engine described above can be an architecture based on Transformer-Transducer or Conformer.

[0254] In some embodiments of this application, such as Figure 3 As shown, the above multimodal segmentation model 202 may include:

[0255] The Speech Feature Analysis Unit 2021 is a convolutional recurrent neural network that can generate the acoustic tensor corresponding to the first audio and video based on the extracted prosodic features of the first audio and video, and establish an acoustic change time series map.

[0256] Semantic Understanding Unit 2022: Used to calculate the vector space similarity of adjacent sentences using a pre-trained language model, and combined with dependency syntax to analyze and detect topic transition nodes;

[0257] Multimodal fusion unit 2023: used to spatiotemporally align acoustic feature change maps (such as temporally sequenced acoustic tensor matrices and / or label sequences) with semantic similarity matrices, and then dynamically adjust the weight allocation of features from different modalities through a gating mechanism;

[0258] Decision Optimization Unit 2024: Employs the Viterbi algorithm to find the globally optimal segmentation scheme within a sliding window, ensuring that the segmentation point simultaneously satisfies the acoustic feature mutation and topic switching conditions; that is, it is used to find the optimal segmentation point for the first audio and video.

[0259] Post-segmentation processing unit 2025: Further verifies the optimization results of decision optimization unit 2024. For example, when the duration of an audio segment exceeds a preset threshold, a recursive call to the segmentation algorithm can be used for secondary segmentation processing.

[0260] In some embodiments of this application, step 104 may include steps 104A to 104D.

[0261] Step 104A: The index information generation device performs temporal modeling on the prosodic features of the first audio and video through the speech feature analysis unit in the multimodal segmentation model, and obtains the temporal acoustic tensor matrix corresponding to the first audio and video.

[0262] In some embodiments of this application, the prosodic features of the first audio and video can be time-series modeled by a speech feature analysis unit to generate a time-series acoustic tensor matrix corresponding to the first audio and video, so as to capture fine-grained acoustic changes of the first audio and video through the time-series acoustic tensor matrix.

[0263] For example, the speech feature analysis unit can be a convolutional recurrent neural network (CRNN). This CRNN can form a temporal structured tensor corresponding to each audio and video frame based on the prosodic features of the audio and video frame sequence, such as the extracted fundamental frequency profile, short-time energy distribution, and formant trajectories of each audio and video frame. This temporal structured tensor can be used to represent the pitch, energy, and first three formants of each frame.

[0264] in, The acoustic feature tensor of the first audio and video at time t is a vector or array that comprehensively describes the multidimensional characteristics of the speech at time t. The fundamental frequency, representing time t, corresponds to the pitch of speech and is an important feature of the speaker's identity and tone. Represents the short-time energy at time t. It reflects the loudness or intensity of speech and helps in detecting stress and silence. The formant frequency of the first audio / video signal at time t is the i-th formant, typically i = 1, 2, or 3. The formant frequency determines the timbre of the vowel and is a key characteristic of the pronunciation.

[0265] It can be understood that the temporal structured tensors corresponding to all audio and video frames in the audio and video frame sequence constitute the temporal acoustic tensor matrix or acoustic change temporal map corresponding to the first audio and video. Through analysis... The change over time t, for example, calculation The first derivative of the α can detect abrupt changes at the acoustic level, such as sudden changes in pitch, sharp increases in energy, or switching of formants. These points often correspond to turn-taking, emphasis, or the start of a new topic.

[0266] Step 104B: The index information generation device performs sentence similarity analysis on the second text through the semantic understanding unit in the multimodal segmentation model to obtain the sentence-level semantic similarity matrix corresponding to the first audio and video.

[0267] In some embodiments of this application, the semantic understanding unit may employ sentence vector representation based on the standard BERT model to construct a sentence-level semantic similarity matrix corresponding to the first audio and video. Among these, adjacent sentence pairs... The cosine similarity can be defined as follows:

[0268]

[0269] in, Sentence and sentences Cosine similarity; and They are sentences and The context embedding vectors, which are respectively the sentence... and The sentence vector representation, where i and j represent different sentences or objects, This indicates the calculation of the Euclidean norm.

[0270] It is understandable that the sentence-level semantic similarity matrix mentioned above, after normalization, can be used as an importance matrix for semantic coherence in the fusion of multimodal features, and the first audio and video can be segmented based on the fusion result. It is understood that "multimodal features" here can include: label sequences and temporally sequenced acoustic tensor matrices.

[0271] In some embodiments of this application, the context embedding vector of a sentence can be determined with the help of the sentence's corresponding label.

[0272] In some embodiments of this application, adjacent statement pairs The similarity between two adjacent sentences is calculated on a semantic dimension. This calculation is independent of the split audio and video frames. It does not mean that each frame corresponds to multiple sentences, but rather that the similarity is calculated between adjacent sentences after speech-to-text conversion.

[0273] Step 104C: The index information generation device determines the globally optimal segmentation point sequence by using the multimodal fusion unit in the multimodal segmentation model, based on the acoustic mutation intensity measured by the first derivative of the temporal acoustic tensor matrix, the semantic coherence changes represented by the sentence-level semantic similarity matrix, and the speaker identity, emotion, and rhythm changes represented by the tag sequence.

[0274] Step 104D: The index information generation device uses the global optimal segmentation point sequence as the content segmentation result.

[0275] It is understandable that the multimodal fusion unit can fuse the temporal acoustic tensor matrix, sentence-level semantic similarity matrix, and label sequence through concatenation or attention-based weighting to form a unified and information-rich fusion feature vector. Then, based on this fusion feature vector, the globally optimal segmentation point sequence is determined, and the globally optimal segmentation point sequence is used as the content segmentation result.

[0276] In some embodiments of this application, the multimodal fusion unit may employ a gating control mechanism to achieve weight scheduling between acoustic and semantic modalities. The "acoustics" in "acoustics and semantic modalities" can be determined by a temporally sequenced acoustic tensor matrix and a label sequence, while the "semantics" can be determined by a sentence-level semantic similarity matrix.

[0277] It should be noted that step 104 is illustrated by the example of the tag sequence also participating in audio and video segmentation. In actual implementation, the tag sequence may not participate in audio and video segmentation. That is, it can be fused by splicing or attention weighting mechanism based only on the temporal acoustic tensor matrix and sentence-level semantic similarity matrix to form a fused feature vector.

[0278] In some embodiments of this application, the multimodal fusion unit can define a fusion activation function at each candidate cut position t. :

[0279]

[0280] in, This represents the segmentation activation at time t, which assesses the necessity of content segmentation at time t. It is the first derivative of the acoustic tensor of the audio / video frame, and can measure the intensity of the acoustic abrupt change corresponding to the audio / video frame. It refers to the semantic similarity between adjacent sentences. This indicates the semantic dissimilarity between adjacent sentences. A higher value indicates lower semantic coherence and a higher probability of topic shift. and These are the control coefficients for the two modalities, and this function is used to jointly evaluate the necessity of content segmentation. Wherein, when When the first threshold is exceeded and the label changes, the current frame is determined as the initial cut-off point. This label change can be understood as a difference from the label of the previous frame, such as a change in the speaker's identity.

[0281] In some embodiments of this application, if the speaker's identity remains unchanged, but the speaker's emotions or rhythm change, this will also serve as a reference for determining the initial cutting point.

[0282] For example, if a person's emotions are continuously rising while speaking, it is appropriate to segment the audio from a larger time dimension to ensure that the segmented audio segments retain the speaker's complete emotional process.

[0283] For example, if a person's tone of voice changes from high to low, that is, an emotional shift has occurred, so the emotion can be considered to meet the segmentation criteria. If the threshold is exceeded, it can also be used as a preliminary cutting point.

[0284] In some embodiments of this application, interjections can also be used to segmentation for interference information tags. For example, if a person says "yes, yes, yes," interjections expressing attitude cannot be segmented, but should be placed in the same audio segment.

[0285] It can be understood that the "candidate cutting position" can be the end position of each audio and video frame, that is, the position between two adjacent audio and video frames can be used as the candidate cutting position.

[0286] The fusion activation function is defined by combining acoustic mutation intensity, semantic similarity between adjacent sentences, and control coefficients.

[0287] In some embodiments of this application, in order to eliminate the missegmentation caused by sudden background noise or short-term semantic breaks, after determining the initial segmentation point, the initial segmentation point can be optimized by a decision optimization unit.

[0288] Specifically, the decision optimization unit can use a dynamic programming mechanism and a global segmentation scoring method that maximizes Viterbi path within a sliding window to re-examine the initial segmentation points in order to output the optimal segmentation sequence.

[0289] For example, if there is noise in the sliding window, the sliding window is moved backward, and the content of the window before and after the movement is analyzed. For example, the split point 1 and 2 are merged, that is, the split point 1 is canceled.

[0290] In some embodiments of this application, the decision optimization unit may use a chapter clustering algorithm to re-examine the initial segmentation points in order to output the optimal paragraph division sequence. The chapter clustering algorithm is a post-processing and analysis method. After obtaining a series of "initial segmented paragraphs", multiple adjacent and semantically similar initial paragraphs are merged or clustered into a larger "chapter" based on the semantic similarity between paragraphs (e.g., they all discuss the same sub-topic).

[0291] In some embodiments of this application, the chapter clustering algorithm can be any of the following: hierarchical clustering algorithm, or a clustering algorithm based on time-constrained DBSCAN.

[0292] For example, taking hierarchical clustering algorithms as an example, the clustering process may include:

[0293] Step 1: Receive the semantic vector representations of N content units (such as BERT sentence vectors).

[0294] Step 2: Calculate the similarity matrix based on the semantic vector representations of the N content units, such as using cosine similarity to calculate the similarity between all paragraph pairs.

[0295] Step 3, Initialization: Treat each content unit as an independent cluster.

[0296] Step 4, Iterative Merging:

[0297] 4.1 Find the two most similar clusters among all current clusters.

[0298] 4.2. Merge these two clusters into a new cluster.

[0299] 4.3 Update the similarity matrix, such as using strategies like single connection, full connection, or average connection.

[0300] Step 5: When the number of all clusters is less than the preset number, output the chapter clustering results, with each cluster corresponding to one chapter.

[0301] In some embodiments of this application, after the decision optimization unit outputs the optimal paragraph segmentation sequence, the post-segmentation processing unit can verify the paragraph duration corresponding to the optimal paragraph segmentation sequence. For units exceeding a preset duration threshold, recursive backtracking is performed to ensure that the final generated content units possess both semantic integrity and structural consistency of podcast information. This segmentation result is then used as input to the structured content generation model for subsequent generation of structured index information such as chapter generation and viewpoint extraction.

[0302] It can be understood that the above N audio and video segments are obtained by segmenting the first audio and video using the segmentation sequence verified by the segmentation post-processing unit.

[0303] Thus, because the segmentation decision for the first audio and video is made by incorporating multimodal features such as acoustic abrupt change intensity, semantic coherence changes, and speaker-related change information in the first audio and video, it is possible to more comprehensively and accurately identify the content boundaries that conform to human perception, and therefore the segmentation results obtained are more accurate and reliable.

[0304] It is understandable that step 104 enables accurate segmentation and semantic division of multimodal content. It allows for precise segmentation and clear semantic division of multimodal content. Step 104 employs a segmentation strategy that integrates acoustic features and language models, combined with semantic boundary detection and chapter clustering algorithms. This effectively solves the problems of semantic fragmentation and paragraph confusion in traditional transcribed texts, achieving automatic extraction and clear division of different themes, topics, and paragraphs in podcasts.

[0305] Thus, by utilizing the prosodic features, tag sequences, and multimodal features such as the second text of the first audio and video, the first audio and video can be segmented into N semantically complete audio and video segments. This provides more semantically complete and logically unified content units for the subsequent generation of structured index information, thereby providing a semantic foundation for the generation of high-quality structured index information.

[0306] In some embodiments of this application, after step 103 above, the index information generation method provided in the embodiments of this application may further include the following step 105.

[0307] Step 105: The index information generation device displays the structured index information and establishes a bidirectional mapping relationship between the text elements in the structured index information and the first audio and video time axis through a dynamic programming shortest path matching algorithm.

[0308] The aforementioned bidirectional mapping relationship can be used to achieve bidirectional jumps between text elements and audio / video clips, as well as synchronized highlighting of corresponding text when playing audio / video.

[0309] In some embodiments of this application, the index information generation device can generate a layout framework containing chapter tags, summary boxes, and index sidebars based on structured index information using a visualization model, and present responsive structured index information to users by parameterizing the structured index information to adapt to the user interface (UI) of different platforms; that is, by using a dynamic programming shortest path matching algorithm, a bidirectional mapping relationship is established between text elements in the structured index information and the first audio and video timeline.

[0310] In some embodiments of this application, such as Figure 5 As shown, the visualization model 205 may include: a spatiotemporal alignment module 51, a hot zone annotation module 52, a jump control module 53, a playback synchronization module 54, and a user annotation module 55;

[0311] Among them, the spatiotemporal alignment module 51 can establish an accurate mapping between the position of text characters and the audio and video timecodes, and the spatiotemporal alignment module 51 can use dynamic programming algorithm to compensate for speech recognition delay.

[0312] The hotspot annotation module 52 can identify key entities based on part-of-speech tagging, such as adding floating explanation boxes for technical terms and names of people.

[0313] The jump control module 53 enables non-linear indexing, allowing the structured index information to support hierarchical jumps and keyword search positioning based on the directory tree; that is, non-linear navigation between text contents is achieved through the jump control module 53.

[0314] The playback synchronization module 54 can automatically highlight the corresponding text during audio and video playback and update the reading progress indicator in real time.

[0315] The user annotation module 55 allows users to add personal notes and generate personalized content indexes.

[0316] It should be noted that the visualization model can receive user input, such as the first input described below.

[0317] In some embodiments of this application, the interactive presentation interface serves as an interactive hub for structured index information facing the user. Its core task is to accurately map the text structure with the original audio and video streams in the time dimension and establish a two-way control relationship.

[0318] In some embodiments of this application, the visualization model can generate operation text elements corresponding to the structured index information based on the structured index information, and display these text elements in the interface; that is, the structured index information is displayed through an interactive presentation interface.

[0319] In some embodiments of this application, the visualization model can establish a bidirectional mapping relationship between structured index information and the first audio and video on the time axis.

[0320] In some embodiments of this application, the citation information generation device can use a spatiotemporal alignment module and a dynamic programming shortest path matching algorithm to precisely match each text element (such as argument sentences and table of contents entries) in the structured index information with the start and end time points on the audio and video timeline, thereby constructing a bidirectional mapping table of "text position and audio / video timecode". This mapping requires compensation for text delay and speech recognition errors.

[0321] In some embodiments of this application, the index information generation device can use dynamic rendering technology to mark hot zones of text paragraphs, and users can trigger precise jumps to corresponding audio and video segments by clicking on text elements.

[0322] In some embodiments of this application, the spatiotemporal alignment module can utilize the text time tags and audio frame markers in the aforementioned structured index information to construct a mapping between the start and end positions of text paragraphs and the audio / video timeline. To address the offset error caused by speech recognition delay, the spatiotemporal alignment module introduces a shortest path matching algorithm based on dynamic programming, defining the time alignment loss function as:

[0323]

[0324] in, This represents the time alignment loss value, which is a scalar value. The smaller the value, the more accurate the alignment between the text paragraph and the audio frame. This represents the theoretical time position of the k-th breakpoint in the text segment. The first term represents the actual frame time corresponding to the audio stream. The first term is the absolute position error term, used to measure position error; the second term... The relative error term is used to measure rhythm offset. The parameter λ is used to adjust the balance between structural consistency and positional accuracy, i.e., rhythm consistency weight; k is the breakpoint index, and L is the total number of breakpoints. L is determined by the above content segmentation results, such as L=N+1, that is, L is the number of boundary points on which the first audio and video is divided into N audio segments. Therefore, L can also be called the number of boundaries to be aligned.

[0325] In some embodiments of this application, the time alignment loss function can achieve seamless mapping from character-level granularity to frame-level timecode through alignment path backtracking operations, providing time anchors for hotspot annotation and jump mechanisms. The hotspot annotation module calls a multi-layer part-of-speech analysis model to identify proper nouns, technical terms, and high-frequency entities from structured content. It combines the contextual discriminability of BERT embedding to construct a floating explanation panel, which is dynamically bound to interactive nodes during view rendering. User note insertion and content customization expansion are realized through the editable Document Object Model (DOM) structure.

[0326] In some embodiments of this application, the jump control module can use the hierarchical directory built in the previous stage as a navigation framework to respond to various user interaction commands (such as keyword search, highlight tag click or chapter node switching), and rely on an efficient index data structure to achieve non-linear and accurate content navigation.

[0327] Specifically, the jump control module establishes two complementary quick location indexes for all navigable units (such as paragraphs, chapters, and keyword anchors):

[0328] (1) Hash index: Calculate a hash value for each keyword, tag or unique content identifier, which is directly mapped to a pointer or address of the target fragment in storage.

[0329] (2) Segment-level skip list: A multi-level skip list index is built for audio / text segments in chronological order, which supports fast interval jumping and positioning in an ordered sequence.

[0330] Through hash indexing and segment-level skip list mechanisms, the jump control module can optimize the time complexity of most navigation operations to an amortized approximate O(1). "Amortized" means that, over the long term or in average use (e.g., handling a large number of user requests), the average time for a single jump operation is close to constant time. Although occasional slightly higher latency may occur due to hash collisions or skip list traversal, these situations account for a very small percentage of the overall operation, and the average cost remains extremely low. "Approximate O(1)" means that, in engineering practice, the time represented by "approximate O(1)" does not increase significantly with the total length of the content or the size of the data, remaining almost constant and extremely short.

[0331] Therefore, after any user interaction is triggered, the jump control module can instantly jump to the target audio and video segment, thus ensuring the high smoothness and real-time responsiveness of the interactive interface and realizing an efficient non-linear content exploration experience.

[0332] In some embodiments of this application, the playback synchronization module may employ an adaptive refresh strategy, automatically adjusting the size and offset direction of the highlighted area window based on the difference between the playback frame rate and the text progress. The offset formula is as follows:

[0333]

[0334] in, This represents the synchronization offset between the text and video corresponding to the current frame, where t represents system time and is a control factor. The refresh rate is dynamically adjusted based on network latency and device frame buffering to ensure a smooth and real-time reading experience. This represents the playback timestamp of the audio / video at system time t, or simply the video timestamp. For example, it indicates that the content at the 120.5-second mark of the first audio / video is being played. This represents the theoretical timestamp corresponding to the currently highlighted text on the interface at system time t, and is simply referred to as the text highlight timestamp. This timestamp comes from the best spatiotemporal mapping table between text, audio, and video generated by the spatiotemporal alignment module. It represents the instantaneous rate of change of the video timestamp relative to the system time, i.e., the playback rate of the current video; This represents the instantaneous rate of change of the text highlight timestamp relative to the system time, i.e., the rate at which the system is currently driving text highlighting.

[0335] In this way, the playback synchronization module can continuously monitor the speed difference between "where the video is actually playing" and "where the text should be highlighted" using the aforementioned offset formula, and immediately generate a correction instruction. The position and scrolling speed of the text highlight window are dynamically adjusted to achieve smooth, real-time synchronization of "audio, visual and text" in the user's eyes, masking the underlying latency and jitter.

[0336] In some embodiments of this application, through the user annotation module, users can provide a personalized interaction mechanism. Users can annotate on the structured index information. Each note or underline action is embedded with a timestamp index and serialized into an index node to participate in the content reordering logic. All annotation behaviors will be synchronously uploaded to the updated model as feedback features for subsequent model optimization.

[0337] It is understandable that, in order to make user interaction smoother and information search more convenient, a two-way synchronization mechanism between audio timeline and text content can be used, combined with interactive functions such as keyword highlighting, paragraph hotspots, and tag jumping, so that users can achieve "listen-read-jump" linked operation, improve the intuitiveness, accuracy and smoothness of content browsing, and be suitable for a variety of application scenarios such as learning, summarizing, and content reconstruction.

[0338] Thus, since structured index information can be displayed, and the structured index information has a bidirectional mapping relationship between the original audio and video timelines, the text-based index information can be precisely associated with the corresponding audio and video positions. Therefore, the playback progress of the first audio and video can be precisely controlled based on the structured index information.

[0339] In some embodiments of this application, after step 105 above, the index information generation method provided in the embodiments of this application may further include the following steps 106 and 107.

[0340] Step 106: The index information generation device receives the user's first input on the target text elements in the structured index information.

[0341] Step 107: The index information generation device responds to the first input and performs the first operation.

[0342] The first operation mentioned above includes any one of the following:

[0343] Control the first audio / video to jump to the playback position corresponding to the target text element;

[0344] Display the text content in the first text that corresponds to the target text element.

[0345] In some embodiments of this application, the aforementioned target text element refers to the basic interactive unit constituting structured index information, which originates from any one of the following: summary, hierarchical directory, and relational mapping network.

[0346] For example, a target text element could be a sentence in the abstract, a chapter title in the table of contents, a viewpoint or argument node in a relational network, or a word, sentence, or paragraph in the target text.

[0347] For example, indexing from text to audio / video. When a user clicks on a key sentence in the summary, the index information generation device can control the first audio / video to jump to the starting time point corresponding to that sentence and start playing, based on the mapping relationship corresponding to the key sentence, thus achieving "click to listen / watch".

[0348] For example, the association between texts can be expanded. When a user clicks on a chapter title (target text element) in the hierarchical directory, the index information generation device can highlight and display all the original text content corresponding to that chapter in the first text in the sidebar or pop-up window, making it convenient for the user to read and verify in depth.

[0349] In some embodiments of this application, when the first audio and video are playing normally, the index information generation device can automatically highlight the corresponding text elements in the structured index information, scroll the table of contents, highlight summary sentences, and update the progress indicator in real time according to the current playback timecode, so as to achieve "what is played is what is seen".

[0350] In some embodiments of this application, the index information generation device can also perform hotspot annotation and interpretation. For text elements such as technical terms and key entities in the index information, the index information generation device can provide instant explanations through a tooltip, which is achieved by integrating entity links and a knowledge base.

[0351] In some embodiments of this application, the index information generation device allows users to make personal notes, underline, or tag on the index information or first text. These annotations are saved and associated with the user account and a specific timestamp to form a personalized knowledge index.

[0352] In some embodiments of this application, the first input may include, but is not limited to: touch input from a user via a finger or stylus to the index information generation device, voice commands input by the user, specific gestures input by the user, or other feasible inputs. The specific input can be determined according to actual usage needs, and this application does not limit it. Specifically, the specific gesture in this application embodiment may be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-tap gesture; the click input in this application embodiment may be a single-click input, a double-tap input, or any number of clicks, and may also be a long-press input or a short-press input.

[0353] In this way, users can control audio and video to jump to the corresponding position or display the relevant original text by inputting text elements in the structured index information. This enables direct and accurate access from semantic index to specific media content, thereby greatly improving the efficiency and convenience of users browsing long audio and video content.

[0354] In some embodiments of this application, the index information generation method provided in this application may further include the following steps 108 and 109.

[0355] Step 108: The index information generation device collects interaction data between the user and the structured index information.

[0356] Step 109: The index information generation device dynamically optimizes the generation strategy of structured index information based on the above-mentioned interactive behavior data.

[0357] In some embodiments of this application, the aforementioned interactive behavior data may include at least one of the following:

[0358] User click intensity on text paragraphs, such as access frequency, duration of user stay on text paragraphs, navigation paths within text paragraphs, such as the relationship between previous and subsequent navigation, and user annotation behaviors, such as the number of notes and underlines, are more representative of any interactive data that characterizes the degree of user attention and usage patterns of each content segment within the indexed information.

[0359] In some embodiments of this application, the aforementioned interactive behavior data can be represented by a multidimensional feature vector, which is used to quantify the user's attention to and usage patterns of each content fragment within the index information.

[0360] In some embodiments of this application, the above-mentioned generation strategy may include a set of adjustable parameters for the final form and content of the structured index information.

[0361] For example: the length and compression ratio of the generated summary, the clustering granularity (number of topic clusters) of the hierarchical directory, and the threshold for displaying the strength of links in the relation mapping network.

[0362] In some embodiments of this application, the index information generation device dynamically optimizes the generation strategy of structured index information by updating the model based on the aforementioned interactive behavior data.

[0363] In some embodiments of this application, the updated model can continuously monitor user interactions, evaluate the quality of structured index information, and automatically adjust the generation strategy of structured index information based on reinforcement learning, forming a continuous learning closed loop of "interaction-evaluation-optimization-verification".

[0364] In some embodiments of this application, combined with Figure 2 ,like Figure 6 As shown, the update model 206 may include a monitoring and evaluation layer 61 and a decision and optimization layer 62. The monitoring and evaluation layer 61 is responsible for continuously collecting data and determining optimization trigger conditions, while the decision and optimization layer 62 is responsible for deciding how to optimize the generation strategy of structured index information.

[0365] In some embodiments of this application, such as Figure 6 As shown, the monitoring and evaluation layer 61 may include a user behavior analysis unit 611 and a content evaluation unit 612.

[0366] The user behavior analysis unit 611 can be used to track multi-dimensional user behavior on the interactive interface in real time, that is, to collect user access data on the above-mentioned structured navigation information, including but not limited to at least one of the following behaviors:

[0367] Click heatmap: Records the click frequency distribution of each paragraph, chapter, and hot zone;

[0368] Dwell time: The average reading / listening time users spend in front of a specific content unit;

[0369] Jump path: Analyze the non-linear navigation sequence of users between different content nodes;

[0370] Notes and annotations: Collects user-added personal notes, underlines, and tags.

[0371] User behavior analysis unit 611 can construct a multi-dimensional access vector for each content segment. ,in, Indicates click popularity. Indicates the duration of stay. Indicates the frequency of jumps. This vector represents user note coverage and is updated by sliding across the time dimension. It is used to calculate the paragraph access saturation index, which is determined when the following conditions are met within a consecutive period. Furthermore, this is accompanied by a significant focus on redirection paths, triggering a content optimization mechanism. In other words, the user behavior analysis unit can aggregate the above behaviors into a multi-dimensional access vector for each content segment i.

[0372] The content evaluation unit 612 is used to quantitatively evaluate the currently generated structured index information, such as evaluating the entropy of the summary information and the clarity of the directory.

[0373] In some embodiments of this application, such as Figure 6 As shown, the decision and optimization layer 62 may include an adaptive optimizer 621. The adaptive optimizer 621 incorporates a Deep Q-Network reinforcement learning model as its policy core. The adaptive optimizer learns the optimal policy by continuously trying actions and observing the results. When the user behavior analysis unit detects that the access saturation index of a specific paragraph continuously exceeds the limit and its corresponding content quality score is lower than a preset threshold, the adaptive optimizer 621 executes an optimization process for the structured index content.

[0374] For example, the adaptive optimizer adjusts the summary generation hyperparameters and reward function through a policy network based on deep Q-learning. The summary generation hyperparameters may include summary length, title abstraction level, and syntactic compression ratio. The reward function is updated based on feedback signals from subsequent user actions, such as reading completion rate and note activity. The policy function follows the Q-learning rule in its core updates.

[0375]

[0376] in, Let represent the expected value obtained by adopting the summary strategy parameter combination 'a' in state 's', where 'r' is the immediate reward calculated from feedback of subsequent interaction behaviors. Given the new state and η as the learning rate, this policy network forms a content update curve that is consistent with changes in user interests during multiple rounds of optimization.

[0377] In some embodiments of this application, to ensure that the integrity of user annotations is not compromised, such as... Figure 6 As shown, the updated model also includes an implementation and verification layer 63, which is responsible for safely executing optimization decisions and evaluating their effectiveness.

[0378] In some embodiments of this application, such as Figure 6 As shown, the implementation and verification layer 63 may include an incremental update unit 631 and an effect verification unit 632.

[0379] The incremental update unit 631 receives action instructions from the adaptive optimizer and performs precise, lossless updates to the target content. Its key responsibility is to ensure that all user-added personal notes, underlines, and annotations are fully preserved when modifying the summary or adjusting the structure, seamlessly migrating this user data to the new version of the content through semantic anchor mapping technology. In other words, it maintains the integrity of existing user annotations during index information optimization.

[0380] The effectiveness verification unit 632 is used to scientifically verify the optimization strategy using an A / B testing framework, comparing user retention rates under different strategies. A portion of user traffic is redirected to the optimized new content (experimental group), while the remaining users continue using the original content (control group). Key metrics such as user retention rate, bounce rate, return frequency, and interaction depth are continuously compared between the two groups. The verification result report will serve as crucial feedback, on the one hand to decide whether to fully release this optimization, and on the other hand as new empirical data fed back to the adaptive optimizer for further training of its reinforcement learning model.

[0381] In some embodiments of this application, the above-described updated model can perform the following operations:

[0382] The user behavior analysis unit tracks click heatmaps, dwell time, and navigation paths to construct multi-dimensional feature vectors.

[0383] The content evaluation unit calculates summary information entropy, table of contents hierarchy clarity, and question-answer pair coverage metrics.

[0384] The adaptive optimizer uses a deep Q-learning framework to dynamically adjust the summary generation length and title abstraction level;

[0385] The incremental update component maintains the integrity of existing user annotations when optimizing content;

[0386] The performance verification module compares user retention rates under different strategies through A / B testing.

[0387] In some embodiments of this application, the update model, as a core module for continuous optimization in the lifecycle of structured content, relies on the full set of user behavior trajectory data recorded in the preceding interactive interface to realize a feedback-driven summary content evolution mechanism.

[0388] For example, the updated model can be constructed by building a multi-dimensional access vector for each content segment in the user behavior analysis unit. ,in Indicates click popularity. Indicates the duration of stay. Indicates the frequency of jumps. This vector represents user note coverage and is updated by sliding across the time dimension. It is used to calculate the paragraph access saturation index, which is determined when the following conditions are met within a consecutive period. Furthermore, it is accompanied by a significant focus on the jump path, triggering the content optimization mechanism.

[0389] Then, the updated model receives the vectorized summary content output by the structured summary through the content evaluation unit and constructs the summary quality evaluation function. :

[0390]

[0391] Where Q represents the overall quality score of the structured summary content. Entropy is used to summarize information and measure the balance of key information distribution. To measure the cluster compactness of the hierarchical organizational structure and the clarity of the hierarchy of the directory; Question and answer coverage reflects the completeness of the structured content's response to user-submitted questions. α, β, and γ are dynamically adjusted weighting factors that represent the relative importance of information entropy, directory clarity, and question and answer coverage in the overall score, respectively.

[0392] in, Entropy is used to measure the balance of key information distribution in an abstract. If the entropy value is too high, the information may be scattered, and if it is too low, the information may be too concentrated. Ideally, it should cover multiple core arguments.

[0393] The quality of a hierarchical directory is typically evaluated by calculating the semantic tightness (high) of paragraphs within the same chapter and the semantic separation (high) between different chapters. The higher the value, the stronger the directory logic;

[0394] This is used to evaluate the completeness of structured content (such as extracted question-and-answer pairs) in responding to potential user questions. For example, if the content mentions a concept, does the system generate corresponding explanatory question-and-answer pairs?

[0395] Specifically, when a quality indicator, such as the aforementioned comprehensive quality score Q, fluctuates beyond the warning threshold, the updated model can invoke an adaptive optimizer to adjust the summary generation hyperparameters and reward function through a deep Q-learning-based policy network. The summary generation hyperparameters may include summary length, title abstraction level, and syntactic compression ratio. The reward function is updated based on feedback signals from subsequent user actions, such as reading completion rate and note activity. The policy function is as follows:

[0396]

[0397] in, Let represent the expected value obtained by adopting the summary strategy parameter combination 'a' in state 's', where 'r' is the immediate reward calculated from feedback of subsequent interaction behaviors. Given the new state and η as the learning rate, this policy network forms a content update curve that is consistent with changes in user interests during multiple rounds of optimization.

[0398] To ensure the integrity of user annotations is not compromised, the incremental update component employs a hierarchical mapping retention strategy. This maps old version annotation anchors to similar semantic segments within the optimized summary structure and generates annotation drift logs for user review. Finally, the performance verification module continuously conducts A / B group experiments, comparing content performance under different strategies based on user retention rate, bounce rate, and return visit frequency. The output results will serve as the experience replay input for the next round of Q-learning. The system's closed-loop optimization process is now complete, and the optimized content structure will be remapped to the interactive presentation interface to form an update push cycle.

[0399] In some embodiments of this application, in addition to employing a deep Q-learning framework, the update model can dynamically adjust the summary generation length and the title abstraction level. Its reward signal is determined by a composite function consisting of information quality scores (such as ROUGE and BERTScore), user behavior metrics (such as completeness rate and note activity), and optional explicit satisfaction feedback, and causal inference methods (such as IPS) can be used to debias the behavior metrics.

[0400] In some embodiments of this application, the index information generation device can construct an intelligent evolutionary closed loop and a summary strategy adaptive mechanism. This mechanism enables the index information generation device to self-optimize and continuously improve. It should be noted that by introducing a user behavior tracking and feedback mechanism, combined with the summary and structure recommendation strategies of a reinforcement learning model, the content generation logic can be self-adjusted and evolved, improving user satisfaction while achieving continuous optimization and refined operation of the system's content output. Through user behavior tracking and feedback, combined with the summary and structure recommendation strategies of a reinforcement learning model, the content generation logic can be self-adjusted and evolved, improving user satisfaction while achieving continuous optimization and refined operation of the system's content output.

[0401] Thus, because user interaction data can be collected and used to dynamically optimize the index information generation strategy through reinforcement learning models, the index information generation process becomes adaptive and can continuously keep up with users' actual usage preferences, thereby maintaining the practicality and effectiveness of the index information in the long term.

[0402] In some embodiments of this application, the index information generation method provided in this application may further include the following step 110.

[0403] Step 110: The index information generation device updates the structured index information using an incremental update mechanism according to the optimized generation strategy.

[0404] The updated structured index information retains user annotations.

[0405] Thus, by employing an incremental update mechanism and retaining user annotations when optimizing and updating structured index information, it is possible to ensure that users' personalized data and work results are not lost while continuously improving structured index information, thereby guaranteeing the consistency of user experience and the security of user assets.

[0406] The index information generation method provided in this application can transform unstructured audio and video dialogues into highly structured, machine-readable knowledge units (such as summaries, tables of contents, and text). Therefore, the application potential of this index information generation method goes far beyond optimizing the user listening experience; it can also serve as a key infrastructure and content creation engine in the next-generation artificial intelligence ecosystem.

[0407] For example, the index information generation method provided in this application generates structured index information that can serve as a high-quality fine-tuning data engine for domain-specific large language models (LLMs). While current large language models perform well in general knowledge, their depth and accuracy in specialized fields (such as medicine, law, and finance) remain insufficient. The structured text generated by this method, such as precisely divided chapters, extracted core arguments, supporting relationships between viewpoints and evidence, and structured question-and-answer pairs, can be directly used for efficient fine-tuning of the basic large model or as high-quality contextual background information in prompting projects. Therefore, accurate and reliable domain-specific AI models can be quickly built at a lower cost, effectively solving the core pain points of difficult acquisition of training data and high annotation costs in specialized fields.

[0408] For example, the structured index information described above is used for automated knowledge graph construction and reasoning: This method initially constructs a logical network between text elements through a relationship mapping component. This capability can be further extended to automatically inject or construct a dynamic, evolvable domain knowledge graph from extracted entities (names, organizations, terms), viewpoints, and their relationships. For instance, by processing hundreds of podcast episodes in a specific domain, the system can automatically construct a relationship graph between experts, theories, and points of contention in that domain, supporting complex knowledge queries and logical reasoning, thus achieving a leap from "information acquisition" to "knowledge discovery."

[0409] For example, this application enables the secondary generation and re-creation of cross-modal content: it can not only parse content, but also drive its re-creation. The structured elements generated by the system (such as chapter titles, key summaries, and highlight introductions) can serve as instructions to drive other generative AI models to perform cross-modal creation.

[0410] For example, automatically generating video trailers: by combining the generated chapter titles and summary text with the corresponding timecodes and inputting them into the video editing model, a short video trailer containing key shots and core ideas can be automatically generated.

[0411] For example, generating digital human broadcasts: feeding a summary and opinion text to a digital human (Avatar) model can generate a virtual anchor to broadcast or interpret the content of this podcast episode, creating a completely new form of content distribution.

[0412] Thus, this application is not only a methodological tool for converting audio and video to text, but also a "knowledge middleware" connecting original dialogue content with advanced AI applications. Through deep analysis and structuring, it unlocks the enormous potential of massive audio and video data in knowledge engineering, model training, and content re-creation, forming a key link in the future intelligent content ecosystem.

[0413] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0414] This application also provides an index information generation system, such as Figure 2 As shown, the index information generation system 200 may include: an audio feature extraction model 201, a tag generation model 202, and a structured content generation model 203;

[0415] The audio feature extraction model 201 is used to obtain the audio and video features of the first audio and video, wherein the audio and video features include at least one of spectral features and prosodic features;

[0416] The tag generation model 202 is used to determine a tag sequence based on the audio and video features. The tag sequence includes tags corresponding to each audio and video frame of the first audio and video, and the tags include at least one of speaker-related tags and interference information tags.

[0417] The structured content generation model 203 is used to generate structured index information of the first audio and video based on the tag sequence and the first text corresponding to the first audio and video.

[0418] In some embodiments of this application, the above-mentioned tag generation model includes a speaker-related tag generation model;

[0419] The speaker-related label generation model described above is used for:

[0420] Based on the spectral features, an embedding vector for at least one speech segment is obtained, and each speech segment includes at least one consecutive audio and video frame in the first audio and video frame.

[0421] Clustering is performed on the at least one embedding vector to obtain the speaker identity label corresponding to each audio and video frame.

[0422] In some embodiments of this application, the above-mentioned label generation model further includes an interference information label generation model;

[0423] The aforementioned interference information label generation model is used to perform context-aware silence detection, tone component recognition, and noise recognition processing on the audio and video features to obtain interference information labels corresponding to each audio and video frame.

[0424] In some embodiments of this application, the first text described above includes N semantically complete content units, where N is a positive integer;

[0425] The structured content generation model described above includes an encoding unit, an argument extraction unit, and a structured content generation unit.

[0426] The above-mentioned encoding unit is used to perform temporal alignment and textual encoding processing on the N content units and the tag sequence to obtain N multidimensional feature encoding sequences that correspond one-to-one with the N content units;

[0427] The aforementioned argument extraction unit is used to perform proposition structured decoding and association mapping processing on the N multidimensional feature encoding sequences to obtain a proposition tensor matrix and a proposition-text association matrix; each proposition tensor in the proposition tensor matrix is ​​used to indicate the semantic information of at least one argument and its corresponding supporting statement in the first text; the proposition-text association matrix is ​​used to indicate the association strength between each proposition tensor and the corresponding supporting statement in the first text;

[0428] The aforementioned structured content generation unit is used to perform structured fusion and organization processing on the first text based on the proposition tensor matrix and the proposition-text association matrix to obtain the structured index information.

[0429] In some embodiments of this application, the structured index information includes a summary; the structured content generation unit may include a summary generation module.

[0430] The above-mentioned summary generation module is used to locate and extract key phrases corresponding to key propositions from the first text based on the proposition-original text association matrix; and to generate coherent text connecting the key phrases according to the proposition tensor matrix to form the summary.

[0431] In some embodiments of this application, the structured index information mentioned above includes a hierarchical directory, and the structured content generation unit mentioned above may include a directory building module;

[0432] The above directory building module is used for:

[0433] Based on the proposition tensor matrix, semantic clustering is performed on the N content units to obtain at least two topic clusters;

[0434] Based on the core arguments of each topic cluster, a summary title for each topic cluster is obtained;

[0435] Based on the proposition-text association matrix, the arguments corresponding to the proposition tensors belonging to the same topic cluster are organized into child nodes under the topic cluster.

[0436] The topic of the first text, the general titles of the at least two topic clusters, and the child nodes under each topic cluster are organized into a hierarchical directory containing at least three logical levels.

[0437] In some embodiments of this application, the directory construction module described above is specifically used for:

[0438] Based on the proposition tensor matrix, semantic clustering is performed on the N content units according to the time constraint method to obtain the at least two topic clusters; or, based on the proposition tensor matrix, semantic clustering is performed on the N content units according to the cross-time constraint method to obtain the at least two topic clusters.

[0439] In some embodiments of this application, the structured index information mentioned above includes a relation mapping network, and the structured content generation unit mentioned above includes a relation mapping module;

[0440] The aforementioned relation mapping module is used for:

[0441] Based on the syntactic features of the statements in the first text, identify the question-type language blocks in the first text;

[0442] Based on the proposition-text correlation matrix and the proposition tensor matrix, the argument-type blocks and their corresponding supporting blocks in the first text are identified, and the answer-type blocks corresponding to each question-type block in the first text are also identified.

[0443] A support relationship link is established between each argument-type block and its corresponding support block, and a question-and-answer relationship link is established between each question-type block and its corresponding answer-type block. A hyperlink path index is generated through graph embedding to obtain the above relationship mapping network.

[0444] In some embodiments of this application, such as Figure 2 As shown, the index information generation system 200 also includes a multimodal segmentation model 204;

[0445] The aforementioned multimodal segmentation model 204 is used to perform multimodal fusion decision processing on the prosodic features, the second text corresponding to the first audio and video, and the tag sequence before the structured content generation model 203 generates the structured index information of the first audio and video based on the tag sequence and the first text corresponding to the first audio and video, so as to obtain the content segmentation result corresponding to the first audio and video.

[0446] The content segmentation result is used to segment the first audio and video into N semantically complete audio and video segments, and the N audio and video segments correspond one-to-one with the N content units.

[0447] In some embodiments of this application, the multimodal segmentation model includes: a speech feature analysis unit, a semantic understanding unit, and a multimodal fusion unit;

[0448] The aforementioned speech feature analysis unit is used to perform temporal modeling on the prosodic features to obtain the temporalized acoustic tensor matrix corresponding to the first audio and video.

[0449] The aforementioned semantic understanding unit is used to perform sentence similarity analysis on the second text to obtain the sentence-level semantic similarity matrix corresponding to the first audio and video.

[0450] The multimodal fusion unit is used to determine the globally optimal segmentation point sequence based on the acoustic abrupt change intensity measured by the first derivative of the acoustic tensor matrix, the semantic coherence change represented by the sentence-level semantic similarity matrix, and the speaker identity, emotion, and rhythm changes represented by the tag sequence; and to use the globally optimal segmentation point sequence as the content segmentation result.

[0451] In some embodiments of this application, such as Figure 2As shown, the index information generation system 200 also includes a visualization model 205;

[0452] The visualization model 205 is used to display the structured index information of the first audio and video after the structured content generation model 203 generates the structured index information of the first audio and video. It also constructs a bidirectional mapping relationship between the text elements in the index information and the timeline of the first audio and video through a dynamic programming shortest path matching algorithm. The bidirectional mapping relationship is used to realize bidirectional jump between text elements and audio and video segments, as well as synchronized highlighting of corresponding text when playing audio and video.

[0453] In some embodiments of this application, the index information generation system described above further includes: an update model 206;

[0454] The aforementioned visualization model is also used to collect data on user interaction behavior with the aforementioned structured navigation information;

[0455] The aforementioned updated model 206 is used to dynamically optimize the generation strategy of the aforementioned structured navigation information based on the aforementioned interaction behavior data through a reinforcement learning model.

[0456] In the index information generation system provided in this application embodiment, since the label sequence containing speaker-related and interference information in the first audio and video can be determined according to the audio spectrum characteristics of the first audio and video, and the label sequence is combined with the text content of the first audio and video to generate index information, it can ensure that the generated index information has contextual relevance. In this way, high-quality index information of audio and video can be obtained, improving the efficiency of users in obtaining key information in audio and video scenarios.

[0457] For other descriptions of the index information generation system provided in the embodiments of this application, please refer to the relevant descriptions of the index information generation method in the above embodiments. To avoid repetition, they will not be repeated here.

[0458] Index information generation in this application embodiment DeviceIt can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or any other device besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application does not specifically limit the scope of the embodiments.

[0459] The index information generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0460] The index information generation device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0461] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above-described index information generation method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0462] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0463] Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0464] The electronic device 1500 includes, but is not limited to, components such as: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 1510.

[0465] Those skilled in the art will understand that the electronic device 1500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0466] It should be understood that, in this embodiment, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042. The GPU 15041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1506 may include a display panel 15061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1507 includes at least one of a touch panel 15071 and other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0467] The memory 1509 can be used to store software programs and various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0468] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.

[0469] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described index information generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0470] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0471] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described index information generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0472] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0473] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described index information generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0474] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0475] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0476] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for generating index information, characterized in that, The method includes: Acquire the audio and video features of the first audio and video, wherein the audio and video features include at least one of spectral features and prosodic features; A tag sequence is determined based on the audio and video features. The tag sequence includes tags corresponding to each audio and video frame of the first audio and video. The tags include at least one of speaker-related tags and interference information tags. Based on the tag sequence and the first text corresponding to the first audio and video, generate structured index information for the first audio and video.

2. The method according to claim 1, characterized in that, The speaker-related tags include speaker identity tags. The step of determining the speaker identity tags based on the audio and video features includes: By generating a speaker-related label model, based on the spectral features, an embedding vector of at least one speech segment is obtained, and each speech segment includes at least one consecutive audio-visual frame in the first audio-visual frame. Clustering is performed on the at least one embedding vector to obtain the speaker identity label corresponding to each audio and video frame.

3. The method according to claim 1 or 2, wherein the step of determining the interference information tag based on the audio and video features includes: By using the interference information label generation model, context-aware silence detection, tone component recognition, and noise recognition processing are performed on the audio and video features to obtain interference information labels corresponding to each audio and video frame.

4. The method according to claim 1, characterized in that, The first text comprises N semantically complete content units, where N is a positive integer; The step of generating structured index information for the first audio / video based on the tag sequence and the first text corresponding to the first audio / video includes: Temporal alignment and text encoding processing are performed on the N content units and the tag sequence to obtain N multidimensional feature encoding sequences that correspond one-to-one with the N content units; The argument extraction unit performs proposition structured decoding and association mapping on the N multidimensional feature encoding sequences to obtain a proposition tensor matrix and a proposition-text association matrix. Each proposition tensor in the proposition tensor matrix is ​​used to indicate the semantic information of at least one argument and its corresponding supporting statement in the first text. The proposition-text association matrix is ​​used to indicate the association strength between each proposition tensor and the corresponding supporting statement in the first text. The structured content generation unit performs structured fusion and organization processing on the first text based on the proposition tensor matrix and the proposition-text association matrix to obtain the structured index information.

5. The method according to claim 4, characterized in that, The structured index information includes a summary; The step of performing structured fusion and organization processing on the first text using the structured content generation unit, based on the proposition tensor matrix and the proposition-text association matrix, to obtain the structured index information includes: The summary generation module in the structured content generation unit locates and extracts key phrases corresponding to key propositions from the first text based on the proposition-original text association matrix. Based on the proposition tensor matrix, a coherent text connecting the key phrases is generated to form the summary.

6. The method according to claim 4, characterized in that, The structured index information includes a hierarchical directory; The steps of performing structured fusion and organization processing on the first text based on the proposition tensor matrix and the proposition-text association matrix to obtain the structured index information include: Based on the proposition tensor matrix, semantic clustering is performed on the N content units to obtain at least two topic clusters; Based on the core arguments of each topic cluster, a summary title for each topic cluster is obtained; Based on the proposition-text association matrix, the arguments corresponding to the proposition tensors belonging to the same topic cluster are organized into child nodes under the topic cluster. The topic of the first text, the general titles of the at least two topic clusters, and the child nodes under each topic cluster are organized into a hierarchical directory containing at least three logical levels.

7. The method according to claim 6, characterized in that, Based on the proposition tensor matrix, semantic clustering is performed on the N content units to obtain at least two topic clusters, including: Based on the proposition tensor matrix, semantic clustering is performed on the N content units according to the time constraint method to obtain the at least two topic clusters; Alternatively, based on the proposition tensor matrix, semantic clustering can be performed on the N content units according to the cross-time constraint method to obtain the at least two topic clusters.

8. The method according to claim 4, characterized in that, The structured index information includes a relational mapping network; The steps of performing structured fusion and organization processing on the first text based on the proposition tensor matrix and the proposition-text association matrix to obtain the structured index information include: Based on the syntactic features of the statements in the first text, identify the question-type language blocks in the first text; Based on the proposition-text correlation matrix and the proposition tensor matrix, the argument-type blocks and their corresponding supporting blocks in the first text are identified, and the answer-type blocks corresponding to each question-type block in the first text are also identified. A support relationship link is established between each argument block and its corresponding support block, and a question-and-answer relationship link is established between each question block and its corresponding answer block. A hyperlink path index is generated through graph embedding to obtain the relationship mapping network.

9. The method according to claim 4, characterized in that, Before generating the structured index information of the first audio / video based on the tag sequence and the first text corresponding to the first audio / video, the method further includes: Using a multimodal segmentation model, multimodal fusion decision processing is performed on the prosodic features, the second text corresponding to the first audio and video, and the tag sequence to obtain the content segmentation result corresponding to the first audio and video; The content segmentation result is used to segment the first audio and video into N semantically complete audio and video segments, and the N audio and video segments correspond one-to-one with the N content units.

10. The method according to claim 9, characterized in that, The process involves using a multimodal segmentation model, based on the prosodic features, the second text corresponding to the first audio / video, and the tag sequence, to perform multimodal fusion decision processing to obtain the content segmentation result corresponding to the first audio / video, including: The prosodic features are temporally modeled using the speech feature analysis unit in the multimodal segmentation model to obtain the temporal acoustic tensor matrix corresponding to the first audio and video. The semantic understanding unit in the multimodal segmentation model is used to perform sentence similarity analysis on the second text to obtain the sentence-level semantic similarity matrix corresponding to the first audio and video. The global optimal segmentation point sequence is determined by the multimodal fusion unit in the multimodal segmentation model based on the acoustic mutation intensity measured by the first derivative of the acoustic tensor matrix, the semantic coherence change represented by the sentence-level semantic similarity matrix, and the speaker identity, emotion, and rhythm changes represented by the tag sequence. The global optimal segmentation point sequence is used as the content segmentation result.

11. The method according to claim 1, characterized in that, After generating the structured index information of the first audio and video, the method further includes: The structured index information is displayed, and a bidirectional mapping relationship between text elements in the index information and the first audio / video timeline is constructed using a dynamic programming shortest path matching algorithm. The bidirectional mapping relationship is used to realize bidirectional jump between text elements and audio / video segments, as well as synchronized highlighting of corresponding text when playing audio / video.

12. The method according to claim 11, characterized in that, The method further includes: Collect data on user interaction behavior with the structured index information; Based on the interaction behavior data, the generation strategy for the structured index information is dynamically optimized.

13. An index information generation system, characterized in that, include: Audio feature extraction model, tag generation model, and structured content generation model; The audio feature extraction model is used to obtain the audio and video features of the first audio and video, wherein the audio and video features include at least one of spectral features and prosodic features; The tag generation model is used to determine a tag sequence based on the audio and video features. The tag sequence includes tags corresponding to each audio and video frame of the first audio and video, and the tags include at least one of speaker-related tags and interference information tags. The structured content generation model is used to generate structured index information of the first audio and video based on the tag sequence and the first text corresponding to the first audio and video.

14. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the index information as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the index information as described in any one of claims 1 to 12.