Audio and video content segmentation and localization method and system

CN122842010APending Publication Date: 2026-09-29SHANDONG SANMU SUMMER INFORMATION & TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610887810.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]现有技术在处理音视频分割任务时,主要依赖固定时长的物理切分或基于视觉像素差异的切镜检测,但此类方法在应对复杂语义场景时表现出明显的片面性

Benefits of technology

1、本发明通过多模态语义融合技术,改变了以往单纯依赖视觉像素差异或固定时长切分的局限性。通过将视觉画面特征与音频文本语义在时间轴上深度对齐,本方案能够精准识别内容层面的主题转换点。这确保了分割后的每一个片段都是一个具有独立含义的逻辑单元,避免了传统方法中经常出现的语音被强行阻断、动作被切分在两个片段中的弊端。本方法在教育教学、会议记录等逻辑性强的视频处理中,片段语义完整率达到极高水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842010A_ABST
    Figure CN122842010A_ABST
Patent Text Reader

Abstract

The application discloses an audio-video content segmentation and positioning method and system, and relates to the technical field of visual language cross-modal positioning. The method comprises the following steps: acquiring audio-video original data and disassembling the audio-video original data into a video frame sequence and an audio stream; extracting a visual semantic feature vector sequence from the video frame sequence, and mapping the audio stream into a text semantic feature vector sequence after converting the audio stream into text through speech recognition; aligning the two feature sequences on a time axis, calculating a semantic consistency score through a sliding window, and extracting a fundamental frequency feature and a duration feature of the audio; dynamically adjusting a boundary early warning threshold according to the fundamental frequency and the duration feature, combining the semantic consistency score, visual cut detection and scene classification constraint judgment logic boundary candidate points; finally, adjusting the segmentation points to a mute interval or an energy minimum value point by using speech pause information, and performing lip shape synchronous detection fine adjustment to a mouth closing moment to determine a final segmentation moment. The application has the effect of improving the semantic integrity and audio-visual coherence of segmented fragments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of visual language cross-modal localization technology, specifically involving an audio and video content segmentation and localization method and system. Background Technology

[0002] With the rapid evolution of multimedia communication and internet technologies, audio and video data have become core carriers of information dissemination, online education, and digital entertainment. Audio and video content segmentation and localization, as fundamental technologies for multimedia content understanding and management, play a crucial role in areas such as massive resource retrieval, automated editing, and intelligent recommendation. As users' demands for interactive content experiences continue to grow, accurately extracting logically meaningful units from complex streaming media data has become a key research topic for improving multimedia resource processing efficiency and knowledge service quality.

[0003] Audio and video content segmentation aims to identify the transition boundaries between different themes or scenes by analyzing the features of raw streaming data, and to achieve precise temporal positioning. This process typically involves spatial feature analysis of video frames and frequency domain processing of audio signals, determining segmentation points by capturing abrupt changes or evolutionary patterns in audiovisual features. Effective segmentation techniques should not only transform long videos into indexable micro-units, but also ensure that the segmented segments possess a high degree of consistency and integrity in audiovisual presentation and narrative logic, laying the foundation for subsequent secondary use and deep retrieval.

[0004] Existing technologies for audio and video segmentation primarily rely on fixed-duration physical segmentation or segmentation detection based on visual pixel differences. However, these methods exhibit significant limitations when dealing with complex semantic scenarios. Fixed-duration segmentation mechanisms are completely detached from content logic, often resulting in the forced interruption of speech sequences or dynamic behaviors, leading to severe fragmentation of semantic information. Simultaneously, single visual feature extraction models struggle to analyze the deep connections between visual images and audio content, easily resulting in misjudgments or omissions in special cases such as voice-over narration or asynchronous audio-visual presentations, leading to significant localization errors. Furthermore, due to the lack of refined monitoring of audio stream pronunciation features and pauses, existing methods cannot smoothly fine-tune logical boundaries at the millisecond scale, resulting in logical lag or auditory discontinuities in the start and end positions of segmented segments, severely impacting the accuracy of automated audio and video resource decomposition. Therefore, a novel audio and video content segmentation and localization scheme is desired. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for audio and video content segmentation and localization, which can effectively solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for audio and video content segmentation and localization includes the following specific steps: Obtain the raw audio and video data to be processed and decompose it into video frame sequences and audio streams; The video frame sequence is subjected to feature parsing to generate a visual semantic feature vector sequence. The audio stream is processed and converted into text content through automatic speech recognition technology, and then mapped into a text semantic feature vector sequence. Align the visual semantic feature vector sequence with the text semantic feature vector sequence on the time axis, calculate the semantic consistency score within each time window using a sliding window, and extract the fundamental frequency feature and duration feature from the audio stream; The first preset threshold for triggering boundary warnings is dynamically adjusted based on the fundamental frequency feature and the duration feature, and logical boundary candidate points are determined by combining the semantic consistency score, visual slicing detection and scene classification constraints. The logical boundary candidate points are finely adjusted. The speech pause information in the audio stream is used to adjust the segmentation point to the silence range or the energy minimum point. Lip-sync detection is performed to finely adjust the segmentation point to the moment when the lips close, and the final segmentation moment is determined.

[0007] Furthermore, feature parsing is performed on the video frame sequence to generate a visual semantic feature vector sequence, specifically including: A visual feature extraction model based on deep residual networks is used to extract spatial features of video frames, and a spatial attention mechanism is used to guide the network to focus on key areas in the image to generate spatial feature vectors. The Lucas-Cannard pyramid algorithm is used to extract optical flow features from video frames. Based on the overall distribution characteristics of the motion vector field, the dynamic semantics of the shot are identified, and motion feature vectors are generated. After L2 normalization of the spatial feature vector and the motion feature vector respectively, a weighted sum is performed and fused to generate the visual semantic feature vector sequence.

[0008] Furthermore, the audio stream is processed and converted into text content using automatic speech recognition technology, then mapped into a sequence of text semantic feature vectors, specifically including: An acoustic model based on a long short-term memory network and a language model based on a recurrent neural network are used to decode the audio stream into text content through a weighted finite state converter and a beam search algorithm. The text content is mapped into a sequence of text semantic feature vectors using a pre-trained language model based on a converter architecture; For words that are not recognized or are misrecognized during speech recognition, the contextual semantic features around the corresponding audio segment are extracted, semantically similar words are searched in an external knowledge base, and the vector representations corresponding to the searched words are used to fill or replace them to ensure the integrity of the semantic feature sequence.

[0009] Furthermore, the fundamental frequency and duration features of the audio stream are extracted, and the first preset threshold for triggering boundary warnings is dynamically adjusted based on the fundamental frequency and duration features. Specifically, this includes: Fundamental frequency features are extracted from the audio stream using autocorrelation or cepstral methods, and duration features are obtained by measuring the duration of syllables or words. When the change in the fundamental frequency feature is detected to exceed a preset pitch change threshold, or when the silence interval in the duration feature is detected to exceed a preset speech rate slowdown threshold, the first preset threshold is multiplied by a reduction coefficient ranging from 0.5 to 0.9 to lower the judgment threshold.

[0010] Furthermore, the visual semantic feature vector sequence and the text semantic feature vector sequence are aligned on the time axis, and the semantic consistency score within each time window is calculated using a sliding window, specifically including: A linear interpolation algorithm is used to unify the sampling frequencies of the visual semantic feature vector sequence and the text semantic feature vector sequence to a preset target frequency; Set a sliding window with a duration of 3 to 5 seconds, and slide it from the beginning of the timeline with a preset step size; At each sliding window position, the cosine similarity of all audiovisual feature pairs within the window is calculated, and the calculated score is used as the semantic consistency score at the center time of that window, thus obtaining a score sequence that changes over time.

[0011] Furthermore, by combining semantic consistency scores, visual slicing detection, and scene classification constraints, candidate points for logical boundaries are determined, specifically including: Calculate the first-order difference of the semantic consistency score sequence. When the first-order difference shows a continuous decreasing trend within a predetermined number of consecutive sampling points, and the absolute value of the semantic consistency score is lower than the dynamically adjusted first preset threshold, a logical boundary warning is triggered. In the warning state, the intersection of the brightness histograms between adjacent keyframes is calculated. When the overlap of the brightness histogram intersection is lower than the second preset threshold, it is confirmed that the shot has switched. When the logical boundary warning is triggered and the camera switch is confirmed, that moment is marked as a logical boundary candidate point.

[0012] Furthermore, by combining semantic consistency scores, visual slicing detection, and scene classification constraints, the determination of logical boundary candidate points also includes: When the logical boundary warning and camera switching conditions are not met simultaneously, obtain scene label change information near the warning time; The weighted confidence score is calculated using the weighted confidence score formula based on the semantic gradient warning indicator variable, histogram difference, and scene label change indicator variable. When the weighted confidence level exceeds the preset confidence level limit, that moment is marked as a logical boundary candidate point; Wherein, the histogram difference is equal to 1 minus the brightness histogram intersection overlap.

[0013] Furthermore, the segmentation point is adjusted to a silent region or an energy minimum point using speech pause information in the audio stream, specifically including: A speech activity detection algorithm is used to monitor the short-time energy, zero-crossing rate, and spectral entropy of the audio stream to distinguish between speech intervals and silence intervals; When the logical boundary candidate point is located within the speech interval, the silent interval within a preset time range is searched backward from the candidate point. If a silent interval with a duration greater than the predetermined silence threshold is found within the search range, the final segmentation position is set at the start time of that silent interval. If no silent interval that meets the duration requirement is found, but multiple short silent segments exist, the local minimum point of audio energy within the search range is calculated, and this point is used as the final segmentation position.

[0014] Furthermore, lip-shape synchronization detection is performed, and the segmentation point is fine-tuned to the moment the lips close, specifically including: For close-up portrait videos, detect the facial area of ​​the person in the video frame and locate the key points of the lips; The opening and closing state of the lips is determined by calculating the vertical distance between key points of the upper and lower lips; Near the fine-tuned segmentation position, the opening and closing state of the lips is analyzed frame by frame. If the current segmentation position is in the open mouth state, the nearest closing mouth moment is searched forward or backward, and the closing moment is determined as the final segmentation position.

[0015] An audio / video content segmentation and localization system, comprising: The data acquisition and decomposition module is used to acquire the raw audio and video data to be processed and decompose it into video frame sequences and audio streams. The visual semantic extraction module is used to perform feature parsing on the video frame sequence and generate a visual semantic feature vector sequence. The audio-text semantic extraction module is used to process the audio stream, convert it into text content through automatic speech recognition technology, and then map it into a sequence of text semantic feature vectors. The semantic consistency calculation module is used to align the visual semantic feature vector sequence with the text semantic feature vector sequence on the time axis, and calculate the semantic consistency score within each time window through a sliding window. A dynamic threshold adjustment module is used to extract the fundamental frequency features and duration features in the audio stream, and dynamically adjust the first preset threshold for triggering boundary warnings based on the fundamental frequency features and duration features. The candidate point determination module is used to determine logical boundary candidate points by combining the semantic consistency score, visual slicing detection and scene classification constraints. The segmentation point fine-tuning module is used to fine-tune the candidate logical boundary points. It uses the speech pause information in the audio stream to adjust the segmentation points to the silence range or the energy minimum point, and performs lip-sync detection to fine-tune the segmentation points to the moment when the lips close, thus determining the final segmentation moment.

[0016] In summary, this application includes at least one of the following beneficial technical effects: 1. This invention overcomes the limitations of previous methods that relied solely on visual pixel differences or fixed-duration segmentation through multimodal semantic fusion technology. By deeply aligning visual image features with audio text semantics along the timeline, this solution can accurately identify content-level topic transition points. This ensures that each segmented fragment is a logical unit with independent meaning, avoiding the drawbacks of traditional methods where speech is forcibly interrupted or actions are split into two segments. This method achieves an extremely high level of semantic integrity in video processing with strong logical requirements, such as educational teaching and meeting recording.

[0017] 2. This invention introduces a fine-tuning mechanism combined with speech activity detection. After identifying large-scale logical boundaries, the segmentation points are fine-tuned to the nearest silent zone through refined analysis of audio energy distribution. This processing method greatly improves the user's listening experience when watching segmented segments, eliminating potential pops or punctuation breaks at the segmentation points. The fine-tuning accuracy reaches the millisecond level, allowing automatically disassembled video segments to be directly used for secondary dissemination and knowledge localization without the need for manual secondary editing.

[0018] 3. This invention achieves efficient processing of massive audio and video resources through the optimization of deep learning models and the application of parallel computing architecture. The automatically generated semantic tagging system not only summarizes the core content of segments but also significantly improves the speed and accuracy of retrieval through vector space indexing. Compared with traditional manual point segmentation, this method significantly improves efficiency, and the retrieval recall rate is significantly increased compared with single keyword matching technology.

[0019] 4. This invention employs a complementary decision logic based on both audiovisual and visual modalities, maintaining a stable recognition rate even in complex editing scenarios such as voice-over narration, asynchronous audio and video, and extremely smooth transitions. When visual features cause interference, the consistency score of the text semantics can play a corrective role; conversely, when the audio quality is poor, the visual cut-off decision provides auxiliary support. This complementary mechanism makes this method exhibit extremely high reliability in audio and video processing across various fields. Attached Figure Description

[0020] Figure 1 A schematic diagram of the overall technical solution for audio and video content segmentation and localization methods; Figure 2 A schematic diagram illustrating the core principles of audiovisual bimodal semantic consistency analysis and logical boundary determination; Figure 3 A flowchart illustrating the logic for calculating the temporal alignment and consistency between video semantic features and text semantic features; Figure 4 A schematic diagram of the multi-level interaction relationship and data flow for determining logical boundary candidate points and fine-tuning speech pauses; Figure 5 This is a flowchart illustrating the logic of semantic tag generation and precise keyword localization based on semantic feature clustering. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with the appendix. Figure 1-5 The present invention will be further described in detail with reference to specific embodiments.

[0022] The first aspect concerns the audio and video content segmentation and localization method disclosed in this application, which is implemented according to the following steps: The first step, S1, involves acquiring the raw audio and video data to be processed and decomposing it into video frame sequences and audio streams that can be used independently by the subsequent feature extraction module. This process needs to address two key issues: first, compatibility with various common multimedia container formats; and second, ensuring strict synchronization of the decomposed audiovisual data on the timeline. To address these issues, this embodiment employs a preprocessing workflow based on multi-threaded demultiplexing and a circular buffer, specifically implemented through the following sub-steps.

[0023] Step S101: Decapsulate and locate the audio and video streams; first, obtain the original audio and video data to be processed, and then parse the input file format using multimedia decapsulation technology. This process is compatible with multiple container formats, such as MP4, MKV, AVI, MOV, FLV, etc.

[0024] The physical offset addresses of video and audio streams within a multimedia file are precisely located by parsing the header index information. Taking the MPEG-4 encapsulation format as an example, it is necessary to parse the moov atom and its internal sub-tables such as trak and stbl to read the position, size, and timestamp information of each data packet.

[0025] In step S102, a multi-threaded demultiplexer is used to push data to the circular buffer. During the parsing process, two independent threads are started, each responsible for the demultiplexing of the video and audio streams. Each thread reads data packets from the physical offset address located in step S101 and pushes them to the corresponding circular buffer. Video data packets are pushed to the video decoding buffer, and audio data packets are pushed to the audio sampling buffer.

[0026] The size of each circular buffer is dynamically set according to the data throughput. The initial size is 2 megabytes. When the buffer occupancy rate reaches 80%, new memory blocks are automatically requested for expansion.

[0027] The two threads ensure synchronization by comparing timestamps in the data packets: if the video packet's timestamp is earlier than the audio packet's, the video thread waits for the audio thread to process it, and vice versa. This multi-threaded architecture ensures strict synchronization of audiovisual data during the parsing phase.

[0028] Step S103: Decode to generate a video frame sequence and an audio stream; decode the data in the video decoding buffer and the audio sampling buffer respectively. The video decoding uses a standard decoder, such as an H.264 or H.265 decoder, to restore the compressed video data packets into a continuous sequence of original image frames. The pixel format of the image frames can be RGB or YUV.

[0029] Audio decoding employs a corresponding audio decoder, such as an AAC or MP3 decoder, to restore the compressed audio data packets to an audio stream in pulse code modulation format. During decoding, the sampling frequency and bit depth of the audio stream are checked, requiring a sampling frequency of no less than 44.1kHz or 48kHz and a bit depth of no less than 16bit; simultaneously, the video frame rate is checked, requiring a minimum of 25fps or 30fps.

[0030] If the input data does not meet these preset thresholds, automatic resampling or frame rate conversion will be performed to ensure the data precision during subsequent feature extraction.

[0031] In summary, step S1 completes the preprocessing of audio and video files in any encapsulation format. Regardless of the input file format, it can be reliably decomposed into temporally synchronized and independent video frame sequences and audio streams. This preprocessing process provides standardized and refined input data for subsequent multimodal feature extraction.

[0032] The next step, S2, involves deep feature analysis of the video frame sequence generated in step S1 to produce a sequence of visual semantic feature vectors that characterize the semantics and dynamic information of the scene. These features will be used for subsequent cross-modal alignment with audio features. This is achieved through the following sub-steps.

[0033] Step S201: Construct and train a visual feature extraction model; This embodiment uses a pre-defined visual feature extraction model to process video frames. This model is based on a deep residual network structure, specifically ResNet-50 or ResNet-101. The main structure of the model includes multiple convolutional layers, pooling layers, residual blocks, and a global average pooling layer located at the end of the network.

[0034] The training process for this model is as follows.

[0035] The training dataset was constructed by collecting a large number of labeled video frame images. The annotation information for each frame image includes the categories of key objects appearing in the scene, such as people, text, and icons, as well as the background scene categories, such as classroom, outdoors, and office. The label for each frame image is a multi-class vector that indicates the presence of both objects and scenes.

[0036] The loss function employs multi-task cross-entropy loss. For each training sample, the model outputs two prediction branches: the first branch predicts the key object category, and the second branch predicts the background scene category. The total loss is the sum of the cross-entropy losses of the two branches. The formula is expressed as: ; Where C1 is the total number of key object categories, y 1c It is the true label of the c-th object category, with a value of 0 or 1, p 1c y is the probability of the c-th object category predicted by the model; C2 is the total number of background scene categories. 2d p is the true label for the d-th scene category. 2d It is the probability of the d-th scene category predicted by the model.

[0037] The training objective of the model is to minimize the value of the aforementioned loss function L, thereby enabling the model to accurately extract high-dimensional feature vectors containing spatial semantic information from video frames. After training, the model's classification output layer is removed, and the feature vector output by the global average pooling layer is taken as the visual semantic feature vector for that frame.

[0038] Step S202: Extract keyframes at an adaptive frequency; to improve processing efficiency, the system does not perform complete feature extraction on every frame. During periods of smooth video content change, the system extracts keyframes at a preset fixed frequency, such as 1 to 3 keyframes per second.

[0039] The system monitors pixel differences between adjacent frames in real time to determine whether the scene has changed drastically. Specifically, it calculates the difference in brightness histograms between the current frame and the previous frame, using the histogram intersection distance as the formula.

[0040] If the difference exceeds a preset switching threshold, such as 0.6, it is determined to be a drastic scene change. In addition, the system also calculates the peak signal-to-noise ratio between adjacent frames. If the peak signal-to-noise ratio is lower than 20 dB, it is also determined to be a drastic scene change.

[0041] When a drastic scene change is detected, the system automatically switches from keyframe mode to frame-by-frame processing mode, extracting features from each consecutive frame to ensure that no important semantic information is lost.

[0042] Step S203: Apply spatial attention mechanism to generate visual semantic feature vectors; when the model processes each frame of image, the internal spatial attention module guides the network to focus on key areas in the image. This embodiment uses a specific and implementable spatial attention mechanism, the calculation process of which is as follows.

[0043] First, the model performs global max pooling and global average pooling on the input feature map, respectively, to obtain two different feature descriptors. Then, these two descriptors are concatenated along the channel dimension to form a new feature map.

[0044] Next, the feature map is convolved by a 7x7 convolutional layer, and then activated by a sigmoid function to generate a two-dimensional attention weight map with the same spatial dimensions as the original feature map. The value at each position in this weight map is between 0 and 1, representing the importance of that position.

[0045] Finally, the original feature map is multiplied element-wise with the generated attention weight map to obtain the weighted feature map. After attention weighting, the model outputs a fixed-dimensional feature vector through a global average pooling layer. The dimension of this vector is preset, such as 512 or 1024. This vector is the visual semantic feature vector of that frame.

[0046] Step S204: Extract optical flow features and identify dynamic semantics; In addition to the spatial features of a single frame, the system also extracts optical flow features from the video to capture the dynamic information of the shot and objects. The optical flow extraction algorithm adopts the Lucas-Carnard pyramid algorithm. This algorithm starts from the highest resolution layer of the image pyramid, calculates the pixel displacement layer by layer, and finally obtains the motion vector field of each pixel, including horizontal and vertical displacement.

[0047] The system identifies the dynamic semantics of the shot based on the overall distribution characteristics of the motion vector field. The specific judgment logic is as follows.

[0048] If the motion vectors of most pixels in the scene point towards the center of the image, and the vector length gradually decreases from the edge to the center, then it is determined that the camera is zooming out.

[0049] If the motion vectors of most pixels point from the center of the image to the edge, and the vector length gradually increases from the center to the edge, then it is determined that the camera is zooming in.

[0050] If the motion vectors are all pointing to the left or right, it is determined to be a camera pan or tilt.

[0051] If the direction of the motion vector shows a rotation pattern centered on a certain point, it is determined to be a camera rotation.

[0052] For complex motions that cannot be classified into the above types, the system marks them as general object motions and does not use them as camera motion semantics.

[0053] Step S205: Weighted fusion of spatial and motion features; Before fusion, the system first performs L2 normalization on both the spatial and motion feature vectors, ensuring that the magnitude of both vectors is 1. The normalization formula is: For any vector X, the normalization result is... ,in Let X denote the Euclidean norm of vector X.

[0054] The system then performs a weighted summation of the normalized spatial feature vector and the motion feature vector to form the final visual feature vector. The fusion formula is: ; Among them, F spatial It is the normalized spatial eigenvector, F motion It is the normalized motion feature vector, w s It is the weighting coefficient of spatial features, w m These are the weighting coefficients for motion features. The values ​​of the two weighting coefficients satisfy a preset ratio, for example, setting w... s =0.7, w m =0.3. The fused feature vector F final This will serve as a complete visual representation of the video stream, used for subsequent cross-modal alignment with audio features.

[0055] In summary, step S2 completes the extraction of visual features from the video frame sequence. First, the trained deep residual network can extract semantically rich spatial features from a single frame image; second, the spatial attention mechanism ensures that the model focuses on key content in the scene; third, the optical flow algorithm supplements the dynamic information of camera movement; finally, weighted fusion combines static and dynamic features into a unified visual semantic feature vector. This vector sequence contains both what is seen and how the scene moves, providing high-quality visual input for subsequent semantic consistency calculations with audio information.

[0056] The next step, S3, processes the audio stream generated in step S1. First, the audio signal is converted into text content using automatic speech recognition technology. Then, this text content is mapped into a high-dimensional sequence of text semantic feature vectors. This feature vector sequence will be used for subsequent cross-modal alignment with visual features. This is implemented through the following sub-steps.

[0057] Step S301: Construct and train the acoustic model; the automatic speech recognition process includes joint decoding of the acoustic model and the language model. The acoustic model is used to map the feature sequence of the audio signal to a phoneme sequence. In this embodiment, the acoustic model is based on a Long Short-Term Memory (LSTM) network structure, but a gated recurrent unit (GRU) structure can also be used as an alternative. The model employs a deep bidirectional LSTM architecture, containing three bidirectional LSTM layers, with each hidden layer containing 512 neurons.

[0058] The training process for the acoustic model is as follows.

[0059] The training dataset is constructed as follows: a large number of labeled audio segments are collected, each segment corresponding to a phoneme-level annotation sequence. The audio data needs to cover multiple speakers, speech rates, and environmental noise conditions. For each audio data point, Mel-frequency cepstral coefficients are first extracted as input, with a feature dimension of 40 and a frame shift of 10 milliseconds. The corresponding annotations are phoneme sequences, using standard phoneme sets such as the International Phonetic Alphabet or phoneme sets specific to the language.

[0060] The loss function is designed using connection-time classification loss. This loss function allows for length mismatches between the input feature sequence and the output phoneme sequence, automatically learning alignment paths. For a training sample with an input feature sequence of length T and an output phoneme sequence of length U, the connection-time classification loss function is defined as the negative logarithm of the sum of probabilities of all possible alignment paths. The formula is expressed as: ; Where x is the input audio feature sequence, T is the time step number of the sequence, and y is the actual phoneme sequence. Let represent the set of all aligned paths π that can be mapped to y by compressing duplicates and removing whitespace labels. The model predicts the alignment path π at time step t. t The probability of [the loss function]. The training objective is to minimize this loss function so that the model can accurately map audio feature sequences to phoneme sequences.

[0061] Step S302: Construct and train a language model; the language model is used to perform lexical and syntactic correction on the phoneme sequence output by the acoustic model, and to find the optimal text decoding result in the lexical space constructed by the weighted finite state converter through the beam search algorithm.

[0062] The language model in this embodiment adopts a recurrent neural network-based language model, specifically using a two-layer LSTM structure, with each hidden layer having a dimension of 1024, and the vocabulary size being the number of commonly used words, such as 50,000 words.

[0063] The training process of the language model is as follows.

[0064] The training dataset is constructed by collecting a large amount of plain text corpus relevant to the target task domain. For example, for educational videos, this involves collecting text data such as textbooks and handouts. The corpus needs to undergo word segmentation and noise reduction, and each training sample is a fixed-length word sequence.

[0065] The loss function uses cross-entropy loss. For each prediction position t, the model predicts the probability distribution of the current word based on the previous t-1 words. The loss function is the average of the prediction errors across all positions. The formula is expressed as: ; Where N is the length of the word sequence in the training samples, w t It is the real word at position t in the sequence. The model predicts w given all preceding words. t The conditional probability of L. The training objective is to minimize L. LM This enables the model to accurately predict the next word, thus providing language-level constraints during the decoding process.

[0066] During the decoding phase, the acoustic model and the language model are jointly decoded using a weighted finite-state converter. The weighted finite-state converter compiles the phoneme probabilities output by the acoustic model, the word probabilities output by the language model, and the pronunciation dictionary mapping into a unified search graph.

[0067] The beam search algorithm maintains multiple candidate paths on the search graph, with the beam width set to 10, and finally outputs the word sequence with the highest score as the recognized text.

[0068] Step S303: Generate a sequence of text semantic feature vectors using a pre-trained language model; after obtaining the text content through speech recognition, the system uses a pre-trained language model based on a converter architecture to map the text content into a sequence of text semantic feature vectors. This embodiment can use either the BERT model or the RoBERTa model.

[0069] The training process of this pre-trained language model is as follows.

[0070] The training dataset is constructed by collecting a large-scale general text corpus, such as Wikipedia, books, and news articles. Each training sample is a continuous sentence or paragraph.

[0071] The pre-training tasks consist of two parts. The first task is masked language modeling, where 15% of the words in the input text are randomly replaced with masked tokens, and the model needs to predict the masked words based on the context. The second task is next sentence prediction, where the model needs to determine whether two sentences have a coherent relationship in the original text.

[0072] The loss function for the masked language model is the cross-entropy loss, which only calculates the prediction error at the masked position. The formula is expressed as: ; in, M is the set of masked locations, and M is the size of this set. It is the real word at position i. It is the probability that the model predicts for the word given a context.

[0073] The loss function for predicting the next sentence is a binary cross-entropy loss, which determines whether the two sentences are coherent. The total loss is the sum of the two losses. The training objective is to minimize the total loss so that the model can capture contextual dependencies of up to 512 characters.

[0074] In practical use, the parameters of the pre-trained model are loaded, and the text content output in step S302 is input into the model in sentences or fixed-length windows. The feature vectors corresponding to each word output by the last layer of the model are extracted, or the pooling vectors of the entire sequence are taken to form a sequence of text semantic feature vectors.

[0075] Step S304: Processing unrecognized and uncommon words; For proper nouns or uncommon words that cannot be correctly recognized during speech recognition, the system initiates an external knowledge base retrieval module for post-processing. The external knowledge base can be Wikipedia, a domain-specific dictionary, or a general knowledge graph.

[0076] The processing procedure is as follows: First, for text locations that fail to be identified, the system extracts the contextual semantic features around the audio segment corresponding to that location.

[0077] Then, a search is conducted in an external knowledge base for terms that are semantically similar to the context. The search method involves encoding the description text of each term in the knowledge base into a vector using the same pre-trained language model, calculating the cosine similarity with the context feature vector, and selecting the top 3 terms with the highest similarity as candidates.

[0078] Finally, the vector representations corresponding to the selected terms are used for filling. The filling method is as follows: if no word was originally identified at the position, the vector of that term is directly inserted; if an incorrect word is identified, the vector of the correct term is used to replace the incorrect vector.

[0079] When a suitable term cannot be found, the system interpolates and fills the gap based on the average of the context word vectors. The specific interpolation formula is: the filled vector is equal to the average of the two adjacent valid vectors, or equal to the average of the entire sentence vector. This process ensures the continuity and integrity of the semantic feature sequence.

[0080] Step S3 completes the transformation from the original audio stream to text semantic feature vectors. First, the jointly trained acoustic and language models achieve high-precision speech recognition through a weighted finite-state converter and a beam search algorithm. Second, the pre-trained converter architecture language model transforms the text content into feature vectors rich in contextual semantics. Finally, the external knowledge base retrieval module addresses the identification of rare words and proper nouns, ensuring the integrity of the semantic features. This text semantic feature vector sequence will then undergo subsequent cross-modal alignment and consistency calculations with the visual feature vectors generated in step S2.

[0081] For step S4, the visual semantic feature vector sequence generated in step S2 is aligned with the text semantic feature vector sequence generated in step S3 on the time axis, and a semantic consistency score is calculated within each time window using a sliding window. This score will be used to subsequently determine logical boundary candidate points. This is implemented through the following sub-steps.

[0082] Step S401: Unify the sampling frequency of the two modalities; because the video frame rate and the text output frequency are inconsistent (the video frame rate is usually 30 frames per second, while the text output frequency is determined by the speech rate, generally 3 to 5 words per second), the sampling points of the two on the time axis cannot be directly correlated. This embodiment uses a linear interpolation algorithm to unify the sampling frequency of the two modalities to a preset target frequency, for example, set to 10 Hz, that is, generating one feature point every 100 milliseconds.

[0083] The specific method of linear interpolation is as follows: For a sequence of visual feature vectors, find the feature vectors of two adjacent original video frames for each target time point, and perform a weighted average according to the reciprocal of the time distance to obtain the visual feature value of the target time point.

[0084] The same process is applied to the text feature vector sequence. The aligned feature sequence is stored in a synchronized memory block, with each time point corresponding to a visual feature vector and a text feature vector, forming an audiovisual feature pair.

[0085] Step S402: Set up a sliding window and initialize parameters; the system sets up a sliding window with a preset time length ranging from 3 to 5 seconds. The sliding step size of the window also uses a preset value, for example, set to 500 milliseconds. Starting from the beginning of the timeline, the system moves the window forward by one step each time, performing subsequent calculations within each window position.

[0086] Step S403: Calculate the semantic consistency score within the window; within each sliding window, the system calculates the semantic consistency score for all audiovisual feature pairs within that window. The score is calculated using cosine similarity, and the specific formula is as follows: ; Where S represents the semantic consistency score of the current sliding window, ranging from 0 to 1. The closer the score is to 1, the more consistent the audiovisual semantics; the closer it is to 0, the less consistent they are. V i T represents the visual feature vector of the i-th sampling point within the sliding window. i This represents the text feature vector corresponding to the sampling point, where n is the total number of sampling points within the sliding window. (Symbol V) i ·T i Represents the dot product of two vectors. This represents the square root of the sum of the squares of the magnitudes of all visual feature vectors within the window. Similarly.

[0087] The system uses the calculated score S as the semantic consistency score at the center of the window, thus obtaining a score sequence that changes over time.

[0088] Step S404: Extract intonation and speech rate features from the audio; while calculating the semantic consistency score, the system also extracts fundamental frequency features and duration features from the audio stream to identify changes in the speaker's intonation and speech rate. Fundamental frequency features reflect the pitch of the sound and are extracted from the audio signal using autocorrelation or cepstral methods, denoted as F0. Duration features are obtained by measuring the duration of each syllable or word.

[0089] When the system detects a significant decrease or increase in the fundamental frequency characteristic, and the magnitude of the change exceeds the preset tone change threshold, it is determined to be a tone shift. When the system detects that the duration of a syllable or word is significantly longer than the speaker's average duration, or that the silence interval between adjacent words has significantly increased, it is determined to be a significant slowdown in speech rate.

[0090] Step S405: Dynamically adjust the judgment threshold of semantic consistency score; In order to adapt to the natural switching rules of content logic, the system will dynamically adjust the first preset threshold used to trigger boundary warning in the subsequent step S5 according to the recognition result of step S404. The original reference value of the threshold is 0.4, and the specific definition is shown in step S502.

[0091] The specific adjustment method is as follows: when a change in tone or a significant slowdown in speech rate is detected, the system multiplies the current first preset threshold by a reduction coefficient, for example, a reduction coefficient of 0.8, thereby lowering the judgment threshold. This means that at points of tone or speech rate change, even if the semantic consistency score does not drop to the original threshold level, it may still be judged as a potential logical boundary. The specific value of the reduction coefficient can be preset according to the application scenario, ranging from 0.5 to 0.9.

[0092] In summary, step S4 achieves precise alignment of visual and textual features along the time axis, generating a score sequence reflecting audiovisual semantic consistency. Linear interpolation addresses the inconsistency in sampling frequencies between the two modalities, while sliding window cosine similarity calculation provides a quantified semantic consistency index. The introduction of intonation and speech rate features allows for adaptive adjustment of the decision threshold, improving the sensitivity and rationality of boundary detection. This score sequence serves as input for step S5, monitoring gradient changes and triggering logical boundary warnings.

[0093] For step S5, based on the semantic consistency score sequence generated in step S4, and combined with visual slicing detection and scene classification constraints, possible logical boundary locations are determined as candidate points. This is specifically achieved through the following sub-steps.

[0094] Step S501: Calculate the first-order difference of the semantic consistency score sequence; the system first obtains the time-varying semantic consistency score sequence generated in step S403. To identify abrupt changes in the score curve, the system calculates the first-order difference of this sequence. For the score S at the k-th time point... k Its first difference value This value reflects the trend and rate of change of the score at the current moment, i.e., the slope of the curve.

[0095] Step S502: Monitor gradient changes and trigger logical boundary warnings; the system continuously monitors changes in the first-order difference value. A logical boundary warning is triggered when both of the following conditions are met simultaneously.

[0096] The first condition is that the first-order difference value exhibits a continuously decreasing trend over a predetermined number of consecutive sampling points. This predetermined number can be set to 3 to 5 sampling points. A continuous decrease means that the first-order difference value at each sampling point is smaller than the first-order difference value at the previous sampling point.

[0097] The second condition is that the absolute value of the semantic consistency score is lower than a first preset threshold. This threshold is a dynamically adjusted value in step S405, with an original reference value of 0.4. If step S405 does not trigger the adjustment, 0.4 is used directly as the threshold.

[0098] When both conditions are met simultaneously, the system marks that moment as an early warning state.

[0099] Step S503: Construct and train a scene recognition model; This embodiment introduces a scene recognition model to determine whether a preset scene change occurs in the video frame. This model is based on a convolutional neural network structure, specifically a lightweight network such as MobileNet or EfficientNet-Lite. The model input is a video frame image, and the output is a scene category label. Predefined scene categories include common scenes such as classroom, office, outdoors, indoor living room, street, and inside a vehicle.

[0100] The training process of the scene recognition model is as follows.

[0101] The training dataset is constructed by collecting a large number of image data with labeled scene categories, with each image corresponding to a scene category label. The data needs to cover all preset scene categories, with each category containing at least 5000 images, and the images come from different shooting angles, lighting conditions, and video sources to ensure the model's generalization ability.

[0102] The loss function employs multi-class cross-entropy loss. For each training sample, the model outputs a probability distribution vector, representing the probability that the input image belongs to each scene category. The loss function formula is: ; Where C3 is the total number of preset scene categories, y c It is the true label of the c-th scene category, with a value of 0 or 1, and all y c The sum of p is 1. c It is the probability of the c-th scene category predicted by the model.

[0103] The training objective of the model is to make the above loss function L scene Minimize the value of , so that the model can accurately identify the scene category to which any input video frame belongs.

[0104] Step S504: Perform visual cut detection; in the warning state, the system combines the histogram differences of the video frames to determine whether a shot change has occurred. Specifically, it calculates the intersection of the luminance histograms between two adjacent keyframes. The luminance histogram is a distribution obtained by dividing the image luminance values ​​into several levels and counting the number of pixels in each level. The formula for calculating the histogram intersection is: for two histograms H1 and H2, the intersection is... , where j iterates through all brightness levels. This intersection value reflects the similarity of brightness distribution between two frames.

[0105] If the overlap of the histogram intersection is lower than a second preset threshold, such as 0.6, then a shot change is confirmed. The lower the overlap, the greater the difference in brightness distribution between the two frames, and the more likely it is to be a shot change point.

[0106] Step S505: Obtain scene label change information; the system uses the scene recognition model trained in step S503 to classify the keyframes near the current warning time. Specifically, it extracts several frames before and after the warning time, for example, 5 frames before and after, and inputs them into the scene recognition model to obtain scene labels. The scene labels of the frames before the warning time are compared with the scene labels of the frames after the warning time. If they are inconsistent, a scene change is determined. For example, a change from an interior classroom scene to an outdoor playground scene.

[0107] Step S506: Calculate the weighted confidence score and determine the logical boundary candidate points; the system comprehensively determines whether a point is a logical boundary candidate point based on the above multiple conditions. The determination logic is as follows.

[0108] First, a basic judgment condition is set: the semantic gradient and the visual camera switching condition are simultaneously satisfied, that is, step S502 triggers the warning and step S504 confirms the camera switching. At this time, this moment is directly marked as a candidate point of the logical boundary.

[0109] Secondly, for cases where both of the above conditions are not met simultaneously but a scene label change is detected in step S505, the system introduces a weighted confidence mechanism. Weighted confidence C conf The calculation formula is: ; Among them, I grad It is a semantic gradient warning indicator variable, which takes a value of 1 when a warning is triggered, and 0 otherwise; I hist This is the overlap of histogram intersections, with values ​​ranging from 0 to 1; 1-I hist Indicates the histogram difference; I scene This is a scene label change indicator variable; it takes a value of 1 when a change occurs, and 0 otherwise. grad w hist w scene These are the weight coefficients for the three conditions, and their sum is 1. An example weight setting is: w grad =0.5, w hist =0.3, w scene =0.2.

[0110] When the weighted confidence level C conf When the confidence level exceeds a preset threshold, for example, a threshold of 0.7, that moment is also marked as a candidate logical boundary point.

[0111] In summary, step S5 completes the transformation from semantic consistency score sequence to logical boundary candidate points. First, a continuous decrease in the score gradient is monitored using first-order difference, and an alert is triggered based on the absolute value of the score. Second, visual cut-off detection is performed using the intersection of brightness histograms. Third, a trained scene recognition model is introduced to obtain scene label change information. Finally, a weighted confidence mechanism integrates multiple conditions, and candidate points are marked when both the alert and cut-off conditions are met, or when the confidence level is sufficiently high. This composite judgment logic effectively avoids misjudgment based on a single feature and improves the accuracy of logical boundary localization.

[0112] Finally, in step S6, the logical boundary candidate points generated in step S5 are finely adjusted. By utilizing the speech pause information in the audio stream, the segmentation points are adjusted to the most comfortable listening position, ultimately determining the actual segmentation time for each segment. This is achieved through the following sub-steps.

[0113] Step S601: Perform speech activity detection; the system uses a speech activity detection algorithm to analyze the audio stream in real time to distinguish between speech intervals and silence intervals. This algorithm simultaneously monitors three characteristic parameters of the audio stream: short-time energy, zero-crossing rate, and spectral entropy.

[0114] Short-time energy reflects the amplitude of an audio signal within a short time window; the energy in the speech region is typically significantly higher than that in the silence region. The system divides the audio stream into fixed-length frames, each lasting 25 milliseconds with a frame shift of 10 milliseconds, and calculates the sum of the squares of the samples within each frame as the short-time energy.

[0115] Zero-crossing rate refers to the number of times a signal waveform crosses the zero level. The zero-crossing rate of a voice signal usually fluctuates within a certain range, while the zero-crossing rate of silence or noise may be higher or lower.

[0116] Spectral entropy measures the degree of disorder in the spectral distribution of an audio signal. Speech signals have lower spectral entropy because their spectral energy is concentrated in a specific frequency band; silence or background noise has higher spectral entropy and a more uniform energy distribution.

[0117] The system combines the three features mentioned above into a feature vector, which is then input into a pre-trained binary classifier. The classifier outputs whether the current frame belongs to a speech frame or a silence frame. Multiple consecutive silence frames form a silence interval, and multiple consecutive speech frames form a speech interval.

[0118] Step S602: Determine whether the logical boundary candidate point falls within the speech interval; for each logical boundary candidate point marked in step S506, the system checks the audio interval type at that moment. If the moment is within a silent interval, no adjustment is needed, and the candidate point is directly used as the final segmentation position.

[0119] If the moment falls within a speech interval, meaning the speaker is currently in the process of speaking, then the fine-tuning logic is triggered, and step S603 is entered.

[0120] Step S603: Search backward for the nearest silent interval; the system starts from the logical boundary candidate point and searches backward for silent intervals within a preset duration. The preset duration of the search range can be set to 2 seconds. Within the search range, the system looks for silent intervals whose duration is greater than a preset silence threshold. The preset value of the silence threshold can be set to 200 milliseconds, meaning that a silent interval with a continuous duration of more than 200 milliseconds is considered a valid cut-off point.

[0121] Step S604: Determine the final segmentation position; the final segmentation position is determined based on the search results, and is handled in three cases.

[0122] The first scenario: A silent interval that meets the search criteria is found within the search range. In this case, the system sets the final split position at the beginning of that silent interval. This ensures that the split occurs at a natural pause, avoiding interruption in the middle of the statement.

[0123] The second scenario: No silent intervals meeting the duration requirement are found within the search range, but multiple short silent segments exist. In this case, the system calculates the local minimum point of audio energy within the search range, i.e., the point in the energy curve that is lower than both its immediate and adjacent points, and uses this point as the final segmentation location. This process ensures that even without complete silence, segmentation can still be performed at relatively quiet points.

[0124] The third scenario: There are neither suitable silent intervals nor obvious energy minimum points within the search range. In this case, the system retains the original logical boundary candidate points as the final segmentation positions without making any adjustments.

[0125] Step S605: Perform lip-sync detection for fine-tuning; for close-up portrait videos, the system additionally performs lip-sync detection to further optimize the segmentation position. The specific steps are as follows.

[0126] First, the system detects the facial region of a person in the video frame and locates key points of the lips, such as the outline points of the upper and lower lips. The opening and closing state of the lips is determined by calculating the vertical distance between these key points. When the vertical distance is greater than a preset opening threshold, the mouth is determined to be open; when it is less than a closing threshold, the mouth is determined to be closed.

[0127] The system analyzes the lip opening and closing state frame by frame near the candidate segmentation position determined in step S604, such as within a range of 500 milliseconds before and after. If the candidate segmentation position is in the middle of an open mouth state, the system automatically fine-tunes forward or backward to find the nearest closed mouth moment as the final segmentation position. The fine-tuning direction prioritizes searching backward; if no closed state is found within 100 milliseconds backward, the system searches forward.

[0128] This lip-sync detection ensures that the cutting position is not in the middle of the character's open mouth, thus avoiding mechanical cutting of the complete sentence and achieving millisecond-level fine-grained correction of the segmentation timing.

[0129] In summary, step S6 completed the fine-tuning of the candidate points for the logical boundary. First, the speech activity detection algorithm accurately distinguishes between speech intervals and silence intervals using three features: short-time energy, zero-crossing rate, and spectral entropy. Second, for candidate points falling into speech intervals, the algorithm searches backward for the nearest silence interval or energy minimum point for adjustment. Finally, in close-up portrait scenes, lip-sync detection further fine-tunes the cutting point to the moment the lips close. This series of fine-tuning operations makes the final segmentation position more natural and coherent both auditorily and visually, avoiding the discomfort caused by cutting in the middle of a sentence or when the mouth is open, as in traditional methods.

[0130] After determining the final segmentation location, the method also includes automatically generating semantic labels for each segment.

[0131] The semantic tag generation process is as follows. First, the system obtains all text semantic feature vectors generated in step S303 within the segment, and uses the K-means clustering algorithm to divide these vectors into several clusters. The number of clusters K is adaptively set according to the segment length, for example, one cluster is set for every 30 seconds of text, with a minimum of one cluster and a maximum of five clusters. Each cluster corresponds to a potential topic.

[0132] Then, for each cluster, the system extracts the corresponding original text segment and calculates the term frequency-inverse document frequency (TF-IDF) value for each word using the TF-IDF algorithm. The TF-IDF formula is: term frequency multiplied by inverse document frequency, where term frequency is the number of times the word appears in the current segment divided by the total number of words in the segment, and inverse document frequency is the logarithm of the total number of documents in the entire corpus divided by the number of documents containing the word. The system selects the 3 to 5 words with the highest TF-IDF values ​​as the keywords for that cluster. Alternatively, the TextRank algorithm can be used to construct a word graph, calculate node weights based on word co-occurrence relationships, and select the words with the highest weights as keywords.

[0133] Finally, the system merges and deduplicates the keywords from each cluster, generating descriptive terms that summarize the theme of the segment as semantic tags. For example, in an educational video, if the text of a segment frequently contains "Newton's Second Law" and "acceleration," and the TF-IDF values ​​of these words are significantly higher than other words, the semantic tag "Physics Teaching - Newton's Second Law" is automatically generated. These tags are associated with the segment's timestamp information and visual feature fingerprint and stored in a vector database.

[0134] On the other hand, the audio and video content segmentation and positioning system disclosed in this application includes: The data acquisition and decomposition module is used to acquire the raw audio and video data to be processed and decompose it into video frame sequences and audio streams. The visual semantic extraction module is used to perform feature parsing on video frame sequences and generate a sequence of visual semantic feature vectors. The audio-text semantic extraction module is used to process audio streams, convert them into text content through automatic speech recognition technology, and then map them into a sequence of text semantic feature vectors. The semantic consistency calculation module is used to align the visual semantic feature vector sequence with the text semantic feature vector sequence on the time axis and calculate the semantic consistency score within each time window through a sliding window. The dynamic threshold adjustment module is used to extract the fundamental frequency features and duration features in the audio stream, and dynamically adjust the first preset threshold used to trigger the boundary warning based on the fundamental frequency features and duration features. The candidate point determination module is used to determine logical boundary candidate points by combining semantic consistency score, visual slicing detection and scene classification constraints; The segmentation point fine-tuning module is used to fine-tune the candidate points of the logical boundary. It uses the speech pause information in the audio stream to adjust the segmentation point to the silence range or the energy minimum point, and performs lip-sync detection to fine-tune the segmentation point to the moment when the lips close, thus determining the final segmentation moment.

[0135] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.

[0136] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for audio and video content segmentation and localization, characterized in that, Includes the following steps: Obtain the raw audio and video data to be processed and decompose it into video frame sequences and audio streams; The video frame sequence is subjected to feature parsing to generate a visual semantic feature vector sequence. The audio stream is processed and converted into text content through automatic speech recognition technology, and then mapped into a text semantic feature vector sequence. Align the visual semantic feature vector sequence with the text semantic feature vector sequence on the time axis, calculate the semantic consistency score within each time window using a sliding window, and extract the fundamental frequency feature and duration feature from the audio stream; The first preset threshold for triggering boundary warnings is dynamically adjusted based on the fundamental frequency feature and the duration feature, and logical boundary candidate points are determined by combining the semantic consistency score, visual slicing detection and scene classification constraints. The logical boundary candidate points are finely adjusted. The speech pause information in the audio stream is used to adjust the segmentation point to the silence range or the energy minimum point. Lip-sync detection is performed to finely adjust the segmentation point to the moment when the lips close, and the final segmentation moment is determined.

2. The method according to claim 1, characterized in that, The video frame sequence is subjected to feature parsing to generate a sequence of visual semantic feature vectors, specifically including: A visual feature extraction model based on deep residual networks is used to extract spatial features of video frames, and a spatial attention mechanism is used to guide the network to focus on key areas in the image to generate spatial feature vectors. The Lucas-Cannard pyramid algorithm is used to extract optical flow features from video frames. Based on the overall distribution characteristics of the motion vector field, the dynamic semantics of the shot are identified, and motion feature vectors are generated. After L2 normalization of the spatial feature vector and the motion feature vector respectively, a weighted sum is performed and fused to generate the visual semantic feature vector sequence.

3. The method according to claim 1, characterized in that, The audio stream is processed and converted into text content using automatic speech recognition technology, then mapped into a sequence of text semantic feature vectors, specifically including: An acoustic model based on a long short-term memory network and a language model based on a recurrent neural network are used to decode the audio stream into text content through a weighted finite state converter and a beam search algorithm. The text content is mapped into a sequence of text semantic feature vectors using a pre-trained language model based on a converter architecture; For words that are not recognized or are misrecognized during speech recognition, the contextual semantic features around the corresponding audio segment are extracted, semantically similar words are searched in an external knowledge base, and the vector representations corresponding to the searched words are used to fill or replace them to ensure the integrity of the semantic feature sequence.

4. The method according to claim 1, characterized in that, Extracting fundamental frequency and duration features from the audio stream, and dynamically adjusting the first preset threshold for triggering boundary warnings based on the fundamental frequency and duration features, specifically including: Fundamental frequency features are extracted from the audio stream using autocorrelation or cepstral methods, and duration features are obtained by measuring the duration of syllables or words. When the change in the fundamental frequency feature is detected to exceed a preset pitch change threshold, or when the silence interval in the duration feature is detected to exceed a preset speech rate slowdown threshold, the first preset threshold is multiplied by a reduction coefficient ranging from 0.5 to 0.9 to lower the judgment threshold.

5. The method according to claim 1, characterized in that, The visual semantic feature vector sequence and the text semantic feature vector sequence are aligned on the time axis, and the semantic consistency score within each time window is calculated using a sliding window, specifically including: A linear interpolation algorithm is used to unify the sampling frequencies of the visual semantic feature vector sequence and the text semantic feature vector sequence to a preset target frequency; Set a sliding window with a duration of 3 to 5 seconds, and slide it from the beginning of the timeline with a preset step size; At each sliding window position, the cosine similarity of all audiovisual feature pairs within the window is calculated, and the calculated score is used as the semantic consistency score at the center time of that window, thus obtaining a score sequence that changes over time.

6. The method according to claim 1, characterized in that, By combining semantic consistency scores, visual slicing detection, and scene classification constraints, candidate points for logical boundaries are determined, specifically including: Calculate the first-order difference of the semantic consistency score sequence. When the first-order difference shows a continuous decreasing trend within a predetermined number of consecutive sampling points, and the absolute value of the semantic consistency score is lower than the dynamically adjusted first preset threshold, a logical boundary warning is triggered. In the warning state, the intersection of the brightness histograms between adjacent keyframes is calculated. When the overlap of the brightness histogram intersection is lower than the second preset threshold, it is confirmed that the shot has switched. When the logical boundary warning is triggered and the camera switch is confirmed, that moment is marked as a logical boundary candidate point.

7. The method according to claim 6, characterized in that, Combining semantic consistency scores, visual slicing detection, and scene classification constraints, the determination of logical boundary candidate points also includes: When the logical boundary warning and camera switching conditions are not met simultaneously, obtain scene label change information near the warning time; The weighted confidence score is calculated using the weighted confidence score formula based on the semantic gradient warning indicator variable, histogram difference, and scene label change indicator variable. When the weighted confidence level exceeds the preset confidence level limit, that moment is marked as a logical boundary candidate point; Wherein, the histogram difference is equal to 1 minus the brightness histogram intersection overlap.

8. The method according to claim 1, characterized in that, Utilizing speech pause information in the audio stream to adjust the segmentation point to a silent region or an energy minimum point, specifically including: A speech activity detection algorithm is used to monitor the short-time energy, zero-crossing rate, and spectral entropy of the audio stream to distinguish between speech intervals and silence intervals; When the logical boundary candidate point is located within the speech interval, the silent interval within a preset time range is searched backward from the candidate point. If a silent interval with a duration greater than the predetermined silence threshold is found within the search range, the final segmentation position is set at the start time of that silent interval. If no silent interval that meets the duration requirement is found, but multiple short silent segments exist, calculate the local minimum point of audio energy within the search range and use this point as the final segmentation position.

9. The method according to claim 1 or 8, characterized in that, Perform lip-shape synchronization detection and fine-tune the segmentation point to the moment of lip closure, specifically including: For close-up portrait videos, detect the facial area of ​​the person in the video frame and locate the key points of the lips; The opening and closing state of the lips is determined by calculating the vertical distance between key points of the upper and lower lips; Near the fine-tuned segmentation position, the opening and closing state of the lips is analyzed frame by frame. If the current segmentation position is in the open mouth state, the nearest closing mouth moment is searched forward or backward, and the closing moment is determined as the final segmentation position.

10. An audio / video content segmentation and positioning system, characterized in that, For performing the method according to any one of claims 1 to 9, comprising: The data acquisition and decomposition module is used to acquire the raw audio and video data to be processed and decompose it into video frame sequences and audio streams. The visual semantic extraction module is used to perform feature parsing on the video frame sequence and generate a visual semantic feature vector sequence. The audio-text semantic extraction module is used to process the audio stream, convert it into text content through automatic speech recognition technology, and then map it into a sequence of text semantic feature vectors. The semantic consistency calculation module is used to align the visual semantic feature vector sequence with the text semantic feature vector sequence on the time axis, and calculate the semantic consistency score within each time window through a sliding window. A dynamic threshold adjustment module is used to extract the fundamental frequency features and duration features in the audio stream, and dynamically adjust the first preset threshold for triggering boundary warnings based on the fundamental frequency features and duration features. The candidate point determination module is used to determine logical boundary candidate points by combining the semantic consistency score, visual slicing detection and scene classification constraints. The segmentation point fine-tuning module is used to fine-tune the candidate logical boundary points. It uses the speech pause information in the audio stream to adjust the segmentation points to the silence range or the energy minimum point, and performs lip-sync detection to fine-tune the segmentation points to the moment when the lips close, thus determining the final segmentation moment.