Intelligent switching optimization method and system for audio and video scenes combined with pattern recognition

By constructing an audio and video scene pattern map and performing dynamic pattern matching, the problems of accidental switching and insufficient adaptability of existing audio and video players during scene transitions are solved, achieving more accurate and smoother scene transitions.

CN120856941BActive Publication Date: 2026-01-30SHENZHEN ZIDOO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511344043.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-01-30
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing audio and video players rely on a single feature for scene switching, leading to accidental and inaccurate switching, and lacking dynamic adaptability, thus failing to meet the intelligent switching needs of complex audio and video content.

Method used

By acquiring real-time audio and video data streams, a current scene pattern map is constructed, which includes audio, video, and audio-video related pattern branches. Dynamic pattern matching is performed in combination with a preset scene pattern template library to generate scene switching trigger signals and transition parameters, thereby achieving multi-dimensional and fine-grained scene recognition and decision-making.

Benefits of technology

It improves the accuracy and smoothness of audio and video player switching in multiple scenarios, adapts to the needs of complex playback scenarios, avoids accidental switching and network fluctuation interference, and ensures audio-visual synchronization and smooth transition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856941B_ABST
    Figure CN120856941B_ABST
Patent Text Reader

Abstract

This application relates to the fields of pattern recognition and audio / video processing technology, and provides a method and system for intelligent switching optimization of audio / video scenes combined with pattern recognition. In this application, real-time audio / video data streams of audio / video units sorted by timestamps are acquired; based on a scene pattern template library, the pattern features of the audio / video data streams are analyzed to construct a current scene pattern map containing audio, video, and related branches; the current scene pattern map is dynamically matched with a standard pattern map to calculate the pattern matching degree sequence and deviation feature sequence; combined with a scene switching decision rule library, scene switching trigger signals, target scene identifiers, and switching transition parameters are generated; based on the above information, switching and optimization are performed to obtain the optimized audio / video output stream. This method improves the accuracy and smoothness of scene switching and can meet the requirements of intelligent switching in complex audio / video playback scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of pattern recognition and audio / video processing technology, and in particular to an intelligent switching optimization method and system for audio / video scenes that combines pattern recognition. Background Technology

[0002] In the field of audio and video processing, intelligent scene switching in audio and video players is a key element in enhancing the user viewing experience and adapting to diverse playback needs. Whether it's local video playback, online streaming, or multi-source content aggregation, various scene transition requirements must be addressed: for example, when switching from a movie to the opening and closing credits, the aspect ratio needs to be automatically adjusted to adapt to the text display; when switching from a static landscape to a fast-moving scene, frame rate and dynamic compensation parameters need to be optimized to avoid stuttering; when switching from a single interview to a multi-person interactive scene, the audio channel allocation needs to be adjusted to highlight the main voice; when switching from a regular 2D video to an HDR scene, color mapping and brightness parameters need to be adapted simultaneously. These transitions require the player to automatically adjust audio and video decoding, rendering, and output parameters to ensure the accuracy, smoothness, and consistency of the content presentation.

[0003] Existing audio and video players have significant limitations in scene transition processing. Most rely solely on a single feature to determine the timing of transitions, such as triggering aspect ratio adjustments based solely on changes in video resolution or channel switching based solely on audio volume thresholds. This ignores the collaborative relationship between multiple audio and video features, easily leading to erroneous transitions. For example, when playing documentaries with sudden loud sound effects, triggering "voice enhancement" solely based on volume thresholds can result in audio distortion; when playing restored old films with short aspect ratio fluctuations, frequently switching picture modes based solely on resolution changes can cause visual interference. While some technologies consider both audio and video features, the processing is fragmented and lacks a structured scene characterization system. For instance, video features are only extracted from the current frame's color histogram, and audio features are only analyzed for frequency distribution, without linking the motion trajectories of consecutive frames to the audio and video content. In complex scenarios such as variety shows with multi-camera transitions or educational videos with picture-in-picture elements, they cannot accurately identify the dominant scene, leading to chaotic transition decisions.

[0004] Meanwhile, existing technology matching lacks dynamic adaptability, often relying on static template comparison. It presets fixed standard scene parameters, simply comparing current features with the template without considering dynamic content changes. For example, when playing a concert video with a gradually changing rhythm, the static template cannot adapt to the continuous changes in rhythm and dynamic range, easily leading to switching delays or false triggers. When playing online videos with dynamic bitrates, temporary feature deviations caused by network fluctuations may be misjudged as scene switching requirements. Furthermore, transitions often use fixed parameters, such as a fixed 0.5-second fade-in / fade-out duration, without dynamically adjusting to scene differences. When switching from a low-frame-rate nostalgic video to a high-frame-rate high-definition video, the fixed transition results in abrupt image transitions, failing to meet the intelligent switching needs of complex audio and video content. Summary of the Invention

[0005] In view of the above, and aiming to at least partially address the shortcomings of existing technologies and bring new solutions to the field of audio transmission, this application provides, in a first aspect, an audio-visual scene intelligent switching optimization method combining pattern recognition, the method comprising:

[0006] Acquire real-time audio and video data streams, wherein the audio and video data streams contain continuous audio and video units arranged in timestamp order, and each audio and video unit contains synchronized audio and video segments;

[0007] Based on a preset scene pattern template library, the audio and video data streams are processed by pattern feature parsing to construct a current scene pattern graph. The current scene pattern graph includes an audio pattern branch, a video pattern branch, and an audio-video associated pattern branch. The audio pattern branch is composed of an audio feature chain, the video pattern branch is composed of a video feature chain, and the audio-video associated pattern branch is composed of an associated feature chain.

[0008] The current scene pattern map is dynamically matched with the standard pattern map in the scene pattern template library to calculate the pattern matching degree sequence and the deviation feature sequence. The pattern matching degree sequence includes audio branch matching degree, video branch matching degree and related branch matching degree. The deviation feature sequence includes audio deviation feature, video deviation feature and related deviation feature.

[0009] Based on the pattern matching degree sequence and the deviation feature sequence, and combined with the preset scene switching decision rule base, a scene switching trigger signal, a target scene identifier, and switching transition parameters are generated.

[0010] Based on the scene switching trigger signal, the target scene identifier, and the switching transition parameters, scene mode switching and optimization processing is performed on the audio and video data stream to obtain an optimized audio and video output stream.

[0011] Secondly, embodiments of this application also provide an intelligent audio-visual scene switching optimization system that combines pattern recognition, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the machine-readable storage medium to realize the intelligent audio-visual scene switching optimization method that combines pattern recognition.

[0012] In summary, the intelligent audio-visual scene switching optimization method and system combining pattern recognition provided in this application acquires real-time data streams of audio and video units sorted by timestamps, and constructs a current scene pattern graph based on a scene pattern template library. This graph includes audio mode branches, video mode branches, and audio-video association mode branches, achieving structured integration of multi-dimensional features of the player's audio and video data streams. This graph not only covers the spectrum, rhythm, and semantic features of audio (adapting to channel adjustment and sound effect optimization scenarios), and the frame images, motion trajectories, and scene structure features of video (adapting to aspect ratio and frame rate compensation scenarios), but also includes time synchronization and content association features between audio and video (adapting to audio-visual collaborative optimization scenarios). Compared to the switching methods of existing players that rely on single or simple feature combinations, this method can more comprehensively depict the essence of the playback scene, providing a reliable feature foundation for subsequent switching decisions and solving the problem of scene misjudgment caused by feature fragmentation in traditional players.

[0013] In the dynamic pattern matching stage, the current scene pattern map is compared with the standard pattern map, and the matching degree and deviation feature sequence of audio, video, and related branches are calculated separately to achieve multi-dimensional and fine-grained comparison. For example, audio branch matching combines spectrum, rhythm, and semantic similarity to accurately identify the differences between "voice scenes" and "sound effect scenes"; video branch matching combines frame images, motion trajectories, and scene structure consistency to accurately distinguish between "static scenes" and "dynamic scenes"; related branch matching focuses on time synchronization accuracy and content association strength to avoid erroneous switching in audio-visual desynchronization scenarios. This multi-branch parallel matching method effectively resists the influence of feature deviations caused by network fluctuations during playback and image fluctuations of older resources on single features, significantly improving the accuracy of scene pattern matching for the player.

[0014] Furthermore, based on pattern matching degree and deviation characteristics, and combined with a scene switching decision rule library, switching trigger signals, target scene identifiers, and transition parameters are generated, overcoming the limitations of fixed threshold decision-making in existing players. By smoothly processing matching degree through a sliding window and identifying significant deviations through peak detection, and dynamically adjusting the decision threshold in conjunction with preset priority rules, the switching timing is made more aligned with playback requirements. For example, during HDR scene switching, it is more sensitive to deviations in video frame image features, and can quickly trigger color parameter adaptation; during the switching transition phase, a transition audio-visual feature chain is generated through feature interpolation based on transition parameters, and synchronization is optimized to avoid abrupt changes in picture and sound caused by fixed transition parameters. Thus, through the synergistic effect of multi-dimensional feature integration, dynamic and accurate matching, intelligent decision-making, and smooth transition, the audio-visual player achieves accuracy, timeliness, and smoothness in switching across multiple scenarios, meeting the high requirements of intelligent scene switching in complex playback scenarios.

[0015] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the above drawings without creative effort.

[0017] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.

[0018] Figure 1 This is a flowchart illustrating an intelligent audio-visual scene switching optimization method that combines pattern recognition, as provided in an embodiment of this application.

[0019] Figure 2 This is a schematic diagram illustrating an application scenario of an intelligent audio-visual scene switching optimization method that combines pattern recognition, as provided in an embodiment of this application.

[0020] Figure 3 This is a schematic diagram of an audio-visual scene intelligent switching optimization system that combines pattern recognition, provided in an embodiment of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0022] Please see Figure 1 and Figure 2 As shown, Figure 1 This is a flowchart illustrating the intelligent audio-visual scene switching optimization method combining pattern recognition provided in this application embodiment. Figure 2 This is a schematic diagram of an audio-visual interaction scenario. The audio-visual interaction scenario may include an audio-visual content processing platform 100 and multiple audio-visual playback terminals 200 for audio-visual content interaction. In this embodiment, the method can be implemented by the audio-visual content processing platform 100 or by the audio-visual playback terminals 200, depending on the actual application scenario. For example, Figure 2 As shown, the audio and video playback terminal can be a computer terminal, mobile phone, tablet computer, or other device with data transmission, analysis, and processing capabilities. The audio and video content processing platform 100 can be a server, server cluster, computer equipment, etc. For example, it can be a backend server for providing audio and video content for the audio and video playback terminal 200 to play. This embodiment does not specifically limit this.

[0023] like Figure 1 As shown, the method includes steps S110-S150, which will be described in detail below.

[0024] The following section uses a specific audio and video playback scenario (such as online concert audio and video content, online live audio and video content, streaming media audio and video content, etc., hereinafter referred to as "scenario") as an example to provide a detailed explanation of the intelligent switching optimization method for audio and video scenarios combined with pattern recognition of the present invention.

[0025] Step S110: Acquire real-time audio and video data streams, wherein the audio and video data streams contain continuous audio and video units arranged in timestamp order, and each audio and video unit contains synchronized audio and video segments.

[0026] In this embodiment, as an example, in a specific audio-visual interaction scenario, audio and video data streams are acquired in real time through acquisition devices set up within the scenario. Acquisition devices include, but are not limited to, cameras and microphones (e.g., cameras and microphones built into the user terminal, or independent cameras and microphones set up in the scenario). These devices operate at a preset sampling frequency, converting the acquired audio and video information into digital signals, and then encapsulating them into continuous audio and video units according to the order of their timestamps. Each audio and video unit corresponds to the same time length, which is a preset fixed duration, and the audio and video segments within each audio and video unit are synchronized in time; that is, the start and end timestamps of the audio segment are consistent with the start and end timestamps of the corresponding video segment.

[0027] Step S120: Based on a preset scene pattern template library, perform pattern feature parsing processing on the audio and video data stream to construct a current scene pattern graph. The current scene pattern graph includes an audio pattern branch, a video pattern branch, and an audio-video associated pattern branch. The audio pattern branch is composed of an audio feature chain, the video pattern branch is composed of a video feature chain, and the audio-video associated pattern branch is composed of an associated feature chain.

[0028] In this embodiment, as an example, a preset scene pattern template library is a database pre-built and stored in the system, containing standard pattern templates for various common audio-visual interaction scenarios. When processing audio-visual data streams, this template library is invoked, and the data streams are parsed for pattern features using it as a reference. Through parsing, features related to audio and video, as well as the correlation features between them, are extracted from the audio-visual data streams, forming audio feature chains, video feature chains, and correlation feature chains, which are then combined to form the current scene pattern map. This map can comprehensively reflect the pattern features of the current audio-visual interaction scenario.

[0029] In this embodiment, step S120 may include sub-steps S121-S125, which will be described in detail below.

[0030] Step S121: Obtain the scene mode template library, which contains at least one standard scene mode template. Each standard scene mode template contains a standard audio mode branch, a standard video mode branch, and a standard association mode branch.

[0031] In this embodiment, as an example, a scene pattern template library is read from the storage module via an internal interface. The standard scene pattern templates in this library are constructed based on a large amount of historical audio and video data and common scene features. Each standard scene pattern template is designed for a specific scene type. Its standard audio pattern branch contains typical audio feature sequences for that scene, its standard video pattern branch contains typical video feature sequences, and its standard association pattern branch contains typical association sequences between audio and video features. For example, a standard scene pattern template might correspond to a multi-person discussion scene. Its standard audio pattern branch might contain feature sequences from multiple different sound sources, its standard video pattern branch might contain video feature sequences of multi-person activities, and its standard association pattern branch might contain association feature sequences between sound sources and corresponding character actions.

[0032] Step S122: Perform audio and video separation processing on the audio and video data stream to obtain independent audio data stream and video data stream. The audio data stream is composed of audio segments in the audio and video unit arranged in timestamp order, and the video data stream is composed of video segments in the audio and video unit arranged in timestamp order.

[0033] In this embodiment, as an example, an audio-video separation algorithm can be used to process the audio and video data streams. By identifying the differences in the encapsulation formats of the audio and video data in the audio-video unit, the two can be separated. The resulting audio data stream retains the timestamp order of the audio segments in the original audio-video unit, and the time information of each audio segment is consistent with the time information in the original audio-video unit. Similarly, the video data stream is also arranged according to the timestamp order of the video segments in the original audio-video unit, ensuring the correspondence between the audio and video data streams in the time dimension, which facilitates subsequent feature parsing.

[0034] Step S123: Perform audio pattern feature parsing processing on the audio data stream to obtain an audio feature chain. Use the audio feature chain as the audio pattern branch of the current scene pattern map. The audio feature chain includes audio spectrum features, audio rhythm features, and audio semantic features distributed along the time axis.

[0035] In this embodiment, as an example, the audio spectral features are first extracted by analyzing the frequency distribution of the audio signal to obtain information such as the energy distribution of different frequency bands. Next, the rhythmic features of the audio are analyzed to identify beat patterns. Then, semantic analysis is performed on the speech portion of the audio to extract the theme and intent. The extracted features are arranged in chronological order to form an audio feature chain. This feature chain serves as an audio pattern branch of the current scene pattern map, reflecting the feature changes of the audio at different points in time.

[0036] In this embodiment, step S123 may include sub-steps S1231-S1235, which will be described in detail below.

[0037] Step S1231: Perform spectrum conversion processing on each audio segment in the audio data stream to convert the time-domain audio signal into frequency-domain spectrum data, extract features from the frequency-domain spectrum data to obtain audio spectrum features, which include spectral energy distribution features and spectral entropy value features.

[0038] In this embodiment, as an example, a spectrum conversion algorithm is used to process each audio segment, converting the original time-domain audio signal, which is variable in time, into frequency-domain spectral data, which is variable in frequency. During the conversion process, operations such as framing and windowing are performed on the audio signal to reduce the impact of spectral leakage. After obtaining the frequency-domain spectral data, the spectral energy distribution characteristics and spectral entropy characteristics are further extracted. The spectral energy distribution characteristics reflect the energy relationship between different frequency components, while the spectral entropy characteristics reflect the degree of disorder in the spectral distribution. These characteristics together constitute the audio spectral features.

[0039] Step S1232: Perform rhythm feature extraction processing on the audio data stream, identify the beat start point and beat interval in the audio signal, calculate the rhythm intensity parameter and rhythm stability parameter based on the beat start point and the beat interval, and combine the rhythm intensity parameter and the rhythm stability parameter into audio rhythm features.

[0040] In this embodiment, as an example, an audio data stream is processed using a rhythm detection algorithm. First, the audio signal is preprocessed to remove noise interference. Then, by analyzing the energy and frequency changes of the signal, the starting point of the beat and the interval between adjacent beats are identified. Based on this information, a rhythm intensity parameter is calculated, reflecting the strength of the beat; simultaneously, a rhythm stability parameter is calculated, reflecting the stability of the beat interval. These two parameters are combined to form the audio rhythm features, describing the rhythmic characteristics of the audio.

[0041] Step S1233: Perform speech activity detection processing on the speech signal in the audio data stream, separate the audio segments containing speech from the non-speech audio segments, perform speech recognition processing on the audio segments containing speech, and obtain text information.

[0042] In this embodiment, as an example, a speech activity detection algorithm is used to process the audio data stream. By analyzing the energy and spectral characteristics of the audio signal, it determines whether an audio segment contains a speech signal, thereby dividing the audio data stream into audio segments containing speech and non-speech audio segments. For audio segments containing speech, a speech recognition algorithm is invoked to process the speech signal into corresponding text information. During the speech recognition process, the cooperation between acoustic and language models improves the accuracy of recognition, ensuring that the obtained text information accurately reflects the speech content.

[0043] Step S1234: Perform semantic analysis on the text information, extract topic keywords and semantic intent features, and combine the topic keywords and semantic intent features into audio semantic features.

[0044] In this embodiment, as an example, the text can first undergo preprocessing operations such as word segmentation and part-of-speech tagging. Then, by analyzing the relationships between words and the structure of sentences, the main keywords in the text can be extracted. These keywords can reflect the core content of the text. At the same time, by analyzing the context of the text, the semantic intent of the scene object can be determined, forming semantic intent features. Combining the main keywords and semantic intent features constitutes audio semantic features, which can reflect the semantic content of the speech in the audio.

[0045] Step S1235: Combine the audio spectral features, audio rhythm features, and audio semantic features in time-stamp order to obtain an audio feature chain; use the audio feature chain as an audio mode branch of the current scene mode graph, where each node in the audio mode branch corresponds to an audio feature combination with a timestamp.

[0046] In this embodiment, as an example, the audio spectral features, audio rhythm features, and audio semantic features corresponding to each time stamp are combined according to the chronological order of the audio segments' timestamps. Each feature combination corresponding to a timestamp forms a node, and multiple nodes are connected chronologically to form an audio feature chain. This audio feature chain is assigned the role of an audio pattern branch of the current scene pattern graph, where each node fully contains all the features of the audio at the corresponding timestamp, facilitating subsequent comparison and analysis with a standard pattern graph.

[0047] Step S124: Perform video mode feature parsing processing on the video data stream to obtain a video feature chain. Use the video feature chain as the video mode branch of the current scene mode map. The video feature chain includes video frame image features, motion trajectory features and scene structure features distributed along the time axis.

[0048] In this embodiment, as an example, the parsing of video data streams can be carried out from multiple dimensions. First, image features of the video frames are extracted, including information such as color and texture. Then, by analyzing the changes between consecutive video frames, motion trajectory features are obtained, reflecting the movement of objects in the scene. Next, the scene structure in the video frames is analyzed, and scene structure features are extracted. The above features are arranged in chronological order to form a video feature chain. This feature chain serves as the video mode branch of the current scene mode graph, which can comprehensively display the feature status of the video at different points in time.

[0049] In this embodiment, step S124 may include sub-steps S1241-S1245, which will be described in detail below.

[0050] Step S1241: Perform frame segmentation processing on each video segment in the video data stream to obtain a continuous video frame sequence. Perform image feature extraction processing on each video frame to obtain video frame image features, which include color distribution features and texture structure features.

[0051] In this embodiment, as an example, each video segment can be divided into frames according to a preset frame rate, decomposing continuous video segments into a series of discrete video frames to form a video frame sequence. For each video frame, an image feature extraction algorithm is used for processing. When extracting color distribution features, the video frame is converted to a specific color space, and the distribution of different color components is statistically analyzed. When extracting texture structure features, parameters reflecting texture characteristics are obtained by analyzing the arrangement patterns of pixels and grayscale changes in the image. These parameters together constitute the image features of the video frame.

[0052] Step S1242: Perform optical flow estimation processing on adjacent video frames in the video frame sequence, calculate pixel-level motion vectors, perform target tracking processing based on the motion vectors, obtain the trajectory coordinate sequence of the moving target, extract features from the trajectory coordinate sequence, and obtain motion trajectory features, which include trajectory direction features and trajectory velocity features.

[0053] In this embodiment, as an example, an optical flow estimation algorithm can be applied to two adjacent video frames in a video frame sequence. This algorithm calculates the motion vector of each pixel—its displacement in the horizontal and vertical directions—by analyzing the grayscale changes of pixels between adjacent frames. Based on these motion vectors, a target tracking algorithm is used to track moving targets in the scene, recording the target's position coordinates in different video frames to form a trajectory coordinate sequence. Analyzing this sequence, the direction information of the trajectory (i.e., trajectory direction features) and the distance the target moves per unit time (i.e., trajectory velocity features) are extracted. These two features together constitute the motion trajectory features.

[0054] Step S1243: Perform scene segmentation processing on the keyframes in the video frame sequence, divide the keyframes into multiple semantic regions, perform spatial relationship analysis on the semantic regions, calculate the region area ratio and region position relationship, and combine the region area ratio and region position relationship into scene structure features.

[0055] In this embodiment, as an example, keyframes are first selected from the video frame sequence. Keyframe selection is based on the degree of change in the video content; when the change exceeds a preset threshold, the frame is identified as a keyframe. Then, scene segmentation processing is performed on the keyframes, using an image segmentation algorithm to divide them into multiple regions with different semantic meanings, such as character regions and background regions. Next, spatial relationship analysis is performed on these semantic regions, calculating the area proportion of each region within the keyframe and the relative positional relationships between different regions, such as top / bottom, left / right, and containment. The region area proportions and positional relationships are combined to form scene structural features.

[0056] Step S1244: Combine the video frame image features, motion trajectory features, and scene structure features in time stamp order to obtain a video feature chain.

[0057] In this embodiment, as an example, the video frame image features, motion trajectory features, and scene structure features corresponding to each timestamp can be integrated according to the timestamp order of the video segments. The feature combination under each timestamp forms an independent unit, and multiple such units are connected in chronological order to form a video feature chain. This feature chain completely records the various features of the video data stream at different time points, so as to facilitate the subsequent construction of video mode branches.

[0058] Step S1245: Use the video feature chain as the video mode branch of the current scene mode graph, where each node in the video mode branch corresponds to a combination of video features with a timestamp.

[0059] In this embodiment, as an example, the constructed video feature chain is given the functionality of a video mode branch in the current scene mode graph. Each unit in the video feature chain corresponds to a timestamp, and the video frame image features, motion trajectory features, and scene structure features contained in each unit together constitute the node of that timestamp in the video mode branch. The above nodes are arranged in order of timestamp, clearly showing the changes in video features over time, which facilitates comparison with the standard video mode branch in the standard scene mode template.

[0060] Step S125: Perform audio-video association pattern feature parsing processing on the audio data stream and the video data stream to obtain an association feature chain. Use the association feature chain as the audio-video association pattern branch of the current scene pattern graph. The association feature chain includes time synchronization features and content association features distributed along the time axis. Perform graph structuring processing on the audio pattern branch, the video pattern branch, and the audio-video association pattern branch to generate a current scene pattern graph containing node layers, connection edges, and weight parameters. The node layers correspond to the feature units in each feature chain, the connection edges represent the association relationship between feature units, and the weight parameters represent the strength of the association relationship.

[0061] In this embodiment, as an example, correlation analysis can be performed on audio and video data streams. First, the time synchronization of the two streams is analyzed to obtain time synchronization features; then, the content correlation is analyzed to obtain content correlation features. These features are arranged in chronological order to form a correlation feature chain, serving as the audio-video correlation pattern branch of the current scene pattern graph. Subsequently, the audio pattern branch, video pattern branch, and audio-video correlation pattern branch can be processed into a graph structure, using the feature units in each feature chain as node layers. Connection edges represent the correlation between different nodes, and weight parameters are assigned to the connection edges based on the tightness of the correlation, thereby generating a complete current scene pattern graph.

[0062] In this embodiment, step S125 may include sub-steps S1251-S1254, which will be described in detail below.

[0063] Step S1251: Extract the audio timestamp sequence of the audio data stream and the video timestamp sequence of the video data stream, perform time alignment processing on the audio timestamp sequence and the video timestamp sequence, calculate the time deviation value and the synchronization fluctuation parameter, and combine the time deviation value and the synchronization fluctuation parameter into a time synchronization feature.

[0064] In this embodiment, as an example, timestamp sequences are extracted from the audio and video data streams respectively. The audio timestamp sequence records the start and end times of each audio segment, and the video timestamp sequence records the start and end times of each video segment. Then, time alignment processing is performed on the two timestamp sequences. By comparing the timestamps of corresponding segments, the time deviation value between them is calculated, i.e., the time difference between the audio segment and the corresponding video segment. Simultaneously, the changes in the time deviation value over a period of time are analyzed to obtain a synchronization fluctuation parameter, reflecting the stability of time synchronization. The time deviation value and the synchronization fluctuation parameter are combined to form the time synchronization feature.

[0065] Step S1252: Perform content matching processing on the audio semantic features and the video frame image features, calculate the semantic relevance parameter, perform dynamic association processing on the audio rhythm features and the motion trajectory features, calculate the rhythm matching parameter, and combine the semantic relevance parameter and the rhythm matching parameter into a content association feature.

[0066] In this embodiment, as an example, audio semantic features are matched with video frame image features. By analyzing the correlation between the content expressed by the audio semantics and the content displayed by the video frame image, a semantic relevance parameter is calculated. The higher the parameter, the better the content match between the two. Simultaneously, audio rhythm features are dynamically correlated with motion trajectory features. By analyzing whether the changes in audio beat and motion trajectory are synchronized, a rhythm matching parameter is calculated. The semantic relevance parameter and the rhythm matching parameter are combined to form a content association feature, reflecting the degree of content association between audio and video.

[0067] Step S1253: Combine the time synchronization feature and the content association feature in the order of timestamps to obtain an association feature chain; use the association feature chain as the audio-video association mode branch of the current scene mode graph, where each node in the audio-video association mode branch corresponds to an audio-video association feature combination with a timestamp; construct cross-branch association edges based on the audio mode branch, the video mode branch, and the nodes with corresponding timestamps in the audio-video association mode branch, where the weight parameters of the cross-branch association edges are determined based on the association strength of the corresponding node features.

[0068] In this embodiment, as an example, the time synchronization features and content association features corresponding to each timestamp can be combined in chronological order to form an association feature chain. This association feature chain serves as the audio-video association mode branch of the current scene mode graph, where each node corresponds to a combination of audio-video association features for a single timestamp. Next, cross-branch association edges are constructed based on the nodes corresponding to the timestamps in the audio mode branch, video mode branch, and audio-video association mode branch. The weight parameters of the association edges are determined based on the association strength between the features of the corresponding nodes; the higher the association strength, the larger the weight parameter, thus reflecting the degree of association between nodes in different branches.

[0069] Step S1254: Perform graph structuring processing on the audio mode branch, the video mode branch, and the audio-video association mode branch to generate a current scene mode graph containing node layers, connection edges, and weight parameters. The node layers correspond to the feature units in each feature chain, the connection edges represent the association relationships between feature units, and the weight parameters represent the strength of the association relationships.

[0070] In this embodiment, as an example, the audio mode branch, video mode branch, and audio-video association mode branch can be processed into a graph structure. First, the feature units of each feature chain in the three branches are mapped to the node layer of the graph, with each feature unit corresponding to an independent node containing the specific feature information of that feature unit. Then, the association relationships between different nodes are analyzed. For nodes with adjacent timestamps within the same branch, connection edges are established based on the continuity and trend of features; for nodes with corresponding timestamps between different branches, cross-branch connection edges are established based on the association of features. Finally, a weight parameter is assigned to each connection edge according to the tightness of the association relationship between nodes. The tighter the association relationship, the larger the weight parameter, thus forming a current scene mode graph containing node layers, connection edges, and weight parameters.

[0071] Step S130: Perform dynamic pattern matching processing on the current scene pattern map and the standard pattern map in the scene pattern template library, and calculate the pattern matching degree sequence and the deviation feature sequence. The pattern matching degree sequence includes audio branch matching degree, video branch matching degree and related branch matching degree, and the deviation feature sequence includes audio deviation feature, video deviation feature and related deviation feature.

[0072] In this embodiment, as an example, the constructed current scene pattern map is dynamically matched with the standard pattern map in the scene pattern template library. During the matching process, the node features and relationships of the two are compared time-stamp by time. The matching degree of audio, video, and related branches is calculated using a specific algorithm to form a pattern matching degree sequence. At the same time, the feature differences between the two on each branch are calculated to form a deviation feature sequence. The above sequence can quantify the similarity and differences between the current scene and the standard scene, providing a basis for subsequent scene switching decisions.

[0073] In this embodiment, step S130 may include sub-steps S131-S136, which will be described in detail below.

[0074] Step S131: Select at least one standard scene mode template corresponding to the current application scenario from the scene mode template library, and extract the standard scene mode graph from the standard scene mode template. The standard scene mode graph includes a standard audio mode branch, a standard video mode branch, and a standard association mode branch.

[0075] In this embodiment, as an example, standard scene pattern templates related to the current scene type can be selected from the scene pattern template library based on the preliminary judgment result of the current audio and video interaction scene. For example, if the current scene is initially judged to be a speech scene, then a standard scene pattern template related to a speech is selected. Then, a standard scene pattern graph is extracted from the selected standard scene pattern template. The standard audio pattern branch, standard video pattern branch, and standard association pattern branch in this graph contain the standard feature sequence and association relationship under this type of scene, respectively.

[0076] Step S132: Perform feature comparison processing on the audio mode branch of the current scene mode map and the standard audio mode branch, calculate the audio feature similarity of the corresponding timestamp node, arrange the audio feature similarity according to the time axis, and obtain the audio branch matching degree.

[0077] In this embodiment, as an example, the audio mode branch of the current scene mode graph is compared with the standard audio mode branch on a time-stamp basis. For each time stamp node, the audio spectral features, audio rhythm features, and audio semantic features of the two are compared. By comprehensively calculating the similarity of the above features, the audio feature similarity of that time stamp node is obtained. The audio feature similarities of all time stamp nodes are arranged in chronological order to form the audio branch matching degree, which reflects the overall matching situation between the current audio mode and the standard audio mode.

[0078] In this embodiment, step S132 may include sub-steps S1321-S1327, which will be described in detail below.

[0079] Step S1321: Extract the audio spectrum features, audio rhythm features, and audio semantic features from the audio mode branches of the current scene mode map to form the current audio feature matrix.

[0080] In this embodiment, as an example, audio spectral features, audio rhythm features, and audio semantic features are extracted from the audio mode branch of the current scene mode graph for each timestamp node. These features are arranged in timestamp order, with each timestamp corresponding to a feature as a row of a matrix, and different types of features as different columns, thus constructing the current audio feature matrix. This matrix completely contains all feature information of the current audio mode branch at each timestamp.

[0081] Step S1322: Extract the standard audio spectral features, standard audio rhythm features, and standard audio semantic features from the standard audio mode branch to form a standard audio feature matrix.

[0082] In this embodiment, as an example, standard audio spectral features, standard audio rhythm features, and standard audio semantic features are extracted from each timestamp node of the standard audio mode branch. Using the same method as constructing the current audio feature matrix, the above standard features are arranged in timestamp order to form a standard audio feature matrix, where each row corresponds to a standard feature of one timestamp, and each column corresponds to a type of standard feature.

[0083] Step S1323: Perform matrix alignment processing on the current audio feature matrix and the standard audio feature matrix to make their time dimension and feature dimension consistent.

[0084] In this embodiment, as an example, the time dimension and feature dimension of the current audio feature matrix and the standard audio feature matrix are first checked to see if they are consistent. If the time dimension is inconsistent, the number of rows in the matrix is ​​adjusted by interpolation or truncation to make the number of timestamps in both matrices the same. If the feature dimensions are inconsistent, corresponding feature columns are added or removed to make the feature types contained in both matrices the same. Through the above processing, the current audio feature matrix and the standard audio feature matrix are made structurally consistent to facilitate subsequent feature comparison.

[0085] Step S1324: Calculate the cosine similarity between the audio spectral features in the current audio feature matrix and the standard audio spectral features to obtain the spectral similarity parameter.

[0086] In this embodiment, as an example, a cosine similarity algorithm is applied to calculate the audio spectral similarity columns corresponding to the timestamps in the current audio feature matrix and the standard audio feature matrix. This algorithm measures the similarity between two feature vectors by calculating the cosine of the angle between them; a value closer to 1 indicates a higher similarity. The calculation results for each timestamp are used as spectral similarity parameters to form a sequence of spectral similarity parameters.

[0087] Step S1325: Perform dynamic time warping on the audio rhythm features in the current audio feature matrix and the standard audio rhythm features to obtain rhythm similarity parameters.

[0088] In this embodiment, as an example, a dynamic time warping algorithm is used to process the audio rhythm feature columns corresponding to the timestamps in the current audio feature matrix and the standard audio feature matrix. By finding the optimal time alignment path, the problem of temporal scaling of audio rhythm is solved. Then, based on the aligned features, the similarity is calculated to obtain the rhythm similarity parameter under each timestamp, forming a rhythm similarity parameter sequence.

[0089] Step S1326: Calculate the semantic similarity between the audio semantic features in the current audio feature matrix and the standard audio semantic features to obtain semantic similarity parameters.

[0090] In this embodiment, as an example, semantic similarity can be calculated for the audio semantic feature columns corresponding to the timestamps in the current audio feature matrix and the standard audio feature matrix. By analyzing the degree of overlap of topic keywords and the degree of matching of semantic intent, a specific semantic similarity measurement method is used to calculate the semantic similarity parameter under each timestamp, forming a semantic similarity parameter sequence.

[0091] Step S1327: Based on the preset feature weight parameters, perform weighted summation on the spectral similarity parameters, the rhythm similarity parameters, and the semantic similarity parameters to obtain the audio feature similarity of each timestamp node; arrange the audio feature similarities in timestamp order to obtain the audio branch matching degree.

[0092] In this embodiment, as an example, different feature weight parameters are pre-set for the spectral similarity parameter, rhythmic similarity parameter, and semantic similarity parameter. These weight parameters are determined based on the importance of different features in audio pattern matching. The three similarity parameters for each timestamp are multiplied by their corresponding weight parameters, and the products are then summed to obtain the audio feature similarity for that timestamp node. The audio feature similarities of all timestamp nodes are arranged in timestamp order to form the audio branch matching degree.

[0093] Step S133: Perform feature comparison processing on the video mode branch of the current scene mode graph and the standard video mode branch, calculate the video feature similarity of the corresponding timestamp node, arrange the video feature similarity according to the time axis, and obtain the video branch matching degree.

[0094] In this embodiment, as an example, a method similar to audio branch comparison is used to perform time-stamp-wise feature comparison between the video mode branch of the current scene mode graph and the standard video mode branch. The similarity of video frame image features, motion trajectory features, and scene structure features is compared separately. The video feature similarity at each timestamp node is calculated by weighted summation. These similarities are then arranged in chronological order to obtain the video branch matching degree, reflecting the matching status between the current video mode and the standard video mode.

[0095] Step S134: Perform feature comparison processing on the audio-visual association mode branch of the current scene mode map and the standard association mode branch, calculate the association feature similarity of the corresponding timestamp nodes, arrange the association feature similarity according to the time axis, and obtain the association branch matching degree.

[0096] In this embodiment, as an example, the audio-video association pattern branch of the current scene pattern graph can be compared with the standard association pattern branch on a time-stamp basis. For each timestamp node, the similarity of time synchronization features and content association features is compared, and the association feature similarity is obtained through comprehensive calculation. The association feature similarities of all timestamp nodes are arranged in chronological order to form the association branch matching degree, which reflects the degree of matching between the current audio-video association pattern and the standard association pattern.

[0097] Step S135: Combine the audio branch matching degree, the video branch matching degree, and the associated branch matching degree into a pattern matching degree sequence.

[0098] In this embodiment, as an example, the audio branch matching degree, video branch matching degree, and associated branch matching degree are integrated in timestamp order to form a pattern matching degree sequence. Each element in this sequence contains the matching degree of the three branches under the corresponding timestamp, which can comprehensively reflect the matching status of the current scene pattern map and the standard scene pattern map on each branch, providing multi-dimensional reference information for subsequent scene switching decisions.

[0099] Step S136: Calculate the feature difference between the audio mode branch of the current scene mode graph and the standard audio mode branch to obtain audio deviation features; calculate the feature difference between the video mode branch of the current scene mode graph and the standard video mode branch to obtain video deviation features; calculate the feature difference between the audio-video association mode branch of the current scene mode graph and the standard association mode branch to obtain association deviation features; combine the audio deviation features, the video deviation features, and the association deviation features into a deviation feature sequence.

[0100] In this embodiment, as an example, the feature differences between the current scene pattern map and the standard scene pattern map in audio, video, and related branches are calculated respectively. For each branch, the difference between the current feature and the standard feature is calculated time-stamped to obtain the deviation features of each branch. Then, the above deviation features are combined into a deviation feature sequence. This sequence can clearly present the differences between the current scene and the standard scene in various aspects, providing a quantitative basis for the degree of difference in scene switching decisions.

[0101] In this embodiment, step S136 may include sub-steps S1361-S1365, which will be described in detail below.

[0102] Step S1361: Calculate the absolute value of the difference between the corresponding elements of the current audio feature matrix and the standard audio feature matrix to obtain the audio feature difference matrix. Perform row vector summation on the audio feature difference matrix to obtain the total audio feature difference of each timestamp node. Arrange the total audio feature difference according to the time axis to obtain the audio deviation feature.

[0103] In this embodiment, as an example, the current audio feature matrix and the standard audio feature matrix can be subtracted element-wise, and then the absolute value is taken to obtain an audio feature difference matrix. Each element in this matrix represents the degree of difference of the features at the corresponding position. Next, each row of the audio feature difference matrix is ​​summed to obtain the total audio feature difference for each timestamp node. This total difference reflects the overall difference of the audio features at that timestamp. The total audio feature differences for all timestamp nodes are arranged in chronological order to form the audio deviation features.

[0104] Step S1362: Extract video frame image features, motion trajectory features, and scene structure features from the video mode branch of the current scene mode map to form a current video feature matrix; extract standard video frame image features, standard motion trajectory features, and standard scene structure features from the standard video mode branch to form a standard video feature matrix.

[0105] In this embodiment, as an example, video frame image features, motion trajectory features, and scene structure features for each timestamp node are extracted from the video mode branch of the current scene mode graph and arranged in timestamp order to form the current video feature matrix. Simultaneously, corresponding standard features are extracted from the standard video mode branch to form a standard video feature matrix. The matrix is ​​constructed in the same way as the current video feature matrix to facilitate subsequent difference calculations.

[0106] Step S1363: Calculate the absolute value of the difference between the corresponding elements of the current video feature matrix and the standard video feature matrix to obtain the video feature difference matrix. Perform row vector summation on the video feature difference matrix to obtain the total video feature difference for each timestamp node. Arrange the total video feature difference according to the time axis to obtain the video deviation feature.

[0107] In this embodiment, as an example, the absolute value of the difference between corresponding elements of the current video feature matrix and the standard video feature matrix can be calculated to obtain a video feature difference matrix. Then, the summation of each row of this matrix is ​​performed to obtain the total difference of video features at each timestamp. This total difference reflects the overall difference of video features at that timestamp. The total difference is arranged in chronological order to form the video deviation feature.

[0108] Step S1364: Extract the time synchronization features and content association features from the audio-visual association mode branch of the current scene mode graph to form the current association feature matrix; extract the standard time synchronization features and standard content association features from the standard association mode branch to form the standard association feature matrix.

[0109] In this embodiment, as an example, time synchronization features and content association features for each timestamp node are extracted from the audio-visual association pattern branch of the current scene pattern graph, and arranged in timestamp order to form the current association feature matrix. Simultaneously, standard time synchronization features and standard content association features are extracted from the standard association pattern branch to form a standard association feature matrix, with the matrix structure consistent with the current association feature matrix.

[0110] Step S1365: Calculate the absolute value of the difference between the corresponding elements of the current association feature matrix and the standard association feature matrix to obtain the association feature difference matrix. Perform row vector summation on the association feature difference matrix to obtain the total association feature difference for each timestamp node. Arrange the total association feature difference according to the time axis to obtain the association deviation feature. Combine the audio deviation feature, the video deviation feature, and the association deviation feature into a deviation feature sequence.

[0111] In this embodiment, as an example, the absolute value of the difference between corresponding elements of the current association feature matrix and the standard association feature matrix is ​​calculated to obtain the association feature difference matrix. Each row of this matrix is ​​summed to obtain the total association feature difference for each timestamp node. These total differences are then arranged along the time axis to form association deviation features. Finally, the audio deviation features, video deviation features, and association deviation features are combined in timestamp order to form a deviation feature sequence.

[0112] Step S140: Based on the pattern matching degree sequence and the deviation feature sequence, and combined with the preset scene switching decision rule library, generate a scene switching trigger signal, a target scene identifier, and switching transition parameters. The switching transition parameters include transition duration features and transition smoothness features.

[0113] In this embodiment, as an example, the pattern matching degree sequence and the deviation feature sequence are matched with rules in the scene switching decision rule base. Based on the matching result, it is determined whether a scene switch is needed. When the switching conditions are met, a scene switching trigger signal is generated, and parameters such as the target scene identifier, transition duration, and smoothness during the switching process are determined. These parameters ensure a smooth and natural scene switching process, meeting the needs of current audio-visual interaction scenarios.

[0114] In this embodiment, step S140 may include sub-steps S141-S148, which will be described in detail below.

[0115] Step S141: Obtain the scene switching decision rule base, which includes matching degree threshold rules, deviation threshold rules and decision priority rules.

[0116] In this embodiment, as an example, a scene switching decision rule base is read from the storage module. This rule base is built based on a large number of scene switching cases and expert experience. Among them, the matching degree threshold rule specifies the minimum threshold for the matching degree of each branch. When the matching degree is lower than the threshold, scene switching may be required. The deviation threshold rule specifies the maximum threshold for the deviation feature of each branch. When the deviation feature exceeds the threshold, scene switching may be required. The decision priority rule specifies the priority judgment order when multiple switching conditions are met simultaneously.

[0117] Step S142: Perform sliding window averaging on the audio branch matching degree, video branch matching degree, and associated branch matching degree in the pattern matching degree sequence to obtain smoothed audio branch matching degree, smoothed video branch matching degree, and smoothed associated branch matching degree.

[0118] In this embodiment, as an example, a sliding window averaging algorithm is used to process the matching degree of each branch in the pattern matching degree sequence. A sliding window of fixed length is set, and the window moves along the time axis. The matching degree value within each window is averaged to obtain the smoothed matching degree at the center timestamp of that window. This processing can reduce noise interference in the matching degree sequence, make the trend of matching degree change smoother, and facilitate more accurate determination of whether the scene needs to be switched.

[0119] Step S143: Compare the smoothed audio branch matching degree with a preset audio matching degree threshold to determine whether the audio branch switching condition is met; compare the smoothed video branch matching degree with a preset video matching degree threshold to determine whether the video branch switching condition is met; compare the smoothed associated branch matching degree with a preset associated matching degree threshold to determine whether the associated branch switching condition is met.

[0120] In this embodiment, as an example, preset audio matching thresholds, video matching thresholds, and association matching thresholds are obtained from the scene switching decision rule base. The smoothed audio branch matching degree, video branch matching degree, and association branch matching degree are compared with their corresponding thresholds. If the smoothed matching degree is lower than the corresponding threshold, the switching condition for that branch is determined to be met; otherwise, it is not. This method provides a preliminary determination of whether each branch needs to be switched.

[0121] Step S144: Perform peak detection processing on the audio deviation features, video deviation features and associated deviation features in the deviation feature sequence, identify deviation peak points that exceed the preset deviation threshold, and determine whether the deviation triggering condition is met based on the number and intensity of the deviation peak points.

[0122] In this embodiment, as an example, peak detection can be performed on the deviation features of each branch in the deviation feature sequence. First, a preset deviation threshold is set, and then the peak points in the deviation features that exceed the threshold are identified by the peak detection algorithm. The position, intensity, and duration of each peak point are recorded. Based on the number of peak points, the total intensity, and the distribution, it is determined whether the deviation triggering condition is met. If the number of peak points is large and the total intensity is large, the deviation triggering condition may be met.

[0123] Step S145: Based on the audio branch switching condition, the video branch switching condition, the associated branch switching condition, and the deviation triggering condition, and in conjunction with the decision priority rule, a comprehensive decision is made. When at least one switching condition is met and a preset decision threshold is reached, a scene switching trigger signal is generated. In this embodiment, as an example, the audio branch switching condition, video branch switching condition, associated branch switching condition, and deviation triggering condition are comprehensively considered.

[0124] Deviation triggering conditions are combined with decision priority rules for comprehensive decision-making. The decision priority rules pre-set the priority of each condition. For example, in a specific audio-visual interaction scenario, the priority of the associated branch switching condition may be higher than the audio branch switching condition and the video branch switching condition, while the priority of the deviation triggering condition may be higher than the associated branch switching condition.

[0125] First, check whether each condition is met. If the audio branch switching condition is met, that is, the matching degree of the smoothed audio branch is lower than the preset audio matching degree threshold, then mark the audio branch as meeting the switching condition; if the video branch switching condition is met, that is, the matching degree of the smoothed video branch is lower than the preset video matching degree threshold, then mark the video branch as meeting the switching condition; if the associated branch switching condition is met, that is, the matching degree of the smoothed associated branch is lower than the preset associated matching degree threshold, then mark the associated branch as meeting the switching condition; if the deviation triggering condition is met, that is, the number and intensity of the deviation peak points reach the preset requirements, then mark the deviation triggering condition as meeting the requirement.

[0126] Then, the satisfied conditions can be prioritized according to decision priority rules. For example, if both the deviation trigger condition and the associated branch switching condition are satisfied simultaneously, and the deviation trigger condition has a higher priority, then the decision is made based on the deviation trigger condition first. A comprehensive score is calculated for each satisfied condition, based on its priority and corresponding matching degree or deviation value. When the comprehensive score reaches a preset decision threshold, a scene switching trigger signal is generated. This preset decision threshold is set based on a large amount of historical data and scene requirements to ensure that the signal is only triggered when the scene truly needs to switch, avoiding accidental switching.

[0127] Step S146: After the scene switching trigger signal is generated, a standard scene mode template that is least different from the current scene mode and meets the application requirements is selected from the scene mode template library, and the identifier of the standard scene mode template is determined as the target scene identifier.

[0128] In this embodiment, as an example, after generating the scene switching trigger signal, a suitable target scene mode template needs to be selected from the scene mode template library. First, the degree of difference between the current scene mode graph and each standard scene mode template in the scene mode template library is calculated. The degree of difference is calculated based on the feature differences between the audio mode branch, video mode branch, and audio-video association mode branch of the current scene mode graph and the corresponding branches of each standard scene mode template.

[0129] Specifically, the feature differences between the current audio mode branch and each standard audio mode branch, the feature differences between the current video mode branch and each standard video mode branch, and the feature differences between the current audio-video association mode branch and each standard association mode branch are calculated separately. Then, these three sums of differences are weighted and summed according to preset weights to obtain the comprehensive difference value between the current scene mode and each standard scene mode template.

[0130] Simultaneously, the system checks whether each standard scenario mode template meets the application requirements, including scenario adaptability and resource consumption. For example, in situations with limited network bandwidth, standard scenario mode templates with lower bandwidth requirements are prioritized. From the standard scenario mode templates that meet the application requirements, the template with the smallest overall difference value is selected; this template is the one with the smallest difference from the current scenario mode. Then, the identifier of this standard scenario mode template is extracted and designated as the target scenario identifier for subsequent scenario switching operations.

[0131] Step S147: Based on the position of the deviation peak point in the deviation feature sequence and the position of the lowest matching degree in the pattern matching degree sequence, determine the switching start timestamp, and generate the transition duration feature according to the switching start timestamp and the preset transition duration range.

[0132] In this embodiment, as an example, the position of the deviation peak point in the deviation feature sequence is first analyzed. The deviation peak point refers to the point where the deviation value exceeds a preset deviation threshold. This point reflects the moment when the current scene mode deviates significantly from the standard scene mode. Simultaneously, the position of the lowest matching degree in the mode matching degree sequence is analyzed. The lowest matching degree position refers to the moment when the audio branch matching degree, video branch matching degree, or associated branch matching degree is the lowest. At this moment, the degree of matching between the current scene mode and the standard scene mode is the lowest.

[0133] Taking into account both the peak deviation location and the lowest matching degree location, the time that best reflects the need for a scenario switch is selected as the switch start timestamp. For example, if a certain peak deviation location also corresponds to a low matching degree, that location is preferentially determined as the switch start timestamp.

[0134] The preset transition duration range is set based on the scenario type and user experience requirements. For example, in scenarios requiring a smooth transition, the transition duration range may be set slightly longer, while in scenarios with high real-time requirements, the transition duration range may be set shorter. Based on the switch start timestamp and the preset transition duration range, an appropriate transition duration can be selected within that range to generate transition duration features. The selection of the transition duration takes into account the changing trends of deviation features and matching degree features. If the deviation is gradually increasing and the matching degree is gradually decreasing, a slightly longer transition duration may be selected to ensure a smooth switch.

[0135] Step S148: Based on the fluctuation of the audio branch matching degree and the video branch matching degree, calculate the smoothness coefficient and use the smoothness coefficient as the transition smoothness feature.

[0136] In this embodiment, as an example, the fluctuation of audio branch matching degree is first analyzed. The degree of fluctuation is measured by calculating the amplitude and frequency of change of audio branch matching degree over a period of time. The larger the amplitude and the higher the frequency of change, the greater the fluctuation of audio branch matching degree; conversely, the smaller the fluctuation. Similarly, the fluctuation of video branch matching degree can be analyzed.

[0137] A smoothness coefficient is calculated based on the fluctuations in audio and video branch matching scores. During the calculation, different weights are assigned according to the degree of fluctuation in each factor. For example, if the fluctuation in video branch matching score has a greater impact on the smoothness of scene transitions, it will be given a higher weight. The smoothness coefficient is inversely proportional to the degree of fluctuation; that is, the smaller the fluctuation, the larger the smoothness coefficient, and vice versa.

[0138] The calculated smoothness coefficient is used as the transition smoothness feature, which guides subsequent audio and video transition processing. The larger the smoothness coefficient, the smoother the transition process needs to be to reduce the abruptness of the switch.

[0139] Step S149: Combine the transition duration feature and the transition smoothness feature into a switching transition parameter.

[0140] In this embodiment, as an example, transition duration features and transition smoothness features are combined according to a preset format to form a switching transition parameter. This parameter contains information on the transition duration and smoothness during scene switching, providing clear guidance for subsequent audio and video scene switching and optimization processing. For example, the switching transition parameter can be represented as a structure containing transition duration features and transition smoothness features, where the transition duration feature specifies the duration of the transition process, and the transition smoothness feature specifies the required smoothness of the transition process.

[0141] Step S150: Based on the scene switching trigger signal, the target scene identifier and the switching transition parameters, perform scene mode switching and optimization processing on the audio and video data stream to obtain the optimized audio and video output stream.

[0142] In this embodiment, as an example, upon receiving a scene switching trigger signal, the system retrieves the corresponding target scene pattern map from the scene pattern template library based on the target scene identifier, and processes the audio and video data streams in conjunction with the switching transition parameters. The processing includes transitioning the audio and video segments within the transition time period to achieve audio-video synchronization, and splicing the audio and video data streams from different time periods to obtain an optimized audio and video output stream. This output stream better adapts to the target scene and improves the user experience.

[0143] In this embodiment, step S150 may include sub-steps S151-S156, which will be described in detail below.

[0144] Step S151: When the scene switching trigger signal is received, the target scene identifier is parsed, and the target scene pattern map corresponding to the target scene identifier is obtained from the scene pattern template library. The target scene pattern map includes a target audio mode branch, a target video mode branch, and a target associated mode branch.

[0145] In this embodiment, as an example, upon receiving a scene switching trigger signal, the target scene identifier is immediately parsed to determine the target scene to which the user needs to switch. Then, by querying the scene pattern template library, the target scene pattern map corresponding to the target scene identifier is found.

[0146] The target audio mode branch in the target scene mode graph contains typical audio feature chains for the target scene. These feature chains are distributed along the time axis and include audio spectrum features, audio rhythm features, and audio semantic features. The target video mode branch contains typical video feature chains for the target scene, including video frame image features, motion trajectory features, and scene structure features. The target association mode branch contains association feature chains between audio and video in the target scene, including time synchronization features and content association features. These branches together constitute the mode features of the target scene, providing a reference for scene switching. Step S152: Analyze the switching transition parameters to obtain transition duration features and transition smoothness features. Determine the switching transition time period based on the transition duration features. The starting point of the switching transition time period is the switching start timestamp, and the length of the switching transition time period is determined by the transition duration features.

[0147] In this embodiment, as an example, the transition parameters can be parsed to extract transition duration and smoothness features. The transition duration feature specifies the required time length for the transition process, while the smoothness feature specifies the required smoothness of the transition process. The transition time period can be determined based on the transition duration feature and the previously determined transition start timestamp. The starting point of the transition time period is the transition start timestamp, and the ending point is the transition start timestamp plus the time length corresponding to the transition duration feature. For example, if the transition start timestamp is t5 and the time length corresponding to the transition duration feature is t_len, then the transition time period is from t5 to t5+t_len. During this time period, the audio and video data streams will undergo transition processing to achieve a smooth transition from the current scene mode to the target scene mode.

[0148] Step S153: Perform audio mode transition processing on the audio segments in the audio and video data stream that are in the switching transition time period. Based on the audio mode branch of the current scene mode map and the target audio mode branch, generate a transition audio feature chain through feature interpolation, and replace the original audio segment with the audio signal corresponding to the transition audio feature chain to obtain the transition audio data stream.

[0149] In this embodiment, as an example, audio mode transition processing is performed on audio segments within the switching transition time period. First, the audio feature chain corresponding to the switching transition time period in the audio mode branch of the current scene mode map and the audio feature chain corresponding to the switching transition time period in the target audio mode branch of the target scene mode map are extracted.

[0150] Then, a transitional audio feature chain is generated between the two audio feature chains using feature interpolation. During the feature interpolation process, the interpolation method is adjusted according to the transition smoothness feature. If the transition smoothness feature requirement is high, the interpolation process will pay more attention to the continuity of feature changes, so that the transitional audio feature chain smoothly transitions from the current audio feature chain to the target audio feature chain.

[0151] After the transition audio feature chain is generated, it is converted into the corresponding audio signal and used to replace the original audio segments in the audio / video data stream during the transition period, thus obtaining the transition audio data stream. The transition audio data stream enables a smooth transition from the current scene to the target scene at the audio level, reducing the discomfort caused by audio switching.

[0152] Step S154: Perform video mode transition processing on the video segments in the audio and video data stream that are in the switching transition time period. Based on the video mode branch of the current scene mode map and the target video mode branch, generate a transition video feature chain through feature interpolation. Replace the original video segment with the video signal corresponding to the transition video feature chain to obtain the transition video data stream.

[0153] In this embodiment, as an example, video mode transition processing can be performed on video segments within the switching transition time period. Similar to audio mode transition processing, the video feature chains corresponding to the switching transition time period are first extracted from the video mode branch of the current scene mode graph and the target video mode branch of the target scene mode graph.

[0154] Next, a transitional video feature chain is generated using feature interpolation. During the generation process, transition smoothness features are also considered to ensure a smooth transition from the current video feature chain to the target video feature chain. For example, for color distribution features in video frame image features, the interpolation process gradually transitions the color from its current distribution state to the target distribution state; for motion trajectory features, the direction and speed of the motion trajectory are gradually adjusted to the target state.

[0155] After the transition video feature chain is generated, it is converted into the corresponding video signal and used to replace the original video segments within the transition time period in the audio and video data stream, resulting in the transition video data stream. The transition video data stream achieves smooth scene switching at the video level.

[0156] Step S155: Perform audio and video synchronization optimization processing on the transition audio data stream and the transition video data stream, and adjust the synchronization deviation based on the transition smoothness characteristics so that the time synchronization characteristics during the transition process meet the preset synchronization requirements.

[0157] In this embodiment, as an example, audio and video synchronization optimization processing can be performed on the transitional audio data stream and the transitional video data stream. First, the time synchronization deviation between the transitional audio data stream and the transitional video data stream is detected, that is, the misalignment between the two in time.

[0158] Then, the synchronization deviation can be adjusted based on the transition smoothness characteristics. If a high degree of smoothness is required, the synchronization deviation can be adjusted more precisely by making small time offset adjustments to the audio or video data to ensure they are highly consistent in time. During the adjustment process, preset synchronization requirements are referenced, which define the maximum allowable deviation range for audio and video synchronization. In this way, through synchronization optimization, the time synchronization characteristics during the transition process meet the preset synchronization requirements, avoiding audio-visual asynchrony and ensuring a good user experience when watching the transition.

[0159] Step S156: The original audio and video data stream before the switching transition period, the transition audio data stream and the transition video data stream, and the target audio and video data stream after the switching transition period are spliced ​​together to obtain the optimized audio and video output stream.

[0160] In this embodiment, as an example, audio and video data streams from different time periods can be spliced ​​together in chronological order. First, the original audio and video data streams before the transition period are retained; this part of the data stream reflects the scene content before the switch. Then, the transition audio and video data streams are spliced ​​after the original audio and video data streams; this part of the data stream realizes the transition from the current scene to the target scene. Finally, the target audio and video data stream after the transition period is spliced ​​after the transition data stream; this part of the data stream reflects the target scene content after the switch.

[0161] During the splicing process, the timing between the various data streams was accurate and seamless, with no overlap or gaps. Simultaneously, audio and video features at the splicing points were smoothed to further reduce any abruptness. After splicing, an optimized audio and video output stream was obtained. This stream smoothly displays the transition from the current scene to the target scene and conforms to the pattern characteristics of the target scene, thus improving the quality of audio and video interaction.

[0162] like Figure 3The diagram shown is a schematic of an audio-visual scene intelligent switching optimization system combining pattern recognition, provided in an embodiment of this application. The audio-visual scene intelligent switching optimization system includes components such as a processor, a machine-readable storage medium, and input / output devices. The machine-readable storage medium is connected to the processor, and is used to store programs, instructions, or code. The processor is used to execute the programs, instructions, or code in the machine-readable storage medium to implement the aforementioned audio-visual scene intelligent switching optimization method combining pattern recognition. The audio-visual scene intelligent switching optimization system can be... Figure 2 The audio and video content provider or receiver in the embodiment may be itself or a component of the audio and video content provider or receiver. This embodiment does not impose specific limitations on these components.

[0163] The machine-readable storage medium may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), etc. The machine-readable storage medium is used to store a program, which the processor executes upon receiving an execution instruction.

[0164] The processor may be an integrated circuit chip with signal processing capabilities. The processor mentioned above can be, but is not limited to, a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.

[0165] In summary, the intelligent audio-visual scene switching optimization method and system combining pattern recognition provided in this application acquires real-time data streams of audio and video units sorted by timestamps, and constructs a current scene pattern graph based on a scene pattern template library. This graph includes audio mode branches, video mode branches, and audio-video association mode branches, achieving structured integration of multi-dimensional features of the player's audio and video data streams. This graph not only covers the spectrum, rhythm, and semantic features of audio (adapting to channel adjustment and sound effect optimization scenarios), and the frame images, motion trajectories, and scene structure features of video (adapting to aspect ratio and frame rate compensation scenarios), but also includes time synchronization and content association features between audio and video (adapting to audio-visual collaborative optimization scenarios). Compared to the switching methods of existing players that rely on single or simple feature combinations, this method can more comprehensively depict the essence of the playback scene, providing a reliable feature foundation for subsequent switching decisions and solving the problem of scene misjudgment caused by feature fragmentation in traditional players.

[0166] In the dynamic pattern matching stage, the current scene pattern map is compared with the standard pattern map, and the matching degree and deviation feature sequence of audio, video, and related branches are calculated separately to achieve multi-dimensional and fine-grained comparison. For example, audio branch matching combines spectrum, rhythm, and semantic similarity to accurately identify the differences between "voice scenes" and "sound effect scenes"; video branch matching combines frame images, motion trajectories, and scene structure consistency to accurately distinguish between "static scenes" and "dynamic scenes"; related branch matching focuses on time synchronization accuracy and content association strength to avoid erroneous switching in audio-visual desynchronization scenarios. This multi-branch parallel matching method effectively resists the influence of feature deviations caused by network fluctuations during playback and image fluctuations of older resources on single features, significantly improving the accuracy of scene pattern matching for the player.

[0167] Furthermore, based on pattern matching degree and deviation characteristics, and combined with a scene switching decision rule library, switching trigger signals, target scene identifiers, and transition parameters are generated, overcoming the limitations of fixed threshold decision-making in existing players. By smoothly processing matching degree through a sliding window and identifying significant deviations through peak detection, and dynamically adjusting the decision threshold in conjunction with preset priority rules, the switching timing is made more aligned with playback requirements. For example, during HDR scene switching, it is more sensitive to deviations in video frame image features, and can quickly trigger color parameter adaptation; during the switching transition phase, a transition audio-visual feature chain is generated through feature interpolation based on transition parameters, and synchronization is optimized to avoid abrupt changes in picture and sound caused by fixed transition parameters. Thus, through the synergistic effect of multi-dimensional feature integration, dynamic and accurate matching, intelligent decision-making, and smooth transition, the audio-visual player achieves accuracy, timeliness, and smoothness in switching across multiple scenarios, meeting the high requirements of intelligent scene switching in complex playback scenarios.

[0168] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The embodiments, implementation methods, and related technical features of this application can be combined and substituted with each other without conflict. The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. An audio-video scene intelligent switching optimization method combined with pattern recognition, characterized in that, The method comprises: acquiring a real-time collected audio and video data stream, the audio and video data stream containing continuous audio and video units arranged in chronological order, each audio and video unit containing synchronized audio and video fragments; based on a preset scene mode template library, performing mode feature analysis processing on the audio and video data stream to construct a current scene mode graph, the current scene mode graph containing audio mode branches, video mode branches and audio-video associated mode branches, the audio mode branches being composed of audio feature chains, the video mode branches being composed of video feature chains, and the audio-video associated mode branches being composed of associated feature chains; performing dynamic mode matching processing on the current scene mode graph and standard mode graphs in the scene mode template library to calculate a mode matching degree sequence and a deviation feature sequence, the mode matching degree sequence containing audio branch matching degrees, video branch matching degrees and associated branch matching degrees, and the deviation feature sequence containing audio deviation features, video deviation features and associated deviation features; based on the mode matching degree sequence and the deviation feature sequence, combining a preset scene switching decision rule library to generate a scene switching trigger signal, a target scene identifier and a switching transition parameter; performing scene mode switching and optimization processing on the audio and video data stream according to the scene switching trigger signal, the target scene identifier and the switching transition parameter to obtain an optimized audio and video output stream.

2. The method for intelligent switching optimization of audio-video scene with pattern recognition of claim 1, wherein, The method comprises: acquiring the scene mode template library, the scene mode template library containing at least one standard scene mode template, each standard scene mode template containing standard audio mode branches, standard video mode branches and standard associated mode branches; performing audio and video separation processing on the audio and video data stream to obtain independent audio and video data streams; performing audio mode feature analysis processing on the audio data stream to obtain an audio feature chain, taking the audio feature chain as the audio mode branches of the current scene mode graph, the audio feature chain containing audio spectrum features, audio rhythm features and audio semantic features distributed along a time axis; performing video mode feature analysis processing on the video data stream to obtain a video feature chain, taking the video feature chain as the video mode branches of the current scene mode graph, the video feature chain containing video frame image features, motion trajectory features and scene structure features distributed along a time axis; performing audio-video associated mode feature analysis processing on the audio and video data streams to obtain an associated feature chain, taking the associated feature chain as the audio-video associated mode branches of the current scene mode graph, the associated feature chain containing time synchronization features and content associated features distributed along a time axis; performing graph structuring processing on the audio mode branches, the video mode branches and the audio-video associated mode branches to generate a current scene mode graph containing node layers, connection edges and weight parameters, the node layers corresponding to feature units in the feature chains.

3. The method for intelligent switching optimization of audio-video scene with pattern recognition of claim 2, wherein, The audio mode feature analysis processing on the audio data stream obtains an audio feature chain, and the audio feature chain is taken as an audio mode branch of a current scene mode graph, and the audio mode branch comprises: Spectrum conversion processing is performed on each audio segment in the audio data stream to convert a time-domain audio signal into frequency-domain spectrum data, and audio spectrum features are obtained by performing feature extraction on the frequency-domain spectrum data; Rhythm feature extraction processing is performed on the audio data stream to identify a beat starting point and a beat interval in the audio signal, and rhythm intensity parameters and rhythm stability parameters are calculated based on the beat starting point and the beat interval, and the rhythm intensity parameters and the rhythm stability parameters are combined into audio rhythm features; Voice activity detection processing is performed on the audio data stream to separate audio segments containing voice and audio segments not containing voice, and text information is obtained by performing voice recognition processing on the audio segments containing voice; Semantic analysis processing is performed on the text information to extract theme keywords and semantic intention features, and the theme keywords and the semantic intention features are combined into audio semantic features; The audio spectrum features, the audio rhythm features, and the audio semantic features are combined in a time stamp order to obtain an audio feature chain; The audio feature chain is taken as an audio mode branch of a current scene mode graph, and each node in the audio mode branch corresponds to an audio feature combination of a time stamp.

4. The method for intelligent switching optimization of audio-video scene with pattern recognition of claim 2, wherein, The video mode feature analysis processing on the video data stream obtains a video feature chain, and the video feature chain is taken as a video mode branch of a current scene mode graph, and the video mode branch comprises: Frame processing is performed on each video segment in the video data stream to obtain a continuous video frame sequence, and image feature extraction processing is performed on each video frame to obtain video frame image features; Optical flow estimation processing is performed on adjacent video frames in the video frame sequence to calculate pixel-level motion vectors, target tracking processing is performed based on the motion vectors to obtain a trajectory coordinate sequence of a moving target, and moving trajectory features are obtained by performing feature extraction on the trajectory coordinate sequence; Scene segmentation processing is performed on key frames in the video frame sequence to divide the key frames into a plurality of semantic regions, spatial relationship analysis is performed on the semantic regions to calculate region area proportions and region position relationships, and the region area proportions and the region position relationships are combined into scene structure features; The video frame image features, the moving trajectory features, and the scene structure features are combined in a time stamp order to obtain a video feature chain; The video feature chain is taken as a video mode branch of a current scene mode graph, and each node in the video mode branch corresponds to a video feature combination of a time stamp.

5. The method for audio-video scene intelligent switching optimization with pattern recognition of claim 2, wherein, The audio-video association mode feature analysis processing on the audio data stream and the video data stream obtains an association feature chain, and the association feature chain is taken as an audio-video association mode branch of a current scene mode graph, and the audio-video association mode branch comprises: extracting an audio timestamp sequence of the audio data stream and a video timestamp sequence of the video data stream, performing time alignment processing on the audio timestamp sequence and the video timestamp sequence, calculating a time offset value and a synchronization fluctuation parameter, and combining the time offset value and the synchronization fluctuation parameter into a time synchronization feature; performing content matching processing on the audio semantic feature and the video frame image feature, calculating a semantic correlation degree parameter, performing dynamic correlation processing on the audio rhythm feature and the motion trajectory feature, calculating a rhythm matching degree parameter, and combining the semantic correlation degree parameter and the rhythm matching degree parameter into a content correlation feature; combining the time synchronization feature and the content correlation feature in a timestamp order to obtain a correlation feature chain; combining the correlation feature chain as an audio-video correlation mode branch of a current scene mode graph, each node in the audio-video correlation mode branch corresponding to an audio-video correlation feature combination of a timestamp; based on the nodes of the corresponding timestamps in the audio mode branch, the video mode branch, and the audio-video correlation mode branch, constructing a cross-branch correlation edge, and a weight parameter of the cross-branch correlation edge being determined based on a correlation strength of a corresponding node feature.

6. The method for intelligent cut optimization of audio-video scene with pattern recognition of claim 1, wherein, The dynamic mode matching processing of the current scene mode graph and the standard mode graph in the scene mode template library to calculate a mode matching degree sequence and a deviation feature sequence includes: selecting at least one standard scene mode template corresponding to a current application scene from the scene mode template library, extracting a standard scene mode graph in the standard scene mode template, and the standard scene mode graph including a standard audio mode branch, a standard video mode branch, and a standard correlation mode branch; performing feature comparison processing on the audio mode branch of the current scene mode graph and the standard audio mode branch, calculating an audio feature similarity of a corresponding timestamp node, arranging the audio feature similarity in a time axis to obtain an audio branch matching degree; performing feature comparison processing on the video mode branch of the current scene mode graph and the standard video mode branch, calculating a video feature similarity of a corresponding timestamp node, arranging the video feature similarity in a time axis to obtain a video branch matching degree; performing feature comparison processing on the audio-video correlation mode branch of the current scene mode graph and the standard correlation mode branch, calculating a correlation feature similarity of a corresponding timestamp node, arranging the correlation feature similarity in a time axis to obtain a correlation branch matching degree; combining the audio branch matching degree, the video branch matching degree, and the correlation branch matching degree into a mode matching degree sequence; calculating a feature difference amount of the audio mode branch of the current scene mode graph and the standard audio mode branch to obtain an audio deviation feature, calculating a feature difference amount of the video mode branch of the current scene mode graph and the standard video mode branch to obtain a video deviation feature, and calculating a feature difference amount of the audio-video correlation mode branch of the current scene mode graph and the standard correlation mode branch to obtain a correlation deviation feature; Combine the audio deviation feature, the video deviation feature and the correlation deviation feature into a deviation feature sequence.

7. The method for intelligent cut optimization of audio-video scene with pattern recognition of claim 1, wherein, The scene switching trigger signal, the target scene identifier and the switching transition parameter are generated based on the mode matching degree sequence and the deviation feature sequence in combination with a preset scene switching decision rule library, including: The scene switching decision rule library is obtained, and the scene switching decision rule library includes a matching degree threshold rule, a deviation threshold rule and a decision priority rule; The audio branch matching degree, the video branch matching degree and the correlation branch matching degree in the mode matching degree sequence are subjected to sliding window average processing to obtain a smoothed audio branch matching degree, a smoothed video branch matching degree and a smoothed correlation branch matching degree; The smoothed audio branch matching degree is compared with a preset audio matching degree threshold to determine whether an audio branch switching condition is met, the smoothed video branch matching degree is compared with a preset video matching degree threshold to determine whether a video branch switching condition is met, and the smoothed correlation branch matching degree is compared with a preset correlation matching degree threshold to determine whether a correlation branch switching condition is met; The audio deviation feature, the video deviation feature and the correlation deviation feature in the deviation feature sequence are subjected to peak detection processing to identify a deviation peak point exceeding a preset deviation threshold, and it is determined whether a deviation trigger condition is met based on the number and intensity of the deviation peak point; Based on the audio branch switching condition, the video branch switching condition, the correlation branch switching condition and the deviation trigger condition, comprehensive decision is made in combination with the decision priority rule, and when at least one switching condition is met and a preset decision threshold is reached, a scene switching trigger signal is generated; When the scene switching trigger signal is generated, a standard scene mode template with the smallest difference from the current scene mode and meeting application requirements is selected from the scene mode template library, and the identifier of the standard scene mode template is determined as the target scene identifier; Based on the deviation peak point position in the deviation feature sequence and the lowest matching degree position in the mode matching degree sequence, a switching start timestamp is determined, and a transition duration feature is generated according to the switching start timestamp and a preset transition duration range; Based on the fluctuation of the audio branch matching degree and the video branch matching degree, a smoothing degree coefficient is calculated, and the smoothing degree coefficient is taken as a transition smoothing degree feature; The transition duration feature and the transition smoothing degree feature are combined into a switching transition parameter.

8. The method for intelligent audio-video scene cut optimization with pattern recognition of claim 1, wherein, The scene mode switching and optimization processing is performed on the audio and video data stream according to the scene switching trigger signal, the target scene identifier and the switching transition parameter to obtain an optimized audio and video output stream, including: When the scene switching trigger signal is received, the target scene identifier is analyzed, and a target scene mode graph corresponding to the target scene identifier is obtained from the scene mode template library, and the target scene mode graph includes a target audio mode branch, a target video mode branch and a target correlation mode branch; The switching transition parameter is parsed to obtain a transition duration feature and a transition smoothness feature, a switching transition time period is determined based on the transition duration feature, a starting point of the switching transition time period is a switching starting timestamp, and a length of the switching transition time period is determined by the transition duration feature; The audio mode transition processing is performed on the audio segment in the audio and video data stream within the switching transition time period, the transition audio feature chain is generated by feature interpolation based on the audio mode branch of the current scene mode graph and the target audio mode branch, the audio signal corresponding to the transition audio feature chain is used to replace the original audio segment, and a transition audio data stream is obtained; The video mode transition processing is performed on the video segment in the audio and video data stream within the switching transition time period, the transition video feature chain is generated by feature interpolation based on the video mode branch of the current scene mode graph and the target video mode branch, the video signal corresponding to the transition video feature chain is used to replace the original video segment, and a transition video data stream is obtained; The audio and video synchronization optimization processing is performed on the transition audio data stream and the transition video data stream, and the synchronization deviation is adjusted based on the transition smoothness feature; The original audio and video data stream before the switching transition time period, the transition audio data stream and the transition video data stream, and the target audio and video data stream after the switching transition time period are spliced to obtain an optimized audio and video output stream.

9. An audio-video scene intelligent switching optimization system combined with pattern recognition, characterized in that, The processor, the machine readable storage medium, the machine readable storage medium and the processor are connected, the machine readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine readable storage medium to realize the audio and video scene intelligent switching optimization method combined with mode recognition in any one of claims 1-8.

10. A computer storage medium, characterized in that, The program, instruction or code is stored, and when the program, instruction or code is executed, the audio and video scene intelligent switching optimization method combined with mode recognition in any one of claims 1-8 is realized.

Citation Information

Patent Citations

  • Intelligent video source switching system and method for director switching station

    CN119316539A

  • Intelligent MV generation method, system and device based on AIGC and medium

    CN119788886A