Video clip automatic clipping and splicing method and system applied to digital multimedia

By extracting the narrative logic anchors of video clips and constructing the semantic flow of the images, the problem of low efficiency and poor quality in video editing and splicing in existing technologies is solved, and efficient and coherent video content generation is achieved.

CN121531210BActive Publication Date: 2026-03-27SHANGHAI MINGQI NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing video editing and splicing methods rely on manual operation, which is inefficient and easily affected by individual skill levels and subjective factors, making it difficult to construct high-quality video content that conforms to narrative logic.

Method used

By acquiring a set of video clips, extracting narrative logic anchors, constructing an overall semantic flow of the video, and generating splicing adaptation rules based on the narrative logic anchors, transition and fusion processing is performed to ensure the narrative logic coherence and semantic consistency of the video content.

Benefits of technology

It improves the efficiency and effectiveness of video editing and splicing, ensuring that the video content is coherent in narrative logic and consistent in visual semantics, generating high-quality, entertaining video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531210B_ABST
    Figure CN121531210B_ABST
Patent Text Reader

Abstract

The application provides a video segment automatic clipping and splicing method and system applied to digital multimedia, relates to the technical field of digital multimedia, and first acquires a video segment set to be clipped, extracts a narrative logic anchor point for each video segment unit to be clipped, and constructs an overall picture semantic flow; generates a splicing adaptation rule for adjacent segment units, performs transition fusion processing to obtain a video segment sequence; finally, the video segment sequence with the fusion transition effect is subjected to integration processing according to the arrangement order of the overall picture semantic flow, a complete digital multimedia video after clipping and splicing is obtained, the video narrative logic can be accurately grasped, a coherent picture semantic flow is constructed, the video after splicing performs well in narrative coherence and picture consistency, and the efficiency and quality of video clipping and splicing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital multimedia, in particular to a video segment automatic clipping and splicing method and system applied to digital multimedia. BACKGROUND

[0002] In the current booming development of digital multimedia, the creation and dissemination of video content are increasingly widespread. Whether it is in the fields of film and television production, advertising, or short video creation, a large amount of video clipping and splicing work is involved. However, the existing video clipping and splicing methods have many shortcomings.

[0003] Traditional video clipping and splicing methods often rely on manual operation. The editor needs to select appropriate segments from a vast amount of video materials based on their own experience and subjective judgment, and manually adjust the order and transition effects between segments. The above method is not only inefficient, but also easily affected by the individual level and subjective factors of the editor, resulting in incoherent narrative logic and unsmooth picture semantics in the edited video.

[0004] Some automatic video clipping and splicing methods can improve efficiency to some extent, but they are usually based on simple rules such as time sequence, picture similarity, etc. for splicing, lacking deep understanding and analysis of video content. The above methods cannot accurately grasp the core content direction of video segments in the overall narrative, and it is difficult to build a picture semantic flow that meets the narrative logic, making the spliced video perform poorly in coherence and consistency, and unable to meet the user's demand for high-quality video content. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a video segment automatic clipping and splicing method and system applied to digital multimedia.

[0006] According to the first aspect of the present application, a video segment automatic clipping and splicing method applied to digital multimedia is provided, which comprises:

[0007] Obtaining a set of digital multimedia video segments to be clipped, which contains a plurality of video segment units with picture content elements, audio content elements and timeline markers;

[0008] performing a narrative logic anchor extraction process on each video segment unit in the set of digital multimedia video segments to be edited, to obtain a narrative logic anchor corresponding to each video segment unit, wherein the narrative logic anchor extraction process comprises: splitting a picture content element of the video segment unit to obtain picture content sub-elements including a character action, a scene environment, and an object form; performing time sequence and spatial analysis on the character action sub-element, the scene environment sub-element, and the object form sub-element respectively to obtain a corresponding character action association group, a scene environment association group, and an object form association group; performing audio feature extraction on an audio content element of the video segment unit, and matching and associating the audio features according to the picture content sub-elements to generate an audio association group; and generating the narrative logic anchor based on a fusion result of the character action association group, the scene environment association group, the object form association group, and the audio association group, the narrative logic anchor being used to represent a core content direction of the video segment unit in the overall narrative.

[0009] performing picture semantic stream construction processing based on the narrative logic anchors corresponding to all video segment units to obtain an overall picture semantic stream corresponding to the set of digital multimedia video segments to be edited, the overall picture semantic stream arranging the video segment units according to the association relationship of the narrative logic anchors;

[0010] performing splicing adaptation rule generation processing on adjacent video segment units in the overall picture semantic stream to obtain a splicing adaptation rule corresponding to the adjacent video segment units, the splicing adaptation rule matching picture semantic stream features and audio semantic stream features of the adjacent video segment units;

[0011] performing transition fusion processing on the adjacent video segment units based on the splicing adaptation rule to obtain a video segment sequence with fusion transition effects;

[0012] performing integration processing on the video segment sequence with fusion transition effects according to the arrangement order of the overall picture semantic stream to obtain a complete edited and spliced digital multimedia video, the complete edited and spliced digital multimedia video maintaining the coherence of the narrative logic anchors and the consistency of the picture semantic stream.

[0013] According to a second aspect of the present application, a video segment automatic editing and splicing system applied to digital multimedia is provided, the video segment automatic editing and splicing system applied to digital multimedia comprising a machine-readable storage medium and a processor, the machine-readable storage medium storing machine-executable instructions, and the processor, when executing the machine-executable instructions, implements the aforementioned video segment automatic editing and splicing method applied to digital multimedia.

[0014] According to a third aspect of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores computer executable instructions, when the computer executable instructions are executed, the foregoing video segment automatic editing and splicing method applied to digital multimedia is implemented.

[0015] According to any one of the above aspects, the technical effect of the present application is that:

[0016] First, a set of video segment units with picture, audio content elements and timeline markers are obtained, by performing narrative logic anchor extraction processing on each video segment unit, the core content direction of each segment in the overall narrative is located, based on the overall picture semantic stream constructed by the narrative logic anchor, the video segment units are arranged according to the association relationship, so that the video content is more coherent and ordered in the narrative, and the splicing adaptation rules are generated for adjacent video segment units, matching the picture and audio semantic stream features, fully considering the coherence of video in two aspects of vision and hearing, avoiding the problems of abrupt picture or discordant audio caused by improper splicing. The transition fusion processing based on the splicing adaptation rules makes the transition between adjacent video segments more natural and smooth, further improving the overall quality of the video. Finally, the video segment sequence with fusion transition effect is integrated in the order of the overall picture semantic stream, and the complete edited and spliced digital multimedia video is obtained, which not only maintains the coherence of the narrative logic anchor, but also ensures the consistency of the picture semantic stream, can present high-quality, logical and enjoyable video content to the user, and greatly improves the efficiency and effect of video editing and splicing. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 The flowchart of the video segment automatic editing and splicing method applied to digital multimedia provided by the embodiment of the present application is shown;

[0018] Figure 2 The component structure diagram of the video segment automatic editing and splicing system applied to digital multimedia provided by the embodiment of the present application is shown. DETAILED DESCRIPTION

[0019] Figure 1 The flowchart of the video segment automatic editing and splicing method and system applied to digital multimedia provided by the embodiment of the present application is shown, and the detailed steps include:

[0020] Step S110: obtaining a set of digital multimedia video segments to be edited, the set of digital multimedia video segments to be edited contains a plurality of video segment units with picture content elements, audio content elements and timeline markers.

[0021] In this embodiment, the application scenario is the automatic editing and splicing of food making tutorial videos in the catering field. The digital multimedia video segments to be edited are derived from raw materials recorded by a professional filming team in a kitchen environment, including food processing, cooking operations, and finished product display. Each video segment unit contains picture content elements captured by a 4K resolution camera, environmental sound effects and explanation audio content elements recorded by a professional microphone, and timeline markers generated based on time codes, which are accurate to the capture time of each frame of picture. During acquisition, the raw materials are uniformly formatted, the picture content elements are converted to H.265 encoding format, the audio content elements are converted to PCM lossless encoding format, and the timeline markers are in SMPTE standard time code format. For possible privacy-sensitive data such as the faces and voices of the filming personnel, a face blurring algorithm based on deep learning is used to perform pixel-level blurring on the face area in the picture, and a voiceprint replacement technology is used to replace the original human voice in the explanation audio with synthetic speech, ensuring that the privacy data is not leaked.

[0022] Step S120: Perform narrative logic anchor point extraction processing on each video segment unit in the set of digital multimedia video segments to be edited, to obtain the narrative logic anchor point corresponding to each video segment unit, which represents the core content direction of the video segment unit in the overall narrative.

[0023] In the catering food making tutorial scenario, the narrative logic anchor point represents the content direction of the key nodes of the cooking process, such as the completion of food preparation, the start of core cooking steps, and the key actions of seasoning and proportioning. By performing multi-dimensional feature extraction and association processing on each video segment unit, the scattered picture and audio information is condensed into an anchor point data structure with clear narrative direction.

[0024] Step S121: Perform splitting processing on the picture content elements of each video segment unit to obtain a set of picture content sub-elements, including character action sub-elements, scene environment sub-elements, and object form sub-elements.

[0025] The picture content splitting model based on the convolutional neural network is adopted to perform feature extraction on each frame of the video segment unit. The human action sub-element is obtained by a human pose estimation network, which includes 17 key skeleton point detection layers and can output three-dimensional coordinate information of human joints such as shoulders, elbows, wrists, hips, knees and ankles, and further determine the action type through joint motion trajectory; the scene environment sub-element is extracted by a semantic segmentation network, which includes a 512-channel feature extraction layer and a multi-scale fusion layer, and can identify 20 types of scene elements such as stoves, sinks and chopping boards in the kitchen scene, and output the bounding box coordinates and pixel ratio of the elements; the object shape sub-element is obtained by a target detection network, which adopts the YOLOv5 architecture and includes a CSPDarknet53 backbone network and a PANet feature fusion layer, and can identify 50 types of cooking-related objects such as knives, pots and food materials, and output the class label, bounding box coordinates and confidence parameter of the object. Through the above processing, the picture content elements are divided into a set of three types of sub-elements, and each type of sub-element includes a feature parameter matrix and timestamp information.

[0026] Step S122: performing continuous frame analysis processing on the human action sub-element in the picture content sub-element set, extracting the continuous action sequence of the human action sub-element in the video segment unit timeline, and arranging the continuous action sequence in the order of timeline marking.

[0027] The continuous frame analysis of the human action sub-element adopts the combination of optical flow estimation and action classification. First, the pixel motion vector between adjacent frames is calculated by the TV-L1 optical flow algorithm to generate an optical flow field matrix, which records the horizontal and vertical motion speed of each pixel point in the picture. Based on the optical flow field matrix, the motion trajectory of the human key skeleton points is tracked, and the displacement and motion direction angle of the skeleton points between each frame are calculated. The motion parameters of the skeleton points in the continuous 30 frames are input into the pre-trained action classification model, which includes 3 layers of LSTM recurrent neural network layers, each layer includes 256 hidden units, adopts dropout regularization technology to prevent overfitting, and outputs classification probabilities of 12 types of cooking actions such as cutting, frying and stirring. The action categories with classification probabilities higher than the preset threshold (0.85) are arranged in order according to the timeline marking sequence to form a continuous action sequence, and each action element in the sequence includes the starting frame index, the ending frame index and the action category label.

[0028] Step S123: performing action feature association processing on the continuous action sequence to combine action features with causal association relationship to form a human action association group, and the causal association relationship is determined based on the logical order of action occurrence and the purpose of action.

[0029] Step S1231: Number each action feature in the continuous action sequence in time line label order to form an ordered action feature list.

[0030] A unique identifier is assigned to each action feature in the continuous action sequence, which is composed of a video segment unit ID and an action sequence number, for example, "V001-A003" represents the 3rd action feature in the 1st video segment unit. In the order of time line labels, the action features are arranged in ascending order of the starting frame index to form an ordered action feature list, which is stored in a double-linked list data structure, and each node contains the category label, time interval, and bone point motion parameters of the action feature.

[0031] Step S1232: Select the first action feature in the ordered action feature list as the starting action feature, and analyze the action purpose direction of the starting action feature, which is determined based on the limb motion direction and motion amplitude of the action.

[0032] The action purpose direction is determined by calculating the motion direction cosine value and displacement amplitude vector of the key bone points in the starting action feature. For example, in the "cutting vegetables" action, the wrist joint relative to the shoulder joint motion direction cosine value in the Z axis (vertical direction) is 0.7, the X axis (horizontal left-right direction) is 0.2, and the Y axis (forward-backward direction) is 0.1, and the displacement amplitude vector length is 0.4 meters, combined with the relative position relationship between the knife and the food material, it is judged that the purpose of the action is "cutting food shape".

[0033] Step S1233: Select the first action feature after the starting action feature from the ordered action feature list as the to-be-associated action feature, and analyze the action purpose direction of the to-be-associated action feature.

[0034] The same analysis method as step S1232 is used to calculate the limb motion direction and amplitude parameters of the to-be-associated action feature to determine its action purpose direction. For example, in the "cutting vegetables" action, the wrist joint motion direction cosine value in the X axis is 0.6, the Y axis is 0.3, and the Z axis is 0.1, and the displacement amplitude vector length is 0.3 meters, and the action purpose direction is "food material position transfer".

[0035] Step S1234: Compare the action purpose direction of the starting action feature with the action purpose direction of the to-be-associated action feature to determine whether the starting action feature and the to-be-associated action feature have a causal association relationship. If the action purpose direction of the to-be-associated action feature is a continuation or result of the action purpose direction of the starting action feature, then the starting action feature and the to-be-associated action feature have a causal association relationship.

[0036] The action purpose direction correlation matrix is constructed, the row dimension of the matrix is the starting action purpose type, the column dimension is the to-be-correlated action purpose type, and the matrix element value represents the correlation strength of the two purpose types. When the correlation strength value is greater than a preset threshold (0.7), it is determined that there is a causal correlation relationship. For example, the correlation strength value of the "food material form cutting" purpose direction and the "food material position transfer" purpose direction is 0.82, so it is determined that the "cutting vegetables" action and the "arranging dishes" action have a causal correlation relationship; the correlation strength value of the "cutting vegetables" action and the "stove ignition" action (purpose direction "heat source activation") is 0.23, and it is determined that there is no causal correlation relationship.

[0037] Step S1235: If the starting action feature and the to-be-correlated action feature have a causal correlation relationship, the starting action feature and the to-be-correlated action feature are combined to form a temporary correlation group, and the to-be-correlated action feature is taken as a new starting action feature.

[0038] The action feature nodes with a causal correlation relationship are connected through a pointer to form a temporary correlation group data structure. The temporary correlation group includes an ordered linked list of action features in the group, a total time interval, a composite action purpose direction, and the like. For example, the "cutting vegetables" and "arranging dishes" action features are combined into a "food material pretreatment-transfer" temporary correlation group, and the new starting action feature is updated to the "arranging dishes" action feature.

[0039] Step S1236: If the starting action feature and the to-be-correlated action feature do not have a causal correlation relationship, the starting action feature is taken as a temporary correlation group alone, and the to-be-correlated action feature is taken as a new starting action feature.

[0040] When the correlation strength value is less than or equal to the preset threshold, the current starting action feature is encapsulated as an independent temporary correlation group and stored in the character action correlation group set, and the to-be-correlated action feature is set as a new starting action feature. The processing procedures of steps S1232 to S1235 are repeatedly executed.

[0041] Step S1237: The above operations of selecting a to-be-correlated action feature, analyzing an action purpose direction, judging a causal correlation relationship, and combining a temporary correlation group are repeated until all action features in the ordered action feature list are processed, and a plurality of temporary correlation groups are obtained.

[0042] By traversing the bidirectional linked list structure of the ordered action feature list, the correlation judgment and combination operation are performed on each action feature in turn until the tail node of the linked list is processed. In the traversal process, a recursive call is used to process the nested correlation relationship, for example, a three-level correlation action sequence of "peeling-cutting-marinating", which is combined into a temporary correlation group through three times of recursion.

[0043] Step S1238: Timeline continuity check processing is performed on the action features in each temporary association group to determine whether the action features in the temporary association group are continuous in the timeline. If so, the structure of the temporary association group is maintained.

[0044] The difference between the end frame index of the first action feature and the start frame index of the second action feature in the temporary association group is checked. If the difference is equal to 1 (indicating that the two actions are continuous in the timeline) or less than a preset maximum interval frame number (5 frames) and there are no other action features in the interval frame pictures, it is determined that the timeline is continuous. For example, if the end frame of the "chopping" action is frame 120 and the start frame of the "arranging" action is frame 121, it is determined that the timeline is continuous.

[0045] Step S1239: If the timeline is not continuous, the action features between the non-continuous action features are searched for in the ordered action feature list. If there are missing action features, the missing action features are supplemented to the temporary association group to make the action features in the temporary association group continuous in the timeline.

[0046] When the interval frame number between the action features is greater than the preset maximum interval frame number, the ordered action feature list is traced back to check whether there are unassociated action features. For example, there is an "igniting" action (frames 101-149) between the "chopping" action (frames 1-100) and the "stir-frying" action (frames 150-200) that is not associated. The "igniting" action is supplemented to the temporary association group by timeline index comparison to form a continuous action sequence of "chopping-igniting-stir-frying".

[0047] Step S12310: If there are no missing action features, the non-continuous temporary association group is split into two independent temporary association groups, and the action features in each temporary association group are continuous in the timeline.

[0048] When the interval frame region is confirmed to have no missing action features through searching, it is determined that the action sequence is naturally interrupted, and the temporary association group splitting operation is performed. The splitting point is set at the start position of the interval frame, and the original temporary association group is divided into two independent temporary association groups in the timeline, which are stored in the association group set.

[0049] Step S12311: Action logic integrity analysis processing is performed on all processed temporary association groups to determine whether the action features in each temporary association group can express an independent action logic. If so, the temporary association group is determined as a character action association group.

[0050] An action logic integrity evaluation model is constructed, which includes three evaluation dimensions of action quantity threshold, time span threshold, and goal direction consistency. When the number of action characteristics in the temporary association group is not less than 3, the time span is not less than 5 seconds, and the average value of the association strength of all action goals is greater than 0.65, it is determined that the action logic is complete. For example, the temporary association group of "food cleaning - peeling - cutting", which contains 3 action characteristics, has a time span of 12 seconds, and an average value of goal direction association strength of 0.78, is determined to be a complete action logic.

[0051] Step S12312: If not, the adjacent temporary association groups with continuous action logic are combined to form a new temporary association group, and the action logic integrity analysis process is performed again until the action logic of the independent action logic is obtained.

[0052] For the temporary association group with incomplete action logic, the goal direction association strength with the adjacent temporary association group is calculated, and the two temporary association groups with the highest association strength are combined. For example, the association strength between "food cleaning" (1 action) and "peeling - cutting" (2 actions) is 0.81, and after combination, a new temporary association group containing 3 actions is formed, and the integrity evaluation is performed again.

[0053] Step S12313: All obtained character action association groups are sorted to form a character action association group set corresponding to each video segment unit, which is used for subsequent generation of narrative logic anchor points.

[0054] The character action association group set is stored in JSON data format, and each association group contains four fields of group ID, action characteristic list, time interval, and core goal direction. In the catering scene example, a character action association group set may contain "meat marinating association group", "vegetable slicing association group", "sauce mixing association group", etc., and each association group corresponds to a sub-task unit in the cooking process.

[0055] Step S124: Perform spatial feature extraction processing on the scene environment sub-element in the picture content sub-element set to obtain the spatial layout feature and scene atmosphere feature of the scene environment sub-element. The spatial layout feature represents the position distribution relationship of the elements in the scene, and the scene atmosphere feature represents the visual style tendency of the scene.

[0056] The scene feature extraction network based on the Transformer architecture is adopted to perform spatial feature extraction on the key frames (extract one frame every 10 frames) of the video segment unit. The spatial layout feature is constructed by calculating the center point coordinates, aspect ratio, and area ratio of the bounding rectangle of the scene elements to form a 128-dimensional feature vector, such as the center point coordinates (x1, y1) of the cooking stove element, the area ratio 0.35, the center point coordinates (x2, y2) of the sink element, and the area ratio 0.20, which represent the position distribution through the relative distance and angle relationship between elements. The scene atmosphere feature is constructed by extracting the mean value, contrast, saturation, brightness, and other parameters of the picture in the HSV color space to form a 64-dimensional feature vector, such as the warm color tone ratio 0.7, the contrast 0.6, and the average brightness 0.55, which represent the visual style tendency through color and light parameters.

[0057] Step S125: Perform association processing on the spatial layout feature and the scene atmosphere feature, combine the elements with fixed position association in the spatial layout feature with the corresponding style tendency in the scene atmosphere feature to form a scene environment association group.

[0058] A scene element position association matrix is constructed, and the matrix elements represent the position association degree of any two scene elements. When the association degree is greater than 0.6, it is determined to have a fixed position association. For example, the position association degree of the cooking stove and the exhaust hood is 0.92, which is a fixed position associated element. At the same time, an atmosphere style tendency mapping table is constructed, which maps the element type in the spatial layout to the corresponding atmosphere feature parameter range, such as the cooking stove element of metal material corresponding to the scene atmosphere feature of high light reflectivity (0.7-0.9). By combining the spatial parameters of the fixed position associated elements with the corresponding atmosphere feature parameters, a scene environment association group is formed, and the data structure includes three parts: element association list, spatial distribution matrix, and atmosphere parameter vector.

[0059] Step S126: Perform morphological change analysis processing on the object form sub-element in the picture content sub-element set, extract the morphological change sequence of the object form sub-element in the video segment unit timeline. The morphological change sequence records the change process of the object form in the order of timeline marking.

[0060] For food materials, kitchen utensils, and other object form sub-elements, a morphological analysis method based on three-dimensional mesh reconstruction is adopted. For the frame sequence containing the target object in the video segment unit, a three-dimensional point cloud model of the object is constructed through a multi-view stereo matching algorithm, and the volume, surface area, curvature, and other geometric parameters of each frame point cloud model are calculated. These parameters are arranged in the order of timeline marking to form a morphological change sequence, and each element in the sequence contains a timestamp, a set of geometric parameters, and a topological structure change marker. For example, the change sequence of a potato from a complete form to a shredded form records the process of gradually decreasing the volume from the initial value, gradually increasing the surface area, and changing the topological structure from a continuous body to a discrete body.

[0061] Step S127: Perform change feature association processing on the morphological change sequence to combine features with continuous change logic to form object morphological association groups, and the continuous change logic is determined based on the chronological order and change amplitude association relationship of object morphological changes.

[0062] Calculate the change amplitude value of adjacent morphological change features. When the difference value of the change amplitude value is less than a preset threshold (0.15), it is determined that there is continuous change logic. For example, in the process of slicing food materials, the thickness parameter changes from 5 mm to 4.8 mm, and then to 4.6 mm, and the change amplitude difference is 0.2 mm, which is less than the threshold, so it is combined to form a "thickness decreasing association group". Combine features with continuous change logic into object morphological association groups, and each association group contains a change starting feature, a change ending feature, a list of intermediate transition features, and a change rate curve.

[0063] Step S128: Perform audio feature extraction processing on the audio content elements of each video segment unit to obtain rhythm features, tone features, and semantic features of the audio content elements, the rhythm features representing the beat interval regularity of the audio, the tone features representing the frequency change trend of the audio, and the semantic features representing the meaning direction of the voice content in the audio.

[0064] Feature extraction is performed on the audio content elements using the librosa audio processing library. The rhythm features are constructed by calculating the spectral flux, zero-crossing rate, beat interval time, etc. of the audio signal to form a 64-dimensional feature vector; the tone features are constructed by extracting the fundamental frequency F0, harmonic frequency, spectral envelope, etc. from the frequency spectrum obtained by STFT transformation to form a 128-dimensional feature vector; and the semantic features are generated by encoding the text content after voice-to-text conversion using a pre-trained BERT model to generate a 768-dimensional context semantic vector. In the catering scene, the rhythm features may correspond to the regularity of the sound interval of cutting vegetables, the tone features correspond to the change of the voice of the explainer, and the semantic features correspond to the meaning of instructions such as "add half a spoon of salt" and "stew for five minutes".

[0065] Step S129: Perform association processing on the rhythm features, tone features, and semantic features to combine the beat intervals in the rhythm features that match the timeline of the picture content sub-elements, the frequency changes in the tone features that match the scene atmosphere features, and the meaning directions in the semantic features that match the character action association groups to form audio association groups.

[0066] An audio-picture time alignment model is constructed, and the beat interval of the rhythm feature is aligned with the timeline mark of the picture content sub-element through a dynamic time warping algorithm. When the matching degree of the beat interval and the action period of the character action sub-element is greater than 0.85, the time correlation is established. An audio-scene atmosphere mapping model is constructed, and the frequency change range of the tone feature is associated with the color temperature range of the scene atmosphere feature. For example, high-frequency tone (above 2000 Hz) corresponds to cold-tone scene (color temperature above 6500K). A semantic-action matching model is constructed, and the semantic feature vector is matched with the purpose pointing vector of the character action correlation group through cosine similarity calculation. When the similarity is greater than 0.7, the semantic correlation is established. The audio features meeting the above correlation conditions are combined into an audio correlation group, which includes three subfields of time alignment parameters, frequency-atmosphere mapping table, and semantic-action matching degree.

[0067] Step S1210: Perform fusion processing on the character action correlation group, scene environment correlation group, object form correlation group, and audio correlation group to obtain a narrative logic anchor point corresponding to each video segment unit. The narrative logic anchor point contains core correlation information of the picture and the audio, and is used for subsequent construction of an overall picture semantic stream.

[0068] The four groups of correlation features are fused by using an attention mechanism to construct a 896-dimensional (128 characters + 192 scenes + 256 objects + 320 audio) narrative logic anchor point vector. The interaction weights between the features in each group are calculated by a multi-head self-attention layer. The weight of the character action correlation group is set to 0.4, the weight of the scene environment correlation group is 0.2, the weight of the object form correlation group is 0.2, and the weight of the audio correlation group is 0.2. After weighted splicing, the dimension is reduced to a 256-dimensional anchor point vector through two fully connected networks. In a catering scene example, a narrative logic anchor point may represent the core content direction of "in a warm-tone kitchen scene, processing potatoes into a filamentous shape through a cutting action, accompanied by regular cutting sound and a 'potato cutting' voice instruction".

[0069] Step S130: Perform picture semantic stream construction processing based on the narrative logic anchor points corresponding to all video segment units to obtain an overall picture semantic stream corresponding to the set of digital multimedia video segments to be edited, wherein the overall picture semantic stream arranges the video segment units according to the correlation relationship of the narrative logic anchor points.

[0070] In the catering food tutorial scene, the overall picture semantic stream embodies the logical order of the cooking process, such as the linear narrative structure of "food procurement - preprocessing - cooking - plating - tasting". By calculating the semantic correlation strength between the narrative logic anchor points, the discrete video segment units are organized into an ordered sequence that conforms to the cooking process, ensuring the logicality and coherence of the tutorial content.

[0071] Step S131: Extract the character action association group, scene environment association group, object form association group and audio association group in each narrative logic anchor point as the basic elements for constructing the picture semantic stream.

[0072] The four groups of association group data are recovered from the narrative logic anchor point vector through feature decoupling operation. The character action association group is reconstructed by the first 128 features of the anchor point vector, the scene environment association group is reconstructed by features 129-320, the object form association group is reconstructed by features 321-576, and the audio association group is reconstructed by features 577-896. Each association group data contains original feature parameters and association strength values for subsequent association matching processing.

[0073] Step S132: Select one of the video segment units as the initial basic element, and perform association matching processing on the character action association group in the initial basic element to filter out the narrative logic anchor points with continued logic from the character action association group of other video segment units.

[0074] The continued logic of the character action association group is judged by a random forest classifier. The classifier input is the action feature sequence, time interval and destination pointing vector of the two association groups, and the output is the continued logic probability. Select the video segment unit corresponding to the starting action of the cooking process (such as "food material washing") as the initial basic element, and compare its character action association group with the character action association groups of all other video segment units to filter out anchor points with continued logic probability greater than 0.75 as candidate continuation units.

[0075] Step S133: Arrange the video segment units corresponding to the narrative logic anchor points with continued logic in the order of the continued sequence of the character action association group after the initial video segment unit to form a preliminary segment sequence.

[0076] Sort the candidate units according to the time interval between the ending action of the character action association group and the starting action of the candidate unit. The smaller the time interval, the closer to the front the candidate unit is arranged. For example, the ending action of the initial unit is "food material washing completed", the starting action of candidate unit A is "peeling start" (time interval 1 second), and the starting action of candidate unit B is "cutting start" (time interval 3 seconds), then the preliminary segment sequence is initial unit - candidate unit A - candidate unit B.

[0077] Step S134: Perform adaptation analysis processing on the scene environment association groups of adjacent video segment units in the preliminary segment sequence to determine whether the spatial layout features and scene atmosphere features of adjacent scene environment association groups have transition logic. If they have transition logic, the arrangement order of adjacent video segment units is maintained.

[0078] Step S1341: Extract the area division information in the spatial layout features of the pre-scene environmental association group, and determine the core area and non-core area in the pre-scene. The core area is the main area where the action of the character or the change of the object form occurs in the scene.

[0079] By calculating the action interaction frequency of the scene elements, the area where the scene element with the most interaction times with the character action sub-element is determined as the core area. For example, in the pre-scene, the interaction frequency of the chopping board element with the chopping action is 0.9, which is higher than that of the stove (0.1) and the sink (0.05), and therefore the area where the chopping board is located is the core area, and the rest is the non-core area.

[0080] Step S1342: Extract the area division information in the spatial layout features of the post-scene environmental association group, and determine the core area and non-core area in the post-scene.

[0081] Using the same method as step S1341, the action interaction frequency of the post-scene elements is calculated to determine the core area. For example, in the post-scene, the interaction frequency of the stove element with the stir-frying action is 0.85, which is determined as the core area.

[0082] Step S1343: Compare the position relationship of the core area of the pre-scene and the core area of the post-scene, determine the overlap ratio of the core area of the pre-scene and the core area of the post-scene in the picture coordinate system, and the corresponding relationship between the exit direction of the core area of the pre-scene and the entrance direction of the core area of the post-scene.

[0083] A picture coordinate system is established, with the origin at the top left corner of the picture, the X-axis to the right, and the Y-axis downward. The intersection area and the union area of the circumscribed rectangle of the pre-core area and the circumscribed rectangle of the post-core area are calculated to obtain the overlap ratio; by analyzing the motion direction vector of the character action sub-element at the boundary of the core area, the exit direction (pre-scene) and the entrance direction (post-scene) are determined, and the direction vector cosine value is calculated to determine the corresponding relationship.

[0084] Step S1344: Extract the visual style elements in the scene atmosphere features of the pre-scene environmental association group, including color saturation distribution, light intensity distribution, and texture density distribution.

[0085] The pre-scene atmosphere feature vector is decomposed into three sub-vectors: color saturation distribution (16 dimensions), light intensity distribution (16 dimensions), and texture density distribution (16 dimensions). Each sub-vector represents the distribution of the corresponding feature in different areas of the picture, for example, the color saturation distribution includes the average saturation of the upper left, upper right, lower left, lower right, and center of the picture.

[0086] Step S1345: Extract the visual style elements in the scene atmosphere features of the post-sequential scene environment association group, including color saturation distribution, light intensity distribution, and texture density distribution.

[0087] Using the same method as step S1344, the post-sequential scene atmosphere feature vector is decomposed into three sub-vectors, and the distribution parameters of the corresponding visual style elements are obtained.

[0088] Step S1346: Compare the color saturation distribution of the pre-sequential scene and the post-sequential scene to determine the change trend of the pre-sequential scene and the post-sequential scene in the main color channel, including the increase or decrease direction of the color saturation and the change range.

[0089] Calculate the average difference of the saturation of the pre-sequential and post-sequential scenes in the red, green, and blue three main color channels. If the difference is positive, it represents an increasing trend, and if it is negative, it represents a decreasing trend. The change range is represented by the ratio of the absolute value of the difference to the pre-sequential average saturation. For example, the red channel saturation difference is 0.1, the pre-sequential average is 0.5, and the change range is 0.2 (20%).

[0090] Step S1347: Compare the light intensity distribution of the pre-sequential scene and the post-sequential scene to determine the change trend of the light intensity of the pre-sequential scene and the post-sequential scene in the main area of the picture, including the increase or decrease direction of the light intensity and the change range.

[0091] Divide the picture into a center area and an edge area, calculate the average difference of the light intensity of the pre-sequential and post-sequential scenes in the two areas, and determine the increase or decrease direction and the change range. The calculation method is the same as the color saturation change trend.

[0092] Step S1348: Compare the texture density distribution of the pre-sequential scene and the post-sequential scene to determine the change trend of the texture density of the pre-sequential scene and the post-sequential scene in the core area and the non-core area, including the increase or decrease direction of the texture density and the change range.

[0093] Calculate the average difference of the texture density of the pre-sequential and post-sequential scenes in the core area and the non-core area (texture density is represented by the proportion of edge pixels output by the edge detection operator), and determine the increase or decrease direction and the change range.

[0094] Step S1349: Based on the position relationship of the core area of the pre-sequential scene and the post-sequential scene, the color saturation distribution change trend, the light intensity distribution change trend, and the texture density distribution change trend, construct the transition logic evaluation dimension.

[0095] Four evaluation dimensions are set, each of which contains a weight coefficient: core area position relationship (weight 0.4), color saturation change (weight 0.2), light and shadow intensity change (weight 0.2), and texture density change (weight 0.2). The evaluation value of each dimension ranges from 0 to 1, and the total transition logic score is obtained by weighted summation.

[0096] Step S13410: Perform the correlation degree description for each evaluation dimension. If the overlap ratio in the core area position relationship meets the preset transition overlap standard and the outlet-inlet direction corresponds, and the change trends of color saturation, light and shadow intensity, and texture density all show gradual changes, then the adjacent scene environment correlation group has transition logic.

[0097] The preset transition overlap standard is that the overlap ratio is greater than or equal to 0.3, and the outlet-inlet direction correspondence standard is that the direction vector cosine value is greater than or equal to 0.6. The gradual change standard is that the change range is less than or equal to 0.3 (30%). When the core area overlap ratio is 0.4 (≥0.3), the direction cosine value is 0.7 (≥0.6), the color change range is 0.2 (≤0.3), the light and shadow change range is 0.15 (≤0.3), and the texture change range is 0.25 (≤0.3), the total transition logic score is 0.4*1+0.2*1+0.2*1+0.2*1=1.0, and it is determined that there is transition logic.

[0098] Step S13411: If the overlap ratio in the core area position relationship does not meet the preset transition overlap standard, or the outlet-inlet direction does not correspond, or the change trends of color saturation, light and shadow intensity, and texture density show jump changes, then the adjacent scene environment correlation group does not have transition logic.

[0099] When the core area overlap ratio is 0.2 (<0.3), or the direction cosine value is 0.5 (<0.6), or the change range of any visual style element is greater than 0.3, it is determined that there is no transition logic. For example, the color change range is 0.4 (>0.3), and the total transition logic score is 0.4*1+0.2*0+0.2*1+0.2*1=0.8, which is still determined as not having transition logic.

[0100] Step S13412: Integrate the correlation degree description of each evaluation dimension to form the judgment result of whether the adjacent scene environment correlation group has transition logic, which is used for subsequent adjustment of the preliminary segment sequence.

[0101] The correlation degree description of the evaluation dimension is quantified as a score (0 or 1), and the total score is obtained by weighted summation. When the total score is greater than or equal to 0.8, it is determined that there is transition logic, otherwise there is not. The judgment result is stored as a Boolean value, True indicating that there is transition logic, and False indicating that there is not.

[0102] Step S135: If the adjacent scene environment association groups do not have transition logic, screen out video segment units from the set of digital multimedia video segments to be edited, which have transition logic with both the front and rear scene environment association groups, insert the video segment units corresponding to the screened out video segment units between the adjacent two video segment units, and update the preliminary segment sequence.

[0103] A candidate library of scene transitions is established to store the scene environment association group features of all video segment units. For adjacent segment pairs (A, B) without transition logic, the candidate library is traversed to screen out segments C that have transition logic with both A and B (transition logic score of A and C ≥ 0.8, transition logic score of C and B ≥ 0.8). In a dining scene, when A is an "outdoor food material procurement scene" and B is a "kitchen cooking scene", a "food material brought into the kitchen scene" is inserted as a transition segment C to naturally transition the scene from outdoor to indoor.

[0104] Step S136: Perform continuity analysis processing on the object form association groups of adjacent video segment units in the updated preliminary segment sequence to determine whether the form change sequence of adjacent object form association groups has continuity logic, and if so, maintain the arrangement order of the adjacent video segment units.

[0105] The dynamic time warping distance (DTW distance) of the form change sequence of adjacent object form association groups is calculated, and the smaller the DTW distance, the stronger the continuity logic. A preset DTW distance threshold is 0.2, and when the DTW distance of adjacent association groups is < 0.2, it is determined that there is continuity logic. For example, the DTW distance of a "potato dicing" association group and a "potato stir-frying" association group is 0.15, which is determined to have continuity logic.

[0106] Step S137: If adjacent object form association groups do not have continuity logic, perform feature adjustment processing on the object form association groups of adjacent video segment units to form continuity logic between the end feature of the object form association group of the former video segment unit and the start feature of the object form association group of the latter video segment unit, and update the preliminary segment sequence again.

[0107] Interpolation adjustment is performed on the object form features, the feature difference ΔF = F2 - F1 between the end feature (F1) of the former association group and the start feature (F2) of the latter association group is calculated, ΔF is evenly distributed to N transition frames, and the feature is smoothly transitioned from F1 to F2 through interframe interpolation. For example, F1 is a potato piece size of 2 cm, F2 is a potato piece size of 1 cm, and N = 10 frames, then each frame is reduced by 0.1 cm to form a continuous change sequence.

[0108] Step S138: Perform rhythm matching processing on the audio association groups of all video segment units in the re-updated preliminary segment sequence. If the beat intervals of the rhythm features of the audio association groups of adjacent video segment units are inconsistent, perform adjustment processing on the rhythm features of the audio association group of one of the video segment units to keep the beat intervals of the rhythm features of the audio association groups of adjacent video segment units consistent.

[0109] An audio rhythm alignment algorithm based on dynamic time warping is used to calculate the beat interval difference of adjacent audio association groups. When the difference is >100 ms, rhythm adjustment is performed. The audio signal is resampled by a time stretching algorithm to adjust the beat interval while keeping the tone unchanged, so that the end beat interval of the former association group and the start beat interval of the latter association group have a difference ≤50 ms. For example, the former audio beat interval is 600 ms and the latter audio beat interval is 750 ms. The latter audio beat interval is adjusted to 600 ms by time compression.

[0110] Step S139: Perform overall narrative logic verification processing on the preliminary segment sequence after rhythm adjustment. By comparing the association relationship of the narrative logic anchors of all video segment units, it is determined whether the preliminary segment sequence conforms to the overall narrative logic. If it does, the preliminary segment sequence is taken as the overall picture semantic stream corresponding to the digital multimedia video segment set to be edited.

[0111] An overall narrative logic evaluation index system is constructed, including action sequence continuity (weight 0.4), scene transition naturalness (weight 0.2), object change rationality (weight 0.2), and audio rhythm consistency (weight 0.2). Each index is scored using a 1-5 point system. When the weighted average score is ≥4 points, it is determined to conform to the overall narrative logic. In a catering scene, when the action sequence completely covers the cooking process, the scene transition is smooth, the food material form changes in accordance with the processing rules, and the explanation audio rhythm is stable, it is determined to conform to the overall narrative logic.

[0112] Step S1310: If the overall narrative logic is not met, the initial basic elements are reselected, and the above-mentioned screening, arrangement, adaptation, adjustment and verification processing are repeated until the overall picture semantic stream that meets the overall narrative logic is obtained.

[0113] When the weighted average score is <4 points, the lowest scoring index item is analyzed, and the video segment unit related to the index is reselected as the initial basic element. For example, when the total score is insufficient due to a "action sequence continuity" score of 2 points, a segment containing core cooking actions is reselected as the initial basic element, and the picture semantic stream construction process is re-executed.

[0114] Step S140: perform a splicing adaptation rule generation process on adjacent video segment units in the overall picture semantic stream to obtain splicing adaptation rules corresponding to the adjacent video segment units, which match the picture semantic stream features and the audio semantic stream features of the adjacent video segment units.

[0115] The splicing adaptation rule is embodied as a specific parameter set guiding the smooth transition between video segments, which, in a catering scenario, includes natural connection of character cutting actions, light and shadow transition of the kitchen scene, continuous change of food material form, smooth superposition of cooking sound effects, and other rule elements. By analyzing the differences in narrative logic anchor features of adjacent segments, structured rule data containing transition types, duration, parameter change curves, and other content is generated.

[0116] Step S141: extract the narrative logic anchors of the two adjacent video segment units in the overall picture semantic stream, and label them as pre-narrative logic anchors and post-narrative logic anchors, respectively.

[0117] Read the IDs of adjacent video segment units from the metadata of the overall picture semantic stream, index to the corresponding narrative logic anchor data according to the IDs, and add "pre" and "post" labels respectively. The pre-anchor corresponds to the earlier segment unit on the timeline, and the post-anchor corresponds to the later segment unit.

[0118] Step S142: perform action connection analysis processing on the character action association groups in the pre-narrative logic anchor and the character action association groups in the post-narrative logic anchor to determine the connection mode of the end features of the pre-character action and the start features of the post-character action, which is determined based on the amplitude change and the direction change of the action.

[0119] Calculate the difference in bone point motion parameters between the end frame of the pre-action and the start frame of the post-action, with the amplitude change represented by the difference in displacement vector module length and the direction change represented by the angle between the motion direction vectors. According to the amplitude change rate (difference / pre-amplitude) and the direction angle value, divide the connection mode type: amplitude gradual connection when amplitude change rate <0.3 and direction angle <30°; direction turning connection when amplitude change rate ≥0.3 and direction angle ≥30°; action pause connection when action interval time >1 second.

[0120] Step S143: generate a character action connection rule according to the connection mode, which is used to regulate the transition mode of the character action in adjacent video segment units.

[0121] Step S1431: determine the type of character action connection mode based on the connection relationship between the end features of the pre-character action and the start features of the post-character action, including amplitude gradual connection, direction turning connection, and action pause connection.

[0122] The type of the connection mode is determined directly by the analysis result of step S142, for example, the amplitude change rate 0.25 (<0.3) and the direction angle 25° (<30°) of the ending feature of the pre-action "cutting vegetables" and the starting feature of the post-action "stir-frying" are determined as the "amplitude gradual connection" type.

[0123] Step S1432: If the type of the connection mode is the amplitude gradual connection, the change interval of the amplitude value of the ending feature of the pre-action and the amplitude value of the starting feature of the post-action is analyzed.

[0124] The amplitude value (A1) of the displacement of the skeleton point of the ending frame of the pre-action and the amplitude value (A2) of the displacement of the skeleton point of the starting frame of the post-action are extracted, and the change interval ΔA = |A2-A1| is calculated. For example, A1 = 0.4 meters, A2 = 0.6 meters, and ΔA = 0.2 meters.

[0125] Step S1433: According to the change interval, the number of steps of the amplitude gradual change is determined, each step corresponds to an incremental or decremental change of the amplitude value, and the amplitude value of the pre-action is transitioned to the amplitude value of the post-action through multiple steps.

[0126] The number of steps N is calculated according to the video frame rate F (unit: fps) and the preset transition time T (unit: seconds), N = F*T. In the catering scene, usually T = 0.5 seconds and F = 30 fps, then N = 15 steps. The amplitude change amount of each step is ΔA / N, which ensures that the gradual change from A1 to A2 is completed in 15 steps.

[0127] Step S1434: The time length corresponding to each step is specified, which matches the frame rate of the video segment unit, so that the amplitude gradual change process presents a smooth effect on the screen.

[0128] The time length t of each step is 1 / F, for example, when F = 30 fps, t = 1 / 30 seconds ≈ 33.33 milliseconds. By controlling the amplitude change amount of each frame, the action gradual change process conforms to the human eye's visual persistence characteristics, avoiding the feeling of lag.

[0129] Step S1435: Based on the number of steps of the amplitude gradual change and the time length of each step, the amplitude gradual connection details are generated, which clearly indicate the amplitude value of the action at each time node.

[0130] Taking the time node t_i = i*t (i = 0, 1,..., N) as the horizontal axis and the amplitude value A_i = A1+i*(ΔA / N) (when increasing) or A_i = A1-i*(ΔA / N) (when decreasing) as the vertical axis, an amplitude change curve is generated. The curve is discretized into N+1 key points, each key point contains a timestamp and a corresponding amplitude value, forming the details data.

[0131] Step S1436: If the connection mode type is direction turning connection, analyze the turning range of the direction angle of the end feature of the previous character action and the direction angle of the start feature of the subsequent character action.

[0132] Extract the motion direction vector of the end frame of the previous action and the motion direction vector of the start frame of the subsequent action, calculate the direction angle θ by the vector dot product formula, and the turning range is the angle value of θ. For example, the previous direction vector is (1, 0, 0) and the subsequent direction vector is (0, 1, 0), θ = 90°, and the turning range is 90°.

[0133] Step S1437: According to the turning range, determine the number of segments of direction turning, each segment corresponding to a turning adjustment of the direction angle, so that the direction angle of the previous character action transitions to the direction angle of the subsequent character action through multiple segments.

[0134] The number of segments M is calculated according to the turning range θ and the preset maximum turning angle α of each segment (usually 15°), M = ceil(θ / α). For example, θ = 90° and α = 15°, then M = 6 segments, each segment turning 15°.

[0135] Step S1438: Define the time length corresponding to each segment, which matches the frame rate of the video segment unit, so that the direction turning process presents a smooth effect on the screen.

[0136] The time length t_s of each segment is T / M, where T is the preset transition time (0.5 seconds). For example, M = 6 segments, t_s = 0.5 / 6 ≈ 0.083 seconds, corresponding to 2-3 frames of pictures (30fps).

[0137] Step S1439: Based on the number of segments of direction turning and the time length of each segment, generate the direction turning connection details, which clearly define the direction angle of the character action corresponding to each time node.

[0138] Take the time node t_j = j*t_s (j = 0, 1,..., M) as the horizontal axis and the direction angle θ_j = j*(θ / M) as the vertical axis to generate the direction change curve. Calculate the three-dimensional direction vector of each time node by the spherical linear interpolation (Slerp) algorithm to ensure that the turning process is smooth and natural.

[0139] Step S14310: If the connection mode type is action pause connection, analyze the stable duration requirement of the end feature of the previous character action and the preparation time requirement of the start feature of the subsequent character action.

[0140] The stability duration requirement is calculated according to the variance of the skeleton point position of the action end frame, and the smaller the variance, the more stable the action, and the shorter the required time, usually 0.2-0.5 seconds; the preparation time requirement is calculated according to the skeleton point acceleration of the start frame of the subsequent action, and the greater the acceleration, the more intense the preparation action, and the longer the required time, usually 0.3-0.7 seconds.

[0141] Step S14311: According to the stability duration requirement and the preparation time requirement, the total duration of the action pause is determined, which is equal to the sum of the stability duration and the preparation time.

[0142] The total duration T_pause=T_stable+T_prepare, for example, T_stable=0.3 seconds, T_prepare=0.4 seconds, T_pause=0.7 seconds.

[0143] Step S14312: Within the total duration of the action pause, the time period when the end feature of the previous character action remains stable and the time period when the start feature of the subsequent character action begins to prepare are specified.

[0144] The time period is divided into a stable period of [0, T_stable) and a preparation period of [T_stable, T_pause]. In the stable period, the skeleton point position remains unchanged, and in the preparation period, the skeleton point gradually transitions from the stable position to the start position of the subsequent action.

[0145] Step S14313: Based on the total duration of the action pause, the stability time period and the preparation time period, the action pause connection details are generated, which clearly specify the character action state corresponding to each time period in the action pause process.

[0146] The details include a set of static skeleton point coordinates in the stable period and a sequence of skeleton point motion parameters in the preparation period, and the preparation period parameter sequence is generated by linear interpolation to ensure smooth transition from static to motion state.

[0147] Step S14314: The amplitude gradual change connection details, the direction turning connection details or the action pause connection details are integrated to form the character action connection rule, which includes the coordination requirements of the joint motion of the character limbs in the connection process, which are set based on the motion law of natural continuity of the character action.

[0148] According to the type of connection mode, the corresponding details are selected as the main body, and the limb coordination requirement parameters are added, such as the motion phase difference between the shoulder joint and the elbow joint needs to be kept at 90°±15°, and the wrist joint motion needs to lag behind the elbow joint motion by 1-2 frames. The coordination requirements are expressed by kinematic constraint equations to ensure that the action conforms to the physiological structure limit of the human body.

[0149] Step S14315: Perform a picture adaptability confirmation process on the generated character motion transition rule, judge whether the transition rule is executable under the corresponding picture parameters by comparing the picture resolution and frame rate of the adjacent video segment units, if executable, determine the final character motion transition rule.

[0150] Check whether the product of the step length / segment number and the frame rate in the transition rule is an integer (ensure that each step / segment corresponds to an integer number of frames), and whether the number of frames within the total duration of the rule matches the picture resolution (more transition frames may be needed under high resolution). For example, N=15 steps, F=30fps, 15*33.33ms=500ms, corresponding to 15 frames of pictures, the executability confirmation is passed.

[0151] Step S14316: If not executable, adjust the step length, segment number or total duration in the transition rule to match the picture parameters, perform the picture adaptability confirmation process again until an executable character motion transition rule is obtained.

[0152] When the product of the step length and the frame rate is not an integer, adjust the step length N'=round(N*F*t) to ensure that N'*t is an integer second. For example, original N=14, F=30fps, t=33.33ms, 14*33.33ms≈466.67ms, adjust N'=15 to make the total duration 500ms, matching the 30fps frame rate.

[0153] Step S144: Perform a scene transition analysis process on the scene environment association groups in the pre-narrative logic anchor point and the scene environment association groups in the post-narrative logic anchor point, determine the overlap relationship between the end region of the spatial layout features of the pre-scene environment and the start region of the spatial layout features of the post-scene environment, and the transition tendency of the pre-scene atmosphere features and the post-scene atmosphere features.

[0154] Calculate the spatial layout feature difference between the last 5 frames of the pre-scene and the first 5 frames of the post-scene, the overlap relationship is represented by the intersection over union (IoU) of the core region, IoU≥0.5 is high overlap, 0.3≤IoU<0.5 is moderate overlap, and IoU<0.3 is low overlap; the transition tendency is represented by the cosine similarity of the scene atmosphere feature vector, similarity≥0.8 is gradual transition, <0.8 is abrupt transition.

[0155] Step S145: Generate a scene environment transition rule according to the overlap relationship and the transition tendency, the scene environment transition rule is used to regulate the transition mode of the scene environment in adjacent video segment units.

[0156] According to the overlap relationship, the transition type is selected: high overlap uses "fade-in and fade-out" transition, moderate overlap uses "wipe" transition, and low overlap uses "dissolve" transition; according to the transition tendency, the transition duration is set: gradual transition is 30 frames (1 second), and sudden transition is 15 frames (0.5 second). The rules include transition type, duration, direction (when wiping), color correction parameters, etc. For example, "transition from left to right, duration 30 frames, color correction matrix [[r11, r12, r13], [r21, r22, r23], [r31, r32, r33]]".

[0157] Step S146: Perform shape continuation analysis processing on the object shape association group in the previous narrative logic anchor point and the object shape association group in the subsequent narrative logic anchor point, determine the association relationship between the end change characteristics of the previous object shape and the start change characteristics of the subsequent object shape, and the association relationship is determined based on the structural characteristics and change speed of the object shape.

[0158] The structural characteristics (volume, surface area, principal axis direction) and change speed (volume change rate per unit time) of the object shape are extracted, and the structural similarity (SSIM) and speed difference between the previous end characteristics and the subsequent start characteristics are calculated. SSIM≥0.7 and speed difference<0.2 are "strong association", 0.5≤SSIM<0.7 and 0.2≤speed difference<0.4 are "moderate association", and SSIM<0.5 or speed difference≥0.4 are "weak association".

[0159] Step S147: Generate object shape continuation rules according to the association relationship, which are used to regulate the transition mode of the object shape in adjacent video segment units.

[0160] When the association is strong, use the "direct gradual change" rule to realize the transition through linear interpolation of shape parameters; when the association is moderate, use the "structure priority gradual change" rule to preferentially ensure smooth transition of structural characteristics; when the association is weak, use the "replacement transition" rule to introduce a new shape generation algorithm. The rules include interpolation function type (linear / nonlinear), key shape parameter control points, transition duration, etc. For example, "volume parameter uses cubic Bezier curve interpolation, control points (t0, v0), (t1, v1), (t2, v2), transition duration 45 frames".

[0161] Step S148: Perform audio fusion analysis processing on the audio association group in the previous narrative logic anchor point and the audio association group in the subsequent narrative logic anchor point, determine the matching relationship between the end beat of the rhythm characteristics of the previous audio and the start beat of the rhythm characteristics of the subsequent audio, and the transition interval of the tonal characteristics of the previous audio and the tonal characteristics of the subsequent audio.

[0162] The beat interval difference and the fundamental frequency F0 difference between the last 2 seconds of the pre-sequence audio and the first 2 seconds of the post-sequence audio are calculated, the matching relationship is represented by the beat synchronization error (≤50 ms is synchronous, >50 ms is asynchronous), and the transition interval is represented by the frequency range (Hz) of the F0 difference. For example, the beat synchronization error is 30 ms, the F0 difference is 50 Hz, and the transition interval is [F0_prev-25, F0_next+25] Hz.

[0163] Step S149: generating an audio fusion rule according to the matching relationship and the transition interval, the audio fusion rule being used to regulate the transition mode of the audio in adjacent video segment units.

[0164] The "cross-fade" rule is used when the beats are synchronized, the volume of the pre-sequence audio is linearly attenuated, the volume of the post-sequence audio is linearly enhanced, and the cross point is at the beat peak; the "beat remapping" rule is used when the beats are asynchronous, and the beats of the post-sequence audio are adjusted by time stretching. The "pitch interpolation" rule is used in the transition interval, and the F0 is smoothly transitioned from F0_prev to F0_next by the PSOLA algorithm. The rule contains parameters such as fade-in and fade-out duration, volume change curve, and pitch adjustment rate, for example, "cross-fade duration 500 ms, volume curve using cosine function, and pitch adjustment rate 10 Hz / 100 ms".

[0165] Step S1410: performing integration processing on the character action connection rule, the scene environment transition rule, the object form continuation rule, and the audio fusion rule to obtain a splicing adaptation rule corresponding to adjacent video segment units, the splicing adaptation rule containing all transition specifications of adjacent video segment units in terms of picture and audio.

[0166] The four groups of rule data are stored as XML format files, the root node is "splicing adaptation rule", the child nodes are "character action connection", "scene environment transition", "object form continuation", and "audio fusion" respectively, and each child node contains three attributes of rule type, parameter list, and time interval. In the dining scene example, a complete splicing adaptation rule may contain specific specifications such as "direction transition connection of potato shredding (15 frames)", "kitchen warm light fade-in and fade-out transition (30 frames)", "cubic Bezier interpolation of food material volume (45 frames)", and "cross-fade of cutting sound (500 ms)".

[0167] Step S150: performing transition fusion processing on adjacent video segment units based on the splicing adaptation rule to obtain a video segment sequence with fusion transition effect.

[0168] The transition fusion processing is a process of converting the splicing adaptation rule into actual video pictures and audio signals, which is manifested as seamless switching between food processing shots, smooth connection of cooking actions, and natural transition of explanation audio in the catering scene. Through digital signal processing technology, accurate adjustment is performed on picture frames and audio sampling points of adjacent segments to generate a new video segment sequence containing transition special effects.

[0169] Step S151: Extract the character action connection rule in the splicing adaptation rule, and determine the transition mode of the character action in the adjacent video segment unit.

[0170] The "rule type" attribute of the "character action connection" sub-node in the splicing adaptation rule XML file is parsed to determine that the transition mode is amplitude fading, direction turning, or action pausing. For example, if the parsing result is "amplitude fading connection", the corresponding connection parameters in the rule are used for subsequent processing.

[0171] Step S152: According to the transition mode of the character action, perform action feature extension processing on the end frame of the character action associated group of the previous video segment unit to generate an action transition frame, and the action feature of the action transition frame is between the end feature of the previous character action and the start feature of the subsequent character action.

[0172] According to the step number N and the amplitude / direction change curve in the connection rule, N transition frames are inserted after the end frame of the previous sequence. The coordinates of the bone points of each transition frame are calculated by the coordinates of the end frame of the previous sequence and the change curve parameters, for example, the coordinates of the ith frame = the end coordinates of the previous sequence + i*(the start coordinates of the subsequent sequence - the end coordinates of the previous sequence) / N. The transition frame picture is synthesized based on the bone point coordinates by the image generation model to ensure the natural and smooth action.

[0173] Step S153: Insert the action transition frame between the end of the previous video segment unit and the start of the subsequent video segment unit to realize the transition fusion of the character action.

[0174] The generated N transition frames are spliced in time sequence at the end of the previous segment, and the start timestamp of the subsequent segment is adjusted so that the transition frame sequence and the previous and subsequent segments form a continuous time line. The insertion of the transition frame is realized by the track nesting technology in the video editing software to ensure that the picture frame sequence number is continuous without gaps.

[0175] Step S154: Extract the scene environment transition rule in the splicing adaptation rule to determine the transition mode of the scene environment in the adjacent video segment unit.

[0176] The "rule type" attribute of the "scene environment transition" sub-node in the splicing adaptation rule XML file is parsed to determine that the transition mode is fade-in and fade-out, cut, or dissolve. For example, if the parsing result is "fade-in and fade-out transition", the transition duration (such as 30 frames) and direction parameters are obtained.

[0177] Step S155: According to the transition mode of the scene environment, spatial feature gradient processing is performed on the end area of the scene environment associated group of the prelude video segment unit to generate a scene transition area, and the spatial layout feature of the scene transition area is between the prelude scene environment and the subsequent scene environment.

[0178] According to the spatial layout feature difference and the scene atmosphere feature difference in the transition rule, the feature parameters of each transition frame are calculated. The spatial layout feature is generated by linear interpolation of element position and size, and the scene atmosphere feature is generated by smooth transition of color and light parameters. For example, the prelude scene color temperature is 5500K, the subsequent scene color temperature is 6500K, and the color temperature of each frame in the 30-frame transition increases by (6500-5500) / 30≈33.3K.

[0179] Step S156: The scene transition area is integrated into the overlapping interval of the end of the prelude video segment unit and the start of the subsequent video segment unit to realize the transition fusion of the scene environment.

[0180] For example, step S1561: determine the end frame sequence of the prelude video segment unit and the start frame sequence of the subsequent video segment unit, select the frame sequence adjacent to the end frame sequence of the prelude video segment unit and the start frame sequence of the subsequent video segment unit on the time line as the frame set of the overlapping interval, and the number of frames of the overlapping interval is determined based on the complexity of the scene transition area.

[0181] The complexity of the scene transition area is represented by the total variance of the feature parameters, and the greater the variance, the higher the complexity, and the more the number of frames of the overlapping interval. Generally, the number of frames corresponding to the transition time is taken as the number of frames of the overlapping interval, for example, 30 frames of transition correspond to 15 frames at the end of the prelude and 15 frames at the start of the subsequent as the overlapping interval.

[0182] Step S1562: Extract the spatial layout feature and the scene atmosphere feature of the scene transition area, and respectively decompose them into a plurality of independently adjustable sub-features, including region boundary sub-feature, color distribution sub-feature, light sub-feature and texture sub-feature.

[0183] The spatial layout feature is decomposed into sub-features such as coordinate set of region boundary, aspect ratio, rotation angle, etc.; and the scene atmosphere feature is decomposed into sub-features such as average value of RGB color channel, shadow area ratio, highlight intensity, texture definition, etc. Each sub-feature is represented by an independent parameter for easy individual adjustment.

[0184] Step S1563: sub-feature adjustment processing is performed on the prelude end frame sequence in the frame set of the overlapping interval, and the region boundary sub-feature of the prelude end frame sequence is gradually adjusted to the region boundary sub-feature of the scene transition area, and the adjustment process is performed in sequence by frame, and the region boundary sub-feature change amount of each frame remains consistent.

[0185] The region boundary coordinates of the kth frame at the end of the prologue (k = 1 to 15) = the prologue original coordinates + (the scene transition region coordinates - the prologue original coordinates) * k / 15, ensuring that the transition from the original feature to the transition region feature is completed within 15 frames. The change amount of each frame = (the scene transition region coordinates - the prologue original coordinates) / 15, maintaining a constant change rate.

[0186] Step S1564: At the same time, the color distribution sub-feature of the sequence of frames at the end of the prologue is gradually adjusted to the color distribution sub-feature of the scene transition region, and the adjustment process is performed in sequence for each frame. The change amount of the color distribution sub-feature of each frame remains consistent.

[0187] The same linear interpolation method as the region boundary sub-feature is used to perform frame-by-frame adjustment on the RGB color channel mean. For example, the prologue R channel mean is 0.6, the transition region R channel mean is 0.8, and the R channel mean of each frame increases by (0.8-0.6) / 15≈0.0133, reaching the transition region feature value after 15 frames.

[0188] Step S1565: The same gradual adjustment process is performed on the light and shadow sub-feature and the texture sub-feature of the sequence of frames at the end of the prologue, so that all sub-features of the sequence of frames at the end of the prologue transition to the corresponding sub-features of the scene transition region through multi-frame adjustment.

[0189] The shadow area proportion and highlight intensity of the light and shadow sub-feature are adjusted, and the texture clarity and direction of the texture sub-feature are adjusted. The adjustment method is consistent with the region boundary and color distribution sub-feature, and linear interpolation and constant change rate are used to ensure that all sub-features are synchronized to complete the transition.

[0190] Step S1566: The sub-feature adjustment process is performed on the sequence of frames at the beginning of the postlude in the overlapping interval, and the region boundary sub-feature of the sequence of frames at the beginning of the postlude is gradually adjusted from the region boundary sub-feature of the scene transition region to the region boundary sub-feature of the postlude scene environment. The adjustment process is performed in sequence for each frame, and the change amount of the region boundary sub-feature of each frame remains consistent.

[0191] The region boundary coordinates of the mth frame at the beginning of the postlude (m = 1 to 15) = the scene transition region coordinates + (the postlude original coordinates - the scene transition region coordinates) * m / 15, which is opposite to the adjustment direction of the prologue end frame, forming a symmetrical transition curve.

[0192] Step S1567: At the same time, the color distribution sub-feature of the sequence of frames at the beginning of the postlude is gradually adjusted from the color distribution sub-feature of the scene transition region to the color distribution sub-feature of the postlude scene environment. The adjustment process is performed in sequence for each frame, and the change amount of the color distribution sub-feature of each frame remains consistent.

[0193] The color sub-feature adjustment method of the post-sequential starting frame is similar to that of the pre-sequential ending frame, starting from the scene transition region feature and gradually transitioning to the post-sequential original feature, ensuring the symmetry of the adjustment on both sides of the transition region.

[0194] Step S1568: Perform the same gradual adjustment process on the light and shadow sub-features and the texture sub-features of the post-sequential starting frame sequence, so that all sub-features of the post-sequential starting frame sequence transition from the corresponding sub-features of the scene transition region to the corresponding sub-features of the post-sequential scene environment through multi-frame adjustment.

[0195] The light and shadow sub-features and the texture sub-features use a symmetric adjustment strategy, matching the adjustment parameters of the pre-sequential ending frame to ensure smooth and continuous changes in scene features throughout the entire overlap interval.

[0196] Step S1569: Combine the adjusted pre-sequential ending frame sequence, the corresponding frame sequence of the scene transition region, and the adjusted post-sequential starting frame sequence in chronological order to form a scene transition frame sequence.

[0197] Splice the pre-sequential adjustment frames (15 frames), the transition region frames (0 frames, features have been integrated into the pre- and post-adjustment frames), and the post-sequential adjustment frames (15 frames) in chronological order to form a complete scene transition frame sequence containing 30 frames. The scene feature parameters of each frame in the sequence change continuously without jumps.

[0198] Step S15610: Replace the original ending frame sequence of the pre-sequential video segment unit and the starting frame sequence of the post-sequential video segment unit with the scene transition frame sequence to achieve the scene environment transition fusion of the pre-sequential video segment unit and the post-sequential video segment unit.

[0199] Delete the pre-sequential ending 15 frames and the post-sequential starting 15 frames in the video editing track and insert the scene transition frame sequence. Ensure that the replaced segments remain synchronized with other track content through timeline locking. After replacement, check the frame number continuity to ensure that there are no duplicate or missing frames.

[0200] Step S15611: Perform a picture coherence check process on the fused scene environment to observe whether the changes in scene environment sub-features in consecutive frames are smooth, confirm the transition fusion effect, and if the changes are smooth, complete the transition fusion of the scene environment.

[0201] Use an inter-frame difference analysis algorithm to calculate the absolute difference of scene feature parameters between adjacent frames. When the difference values of all sub-features are < a preset threshold (such as region boundary < 2 pixels, color mean < 0.02, light and shadow ratio < 0.05), it is determined that the changes are smooth and the transition fusion effect is qualified.

[0202] Step S15612: If the change is not smooth, adjust the sub-feature adjustment amount of the sequence of the end of the previous sequence and the sequence of the start of the subsequent sequence, increase the number of adjustment frames, and perform the sub-feature adjustment and the picture continuity check process again until the sub-feature change of the scene environment is smooth, and the transition fusion of the scene environment is completed.

[0203] When the difference of a certain sub-feature is greater than the threshold, the number of overlapping interval frames is increased (e.g., from 30 frames to 45 frames), the adjustment amount per frame is reduced, and the sub-feature adjustment is performed again. For example, the original region boundary is adjusted by 2 pixels per frame, the difference is 3 pixels which is greater than 2 pixels, the number of frames is increased to 45, and the adjustment amount per frame is 1.33 pixels, so that the difference is less than 2 pixels.

[0204] Step S157: Extract the object shape continuation rule in the splicing adaptation rule to determine the transition mode of the object shape in the adjacent video segment unit.

[0205] Parse the "rule type" attribute of the "object shape continuation" sub-node in the splicing adaptation rule XML file to determine the transition mode as direct gradient, structure priority gradient, or replacement transition. For example, the parsing result is "structure priority gradient", and the interpolation function type and key control point parameters are obtained.

[0206] Step S158: According to the transition mode of the object shape, perform shape gradient processing on the end change feature of the object shape associated group of the previous video segment unit to generate a shape transition sequence, and the shape change feature of the shape transition sequence is between the end change feature of the previous object shape and the start change feature of the subsequent object shape.

[0207] According to the shape parameter change curve (such as a cubic Bezier curve) in the continuation rule, the object shape parameters (volume, surface area, principal axis direction, etc.) of each frame in the transition sequence are calculated. Through a three-dimensional modeling software, the object mesh model is generated based on the parameters, and the picture frames of the shape transition sequence are rendered. For example, the transition of the food material from a block to a filament, the volume parameter changes from large to small according to the curve, and the surface area parameter changes from small to large.

[0208] Step S159: Insert the shape transition sequence between the end of the previous video segment unit and the start of the subsequent video segment unit to realize the transition fusion of the object shape.

[0209] Similar to the action transition frame insertion method, the N frames of the shape transition sequence are spliced at the end of the previous segment, and the start timestamp of the subsequent segment is adjusted. Through video synthesis technology, the transition sequence picture and the original segment picture are fused to ensure the consistency of the position and lighting of the object in the scene.

[0210] Step S1510: Extract the audio fusion rule in the splicing adaptation rule to determine the transition mode of the audio in the adjacent video segment unit.

[0211] Parse the "rule type" attribute of the "audio fusion" child node in the splicing adaptation rule XML file to determine the transition mode as cross-fade, beat remapping, or pitch interpolation. For example, the parsing result is "cross-fade", and the fade-in and fade-out duration and volume curve parameters are obtained.

[0212] Step S1511: According to the transition mode of the audio, perform rhythm fade processing on the last beat of the audio associated group of the prequel video segment unit, and perform pitch transition processing on the pitch feature of the prequel audio, to generate an audio transition segment, the rhythm feature and the pitch feature of which are between the prequel audio and the sequel audio.

[0213] According to the beat adjustment parameter in the fusion rule, perform time stretching or compression on the end of the prequel audio to gradually approach the beat interval of the sequel audio; according to the pitch interpolation parameter, adjust the fundamental frequency F0 of the end of the prequel audio through the PitchShift function of the audio editing software to make it change smoothly along the transition interval. The audio segment after rhythm and pitch adjustment is used as the audio transition segment, the duration of which matches the picture transition frame sequence.

[0214] Step S1512: Insert the audio transition segment between the end of the audio of the prequel video segment unit and the beginning of the audio of the sequel video segment unit to realize the transition fusion of the audio.

[0215] In the audio editing track, delete T seconds (T is the transition segment duration) at the end of the prequel audio and the beginning of the sequel audio, and insert the audio transition segment. Control the volume change of the transition segment through the volume automation curve, linearly attenuate the prequel part, and linearly enhance the sequel part to ensure that there is no obvious discontinuity in hearing.

[0216] Step S1513: After completing the transition fusion of the audio, scene environment, object form, and character action, combine the adjacent video segment units containing transition frames, transition regions, transition sequences, and transition segments to form a segment pair with fusion transition effect.

[0217] Combine the processed prequel segment, transition element, and sequel segment in chronological order to form a new segment pair containing transition effects. The timeline of the segment pair starts at the original start timestamp of the prequel segment and ends at the original end timestamp of the sequel segment + transition duration, ensuring accurate total duration.

[0218] Step S1514: Repeat the above transition fusion processing for all adjacent video segment units in the overall picture semantic stream to obtain a video segment sequence with fusion transition effect.

[0219] The processing of steps S151-S1513 is performed in sequence for all adjacent segment pairs in the whole picture semantic stream to generate a video segment sequence containing all transition effects. Each segment in the sequence contains original content and transition content, which are concatenated in time sequence to form a continuous video data stream.

[0220] Step S160: The video segment sequence with the fused transition effects is subjected to integration processing in the arrangement order of the whole picture semantic stream to obtain a complete digital multimedia video after editing and splicing, which maintains the coherence of narrative logic anchors and the consistency of picture semantic streams.

[0221] The integration processing is a process of final normalization and optimization of the video segment sequence after fusion of the transition effects, including steps such as time axis unification, parameter standardization, and content integrity verification, to generate a cooking tutorial video that meets the broadcast standards in the catering scenario, ensuring that the audience can clearly understand the logical relationship of the cooking process.

[0222] Step S161: The timeline markers of each video segment unit in the video segment sequence with the fused transition effects are extracted to determine the start time point and the end time point of each video segment unit.

[0223] The timeline marker information of each segment is read from the metadata of the video segment sequence, the start time point is the time code of the first frame of the segment, and the end time point is the time code of the last frame of the segment. For example, a segment contains 150 frames of pictures, the frame rate is 30 fps, the start time code is 00:01:20:00, and the end time code is 00:01:25:00 (150 / 30=5 seconds).

[0224] Step S162: Based on the arrangement order of the whole picture semantic stream, the start time point and the end time point of each video segment unit are integrated in sequence to generate a preliminary time axis, which records the time positions of all video segment units in the whole video.

[0225] The start time point of the first segment is taken as the 0 time of the whole video, and the start time point of the subsequent segment is taken as the end time point of the previous segment, and the preliminary time axis is generated in sequence. The time axis adopts a linear list data structure, and each element contains three fields of segment ID, start time, end time, and duration.

[0226] Step S163: The time connection relationship of adjacent video segment units in the preliminary time axis is checked to confirm whether the end time point of the previous video segment unit and the start time point of the subsequent video segment unit are continuous. If they are continuous, the structure of the preliminary time axis is maintained.

[0227] Calculate the difference between the end time and the start time of the adjacent segments, and determine that the time is continuous when the difference = 0. For example, the end time of the previous segment is 00:01:25:00, the start time of the next segment is 00:01:25:00, the difference is 0, and the time connection is continuous.

[0228] Step S164: If not continuous, analyze the cause of the time interval, and if the time is extended due to the insertion of transition frames, transition areas, etc. in the transition fusion processing, adjust the start time point of the next video segment unit to make the time axis continuous, and update the preliminary time axis.

[0229] When the transition fusion processing inserts N transition frames, causing the length of the previous segment to increase by N / F seconds, the start time point of the next segment is = the original start time point + N / F seconds. For example, 15 frames of transition frames are inserted, F = 30fps, the length increases by 0.5 seconds, and the start time point of the next segment is adjusted from 00:01:25:00 to 00:01:25:15 (time code format), ensuring that the time axis is continuous.

[0230] Step S165: Extract the picture parameters of each video segment unit in the video segment sequence with fusion transition effect, including picture resolution, frame rate and color space, and confirm whether the picture parameters of all video segment units are consistent.

[0231] Read the encoding information of each video segment to obtain the picture resolution (such as 3840x2160), frame rate (such as 30fps), and color space (such as Rec.709) parameters, and compare whether the parameter values of all segments are exactly the same. If any parameter is different, it is determined that the picture parameters are inconsistent.

[0232] Step S166: If there is a video segment unit with inconsistent picture parameters, perform uniform processing on the picture parameters of the video segment unit, and adjust its picture resolution, frame rate and color space to the same parameters as other video segment units, so that the overall picture parameters remain uniform.

[0233] Use video transcoding software to perform transcoding processing on the segments with inconsistent parameters: adjust the resolution to the target resolution through bicubic interpolation algorithm; adjust the frame rate to the target frame rate through frame interpolation (increase frames) or frame discarding (reduce frames); and convert the color space to the target color space through the color gamut conversion matrix. For example, transcode a segment of 1080p / 25fps / Rec.601 to 4K / 30fps / Rec.709 to match the parameters of other segments.

[0234] Step S167: Perform audio parameter unification processing on the video segment sequence with unified picture parameters and fusion transition effects, extract the audio parameters of each video segment unit, including sampling rate, channel number and bit rate, adjust the audio parameters of all video segment units to consistent parameters, and keep the overall audio parameters uniform.

[0235] Perform standardization processing on the audio parameters using audio editing software: adjust the sampling rate to 48 kHz through resampling algorithm; convert mono to stereo or stereo to mono through matrix encoding; adjust the bit rate to 24bit / 48kHz through lossless encoding. Ensure that the audio parameters of all segments meet the broadcast standards.

[0236] Step S168: Perform data integration processing on the video segment sequence with unified parameters and fusion transition effects according to the updated preliminary timeline, merge the picture data and audio data of all video segment units in chronological order to generate preliminary complete video data.

[0237] Use the "sequence merge" function of video editing software to splice all video segments into a complete sequence according to the segment order of the updated timeline. Picture data is merged through frame splicing technology, and audio data is merged through sample point splicing technology to ensure continuous data without gaps. The generated preliminary complete video data is stored in ProRes4444 encoding format, preserving the post-coloring space.

[0238] Step S169: Perform overall narrative logic check processing on the preliminary complete video data, play the preliminary complete video data to confirm the coherence of the overall narrative logic anchor points and the consistency of the picture semantic flow, if consistent, determine the preliminary complete video data as the complete digital multimedia video after editing and splicing.

[0239] Organizational professionals review the preliminary complete video according to the cooking process logic, check whether the process of food material processing-cooking operation-finished product display is coherent, whether the key steps are missing, and whether the transition effect is natural. When the audit score is ≥90 points (100 points), it is determined to be coherent and consistent, passing the narrative logic check.

[0240] Step S1610: If there is a narrative logic anchor point break or picture semantic flow inconsistency, analyze the video segment unit interval where the problem lies, re-perform splicing adaptation rule generation and transition fusion processing on the video segment unit interval of the video segment unit, and perform data integration and narrative logic check processing again until a complete digital multimedia video after editing and splicing with coherent narrative logic and consistent picture semantic flow is obtained.

[0241] When the review finds that "food marinating" and "stir-frying in the pot" are missing the key anchor point "oil temperature judgment", the problem interval is located, and the segment containing "oil temperature judgment" is inserted from the to-be-edited set again to regenerate the splicing adaptation rule and transition fusion processing, and the integration and checking are performed again. Repeat the process until the review score is ≥ 90 points, and finally generate a cooking tutorial video that meets the narrative logic.

[0242] Figure 2 A video segment automatic clipping and splicing system 100 applied to digital multimedia provided in an embodiment of the present application is shown, which includes a processor 1001 and a memory 1003 and program code stored in the memory 1003. The processor 1001 executes the above-mentioned program code to realize the steps of the video segment automatic clipping and splicing method applied to digital multimedia. The processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the video segment automatic clipping and splicing system 100 applied to digital multimedia can also include a transceiver 1004, which can be used for data interaction between the video segment automatic clipping and splicing system applied to digital multimedia and other video segment automatic clipping and splicing systems applied to digital multimedia, such as data sending and / or data receiving, etc. It should be noted that the transceiver 1004 is not limited to one in actual scheduling, and the structure of the video segment automatic clipping and splicing system 100 applied to digital multimedia does not constitute a limitation on the embodiments of the present application.

[0243] The memory 1003 is used to store the program code for executing the embodiments of the present application, and is controlled by the processor 1001 to execute. The processor 1001 is used to execute the program code stored in the memory 1003 to realize the steps shown in the foregoing method embodiments.

[0244] The embodiments of the present application provide a computer readable storage medium, which stores program code. The program code is executed by a processor to realize the steps and corresponding contents of the foregoing method embodiments.

[0245] The above is only an optional implementation of some implementation scenarios of the present application. It should be noted that for those skilled in the art, other similar implementation means according to the technical concept of the present application can be used without departing from the technical concept of the present application, and such implementation also falls within the protection scope of the embodiments of the present application.

Claims

1. A method for automatically editing and splicing video clips in digital multimedia, characterized in that, The method includes: Obtain a set of digital multimedia video clips to be edited, wherein the set of digital multimedia video clips to be edited contains multiple video clip units with visual content elements, audio content elements and timeline markers; Narrative logic anchor point extraction processing is performed on each video segment unit in the set of digital multimedia video segments to be edited, to obtain the narrative logic anchor point corresponding to each video segment unit. The narrative logic anchor point extraction processing includes: splitting the screen content elements of the video segment unit to obtain screen content sub-elements containing character actions, scene environment, and object shapes; performing temporal and spatial analysis on the character action sub-elements, scene environment sub-elements, and object shape sub-elements respectively to obtain corresponding character action association groups, scene environment association groups, and object shape association groups; extracting audio features from the audio content elements of the video segment unit, and matching and associating the audio features according to the screen content sub-elements to generate audio association groups; and generating the narrative logic anchor point based on the fusion result of the character action association group, scene environment association group, object shape association group, and audio association group. The narrative logic anchor point is used to characterize the core content orientation of the video segment unit in the overall narrative. Based on the narrative logic anchor points corresponding to all video segment units, the image semantic flow construction process is performed to obtain the overall image semantic flow corresponding to the set of digital multimedia video segments to be edited. The overall image semantic flow arranges the video segment units according to the relationship of the narrative logic anchor points. For adjacent video segment units in the overall picture semantic stream, a splicing adaptation rule generation process is performed to obtain splicing adaptation rules corresponding to adjacent video segment units. The splicing adaptation rules match the picture semantic stream features and audio semantic stream features of adjacent video segment units. Based on the splicing adaptation rules, a transition fusion process is performed on adjacent video segment units to obtain a video segment sequence with a fused transition effect. The video segment sequence with the fusion transition effect is integrated according to the overall semantic flow of the picture to obtain a fully edited and spliced ​​digital multimedia video. The fully edited and spliced ​​digital multimedia video maintains the coherence of the narrative logic anchor point and the consistency of the picture semantic flow.

2. The automatic video clip editing and splicing method for digital multimedia according to claim 1, characterized in that, The step of performing narrative logic anchor point extraction processing on each video segment unit in the set of digital multimedia video segments to be edited, to obtain the narrative logic anchor point corresponding to each video segment unit, includes: The screen content elements of each video segment unit are split to obtain a set of screen content sub-elements, which includes character action sub-elements, scene environment sub-elements, and object shape sub-elements; Perform continuous frame analysis processing on the character action sub-elements in the set of sub-elements of the screen content, and extract the continuous action sequence of the character action sub-elements within the timeline of the video segment unit. The continuous action sequence is arranged in the order of the timeline markers. The continuous action sequence is processed to perform action feature association processing, and action features with causal relationships are combined to form a character action association group. The causal relationship is determined based on the logic of the sequence of action occurrence and the purpose of the action. Spatial feature extraction processing is performed on the scene environment sub-elements in the set of scene content sub-elements to obtain the spatial layout features and scene atmosphere features of the scene environment sub-elements. The spatial layout features represent the positional distribution relationship of elements in the scene, and the scene atmosphere features represent the visual style tendency of the scene. The spatial layout features and scene atmosphere features are associated, and the elements with fixed positional associations in the spatial layout features are combined with the corresponding style tendencies in the scene atmosphere features to form scene environment association groups. A morphological change analysis process is performed on the object shape sub-elements in the set of sub-elements of the image content to extract the morphological change sequence of the object shape sub-elements within the timeline of the video segment unit. The morphological change sequence records the change process of the object shape in the order marked on the timeline. The morphological change sequence is subjected to change feature association processing, and features with continuous change logic are combined to form object morphological association groups. The continuous change logic is determined based on the relationship between the order of object morphological changes and the magnitude of changes. Audio feature extraction processing is performed on the audio content elements of each video segment unit to obtain the rhythm features, pitch features and semantic features of the audio content elements. The rhythm features represent the beat interval pattern of the audio, the pitch features represent the frequency change trend of the audio, and the semantic features represent the meaning of the speech content in the audio. The rhythm features, pitch features and semantic features are associated with each other, and the rhythm features that match the timeline of the sub-elements of the screen content, the pitch features that match the frequency changes that match the scene atmosphere features, and the semantic features that match the meaning of the character action association group are combined to form an audio association group. The character action association group, scene environment association group, object shape association group and audio association group are fused together to obtain the narrative logic anchor point corresponding to each video segment unit. The narrative logic anchor point contains the core association information between the picture and the audio, which is used to construct the overall picture semantic flow in the future.

3. The automatic video clip editing and splicing method for digital multimedia according to claim 1, characterized in that, The process of constructing the semantic flow of the video clips based on the narrative logic anchor points corresponding to all video clip units yields the overall semantic flow of the set of digital multimedia video clips to be edited, including: Extract the character action association group, scene environment association group, object form association group and audio association group from each narrative logic anchor point as the basic elements for constructing the semantic flow of the image; Select the narrative logic anchor point of one video segment unit as the initial basic element, perform association matching processing on the character action association group in the initial basic element, and filter out the narrative logic anchor points with continuous logic of character action association group from the narrative logic anchor points of other video segment units. The selected video clip units corresponding to the narrative logic anchors with continuity logic are arranged after the initial video clip units according to the continuity order of the character action association groups, forming a preliminary clip sequence; Adaptation analysis is performed on the scene environment association group of adjacent video segment units in the preliminary segment sequence to determine whether the spatial layout features and scene atmosphere features of adjacent scene environment association groups have transition logic. If there is transition logic, the arrangement order of adjacent video segment units is maintained. If the adjacent scene environment association group does not have transition logic, then select video segment units from the set of digital multimedia video segments to be edited that have transition logic with the two scene environment association groups before and after, insert the corresponding selected video segment units between two adjacent video segment units, and update the initial segment sequence. Continuity analysis is performed on the object shape association groups of adjacent video segment units in the updated preliminary segment sequence to determine whether the shape change sequence of adjacent object shape association groups has continuous logic. If it has continuous logic, the arrangement order of adjacent video segment units is maintained. If adjacent object shape association groups do not have continuous logic, feature adjustment processing is performed on the object shape association groups of adjacent video segment units to make the end feature of the object shape association group of the previous video segment unit and the beginning feature of the object shape association group of the next video segment unit form continuous logic, and the initial segment sequence is updated again. Rhythm matching is performed on the audio association groups of all video segment units in the newly updated initial segment sequence. If the beat intervals of the rhythm features of the audio association groups of adjacent video segment units are inconsistent, the rhythm features of the audio association groups of one of the video segment units are adjusted to make the beat intervals of the rhythm features of the audio association groups of adjacent video segment units consistent. The initial segment sequence after rhythm adjustment is subjected to overall narrative logic verification processing. By comparing the relationship between the narrative logic anchor points of all video segment units, it is determined whether the initial segment sequence conforms to the overall narrative logic. If it does, the initial segment sequence is used as the overall picture semantic flow corresponding to the set of digital multimedia video segments to be edited. If it does not conform to the overall narrative logic, then the initial basic elements are reselected, and the above-mentioned filtering, arranging, adaptation, adjustment and verification processes are repeated until an overall visual semantic flow that conforms to the overall narrative logic is obtained.

4. The automatic video clip editing and splicing method for digital multimedia according to claim 1, characterized in that, The step of performing splicing adaptation rule generation processing on adjacent video segment units in the overall image semantic stream to obtain splicing adaptation rules corresponding to adjacent video segment units includes: Extract the narrative logic anchor points of two adjacent video segment units in the overall semantic flow of the picture, and mark them as the preceding narrative logic anchor point and the following narrative logic anchor point, respectively. Action connection analysis is performed on the character action association group in the preceding narrative logic anchor point and the character action association group in the following narrative logic anchor point to determine the connection method between the ending feature of the preceding character action and the starting feature of the following character action. The connection method is determined based on the changes in the amplitude and direction of the action. Based on the aforementioned connection method, character action connection rules are generated, which are used to standardize the transition method of character actions in adjacent video segment units; Scene transition analysis is performed on the scene environment association group in the preceding narrative logic anchor point and the scene environment association group in the following narrative logic anchor point to determine the overlap relationship between the end area of ​​the spatial layout features of the preceding scene environment and the starting area of ​​the spatial layout features of the following scene environment, as well as the transition tendency between the atmosphere features of the preceding scene and the atmosphere features of the following scene. Scene environment transition rules are generated based on the overlapping relationship and transition tendency. These scene environment transition rules are used to regulate the transition mode of scene environment in adjacent video segment units. A morphological continuity analysis is performed on the object morphology association group in the preceding narrative logic anchor point and the object morphology association group in the following narrative logic anchor point to determine the correlation between the final change characteristics of the preceding object morphology and the initial change characteristics of the following object morphology. The correlation is determined based on the structural characteristics and change rate of the object morphology. Based on the aforementioned relationship, object shape continuation rules are generated, which are used to regulate the transition of object shapes in adjacent video segment units; Audio fusion analysis is performed on the audio association groups in the preceding narrative logic anchor points and the audio association groups in the following narrative logic anchor points to determine the matching relationship between the last beat of the rhythmic features of the preceding audio and the starting beat of the rhythmic features of the following audio, as well as the transition interval between the pitch features of the preceding audio and the pitch features of the following audio. Based on the matching relationship and transition interval, an audio fusion rule is generated. The audio fusion rule is used to standardize the transition mode of audio in adjacent video segment units. The character action connection rules, scene environment transition rules, object form continuation rules and audio fusion rules are integrated and processed to obtain the splicing adaptation rules corresponding to adjacent video segment units. The splicing adaptation rules include all transition specifications of adjacent video segment units in terms of picture and audio.

5. The automatic video clip editing and splicing method applied to digital multimedia according to claim 1, characterized in that, The step of performing transition fusion processing on adjacent video segment units based on the splicing adaptation rules to obtain a video segment sequence with fused transition effects includes: Extract the character action connection rules from the splicing adaptation rules to determine the transition method of character actions in adjacent video segment units; Based on the transition method of the character's actions, the last frame of the character action association group of the preceding video segment unit is subjected to action feature extension processing to generate an action transition frame. The action feature of the action transition frame is between the end feature of the preceding character action and the start feature of the following character action. The motion transition frame is inserted between the end of the preceding video segment unit and the beginning of the following video segment unit to achieve a smooth transition of the character's actions. Extract scene environment transition rules from the splicing adaptation rules to determine the transition mode of scene environment in adjacent video segment units; According to the transition method of the scene environment, spatial feature gradation processing is performed on the end area of ​​the scene environment association group of the preceding video segment unit to generate a scene transition area. The spatial layout features and scene atmosphere features of the scene transition area are between the preceding scene environment and the following scene environment. The scene transition area is integrated into the overlapping area between the end of the preceding video segment unit and the beginning of the following video segment unit to achieve scene environment transition and fusion. Extract the object shape continuation rules from the splicing adaptation rules to determine the transition method of object shapes in adjacent video segment units; Based on the transition method of the object shape, the shape gradation processing is performed on the end change features of the object shape association group of the preceding video segment unit to generate a shape transition sequence. The shape change features of the shape transition sequence are between the end change features of the preceding object shape and the beginning change features of the subsequent object shape. The morphological transition sequence is inserted between the end of the preceding video segment unit and the beginning of the following video segment unit to achieve the transition and fusion of the object's morphology. Extract the audio fusion rules from the splicing adaptation rules to determine the audio transition method in adjacent video segment units; According to the audio transition method, the rhythm gradient processing is performed on the last beat of the audio association group of the preceding video segment unit, and the pitch transition processing is performed on the pitch features of the preceding audio to generate an audio transition segment. The rhythm features and pitch features of the audio transition segment are between the preceding audio and the following audio. The audio transition segment is inserted between the end of the audio of the preceding video segment unit and the beginning of the audio of the following video segment unit to achieve audio transition and fusion. After completing the transition and fusion of character movements, scene environment, object shape and audio, adjacent video segment units containing transition frames, transition regions, transition sequences and transition segments are combined to form a segment pair with a fused transition effect. The above transition fusion process is repeated for all adjacent video segment units in the overall semantic stream to obtain a video segment sequence with fused transition effect.

6. The automatic video clip editing and splicing method for digital multimedia according to claim 2, characterized in that, The step of performing action feature association processing on the continuous action sequence, combining action features with causal relationships to form character action association groups, includes: Each action feature in a continuous action sequence is numbered according to the timeline markers to form an ordered list of action features; The first action feature in the ordered action feature list is selected as the starting action feature, and the action purpose of the starting action feature is analyzed. The action purpose is determined based on the limb movement direction and amplitude of the action. Select the first action feature after the initial action feature from the ordered action feature list as the action feature to be associated, and analyze the action purpose of the action feature to be associated. By comparing the action purpose of the initial action feature with that of the action purpose of the action feature to be associated, it is determined whether the initial action feature and the action feature to be associated have a causal relationship. If the action purpose of the action feature to be associated is a continuation or result of the action purpose of the initial action feature, then the initial action feature and the action feature to be associated have a causal relationship. If the initial action feature and the action feature to be associated have a causal relationship, then the initial action feature and the action feature to be associated are combined to form a temporary association group, and the action feature to be associated is used as the new initial action feature; If the initial action feature and the action feature to be associated do not have a causal relationship, then the initial action feature is treated as a separate temporary association group, and the action feature to be associated is treated as a new initial action feature. Repeat the above steps of selecting action features to be associated, analyzing the action purpose, judging the causal relationship, and combining temporary association groups until all action features in the ordered action feature list have been processed, resulting in multiple temporary association groups. Perform a timeline continuity check on the action features in each temporary association group to confirm whether the action features in the temporary association group are continuous on the timeline markers. If they are continuous, maintain the structure of the temporary association group. If they are not continuous, then search for the action features located between the discontinuous action features in the ordered action feature list. If there are any missing action features, then add the missing action features to the temporary association group so that the action features in the temporary association group are continuous on the timeline marker. If there are no missing action features, the discontinuous temporary association group is split into two independent temporary association groups, and the action features in each temporary association group are continuous on the timeline marker. Perform action logic integrity analysis on all processed temporary association groups to determine whether the action features in each temporary association group can fully express an independent action logic. If so, the temporary association group is identified as a character action association group. If not, then the adjacent temporary associations with continuous action logic are merged to form a new temporary association group, and the action logic integrity analysis is performed again until a character action association group that can fully express independent action logic is obtained. All obtained character action association groups are organized to form a set of character action association groups corresponding to each video segment unit, which are used to generate narrative logic anchors in the future.

7. The automatic video clip editing and splicing method for digital multimedia according to claim 3, characterized in that, The determination of whether the spatial layout features and scene atmosphere features of adjacent scene environment association groups have transition logic includes: Extract the region division information from the spatial layout features of the preceding scene environment association group to determine the core region and non-core region in the preceding scene. The core region is the main area where character actions or object shape changes occur in the scene. Extract the region division information from the spatial layout features of the subsequent scene environment association group to determine the core and non-core regions in the subsequent scene. By comparing the positional relationship between the core areas of the preceding scene and the core areas of the following scene, the overlap ratio between the core areas of the preceding scene and the core areas of the following scene in the screen coordinate system is determined, as well as the correspondence between the exit direction of the core area of ​​the preceding scene and the entrance direction of the core area of ​​the following scene. Visual style elements, including color saturation distribution, light and shadow intensity distribution, and texture density distribution, are extracted from the scene atmosphere features of the preceding scene environment association group. Visual style elements, including color saturation distribution, light and shadow intensity distribution, and texture density distribution, are extracted from the scene atmosphere features of the subsequent scene environment association group. By comparing the color saturation distribution of the preceding and subsequent scenes, the changing trends of the preceding and subsequent scenes in the main color channels are determined. The changing trends include the direction and range of increase or decrease in color saturation. By comparing the light and shadow intensity distribution of the preceding and following scenes, the changing trends of light and shadow intensity in the main areas of the image are determined, including the direction of increase and decrease and the range of change of light and shadow intensity. By comparing the texture density distribution of the preceding and following scenes, the changing trends of texture density in the core and non-core regions of the preceding and following scenes are determined, including the direction of increase and decrease and the range of change of texture density. Based on the core area positional relationship, color saturation distribution change trend, light and shadow intensity distribution change trend, and texture density distribution change trend of the preceding and following scenes, a transition logic evaluation dimension is constructed. For each evaluation dimension, the degree of correlation is described. If the overlap ratio in the core area location relationship meets the preset transition overlap standard and the exit and entrance directions correspond, and the color saturation, light and shadow intensity, and texture density all show a gradual change trend, then the adjacent scene environment association group has transition logic. If the overlap ratio in the core area location relationship does not meet the preset transition overlap standard or the exit and entrance directions do not correspond, or the color saturation, light and shadow intensity, and texture density change trends show a jump change, then the adjacent scene environment association group does not have transition logic. The correlation descriptions of each evaluation dimension are integrated to form a judgment result on whether adjacent scene environment association groups have transition logic, which is used to adjust the initial segment sequence later.

8. The automatic video clip editing and splicing method for digital multimedia according to claim 4, characterized in that, The step of generating character action connection rules based on the connection method includes: The types of character action transitions are determined based on the connection relationship between the ending characteristics of the preceding character action and the starting characteristics of the following character action. The transition types include amplitude gradual transition, directional change transition, and action pause transition. If the transition method is amplitude gradual transition, then analyze the range of change between the amplitude value of the ending feature of the preceding character's action and the amplitude value of the starting feature of the following character's action. Based on the range of change, determine the number of steps for the gradual change in amplitude. Each step corresponds to a one-time increase or decrease in amplitude value, so that the amplitude value of the preceding character's action transitions to the amplitude value of the subsequent character's action through multiple steps. The time length corresponding to each step is specified, and this time length is matched with the frame rate of the video segment unit to make the amplitude gradation process appear smooth on the screen; Based on the number of steps and the duration of each step in amplitude gradation, amplitude gradation transition rules are generated, which specify the amplitude value of the character's action at each time point. If the connection method type is directional turning connection, then analyze the turning range of the direction angle of the ending feature of the preceding character's action and the direction angle of the starting feature of the following character's action. Based on the turning range, the number of directional turning segments is determined, and each segment corresponds to one turning adjustment of the directional angle, so that the directional angle of the preceding character's action is transitioned to the directional angle of the subsequent character's action through multiple segments. The time length corresponding to each segment is specified, and this time length is matched with the frame rate of the video segment unit to make the direction turning process appear smooth on the screen; Based on the number of segments for directional turning and the time length of each segment, directional turning connection rules are generated, which specify the character's action direction angle corresponding to each time node. If the connection method type is action pause connection, then analyze the stable duration requirement of the ending feature of the preceding character's action and the preparation time requirement of the starting feature of the following character's action. Based on the stability duration requirement and the preparation time requirement, the total duration of the action pause is determined, which is equal to the sum of the stability duration and the preparation time. Within the total duration of the pause, the time period during which the ending feature of the preceding character's action remains stable, and the time period during which the starting feature of the following character's action begins to prepare are specified. Based on the total duration of the action pause, the stable time period, and the preparation time period, action pause connection rules are generated, which specify the character's action state corresponding to each time period during the action pause process. The rules for transitioning between amplitude gradients, directions, or pauses are integrated to form the rules for transitioning between character movements. These rules include coordination requirements for the movement of the character's limb joints during the transition process. These coordination requirements are set based on the natural and continuous movement patterns of the character's movements. The generated character action connection rules are subjected to screen adaptation confirmation processing. By comparing the screen resolution and frame rate of the connection rules with those of adjacent video segment units, it is determined whether the connection rules can be executed under the corresponding screen parameters. If they can be executed, the final character action connection rules are determined. If it cannot be executed, adjust the step size, number of segments, or total duration in the transition rules to match the transition rules with the screen parameters, and perform the screen compatibility confirmation process again until an executable character action transition rule is obtained.

9. The automatic video clip editing and splicing method for digital multimedia according to claim 1, characterized in that, The process of integrating the video segment sequence with the fused transition effect according to the overall semantic flow of the screen to obtain a fully edited and spliced ​​digital multimedia video includes: Extract the timeline markers of each video segment unit in the video segment sequence with the blended transition effect, and determine the start and end time points of each video segment unit; Based on the arrangement order of the semantic flow of the overall picture, the start time and end time of each video segment unit are integrated in sequence to generate a preliminary timeline. The preliminary timeline records the time position of all video segment units in the overall video. Check the time connection between adjacent video segment units in the preliminary timeline to confirm whether the end time of the previous video segment unit is continuous with the start time of the next video segment unit. If they are continuous, maintain the structure of the preliminary timeline. If the time interval is not continuous, analyze the cause of the time interval. If the time is extended due to the insertion of transition frames or transition regions in the transition fusion process, adjust the start time point of the next video segment unit to make the time axis continuous and update the preliminary time axis. Extract the image parameters of each video segment unit in the video segment sequence with blended transition effects, including image resolution, frame rate and color space, and confirm whether the image parameters of all video segment units are consistent; If there are video clip units with inconsistent image parameters, perform unified processing on the image parameters of the video clip unit, and adjust its image resolution, frame rate and color space to be consistent with the parameters of other video clip units, so as to keep the overall image parameters consistent. The audio parameters of the video clip sequence after the unified picture parameters are merged and the transition effect is processed to unify the audio parameters. The audio parameters of each video clip unit are extracted, including the sampling rate, number of channels and bit rate. The audio parameters of all video clip units are adjusted to be consistent, so that the overall audio parameters are kept consistent. The video clip sequence with unified parameters and blended transition effects is processed by data integration according to the updated preliminary timeline. The picture data and audio data of all video clip units are merged in timeline order to generate preliminary complete video data. Perform an overall narrative logic check on the preliminary complete video data. By playing the preliminary complete video data, confirm the coherence of the overall narrative logic anchor points and the consistency of the semantic flow of the screen. If they are coherent and consistent, then the preliminary complete video data is determined to be a digital multimedia video after complete editing and splicing. If there are broken narrative logic anchor points or discontinuous semantic flow in the images, analyze the video segment unit interval where the problem occurs, re-execute splicing adaptation rule generation and transition fusion processing on the video segment units in that video segment unit interval, and re-execute data integration and narrative logic check processing until a complete edited and spliced ​​digital multimedia video with coherent narrative logic and consistent semantic flow in the images is obtained.

10. An automatic video clip editing and splicing system for digital multimedia, characterized in that, The method includes a processor and a computer-readable storage medium storing machine-executable instructions that, when executed by the processor, implement the automatic video clip editing and splicing method for digital multimedia as described in any one of claims 1-9.

Citation Information

Patent Citations

  • High-quality video content automatic generation method and related equipment

    CN120050487A

  • Video intelligent self-adaptive editing method and system based on deep learning

    CN120730096A