A video intelligent decomposition method and system based on spatiotemporal causal emergent analysis

The video intelligent decomposition method based on spatiotemporal causal emergence analysis solves the problem of in-depth modeling of the causal structure and semantic evolution of events in videos, realizes hierarchical decomposition and structured representation of videos, and improves the efficiency of video understanding and event reasoning.

CN122135268APending Publication Date: 2026-06-02SHANGHAI YISIJUNMU INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI YISIJUNMU INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-02-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing video decomposition methods lack in-depth modeling of the causal structure and semantic evolution of events in videos, making it difficult to capture cross-level emergent phenomena. Furthermore, they are insufficient in depicting the interaction relationships between multiple objects and actions in complex dynamic scenes, resulting in limitations in long video understanding and event reasoning performance.

Method used

By constructing a spatiotemporal feature field, performing feature coherent state evolution and causal emergence analysis, identifying semantic vortices and performing interaction analysis, and outputting hierarchical decomposition results, including structured representations of large paragraphs, sub-stages and keyframes.

Benefits of technology

It effectively captures semantically consistent regions and boundaries in videos, automatically detects important semantic transition moments, improves the accuracy of semantic segmentation, enhances the understanding of event evolution logic, supports the parsing of complex interactive behaviors, and improves the efficiency and interpretability of video summarization, retrieval, and annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135268A_ABST
    Figure CN122135268A_ABST
Patent Text Reader

Abstract

This invention relates to a video intelligent decomposition method and system based on spatiotemporal causal emergence analysis, belonging to the field of video understanding. The method includes: inputting the original video file and constructing a continuous spatiotemporal feature field; performing feature coherence state evolution based on the spatiotemporal feature field to obtain a coherence topology graph; performing causal emergence analysis based on the spatiotemporal feature field to obtain causal critical points; identifying semantic vortices based on the spatiotemporal feature field and the coherence topology graph to obtain a series of vortex parameters; performing interaction analysis based on the vortex parameters to output an interaction matrix; and outputting hierarchical decomposition results based on the coherence topology graph, causal critical points, vortex parameters, and interaction matrix. This invention achieves intelligent video decomposition that integrates spatiotemporal features, multi-scale causal analysis, and dynamic semantic modeling to better support high-level video understanding and semantic content organization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video understanding technology, specifically relating to a video intelligent decomposition method and system based on spatiotemporal causal emergence analysis. Background Technology

[0002] In recent years, video content analysis technology has been widely applied in fields such as security monitoring, human-computer interaction, media retrieval, and autonomous driving. Traditional video decomposition methods mainly rely on techniques such as inter-frame differencing, optical flow analysis, scene detection, and object tracking, such as shot boundary detection based on color histograms, action segmentation based on motion vectors, and scene classification and behavior recognition methods based on deep learning. However, these methods still have significant shortcomings: First, most methods focus on low-level features of visual appearance or motion patterns, lacking in-depth modeling of the causal structure and semantic evolution of events in videos; second, traditional segmentation methods are often based on thresholding or clustering, which is insufficient for characterizing the interaction relationships of multiple objects and actions in complex dynamic scenes; third, existing methods mostly use single-timescale analysis, making it difficult to capture cross-level emergent phenomena from micro-actions to macro-semantics. In addition, most methods fail to model continuous events in videos as semantic units with inherent dynamic characteristics, resulting in limited performance in tasks such as long video understanding, event reasoning, and structured summarization. Summary of the Invention

[0003] To address the aforementioned problems in the existing technology, this invention provides a video intelligent decomposition method and system based on spatiotemporal causal emergence analysis.

[0004] The objective of this invention can be achieved through the following technical solutions: A video intelligent decomposition method based on spatiotemporal causal emergence analysis, the implementation of which includes the following steps: Step S1: Input the original video file and construct a continuous spatiotemporal feature field; Step S2: Based on the spatiotemporal feature field, perform feature coherence state evolution to obtain a coherence topology map; Step S3: Perform causal emergence analysis based on the spatiotemporal characteristic field to obtain the causal critical point; Step S4: Identify semantic vortices based on the spatiotemporal feature field and the coherence topology map to obtain a series of vortex parameters; Step S5: Perform interaction analysis based on the vortex parameters and output the interaction matrix; Step S6: Output hierarchical decomposition results based on the coherence topology graph, the causal critical point, the vortex parameter, and the interaction matrix.

[0005] Preferably, the construction of the spatiotemporal feature field in step S1 specifically involves: Import the original video file and extract the feature vector F(x,y,t,d) for each frame, where d is the feature dimension; The spatiotemporal feature field is obtained by iteratively updating the feature diffusion equation. Mathematically described ,in, For time t The spatiotemporal characteristic field at that location, For time t The feature vector at that location, The diffusion coefficient is... for Pilates operators, The attraction coefficient, It is a significance function.

[0006] Preferably, the evolution of the characteristic coherent state in step S2 specifically includes: For each point in the spatiotemporal feature field, its neighborhood N(x,y,t) is taken, and the average coherence between that point and all its neighboring points is calculated as the local coherence of that point. Mathematically, this is described as follows: ,in, For time t The local coherence of a point, where N is the number of neighboring points. Let i be the spatiotemporal feature field at the i-th neighboring point. For time t The spatiotemporal characteristic field at that location, For vector dot product, It is the vector norm; Based on the local coherence, the spatiotemporal domain is divided into multiple regions to obtain the coherence topology map.

[0007] Preferably, the mathematical description of the causal emergence analysis in step S3 is as follows: Define time scales, obtain variables at each time scale based on the spatiotemporal feature field, and measure the degree of interdependence between two variables by calculating mutual information; The cross-scale causal flow is calculated and mathematically described as follows: ,in, For cross-scale causal flow between variables A and B, Let M be the time delay, and M be the set of all possible mediator variables. Let A and t be variables at time t. Mutual information between variables B at time t. Let C be the mediator variable at time t. Mutual information between variables B at time points; Obtain causal flow time series It records the causal flow strength for each pair of variables; The phase transition intensity is calculated based on the causal flow time series from the micro to the macroscopic level, mathematically described as follows: ,in, Let be the phase transition intensity at time t. for Cross-scale causal flow of time, for Cross-scale causal flow of time, for Cross-scale causal flow of time, Let V be the variance of the CF sequence; when When the value exceeds a preset threshold, t is marked as the causal critical point.

[0008] Preferably, the identification of semantic vortexes in step S4 specifically involves: For each location, the first two dimensions of the spatiotemporal feature field are taken, and the velocity field of the spatiotemporal feature field in these two dimensions is denoted as... and And calculate the curl of the velocity field at that location; Based on the coherence topology diagram, the screening criteria for vortex centers are defined as follows: the curl is a local maximum and the local coherence is greater than a preset threshold; a series of vortex centers are screened out and their intensity and radius of influence are recorded; Output a series of vortex parameters, including center position, intensity, phase, radius of influence, and lifetime.

[0009] Preferably, the interaction analysis in step S5 specifically includes: The interaction potential energy between the vortices is calculated and mathematically described as follows: ,in, Let be the potential energy of the interaction between vortex i and vortex j. The gravitational constant, Let i be the intensity of vortex i. Let j be the intensity of the vortex. Let be the normalized distance between vortex i and vortex j. To prevent small constants from being divided by zero, For phase coupling strength, For the attenuation length, Let i be the phase of vortex i. Let be the phase of vortex j; The interaction matrix is ​​output based on the interaction potential energy.

[0010] Preferably, step S6 specifically comprises: At the macro level, the video is decomposed into large segments based on the causal critical point; within each large segment, it is divided into multiple sub-stages based on the vortex parameter and the interaction matrix; in each sub-stage, local maxima points in the coherence topology graph are found as keyframes.

[0011] A video intelligent decomposition system based on spatiotemporal causal emergence analysis is used to execute the video intelligent decomposition method based on spatiotemporal causal emergence analysis described above. It includes a feature field construction module, a feature coherence state evolution module, a causal emergence analysis module, a semantic vortex recognition module, an interaction analysis module, and a hierarchical decomposition module. The feature field construction module is used to input the original video file and construct a continuous spatiotemporal feature field; The characteristic coherence state evolution module is used to perform characteristic coherence state evolution based on the spatiotemporal characteristic field to obtain a coherence degree topology map; The causal emergence analysis module is used to perform causal emergence analysis based on the spatiotemporal characteristic field to obtain the causal critical point; The semantic vortex recognition module is used to identify semantic vortices based on the spatiotemporal feature field and the coherence topology map, and obtain a series of vortex parameters; The interaction analysis module is used to perform interaction analysis based on the vortex parameters and output the interaction matrix. The hierarchical decomposition module is used to output hierarchical decomposition results based on the coherence topology graph, the causal critical point, the vortex parameter, and the interaction matrix.

[0012] The beneficial effects of this invention are as follows: (1) By constructing a spatiotemporal feature field and performing feature coherence state evolution, it is possible to effectively capture semantically consistent regions and boundaries in the video, thereby improving the accuracy of semantic segmentation; (2) By identifying cross-scale causal flows and causal critical points through causal emergence analysis, important semantic turning points in videos can be automatically detected, enhancing the understanding of the logic of event evolution; (3) Through semantic vortex modeling and interaction analysis, continuous actions or events can be represented as dynamic units with physical meaning, and the attraction, repulsion or synchronization relationship between them can be characterized, thereby supporting the analysis of complex interactive behaviors; (4) By integrating coherence topology, causal critical point, vortex parameter and interaction matrix, the video is decomposed hierarchically and outputs a structured representation including large segments, sub-stages and keyframes, which significantly improves the efficiency and interpretability of tasks such as video summarization, retrieval, annotation and narrative understanding. Attached Figure Description

[0013] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0014] Figure 1 This is a flowchart of the steps of a video intelligent decomposition method based on spatiotemporal causal emergence analysis according to the present invention. Detailed Implementation

[0015] To better understand the invention, various aspects of the invention will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely illustrative of exemplary embodiments of the invention and are not intended to limit the scope of the invention in any way. Throughout the specification, the expression "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, the terms "approximately," "about," and similar terms are used as expressions of approximation, not as expressions of degree, and are intended to describe inherent deviations in measured or calculated values ​​that will be recognized by those skilled in the art. Furthermore, the order in which the steps are described in this invention does not necessarily indicate the order in which these steps occur in actual operation, unless otherwise expressly defined or deduced from the context.

[0016] It should also be understood that expressions such as "comprising," "including," "having," "containing," and / or "comprising" are open-ended rather than closed-ended expressions in this specification, indicating the presence of the stated features, elements, and / or components, but not excluding the presence of one or more other features, elements, components, and / or combinations thereof. Furthermore, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features, not just individual elements in the list. Additionally, when describing embodiments of the invention, the word "may" is used to mean "one or more embodiments of the invention." And the term "exemplary" is intended to refer to examples or illustrations.

[0017] Unless otherwise specified, all terms used herein (including engineering and technical terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that, unless expressly stated herein, terms defined in common dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not in an idealized or overly formalized sense.

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other. The invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] Example 1: Please see Figure 1 A video intelligent decomposition method based on spatiotemporal causal emergence analysis includes: Step S1: Input the original video file and construct a continuous spatiotemporal feature field; Step S2: Based on the spatiotemporal feature field, perform feature coherence state evolution to obtain a coherence topology map, which marks which regions are highly consistent internally (semantic condensation regions) and which are boundaries (semantic boundaries). Step S3: Perform causal emergence analysis based on the spatiotemporal characteristic field to obtain the causal critical point; Step S4: Identify semantic vortices (continuous actions or events in the video) based on the spatiotemporal feature field and the coherence topology graph, and obtain a series of vortex parameters; Step S5: Perform interaction analysis based on the vortex parameters (determine whether the vortices attract, repel, or are independent), and output the interaction matrix; Step S6: Output hierarchical decomposition results based on the coherence topology graph, the causal critical point, the vortex parameter, and the interaction matrix.

[0020] In this embodiment, the construction of the spatiotemporal feature field is specifically as follows: S101: Import the original video file and treat the original video as four-dimensional data: width (x), height (y), time (t), and color channels (RGB). Extract the feature vector F(x,y,t,d) of each frame image through a pre-trained deep neural network, where d is the feature dimension and the feature vector includes object features, motion features, color features, and texture features, etc. S102: Obtain the spatiotemporal feature field by iteratively updating the feature diffusion equation. This makes the features smooth in space and time, while also clustering towards salient regions, mathematically described as... ,in, At time t The spatiotemporal characteristic field at that location, At time t The eigenvector at that location (i.e., the initial one) ), The diffusion coefficient is... for Pilates operators, The attraction coefficient, The significance function is 0-1. Example: Upload a video of tomatoes. Assume that at the initial moment, the chopping action begins, and the feature value is high at the knife's position. After diffusion, the knife's feature expands to the surrounding pixels, and the value is high in the motion region. Finally, the feature field forms a high-value region along the knife's motion trajectory.

[0021] In this embodiment, the evolution of the characteristic coherent state specifically refers to: S201: For each point in the spatiotemporal feature field, take its neighborhood N(x,y,t), and calculate the average coherence between that point and all its neighboring points as the local coherence of that point. Mathematically, this is described as follows: ,in, For time t The local coherence of a point, where N is the number of neighboring points. Let i be the spatiotemporal feature field at the i-th neighboring point. For time t The spatiotemporal characteristic field at that location, For vector dot product, It is the vector norm; S202: Based on the local coherence, the spatiotemporal domain is divided into multiple regions. Points within each region have high coherence, while the coherence between regions is low, resulting in the coherence topology map. Continuing the previous example, for the same frame, the spatial variation of coherence can distinguish different regions in the image. For example, the coherence of the knife's motion trajectory region may reach 0.9, the coherence of the stationary tomato region may reach 0.8, and the coherence of the background region may only be 0.3. For the same spatial location, the variation of local coherence over time can distinguish the semantic continuity of that location at different times.

[0022] In this embodiment, the mathematical description of the causal emergence analysis is as follows: S301: Define time scales, including but not limited to microscale (sampling at the original frame rate, where variables are feature vectors at each spatial location), mesoscale (sampling at a lower frame rate, where variables are features obtained through spatial pooling), and macroscale (sampling at an even lower frequency, where variables are global features of the entire frame). Based on the spatiotemporal feature field, obtain variables at each of the aforementioned time scales. Measure the degree of interdependence between two variables by calculating mutual information, mathematically described as follows: ,in, For mutual information between variables A and B, Let be the probability that variable A takes the value a. Let b be the probability that variable B takes the value b. Let A be the joint probability that variable A takes the value a and variable B takes the value b; S302: The cross-scale causal flow is calculated and mathematically described as follows: ,in, For cross-scale causal flow between variables A and B, M represents the time delay, and M is the set of all possible mediating variables (variables that are in the middle of the causal path). Let A and t be variables at time t. Mutual information between variables B at time t. Let C be the mediator variable at time t. The mutual information between variables B at any given time, for example, micro-variable A may affect meso-variable C, and then C may affect macro-variable B; that is, from the total information flow from A to B, subtract the information flow through the strongest mediator variable C, what remains is the direct causal flow; S303: Obtain the causal flow time series It records the causal flow strength of each pair of variables (micro to macro, meso to macro, etc.); Example: Given a cooking video, we focus on how hand movements (micro) affect the cooking stage (macro). Micro variable A: the amplitude of hand movements per frame; meso variable C: the intensity of the cutting motion per second; macro variable B: the cooking stage score every 10 seconds; Assuming the mutual information between hand movements and the cooking stage is 1.2 bits, and the mutual information between the cutting motion intensity and the cooking stage is 0.9 bits, then the cross-scale causal flow between A and B is 0.3 bits. That is, the total information contribution of hand movements to the cooking stage is 1.2 bits, of which 0.9 bits are passed through the mediating variable of the cutting motion intensity, and the remaining 0.3 bits are the direct influence of hand movements on the cooking stage; S304: Calculate the phase transition intensity (which best reflects the fundamental changes in video semantics) based on the causal flow time series from micro to macro levels. Mathematically, this is described as follows: ,in, Let be the phase transition intensity at time t. for Cross-scale causal flow of time, for Cross-scale causal flow of time, for Cross-scale causal flow of time, Let V be the variance of the CF sequence; when When the value exceeds a preset threshold, t is marked as the causal critical point.

[0023] In this embodiment, the identification of the semantic vortex specifically involves: S401: For each location, take the first two dimensions of the spatiotemporal feature field (two fixed dimensions can be used, such as the first and second dimensions, or the mean of the eigenvectors can be used), and calculate the velocity field (rate of change with time) of the spatiotemporal feature field in these two dimensions, denoted as... and And calculate the curl of the velocity field at that location (in three-dimensional space, curl is a vector, but here we only care about its magnitude); S402: Based on the coherence topology map, define the screening conditions for vortex centers: the curl is a local maximum and the local coherence is greater than a preset threshold (vortices are more likely to appear in areas with high coherence); screen out a series of vortex centers and record the intensity (curl value) and the radius of influence (the distance from the center until the curl value is lower than a certain threshold or the coherence is lower than the threshold). S403: Outputs a series of vortex parameters, including center position, intensity, phase, radius of influence, and lifetime (start time and end time).

[0024] In this embodiment, the interaction analysis specifically includes: The interaction potential energy between the vortices is calculated and mathematically described as follows: ,in, Let be the potential energy of the interaction between vortex i and vortex j. The gravitational constant, Let i be the intensity of vortex i. Let j be the intensity of the vortex. Let be the normalized distance between vortex i and vortex j. To prevent small constants from being divided by zero, For phase coupling strength, For the attenuation length, Let i be the phase of vortex i. Let be the phase of vortex j; The interaction matrix is ​​output based on the interaction potential energy.

[0025] In this embodiment, step S6 can be implemented through the following steps: At a macroscopic level, the video is decomposed into large segments based on the causal critical point; within each large segment, it is further divided into multiple sub-stages based on the vortex parameters and the interaction matrix. For example, if two vortices... If the coherence is less than -0.5 (strong attraction) and their lifecycles overlap, then they belong to the same semantic unit (sub-stage); in each sub-stage, the local maximum point (i.e. the moment with the highest consistency) in the coherence topology graph is found as the keyframe.

[0026] Example 2: A video intelligent decomposition system based on spatiotemporal causal emergence analysis includes a feature field construction module, a feature coherence state evolution module, a causal emergence analysis module, a semantic vortex recognition module, an interaction analysis module, and a hierarchical decomposition module. The feature field construction module is used to input the original video file and construct a continuous spatiotemporal feature field; The feature coherence state evolution module is used to perform feature coherence state evolution based on the spatiotemporal feature field to obtain a coherence topology map, which marks which regions are highly consistent internally (semantic condensation regions) and which are boundaries (semantic boundaries). The causal emergence analysis module is used to perform causal emergence analysis based on the spatiotemporal characteristic field to obtain the causal critical point; The semantic vortex recognition module is used to identify semantic vortices (continuous actions or events in the video) based on the spatiotemporal feature field and the coherence topology graph, and obtain a series of vortex parameters. The interaction analysis module is used to perform interaction analysis based on the vortex parameters (determining whether the vortices attract, repel, or are independent) and output the interaction matrix. The hierarchical decomposition module is used to output hierarchical decomposition results based on the coherence topology graph, the causal critical point, the vortex parameter, and the interaction matrix.

[0027] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A video intelligent decomposition method based on spatiotemporal causal emergence analysis, characterized in that, Includes the following steps: Step S1: Input the original video file and construct a continuous spatiotemporal feature field; Step S2: Based on the spatiotemporal feature field, perform feature coherence state evolution to obtain a coherence topology map; Step S3: Perform causal emergence analysis based on the spatiotemporal characteristic field to obtain the causal critical point; Step S4: Identify semantic vortices based on the spatiotemporal feature field and the coherence topology map to obtain a series of vortex parameters; Step S5: Perform interaction analysis based on the vortex parameters and output the interaction matrix; Step S6: Output hierarchical decomposition results based on the coherence topology graph, the causal critical point, the vortex parameter, and the interaction matrix.

2. The video intelligent decomposition method based on spatiotemporal causal emergence analysis according to claim 1, characterized in that, The construction of the spatiotemporal feature field in step S1 is specifically as follows: Import the original video file and extract the feature vector F(x,y,t,d) for each frame, where d is the feature dimension; The spatiotemporal feature field is obtained by iteratively updating the feature diffusion equation. Mathematically described ,in, At time t The spatiotemporal characteristic field at that location, At time t The feature vector at that location, The diffusion coefficient is... for Pilates operators, The attraction coefficient, It is a significance function.

3. The video intelligent decomposition method based on spatiotemporal causal emergence analysis according to claim 1, characterized in that, The evolution of the characteristic coherent state in step S2 specifically refers to: For each point in the spatiotemporal feature field, its neighborhood N(x,y,t) is taken, and the average coherence between that point and all its neighboring points is calculated as the local coherence of that point. Mathematically, this is described as follows: ,in, At time t The local coherence of a point, where N is the number of neighboring points. Let i be the spatiotemporal feature field at the i-th neighboring point. At time t The spatiotemporal characteristic field at that location, For vector dot product, It is the vector norm; Based on the local coherence, the spatiotemporal domain is divided into multiple regions to obtain the coherence topology map.

4. The video intelligent decomposition method based on spatiotemporal causal emergence analysis according to claim 1, characterized in that, The mathematical description of the causal emergence analysis in step S3 is as follows: Define time scales, obtain variables at each time scale based on the spatiotemporal feature field, and measure the degree of interdependence between two variables by calculating mutual information; The cross-scale causal flow is calculated and mathematically described as follows: ,in, For cross-scale causal flow between variables A and B, Let M be the time delay, and M be the set of all possible mediator variables. Let A and t be variables at time t. Mutual information between variables B at time t. Let C be the mediator variable at time t. Mutual information between variables B at time points; Obtain causal flow time series It records the causal flow strength for each pair of variables; The phase transition intensity is calculated based on the causal flow time series from the micro to the macroscopic level, mathematically described as follows: ,in, Let be the phase transition intensity at time t. for Cross-scale causal flow of time, for Cross-scale causal flow of time, for Cross-scale causal flow of time, Let V be the variance of the CF sequence; when When the value exceeds a preset threshold, t is marked as the causal critical point.

5. The video intelligent decomposition method based on spatiotemporal causal emergence analysis according to claim 1, characterized in that, The identification of semantic vortex in step S4 specifically involves: For each location, the first two dimensions of the spatiotemporal feature field are taken, and the velocity field of the spatiotemporal feature field in these two dimensions is denoted as... and And calculate the curl of the velocity field at that location; Based on the coherence topology diagram, the screening criteria for vortex centers are defined as follows: the curl is a local maximum and the local coherence is greater than a preset threshold; a series of vortex centers are screened out and their intensity and radius of influence are recorded; Output a series of vortex parameters, including center position, intensity, phase, radius of influence, and lifetime.

6. The video intelligent decomposition method based on spatiotemporal causal emergence analysis according to claim 1, characterized in that, The interaction analysis in step S5 specifically includes: The interaction potential energy between the vortices is calculated and mathematically described as follows: ,in, Let be the potential energy of the interaction between vortex i and vortex j. The gravitational constant, Let i be the intensity of vortex i. Let j be the intensity of the vortex. Let be the normalized distance between vortex i and vortex j. To prevent small constants from being divided by zero, For phase coupling strength, For the attenuation length, Let i be the phase of vortex i. Let be the phase of vortex j; The interaction matrix is ​​output based on the interaction potential energy.

7. The video intelligent decomposition method based on spatiotemporal causal emergence analysis according to claim 1, characterized in that, Step S6 specifically involves: At the macro level, the video is decomposed into large segments based on the causal critical point; within each large segment, it is divided into multiple sub-stages based on the vortex parameter and the interaction matrix; in each sub-stage, local maxima points in the coherence topology graph are found as keyframes.

8. A video intelligent decomposition system based on spatiotemporal causal emergence analysis, characterized in that, The system is applied to the video intelligent decomposition method based on spatiotemporal causal emergence analysis as described in any one of claims 1-7. It includes a feature field construction module, a feature coherent state evolution module, a causal emergence analysis module, a semantic vortex recognition module, an interaction analysis module, and a hierarchical decomposition module; The feature field construction module is used to input the original video file and construct a continuous spatiotemporal feature field; The characteristic coherence state evolution module is used to perform characteristic coherence state evolution based on the spatiotemporal characteristic field to obtain a coherence degree topology map; The causal emergence analysis module is used to perform causal emergence analysis based on the spatiotemporal characteristic field to obtain the causal critical point; The semantic vortex recognition module is used to identify semantic vortices based on the spatiotemporal feature field and the coherence topology map, and obtain a series of vortex parameters; The interaction analysis module is used to perform interaction analysis based on the vortex parameters and output the interaction matrix. The hierarchical decomposition module is used to output hierarchical decomposition results based on the coherence topology graph, the causal critical point, the vortex parameter, and the interaction matrix.