Variable-gravity particle experiment process detection method based on double-flow difference

Through the experimental process detection method of variable gravity particles based on dual-stream differential, sliding window sampling and spatiotemporal feature pyramid fusion technology, the problem of low stage process identification efficiency in experimental video data is solved, and efficient and accurate process detection is achieved.

CN119963932AActive Publication Date: 2025-05-09TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510444636.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately determine the multiple stages of the experiment of spatially variable gravity particle material based on experimental video data, resulting in inefficient data processing.

Method used

Using the experimental process detection method of variable gravity particles based on dual-stream differential, by obtaining experimental video data, performing sliding window sampling, extracting the feature vectors of RGB stream and RGB differential stream, constructing a spatiotemporal feature pyramid, and performing weighted fusion to determine the category and time range of the experimental process.

Benefits of technology

It realizes rapid and accurate identification of experimental processes, improves data processing efficiency, and better senses the timing changes of the experimental process by introducing RGB differential flow, improving the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963932A_ABST
    Figure CN119963932A_ABST
Patent Text Reader

Abstract

The invention provides a variable gravity particle experiment process detection method based on double-flow difference, and relates to the technical field of video data processing. According to the method provided by the invention, a plurality of video samples are determined through a sliding window sampling mode according to experiment video data of a variable gravity particle experiment, then spatial-temporal feature pyramids are constructed according to RGB streams and RGB difference streams of each video sample, and then a fused spatial-temporal feature pyramid is determined according to the spatial-temporal feature pyramids; the fused spatiotemporal feature pyramid comprises a plurality of spatiotemporal feature vectors with different time scales, determining a process detection result of each video sample according to the fused spatiotemporal feature pyramid, and finally obtaining process categories included in the experiment video data of the variable gravity particle experiment and starting time and ending time of each process category. The method can quickly and accurately determine a plurality of stage processes included in the space variable gravity granular material experiment according to the experiment video data, and improves the data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data processing, and in particular to a variable gravity particle experimental process detection method based on dual-stream difference. Background Art

[0002] The variable gravity particle experiment is a space variable gravity particle material experiment, which is a scientific experiment to study the behavior of granular materials in a microgravity or variable gravity environment. Granular materials exhibit specific flow, accumulation, compression and other characteristics in the Earth's gravity environment, but their behavior may change significantly in microgravity or different gravity environments. Since the experiment lasts for a long time and includes multiple different stages, it is necessary to detect and identify each stage of the experiment based on the experimental video data, so that technicians can quickly and accurately identify specific stages from the experimental video, thereby providing strong data support for scientific researchers, assisting them in subsequent data analysis and promoting the progress of scientific research.

[0003] Therefore, there is an urgent need for a variable gravity particle experimental process detection method based on dual-stream differential, which can quickly and accurately determine the multiple stage processes included in the space variable gravity particle material experiment based on the experimental video data and improve data processing efficiency. Summary of the invention

[0004] The embodiment of the present invention provides a variable gravity particle experiment process detection method based on dual-stream differential, which can quickly and accurately determine the multiple stage processes included in the space variable gravity particle material experiment according to the experimental video data, thereby improving data processing efficiency.

[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions: In a first aspect, a method for detecting a variable gravity particle experiment process based on dual-stream difference is provided, the method comprising: obtaining experimental video data of the variable gravity particle experiment; based on a preset step size, performing sliding window sampling processing on the experimental video data at a preset sampling frequency to obtain multiple video samples, each video sample comprising an RGB stream and an RGB difference stream of T frame images, wherein the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1; for any video sample, performing feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L with different semantic levels corresponding to the RGB stream 0 and the eigenvector L 1 ; and the feature vectors P with different semantic levels corresponding to the RGB difference stream 0 and the eigenvector P 1 ; According to the feature vector L 0 , eigenvector L 1 , eigenvector P 0and the eigenvector P 1 A spatiotemporal feature pyramid of RGB stream and RGB difference stream is constructed, wherein the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales. A weighted fusion is performed on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each video sample. A process detection result of the experimental video data is determined through a trained process detection unit according to the fused spatiotemporal feature pyramid corresponding to each video sample, wherein the process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.

[0006] In a possible implementation of the first aspect, based on a preset step size, a sliding window sampling process is performed on the experimental video data at a preset sampling frequency to obtain a plurality of video samples, including: downsampling the experimental video data at a preset sampling frequency f, and performing a sliding window sampling process based on a preset step size s to obtain a plurality of samples including T frame images, wherein s is less than T, and s and f are positive integers; determining an RGB stream of each sample including the T frame image; performing a differential process on the subsequent frame and the previous frame image of each two adjacent frames of the RGB stream of each sample including the T frame image to obtain an RGB differential stream of each sample including the T frame image, so as to obtain a plurality of video samples; RGB Stream The data structure is: ; RGB differential flow The data structure is: ; Among them, H is the resolution height of the image, W is the resolution width of the image, and 3 is the number of channels of the image.

[0007] In a possible implementation manner of the first aspect, for any video sample, a feature extraction operation is performed on the RGB stream and the RGB difference stream to obtain feature vectors L having different semantic levels corresponding to the RGB stream. 0 and the eigenvector L 1 ; and the feature vectors P with different semantic levels corresponding to the RGB difference stream 0 and the eigenvector P 1 , including: for any video sample, input the RGB stream into the dilated 3D convolutional network and output the feature vector A through the first node 0 , output feature vector A through the second node 1 ; For the feature vector A 0 and the eigenvector A 1Perform non-local operations and 3D convolution operations respectively to obtain the feature vector L 0 and the eigenvector L 1 ; For any video sample, the RGB difference stream is input into the dilated 3D convolutional network and the feature vector B is output through the first node 0 , output feature vector B through the second node 1 ; For the feature vector B 0 and the eigenvector B 1 Perform non-local operations and 3D convolution operations respectively to obtain the feature vector P 0 and the eigenvector P 1 ; Eigenvector L 0 The data structure is: L 0 ; Eigenvector L 1 The data structure is: L 1 ; Eigenvector P 0 The data structure is: P 0 ; Eigenvector P 1 The data structure is: P 1 ; Where C is the number of channels.

[0008] In a possible implementation of the first aspect, the spatiotemporal feature pyramid of the RGB stream includes n spatiotemporal feature vectors with different time scales: The formula for determining is: ; ; ; L 0 , L 1 ; When i is greater than 0, ; The spatiotemporal feature pyramid of the RGB difference stream includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; ; P0 , P 1 ; When i is greater than 0, ; in, It is the time series change information; is the weighted feature vector of the RGB stream; is the activation function, is a 1D convolution operation with a convolution kernel of size 1, is a linear rectification function; are the learnable weight parameters of convolution kernels of different sizes, 1D convolution operation with kernel size k, k=1, 3, 5 and 7; is a 1D convolution operation with a convolution kernel of size 3; LN refers to processing through the LayerNorm layer; ,and =1,2,...,n; Spatiotemporal feature vector and the spatiotemporal eigenvector The time scale is: ; ; .

[0009] In a possible implementation of the first aspect, the n spatiotemporal feature vectors with different time scales included in the fusion spatiotemporal feature pyramid are The formula for determining is: ; in, and is a learnable weight parameter.

[0010] In a possible implementation of the first aspect, the process detection unit includes a coarse detection subunit and a fine detection subunit: the coarse detection subunit is used to determine multiple proposals based on a fused spatiotemporal feature pyramid corresponding to each video sample, each proposal including a left and right boundary distance and a process category; the fine detection subunit is used to determine the boundary features of each proposal based on a saliency refinement algorithm, optimize the boundary position of each proposal based on the boundary features of each proposal, and obtain the process detection result included in each video sample.

[0011] In a possible implementation of the first aspect, before determining the process detection result of the experimental video data through a trained process detection unit according to the fused spatiotemporal feature pyramid corresponding to each video sample, the method further includes: obtaining a training sample set of a variable gravity particle experiment, the training sample set including multiple training samples, each training sample including a video sample and a process category; constructing a target loss function; and based on the target loss function, iteratively training the process detection unit to obtain a trained process detection unit.

[0012] In a possible implementation of the first aspect, the target loss function for: ; in, The classification loss function of the coarse detection subunit, Regression loss function for coarse detection subunits; The classification loss function of the fine detection subunit, Regression loss function for fine-grained detection of subunits.

[0013] The beneficial effects of the present invention are as follows: the method provided by the present invention determines multiple video samples by sliding window sampling based on the experimental video data of the variable gravity particle experiment, and then constructs a spatiotemporal feature pyramid based on the RGB stream and the RGB difference stream of each video sample, and then constructs a spatiotemporal feature pyramid based on the RGB stream and the RGB difference stream for weighted fusion to obtain a fused spatiotemporal feature pyramid, and finally determines the process category of each video sample based on the multiple spatiotemporal feature vectors included in the fused spatiotemporal feature pyramid, and then obtains the process category included in the experimental video data of the variable gravity particle experiment, as well as the start time and end time of each process category. The method provided by the present invention can quickly and accurately determine the multiple stage processes included in the spatial variable gravity particle material experiment based on the experimental video data, thereby improving data processing efficiency. On the other hand, the method provided by the present invention introduces RGB differential stream, and can construct RGB differential stream to replace optical stream according to the characteristic that the background change of experimental video data is not large, so as to better perceive the temporal changes of experimental process categories; and, the method provided by the present invention constructs the spatiotemporal feature pyramid of RGB stream and RGB differential stream based on the temporal change information of RGB differential stream, which can effectively strengthen the learning on RGB stream, so as to improve the perception ability of experimental process on RGB stream; the spatiotemporal feature pyramid provided by the present invention includes feature vectors of multiple time scales, and uses learnable weighted fusion to construct a fused spatiotemporal feature pyramid, so as to capture richer spatiotemporal semantic information, which can effectively improve the accuracy of detection and meet the usage needs of technicians in different usage scenarios.

[0014] In a second aspect, the present invention provides a variable gravity particle experiment process detection system based on dual-stream differential, the system comprising: a data acquisition unit, used to acquire experimental video data of the variable gravity particle experiment; a sample sampling unit, used to perform sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size, to obtain multiple video samples, each video sample comprising an RGB stream and an RGB differential stream of T frame images, wherein the RGB differential stream is obtained by differential processing of the RGB stream, and T is a positive integer greater than 1; a feature extraction unit, used to perform feature extraction operations on the RGB stream and the RGB differential stream for any video sample, to obtain feature vectors L corresponding to the RGB stream with different semantic levels 0 and the eigenvector L 1 ; and the feature vectors P with different semantic levels corresponding to the RGB difference stream 0 and the eigenvector P 1 ; Feature construction unit, used to construct the feature vector L 0 , eigenvector L 1 , eigenvector P 0 and the eigenvector P 1 A spatiotemporal feature pyramid of RGB stream and RGB difference stream is constructed, wherein the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales; a feature fusion unit is used to perform weighted fusion on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream corresponding to each video sample, so as to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each video sample; a detection unit is used to determine a process detection result of the experimental video data according to the process detection unit completed through training based on the fused spatiotemporal feature pyramid corresponding to each video sample, wherein the process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.

[0015] According to a third aspect, an electronic device is provided, comprising a memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code comprises computer instructions, and when the computer instructions are executed by the processor, the electronic device executes a method as in any implementation of the first aspect.

[0016] According to a fourth aspect, a computer-readable storage medium is provided, comprising computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method in any implementation of the first aspect.

[0017] According to a fifth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to execute the method in any implementation of the first aspect.

[0018] It can be understood that the beneficial effects that can be achieved by the system of the second aspect, the electronic device of the third aspect, the computer-readable storage medium of the fourth aspect, and the computer program product of the fifth aspect provided above can be referred to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic diagram of a process of a variable gravity particle experimental process detection method based on dual flow difference provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention; Figure 3 A flow chart of a variable gravity particle experimental process detection method based on dual flow difference provided in an embodiment of the present invention; Figure 4 A flow chart of another variable gravity particle experimental process detection method based on dual flow difference provided by an embodiment of the present invention; Figure 5 A schematic diagram of a process for generating a feature vector provided by an embodiment of the present invention; Figure 6 A schematic diagram of a process for generating a spatiotemporal feature pyramid provided by an embodiment of the present invention; Figure 7 A schematic diagram of the structure of a change information guidance module provided by an embodiment of the present invention; Figure 8 A schematic diagram of a process for generating a fused spatiotemporal feature pyramid provided by an embodiment of the present invention; Fig. 9 A flow chart of another variable gravity particle experimental process detection method based on dual flow difference provided in an embodiment of the present invention; Fig.10 A schematic diagram of the process categories and duration of a spatial variable gravity granular material experiment for a validation data set provided by an embodiment of the present invention; Fig.11 A schematic diagram of the structure of a detection system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The technical solution in the embodiment of the present invention will be described below in conjunction with the accompanying drawings in the embodiment of the present invention. In the description of the present invention, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B can represent A or B; the "or" in the present invention is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In addition, in the description of the present invention, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items.

[0021] In addition, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish the same or similar items with substantially the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference.

[0022] Meanwhile, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being better or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0023] The variable gravity particle experiment is a space variable gravity particle material experiment, which is a scientific experiment to study the behavior of granular materials in a microgravity or variable gravity environment. Granular materials exhibit specific flow, accumulation, compression and other characteristics in the Earth's gravity environment, but their behavior may change significantly in microgravity or different gravity environments. Since the experiment lasts for a long time and includes multiple different stages, it is necessary to detect and identify each stage of the experiment based on the experimental video data, so that technicians can quickly and accurately identify specific stages from the experimental video, thereby providing strong data support for scientific researchers, assisting them in subsequent data analysis and promoting the progress of scientific research.

[0024] Therefore, there is an urgent need for a variable gravity particle experimental process detection method based on dual-stream differential, which can quickly and accurately determine the multiple stage processes included in the space variable gravity particle material experiment based on the experimental video data and improve data processing efficiency.

[0025] In view of this, an embodiment of the present invention provides a variable gravity particle experiment process detection method based on dual-stream difference, the method comprising: obtaining experimental video data of the variable gravity particle experiment; based on a preset step size, performing sliding window sampling processing on the experimental video data at a preset sampling frequency to obtain multiple video samples, each video sample comprising an RGB stream and an RGB difference stream of T frame images, wherein the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1; for any video sample, performing feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L with different semantic levels corresponding to the RGB stream 0 and the eigenvector L 1 ; and the feature vectors P with different semantic levels corresponding to the RGB difference stream 0 and the eigenvector P 1 ; According to the feature vector L 0 , eigenvector L 1 , eigenvector P 0 and the eigenvector P 1 A spatiotemporal feature pyramid of RGB stream and RGB difference stream is constructed, wherein the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales. A weighted fusion is performed on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each video sample. A process detection result of the experimental video data is determined through a trained process detection unit according to the fused spatiotemporal feature pyramid corresponding to each video sample, wherein the process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.

[0026] The method provided by the present invention determines multiple video samples through sliding window sampling based on the experimental video data of the variable gravity particle experiment, and then constructs a spatiotemporal feature pyramid based on the RGB stream and the RGB difference stream of each video sample, and then constructs a spatiotemporal feature pyramid based on the RGB stream and the RGB difference stream for weighted fusion to obtain a fused spatiotemporal feature pyramid, and finally determines the process category of each video sample based on the multiple spatiotemporal feature vectors included in the fused spatiotemporal feature pyramid, and then obtains the process category included in the experimental video data of the variable gravity particle experiment, as well as the start time and end time of each process category. The method provided by the present invention can quickly and accurately determine the multiple stage processes included in the spatial variable gravity particle material experiment based on the experimental video data, thereby improving data processing efficiency. On the other hand, the method provided by the present invention introduces RGB differential stream, and can construct RGB differential stream to replace optical stream according to the characteristic that the background change of experimental video data is not large, so as to better perceive the temporal changes of experimental process categories; and, the method provided by the present invention constructs the spatiotemporal feature pyramid of RGB stream and RGB differential stream based on the temporal change information of RGB differential stream, which can effectively strengthen the learning on RGB stream, so as to improve the perception ability of experimental process on RGB stream; the spatiotemporal feature pyramid provided by the present invention includes feature vectors of multiple time scales, and uses learnable weighted fusion to construct a fused spatiotemporal feature pyramid, so as to capture richer spatiotemporal semantic information, which can effectively improve the accuracy of detection and meet the usage needs of technicians in different usage scenarios.

[0027] In some embodiments, a variable gravity particle experimental process detection method based on dual-flow differential provided by an embodiment of the present invention may be executed by a variable gravity particle experimental process detection system 100 based on dual-flow differential (hereinafter referred to as detection system 100 ).

[0028] For example, see Figure 1 , Figure 1A process schematic diagram of a variable gravity particle experiment process detection method based on dual stream difference provided by an embodiment of the present invention, first, obtain the experimental video data of the variable gravity particle experiment, then sample the experimental video data to obtain n video samples including RGB streams of T frame images, then perform differential processing on the RGB stream of each video sample to obtain the RGB differential stream of each video sample. For the i-th video sample, obtain the feature vectors of the RGB stream and the RGB differential stream through the feature extraction network according to the RGB stream and the RGB differential stream of the i-th video sample, then construct the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB differential stream according to the feature vectors of the RGB stream and the RGB differential stream based on the temporal change information carried by the RGB differential stream, each spatiotemporal feature pyramid includes multiple feature vectors with different time scales, then perform weighted fusion on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB differential stream to obtain the fused spatiotemporal feature pyramid of the i-th video sample, and finally determine the process detection result through the process detection unit according to the fused spatiotemporal feature pyramid of each video sample.

[0029] As an example, the detection system 100 can be any electronic device 200 with data processing capabilities, such as a general-purpose computer, a personal computer, a laptop computer, a switch or a tablet computer, etc. The specific implementation method of the detection system 100 is not limited here.

[0030] Figure 2 The hardware structure diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device 200 includes a processor 210, a memory 220 and a communication interface 230.

[0031] The processor 210 may include one or more processing cores. The processor 210 uses various interfaces and lines to connect various parts in the electronic device 200, and executes various functions and processes data of the electronic device 200 by running or executing instructions, programs, code sets or instruction sets stored in the memory 220, and calling data stored in the memory 220. Optionally, the processor 210 can be implemented in at least one hardware form of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA).

[0032] The memory 220 may include a random access memory (RAl) or a read-only memory (ROL). Optionally, the memory 220 includes a non-transitory computer-readable storage medium (non-transitory colputer-readable storage lediul). The memory 220 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 220 may include a program storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a data acquisition function, a feature extraction function, and a process detection function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.

[0033] The communication interface 230 is used to communicate with other devices, equipment or communication networks, such as data storage devices, image processing equipment or Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0034] In physical implementation, the above-mentioned components (such as processor 210, memory 220 and communication interface 230) can be components in the same device (such as a laptop computer). Alternatively, at least two of the components can be set in the same device, that is, as different components in a device, such as a deployment method similar to devices or components in a distributed system.

[0035] It is to be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 200. In other embodiments of the present invention, the electronic device 200 may include more or fewer components than those illustrated, or combine certain components, or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0036] A variable gravity particle experimental process detection method based on dual flow difference provided by an embodiment of the present invention is described below in conjunction with the accompanying drawings.

[0037] Figure 3 A flow chart of a variable gravity particle experimental process detection method based on dual flow difference provided by an embodiment of the present invention. Optionally, the method can be Figure 1 The detection system 100 shown, that is, Figure 2 The electronic device 200 shown is executed. The method may include the following steps: S1. Obtain experimental video data of variable gravity particle experiment.

[0038] Specifically, the experimental video data includes multiple consecutive frames of images.

[0039] S2. Based on a preset step size, a sliding window sampling process is performed on the experimental video data at a preset sampling frequency to obtain multiple video samples, each of which includes an RGB stream and an RGB difference stream of T frame images.

[0040] The RGB differential stream is obtained by performing differential processing on the RGB stream, and T is a positive integer greater than 1.

[0041] Specifically, in the video understanding task, RGB stream and RGB difference (RGBDifference) stream are two forms of input data, which capture the static appearance information and dynamic motion information in the video respectively. The RGB stream is a sequence of raw RGB pixel values ​​of multiple consecutive frames. Each frame of the image is a three-dimensional tensor (height × width × number of channels), and the number of channels is usually 3 (R, G, B). Therefore, the RGB stream of each video sample is a four-dimensional tensor (time sequence length T × height × width × number of channels), and the time sequence length is the number of image frames included in the video sample. The RGB difference stream is obtained by calculating the difference between every two adjacent frames (the next frame minus the previous frame), which represents the motion information of pixels between different frames. The RGB difference stream of each video sample is a four-dimensional tensor (time sequence length T-1 × height × width × number of channels).

[0042] In one example, T is 256.

[0043] It can also be understood that the method provided by the present invention first downsamples the continuous multiple frame images included in the experimental video data at a preset sampling frequency, and then performs sliding window sampling processing on the multiple frame images that have completed the downsampling to obtain window segments of equal length, each window segment includes T frame images, and then the RGB stream of the T frame image is differentially processed to obtain an RGB differential stream. In this way, we get a video sample, and each video sample is the RGB stream and RGB differential stream corresponding to the T frame image.

[0044] In one possible implementation, see Figure 4 The above S2 specifically includes the following steps: S21, down-sampling the experimental video data at a preset sampling frequency f, and performing sliding window sampling processing based on a preset step size s, to obtain a plurality of samples including T frame images.

[0045] Wherein, s is less than T, and s and f are positive integers.

[0046] Exemplarily, the preset step size is 25 and the sampling frequency is 3 frames / second.

[0047] S22. Determine the RGB stream of each sample including the T frame image.

[0048] RGB Stream The data structure is: ; Among them, H is the resolution height of the image, W is the resolution width of the image, and 3 is the number of channels of the image.

[0049] S23, performing differential processing on the subsequent frame and the previous frame image of each two adjacent frames of the RGB stream of each sample including the T frame image, obtaining the RGB differential stream of each sample including the T frame image, so as to obtain multiple video samples.

[0050] RGB differential flow The data structure is: .

[0051] It can also be understood that after acquiring the experimental video data, the method provided by the present invention first performs a sliding window sampling operation to obtain multiple samples, each sample includes an RGB stream of a T-frame image, and then performs differential processing on the RGB stream of each sample to obtain an RGB differential stream corresponding to each RGB stream. In this way, multiple video samples are obtained, each video sample includes an RGB stream and an RGB differential stream of a T-frame image.

[0052] The method provided by the embodiment of the present invention performs sliding window sampling processing on the experimental video data of the variable gravity particle experiment. When the data volume of the experimental video data of the variable gravity particle experiment is large, the redundancy of the RGB stream can be reduced, thereby significantly reducing the data volume, reducing the demand for computing resources and storage space, and improving processing efficiency.

[0053] S3. For any video sample, perform feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L with different semantic levels corresponding to the RGB stream. 0 and the eigenvector L 1 ; and the feature vectors P with different semantic levels corresponding to the RGB difference stream 0 and the eigenvector P 1 ; In a possible implementation, the above S3 specifically includes the following steps: For any video sample, the RGB stream is input into the dilated 3D convolutional network and the feature vector A is output through the first node. 0 , output feature vector A through the second node 1 ; For the feature vector A 0 and the eigenvector A 1 Perform non-local operations and 3D convolution operations respectively to obtain the feature vector L 0 and the eigenvector L 1; For any video sample, the RGB difference stream is input into the dilated 3D convolutional network and the feature vector B is output through the first node 0 , output feature vector B through the second node 1 ; For the feature vector B 0 and the eigenvector B 1 Perform non-local operations and 3D convolution operations respectively to obtain the feature vector P 0 and the eigenvector P 1 ; Eigenvector L 0 The data structure is: L 0 ; Eigenvector L 1 The data structure is: L 1 ; Eigenvector P 0 The data structure is: P 0 ; Eigenvector P 1 The data structure is: P 1 ; Where C is the number of channels.

[0054] In one example, the feature vector A 0 The data structure is: A 0 ; Eigenvector A 1 The data structure is: A 1 ; Eigenvector B 0 The data structure is: B 0 Eigenvector B 1 The data structure is: B 1 ; In one example, the number of channels C is 512 and T is 256.

[0055] For details, see Figure 5 , Figure 5This is a schematic diagram of a process for determining a feature vector according to an embodiment of the present invention. The first node is the C4 (Mixed_4f) node of the RGB branch of the dilated 3D convolutional network, and the second node is the C5 (Mixed_5c) node. The RGB stream is input into the dilated 3D convolutional (I3D) network, and the feature vector A is output through the C4 node. 0 , output feature vector A through C5 node 1 , and then the feature vector A is transformed through the Non-Local module and the 3D convolution module 0 Perform non-local operations and convolution operations to obtain the feature vector L 0 , through the Non-Local module and the 3D convolution module, the feature vector A 1 Perform non-local operations and convolution operations to obtain the feature vector L 1 On the other hand, the RGB difference stream is input into the dilated 3D convolutional network, and the feature vector B is output through the C4 node. 0 , output feature vector B through C5 node 1 , and then the feature vector B is transformed through the Non-Local module and the 3D convolution module 0 Perform non-local operations and convolution operations to obtain the feature vector P 0 , through the Non-Local module and the 3D convolution module, the feature vector B 1 Perform non-local operations and convolution operations to obtain the feature vector P 1 .

[0056] It should be understood that the Non-Local module is a neural network module used to capture long-distance dependencies. The Non-Local module is widely used in tasks such as video understanding, image segmentation, and target detection. It can effectively model the global relationship between pixels (or feature points) rather than being limited to local neighborhoods.

[0057] It should be noted that the above method for determining the feature vector is only an exemplary description. The method provided in the embodiment of the present invention can also use other feature extraction networks to extract features from the RGB stream and the RGB difference stream, and the embodiment of the present invention does not impose any special restrictions on this.

[0058] S4. According to the feature vector L 0 , eigenvector L 1 , eigenvector P 0 and the eigenvector P 1 Construct spatiotemporal feature pyramids of RGB stream and RGB difference stream.

[0059] Specifically, the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales. That is, the spatiotemporal feature pyramid of the RGB stream includes n spatiotemporal feature vectors with different time scales. , the spatiotemporal feature pyramid of the RGB difference stream includes n spatiotemporal feature vectors with different time scales , i=0, 1, 2, ..., n.

[0060] In one possible implementation, the spatiotemporal feature pyramid of the RGB stream includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; ; ; L 0 , L 1 ; When i is greater than 0, ; The spatiotemporal feature pyramid of the RGB difference stream includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; ; P 0 , P 1 ; When i is greater than 0, ; in, It is the time series change information; is the weighted feature vector of the RGB stream; is the activation function, is a 1D convolution operation with a convolution kernel of size 1, is a linear rectification function; are the learnable weight parameters of convolution kernels of different sizes, 1D convolution operation with kernel size k, k=1, 3, 5 and 7; is a 1D convolution operation with a convolution kernel of size 3; LN refers to processing through the LayerNorm layer; ,and =1,2,...,n; Spatiotemporal feature vector and the spatiotemporal eigenvector The time scale is: ; ; .

[0061] The following is an example of n spatiotemporal feature vectors with different time scales provided by the embodiment of the present invention. Explain how to determine the.

[0062] For example, see Figure 6 , the feature vector L corresponding to the RGB stream 0 (equivalent to ) and the eigenvector L 1 (equivalent to ), the i-th eigenvector Through a change information guidance module shared with the RGB differential stream , and obtain the spatiotemporal features enhanced by temporal changes (It can also be understood as the i-th layer of the spatiotemporal feature pyramid), and then undergoes a 1D convolution with a convolution kernel size of 3 ( Except for this), we can get the spatiotemporal features of the next layer. Repeat this transformation to obtain the RGB spatiotemporal feature pyramid after the RGB differential stream is enhanced. ,Right now: ; 1; in , It is The change information shared by the layer RGB spatiotemporal features and the RGB differential spatiotemporal features guides the module.

[0063] Similarly, the spatiotemporal feature pyramid of the RGB difference stream can be constructed .

[0064] When constructing the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream, each layer of original features and Need to go through the change information guidance module first , and obtain the spatiotemporal features after temporal variation enhancement and .

[0065] See also Figure 7 , Figure 7 A schematic diagram of the working process of a change information guidance module provided by an embodiment of the present invention. In the CIGM module, a shared structure similar to the Inception network is designed for the RGB stream and the RGB difference stream, and convolution kernels of different sizes with learnable weights are used to alleviate the difference in time scale between long and short processes, namely: In the CIGM module, a shared Inception network-like structure is designed for the RGB stream and the RGB differential stream, using learnable weighted convolution kernels of different sizes to alleviate the difference in time scale between long and short processes, namely: ; in, It refers to the weighted features after multiple convolution kernels. Refers to different convolution kernel sizes, refers to the LayerNorm layer, and is the learnable weight parameter corresponding to the convolution, initialized to 1. In addition, when constructing the CIGM module, if the convolution kernel size Greater than the current input feature , then this convolutional layer is not constructed.

[0066] The core of the CIGM module is to use only the RGB differential flow to calculate the timing change information , in order to utilize the more prominent temporal change information in the RGB differential stream and enhance the learning of the RGB stream in the corresponding time dimension: Among them, the timing change information First, the local changes are highlighted by the maximum pooling, and then the 1D convolution with a convolution kernel size of 1 and the Sigmoid function are used to calculate, and then the spatiotemporal features enhanced by CIGM are obtained. and .

[0067] S5, performing weighted fusion on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each video sample; The fusion spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; in, and is a learnable weight parameter.

[0068] For details, see Figure 8 , the spatiotemporal feature pyramid of the RGB stream corresponding to each video sample Spatiotemporal feature pyramid of RGB difference stream Perform weighted fusion to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each video sample Among them, the fusion of spatiotemporal feature pyramid It includes multiple spatiotemporal feature vectors with different time scales. In this example, the spatiotemporal feature pyramid is fused. Includes 5 spatiotemporal feature vectors with different time scales.

[0069] S6. Determine the process detection result of the experimental video data through the trained process detection unit according to the fused spatiotemporal feature pyramid corresponding to each video sample.

[0070] The process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.

[0071] Specifically, the trained process detection unit determines the process detection result of each video sample according to the fused spatiotemporal feature pyramid corresponding to each video sample, and then summarizes and analyzes the process detection results of all video samples to obtain the process detection results of the experimental video data.

[0072] Optionally, the process detection unit includes a coarse detection subunit and a fine detection subunit: the coarse detection subunit is used to determine multiple proposals based on the fused spatiotemporal feature pyramid corresponding to each video sample, each proposal including left and right boundary distances and a process category; the fine detection subunit is used to determine the boundary features of each proposal based on a saliency refinement algorithm, optimize the boundary position of each proposal based on the boundary features of each proposal, and obtain the process detection results included in each video sample.

[0073] From the above S1-S6, it can be seen that the method provided by the present invention determines multiple video samples according to the experimental video data of the variable gravity particle experiment by means of sliding window sampling, and then constructs a spatiotemporal feature pyramid according to the RGB stream and the RGB difference stream of each video sample, and then constructs a spatiotemporal feature pyramid according to the RGB stream and the RGB difference stream for weighted fusion to obtain a fused spatiotemporal feature pyramid, and finally determines the process category of each video sample according to the multiple spatiotemporal feature vectors included in the fused spatiotemporal feature pyramid, and then obtains the process category included in the experimental video data of the variable gravity particle experiment, as well as the start time and end time of each process category. The method provided by the present invention can quickly and accurately determine the multiple stage processes included in the spatial variable gravity particle material experiment according to the experimental video data, thereby improving data processing efficiency. On the other hand, the method provided by the present invention introduces RGB differential stream, and can construct RGB differential stream to replace optical stream according to the characteristic that the background change of experimental video data is not large, so as to better perceive the temporal changes of experimental process categories; and, the method provided by the present invention constructs the spatiotemporal feature pyramid of RGB stream and RGB differential stream based on the temporal change information of RGB differential stream, which can effectively strengthen the learning on RGB stream, so as to improve the perception ability of experimental process on RGB stream; the spatiotemporal feature pyramid provided by the present invention includes feature vectors of multiple time scales, and uses learnable weighted fusion to construct a fused spatiotemporal feature pyramid, so as to capture richer spatiotemporal semantic information, which can effectively improve the accuracy of detection and meet the usage needs of technicians in different usage scenarios.

[0074] In some embodiments, see Fig. 9 Before the above S6, the method provided by the embodiment of the present invention further includes: S81. Obtain a training sample set of a variable gravity particle experiment, where the training sample set includes multiple training samples, and each training sample includes a video sample and a process category.

[0075] S82. Construct a target loss function.

[0076] Among them, the target loss function for: ; in, The classification loss function of the coarse detection subunit, Regression loss function for coarse detection subunits; The classification loss function of the fine detection subunit, Regression loss function for fine-grained detection of subunits.

[0077] Specifically, the classification loss uses the Focal Loss function as the training target; the regression loss uses the T DIoU loss in the fast and slow two-stream network. The regression branch of the fine detection subunit uses the Smooth L1 Loss function to predict the boundary offset. The localization quality of the fine detection subunit uses the binary cross entropy loss (BCE Loss) function as the training target.

[0078] S83. Based on the target loss function, iteratively train the process detection unit to obtain a trained process detection unit.

[0079] The method provided by the embodiment of the present invention can effectively solve the problem of imbalance between training samples of different process categories and predicted positive and negative samples by constructing a target loss function, and effectively improve the accuracy and robustness of the process detection unit.

[0080] The beneficial effects of a variable gravity particle experimental process detection method based on dual-flow difference provided by an embodiment of the present invention are exemplarily described below with reference to an example.

[0081] Exemplarily, the method provided by the embodiment of the present invention verifies the beneficial effects of the method provided by the embodiment of the present invention through a verification data set, which is composed of experimental videos of the variable gravity experimental cabinet granular material warehouse A collected by a panoramic camera located in the space laboratory, including 207 video clips, each of which captures rich experimental process details at a resolution of 1920×1080 and a frame rate of 25 frames per second, and the average video length reaches 9.4 minutes. In the 207 video clips, a total of 9 different types of process categories are marked, totaling 939 process instances, and the average duration of each instance is approximately 29 seconds. In addition, there is no temporal overlap between all instances.

[0082] See also Fig.10 , Fig.10 A schematic diagram of the process categories and durations of a spatial variable gravity granular material experiment for a validation data set provided by an embodiment of the present invention. The validation data set is mostly long processes, and the average duration of some processes is significantly longer. For example, the average duration of the baffle vibration process can reach 50 seconds. At the same time, some process durations also vary significantly. Compared with the baffle vibration, the average duration of the left movement of the baffle before vibration is only 1.5 seconds.

[0083] In this example, the method provided by the embodiment of the present invention is verified based on the evaluation indicators Mean Average Precision (mAP) and Average Precision (AP), and the mAP value with a step size of 0.1 on tIoU=[0.3,0.7] is compared, where tIoU refers to the temporal intersection-over-union ratio.

[0084] Average precision is an evaluation indicator for a single process category. It is used to measure the average precision at different recall levels, while mAP is the average of APs for all process categories. Specifically, when calculating AP, the prediction results are sorted according to their confidence, and then the precision at different recall thresholds is calculated. Finally, the average precision is obtained by integration or interpolation.

[0085] AP is usually calculated using the 11-point interpolation method, that is, the maximum precision values ​​at the 11 points with recall rates of 0, 0.1, 0.2, ..., 1 are averaged, and mAP is the average of the AP of all process categories, that is: Where C is the number of process categories.

[0086] tIoU is used to measure the overlap between the predicted experimental process fragment and the actual experimental process fragment. It is the ratio of the intersection area of ​​the predicted experimental process fragment and the actual experimental process fragment at time t to the union area. The higher the tIoU value, the closer the predicted result is to the actual result. The calculation of tIoU is: ; When calculating mAP, different tIoU thresholds are usually set. The prediction is considered correct only when the tIoU between the predicted result and the true result is greater than or equal to the threshold.

[0087] Specifically, the embodiment of the present invention is compared with the TP2N* method and the baseline method AFSD in the related art. The TP2N* method is a method for process prediction through fast and slow dual-frequency data sampling and dual-path spatiotemporal pyramid network (TP2N) and other strategies, and the baseline method AFSD is a method that combines the I3D network with the anchor-free method. The results of the comparative experiment are shown in Table 1. TP2N* is the experimental result of the fast and slow dual-stream network under dual-frequency (13fps and 3fps) data sampling, and the rest are experimental results under a single sampling frequency (3fps). As can be seen from Table 1, the method proposed in the present invention achieves an mAP detection result of 89.28% under single-frequency (3fps) data sampling, and has achieved significant improvement under different tIoUs, which is about 7% higher than the baseline method AFSD and about 1.4% higher than the TP2N* method. It shows that the method proposed in the present invention can obtain richer spatiotemporal feature information by using the dual-stream differential fusion method, capture more significant temporal dynamic change information, thereby achieving better recognition and positioning of the experimental process, and proving the effectiveness of the technical solution of the present invention.

[0088] Table 1 Furthermore, in order to explore the impact of different data input modes on detection accuracy, a data input mode comparison experiment was conducted on the space variable gravity granular material experimental video dataset. The experimental results are shown in Table 2.

[0089] The method proposed in the present invention uses the data mode input of RGB stream and RGB difference stream fusion (RGB+RGB Difference), and the AP under different tIoU thresholds is significantly improved, and the detection accuracy mAP reaches 89.28%. Compared with the baseline method using only RGB stream, mAP is improved by about 7%, compared with the baseline method using only RGB difference stream (RGB Difference), mAP is improved by about 5.5%, and compared with the baseline method using RGB stream and optical flow fusion (RGB+Optical Flow), mAP is improved by about 5.84%, which shows that using the data mode input of RGB stream and RGB difference stream fusion (RGB+RGB Difference), using RGB difference stream (RGB Difference) to perceive the dynamic changes of timing, and guiding the RGB stream to focus on learning the experimental content of the change interval is more effective, which proves the effectiveness of the technical solution of the present invention.

[0090] Table 2 Furthermore, in order to verify the role of each unit module included in the detection system 100 in the technical solution of the present invention, a module ablation experiment was carried out on a spatial variable gravity granular material experimental video dataset.

[0091] Table 3 The results of the module ablation experiment are shown in Table 3, where the Baseline method refers to the AFSD method, Diff refers to RGBDifference, and the use of RGB and Diff at the same time represents the use of dual-stream fusion. It can be seen that by fusing the RGB stream with the RGB differential stream on the basis of the baseline method, a significant improvement can be achieved, which is about 7% higher than the baseline method using only the RGB stream, and about 5.5% higher than the baseline method using only the RGB differential stream, indicating that dual-stream differential fusion is helpful for process detection. On the basis of using the RGB stream and the RGB differential stream fusion, the change information guidance module (CIGM) designed in the technical solution of the present invention uses the significant temporal change information on the differential stream to strengthen the learning on the RGB stream, which is improved by about 2%, indicating the effectiveness of the CIGM module.

[0092] Furthermore, in order to explore the influence of the number of CIGM modules on the detection accuracy in the technical solution of the present invention, an ablation experiment of the number of CIGM modules was carried out on the validation data set. The experimental results are shown in Table 4, from which it can be seen that no matter how many CIGM modules there are, as long as the CIGM modules are used, the mAP of the detection result is higher than the result without using the CIGM modules (87.02%). Among them, when the number of CIGM modules is 3, the performance of the method provided by the present invention is the best, and its detection accuracy mAP reaches 89.28%. The number of CIGM modules is not the more the better. With the increase of the number of CIGM modules, the detection accuracy mAP of the model shows a trend of first rising and then falling as a whole, but it is higher than the result without using the CIGM modules, which illustrates the effectiveness of the CIGM modules proposed in the technical solution of the present invention.

[0093] Table 4 The above mainly introduces the scheme of the embodiment of the present invention from the perspective of the method. It is understandable that, in order to realize the above functions, the detection system 100 includes at least one of the hardware structure and software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiment of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present invention.

[0094] The embodiment of the present invention can divide the detection system 100 into functional units according to the above method example. For example, the detection system 100 can be divided into functional units corresponding to various functions, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0095] For example, Fig.11 The hardware structure diagram of a detection system provided by an embodiment of the present invention is shown. The detection system 100 includes: a data acquisition unit 110, which is used to acquire experimental video data of a variable gravity particle experiment; a sample sampling unit 120, which is used to perform sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain multiple video samples, each video sample includes an RGB stream and an RGB difference stream of T frame images, wherein the RGB difference stream is obtained by performing differential processing on the RGB stream, and T is a positive integer greater than 1; a feature extraction unit 130, which is used to perform feature extraction operations on the RGB stream and the RGB difference stream for any video sample, and obtain a feature vector L with different semantic levels corresponding to the RGB stream. 0 and the eigenvector L 1 ; and the feature vectors P with different semantic levels corresponding to the RGB difference stream 0 and the eigenvector P 1 ; Feature construction unit 140, for 0 , eigenvector L 1 , eigenvector P 0 and the eigenvector P 1 A spatiotemporal feature pyramid of an RGB stream and an RGB differential stream is constructed, wherein the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales. A feature fusion unit 150 is used to perform weighted fusion on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB differential stream corresponding to each video sample, so as to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each video sample. A detection unit 160 is used to determine a process detection result of the experimental video data according to the process detection unit trained based on the fused spatiotemporal feature pyramid corresponding to each video sample, wherein the process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.

[0096] It should be understood that the specific description of the above optional methods can refer to the above method embodiments, which will not be repeated here. In addition, the explanation of any detection system 100 provided above and the description of the beneficial effects can refer to the above corresponding method embodiments, which will not be repeated here.

[0097] The embodiment of the present invention further provides a computer-readable storage medium, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the methods of the above embodiments. For the explanation of the relevant contents and the description of the beneficial effects in any of the above-mentioned computer-readable storage media, reference can be made to the above-mentioned corresponding embodiments, which will not be repeated here.

[0098] The embodiment of the present invention further provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the above detection system 100. Optionally, the functions supported by the chip can be referred to above and will not be described in detail here.

[0099] Those skilled in the art will appreciate that all or part of the steps of the above embodiments can be implemented by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, a specific circuit structure (application specific integrated circuit, ASIC), a microprocessor (digital signal processor, DSP), a field programmable gate array (field prograllable gatearray, FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof.

[0100] The embodiment of the present invention also provides a computer program product including instructions, when the instructions are executed on a computer, the computer executes any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD), etc.

[0101] It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiments of the present invention, such as but not limited to the above-mentioned memories, computer-readable storage media, and communication chips, etc., are all non-transitory. Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein the communication medium includes any medium that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0102] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A variable gravity particle experimental process detection method based on dual flow difference, characterized in that: The method comprises: Acquire experimental video data of variable gravity particle experiments; Based on a preset step size, the experimental video data is subjected to sliding window sampling processing at a preset sampling frequency to obtain a plurality of video samples, each of which includes an RGB stream and an RGB difference stream of T frame images, wherein the RGB difference stream is obtained by performing a difference process on the RGB stream, and T is a positive integer greater than 1; For any of the video samples, feature extraction operations are performed on the RGB stream and the RGB difference stream to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; Constructing a spatiotemporal feature pyramid of the RGB stream and the RGB difference stream according to the feature vector L0, the feature vector L1, the feature vector P0, and the feature vector P1, wherein the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales; Performing weighted fusion on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream corresponding to each of the video samples to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each of the video samples; A process detection result of the experimental video data is determined by a trained process detection unit according to the fused spatiotemporal feature pyramid corresponding to each of the video samples. The process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.

2. The method according to claim 1, characterized in that Based on the preset step length, the experimental video data is subjected to sliding window sampling processing at a preset sampling frequency to obtain a plurality of video samples, including: Downsampling the experimental video data at a preset sampling frequency f, and performing sliding window sampling processing based on a preset step size s, to obtain a plurality of samples including T frame images, wherein s is less than T, and s and f are positive integers; Determine the RGB stream of each sample including the T frame image; Performing differential processing on the next frame and the previous frame of each two adjacent frames of the RGB stream of each sample including the T frame image to obtain the RGB differential stream of each sample including the T frame image, so as to obtain the multiple video samples; RGB Stream The data structure is: ; RGB differential flow The data structure is: ; Among them, H is the resolution height of the image, W is the resolution width of the image, and 3 is the number of channels of the image.

3. The method according to claim 2, characterized in that For any of the video samples, a feature extraction operation is performed on the RGB stream and the RGB difference stream to obtain a feature vector L0 and a feature vector L1 with different semantic levels corresponding to the RGB stream; and a feature vector P0 and a feature vector P1 with different semantic levels corresponding to the RGB difference stream, including: For any of the video samples, the RGB stream is input into the dilated 3D convolutional network to output a feature vector A0 through the first node and a feature vector A1 through the second node; Perform non-local operation and 3D convolution operation on feature vector A0 and feature vector A1 respectively to obtain feature vector L0 and feature vector L1; For any of the video samples, the RGB difference stream is input into the dilated 3D convolutional network to output a feature vector B0 through the first node and a feature vector B1 through the second node; Perform non-local operation and 3D convolution operation on eigenvector B0 and eigenvector B1 respectively to obtain eigenvector P0 and eigenvector P1; The data structure of the feature vector L0 is: L0 ; The data structure of the feature vector L1 is: L1 ; The data structure of the feature vector P0 is: P0 ; The data structure of the feature vector P1 is: P1 ; Where C is the number of channels.

4. The method according to claim 3, characterized in that The spatiotemporal feature pyramid of the RGB stream includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; ; ; L0, L1; When i is greater than 0, ; The spatiotemporal feature pyramid of the RGB difference stream includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; ; P0, P1; When i is greater than 0, ; in, It is the time series change information; is the weighted feature vector of the RGB stream; is the activation function, is a 1D convolution operation with a convolution kernel of size 1, is a linear rectification function; are the learnable weight parameters of convolution kernels of different sizes, 1D convolution operation with kernel size k, k=1, 3, 5 and 7; is a 1D convolution operation with a convolution kernel of size 3; LN refers to processing through the LayerNorm layer; ,and =1,2,...,n; Spatiotemporal feature vector and the spatiotemporal eigenvector The time scale is: ; ; 。 5. The method according to claim 4, characterized in that The fused spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales The formula for determining is: ; in, and is a learnable weight parameter.

6. The method according to claim 5, characterized in that The process detection unit includes a coarse detection subunit and a fine detection subunit: The rough detection subunit is used to determine a plurality of proposals according to the fused spatiotemporal feature pyramid corresponding to each of the video samples, each proposal including left and right boundary distances and a process category; The fine detection subunit is used to determine the boundary features of each proposal based on the saliency refinement algorithm, optimize the boundary position of each proposal based on the boundary features of each proposal, and obtain the process detection result included in each of the video samples.

7. The method according to claim 6, characterized in that Before determining the process detection result of the experimental video data by the process detection unit completed through training according to the fused spatiotemporal feature pyramid corresponding to each of the video samples, the method further includes: Acquire a training sample set of a variable gravity particle experiment, wherein the training sample set includes a plurality of training samples, and each training sample includes a video sample and a process category; Construct the target loss function; Based on the target loss function, the process detection unit is iteratively trained to obtain a trained process detection unit.

8. The method according to claim 7, characterized in that The objective loss function for: ; in, The classification loss function of the coarse detection subunit, Regression loss function for coarse detection subunits; The classification loss function of the fine detection subunit, Regression loss function for fine-grained detection of subunits.

9. A variable gravity particle experimental process detection system based on dual flow difference, characterized in that: The system comprises: A data acquisition unit, used to acquire experimental video data of variable gravity particle experiments; A sample sampling unit, configured to perform sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples, each of which includes an RGB stream and an RGB difference stream of T frame images, wherein the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1; A feature extraction unit, for performing feature extraction operations on the RGB stream and the RGB difference stream for any of the video samples, to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; A feature construction unit, used to construct a spatiotemporal feature pyramid of the RGB stream and the RGB difference stream according to the feature vector L0, the feature vector L1, the feature vector P0 and the feature vector P1, wherein the spatiotemporal feature pyramid includes n spatiotemporal feature vectors with different time scales; A feature fusion unit, used to perform weighted fusion on the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream corresponding to each of the video samples, to obtain a fused spatiotemporal feature pyramid of n spatiotemporal feature vectors with different time scales corresponding to each of the video samples; A detection unit is used to determine the process detection result of the experimental video data according to the process detection unit trained according to the fused spatiotemporal feature pyramid corresponding to each of the video samples, and the process detection result of the experimental video data includes at least one process category, and the start time and end time corresponding to each process category.

10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the variable gravity particle experimental process detection method based on dual-flow differential as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Video recognition method based on space-time pyramid network

    CN107909041A

  • Fine-grained video action recognition method based on hierarchical structure

    CN113139467A

  • Video saliency detection method based on space-time double-flow pyramid network architecture

    CN114882405A

  • Lightweight violent behavior identification method in monitoring scene

    CN115690907A

  • Centrifugal pump rotor fault diagnosis method employing cwgan-GP and two-stream CNN models

    WO2025015797A1