A Detection Method for the Process of Variable Gravity Particle Experiments Based on Dual-Flow Difference
The spatiotemporal feature pyramid is constructed through the dual-stream differential method, which solves the problem of inefficiency of multi-stage process identification in variable gravity particle experimental video data, and realizes fast and accurate process detection and data processing.
Patent Information
- Application Number
- CN202510444636.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The prior art is difficult to quickly and accurately identify multiple stage processes from variable gravity particle experimental video data, resulting in inefficient data processing.
Using a dual-stream differential method, the RGB stream and RGB differential stream are obtained through sliding window sampling, a spatiotemporal feature pyramid is constructed and weighted fusion is performed. Combined with the process detection unit completed by the training, the process detection results of the experimental video data are determined.
The multiple stages of the experiment of variable gravity particles are realized quickly and accurately, which improves data processing efficiency, and enhances the perception of timing changes of the experimental process through RGB differential flow, improving the accuracy of detection.
Smart Images

Figure CN119963932B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video data processing, and in particular to a method for detecting the process of variable gravity particle experiments based on dual-stream difference. Background Art
[0002] The variable gravity particle experiment, namely the space variable gravity particle material experiment, is a scientific experiment for studying the behavior of particle materials in a microgravity or variable gravity environment. Particle materials exhibit specific flow, stacking, compression and other characteristics under the earth's gravity environment, but their behavior may change significantly in a microgravity or different gravity environments. Since the experiment lasts for a long time and includes multiple different stage processes, it is necessary to detect and identify each stage process included in the experiment according to the experimental video data, so that technicians can quickly and accurately identify specific stage processes from the experimental video, thereby providing strong data support for scientific researchers, assisting them in subsequent data analysis work, and promoting the scientific research process.
[0003] Therefore, there is an urgent need for a method for detecting the process of variable gravity particle experiments based on dual-stream difference, which can quickly and accurately determine multiple stage processes included in the space variable gravity particle material experiment according to the experimental video data, and improve the data processing efficiency. Summary of the Invention
[0004] An embodiment of the present invention provides a method for detecting the process of variable gravity particle experiments based on dual-stream difference, which can quickly and accurately determine multiple stage processes included in the space variable gravity particle material experiment according to the experimental video data, and improve the data processing efficiency.
[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:
[0006] In a first aspect, a method for detecting the process of variable gravity particle experiments based on dual-stream difference is provided. The method includes: obtaining experimental video data of variable gravity particle experiments; performing sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples, each video sample including an RGB stream and an RGB difference stream of T frames of images, where the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1; for any video sample, performing feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; constructing a spatio-temporal feature pyramid for the RGB stream and the RGB difference stream according to the feature vectors L0, L1, P0, and P1, the spatio-temporal feature pyramid including n spatio-temporal feature vectors with different time scales; performing weighted fusion on the spatio-temporal feature pyramid of the RGB stream and the spatio-temporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample; determining the process detection result of the experimental video data according to the fused spatio-temporal feature pyramid corresponding to each video sample through a trained process detection unit, the process detection result of the experimental video data including at least one process category, and the start time and end time corresponding to each process category.
[0007] In a possible implementation manner of the first aspect, performing sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples includes: downsampling the experimental video data at a preset sampling frequency f, and performing sliding window sampling processing based on a preset step size s to obtain a plurality of samples including T frames of images, where s is less than T, and s and f are positive integers; determining the RGB stream of each sample including T frames of images; performing difference processing on the subsequent frame and the previous frame image of every two adjacent frames of the RGB stream of each sample including T frames of images to obtain the RGB difference stream of each sample including T frames of images, so as to obtain a plurality of video samples;
[0008] The data structure of the RGB stream is: ;
[0009] ;
[0010] The data structure of the RGB difference stream is: ;
[0011] ;
[0012] where H is the resolution height of the image, W is the resolution width of the image, and 3 is the number of channels of the image.
[0013] In a possible implementation of the first aspect, for any video sample, perform feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream, including: for any video sample, input the RGB stream into the dilated 3D convolutional network to output the feature vector A0 through the first node and output the feature vector A1 through the second node; perform non-local operations and 3D convolutional operations on the feature vectors A0 and A1 respectively to obtain the feature vectors L0 and L1; for any video sample, input the RGB difference stream into the dilated 3D convolutional network to output the feature vector B0 through the first node and output the feature vector B1 through the second node; perform non-local operations and 3D convolutional operations on the feature vectors B0 and B1 respectively to obtain the feature vectors P0 and P1;
[0014] The data structure of the feature vector L0 is:
[0015] L0 ;
[0016] The data structure of the feature vector L1 is:
[0017] L1 ;
[0018] The data structure of the feature vector P0 is:
[0019] P0 ;
[0020] The data structure of the feature vector P1 is:
[0021] P1 ;
[0022] where C is the number of channels.
[0023] In a possible implementation of the first aspect, the determination formula for the n spatio-temporal feature vectors with different time scales included in the spatio-temporal feature pyramid of the RGB stream is:
[0024] ;
[0025] ;
[0026] ;
[0027] L0, L1;
[0028] When i > 0, ;
[0029] The spatio-temporal feature pyramid of the RGB difference flow includes n spatio-temporal feature vectors with different time scales The determination formula is:
[0030] ;
[0031] ;
[0032] P0, P1;
[0033] When i > 0, ;
[0034] Wherein, is the temporal variation information; is the weighted feature vector of the RGB flow; is the activation function, is the 1D convolution operation of a convolution kernel of size 1, is the rectified linear unit function; is the learnable weight parameter of convolution kernels of different sizes, is the 1D convolution operation of a convolution kernel of size k, k = 1, 3, 5, and 7; is the 1D convolution operation of a convolution kernel of size 3; LN means processing through the LayerNorm layer; , and = 1, 2,..., n;
[0035] Spatio-temporal feature vector and spatio-temporal feature vector The time scales are:
[0036] ;
[0037] ;
[0038] .
[0039] In a possible implementation manner of the first aspect, the determination formula for fusing the n spatio-temporal feature vectors with different time scales included in the spatio-temporal feature pyramid is:
[0040] ;
[0041] Wherein, and are learnable weight parameters.
[0042] In a possible implementation of the first aspect, the process detection unit includes a rough detection subunit and a fine detection subunit: the rough detection subunit is configured to determine a plurality of proposals according to the fused spatio-temporal feature pyramid corresponding to each video sample, each proposal including the left and right boundary distances and the process category; the fine detection subunit is configured to determine the boundary features of each proposal based on the saliency refinement algorithm, and optimize the boundary positions of each proposal based on the boundary features of each proposal to obtain the process detection result included in each video sample.
[0043] In a possible implementation of the first aspect, before determining the process detection result of the experimental video data by the process detection unit completed through training according to the fused spatio-temporal feature pyramid corresponding to each video sample, the above method further includes: obtaining a training sample set of the variable gravity particle experiment, the training sample set including a plurality of training samples, each training sample including a video sample and a process category; constructing an objective loss function; and iteratively training the process detection unit based on the objective loss function to obtain the process detection unit completed through training.
[0044] In a possible implementation of the first aspect, the objective loss function is:
[0045] ;
[0046] where the classification loss function of the rough detection subunit, the regression loss function of the rough detection subunit; the classification loss function of the fine detection subunit, the regression loss function of the fine detection subunit.
[0047] The beneficial effects of the present invention are as follows: The method provided by the present invention determines multiple video samples by means of sliding window sampling according to the experimental video data of the variable gravity particle experiment, then constructs a spatio-temporal feature pyramid based on the RGB stream and the RGB difference stream of each video sample, and then performs weighted fusion on the spatio-temporal feature pyramids constructed based on the RGB stream and the RGB difference stream to obtain a fused spatio-temporal feature pyramid. Finally, the process category of each video sample is determined according to the multiple spatio-temporal feature vectors included in the fused spatio-temporal feature pyramid, and then the process category included in the experimental video data of the variable gravity particle experiment, as well as the start time and end time of each process category, are obtained. The method provided by the present invention can quickly and accurately determine the multiple stage processes included in the spatial variable gravity particle material experiment according to the experimental video data, improving the data processing efficiency. On the other hand, by introducing the RGB difference stream, the method provided by the present invention can construct the RGB difference stream to replace the optical flow according to the characteristic that the background of the experimental video data changes little, so as to better perceive the temporal variation of the experimental process category; moreover, the method provided by the present invention constructs the spatio-temporal feature pyramids of the RGB stream and the RGB difference stream based on the temporal variation information of the RGB difference stream, which can effectively strengthen the learning on the RGB stream, thereby improving the perception ability of the experimental process on the RGB stream; the spatio-temporal feature pyramid provided by the present invention includes feature vectors of multiple different time scales, and a learnable weighted fusion is used to construct the fused spatio-temporal feature pyramid, so as to capture richer spatio-temporal semantic information, which can effectively improve the detection accuracy and meet the usage requirements of technicians in different usage scenarios.
[0048] Second aspect, the present invention provides a variable gravity particle experiment process detection system based on dual-stream difference. The above system includes: a data acquisition unit for acquiring experimental video data of a variable gravity particle experiment; a sample sampling unit for performing sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step length to obtain a plurality of video samples, each video sample including an RGB stream and an RGB difference stream of T frame images, wherein the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1; a feature extraction unit for, for any video sample, performing feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream, and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; a feature construction unit for constructing a spatio-temporal feature pyramid of the RGB stream and the RGB difference stream according to the feature vectors L0, L1, P0, and P1, the spatio-temporal feature pyramid including n spatio-temporal feature vectors with different time scales; a feature fusion unit for performing weighted fusion on the spatio-temporal feature pyramid of the RGB stream and the spatio-temporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample; a detection unit for determining the process detection result of the experimental video data through a trained process detection unit according to the fused spatio-temporal feature pyramid corresponding to each video sample, the process detection result of the experimental video data including at least one process category, and the start time and end time corresponding to each process category.
[0049] Third aspect, there is provided an electronic device including a memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device is caused to execute the method in any implementation manner of the first aspect.
[0050] Fourth aspect, there is provided a computer-readable storage medium including computer instructions, and when the computer instructions are run on an electronic device, the electronic device is caused to execute the method in any implementation manner of the first aspect.
[0051] Fifth aspect, there is provided a computer program product, and when the computer program product is run on a computer, the computer is caused to execute the method in any implementation manner of the first aspect.
[0052] It can be understood that the beneficial effects that can be achieved by the system in the second aspect, the electronic device in the third aspect, the computer-readable storage medium in the fourth aspect, and the computer program product in the fifth aspect provided above can refer to the beneficial effects in the first aspect and any possible design manner thereof, which will not be elaborated here. Brief Description of the Drawings
[0053] Figure 1 It is a process schematic diagram of a method for detecting the process of a variable gravity particle experiment based on dual - stream difference provided by an embodiment of the present invention;
[0054] Figure 2 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention;
[0055] Figure 3 It is a flowchart of a method for detecting the process of a variable gravity particle experiment based on dual - stream difference provided by an embodiment of the present invention;
[0056] Figure 4 It is a flowchart of another method for detecting the process of a variable gravity particle experiment based on dual - stream difference provided by an embodiment of the present invention;
[0057] Figure 5 It is a process schematic diagram of generating a feature vector provided by an embodiment of the present invention;
[0058] Figure 6 It is a process schematic diagram of generating a spatio - temporal feature pyramid provided by an embodiment of the present invention;
[0059] Figure 7 It is a structural schematic diagram of a change information guiding module provided by an embodiment of the present invention;
[0060] Figure 8 It is a process schematic diagram of generating a fused spatio - temporal feature pyramid provided by an embodiment of the present invention;
[0061] Figure 9 It is a flowchart of yet another method for detecting the process of a variable gravity particle experiment based on dual - stream difference provided by an embodiment of the present invention;
[0062] Figure 10 It is a schematic diagram for verifying the process category and duration of a spatial variable gravity particle material experiment of a verification dataset provided by an embodiment of the present invention;
[0063] Figure 11 It is a structural schematic diagram of a detection system provided by an embodiment of the present invention. Detailed Description of the Embodiment
[0064] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention. Among them, in the description of the present invention, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B. The "or" in the present invention is merely a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. And, in the description of the present invention, unless otherwise specified, "a plurality" means two or more than two. "At least one (piece)" or similar expressions thereof refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces).
[0065] In addition, in order to facilitate a clear description of the technical solutions in the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily limit being different from each other.
[0066] At the same time, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being superior or more advantageous than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0067] The variable gravity particle experiment, namely the space variable gravity particle material experiment, is a scientific experiment to study the behavior of particle materials in a microgravity or variable gravity environment. Particle materials exhibit specific flow, packing, compression and other characteristics in the earth's gravity environment, but their behavior may change significantly in a microgravity or different gravity environment. Since the experiment takes a long time and includes multiple different stage processes, it is necessary to detect and identify each stage process included in the experiment according to the experimental video data, so that technicians can quickly and accurately identify specific stage processes from the experimental video, thereby providing strong data support for scientific researchers to assist them in subsequent data analysis work and promoting the scientific research process.
[0068] Therefore, there is an urgent need for a method for detecting the variable gravity particle experiment process based on double-flow difference, which can quickly and accurately determine the multiple stage processes included in the space variable gravity particle material experiment according to the experimental video data and improve the data processing efficiency.
[0069] In view of this, an embodiment of the present invention provides a method for detecting the process of a variable gravity particle experiment based on dual-stream difference. The above method includes: obtaining experimental video data of a variable gravity particle experiment; performing sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples, each video sample including the RGB stream and the RGB difference stream of T frame images, where the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1; for any video sample, performing feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; constructing a spatio-temporal feature pyramid of the RGB stream and the RGB difference stream according to the feature vectors L0, L1, P0, and P1, the spatio-temporal feature pyramid including n spatio-temporal feature vectors with different time scales; performing weighted fusion on the spatio-temporal feature pyramid of the RGB stream and the spatio-temporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample; determining the process detection result of the experimental video data according to the fused spatio-temporal feature pyramid corresponding to each video sample through a trained process detection unit, the process detection result of the experimental video data including at least one process category, and the start time and end time corresponding to each process category.
[0070] The method provided by the present invention determines multiple video samples by means of sliding window sampling according to the experimental video data of the variable gravity particle experiment, then constructs a spatio-temporal feature pyramid based on the RGB stream and the RGB difference stream of each video sample, then performs weighted fusion on the spatio-temporal feature pyramids constructed based on the RGB stream and the RGB difference stream to obtain a fused spatio-temporal feature pyramid, and finally determines the process category of each video sample according to the multiple spatio-temporal feature vectors included in the fused spatio-temporal feature pyramid, and then obtains the process category included in the experimental video data of the variable gravity particle experiment, as well as the start time and end time of each process category. The method provided by the present invention can quickly and accurately determine the multiple stage processes included in the spatial variable gravity particle material experiment according to the experimental video data, improving the data processing efficiency. On the other hand, by introducing the RGB difference stream, the method provided by the present invention can construct the RGB difference stream to replace the optical flow according to the characteristic that the background of the experimental video data changes little, so as to better perceive the temporal change of the experimental process category; moreover, the method provided by the present invention constructs the spatio-temporal feature pyramids of the RGB stream and the RGB difference stream based on the temporal change information of the RGB difference stream, which can effectively strengthen the learning on the RGB stream, thereby enhancing the perception ability of the experimental process on the RGB stream; the spatio-temporal feature pyramid provided by the present invention includes feature vectors of multiple different time scales, and a learnable weighted fusion is used to construct the fused spatio-temporal feature pyramid, so as to capture richer spatio-temporal semantic information, which can effectively improve the detection accuracy and meet the usage requirements of technicians in different usage scenarios.
[0071] In some embodiments, a method for detecting the process of a variable gravity particle experiment based on dual-stream difference provided by an embodiment of the present invention can be executed by a system 100 for detecting the process of a variable gravity particle experiment based on dual-stream difference (hereinafter referred to as the detection system 100).
[0072] Exemplarily, refer to Figure 1 , Figure 1It is a process schematic diagram of a method for detecting the process of variable gravity particle experiment based on double - stream difference provided by an embodiment of the present invention. First, obtain the experimental video data of the variable gravity particle experiment. Then, sample the experimental video data to obtain n video samples including RGB streams of T - frame images. Then, perform differential processing on the RGB streams of each video sample to obtain the RGB difference streams of each video sample. For the i - th video sample, based on the RGB stream and the RGB difference stream of the i - th video sample, obtain the feature vectors of the RGB stream and the RGB difference stream through a feature extraction network. Then, based on the temporal change information carried by the RGB difference stream, construct a spatio - temporal feature pyramid of the RGB stream and a spatio - temporal feature pyramid of the RGB difference stream according to the feature vectors of the RGB stream and the RGB difference stream. Each spatio - temporal feature pyramid includes multiple feature vectors with different time scales. Then, perform weighted fusion on the spatio - temporal feature pyramid of the RGB stream and the spatio - temporal feature pyramid of the RGB difference stream to obtain the fused spatio - temporal feature pyramid of the i - th video sample. Finally, determine the process detection result through a process detection unit according to the fused spatio - temporal feature pyramid of each video sample.
[0073] As an example, the detection system 100 can be any electronic device 200 with data - processing capabilities, such as a general - purpose computer, a personal computer, a laptop computer, a switch, or a tablet computer, etc. The specific implementation manner of the detection system 100 is not limited here.
[0074] Figure 2 It shows a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. The electronic device 200 includes a processor 210, a memory 220, and a communication interface 230.
[0075] The processor 210 may include one or more processing cores. The processor 210 connects various parts within the electronic device 200 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 220, and by calling data stored in the memory 220, the processor 210 executes various functions of the electronic device 200 and processes data. Optionally, the processor 210 may be implemented in at least one hardware form of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a field - programmable gate array (FPGA), or a programmable logic array (PLA).
[0076] The memory 220 may include a random access memory (RAM), or may include a read-only memory (ROM). Optionally, the memory 220 includes a non-transitory computer-readable storage medium. The memory 220 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 220 may include a storage program area. Among them, the storage program area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a data acquisition function, a feature extraction function, and a process detection function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.
[0077] The communication interface 230 is used to communicate with other devices, equipment, or communication networks, such as data storage devices, image processing devices, or Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0078] In terms of physical implementation, the above-mentioned various devices (such as the processor 210, the memory 220, and the communication interface 230) can be respectively devices in the same device (such as a laptop computer). Or, at least two of them can be arranged in the same device, that is, as different devices in a device, similar to the deployment method of devices or components in a distributed system.
[0079] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 200. In other embodiments of the present invention, the electronic device 200 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0080] The following will describe a method for detecting the process of a variable gravity particle experiment based on dual-stream difference provided by an embodiment of the present invention with reference to the accompanying drawings of the specification.
[0081] Figure 3 It is a flowchart of a method for detecting the process of a variable gravity particle experiment based on dual-stream difference provided by an embodiment of the present invention. Optionally, this method can be Figure 1 executed by the detection system 100 shown, that is, Figure 2 the electronic device 200 shown. This method may include the following steps:
[0082] S1. Obtain the experimental video data of the variable gravity particle experiment.
[0083] Specifically, the experimental video data includes multiple consecutive frames of images.
[0084] S2. Perform sliding window sampling on the experimental video data at a preset sampling frequency based on a preset step size to obtain multiple video samples. Each video sample includes the RGB stream and the RGB difference stream of T frames of images.
[0085] Among them, the RGB difference stream is obtained by performing difference processing on the RGB stream, and T is a positive integer greater than 1.
[0086] Specifically, in video understanding tasks, the RGB stream and the RGB difference (RGBDifference) stream are two forms of input data, which respectively capture the static appearance information and dynamic motion information in the video. The RGB stream is a sequence composed of the original RGB pixel values of multiple consecutive frames of images. Each frame of image is a three-dimensional tensor (height × width × number of channels), and the number of channels is usually 3 (R, G, B). Therefore, the RGB stream of each video sample is a four-dimensional tensor (temporal length T × height × width × number of channels), and the temporal length is the number of image frames included in the video sample. The RGB difference stream is obtained by calculating the difference between every two adjacent frames (the latter frame minus the former frame), representing the motion information of pixels between different frames. The RGB difference stream of each video sample is a four-dimensional tensor (temporal length T - 1 × height × width × number of channels).
[0087] In one example, T is 256.
[0088] It can also be understood that the method provided by the present invention first downsamples the multiple consecutive frames of images included in the experimental video data at a preset sampling frequency, and then performs sliding window sampling on the downsampled multiple frames of images to obtain window segments of equal length. Each window segment includes T frames of images. Then, difference processing is performed on the RGB stream of the T frames of images to obtain the RGB difference stream. In this way, we obtain a video sample, and each video sample is the RGB stream and the RGB difference stream corresponding to T frames of images.
[0089] In one possible implementation manner, referring to Figure 4 , the above S2 specifically includes the following steps:
[0090] S21. Downsample the experimental video data at a preset sampling frequency f, and perform sliding window sampling processing based on a preset step size s to obtain multiple samples including T frames of images.
[0091] Among them, s is less than T, and s and f are positive integers.
[0092] Exemplarily, the preset step size is 25, and the sampling frequency is 3 frames per second.
[0093] S22. Determine the RGB stream of each sample including T-frame images.
[0094] RGB stream has the following data structure:
[0095] ;
[0096] where H is the height of the image resolution, W is the width of the image resolution, and 3 is the number of channels of the image.
[0097] S23. Perform differential processing on the subsequent frame and the previous frame images of every two adjacent frames of the RGB stream of each sample including T-frame images to obtain the RGB differential stream of each sample including T-frame images, so as to obtain multiple video samples.
[0098] RGB differential stream has the following data structure:
[0099] .
[0100] It can also be understood that: after obtaining the experimental video data, the method provided by the present invention first performs a sliding window sampling operation to obtain multiple samples, each sample includes the RGB stream of T-frame images, and then performs differential processing on the RGB stream of each sample to obtain the RGB differential stream corresponding to each RGB stream. In this way, multiple video samples are obtained, and each video sample includes the RGB stream and the RGB differential stream of T-frame images.
[0101] The method provided by the embodiments of the present invention can reduce the redundancy of the RGB stream in the case of a large amount of experimental video data of the variable gravity particle experiment through sliding window sampling processing of the experimental video data of the variable gravity particle experiment, and thus can significantly reduce the amount of data, reduce the requirements for computing resources and storage space, and improve the processing efficiency.
[0102] S3. For any video sample, perform feature extraction operations on the RGB stream and the RGB differential stream to obtain the feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and the feature vectors P0 and P1 with different semantic levels corresponding to the RGB differential stream;
[0103] In a possible implementation manner, the above S3 specifically includes the following steps:
[0104] For any video sample, input the RGB stream into the dilated 3D convolutional network, and output the feature vector A0 through the first node and the feature vector A1 through the second node; perform non-local operations and 3D convolutional operations on the feature vector A0 and the feature vector A1 respectively to obtain the feature vector L0 and the feature vector L1; for any video sample, input the RGB difference stream into the dilated 3D convolutional network, and output the feature vector B0 through the first node and the feature vector B1 through the second node; perform non-local operations and 3D convolutional operations on the feature vector B0 and the feature vector B1 respectively to obtain the feature vector P0 and the feature vector P1;
[0105] The data structure of the feature vector L0 is:
[0106] L0 ;
[0107] The data structure of the feature vector L1 is:
[0108] L1 ;
[0109] The data structure of the feature vector P0 is:
[0110] P0 ;
[0111] The data structure of the feature vector P1 is:
[0112] P1 ;
[0113] where C is the number of channels.
[0114] In an example, the data structure of the feature vector A0 is:
[0115] A0 ;
[0116] The data structure of the feature vector A1 is:
[0117] A1 ;
[0118] The data structure of the feature vector B0 is:
[0119] B0
[0120] The data structure of the feature vector B1 is:
[0121] B1 ;
[0122] In an example, the number of channels C is 512 and T is 256.
[0123] Specifically, refer to Figure 5 , Figure 5A schematic diagram of a process for determining feature vectors shown in an embodiment of the present invention. The first node is the C4 (Mixed_4f) node of the RGB branch of the dilated 3D convolutional network, and the second node is the C5 (Mixed_5c) node. The RGB stream is input into the dilated 3D convolutional (I3D) network, and the feature vector A0 is output through the C4 node, and the feature vector A1 is output through the C5 node. Then, non-local operations and convolutional operations are performed on the feature vector A0 through the Non-Local module and the 3D convolutional module to obtain the feature vector L0. Non-local operations and convolutional operations are performed on the feature vector A1 through the Non-Local module and the 3D convolutional module to obtain the feature vector L1. On the other hand, the RGB difference stream is input into the dilated 3D convolutional network, the feature vector B0 is output through the C4 node, and the feature vector B1 is output through the C5 node. Then, non-local operations and convolutional operations are performed on the feature vector B0 through the Non-Local module and the 3D convolutional module to obtain the feature vector P0. Non-local operations and convolutional operations are performed on the feature vector B1 through the Non-Local module and the 3D convolutional module to obtain the feature vector P1.
[0124] It should be understood that the Non-Local module is a neural network module for capturing long-range dependencies. The Non-Local module is widely used in tasks such as video understanding, image segmentation, and object detection, and can effectively model the global relationship between pixels (or feature points), not limited to the local neighborhood.
[0125] It should be noted that the above method for determining feature vectors is only an exemplary illustration. The method provided in the embodiments of the present invention can also use other feature extraction networks to extract features from the RGB stream and the RGB difference stream, and the embodiments of the present invention do not make special limitations on this.
[0126] S4. Construct a spatio-temporal feature pyramid for the RGB stream and the RGB difference stream according to the feature vector L0, the feature vector L1, the feature vector P0, and the feature vector P1.
[0127] Specifically, the spatio-temporal feature pyramid includes n spatio-temporal feature vectors with different time scales. That is to say, the spatio-temporal feature pyramid of the RGB stream includes n spatio-temporal feature vectors with different time scales , and the spatio-temporal feature pyramid of the RGB difference stream includes n spatio-temporal feature vectors with different time scales , where i = 0, 1, 2,..., n.
[0128] In a possible implementation manner, the determination formula for the n spatio-temporal feature vectors with different time scales included in the spatio-temporal feature pyramid of the RGB stream is as follows:
[0129] ;
[0130] ;
[0131] ;
[0132] L0, L1;
[0133] When i is greater than 0, ;
[0134] The spatio-temporal feature pyramid of the RGB difference flow includes n spatio-temporal feature vectors with different time scales The determination formula is:
[0135] ;
[0136] ;
[0137] P0, P1;
[0138] When i is greater than 0, ;
[0139] Among them, is the temporal variation information; is the weighted feature vector of the RGB flow; is the activation function, is the 1D convolution operation with a convolution kernel of size 1, is the rectified linear unit; is the learnable weight parameter of convolution kernels of different sizes, is the 1D convolution operation with a convolution kernel of size k, where k = 1, 3, 5, and 7; is the 1D convolution operation with a convolution kernel of size 3; LN means processing through the LayerNorm layer; , and = 1, 2,..., n;
[0140] The spatio-temporal feature vector and the spatio-temporal feature vector The time scales are:
[0141] ;
[0142] ;
[0143] .
[0144] The following is an example of n spatiotemporal feature vectors with different time scales provided by the embodiment of the present invention. Explain how to determine the.
[0145] For example, see Figure 6 , the feature vector L0 corresponding to the RGB stream (equivalent to ) and the eigenvector L1 (equivalent to ), the i-th eigenvector Through a change information guidance module shared with the RGB differential stream , and obtain the spatiotemporal features enhanced by temporal changes (It can also be understood as the i-th layer of the spatiotemporal feature pyramid), and then undergoes a 1D convolution with a convolution kernel size of 3 ( Except for this), we can get the spatiotemporal features of the next layer. Repeat this transformation to obtain the RGB spatiotemporal feature pyramid after the RGB differential stream is enhanced. ,Right now:
[0146] ;
[0147] 1;
[0148] in , It is The change information shared by the layer RGB spatiotemporal features and the RGB differential spatiotemporal features guides the module.
[0149] Similarly, the spatiotemporal feature pyramid of the RGB difference stream can be constructed .
[0150] When constructing the spatiotemporal feature pyramid of the RGB stream and the spatiotemporal feature pyramid of the RGB difference stream, each layer of original features and Need to go through the change information guidance module first , and obtain the spatiotemporal features after temporal variation enhancement and .
[0151] See also Figure 7 , Figure 7 A schematic diagram of the working process of a change information guidance module provided by an embodiment of the present invention. In the CIGM module, a shared structure similar to the Inception network is designed for the RGB stream and the RGB difference stream, and convolution kernels of different sizes with learnable weights are used to alleviate the difference in time scale between long and short processes, namely:
[0152] In the CIGM module, a shared Inception-like network structure is designed for the RGB stream and the RGB difference stream. Different-sized convolutional kernels with learnable weights are used to alleviate the differences in the time scale between long and short processes, that is:
[0153] ;
[0154] Among them, refers to the weighted feature after multiple convolutional kernels, refers to different convolutional kernel sizes, refers to the LayerNorm layer, and is the learnable weight parameter corresponding to the convolution, initialized to 1. Additionally, when constructing the CIGM module, if the convolutional kernel size is greater than the of the current input feature, then this convolutional layer is not constructed.
[0155] The core of the CIGM module is to calculate the temporal change information only using the RGB difference stream , in order to utilize the more prominent temporal change information in the RGB difference stream and enhance the learning of the RGB stream in the corresponding time dimension:
[0156]
[0157]
[0158]
[0159] Among them, the temporal change information first passes through max pooling to highlight local changes, and then is calculated using a 1D convolution with a convolutional kernel size of 1 and the Sigmoid function, and then the spatio-temporal feature after being enhanced by CIGM is obtained, along with .
[0160] S5. Perform weighted fusion on the spatio-temporal feature pyramids of the RGB stream and the spatio-temporal feature pyramids of the RGB difference stream corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample;
[0161] The n spatio-temporal feature vectors included in the fused spatio-temporal feature pyramid are determined by the formula:
[0162] ;
[0163] Among them, and are learnable weight parameters.
[0164] Specifically, refer to Figure 8 , and perform weighted fusion on the spatio-temporal feature pyramid of the RGB stream corresponding to each video sample and the spatio-temporal feature pyramid of the RGB difference stream to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample . Among them, the fused spatio-temporal feature pyramid includes multiple spatio-temporal feature vectors with different time scales. In this example, the fused spatio-temporal feature pyramid includes 5 spatio-temporal feature vectors with different time scales.
[0165] S6. Determine the process detection result of the experimental video data according to the fused spatio-temporal feature pyramid corresponding to each video sample through the trained process detection unit.
[0166] Among them, the process detection result of the experimental video data includes at least one process category, as well as the start time and end time corresponding to each process category.
[0167] Specifically, the trained process detection unit determines the process detection result of each video sample according to the fused spatio-temporal feature pyramid corresponding to each video sample. Then, summarize and analyze the process detection results of all video samples to obtain the process detection result of the experimental video data.
[0168] Optionally, the process detection unit includes a rough detection subunit and a fine detection subunit: the rough detection subunit is used to determine multiple proposals according to the fused spatio-temporal feature pyramid corresponding to each video sample, and each proposal includes the left and right boundary distances and the process category; the fine detection subunit is used to determine the boundary features of each proposal based on the saliency refinement algorithm, and optimize the boundary positions of each proposal based on the boundary features of each proposal to obtain the process detection result included in each video sample.
[0169] As can be seen from the above S1 - S6, the method provided by the present invention determines multiple video samples by means of sliding window sampling according to the experimental video data of the variable gravity particle experiment, then constructs a spatio - temporal feature pyramid based on the RGB stream and the RGB difference stream of each video sample, then performs weighted fusion on the spatio - temporal feature pyramids constructed according to the RGB stream and the RGB difference stream to obtain a fused spatio - temporal feature pyramid, and finally determines the process category of each video sample according to the multiple spatio - temporal feature vectors included in the fused spatio - temporal feature pyramid, and then obtains the process categories included in the experimental video data of the variable gravity particle experiment, as well as the start time and end time of each process category. The method provided by the present invention can quickly and accurately determine the multiple stage processes included in the spatial variable gravity particle material experiment according to the experimental video data, improving the data processing efficiency. On the other hand, by introducing the RGB difference stream, the method provided by the present invention can construct the RGB difference stream to replace the optical flow according to the characteristic that the background of the experimental video data changes little, so as to better perceive the temporal changes of the experimental process categories; moreover, the method provided by the present invention constructs the spatio - temporal feature pyramids of the RGB stream and the RGB difference stream based on the temporal change information of the RGB difference stream, which can effectively strengthen the learning on the RGB stream, thereby improving the perception ability of the experimental process on the RGB stream; the spatio - temporal feature pyramid provided by the present invention includes feature vectors of multiple different time scales, and uses learnable weighted fusion to construct the fused spatio - temporal feature pyramid, so as to capture richer spatio - temporal semantic information, which can effectively improve the detection accuracy and meet the usage requirements of technicians in different usage scenarios.
[0170] In some embodiments, referring to Figure 9 , before the above S6, the method provided by the embodiments of the present invention further includes:
[0171] S81. Obtain a training sample set for the variable gravity particle experiment, where the training sample set includes multiple training samples, and each training sample includes a video sample and a process category.
[0172] S82. Construct a target loss function.
[0173] Among them, the target loss function is:
[0174] ;
[0175] Among them, the classification loss function of the rough detection subunit, the regression loss function of the rough detection subunit; the classification loss function of the fine detection subunit, the regression loss function of the fine detection subunit.
[0176] Specifically, the classification loss uses the Focal Loss function as the training objective; the regression loss uses the T DIoU loss in the fast and slow dual-stream network. The regression of the fine detection subunit uses the Smooth L1 Loss function for s branches to predict the boundary offset. The binary cross-entropy loss (BCE Loss) function is used as the training objective for the positioning quality of the fine detection subunit.
[0177] S83. Based on the objective loss function, the process detection unit is iteratively trained to obtain a trained process detection unit.
[0178] The method provided by the embodiment of the present invention can effectively solve the problem of imbalance between training samples of different process categories and predicted positive and negative samples by constructing an objective loss function, and effectively improve the accuracy and robustness of the process detection unit.
[0179] The beneficial effects of a method for detecting the variable gravity particle experiment process based on dual-stream difference provided by the embodiment of the present invention are exemplarily described below with an example.
[0180] Exemplarily, the beneficial effects of the method provided by the embodiment of the present invention are verified through a validation dataset. The validation dataset is composed of experimental videos of the particle material A bin of the variable gravity experiment cabinet collected by the panoramic camera in the space laboratory, including 207 video segments. Each video captures rich experimental process details at a resolution of 1920×1080 and a frame rate of 25 frames per second, and the average video duration reaches 9.4 minutes. Among the 207 video segments, a total of 9 different types of process categories are labeled, with a total of 939 process instances. The average duration of each instance is about 29 seconds. In addition, there is no time overlap between all instances.
[0181] See Figure 10 , Figure 10 It is a schematic diagram of the process categories and durations of the space variable gravity particle material experiment of a validation dataset provided by the embodiment of the present invention. The validation dataset has more long processes, and the average duration of some processes is significantly longer. For example, the average duration of the baffle vibration process can reach 50 seconds. At the same time, there are also significant differences in the durations of some processes. Compared with the baffle vibration, the average duration of the left shift of the baffle before vibration is only 1.5 seconds.
[0182] In this example, the method provided by the embodiments of the present invention is verified based on the evaluation metrics Mean Average Precision (mAP) and Average Precision (AP). The mAP values with a step size of 0.1 on tIoU = [0.3, 0.7] are used for comparison, where tIoU refers to the intersection over union in terms of time.
[0183] Average Precision is an evaluation metric for a single process category. Average Precision is used to measure the average precision at different recall rate levels, while mAP is the average of the APs for all process categories. Specifically, when calculating AP, the prediction results are sorted according to their confidence levels, then the precision at different recall rate thresholds is calculated, and finally the average precision is obtained through integration or interpolation.
[0184] AP is usually calculated using the 11 - point interpolation method, that is, the average of the maximum precision values at 11 points where the recall rates are 0, 0.1, 0.2,..., 1. mAP is the average of the APs for all process categories, that is:
[0185]
[0186] where C is the number of process categories.
[0187] tIoU is used to measure the overlap degree between the predicted experimental process segment and the true experimental process segment. It is the ratio of the intersection area to the union area of the predicted experimental process segment and the true experimental process segment at time t. The higher the tIoU value, the closer the prediction result is to the true result. The calculation of tIoU is:
[0188] ;
[0189] When calculating mAP, different tIoU thresholds are usually set. Only when the tIoU between the prediction result and the true result is greater than or equal to the threshold, the prediction is considered correct.
[0190] Specifically, a comparative experiment was conducted between the embodiments of the present invention and the TP2N* method and the baseline method AFSD in the related art. Among them, the TP2N* method is a method for process prediction through strategies such as fast and slow dual-frequency data sampling and a dual-path spatio-temporal pyramid network (TP2N), and the baseline method AFSD is a method that combines the I3D network with an anchor-free method. The results of the comparative experiment are shown in Table 1. TP2N* is the experimental result of the fast and slow dual-stream network under dual-frequency (13fps and 3fps) data sampling, and the rest are experimental results under a single sampling frequency (3fps). As can be seen from Table 1, the method proposed by the present invention obtains an mAP detection result of 89.28% under single-frequency (3fps) data sampling, and significant improvements are achieved under different tIoUs. It is about 7% higher than the baseline method AFSD and about 1.4% higher than the TP2N* method, indicating that the method proposed by the present invention can obtain richer spatio-temporal feature information and capture more significant temporal dynamic change information by using the dual-stream difference fusion method, so as to better identify and locate the experimental process, proving the effectiveness of the technical solution of the present invention.
[0191] Table 1
[0192]
[0193] Furthermore, to explore the influence of different data input modes on the detection accuracy, a comparative experiment on data input modes was conducted on the experimental video dataset of spatially variable gravity granular materials. The experimental results are shown in Table 2.
[0194] The method proposed by the present invention uses a data mode input of RGB stream and RGB difference stream fusion (RGB+RGB Difference). The AP under different tIoU thresholds is significantly improved, and the detection accuracy mAP reaches 89.28%. Compared with the baseline method that only uses the RGB stream, the mAP is increased by about 7%. Compared with the baseline method that only uses the RGB difference stream (RGB Difference), the mAP is increased by about 5.5%. Compared with the baseline method that uses RGB stream and optical flow fusion (RGB+Optical Flow), the mAP is increased by about 5.84%. It shows that using the data mode input of RGB stream and RGB difference stream fusion (RGB+RGB Difference), and using the RGB difference stream (RGB Difference) to perceive the temporal dynamic changes and guide the RGB stream to focus on learning the experimental content in the changing interval is more effective, proving the effectiveness of the technical solution of the present invention.
[0195] Table 2
[0196]
[0197] Furthermore, to verify the functions of the various unit modules included in the detection system 100 in the technical solution of the present invention, a module ablation experiment was conducted on the experimental video dataset of spatially variable gravity granular materials.
[0198] Table 3
[0199]
[0200] The results of the module ablation experiment are shown in Table 3, where the Baseline method refers to the AFSD method, Diff refers to RGB Difference, and the simultaneous use of RGB and Diff represents the use of two-stream fusion. It can be seen that by fusing the RGB stream with the RGB difference stream on the basis of the baseline method, a relatively large improvement can be obtained. Compared with the baseline method that only uses the RGB stream, it is improved by about 7%, and compared with the baseline method that only uses the RGB difference stream, it is improved by about 5.5%, indicating that two-stream difference fusion is helpful for process detection. On the basis of using the fusion of the RGB stream and the RGB difference stream, the Change Information Guidance Module (CIGM) designed in the technical solution of the present invention uses the significant temporal change information on the difference stream to strengthen the learning on the RGB stream, and is further improved by about 2%, indicating the effectiveness of the CIGM module.
[0201] Furthermore, to explore the influence of the number of CIGM modules in the technical solution of the present invention on the detection accuracy, an ablation experiment on the number of CIGM modules was conducted on the validation dataset. The experimental results are shown in Table 4. It can be seen from it that no matter how many CIGM modules there are, as long as the CIGM module is used, the mAP of its detection result is higher than that of the result without using the CIGM module (87.02%). Among them, when the number of CIGM modules is 3, the performance of the method provided by the present invention is the best, and its detection accuracy mAP reaches 89.28%. The number of CIGM modules is not the more the better. As the number of CIGM modules increases, although the detection accuracy mAP of the model generally shows a trend of first rising and then falling, it is higher than the result without using the CIGM module, indicating the effectiveness of the CIGM module proposed in the technical solution of the present invention.
[0202] Table 4
[0203]
[0204] The above mainly introduces the solutions of the embodiments of the present invention from the perspective of methods. It can be understood that in order for the detection system 100 to implement the above functions, it includes at least one of the corresponding hardware structures and software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.
[0205] The embodiments of the present invention can divide the detection system 100 into functional units according to the above method examples. For example, the detection system 100 can be divided into respective functional units corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiments of the present invention is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0206] Exemplarily, Figure 11The figure shows a schematic hardware structure of a detection system provided by an embodiment of the present invention. The detection system 100 includes: a data acquisition unit 110, configured to acquire experimental video data of a variable gravity particle experiment; a sample sampling unit 120, configured to perform a sliding window sampling process on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples, each video sample including an RGB stream and an RGB difference stream of T frames of images, where the RGB difference stream is obtained by performing a difference process on the RGB stream, and T is a positive integer greater than 1; a feature extraction unit 130, configured to perform a feature extraction operation on the RGB stream and the RGB difference stream for any video sample to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream, and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; a feature construction unit 140, configured to construct a spatio-temporal feature pyramid of the RGB stream and the RGB difference stream according to the feature vectors L0, L1, P0, and P1, the spatio-temporal feature pyramid including n spatio-temporal feature vectors with different time scales; a feature fusion unit 150, configured to perform weighted fusion on the spatio-temporal feature pyramid of the RGB stream and the spatio-temporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample; and a detection unit 160, configured to determine a process detection result of the experimental video data through a trained process detection unit according to the fused spatio-temporal feature pyramid corresponding to each video sample, the process detection result of the experimental video data including at least one process category, and a start time and an end time corresponding to each process category.
[0207] It should be understood that for the specific descriptions of the above optional manners, reference may be made to the foregoing method embodiments, which will not be elaborated herein. In addition, for the explanations and beneficial effects descriptions of any of the above detection systems 100, reference may be made to the corresponding method embodiments above, which will not be elaborated.
[0208] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the methods of the above various embodiments. For the explanations and beneficial effects descriptions of the relevant contents in any of the above computer-readable storage media, reference may be made to the corresponding embodiments above, which will not be elaborated herein.
[0209] An embodiment of the present invention further provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the above detection system 100. Optionally, the functions supported by the chip may refer to the above, which will not be elaborated herein.
[0210] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
[0211] The embodiments of the present invention also provide a computer program product containing instructions. When the instructions run on a computer, the computer is made to execute any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.
[0212] It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided by the embodiments of the present invention, such as but not limited to, the above-mentioned memory, computer-readable storage medium, communication chip, etc., are all non-transitory. Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. The computer-readable storage medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0213] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A detection method for the variable gravity particle experiment process based on double - flow difference, characterized in that, The method includes: Obtaining experimental video data of variable gravity particle experiments; Based on a preset step size, performing sliding window sampling processing on the experimental video data at a preset sampling frequency to obtain a plurality of video samples, each of the video samples including an RGB stream and an RGB difference stream of T frames of images, where the RGB difference stream is obtained by performing differential processing on the RGB stream, and T is a positive integer greater than 1; For any one of the video samples, performing feature extraction operations on the RGB stream and the RGB difference stream to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream; Constructing a spatio-temporal feature pyramid of the RGB stream and the RGB difference stream according to the feature vectors L0, L1, P0, and P1, the spatio-temporal feature pyramid including n spatio-temporal feature vectors with different time scales; Performing weighted fusion on the spatio-temporal feature pyramid of the RGB stream and the spatio-temporal feature pyramid of the RGB difference stream corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample; Determining a process detection result of the experimental video data according to the fused spatio-temporal feature pyramid corresponding to each video sample through a process detection unit completed by training, the process detection result of the experimental video data including at least one process category, and a start time and an end time corresponding to each process category.
2. The method according to claim 1, characterized in that, The performing sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples includes: Performing downsampling on the experimental video data at a preset sampling frequency f, and performing sliding window sampling processing based on a preset step size s to obtain a plurality of samples including T frames of images, where s is less than T, and s and f are positive integers; Determining the RGB stream of each of the samples including T frames of images; Performing differential processing on the subsequent frame and the previous frame image of every two adjacent frames of the RGB stream of each of the samples including T frames of images to obtain an RGB difference stream of each of the samples including T frames of images, so as to obtain the plurality of video samples; RGB stream The data structure is as follows: ; RGB differential flow The data structure is as follows: ; Wherein, H is the resolution height of the image, W is the resolution width of the image, and 3 is the number of channels of the image.
3. The method according to claim 2, wherein The performing feature extraction operations on the RGB stream and the RGB difference stream for any one of the video samples to obtain feature vectors L0 and L1 with different semantic levels corresponding to the RGB stream; and feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference stream includes: For any one of the video samples, inputting the RGB stream into an inflated 3D convolutional network to output a feature vector A0 through a first node and output a feature vector A1 through a second node; Performing non-local operations and 3D convolutional operations on the feature vector A0 and the feature vector A1 respectively to obtain the feature vector L0 and the feature vector L1; For any of the video samples, input the RGB difference flow into the dilated 3D convolutional network, and output the feature vector B0 through the first node and the feature vector B1 through the second node; Perform non-local operations and 3D convolutional operations on the feature vector B0 and the feature vector B1 respectively to obtain the feature vector P0 and the feature vector P1; The data structure of the feature vector L0 is: L0 ; The data structure of the feature vector L1 is: L1 ; The data structure of the feature vector P0 is: P0 ; The data structure of the feature vector P1 is: P1 ; Where C is the number of channels.
4. The method according to claim 3, wherein The spatio-temporal feature pyramid of the RGB stream includes n spatio-temporal feature vectors with different time scales The determination formula is as follows: ; ; ; L0, L1; When i is greater than 0, ; The spatio-temporal feature pyramid of the RGB differential flow includes n spatio-temporal feature vectors with different time scales The determination formula is as follows: ; ; P0, P1; When i is greater than 0, ; Among them, is the time series change information; is the weighted feature vector of the RGB stream; is the activation function, is the 1D convolution operation of the convolutional kernel with size 1, is the rectified linear unit; is the learnable weight parameter of convolutional kernels of different sizes, is the 1D convolution operation of the convolutional kernel with size k, where k = 1, 3, 5, and 7; is the 1D convolution operation of the convolutional kernel with size 3; LN means processing through the LayerNorm layer; , and = 1, 2,..., n; Spatio-temporal feature vector and spatio-temporal feature vector have a time scale of: ; ; 。 5. The method according to claim 4, characterized in that The n spatio-temporal feature vectors with different time scales included in the fused spatio-temporal feature pyramid are determined by the following formula: ; Among them, and are learnable weight parameters.
6. The method according to claim 5, wherein The process detection unit includes a rough detection subunit and a fine detection subunit: The rough detection subunit is used to determine a plurality of proposals according to the fused spatio-temporal feature pyramid corresponding to each video sample, and each proposal includes the left and right boundary distances and the process category; The fine detection subunit is used to determine the boundary features of each proposal based on the saliency refinement algorithm, and optimize the boundary positions of each proposal based on the boundary features of each proposal to obtain the process detection results included in each video sample.
7. The method according to claim 6, wherein Before the process detection unit determines the process detection results of the experimental video data through training according to the fused spatio-temporal feature pyramid corresponding to each video sample, the method further includes: Obtain a training sample set of variable gravity particle experiments, where the training sample set includes a plurality of training samples, and each training sample includes a video sample and a process category; Construct an objective loss function; Based on the objective loss function, perform iterative training on the process detection unit to obtain a trained process detection unit.
8. The method according to claim 7, wherein The target loss function is as follows: ; Among them, the classification loss function of the rough detection subunit, the regression loss function of the rough detection subunit; the classification loss function of the fine detection subunit, the regression loss function of the fine detection subunit.
9. A variable gravity particle experiment process detection system based on double-stream difference, characterized in that The system includes: A data acquisition unit for acquiring experimental video data of variable gravity particle experiments; A sample sampling unit for performing sliding window sampling processing on the experimental video data at a preset sampling frequency based on a preset step size to obtain a plurality of video samples, and each video sample includes the RGB flow and the RGB difference flow of T frame images, where the RGB difference flow is obtained by performing differential processing on the RGB flow, and T is a positive integer greater than 1; A feature extraction unit for, for any of the video samples, performing feature extraction operations on the RGB flow and the RGB difference flow to obtain the feature vectors L0 and L1 with different semantic levels corresponding to the RGB flow; and the feature vectors P0 and P1 with different semantic levels corresponding to the RGB difference flow; A feature construction unit for constructing a spatio-temporal feature pyramid of the RGB flow and the RGB difference flow according to the feature vectors L0, L1, P0, and P1, where the spatio-temporal feature pyramid includes n spatio-temporal feature vectors with different time scales; A feature fusion unit for performing weighted fusion on the spatio-temporal feature pyramid of the RGB flow and the spatio-temporal feature pyramid of the RGB difference flow corresponding to each video sample to obtain a fused spatio-temporal feature pyramid of n spatio-temporal feature vectors with different time scales corresponding to each video sample; A detection unit, configured to determine a process detection result of the experimental video data according to the fusion spatio-temporal feature pyramid corresponding to each of the video samples through a process detection unit completed by training, where the process detection result of the experimental video data includes at least one process category, and a start time and an end time corresponding to each process category.
10. An electronic device, characterized in that, Comprising: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the variable gravity particle experiment process detection method based on two-stream difference as described in any one of claims 1-8.
Citation Information
Patent Citations
Video recognition method based on space-time pyramid network
CN107909041A
Fine-grained video action recognition method based on hierarchical structure
CN113139467A