A line inspection channel crowd behavior recognition method and system based on multi-modal data fusion
By using multimodal data fusion technology, video streams and 3D point cloud streams collected by visible light cameras and depth cameras are combined with orthogonalization processing of graph neural networks to solve the problem of accuracy and reliability of neural network recognition in line inspection channels, and achieve more stable behavior detection.
Patent Information
- Application Number
- CN202511353473.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-09-22
AI Technical Summary
In the line inspection channel, existing technologies struggle to guarantee the accuracy and reliability of neural network recognition in deep structures.
A multimodal data fusion method is adopted, which uses a visible light camera and a depth camera to acquire video streams and 3D point cloud streams respectively. The processing device extracts time-series image feature sequences and time-series 3D point cloud feature sequences, and performs orthogonal aggregation processing through graph neural networks to improve the accuracy and reliability of behavior detection.
It enhances the stability and reliability of behavior detection in line inspection channels, reduces information overlap, strengthens complementarity, and improves the accuracy of detection results.
Smart Images

Figure CN120853224B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a line inspection passage crowd behavior recognition method and system based on multi-modal data fusion. BACKGROUND
[0002] With the acceleration of urbanization and the continuous improvement of infrastructure construction, the safe operation of power lines is crucial to ensure the normal operation of the city. As an important part of the power line, the safety of the line inspection passage directly affects the stability and reliability of the power system. However, traditional inspection methods such as manual inspection or simple monitoring equipment have been unable to meet the current demand for efficiency, accuracy and safety. Therefore, it is particularly important to introduce advanced technical means for line inspection passage crowd behavior recognition.
[0003] Artificial intelligence and computer vision technology: In recent years, the rapid development of artificial intelligence (AI) and computer vision technology has provided new possibilities for solving object recognition and behavior analysis in complex scenarios. Deep learning algorithms, especially convolutional neural networks (CNN), have made significant achievements in image recognition, making it possible to accurately identify specific objects and behaviors from massive video data. Therefore, in some current technologies, line inspection passage crowd behavior can be identified by combining artificial intelligence to improve monitoring efficiency and reduce monitoring costs.
[0004] However, due to the longitudinal structure of the line inspection passage, how to ensure the recognition accuracy and reliability of the neural network under this structure is a problem in current research. SUMMARY
[0005] The embodiments of the present application provide a line inspection passage crowd behavior recognition method and system based on multi-modal data fusion to improve the accuracy of crowd behavior recognition in line inspection passage scenarios.
[0006] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:
[0007] In a first aspect, a line inspection passage crowd behavior recognition method based on multi-modal data fusion is provided, applied to a processing device, the method comprising: the processing device acquiring a video stream of the line inspection passage collected by a visible light camera; the processing device acquiring a 3D point cloud stream of the line inspection passage collected by a depth camera; the processing device extracting a time sequence image feature sequence from the video stream and a time sequence 3D point cloud feature sequence from the 3D point cloud stream, wherein the time sequence image feature sequence and the time sequence 3D point cloud feature sequence are feature sequences of the same target time; the processing device performing aggregation processing on the time sequence image feature sequence and the time sequence 3D point cloud feature sequence by a graph neural network to obtain a behavior detection result of the line inspection passage at the target time.
[0008] In a possible implementation, the sequence of time-series image features is an initial sequence of time-series image features, and the sequence of time-series 3D point cloud features is an initial sequence of time-series 3D point cloud features; the processing device performs, on the sequence of time-series image features and the sequence of time-series 3D point cloud features, an aggregation processing by a graph neural network to obtain a behavior detection result of the line inspection channel at the target time, including: the processing device aligns the initial sequence of time-series image features and the initial sequence of time-series 3D point cloud features in terms of feature sequence length to obtain an aligned sequence of time-series image features and an aligned sequence of time-series 3D point cloud features; the processing device orthogonally assigns the aligned sequence of time-series image features and the aligned sequence of time-series 3D point cloud features by using an orthogonal sequence template to obtain a sequence of time-series image features tending to be orthogonal and a sequence of time-series 3D point cloud features tending to be orthogonal; and the processing device takes the sequence of time-series image features tending to be orthogonal as graph nodes and takes the sequence of time-series 3D point cloud features tending to be orthogonal as edges, and performs, by using the graph neural network, an aggregation processing to obtain the detection result.
[0009] In a possible implementation, the orthogonal sequence template includes a group of sequences, and any two sequences in the group of sequences are orthogonal; the processing device orthogonally assigns the aligned sequence of time-series image features and the aligned sequence of time-series 3D point cloud features by using the orthogonal sequence template to obtain the sequence of time-series image features tending to be orthogonal and the sequence of time-series 3D point cloud features tending to be orthogonal, including: the processing device selects a first sequence from the group of sequences to orthogonally assign the aligned sequence of time-series image features to obtain the sequence of time-series image features tending to be orthogonal, and selects a second sequence from the group of sequences to orthogonally assign the aligned sequence of time-series 3D point cloud features to obtain the sequence of time-series 3D point cloud features tending to be orthogonal; the first sequence has the same length as the aligned sequence of time-series image features, and the second sequence has the same length as the aligned sequence of time-series 3D point cloud features; and the tending to be orthogonal refers to the sequence of time-series image features tending to be orthogonal and the sequence of time-series 3D point cloud features tending to be orthogonal being orthogonal to each other.
[0010] In a possible implementation, the first sequence and the second sequence have a length of 96.
[0011] The first sequence orthogonally assigning the aligned sequence of time-series image features is as follows:
[0012] In a possible implementation, the first sequence and the second sequence have a length of 96.
[0013] ;
[0014] wherein, is the i th feature vector in the aligned sequence of time-series image features, is the i th vector in the first sequence, is the i th vector in the first sequence, is the i th vector in the first sequence, The first element in a temporally orthogonal image feature sequence represents the orthogonal element in the sequence. 1 eigenvector;
[0015] The second sequence represents the aligned temporal 3D point cloud feature sequence orthogonally assigned as follows:
[0016] ;
[0017] in, For the alignment of the temporal 3D point cloud feature sequence, the first... 1 eigenvector For the second sequence A vector, The first term in the temporally orthogonal 3D point cloud feature sequence represents the first term. 1 eigenvector.
[0018] In one possible implementation, the orthogonal sequence template includes a first sequence group and a second sequence group, where any two sequences in the first or second sequence group are orthogonal, and any sequence in the first sequence group is quasi-orthogonal to any sequence in the second sequence group. The processing device performs orthogonal assignment on the aligned temporal image feature sequence and the aligned temporal 3D point cloud feature sequence using the orthogonal sequence template to obtain a near-orthogonal temporal image feature sequence and a near-orthogonal temporal 3D point cloud feature sequence. This includes: the processing device equally dividing the aligned temporal image feature sequence into a first temporal image feature subsequence and a second temporal image feature subsequence; and the processing device equally dividing the aligned temporal 3D point cloud feature sequence into a first temporal 3D point cloud feature subsequence and a second temporal 3D point cloud feature subsequence; the processing device uses the first sequence from the first sequence group to orthogonally assign values to the first temporal image feature subsequence to obtain a near-orthogonal temporal image feature sequence and a near-orthogonal temporal 3D point cloud feature sequence. The processing device orthogonally assigns a first temporal image feature subsequence to the first temporal image feature subsequence using the second sequence in the first sequence group, and orthogonally assigns a third sequence in the second sequence group to the second temporal image feature subsequence to obtain a tend-orthogonal second temporal image feature subsequence. The orthogonal first temporal image feature subsequence is then concatenated with the second sequence in the second sequence group to obtain a tend-orthogonal temporal image feature sequence.
[0019] In one possible implementation, the lengths of the first sequence, the second sequence, the third sequence and the fourth sequence are all 48, the first sequence and the second sequence are even-indexed Zadoff-Chu sequences, and the third sequence and the fourth sequence are odd-indexed Zadoff-Chu sequences; the length of the aligned time sequence image feature sequence and the length of the aligned time sequence 3D point cloud feature sequence are both 96;
[0020] The first sequence orthogonally assigns the first time sequence image feature subsequence as follows:
[0021] ;
[0022] wherein, is the i th feature vector in the first time sequence image feature subsequence, is the i th vector in the first sequence, denotes the i th feature vector in the first time sequence image feature subsequence tending to be orthogonal;
[0023] The second sequence orthogonally assigns the first time sequence 3D point cloud feature subsequence as follows:
[0024] ;
[0025] wherein, is the i th feature vector in the first time sequence 3D point cloud feature subsequence, is the i th vector in the second sequence, denotes the i th feature vector in the first time sequence 3D point cloud feature subsequence tending to be orthogonal;
[0026] The third sequence orthogonally assigns the second time sequence image feature subsequence as follows:
[0027] ;
[0028] wherein, is the i th feature vector in the second time sequence image feature subsequence, is the i th vector in the third sequence, denotes the i th feature vector in the second time sequence image feature subsequence tending to be orthogonal;
[0029] The second sequence orthogonally assigns the first time sequence 3D point cloud feature subsequence as follows:
[0030] ;
[0031] wherein, is the i-th feature vector in the second time-sequential 3D point cloud feature sequence, is the i-th vector in the fourth sequence, is the i-th feature vector in the second time-sequential 3D point cloud feature sub-sequence, is the i-th vector in the fourth sequence, is the i-th feature vector in the second time-sequential 3D point cloud feature sub-sequence.
[0032] In a possible implementation, the processing device aligns the length of the initial time-sequential image feature sequence and the initial time-sequential 3D point cloud feature sequence to obtain an aligned time-sequential image feature sequence and an aligned time-sequential 3D point cloud feature sequence, including: the processing device performs vector fusion on randomly selected feature vectors in the initial time-sequential image feature sequence to obtain the aligned time-sequential image feature sequence, and performs vector fusion on randomly selected feature vectors in the initial time-sequential 3D point cloud feature sequence to obtain the aligned time-sequential 3D point cloud feature sequence, the length of the aligned time-sequential image feature sequence and the length of the aligned time-sequential 3D point cloud feature sequence are both preset feature sequence lengths.
[0033] In a possible implementation, the processing device extracts the time-sequential image feature sequence from the video stream and extracts the time-sequential 3D point cloud feature sequence from the 3D point cloud stream, including: the processing device extracts the image video frame at the target time from the video stream and extracts the 3D point cloud data at the target time from the 3D point cloud stream; the processing device extracts the time-sequential image feature sequence from the image video frame and extracts the time-sequential 3D point cloud feature sequence from the 3D point cloud data.
[0034] In a possible implementation, the processing device extracts the image video frame at the target time from the video stream and extracts the 3D point cloud data at the target time from the 3D point cloud stream, including: the processing device determines the latest two timeout time points of a preset periodic timer, and the timeout time point is the target time; the processing device extracts the image video frame closest to the timeout time point from the video stream and extracts the 3D point cloud data closest to the timeout time point from the 3D point cloud stream.
[0035] In a possible implementation, the visible light camera and the depth camera have the same shooting direction, both along the longitudinal direction of the line inspection channel; thus, the processing device extracts a time sequence of image features from the image video frames and a time sequence of 3D point cloud features from the 3D point cloud data, including: the processing device performs pixel edge extraction on the image video frames to obtain sub-images containing only objects in the image video frames; the processing device extracts 3D point cloud sub-data corresponding to the coordinate range from the 3D point cloud data according to the coordinate range of the sub-images in the image video frames; the processing device extracts features of the sub-images by using a lightweight convolutional neural network and a conversion model encoder to obtain the time sequence of image features; and the processing device extracts features of the 3D point cloud sub-data by using a point cloud semantic segmentation neural network and a three-dimensional convolution-long short-term memory network to obtain the time sequence of 3D point cloud features.
[0036] In a possible implementation, the 3D point cloud data includes a position parameter of the 3D point cloud; and the processing device extracts 3D point cloud sub-data corresponding to the coordinate range from the 3D point cloud data according to the coordinate range of the sub-images in the image video frames, including: the processing device maps the 3D point cloud data to a plane space according to the position parameter of the 3D point cloud to obtain 2D point cloud data; the plane space is a plane on which the image video frames are located, and the plane space has the same size as the image video frames; the processing device determines 2D point sub-cloud data located in the coordinate range from the 2D point cloud data; and the processing device inversely maps the 2D point sub-cloud data to a 3D space to obtain the 3D point cloud sub-data.
[0037] In a second aspect, a line inspection channel crowd behavior recognition system based on multi-modal data fusion is provided, including a visible light camera, a depth camera, and a processing device; the processing device is configured to: the processing device acquires a video stream of the line inspection channel collected by the visible light camera; the processing device acquires a 3D point cloud stream of the line inspection channel collected by the depth camera; the processing device extracts a time sequence of image features from the video stream and a time sequence of 3D point cloud features from the 3D point cloud stream, where the time sequence of image features and the time sequence of 3D point cloud features are feature sequences of the same target moment; and the processing device obtains a behavior detection result of the line inspection channel at the target moment by performing GNN aggregation processing tending to orthogonality on the time sequence of image features and the time sequence of 3D point cloud features.
[0038] The processing device is further specifically configured to perform the method in the first aspect.
[0039] In a third aspect, a computer-readable storage medium is provided, including: a computer program or instructions; when the computer program or instructions run on a computer, the computer program or instructions cause the computer to perform the method in the first aspect.
[0040] In a fourth aspect, a computer program product is provided, comprising a computer program or instructions which, when run on a computer, cause the computer to perform the method of the first aspect.
[0041] In summary, based on the video stream of the line inspection channel collected by the visible light camera and the 3D point cloud stream of the line inspection channel collected by the depth camera, the processing device can extract the time sequence image feature sequence and the time sequence 3D point cloud feature sequence located at the same target moment, respectively. Thus, the processing device can perform the aggregation processing by the graph neural network on the time sequence image feature sequence and the time sequence 3D point cloud feature sequence to reduce the information overlap and enhance the complementarity, so that the behavior detection result of the line inspection channel at the target moment is more stable, i.e., can have higher detection failure stability and reliability. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 A framework schematic diagram of the line inspection channel crowd behavior recognition system based on multi-modal data fusion provided by the embodiments of the present application is provided.
[0043] Figure 2 A flowchart schematic diagram of the line inspection channel crowd behavior recognition method based on multi-modal data fusion provided by the embodiments of the present application is provided.
[0044] Figure 3 A structure schematic diagram of the processing device provided by the embodiments of the present application is provided. DETAILED DESCRIPTION
[0045] The technical solutions in the present application will be described below with reference to the drawings.
[0046] The "predefined" or "preconfigured" can be realized by pre-storing corresponding codes, tables or other means for indicating related information in the device, and the embodiments of the present application do not limit the specific implementation manner. Wherein, "storing" can mean storing in one or more memories. The one or more memories can be separately arranged or integrated in the encoder or decoder, processor, or communication device. The one or more memories can be partially separately arranged and partially integrated in the decoder, processor, or communication device. The type of the memory can be any form of storage medium, and the embodiments of the present application do not limit this.
[0047] The "protocol" involved in the embodiments of the present application can refer to the protocol family in the communication field, the standard protocol similar to the frame structure of the protocol family, or the related protocol applied to the future communication system, and the embodiments of the present application do not make specific limitations.
[0048] In the embodiments of the present application, "when", "in the case of", "if" and the like all refer to the device making corresponding processing under certain objective circumstances, and are not limited to time, and do not require the device to have a judgment action when implemented, nor does it mean that there are other limitations.
[0049] In the description of the embodiments of the present application, unless otherwise specified, " / " represents that the objects before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the embodiments of the present application is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. In addition, in the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or the like means any combination of the items, including any combination of single item or multiple items. For example, at least one of a, b or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second", and the like are used to distinguish the same items or similar items with basically the same function and role. Those skilled in the art can understand that "first", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different. At the same time, in the embodiments of the present application, "exemplary" or "for example" means to serve as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, "exemplary" or "for example" is used to present the relevant concept in a specific manner, for understanding.
[0050] In order to facilitate understanding of the embodiments of the present application, first take the line inspection passage crowd behavior recognition system based on multi-modal data fusion shown in Figure 1 The embodiments of the present application are described in detail. The exemplary, Figure 1 The architecture of a line inspection passage crowd behavior recognition system based on multi-modal data fusion suitable for the method provided by the embodiments of the present application is shown.
[0051] As shown in Figure 1 The line inspection passage crowd behavior recognition system based on multi-modal data fusion includes a visible light camera, a depth camera, and a processing device.
[0052] The visible light camera can be a conventional monitoring probe, a monitoring camera, etc., such as a DH-IPC-HFW3449T-AS4MP starlight level cylinder machine, a UNV-IPC2122SR3-LA28 2.8mm wide-angle gun machine, etc., and the specific model is not limited, and is only some examples.
[0053] The depth camera can be a 3D camera, or a 3D camera, such as Astra+ embedded 3D camera, or Intel RealSense D455, etc., and the specific model is not limited, and is only some examples.
[0054] In the system, the shooting directions of the visible light camera and the depth camera are the same, both along the longitudinal direction of the line inspection channel. For example, for a straight channel of the line inspection channel, the visible light camera and the depth camera can be arranged above the entrance of the line inspection channel, both shooting along the longitudinal direction of the line inspection channel, so as to shoot the whole situation of the line inspection channel. For another example, the visible light camera and the depth camera can be arranged above the corner of the line inspection channel, to shoot along the longitudinal direction of the line inspection channel through wide-angle respectively. It should be noted that the scale range of the shooting of the visible light camera and the depth camera is the same, i.e. the range of the picture shot by the two is consistent, such as the visible light camera shoots region A, and the depth camera also shoots region A, to realize the subsequent feature aggregation processing in the present application.
[0055] In addition, the visible light camera and the depth camera can be deployed on the inspection platform and connected with the on-board industrial computer of the inspection platform, and the on-board industrial computer is connected with a communication module, to finally send the data collected by the visible light camera and the depth camera to the processing device through the communication module.
[0056] The processing device can be a terminal with data processing capability, which can also be referred to as user equipment (UE), access terminal, subscriber unit, subscriber station, mobile station (MS), mobile, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device. The terminal in the embodiments of the present application can be a mobile phone, a cellular phone, a smart phone, a Pad, a wireless data card, a personal digital assistant (PDA), a wireless modem, a handset, a laptop computer, a machine type communication (MTC) terminal, a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical treatment, etc.
[0057] It should be understood that the line inspection passage crowd behavior recognition system based on multi-modal data fusion can perform the method in the embodiments of the present application, such as the method shown in the following Figure 2 The method, please refer to the relevant description below.
[0058] Please refer to Figure 2 The embodiments of the present application provide a line inspection passage crowd behavior recognition method based on multi-modal data fusion. The method can be performed by a processing device. The flow of the method includes:
[0059] S201, the processing device acquires a video stream of a line inspection passage collected by a visible light camera.
[0060] S202, the processing device acquires a 3D point cloud stream of the line inspection passage collected by a depth camera.
[0061] According to the system architecture shown in Figure 1 It can be known from S201-S202 that the visible light camera and the depth camera can send the data collected by themselves to the processing device through the communication module. Therefore, the processing device can receive the video stream of the line inspection passage collected by the visible light camera from the communication module, and receive the 3D point cloud stream of the line inspection passage collected by the depth camera from the communication module.
[0062] The video stream can include continuous multiple image video frames, and a timestamp corresponding to each image video frame, i.e., the time when each image video frame is captured. The 3D point cloud stream can include continuous multiple 3D point cloud data, each 3D point cloud data including the coordinate points of each point cloud in space collected by the depth camera once, the motion vector of each point cloud, etc., which can collectively describe the shape, motion state, etc. of the object in the middle line inspection channel, which can be a person or an environmental article, etc. Each 3D point cloud data also includes its own timestamp, i.e., the time when each 3D point cloud data is captured.
[0063] In S203, the processing device extracts a time sequence image feature sequence from the video stream, and extracts a time sequence 3D point cloud feature sequence from the 3D point cloud stream.
[0064] The time sequence image feature sequence and the time sequence 3D point cloud feature sequence are feature sequences of the same target moment.
[0065] For example, the processing device extracts an image video frame of the target moment from the video stream, and extracts a 3D point cloud data of the target moment from the 3D point cloud stream. Specifically, the processing device can determine the latest two timeout moments of the preset periodic timer, and the timeout moment is the target moment, i.e., the target moment can also be two moments, and the interval is one period of time, such as setting the period of time to 1 second, 1.5 seconds, etc. Then the timer will timeout every 1 second, 1.5 seconds. The processing device can extract the image video frame closest to the timeout moment from the video stream, i.e., find the image video frame with the timestamp closest to the timeout moment, i.e., two image video frames, to retain the time sequence feature. In addition, the processing device can also extract the 3D point cloud data closest to the timeout moment from the 3D point cloud stream, i.e., find the 3D point cloud data with the timestamp closest to the timeout moment, i.e., two 3D point cloud data, to retain the time sequence feature. For convenience of description, the two image video frames will be collectively referred to as image video frames, and the two 3D point cloud data will be collectively referred to as 3D point cloud data.
[0066] The processing device can extract a time-series image feature sequence from the image video frame and a time-series 3D point cloud feature sequence from the 3D point cloud data. Specifically, the processing device can perform pixel edge extraction on the image video frame to obtain a sub-image containing only the object in the image video frame. The pixel edge extraction can be binary or grayscale processing to extract the object from the background in the image video frame. The processing device can extract 3D point cloud sub-data corresponding to the coordinate range of the sub-image in the image video frame from the 3D point cloud data. In a specific implementation, the processing device can map the 3D point cloud data to a plane space according to the position parameters of the 3D point cloud to obtain 2D point cloud data. For example, the processing device can hide the depth coordinates in the position parameters of the 3D point cloud, and obtain 2D point cloud data. Since the plane space is the plane where the image video frame is located, the size of the plane space is the same as the size of the image video frame. The so-called plane space can be understood as a two-dimensional plane formed by the shooting area. On the basis of the same shooting area of the depth camera and the visible light camera, the size of the plane space is the size of the image video frame. The coordinates of the 2D point cloud data with hidden depth coordinates in the plane space are consistent with the coordinates of the pixel points in the image video frame. On this basis, the processing device can determine 2D point sub-cloud data located in the coordinate range from the 2D point cloud data, and inversely map the 2D point sub-cloud data to the 3D space (e.g., restore the hidden depth coordinates) to obtain 3D point cloud sub-data.
[0067] The processing device extracts features of the sub-image through a lightweight convolutional neural network (e.g., MobileNetV3) and a conversion model encoder (Transformer Encoder) to obtain the time-series image feature sequence. The processing device also extracts features of the 3D point cloud sub-data through a point cloud semantic segmentation neural network (RandLA-Net) and a three-dimensional convolution-long short-term memory network (e.g., 3D Conv-LSTM) to obtain the time-series 3D point cloud feature sequence. For convenience of description, the time-series image feature sequence is referred to as an initial time-series image feature sequence, and the time-series 3D point cloud feature sequence is referred to as an initial time-series 3D point cloud feature sequence. The following describes S204 in detail.
[0068] In S204, the processing device performs aggregation processing on the time-series image feature sequence and the time-series 3D point cloud feature sequence by using a graph neural network to tend to orthogonalization, and obtains a behavior detection result of the line inspection channel at the target time.
[0069] Step 1: The processing device can align the feature sequence lengths of the initial time-series image feature sequence and the initial time-series 3D point cloud feature sequence to obtain an aligned time-series image feature sequence and an aligned time-series 3D point cloud feature sequence.
[0070] For example, the processing device can perform vector fusion on randomly selected feature vectors in the initial time sequence image feature sequence to obtain an aligned time sequence image feature sequence, and perform vector fusion on randomly selected feature vectors in the initial time sequence 3D point cloud feature sequence to obtain an aligned time sequence 3D point cloud feature sequence, the length of the aligned time sequence image feature sequence and the length of the aligned time sequence 3D point cloud feature sequence are both preset feature sequence lengths. The preset feature sequence length can be 96, 128, 256, etc. Taking 96 as an example, it means including a 96-dimensional vector, or a 1*96-dimensional vector. For ease of understanding, taking 96 as an example, for example, the initial time sequence image feature sequence is a 100-dimensional vector, 8 vectors are randomly selected therefrom, and sequentially fused, such as vector 1 and vector 2 fused into a vector, vector 3 and vector 4 fused into a vector, vector 5 and vector 6 fused into a vector, and vector 7 and vector 8 fused into a vector, to obtain a 1*96-dimensional vector.
[0071] Of course, the three-dimensional convolution-long short-term memory network and the conversion model encoder described above can also be configured to output feature sequences of a preset feature sequence length, so that the alignment process is not needed.
[0072] Step 2: The processing device performs orthogonal assignment on the aligned time sequence image feature sequence and the aligned time sequence 3D point cloud feature sequence through the orthogonal sequence template to obtain a quasi-orthogonal time sequence image feature sequence and a quasi-orthogonal time sequence 3D point cloud feature sequence.
[0073] In step 2, if the template length of the orthogonal assignment is the same as the preset feature sequence length, the following mode 1 is used, and if the template length of the orthogonal assignment is different from the preset feature sequence length, such as the preset feature sequence length being twice the template length, the following mode is used.
[0074] It should be understood that through the orthogonal assignment of the quasi-orthogonal processing, the features of the two modalities can be made as linearly independent (orthogonal) as possible in the vector space, ensuring that the information overlap is reduced (such as avoiding the use of point cloud to repeatedly describe the texture captured by the image), and the complementarity can also be enhanced (such as the point cloud supplementing the depth information of the image, and the image supplementing the semantic information of the point cloud). The following will introduce mode 1 and mode 2 in detail.
[0075] Mode 1:
[0076] The orthogonal sequence template includes a set of sequences, and any two sequences in the set of sequences are orthogonal; the processing device orthogonally assigns a first sequence to the aligned time sequence image feature sequence to obtain a quasi-orthogonal time sequence image feature sequence, and orthogonally assigns a second sequence to the aligned time sequence 3D point cloud feature sequence to obtain a quasi-orthogonal time sequence 3D point cloud feature sequence. The first sequence has the same length as the aligned time sequence image feature sequence, and the second sequence has the same length as the aligned time sequence 3D point cloud feature sequence; quasi-orthogonal refers to quasi-orthogonal between the quasi-orthogonal time sequence image feature sequence and the quasi-orthogonal time sequence 3D point cloud feature sequence, in other words, the time sequence image feature sequence and the time sequence 3D point cloud feature sequence are made to be linearly independent as much as possible.
[0077] Specifically, taking the first sequence and the second sequence as Zadoff-Chu and the length of the first sequence and the second sequence as 96 as an example.
[0078] The first sequence orthogonally assigns the aligned time sequence image feature sequence as follows:
[0079] ;
[0080] Wherein, is the i-th feature vector in the aligned time sequence image feature sequence, is the i-th vector in the first sequence, represents the i-th feature vector in the quasi-orthogonal time sequence image feature sequence; The second sequence orthogonally assigns the aligned time sequence 3D point cloud feature sequence as follows:
[0081] ;
[0082] ;
[0083] Wherein, is the i-th feature vector in the aligned time sequence 3D point cloud feature sequence, is the i-th vector in the second sequence, represents the i-th feature vector in the quasi-orthogonal time sequence 3D point cloud feature sequence. Wherein, the frequency interval K = |3-5| = 2 of the first sequence and the second sequence satisfies the orthogonal condition j, that is, K is an integer and K is equal to 0 mod 96, so as to ensure that the frequency difference guarantees a complete phase rotation period in a 96-point period.
[0084] Wherein, the frequency interval K = |3-5| = 2 of the first sequence and the second sequence satisfies the orthogonal condition j, that is, K is an integer and K is equal to 0 mod 96, so as to ensure that the frequency difference guarantees a complete phase rotation period in a 96-point period.
[0085] It can be seen that each time series 3D point cloud feature vector or each time series image feature is realized by superimposing the corresponding vector to assign value, since the first sequence and the second sequence are strictly orthogonal sequences, but the assignment to the time series image feature sequence and the time series 3D point cloud feature sequence does not guarantee that the time series image feature sequence and the time series 3D point cloud feature sequence are orthogonal, but can make the two tend to be orthogonal as much as possible, reducing linear independence.
[0086] Mode 2:
[0087] The orthogonal sequence template includes a first sequence group and a second sequence group, any two sequences in the first sequence group or the second sequence group are orthogonal, and any sequence in the first sequence group is quasi-orthogonal to any sequence in the second sequence group. It should be understood that quasi-orthogonal means that the modulus of the inner product of the sequences in the group is as small as possible, such as less than 0.2.
[0088] Therefore, the processing device divides the aligned time series image feature sequence into a first time series image feature subsequence and a second time series image feature subsequence, and divides the aligned time series 3D point cloud feature sequence into a first time series 3D point cloud feature subsequence and a second time series 3D point cloud feature subsequence. In addition, the processing device can use a first sequence in the first sequence group to orthogonally assign the first time series image feature subsequence to obtain a quasi-orthogonal first time series image feature subsequence, and use a third sequence in the second sequence group to orthogonally assign the second time series image feature subsequence to obtain a quasi-orthogonal second time series image feature subsequence, and splice the quasi-orthogonal first time series image feature subsequence and the quasi-orthogonal first time series image feature subsequence to obtain a quasi-orthogonal time series image feature sequence. In addition, the processing device can also use a second sequence in the first sequence group to orthogonally assign the first time series 3D point cloud feature subsequence to obtain a quasi-orthogonal first time series 3D point cloud feature subsequence, and use a fourth sequence in the second sequence group to orthogonally assign the second time series 3D point cloud feature subsequence to obtain a quasi-orthogonal second time series 3D point cloud feature subsequence, and splice the quasi-orthogonal first time series 3D point cloud feature subsequence and the quasi-orthogonal first time series 3D point cloud feature subsequence to obtain a quasi-orthogonal time series 3D point cloud feature sequence.
[0089] Specifically, the lengths of the first sequence, the second sequence, the third sequence, and the fourth sequence are all 48, the first sequence and the second sequence are even-indexed Zadoff-Chu sequences, and the third sequence and the fourth sequence are odd-indexed Zadoff-Chu sequences; taking the lengths of the aligned time series image feature sequence and the aligned time series 3D point cloud feature sequence as 96 as an example.
[0090] The first sequence orthogonally assigns the first time series image feature subsequence as follows:
[0091] ;
[0092] in, The first time-series image feature subsequence 1 eigenvector For the first sequence A vector, The first time-series image feature subsequence that is orthogonal represents the first time-series image feature subsequence. 1 eigenvector;
[0093] The second sequence is orthogonally assigned to the first temporal 3D point cloud feature subsequence as follows:
[0094] ;
[0095] in, The first time series 3D point cloud feature sequence is the first time series 3D point cloud feature sequence. 1 eigenvector For the second sequence A vector, The first temporal 3D point cloud feature subsequence that represents the orthogonal subsequence is the first... 1 eigenvector;
[0096] The third sequence is orthogonally assigned to the feature subsequence of the second time-series image as follows:
[0097] ;
[0098] in, The first in the second time-series image feature subsequence 1 eigenvector For the third sequence A vector, The second time-series image feature subsequence that represents orthogonality 1 eigenvector;
[0099] The second sequence is orthogonally assigned to the first temporal 3D point cloud feature subsequence as follows:
[0100] ;
[0101] in, The first in the second temporal 3D point cloud feature sequence 1 eigenvector For the fourth sequence A vector, The second temporal 3D point cloud feature subsequence represents the orthogonal subsequence of the first time series. 1 eigenvector.
[0102] Wherein, the inner product of the first sequence and the third sequence, or the first sequence and the fourth sequence, or the second sequence and the third sequence, or the second sequence and the fourth sequence is 0.1444, less than 0.2, which meets the inter-group quasi-orthogonal.
[0103] It should be understood that the first and second time sequence image feature subsequences are assigned with the inter-group quasi-orthogonal sequence to retain the original correlation between them, and the same applies to the first and second time sequence 3D point cloud feature subsequences, which will not be described here.
[0104] In step 3, the processing device regards the quasi-orthogonal time sequence image feature sequence as a graph node and the quasi-orthogonal time sequence 3D point cloud feature sequence as an edge, performs aggregation processing through a graph neural network (such as GNN), and obtains a detection result. For example, the processing device labels the quasi-orthogonal time sequence image feature sequence as an edge and the quasi-orthogonal time sequence 3D point cloud feature sequence as an edge, and then inputs them into the GNN, which performs feature recognition processing after aggregation. The GNN can then output a detection result. At this time, due to the weakening of the linear correlation between the quasi-orthogonal time sequence 3D point cloud feature sequence and the quasi-orthogonal time sequence image feature sequence, the aggregation processing of the GNN can be more robust, and the final detection result can be more accurate and stable. The detection result can indicate the behavior of the crowd and identify whether there is abnormal behavior, such as fighting, falling, not walking on the designated route, etc.
[0105] In summary, based on the video stream of the line inspection channel collected by the visible light camera and the 3D point cloud stream of the line inspection channel collected by the depth camera, the processing device can extract the time sequence image feature sequence and the time sequence 3D point cloud feature sequence located at the same target time from them, respectively. Therefore, the processing device can perform aggregation processing on the time sequence image feature sequence and the time sequence 3D point cloud feature sequence through a graph neural network to tend to be orthogonal, reduce information overlap, and enhance complementarity, so that the final behavior detection result of the line inspection channel at the target time is more stable, i.e., has higher detection failure stability and reliability.
[0106] Figure 3 A structural schematic diagram of a processing device provided by an embodiment of the present application is shown. The processing device can be a terminal device, a chip (system) or other components or assemblies that can be provided in the terminal device. As shown in the figure, the processing device 400 can include a processor 401. Optionally, the processing device 400 can also include a memory 402 and / or a transceiver 403. The processor 401 is coupled with the memory 402 and the transceiver 403, which can be connected through a communication bus. In addition, the processing device 400 can also be a chip, which includes the processor 401. At this time, the transceiver can be an input / output interface of the chip. Figure 3
[0107] The following describes the method in detail. Figure 3 The various components of the processing device 400 are described in detail as follows.
[0108] The processor 401 is the control center of the processing device 400, which can be one processor or a plurality of processing elements. For example, the processor 401 is one or more central processing units (CPUs), application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of the present application, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).
[0109] Optionally, the processor 401 can perform various functions of the processing device 400 by running or executing software programs stored in the memory 402 and calling data stored in the memory 402, such as executing the above-described method. Figure 2
[0110] In a specific implementation, as an embodiment, the processor 401 can include one or more CPUs, such as the CPU0 and CPU1 shown in FIG. 1. Figure 3
[0111] In a specific implementation, as an embodiment, the processing device 400 can also include a plurality of processors. Each of the processors can be a single-CPU or a multi-CPU. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer programs or instructions).
[0112] The memory 402 is used to store software programs for implementing the schemes of the present application, and is controlled by the processor 401 to perform, and the specific implementation manner can refer to the above method embodiments, which will not be described here.
[0113] Optionally, the memory 402 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 402 can be integrated with the processor 401 or exist independently and be coupled to the processor 401 through an interface circuit (not shown in the figure) of the processing device 400, and embodiments of the present application do not make a specific limitation hereon. Figure 3
[0114] The transceiver 403 is configured to communicate with other processing devices. For example, the processing device 400 is a terminal device, and the transceiver 403 can be configured to communicate with a network device or another terminal device. For another example, the processing device 400 is a network device, and the transceiver 403 can be configured to communicate with a terminal device or another network device.
[0115] Optionally, the transceiver 403 can include a receiver and a transmitter (not shown in the figure) separately. The receiver is configured to implement the receiving function, and the transmitter is configured to implement the transmitting function. Figure 3
[0116] Optionally, the transceiver 403 can be integrated with the processor 401 or exist independently and be coupled to the processor 401 through an interface circuit (not shown in the figure) of the processing device 400, and embodiments of the present application do not make a specific limitation hereon. Figure 3
[0117] It can be understood that the structure of the processing device 400 shown in the figure does not constitute a limitation on the processing device, and an actual processing device can include more or fewer components than those shown in the figure, or combine certain components, or different component arrangements. Figure 3
[0118] In addition, the technical effects of the processing device 400 can refer to the technical effects of the methods described in the above method embodiments, which will not be described here again.
[0119] It should be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0120] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, but not by way of limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (direct rambus RAM, DR RAM).
[0121] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0122] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects, but can also represent an "and / or" relationship, which can be understood in the context before and after.
[0123] In this application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0124] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0125] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0127] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0128] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0129] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0130] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0131] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A line inspection channel crowd behavior recognition method based on multi-modal data fusion, characterized in that, The method is applied to a processing device, and the method comprises: The processing device acquires a video stream of a line inspection channel collected by a visible light camera; The processing device acquires a 3D point cloud stream of the line inspection channel collected by a depth camera; The processing device extracts a time sequence image feature sequence from the video stream and a time sequence 3D point cloud feature sequence from the 3D point cloud stream, wherein the time sequence image feature sequence and the time sequence 3D point cloud feature sequence are feature sequences of the same target moment; The processing device performs an aggregation process tending to orthogonalization on the time sequence image feature sequence and the time sequence 3D point cloud feature sequence by using a graph neural network to obtain a behavior detection result of the line inspection channel at the target moment; The time sequence image feature sequence is an initial time sequence image feature sequence, and the time sequence 3D point cloud feature sequence is an initial time sequence 3D point cloud feature sequence; the processing device performs an aggregation process tending to orthogonalization on the time sequence image feature sequence and the time sequence 3D point cloud feature sequence by using a graph neural network to obtain a behavior detection result of the line inspection channel at the target moment, which comprises: The processing device aligns the initial time sequence image feature sequence and the initial time sequence 3D point cloud feature sequence in terms of feature sequence length to obtain an aligned time sequence image feature sequence and an aligned time sequence 3D point cloud feature sequence; The processing device orthogonally assigns the aligned time sequence image feature sequence and the aligned time sequence 3D point cloud feature sequence by using an orthogonal sequence template to obtain a time sequence image feature sequence tending to orthogonality and a time sequence 3D point cloud feature sequence tending to orthogonality; The processing device regards the time sequence image feature sequence tending to orthogonality as a graph node and the time sequence 3D point cloud feature sequence tending to orthogonality as an edge, and performs an aggregation process by using a graph neural network to obtain the detection result; The orthogonal sequence template comprises a group of sequences, and any two sequences in the group of sequences are orthogonal; the processing device orthogonally assigns the aligned time sequence image feature sequence and the aligned time sequence 3D point cloud feature sequence by using an orthogonal sequence template to obtain a time sequence image feature sequence tending to orthogonality and a time sequence 3D point cloud feature sequence tending to orthogonality, which comprises: The processing device selects a first sequence from the group of sequences to orthogonally assign the aligned time sequence image feature sequence to obtain the time sequence image feature sequence tending to orthogonality, and selects a second sequence from the group of sequences to orthogonally assign the aligned time sequence 3D point cloud feature sequence to obtain the time sequence 3D point cloud feature sequence tending to orthogonality; wherein the first sequence has the same length as the aligned time sequence image feature sequence, and the second sequence has the same length as the aligned time sequence 3D point cloud feature sequence; tending to orthogonality means that the time sequence image feature sequence tending to orthogonality and the time sequence 3D point cloud feature sequence tending to orthogonality tend to be orthogonal; the lengths of the first sequence and the second sequence are 96, orthogonally assigning the aligned time sequence image feature sequence by using the first sequence is represented as follows: wherein Frbg i is the i-th feature vector in the aligned sequence of time-series image features, is the i-th vector in the first sequence, Frbg i ′ denotes the i-th feature vector in the sequence of time-series image features trending towards orthogonality; The second sequence orthogonally assigns the aligned time sequence 3D point cloud feature sequence, and the orthogonally assigned time sequence 3D point cloud feature sequence is represented as follows: wherein Fpts i is the i-th feature vector in the aligned time-sequential 3D point cloud feature sequence, is the i-th vector in the second sequence, Fpts i ′ denotes the i-th feature vector in the quasi-orthogonal time-sequential 3D point cloud feature sequence.
2. The method of claim 1, wherein, The orthogonal sequence template includes a first sequence group and a second sequence group, any two sequences in the first sequence group or the second sequence group are orthogonal, and any sequence in the first sequence group is quasi-orthogonal to any sequence in the second sequence group; the processing device orthogonally assigns the aligned time sequence image feature sequence and the aligned time sequence 3D point cloud feature sequence by using the orthogonal sequence template to obtain a quasi-orthogonal time sequence image feature sequence and a quasi-orthogonal time sequence 3D point cloud feature sequence, including: The processing device divides the aligned time sequence image feature sequence into a first time sequence image feature subsequence and a second time sequence image feature subsequence, and divides the aligned time sequence 3D point cloud feature sequence into a first time sequence 3D point cloud feature subsequence and a second time sequence 3D point cloud feature subsequence; The processing device orthogonally assigns the first time sequence image feature subsequence by using a first sequence in the first sequence group to obtain a quasi-orthogonal first time sequence image feature subsequence, orthogonally assigns the second time sequence image feature subsequence by using a third sequence in the second sequence group to obtain a quasi-orthogonal second time sequence image feature subsequence, and splices the quasi-orthogonal first time sequence image feature subsequence and the quasi-orthogonal first time sequence image feature subsequence to obtain the quasi-orthogonal time sequence image feature sequence; The processing device orthogonally assigns the first time sequence 3D point cloud feature subsequence by using a second sequence in the first sequence group to obtain a quasi-orthogonal first time sequence 3D point cloud feature subsequence, orthogonally assigns the second time sequence 3D point cloud feature subsequence by using a fourth sequence in the second sequence group to obtain a quasi-orthogonal second time sequence 3D point cloud feature subsequence, and splices the quasi-orthogonal first time sequence 3D point cloud feature subsequence and the quasi-orthogonal first time sequence 3D point cloud feature subsequence to obtain the quasi-orthogonal time sequence 3D point cloud feature sequence.
3. The method of claim 2, wherein, The lengths of the first sequence, the second sequence, the third sequence, and the fourth sequence are all 48, the first sequence and the second sequence are even-indexed Zadoff-Chu sequences, and the third sequence and the fourth sequence are odd-indexed Zadoff-Chu sequences; the lengths of the aligned time sequence image feature sequence and the aligned time sequence 3D point cloud feature sequence are both 96; The first sequence orthogonally assigns the first time sequence image feature subsequence and is represented as follows: wherein F1rbg i is the i-th feature vector in the first time-sequential image feature sub-sequence, is the i-th vector in the first sequence, F1rbg i ′ denotes the i-th feature vector in the first time-sequential image feature sub-sequence that tends to be orthogonal. The second sequence orthogonally assigns the first time sequence 3D point cloud feature subsequence and is represented as follows: wherein F1pts i is the i-th feature vector in the first time-sequential 3D point cloud feature sequence, is the i-th vector in the second sequence, F1pts i ′ denotes the i-th feature vector in the first time-sequential 3D point cloud feature sub-sequence that tends to be orthogonal. The third sequence orthogonally assigns the second time sequence image feature subsequence and is represented as follows: wherein F2rbgis the i-th feature vector in the second time-sequential image feature subsequence, i is the i-th vector in the third sequence, F2rbgis the i-th feature vector in the second time-sequential image feature subsequence, is the i-th vector in the third sequence, F2rbgis the i-th feature vector in the second time-sequential image feature subsequence, i ′ denotes the i-th feature vector in the second time-sequential image feature subsequence that tends to be orthogonal. The second sequence orthogonally assigns the first time sequence 3D point cloud feature subsequence and is represented as follows: wherein F2pts i is the i-th feature vector in the second time-sequential 3D point cloud feature sequence, is the i-th vector in the fourth sequence, F2pts i ′ denotes the i-th feature vector in the second time-sequential 3D point cloud feature sub-sequence that tends to be orthogonal.
4. The method according to any one of claims 1 to 3, characterized in that, The processing device extracts a time sequence image feature sequence from the video stream and a time sequence 3D point cloud feature sequence from the 3D point cloud stream, including: The processing device extracts an image video frame of the target moment from the video stream and extracts 3D point cloud data of the target moment from the 3D point cloud stream; The processing device extracts the time sequence image feature sequence from the image video frame and extracts the time sequence 3D point cloud feature sequence from the 3D point cloud data.
5. The method of claim 4, wherein, The shooting directions of the visible light camera and the depth camera are the same, both being along the longitudinal direction of the line inspection channel; thus, the processing device extracts the time sequence image feature sequence from the image video frame and extracts the time sequence 3D point cloud feature sequence from the 3D point cloud data, including: The processing device performs pixel edge extraction on the image video frame to obtain a sub-image containing only objects in the image video frame; The processing device extracts 3D point cloud sub-data corresponding to the coordinate range from the 3D point cloud data according to the coordinate range of the sub-image in the image video frame; The processing device extracts features of the sub-image by a lightweight convolutional neural network and a conversion model encoder to obtain the time sequence image feature sequence; The processing device extracts features of the 3D point cloud sub-data by a point cloud semantic segmentation neural network and a three-dimensional convolution-long short-term memory network to obtain the time sequence 3D point cloud feature sequence.
6. The method of claim 5, wherein, The 3D point cloud data includes position parameters of 3D point clouds; the processing device extracts 3D point cloud sub-data corresponding to the coordinate range from the 3D point cloud data according to the coordinate range of the sub-image in the image video frame, including: The processing device maps the 3D point cloud data to a plane space according to the position parameters of the 3D point clouds to obtain 2D point cloud data; the plane space is a plane on which the image video frame is located, and the size of the plane space is the same as that of the image video frame; The processing device determines 2D point sub-cloud data located in the coordinate range from the 2D point cloud data; The processing device inversely maps the 2D point sub-cloud data to a 3D space to obtain the 3D point cloud sub-data.
Citation Information
Patent Citations
Communication radiation source cross-mode identification method based on multi-mode information fusion
CN115952466A
Intelligent escalator passenger behavior detection system and method based on multi-mode perception
CN120217300A
Autism symptom identification system based on multivariate emotion path analysis
CN120345898A