Image-signal bimodal fusion truck scale abnormal behavior detection method and system

Through the detection method of image-signal dual-modal fusion, the CBAM attention mechanism and Transformer encoding module are used to extract and fuse video images and sensor signals, which solves the problem of low recognition accuracy in the prior art and achieves higher accuracy and robustness in abnormal behavior recognition.

CN120180193APending Publication Date: 2025-06-20JINAN JINZHONG ELECTRONICS SCALE +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510335132.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing abnormal behavior recognition methods rely on a single data source, making it difficult to fully mine the correlation information between video images and sensor signals, resulting in low recognition accuracy, especially in complex scenarios that are prone to misjudgment or misjudgment.

Method used

Using the detection method of image-signal dual-mode fusion, the CBAM attention mechanism and Transformer encoding module are introduced to feature extraction of optical flow diagrams and weighing signals after superimposing single-frame images and multi-frame images, and feature-level fusion is performed to perform behavior classification.

Benefits of technology

It significantly improves the accuracy of abnormal behavior recognition, reduces the misidentification of pseudo-abnormal actions, and enhances the adaptability and robustness of the system in complex industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180193A_ABST
    Figure CN120180193A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of abnormal behavior recognition. The invention provides an image-signal dual-mode fusion motor truck scale abnormal behavior detection method and system. The method comprises the following steps: extracting a single-frame image, an optical flow diagram formed by superposing multiple frames of images and a corresponding weighing signal from a monitoring video; obtaining image features of the single-frame image according to the plurality of sequentially connected convolution blocks embedded with the CBAM attention mechanism; according to the plurality of Mamba modules connected in sequence, obtaining optical flow characteristics of the optical flow diagram; carrying out self-attention processing and feedforward neural network processing on the weighing signal in sequence to obtain a weighing feature; carrying out feature level fusion on the image features, the optical flow features and the weighing features to obtain fusion features, and carrying out behavior classification according to the fusion features to obtain a behavior classification result; according to the method, the CBAM attention mechanism and the Transform coding module are introduced, so that the accuracy of behavior recognition is effectively improved, and misrecognition of pseudo abnormal actions is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal behavior recognition, and particularly to a method and system for detecting abnormal behaviors of a weighbridge by image-signal dual-modal fusion. Background Art

[0002] The statements in this part only provide background technologies related to the present invention and do not necessarily constitute prior arts.

[0003] At the operation site of a weighbridge (weighing scale), there are usually a large number of personnel and vehicle activities, with a complex scene and high safety risks. If abnormal behaviors occur during the operation of the staff (such as illegal operations, damage to the scale body, interference with weighing signals, etc.), it may not only lead to inaccurate weighing results but also cause safety accidents.

[0004] The existing abnormal behavior recognition solutions have the following problems: (1) The existing abnormal behavior recognition methods often rely on a single data source (such as a single-frame image or a single sensor signal), and cannot fully mine the correlation information between video images and sensor signals. The single-modal processing method is difficult to handle the changing situations in complex scenes, resulting in low accuracy in recognizing abnormal behaviors, especially prone to misjudgment or missed judgment in complex scenes; (2) When traditional algorithms process video sequences, they fail to effectively utilize the spatio-temporal dynamic changes in optical flow maps or temporal data, resulting in weak recognition ability for fast-moving or occluded behaviors and difficulty in accurately capturing the subtle changes during the movement process; (3) Traditional convolutional neural networks usually lack adaptive attention to important regions in image recognition, resulting in the network being easily interfered by background noise in complex scenes and unable to effectively focus on the key information in the image, thereby affecting the recognition effect of abnormal behaviors. Summary of the Invention

[0005] To solve the deficiencies of the existing technologies, the present invention provides a method and system for detecting abnormal behaviors of a weighbridge by image-signal dual-modal fusion. By introducing the CBAM attention mechanism and the Transformer encoding module, feature extraction is respectively performed on single-frame images, optical flow maps after superposition of multiple frames of images, and corresponding weighing signals, effectively improving the accuracy of behavior recognition and reducing the misrecognition of pseudo-abnormal actions.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a method for detecting abnormal behaviors of a weighbridge by image-signal dual-modal fusion.

[0008] A method for detecting abnormal behaviors of a weighbridge by image-signal dual-modal fusion includes the following processes:

[0009] Extract single-frame images, optical flow maps after superimposing multiple frames of images, and corresponding weighing signals from the surveillance video;

[0010] Obtain image features of the single-frame image according to multiple successively connected convolutional blocks embedded with the CBAM attention mechanism;

[0011] Obtain optical flow features of the optical flow map according to multiple successively connected Mamba modules;

[0012] After successively processing the weighing signal through self-attention processing and a feed-forward neural network, obtain weighing features;

[0013] Perform feature-level fusion on the image features, the optical flow features, and the weighing features to obtain fusion features, and perform behavior classification according to the fusion features to obtain a behavior classification result.

[0014] As a further limitation of the first aspect of the present invention, the CBAM attention mechanism includes: a channel attention module and a spatial attention module;

[0015] In the channel attention module, according to the initial operation result of the convolutional block, obtain channel information through first global average pooling and first global max pooling; according to the channel information, generate channel weights through a shared fully connected layer; generate and enhance channel attention weights according to the channel weights;

[0016] In the spatial attention module, perform second global average pooling and second global max pooling on the enhanced result of the channel attention weight, splice the results of the second global average pooling and the second global max pooling and then generate spatial weights through convolution, and enhance the spatial weights to obtain the output of the CBAM attention mechanism.

[0017] As a further limitation of the first aspect of the present invention, obtain channel information through first global average pooling and first global max pooling: F avg =GAP(F conv ), F max =GMP(F conv );

[0018] Generate channel weights through a shared fully connected layer: F MLP =W2·ReLU(W1·F avg )+W2·ReLU(W1·F max );

[0019] Generate and enhance channel attention weights: M channel =Sigmoid(F MLP ), F channel =F conv ·M c hanne l;

[0020] Perform second global average pooling and second global max pooling on the channel dimension:

[0021] Generate spatial weights through convolution after concatenation:

[0022] Spatial weight enhancement: F spatial = F channel · M spatial ;

[0023] The output feature of the CBAM attention mechanism: F out = F spatial .

[0024] As a further limitation of the first aspect of the present invention, the input of the first convolutional block is the single-frame image, and the input of the subsequent convolutional blocks is the output of the previous convolutional block.

[0025] As a further limitation of the first aspect of the present invention, the Mamba module includes: performing an initial linear transformation on the input, performing depth convolution on the result of the initial linear transformation, performing matrix multiplication on the result of the depth convolution, and performing residual connection on the result of the matrix multiplication and the input to obtain the output of the Mamba module.

[0026] As a further limitation of the first aspect of the present invention, the weighing signal is sequentially subjected to self-attention processing, including:

[0027] Performing feed-forward neural network processing, including: F S = FFN(Attention(Q, K, V)), where Q is the query vector, K is the key vector, and V is the value vector.

[0028] In a second aspect, the present invention provides a vehicle scale abnormal behavior detection system for image-signal bimodal fusion.

[0029] A vehicle scale abnormal behavior detection system for image-signal bimodal fusion, including:

[0030] A data acquisition unit configured to: extract a single-frame image, an optical flow map after superposition of multiple frames of images, and a corresponding weighing signal from a monitoring video;

[0031] An image feature extraction unit configured to: obtain the image feature of the single-frame image according to a plurality of sequentially connected convolutional blocks embedded with the CBAM attention mechanism;

[0032] An optical flow feature extraction unit, configured to: obtain the optical flow features of an optical flow map according to a plurality of sequentially connected Mamba modules;

[0033] A weighing feature extraction unit, configured to: obtain weighing features after sequentially processing the weighing signals through self-attention processing and a feed-forward neural network;

[0034] An abnormal behavior classification unit, configured to: perform feature-level fusion on the image features, the optical flow features, and the weighing features to obtain fused features, and perform behavior classification according to the fused features to obtain a behavior classification result.

[0035] In a third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium;

[0036] The processor is adapted to execute a computer program;

[0037] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method for detecting abnormal behaviors of a weighbridge with image-signal dual-modal fusion as described in the first aspect of the present invention.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the method for detecting abnormal behaviors of a weighbridge with image-signal dual-modal fusion as described in the first aspect of the present invention.

[0039] In a fifth aspect, the present invention provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method for detecting abnormal behaviors of a weighbridge with image-signal dual-modal fusion as described in the first aspect of the present invention.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] 1. The present invention innovatively proposes a method for detecting abnormal behaviors of a weighbridge with image-signal dual-modal fusion. By introducing the CBAM attention mechanism and the Transformer encoding module, feature extraction is respectively performed on a single-frame image, an optical flow map after superimposing multiple frames of images, and the corresponding weighing signals, effectively improving the accuracy of behavior recognition and reducing the misrecognition of pseudo-abnormal actions.

[0042] 2. The present invention combines video image data with real-time sensor signals (such as weighing signals) for abnormal behavior recognition, and innovatively fuses the features of two heterogeneous data sources. Traditional behavior recognition methods mostly rely on single visual information or signal data, while the present invention, by fusing these two types of data, conducts joint modeling in a multi-modal feature space, significantly improving the recognition accuracy and robustness of the system for abnormal behavior. This fusion can effectively make up for the deficiencies of a single modality in complex backgrounds or occlusion situations, and at the same time enhance the adaptability in complex industrial scenarios.

[0043] 3. The present invention introduces the CBAM attention mechanism into the convolutional neural network (CNN) to enhance the attention to key regions of images. Compared with traditional convolutional networks, CBAM adaptively adjusts the weights of image features through channel and spatial attention mechanisms, enabling the network to dynamically focus on important information in the image and suppressing the interference of irrelevant backgrounds. Especially in complex scenarios, it can better identify subtle abnormal behaviors.

[0044] 4. The present invention uses the Mamba module to process optical flow features, which is an innovative extension of traditional deep learning feature extraction methods. The Mamba module enhances the expression ability of the model in processing spatio-temporal dynamic information in video sequences through technical means such as multi-stage depthwise separable convolution (DW Conv) and matrix multiplication (SS2D). Especially for behavior scenarios with fast motion changes or occlusions, the Mamba module can capture more fine-grained dynamic changes. This innovative design enables the system to have a more comprehensive understanding of the temporal continuity and spatial distribution of workers' actions.

[0045] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become apparent from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The schematic diagrams in the specification forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0047] Figure 1 Schematic flowchart of the method for detecting abnormal behaviors of a weighbridge with image-signal bimodal fusion provided in Embodiment 1 of the present invention;

[0048] Figure 2 Schematic principle diagram of the method for detecting abnormal behaviors of a weighbridge with image-signal bimodal fusion provided in Embodiment 1 of the present invention;

[0049] Figure 3 Schematic diagram of the system for detecting abnormal behaviors of a weighbridge with image-signal bimodal fusion provided in Embodiment 2 of the present invention;

[0050] Figure 4 A schematic diagram of the computer device provided in Embodiment 3 of the present invention. Specific implementation manners

[0051] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0052] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0053] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0054] Embodiment 1:

[0055] As described in the background art, due to the lack of in-depth understanding of specific working scenarios and staff behaviors, existing algorithms are prone to misjudge normal operations or specific habits of staff as abnormal behaviors, resulting in unnecessary false alarms. Especially for staff behaviors under complex environmental conditions, some "pseudo-abnormal actions" made by staff may be judged as abnormal behaviors (here, pseudo-abnormal actions include: staff simply touching the weighbridge, or staff passing by the weighbridge with destructive tools in hand). Moreover, single visual information is limited in recognition effect due to reasons such as scene occlusion and complex background, and cannot accurately detect abnormal behaviors. Therefore, how to more accurately identify abnormal actions is a technical problem that needs to be solved urgently at present. In view of this, this implementation proposes a method for detecting abnormal behaviors of weighbridges through image-signal bimodal fusion. By introducing the CBAM attention mechanism and the Transformer encoding module, feature extraction is performed on single-frame images, the optical flow map after superimposing multiple frames of images, and the corresponding weighing signals respectively, effectively improving the accuracy of behavior recognition and reducing the misrecognition of pseudo-abnormal actions.

[0056] The following first briefly introduces the technical terms and related concepts involved in this processing solution, where:

[0057] The Mamba module is an efficient and innovative selective structure state space model. It uses a selective state space to quickly process long sequences, combines multiple modes, and supports high efficiency in resolution and practicality. The Mamba module has a linear time complexity, which makes it more efficient than the Transformer when processing long sequence data;

[0058] Optical Flow refers to the instantaneous velocity field of pixel motion in an image, that is, the visual manifestation of the motion of objects in the image. It describes the position change of pixel points in the image from one moment to another. The optical flow field is a two-dimensional vector field, and each vector represents the motion direction and speed of each pixel point in the image. Optical flow calculation algorithms include traditional optical flow algorithms (such as Lucas-Kanade algorithm, Horn-Schunck algorithm, Farneback algorithm, etc.) and deep learning-based optical flow algorithms (such as FlowNet, PWC-Net, RAFT, etc.). When choosing an optical flow calculation algorithm, factors such as application scenarios, computing resources, and real-time requirements need to be considered. For example, in real-time video processing, an algorithm with higher computing efficiency may need to be selected; while in scenarios that require high-precision optical flow estimation, a deep learning-based algorithm may need to be selected;

[0059] Linear transformation is a basic concept in linear algebra, which describes the linear mapping from one vector space to another. In the Mamba module, linear transformation is mainly used to process the input data for subsequent calculation of the state space model;

[0060] Global average pooling is a special pooling operation that performs average pooling on the entire feature map to obtain a global feature description. Specifically, for each channel's feature map, global average pooling calculates the average value of all pixels on the feature map and then uses this average value as the feature representation of this channel. In this way, each channel will obtain a corresponding feature value, and these feature values together constitute the feature vector after global average pooling;

[0061] Global max pooling is a special pooling operation that involves finding the maximum value in each channel of the entire feature map (Feature Map) and using this maximum value as the feature representation of this channel. Specifically, for the input feature map, the global max pooling layer will traverse each channel and select the value with the largest response from it. Finally, a one-dimensional vector with the same number of channels as the input is generated, and each element of this vector represents the maximum activation value of the corresponding channel;

[0062] The CBAM (Convolutional Block Attention Module) attention mechanism is a module that combines channel attention and spatial attention, aiming to enhance the feature representation ability of convolutional neural networks (CNNs); the CBAM attention mechanism can be applied to various CNN-based computer vision tasks, such as image classification, object detection, semantic segmentation, etc. By introducing CBAM, the model performance in these tasks can usually be significantly improved. Specifically, CBAM can help the model capture important features and detailed information in the input image more accurately, thereby improving the recognition accuracy and robustness of the model;

[0063] Transformer is a deep learning model architecture based on the attention mechanism, mainly used for sequence-to-sequence learning in natural language processing (NLP) tasks. The Transformer model mainly consists of an input part, multiple layers of encoders, multiple layers of decoders, and an output part; the core of Transformer is the self-attention mechanism, which allows the model to simultaneously focus on information from different positions when processing the input sequence. The self-attention mechanism generates the output vector of the current position by calculating the similarity or attention weights between the current position and other positions, and then weighted summing the vectors of other positions according to these weights.

[0064] The above method for detecting abnormal behaviors of weighbridges by image-signal bimodal fusion, such as Figure 1 and Figure 2 shown, includes the following processes:

[0065] S1: Input data.

[0066] Image data: Extract a single-frame image X and a multi-frame superimposed optical flow map O from the surveillance video;

[0067] Signal data: The weighing signal S collected in real time, expressed as time-series data.

[0068] S2: Image feature extraction.

[0069] The image data X is input into the convolutional module to extract features, and each convolutional block embeds the CBAM attention mechanism. The specific process is as follows:

[0070] S2.1: Convolution operation.

[0071] Initial operation of each convolutional block:

[0072] F conv = ReLU(BatchNorm(Conv(F in ))) (1);

[0073] S2.2: CBAM Attention Mechanism.

[0074] The CBAM attention mechanism includes a channel attention module and a spatial attention module. In this invention, the CBAM attention mechanism is introduced into the convolutional neural network (CNN) to enhance the attention to key regions of images. Compared with traditional convolutional networks, CBAM adaptively adjusts the weights of image features through channel and spatial attention mechanisms, enabling the network to dynamically focus on important information in the image and suppressing the interference of irrelevant backgrounds. Especially in complex scenarios, it can better identify subtle abnormal behaviors.

[0075] In this implementation, the channel attention module, specifically, includes:

[0076] Obtain channel information through global average pooling (GAP) and global max pooling (GMP):

[0077] F avg = GAP(F conv ), F max = GMP(F conv ) (2);

[0078] Generate channel weights through a shared fully connected layer:

[0079] F MLP = W2·ReLU(W1·F avg ) + W2·ReLU(W1·F max ) (3);

[0080] Generate and enhance channel attention weights:

[0081] M channel = Sigmoid(F MLP ), F channel = F conv ·M channel (4);

[0082] The spatial attention module, specifically, includes:

[0083] Perform global average pooling and max pooling on the channel dimension:

[0084]

[0085] Generate spatial weights through convolution after concatenation:

[0086]

[0087] Spatial weight enhancement:

[0088] F spatial = F channel · M spatial (7);

[0089] Final output feature:

[0090] F out = F spatial (8);

[0091] After passing through multiple convolutional blocks, the image feature is obtained:

[0092] F X = CBAMBlock n (…CBAMBlock1(X)) (9);

[0093] S3: Optical flow feature extraction (implemented by the Mamba module).

[0094] The optical flow map O is input into multiple Mamba modules. The present invention uses the Mamba module to process optical flow features, which is an innovative extension of traditional deep learning feature extraction methods. The Mamba module enhances the model's expressive ability in processing spatio-temporal dynamic information in video sequences through technical means such as multi-stage depthwise separable convolution (DW Conv) and matrix multiplication (SS2D). Especially for behavior scenarios with fast motion changes or occlusions, the Mamba module can capture more fine-grained dynamic changes. This innovative design enables the system to have a more comprehensive understanding of the temporal continuity and spatial distribution of workers' actions.

[0095] In this implementation, the calculation of each Mamba module is as follows:

[0096] S3.1: Initial linear transformation:

[0097] F linearl = W1F in + b1 (10);

[0098] S3.2: Depthwise convolution (DW Conv):

[0099] F dwconv = DWConv(F linearl ) (11);

[0100] S3.3: Matrix multiplication (SS2D):

[0101] F ss2d = LayerNorm(SS2D(F dwconv )) (12);

[0102] S3.4: Residual connection:

[0103] F out = V ss2d + F in (13);

[0104] After passing through multiple Mamba modules, the optical flow features are obtained:

[0105] F O = MambaBlock n (…MambaBlock1(O)) (14);

[0106] S4: Signal feature encoding.

[0107] The signal data S is input into the Transformer encoding module:

[0108] S4.1: Self-attention mechanism:

[0109]

[0110] S4.2: Feed-forward neural network (FFN):

[0111] F S = FFN(Attention(Q, K, V)) (16);

[0112] S5: Feature fusion and classification.

[0113] The image feature F X , the optical flow feature F O and the signal feature F S are subjected to feature-level fusion:

[0114] F fusion = FC(Concat(F X , F O , F S )) (17);

[0115] The classification layer outputs the behavior category:

[0116] Y = Softmax(W cls · F fusion + b cls ) (18).

[0117] The present invention innovatively fuses the characteristics of two heterogeneous data sources, namely video image data and real-time sensor signals (such as weighing signals), conducts joint modeling in a multi-modal feature space, significantly improves the recognition accuracy and robustness of the system for abnormal behaviors. This fusion can effectively make up for the deficiencies of a single modality in complex backgrounds or occlusion situations, and at the same time enhance the adaptability in complex industrial scenarios.

[0118] Embodiment 2:

[0119] As Figure 3 shown, this implementation provides a vehicle scale abnormal behavior detection system with image-signal dual-modal fusion, including:

[0120] A data acquisition unit, configured to: extract a single-frame image, an optical flow map after superimposing multiple frames of images, and the corresponding weighing signal from a monitoring video;

[0121] An image feature extraction unit, configured to: obtain the image features of a single-frame image according to a plurality of successively connected convolutional blocks embedded with a CBAM attention mechanism;

[0122] An optical flow feature extraction unit, configured to: obtain the optical flow features of the optical flow map according to a plurality of successively connected Mamba modules;

[0123] A weighing feature extraction unit, configured to: after sequentially processing the weighing signal through self-attention processing and a feed-forward neural network, obtain weighing features;

[0124] An abnormal behavior classification unit, configured to: perform feature-level fusion on the image features, the optical flow features, and the weighing features to obtain fusion features, and perform behavior classification according to the fusion features to obtain a behavior classification result.

[0125] The specific working processes of the above-mentioned various units can be seen in the introduction in Embodiment 1, and will not be elaborated here.

[0126] It can be understood that the above-mentioned various units can be respectively or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with functional division to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the system may also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0127] According to another embodiment of the present application, the system described in this embodiment can be constructed, and the method of Embodiment 1 of the present application can be implemented by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), and a Read-Only Memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0128] Embodiment 3:

[0129] As Figure 4 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.

[0130] Among them, the communication interface 1002 is used to receive and send data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used to store a computer program, and the computer program includes program instructions. The processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.

[0131] The processor 1001 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the electronic device, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0132] The processor 1001 is configured to execute the following process:

[0133] Extract a single-frame image, an optical flow map after superimposing multiple frames of images, and the corresponding weighing signal from the monitoring video;

[0134] Obtain the image features of the single-frame image according to a plurality of successively connected convolutional blocks embedded with the CBAM attention mechanism;

[0135] Obtain the optical flow features of the optical flow map according to a plurality of successively connected Mamba modules;

[0136] After successively processing the weighing signal through self-attention processing and a feed-forward neural network, obtain the weighing features;

[0137] Perform feature-level fusion on the image features, the optical flow features, and the weighing features to obtain fused features, and perform behavior classification based on the fused features to obtain a behavior classification result.

[0138] For the specific working process, see the introduction in Embodiment 1 and will not be elaborated here.

[0139] Embodiment 4:

[0140] This implementation provides a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in an electronic device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.

[0141] Moreover, one or more instructions suitable for being loaded and executed by a processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0142] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the following process:

[0143] Extract single-frame images, optical flow maps after superimposing multiple frames of images, and corresponding weighing signals from the surveillance video;

[0144] Obtain the image features of the single-frame image according to a plurality of successively connected convolutional blocks with an embedded CBAM attention mechanism;

[0145] Obtain the optical flow features of the optical flow map according to a plurality of successively connected Mamba modules;

[0146] After successively processing the weighing signal through self-attention processing and a feedforward neural network, obtain the weighing features;

[0147] Perform feature-level fusion on the image features, the optical flow features, and the weighing features to obtain fused features, and perform behavior classification based on the fused features to obtain a behavior classification result.

[0148] For the specific working process, see the introduction in Embodiment 1 and will not be elaborated here.

[0149] Example 5:

[0150] This implementation provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the following process:

[0151] Extract a single-frame image, an optical flow map after superimposing multiple frames of images, and the corresponding weighing signal from the monitoring video;

[0152] Obtain the image features of the single-frame image according to a plurality of successively connected convolutional blocks embedded with the CBAM attention mechanism;

[0153] Obtain the optical flow features of the optical flow map according to a plurality of successively connected Mamba modules;

[0154] After the weighing signal is successively processed by self-attention processing and a feed-forward neural network, obtain weighing features;

[0155] Perform feature-level fusion on the image features, the optical flow features, and the weighing features to obtain fusion features, and perform behavior classification according to the fusion features to obtain a behavior classification result.

[0156] For the specific working process, see the introduction in Example 1 and will not be elaborated here.

[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0158] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data processing device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A vehicle scale abnormal behavior detection method based on image-signal dual-modal fusion, characterized in that: The process includes: Extract single-frame images, optical flow maps after superposition of multiple frames, and corresponding weighing signals from surveillance videos; The image features of a single frame image are obtained according to multiple sequentially connected convolution blocks embedded with the CBAM attention mechanism; According to multiple sequentially connected Mamba modules, the optical flow features of the optical flow graph are obtained; The weighing signal is processed by self-attention processing and feedforward neural network processing in sequence to obtain a weighing feature; The image features, the optical flow features and the weighing features are fused at feature level to obtain fused features, and behavior classification is performed according to the fused features to obtain a behavior classification result.

2. The abnormal behavior detection method of a vehicle scale by image-signal dual-modal fusion as claimed in claim 1, characterized in that: CBAM attention mechanism, including: channel attention module and spatial attention module; In the channel attention module, according to the initial operation result of the convolution block, channel information is obtained through the first global average pooling and the first global maximum pooling; according to the channel information, channel weights are generated through a shared fully connected layer; according to the channel weights, channel attention weights are generated and enhanced; In the spatial attention module, the enhanced results of the channel attention weights are subjected to a second global average pooling and a second global maximum pooling, the results of the second global average pooling and the second global maximum pooling are concatenated and then convolution is performed to generate spatial weights, the spatial weights are enhanced, and the output of the CBAM attention mechanism is obtained.

3. The abnormal behavior detection method of a vehicle scale by image-signal dual-modal fusion as claimed in claim 2 is characterized in that: Obtain channel information through the first global average pooling and the first global maximum pooling: F avg =GAP(F conv ),F max =GMP(F conv ); Generate channel weights through a shared fully connected layer: F MLP =W2·ReLU(W1·F avg )+W2·ReLU(W1·F max ); Generate channel attention weights and enhance: M channel =Sigmoid(F MLP ),F channel =F conv ·M c h anne l; Perform the second global average pooling and the second global maximum pooling on the channel dimension: After concatenation, spatial weights are generated through convolution: Spatial weight enhancement: F spatial =F channel ·M spatial ; Output features of CBAM attention mechanism: F out =F spatial .

4. The abnormal behavior detection method of a vehicle scale by image-signal dual-modal fusion as claimed in claim 2 or 3, characterized in that: The input of the first convolution block is the single frame image, and the subsequent convolution blocks use the output of the previous convolution block as input.

5. The abnormal behavior detection method of a vehicle scale by image-signal dual-modal fusion as claimed in claim 1, characterized in that: The Mamba module includes: performing an initial linear transformation on the input, performing a depth convolution on the result of the initial linear transformation, performing a matrix multiplication on the result of the depth convolution, performing a residual connection between the result of the matrix multiplication and the input, and obtaining the output of the Mamba module.

6. The abnormal behavior detection method of a vehicle scale by image-signal dual-modal fusion as claimed in claim 1, characterized in that: The weighing signal is processed by self-attention in sequence, including: Q = W Q S,K=W K S, Perform feedforward neural network processing, including: F S =FFN(Attention(Q,K,V)), Q is the query vector, K is the key vector, and V is the value vector.

7. An image-signal dual-modal fusion vehicle scale abnormal behavior detection system, characterized in that: include: The data acquisition unit is configured to: extract a single-frame image, an optical flow map after superposition of multiple-frame images, and a corresponding weighing signal from the surveillance video; The image feature extraction unit is configured to: obtain image features of a single frame image according to a plurality of sequentially connected convolution blocks embedded with a CBAM attention mechanism; The optical flow feature extraction unit is configured to: obtain the optical flow features of the optical flow graph according to a plurality of sequentially connected Mamba modules; The weighing feature extraction unit is configured to: process the weighing signal through self-attention processing and feedforward neural network processing in sequence to obtain a weighing feature; The abnormal behavior classification unit is configured to: perform feature-level fusion on the image feature, the optical flow feature and the weighing feature to obtain a fusion feature, and perform behavior classification according to the fusion feature to obtain a behavior classification result.

8. A computer device, characterized in that: include: a processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the method for detecting abnormal behavior of a vehicle scale by image-signal dual-modal fusion as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the abnormal behavior detection method of a vehicle scale using image-signal dual-modal fusion as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the abnormal behavior detection method of a vehicle scale using image-signal dual-modal fusion as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Method, device and equipment for identifying process integrity of basin-type insulator assembly

    CN121259922A