Multi-frame optical flow fused truck scale worker abnormal behavior intelligent identification method and system

Through the multi-frame optical flow fusion and multi-modal information fusion analysis method, combined with the convolution module, Mamba module and Transformer classification network model, the problems of low accuracy and poor stability of abnormal behavior recognition in the prior art are solved, and more accurate and reliable abnormal behavior detection is achieved.

CN120180272APending Publication Date: 2025-06-20JINAN JINZHONG ELECTRONICS SCALE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335130.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has problems in the recognition of abnormal behavior of automobile scale workers in the recognition of low accuracy, poor stability, insufficient utilization of space-time correlation information of multi-frame video data, and misidentification caused by failure to combine weighing signals.

Method used

The multi-frame optical flow fusion method is adopted to process the video frames and perform multi-modal information fusion analysis with the weighing signal. The convolution module and the Mamba module are used to extract features, and behavior recognition is performed through the Transformer classification network model.

Benefits of technology

It significantly improves the accuracy and reliability of abnormal behavior recognition, can more effectively distinguish between normal and abnormal behavior, and reduces the misidentification of pseudo-abnormal actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180272A_ABST
    Figure CN120180272A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing. The invention provides a multi-frame optical flow fused truck scale worker abnormal behavior intelligent identification method and system, and the method comprises the steps: calculating optical flows for adjacent monitoring video frame images, stacking each monitoring video frame image with the corresponding optical flow, and obtaining a multi-channel input sequence; respectively inputting the multi-channel input sequence into a convolution module branch and a Mama module branch, carrying out splicing and convolution operation on the output characteristics of the convolution module branch and the output characteristics of the Mama module branch, and fusing the output characteristics of the convolution module branch and the output characteristics of the Mama module branch with a weighing instrument signal to obtain fusion characteristics; according to the fusion features and a behavior classification network model, obtaining a behavior category, and if it is judged that the behavior category is an abnormal category, triggering an alarm and storing a corresponding monitoring video clip; if it is judged that the behavior category is a normal category, the behavior category obtained through recognition is not processed or saved; according to the invention, more accurate and more real-time abnormal behavior detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an intelligent recognition method and system for abnormal behaviors of weighbridge workers based on multi-frame optical flow fusion. Background Art

[0002] The statements in this part only provide background technologies related to the present invention, and do not necessarily constitute prior arts.

[0003] A weighbridge (platform scale) is a high-precision weighing device widely used in factories, warehouses, ports, mines and other places to measure the total weight of vehicles and the goods they carry. It is the cornerstone for achieving accurate measurement of goods and efficient logistics management. The working principle of a weighbridge is based on strain gauge load cell technology. When a vehicle stops on the weighing platform, gravity causes the elastic body on the load cell to deform, which leads to an imbalance in the bridge impedance of the strain gauge attached to the elastic body, and then an electrical signal proportional to the weight value is output. These electrical signals are amplified, converted and processed, and finally data such as weight is displayed.

[0004] There are a large number of personnel and vehicle activities at the operation site of a weighbridge (platform scale). If workers exhibit abnormal behaviors during the operation process (such as illegal operations, damaging the scale body, interfering with weighing signals, etc.), it will lead to inaccurate weighing results or pose safety hazards. Therefore, the traditional method relying on manual monitoring is prone to negligence and low efficiency; while using the video behavior recognition technology of deep learning, the action states of workers can be automatically recognized in the monitoring video stream.

[0005] However, there are the following problems in simply using existing algorithms for abnormal behavior recognition: (1) Existing supervised abnormal detection methods usually have low accuracy in global abnormal detection of complex scenes and are difficult to effectively capture abnormal behaviors in the overall scene; (2) In the face of complex and changeable on-site situations, existing algorithms relying on single visual features may experience performance fluctuations due to problems such as noise interference, incomplete data or abnormal data, affecting the stability and reliability of abnormal behavior recognition; (3) Existing algorithms do not fully utilize the spatio-temporal correlation information in multi-frame video data, and may ignore the dynamic changes between front and back frames, resulting in abnormal behaviors not being accurately detected; (4) Existing solutions simply rely on image processing algorithms for abnormal recognition without combining weighing signals, resulting in many misrecognition cases. Summary of the Invention

[0006] To solve the deficiencies of the existing technology, the present invention provides an intelligent recognition method and system for abnormal behaviors of weighbridge workers based on multi-frame optical flow fusion. By performing optical flow processing on video frames and conducting multi-modal information fusion analysis with weighing signals, misrecognition of "pseudo-abnormal actions" is avoided, and more accurate and real-time abnormal behavior detection is achieved.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides an intelligent recognition method for abnormal behaviors of weighbridge workers based on multi-frame optical flow fusion.

[0009] An intelligent recognition method for abnormal behaviors of weighbridge workers based on multi-frame optical flow fusion includes the following processes:

[0010] Obtain a continuous plurality of monitoring video frame images and the corresponding weighing instrument signals for each video frame image;

[0011] Calculate the optical flow for adjacent monitoring video frame images, and stack each monitoring video frame image with the corresponding optical flow to obtain a multi-channel input sequence;

[0012] Input the multi-channel input sequence into the convolutional module branch and the Mamba module branch respectively. After the output features of the convolutional module branch and the output features of the Mamba module branch are concatenated and convolved, they are fused with the weighing instrument signal to obtain fused features;

[0013] According to the fused features and the behavior classification network model, obtain the behavior category. If it is determined that the behavior category is an abnormal category, trigger an alarm and save the corresponding monitoring video segment; if it is determined that the behavior category is a normal category, do not process it or save the recognized behavior category.

[0014] As a further limitation of the first aspect of the present invention, calculating the optical flow for adjacent monitoring video frame images and stacking each monitoring video frame image with the corresponding optical flow to obtain a multi-channel input sequence includes:

[0015] Let X = {x1, x2, …, x T} represent a continuous plurality of monitoring video frame images;

[0016] Calculate the optical flow φ t , x t+1 ) = F(x t , x t , x t+1 ) for adjacent frames (x

[0017] Stack the original frame with the optical flow to obtain a multi-channel input: X t ′ = Concat(x t , φ t ), t = 1, 2, …, T - 1, and then obtain a multi-channel input sequence

[0018] As a further limitation of the first aspect of the present invention, the convolutional module branch includes multiple convolutional blocks connected in sequence. represents the output of the l-th convolutional block. represents the output of the (l - 1)-th convolutional block. represents the weight of the l-th convolutional block. represents the bias parameter of the l-th convolutional block. * represents the convolution operation, and σ(·) is the non-linear activation.

[0019] As a further limitation of the first aspect of the present invention, the Mamba module branch includes multiple Mamba modules connected in sequence. Any one Mamba module includes: the input feature passes through the first normalization, the first linear transformation, the depthwise separable convolution, the adaptive space-to-space transformation, the second normalization, and the second linear transformation in sequence to obtain the first feature, and the first feature is fused with the input feature to obtain the output of this Mamba module.

[0020] As a further limitation of the first aspect of the present invention, fusing the first feature with the input feature includes: h (l) = h (l-1) + αz′, where h (l-1) is the input feature, h (l) is the output of the l-th Mamba module, z′ is the output of the second linear transformation, and α can be a learnable coefficient.

[0021] As a further limitation of the first aspect of the present invention, obtaining the behavior category according to the fusion feature and the behavior Transformer classification network model includes:

[0022] Let Z final represent the final output of the Transformer classification network model, and project it to the classification space through a fully connected layer: o = FC(Z final ), and obtain the probability distribution of the abnormal behavior category through Softmax: P(Y∣X,S) = Softmax(o), where Y represents the output behavior category;

[0023] Select the category corresponding to the maximum probability as the recognized behavior category according to the probability distribution P(Y∣X,S): where X and S represent the sequence of consecutive multiple surveillance video frame images and the weighing signal sequence respectively.

[0024] In the second aspect, the present invention provides an intelligent recognition system for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion.

[0025] An intelligent recognition system for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion includes:

[0026] A data acquisition unit, configured to: acquire a continuous plurality of monitored video frame images and the weighing instrument signals corresponding to each video frame image;

[0027] An image processing unit, configured to: calculate the optical flow for adjacent monitored video frame images, stack each monitored video frame image with the corresponding optical flow to obtain a multi-channel input sequence;

[0028] A feature extraction unit, configured to: respectively input the multi-channel input sequence into a convolutional module branch and a Mamba module branch, after the output features of the convolutional module branch and the output features of the Mamba module branch are concatenated and subjected to a convolutional operation, fuse with the weighing instrument signal to obtain a fused feature;

[0029] A behavior classification unit, configured to: obtain a behavior category according to the fused feature and a behavior classification network model, if it is determined that the behavior category is an abnormal category, trigger an alarm and save the corresponding monitored video segment; if it is determined that the behavior category is a normal category, do not process or save the recognized behavior category.

[0030] In a third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium;

[0031] The processor is adapted to execute a computer program;

[0032] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the intelligent recognition method for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion as described in the first aspect of the present invention.

[0033] In a fourth aspect, the present invention provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the intelligent recognition method for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion as described in the first aspect of the present invention.

[0034] In a fifth aspect, the present invention provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the intelligent recognition method for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion as described in the first aspect of the present invention.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] 1. The present invention significantly improves the accuracy and reliability of abnormal behavior recognition by introducing weighing signals (including weight values, instrument status, etc.) and fusing visual data and weighing data in the same feature space. The weighing signal provides the correlation between behavior and weighing operation, enabling the system to more effectively distinguish normal and abnormal behaviors in complex scenarios.

[0037] 2. In the Mamba module of the present invention, the input features sequentially pass through the first normalization, the first linear transformation, the depthwise separable convolution, the adaptive spatial-spatial transformation, the second normalization, and the second linear transformation to obtain the first feature. The first feature is fused with the input feature to obtain the output of the Mamba module. This design enables stronger expressive ability in processing spatio-temporal features and can capture finer-grained dynamic changes and abnormal patterns.

[0038] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings constituting a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0040] Figure 1 It is a schematic flow chart of the intelligent recognition method for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion provided in Embodiment 1 of the present invention;

[0041] Figure 2 It is a schematic principle diagram of the intelligent recognition method for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion provided in Embodiment 1 of the present invention;

[0042] Figure 3 It is a schematic diagram of a system for intelligent recognition of abnormal behaviors of weighbridge workers with multi-frame optical flow fusion provided in Embodiment 2 of the present invention;

[0043] Figure 4 It is a schematic diagram of a computer device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0045] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0046] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0047] Embodiment 1:

[0048] As described in the background art, due to the lack of in-depth understanding of specific working scenarios and staff behaviors, existing algorithms are prone to misjudging normal operations or specific habits of staff as abnormal behaviors, resulting in unnecessary false alarms. Especially for staff behaviors under complex environmental conditions, some "pseudo-abnormal actions" of staff may be judged as abnormal behaviors (here, pseudo-abnormal actions include: staff simply touching the weighbridge, or staff passing by the weighbridge with destructive tools in hand, etc.). Therefore, how to more accurately identify abnormal actions is a technical problem that urgently needs to be solved at present. In view of this, this implementation proposes a weighbridge worker abnormal behavior recognition system and method that combines multi-frame stacked optical flow, convolutional network, Mamba module, and Transformer attention mechanism. By comprehensively analyzing multi-modal information such as video frames and weighing signal S, more accurate and real-time abnormal behavior detection is achieved. First, the technical terms and related concepts involved in this processing solution are briefly introduced below, where:

[0049] The Mamba module is an efficient and innovative selective structural state space model. It uses a selective state space to quickly process long sequences, combines multiple modes, and supports high efficiency in resolution and practicality. The Mamba module has a linear time complexity, which makes it more efficient than Transformer when processing long sequence data;

[0050] Optical flow refers to the instantaneous velocity field of pixel motion in an image, that is, the visual manifestation of the motion of an object in the image. It describes the position change of pixel points in the image from one moment to another. The optical flow field is a two-dimensional vector field, and each vector represents the motion direction and speed of each pixel point in the image;

[0051] Linear transformation is a basic concept in linear algebra. It describes a linear mapping from one vector space to another vector space. In the Mamba module, linear transformation is mainly used to process input data for subsequent calculations by the state space model;

[0052] The depthwise separable convolution in the Mamba module is a lightweight convolution operation. While maintaining good feature extraction ability, it can significantly reduce the number of model parameters and computational complexity;

[0053] Adaptive space - space transformation refers to the process of automatically adjusting transformation parameters or models according to the spatial characteristics of input data to achieve the transformation from one spatial representation to another. This transformation can be linear or non - linear. In adaptive space - space transformation, the following steps are usually involved: input data pre - processing, feature extraction, adaptive transformation, output and post - processing;

[0054] Transformer is a deep - learning model architecture based on the attention mechanism, mainly used for sequence - to - sequence learning in natural language processing (NLP) tasks. The Transformer model mainly consists of an input part, multiple layers of encoders, multiple layers of decoders, and an output part; The core of Transformer is the self - attention mechanism, which allows the model to simultaneously focus on information from different positions when processing the input sequence. The self - attention mechanism generates the output vector of the current position by calculating the similarity or attention weights between the current position and other positions, and then performing a weighted sum of the vectors of other positions according to these weights.

[0055] Specific methods, such as Figure 1 and Figure 2 shown, include the following process:

[0056] S1: Input data, the image sequence X and the weighing signal S.

[0057] Let X = {x1, x2, …, x T} represent the frame sequence (with length T) taken from the surveillance video;

[0058] Let S = {s1, s2, …, s T} represent the weighing instrument signals corresponding to each video frame (or time segment), which may include weight values, instrument status, etc.

[0059] S2: Generation of multi - frame stacked optical flow.

[0060] Calculate the optical flow for adjacent frames (x t , x t+1 ):

[0061] φ t = F(x t , x t+1 ) (1);

[0062] where F(·) is the optical flow estimation algorithm, stack the original frame and the optical flow to obtain a multi - channel input:

[0063] X′t = Concat(x t , φ t ), t = 1, 2, …, T - 1 (2);

[0064] For simplicity of expression, multiple frames stacked in time series can be organized as a multi-channel input sequence:

[0065]

[0066] S3: The convolutional module (Conv Block) and the Mamba module are extracted in parallel.

[0067] Input into a multi-layer convolutional module (i.e., the convolutional module branch) and a multi-layer Mamba module (i.e., the Mamba module branch) to obtain spatio-temporal feature representations at different scales.

[0068] In this implementation, optionally, the expression of the convolutional module can be:

[0069]

[0070] where represents the output of the l-th convolutional block, represents the output of the (l - 1)-th convolutional block, represents the weight of the l-th convolutional block, represents the bias parameter of the l-th convolutional block, * represents the convolution operation, and σ(·) is the non-linear activation.

[0071] In this implementation, an innovative Mamba module is designed, enabling the system to have stronger expressive power in processing spatio-temporal features and be able to capture finer-grained dynamic changes and abnormal patterns.

[0072] Specifically, let the input feature be h (l-1) , then the Mamba module can be expressed as:

[0073] Normalization:

[0074] u = LN(h (l-1) ) (5);

[0075] Linear transformation:

[0076] v = Linear(u) (6);

[0077] Depthwise separable convolution:

[0078] w = DWConv(v) (7);

[0079] Adaptive spatial-spatial transformation:

[0080] z = SS2D(w) (8);

[0081] Re - normalization and linear transformation:

[0082] z' = Linear(LN(z)) (9);

[0083] Residual or parallel path:

[0084] h (l) = h (l-1) + αz' (10);

[0085] where α can be a learnable coefficient or directly 1, and the final output h (l) is used as the input for the next stage or the next Mamba module; the output of the last convolutional block and the output of the last Mamba module are concatenated and then convolved to obtain the image feature h.

[0086] S4: Signal fusion.

[0087] Concatenate the feature s corresponding to the weighing signal S (which can be obtained through simple MLP or one - dimensional convolution) with the image feature h:

[0088] f = Concat(h, s) (11);

[0089] This can fuse visual features and weighing signals in the same feature space, facilitating subsequent temporal modeling.

[0090] S5: Transformer temporal context modeling.

[0091] Regard the fused features {f1, f2,..., f T-1} as a sequence and input it into the Transformer;

[0092] The formula for standard multi - head attention is as follows:

[0093]

[0094] where Q is the query vector, K is the key vector, V is the value vector, and Q, K, V are obtained by linear transformation from the output f of the previous layer respectively:

[0095] Q = fW Q ; K = fW K , V = fW V (13);

[0096] where d k is the dimension of the key vector K.

[0097] After passing through several Transformer encoder layers, high-dimensional temporal features Z are obtained.

[0098] S6: Fully connected and Softmax output.

[0099] Let Z final represent the final output of the Transformer module, and project it into the classification space through a fully connected layer (FC):

[0100] o = FC(Z final ) (14);

[0101] Finally, obtain the probability distribution of the abnormal behavior category through Softmax:

[0102] P(Y|X,S) = Softmax(o) (15);

[0103] where Y represents the output behavior category.

[0104] S7: Decision output.

[0105] Select the category corresponding to the maximum probability according to the probability distribution P(Y|X,S):

[0106]

[0107] If it is determined to be an abnormal category, trigger an alarm and save the corresponding segment; if it is normal, do nothing or only save the recognition result.

[0108] Embodiment 2:

[0109] As Figure 3 shown, this implementation provides an intelligent recognition system for abnormal behaviors of weighbridge workers with multi-frame optical flow fusion, including:

[0110] A data acquisition unit, configured to: acquire a continuous plurality of monitoring video frame images and the weighing instrument signals corresponding to each video frame image;

[0111] An image processing unit, configured to: calculate the optical flow for adjacent monitoring video frame images, stack each monitoring video frame image with the corresponding optical flow to obtain a multi-channel input sequence;

[0112] A feature extraction unit, configured to: respectively input the multi-channel input sequence into a convolutional module branch and a Mamba module branch, and after splicing and convolutional operations on the output features of the convolutional module branch and the output features of the Mamba module branch, fuse them with the weighing instrument signals to obtain fused features;

[0113] The behavior classification unit is configured to: obtain a behavior category according to the fusion feature and the behavior classification network model. If it is determined that the behavior category is an abnormal category, an alarm is triggered and the corresponding monitored video segment is saved; if it is determined that the behavior category is a normal category, no processing is performed or the recognized behavior category is saved.

[0114] The specific working processes of the above units are described in Embodiment 1 and will not be elaborated here.

[0115] It can be understood that the above units can be separately or wholly combined into one or several other units to form, or some of them can be further split into multiple smaller units with functional division to form, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of this application, the system can also include other units. In practical applications, these functions can also be assisted by other units and can be implemented by the cooperation of multiple units.

[0116] According to another embodiment of this application, the system described in this embodiment can be constructed by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0117] Embodiment 3:

[0118] As Figure 4 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.

[0119] Among them, the communication interface 1002 is used to receive and send data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used to store computer programs, and the computer programs include program instructions. The processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.

[0120] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0121] The processor 1001 is configured to execute the following process:

[0122] Obtain a continuous plurality of monitored video frame images and the weighing instrument signals corresponding to each video frame image;

[0123] Calculate the optical flow for adjacent monitored video frame images, stack each monitored video frame image with the corresponding optical flow to obtain a multi-channel input sequence;

[0124] Input the multi-channel input sequence into the convolutional module branch and the Mamba module branch respectively. After the output features of the convolutional module branch and the output features of the Mamba module branch are spliced and convolved, they are fused with the weighing instrument signal to obtain fused features;

[0125] According to the fused features and the behavior classification network model, obtain the behavior category. If it is determined that the behavior category is an abnormal category, trigger an alarm and save the corresponding monitored video segment; if it is determined that the behavior category is a normal category, do nothing or save the recognized behavior category.

[0126] The specific working process is described in the introduction of Embodiment 1 and will not be elaborated here.

[0127] Embodiment 4:

[0128] This implementation provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the electronic device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.

[0129] Also, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0130] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes the one or more instructions stored in the computer-readable storage medium to implement the following process:

[0131] Obtain a continuous plurality of monitored video frame images and the weighing instrument signals corresponding to each video frame image;

[0132] Calculate the optical flow for adjacent monitored video frame images, stack each monitored video frame image with the corresponding optical flow to obtain a multi-channel input sequence;

[0133] Input the multi-channel input sequence into the convolutional module branch and the Mamba module branch respectively. After the output features of the convolutional module branch and the output features of the Mamba module branch are concatenated and convolved, they are fused with the weighing instrument signal to obtain fused features;

[0134] According to the fused features and the behavior classification network model, obtain the behavior category. If it is determined that the behavior category is an abnormal category, trigger an alarm and save the corresponding monitored video segment; if it is determined that the behavior category is a normal category, do nothing or save the recognized behavior category.

[0135] For the specific working process, see the introduction in Embodiment 1 and will not be elaborated here.

[0136] Embodiment 5:

[0137] This implementation provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the following process:

[0138] Obtain a continuous plurality of monitored video frame images and the weighing instrument signals corresponding to each video frame image;

[0139] Calculate the optical flow for adjacent monitored video frame images, stack each monitored video frame image with the corresponding optical flow to obtain a multi-channel input sequence;

[0140] Input the multi-channel input sequences into the convolutional module branch and the Mamba module branch respectively. After the output features of the convolutional module branch and the output features of the Mamba module branch are concatenated and convolved, they are fused with the weighing instrument signal to obtain fused features;

[0141] According to the fused features and the behavior classification network model, obtain the behavior category. If it is determined that the behavior category is an abnormal category, trigger an alarm and save the corresponding monitored video segment; if it is determined that the behavior category is a normal category, do nothing or save the recognized behavior category.

[0142] The specific working process is as described in Embodiment 1 and will not be elaborated here.

[0143] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0144] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0145] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for intelligently identifying abnormal behavior of truck scale workers based on multi-frame optical flow fusion, characterized in that: The process includes: Obtain multiple consecutive monitoring video frame images and the weighing instrument signal corresponding to each video frame image; Calculate the optical flow of adjacent surveillance video frame images, stack each surveillance video frame image with the corresponding optical flow to obtain a multi-channel input sequence; The multi-channel input sequence is input into the convolution module branch and the Mamba module branch respectively, and the output features of the convolution module branch and the output features of the Mamba module branch are fused with the weighing instrument signal after splicing and convolution operations to obtain fused features; According to the fusion features and the behavior classification network model, the behavior category is obtained. If the behavior category is determined to be an abnormal category, an alarm is triggered and the corresponding surveillance video clip is saved; if the behavior category is determined to be a normal category, no processing is performed or the identified behavior category is saved.

2. The method for intelligently identifying abnormal behavior of truck scale workers based on multi-frame optical flow fusion as claimed in claim 1 is characterized in that: Calculate the optical flow of adjacent surveillance video frame images, stack each surveillance video frame image with the corresponding optical flow, and obtain a multi-channel input sequence, including: Let X = {x1, x2, …, x T } represents multiple consecutive surveillance video frame images; For adjacent frames (x t ,x t+1 ) Calculate the optical flow φ t =F(x t ,x t+1 ), where F(·) is the optical flow estimation algorithm; Stack the original frame with the optical flow to get a multi-channel input: X t ′=Concat(x t ,φ t ),t=1,2,…,T-1, and then get the multi-channel input sequence 3. The method for intelligently identifying abnormal behavior of truck scale workers by multi-frame optical flow fusion as claimed in claim 1 is characterized in that: The convolutional module branch includes multiple layers of sequentially connected convolutional blocks. represents the output of the l-th convolutional block, represents the output of the l-1th convolutional block, represents the weight of the l-th layer convolutional block, represents the bias parameter of the l-th convolutional block, * represents the convolution operation, and σ(·) is the nonlinear activation.

4. The method for intelligently identifying abnormal behavior of truck scale workers by multi-frame optical flow fusion as claimed in claim 1 is characterized in that: The Mamba module branch includes multiple Mamba modules connected in sequence in multiple layers, and any Mamba module includes: the input feature is sequentially subjected to first normalization, first linear transformation, depthwise separable convolution, adaptive space-space conversion, second normalization and second linear transformation to obtain a first feature, and the first feature is fused with the input feature to obtain the output of this Mamba module.

5. The method for intelligently identifying abnormal behavior of truck scale workers by multi-frame optical flow fusion as claimed in claim 4 is characterized in that: The first feature is combined with the input feature, comprising: h (l) =h (l-1) +αz′, where h (l-1) is the input feature, h (l) is the output of the l-th layer Mamba module, z′ is the output of the second linear transformation, and α is a learnable coefficient.

6. The method for intelligently identifying abnormal behavior of truck scale workers by multi-frame optical flow fusion according to any one of claims 1 to 5, characterized in that: According to the fusion features and the behavior Transformer classification network model, the behavior categories are obtained, including: Order Z final Represents the final output of the Transformer classification network model, which is projected into the classification space through the fully connected layer: o = FC(Z final ), and obtain the probability distribution of abnormal behavior categories through Softmax: P(Y|X,S)=Softmax(o), where Y represents the output behavior category; According to the probability distribution P(Y|X,S), the category corresponding to the maximum probability is selected as the identified behavior category: Wherein, X and S represent a sequence of multiple consecutive surveillance video frames and a sequence of weighing signals, respectively.

7. A multi-frame optical flow fusion intelligent recognition system for abnormal behavior of truck scale workers, characterized by: include: The data acquisition unit is configured to: acquire a plurality of continuous monitoring video frame images and a weighing instrument signal corresponding to each video frame image; The image processing unit is configured to: calculate optical flows for adjacent surveillance video frame images, and stack each surveillance video frame image with the corresponding optical flow to obtain a multi-channel input sequence; The feature extraction unit is configured to: input the multi-channel input sequence into the convolution module branch and the Mamba module branch respectively, and after splicing and convolution operations, the output features of the convolution module branch and the output features of the Mamba module branch are fused with the weighing instrument signal to obtain fused features; The behavior classification unit is configured to: obtain the behavior category based on the fusion features and the behavior classification network model; if the behavior category is determined to be an abnormal category, trigger an alarm and save the corresponding surveillance video clip; if the behavior category is determined to be a normal category, do not process it or save the identified behavior category.

8. A computer device, characterized in that: include: A processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the method for intelligently identifying abnormal behaviors of truck scale workers by multi-frame optical flow fusion as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is suitable for being loaded by a processor and executing the method for intelligently identifying abnormal behaviors of truck scale workers by multi-frame optical flow fusion as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the method for intelligently identifying abnormal behavior of truck scale workers by multi-frame optical flow fusion as described in any one of claims 1 to 6 is implemented.