Tunnel construction video real-time abnormal behavior identification method and system

By setting up a wide-angle main camera and telephoto auxiliary camera in the tunnel construction area, a complementary field of view is formed, combined with a multi-head attention convolution network and behavior prediction classification model, abnormal behaviors in tunnel construction are identified, and the problems of poor monitoring accuracy and high missed detection rate caused by blind spot field in tunnel construction are solved, and efficient safety monitoring is achieved.

CN120340138AActive Publication Date: 2025-07-18POLY CHANGDA ENGINEERING CO LTD

Patent Information

Application Number
CN202510814872.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Due to blind spot visual problems during tunnel construction, the monitoring accuracy is poor and the missed detection rate is high, making it difficult to ensure construction safety.

Method used

A wide-angle main camera and telephoto auxiliary camera are set up in the tunnel construction area, and the field of view is combined to form a complementary field of view, target personnel are detected and tracked, and behavioral characteristics are extracted using a convolutional network of multi-head attention, and a pre-trained behavior prediction classification model is input to identify abnormal behaviors.

Benefits of technology

It effectively solves the blind spot vision problem, improves the accuracy of real-time monitoring of tunnel construction videos, avoids missed target inspections, and ensures construction safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340138A_ABST
    Figure CN120340138A_ABST
Patent Text Reader

Abstract

The invention discloses a tunnel construction video real-time abnormal behavior identification method and system, and the method comprises the steps: constructing a wide-angle field of view and a long-focus field of view through a wide-angle video image and a long-focus video image which are shot by a wide-angle main camera and a long-focus auxiliary camera respectively, and carrying out the field of view combination of the wide-angle field of view and the long-focus field of view. Meanwhile, target detection is carried out on a target person in the long-focus video image, the target person obtained through target detection is tracked, a motion area of the target person is obtained, if the motion area does not completely fall into the view field complementary area, it is indicated that a motion blind area exists, supplementary shooting is carried out on the motion blind area through a long-focus auxiliary camera, and the motion blind area does not completely fall into the view field complementary area. And the behavior characteristics of the target person in the field-of-view complementary region after updating are extracted based on the multi-attention convolutional network, the behavior characteristics are inputted to a pre-trained behavior prediction classification model, and whether the behavior of the target person is abnormal is identified, so that the blind area field-of-view problem is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method and system for real-time abnormal behavior recognition of tunnel construction videos. Background Art

[0002] Due to the narrow and relatively enclosed tunnel construction environment, there is a large blind spot field of view, making it difficult to timely detect potential hazards in tunnel construction. Currently, in tunnel construction monitoring, multiple cameras are mostly used for shooting. However, due to the relatively fixed field of view of the cameras, it is still difficult to solve the problem of the blind spot field of view. At the same time, if a large number of cameras are used for shooting, the construction cost will inevitably increase. Therefore, the current tunnel construction monitoring method has poor monitoring accuracy and a high missed detection rate due to the blind spot field of view problem, making it difficult to ensure the normal operation of tunnel construction. Summary of the Invention

[0003] In view of this, the present invention provides a method and system for real-time abnormal behavior recognition of tunnel construction videos, which solves the technical problem that the current tunnel construction monitoring method has poor monitoring accuracy and a high missed detection rate due to the blind spot field of view problem.

[0004] The first aspect of the present invention provides a method for real-time abnormal behavior recognition of tunnel construction videos. A wide-angle main camera and a telephoto auxiliary camera are set in any one of multiple target construction areas in the tunnel, including: Collecting wide-angle video images and telephoto video images in the target construction area through the wide-angle main camera and the telephoto auxiliary camera respectively; Determining a wide-angle field of view and a telephoto field of view according to the wide-angle video images and the telephoto video images, and performing field of view combination on the wide-angle field of view and the telephoto field of view to obtain a field of view complementary area; Performing target detection on target personnel in the telephoto video images, and tracking the target personnel detected by the target detection to obtain the movement area of the target personnel; Judging whether the movement area completely falls within the field of view complementary area; When the movement area does not completely fall within the field of view complementary area, the telephoto auxiliary camera is used to supplement the shooting of the movement blind area, and the supplemented movement blind area is updated to the field of view complementary area; wherein, the movement blind area is the part of the movement area that does not fall within the field of view complementary area; Extracting the behavior features of the target personnel in the updated field of view complementary area based on a convolutional network with multi-head attention; Inputting the behavior features into a pre-trained behavior prediction classification model to identify whether the behavior of the target personnel is abnormal.

[0005] Optionally, determining a wide-angle field of view and a telephoto field of view based on the wide-angle video image and the telephoto video image, and performing field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary region, includes: Based on the wide-angle video image, establish a wide-angle field of view in a polar coordinate system with the optical center as the origin; Based on the telephoto video image, establish a telephoto field of view in a rectangular coordinate system with the optical center as the origin; Based on the wide-angle field of view and the telephoto field of view, determine an overlapping region and a non-overlapping region; Based on the area ratio of the overlapping region and the non-overlapping region, adjust the focal length parameter of the wide-angle main camera, and update the wide-angle field of view according to the adjusted focal length parameter until the area ratio is greater than a preset area ratio threshold, to obtain an updated wide-angle field of view; Based on the updated wide-angle field of view, re-determine the overlapping region and the non-overlapping region, and integrate the overlapping region and the non-overlapping region into the field-of-view complementary region.

[0006] Optionally, the non-overlapping region includes a wide-angle exclusive region and a telephoto exclusive region; The determining an overlapping region and a non-overlapping region based on the wide-angle field of view and the telephoto field of view includes: Convert the updated wide-angle field of view to a rectangular coordinate system, and fuse the overlapping region of the wide-angle field of view and the telephoto field of view in the rectangular coordinate system through weighted averaging to obtain a fused overlapping region; Map each pixel point in the wide-angle exclusive region to the rectangular coordinate system through an initial mapping matrix to obtain a first pixel point; and map each pixel point in the telephoto exclusive region to the polar coordinate system through the initial mapping matrix to obtain a second pixel point; Perform weighted fusion on the first pixel point and the second pixel point through an initial fusion weight to obtain an initially fused non-overlapping region; Based on the initially fused non-overlapping region, construct a global energy constraint function; wherein, the global energy constraint function is used to constrain the smoothness and consistency in the image fusion process; Through optimizing and solving the global energy constraint function, obtain a fusion weight and a mapping matrix under minimizing the global energy constraint function; Based on the fusion weight and the mapping matrix, optimize the initially fused non-overlapping region to obtain an optimized non-overlapping region.

[0007] Optionally, performing target detection on a target person in the telephoto video image, and tracking the detected target person to obtain a motion region of the target person, includes: Use the YOLO object detection algorithm to detect the target personnel in the long - focal - length video image, and obtain the position information and bounding box of the target personnel; According to the position information and the bounding box, use the optical flow method to calculate the optical flow vector of the target personnel between adjacent frames in the long - focal - length video image; According to the optical flow vector, track the bounding box displacement trajectory points of the detected target personnel, and determine the movement trajectory of the target personnel according to the bounding box displacement; wherein, the movement trajectory includes a plurality of bounding box displacement trajectory points; According to the movement trajectory, perform dilation calculation on each of the bounding box displacement trajectory points to obtain the dilation neighborhood of each of the bounding box displacement trajectory points; Perform connected - component processing on the dilation neighborhood of each of the bounding box displacement trajectory points to obtain the dilated motion region, and screen the effective regions in the dilated motion region through a density threshold to obtain the motion region of the target personnel.

[0008] Optionally, the performing dilation calculation on each of the bounding box displacement trajectory points according to the movement trajectory to obtain the dilation neighborhood of each of the bounding box displacement trajectory points includes: According to each of the bounding box displacement trajectory points, convert the bounding box displacement trajectory point to a spatio - temporal point in the spatio - temporal domain, and determine the density of the spatio - temporal point; Determine the dilation amplitude of the bounding box displacement trajectory point according to the density of the spatio - temporal point; Determine the dilation radius of the bounding box displacement trajectory point according to the dilation amplitude and the velocity vector in the optical flow vector; Determine the dilation neighborhood corresponding to the bounding box displacement trajectory point according to the dilation radius; wherein, the dilation neighborhood is a circular region drawn with the bounding box displacement trajectory point as the center and the dilation radius.

[0009] Optionally, the extracting the behavior features of the target personnel in the updated field - of - view complementary region by the convolutional network based on multi - head attention includes: Input the target personnel image in the updated field - of - view complementary region into the convolutional network model based on multi - head attention; wherein, the convolutional network model based on multi - head attention includes a plurality of self - attention heads including a convolutional network; Calculate the correlation weights between different trajectory points in the target personnel image through each of the self - attention heads to obtain a correlation weight matrix; Extract the features of the target personnel at each of the trajectory points from the target personnel image through the convolutional network, and perform weighted summation on all the features in the target personnel image according to the correlation weight matrix to obtain a weighted feature map; Concatenate the weighted feature maps and process them through a fully connected layer to obtain the behavioral feature vector of the target person; wherein, the behavioral feature vector is used to characterize the behavioral pattern of the target person.

[0010] Optionally, this method further includes: Collect historical behavioral feature samples of the workers in the tunnel construction video and the corresponding behavioral anomaly categories of the historical behavioral feature samples; Construct a training set and a test set according to the historical behavioral feature samples and the corresponding behavioral anomaly categories of the historical behavioral feature samples; Construct an initial behavioral prediction classification model, and the initial behavioral prediction classification model uses a deep learning algorithm; Use the training set to train the behavioral prediction classification model to obtain a behavioral prediction classification model; Use the test set to test the behavioral prediction classification model, and optimize the network parameters of the behavioral prediction classification model according to the test results to obtain an optimized behavioral prediction classification model.

[0011] In a second aspect, the present invention also provides a real-time abnormal behavior recognition system for tunnel construction videos. A wide-angle main camera and a telephoto auxiliary camera are set in any one of multiple target construction areas in the tunnel, including: An image acquisition module, configured to respectively collect wide-angle video images and telephoto video images in the target construction area through the wide-angle main camera and the telephoto auxiliary camera; A field of view combination module, configured to determine a wide-angle field of view and a telephoto field of view according to the wide-angle video image and the telephoto video image, and perform field of view combination on the wide-angle field of view and the telephoto field of view to obtain a field of view complementary area; A moving area determination module, configured to perform target detection on the target person in the telephoto video image and track the detected target person to obtain the moving area of the target person; An area judgment module, configured to judge whether the moving area completely falls within the field of view complementary area; A blind area supplementary shooting module, configured to, when the moving area does not completely fall within the field of view complementary area, perform supplementary shooting on the moving blind area through the telephoto auxiliary camera and update the supplemented moving blind area to the field of view complementary area; wherein, the moving blind area is the partial moving area where the moving area does not fall within the field of view complementary area; A feature extraction module, configured to extract the behavioral features of the target person in the updated field of view complementary area based on a convolutional network with multi-head attention; The abnormal behavior recognition module is used to input the behavior features into a pre-trained behavior prediction classification model to identify whether the behavior of the target person is abnormal.

[0012] In a third aspect, the present invention further provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the tunnel construction video real-time abnormal behavior recognition method as described in the first aspect.

[0013] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the tunnel construction video real-time abnormal behavior recognition method as described in the first aspect are implemented.

[0014] As can be seen from the above technical solutions, the present invention takes advantage of the respective advantages of the wide-angle main camera and the telephoto auxiliary camera, constructs a wide-angle field of view and a telephoto field of view from the wide-angle video image and the telephoto video image captured by the two respectively, and performs field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary area, so as to form a complete and continuous field-of-view complementary area. At the same time, target detection is performed on the target person in the telephoto video image, and the detected target person is tracked to obtain the movement area of the target person, and it is judged whether the movement area completely falls within the field-of-view complementary area. If the movement area does not completely fall within the field-of-view complementary area, it means that there is a movement blind area, and the telephoto auxiliary camera is used to supplement the shooting of the movement blind area, and the supplemented movement blind area is updated to the field-of-view complementary area. The behavior features of the target person in the updated field-of-view complementary area are also extracted based on the convolutional network with multi-head attention, and the behavior features are input into a pre-trained behavior prediction classification model to identify whether the behavior of the target person is abnormal, thus effectively solving the problem of blind area vision, improving the accuracy of real-time monitoring of tunnel construction videos, and effectively avoiding target missing detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a flowchart of a tunnel construction video real-time abnormal behavior recognition method provided by an embodiment of the present invention; Figure 2 It is a schematic structural diagram of a tunnel construction video real-time abnormal behavior recognition system provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0017] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] An embodiment of the present application provides a method for real-time abnormal behavior recognition in tunnel construction videos. A wide-angle main camera and a long-focus auxiliary camera are set in any one of multiple target construction areas in the tunnel.

[0019] Among them, since the tunnel is relatively long and narrow, the present application divides the tunnel into multiple target construction areas along the length direction, and sets a wide-angle main camera and a long-focus auxiliary camera in each area. Among them, the wide-angle main camera is responsible for capturing video images in a large range to cover the entire target construction area and ensure the broadness of the field of view; while the long-focus auxiliary camera focuses on capturing details at a long distance for more refined observation of specific areas or targets. Through the collaborative work of the wide-angle main camera and the long-focus auxiliary camera, comprehensive monitoring of the tunnel construction area can be achieved, effectively reducing the problem of blind spot vision.

[0020] In specific implementation, as Figure 1 shown, an embodiment of the present application provides a method for real-time abnormal behavior recognition in tunnel construction videos, including: 101. Respectively collect wide-angle video images and long-focus video images in the target construction area through the wide-angle main camera and the long-focus auxiliary camera.

[0021] Among them, the position arrangements of the wide-angle main camera and the long-focus auxiliary camera are flexibly adjusted according to the actual situation of the tunnel construction area to ensure that the wide-angle video images and the long-focus video images can comprehensively cover the target construction area and effectively reduce the blind spot vision. The wide-angle main camera is usually arranged at a higher position to obtain a wider field of view range, while the long-focus auxiliary camera is arranged at a position where key details can be captured according to specific monitoring requirements. Such an arrangement method enables the wide-angle main camera and the long-focus auxiliary camera to complement each other and jointly achieve comprehensive monitoring of the tunnel construction area.

[0022] 102. Determine the wide-angle field of view and the long-focus field of view according to the wide-angle video images and the long-focus video images, and perform field-of-view combination on the wide-angle field of view and the long-focus field of view to obtain a field-of-view complementary area.

[0023] Among them, the wide-angle field of view is the field of view formed by the video image captured by the wide-angle main camera, which is characterized by a wide field of view and can cover a large construction area; while the telephoto field of view is the field of view formed by the video image captured by the telephoto auxiliary camera, and its advantage is that it can capture details at a long distance and conduct fine observation on specific targets. By combining the wide-angle field of view and the telephoto field of view, a field of view complementary area that is both wide and fine can be obtained. This area not only covers the entire target construction area but also can clearly present key details, thereby effectively improving the accuracy and comprehensiveness of monitoring. Generally, the wide-angle field of view is wider than the telephoto field of view.

[0024] During the process of field of view combination, it is necessary to consider the overlapping area and non-overlapping area of the wide-angle field of view and the telephoto field of view. By organically combining the information of the two fields of view, a complete and continuous field of view complementary area is formed. In this way, comprehensive monitoring of the tunnel construction area can be realized, effectively reducing the problem of blind spot field of view and improving construction safety and efficiency.

[0025] 103. Perform target detection on the target personnel in the telephoto video image, and track the target personnel detected by the target detection to obtain the movement area of the target personnel.

[0026] Among them, the target detection algorithm used for target detection can adopt deep learning algorithms such as the YOLO (You Only Look Once) series of algorithms. This algorithm has the advantages of fast detection speed and high accuracy and is suitable for real-time detection of target personnel in tunnel construction videos. During the target detection process, the algorithm will traverse each pixel point in the video image and judge whether the pixel point belongs to the target personnel by calculating the eigenvalue. Once the target personnel are detected, the algorithm will immediately record their position information and bounding boxes.

[0027] Target tracking is to continuously track the detected target personnel to obtain their movement trajectories. In the present invention, the optical flow method can be adopted to achieve target tracking. The optical flow method is a method based on the spatio-temporal gradient of pixel intensity data in an image sequence and can calculate the movement information of the target in the image. By calculating the optical flow vectors of the target personnel between adjacent frames, the displacement trajectory points of the bounding boxes of the target personnel can be tracked, thereby obtaining their movement trajectories. The acquisition of movement trajectories is crucial for subsequent abnormal behavior recognition because it can reflect the movement state and behavior patterns of the target personnel.

[0028] 104. Judge whether the movement area completely falls within the field of view complementary area.

[0029] 105. When the moving area does not completely fall within the field-of-view complementary area, the long-focus auxiliary camera is used to supplement the shooting of the moving blind area, and the supplemented moving blind area is updated to the field-of-view complementary area; wherein, the moving blind area is the part of the moving area that does not fall within the field-of-view complementary area. The

[0030] Among them, if the moving area does not completely fall within the field-of-view complementary area, it means that there is a moving blind area, that is, part of the movement trajectory of the target person is not monitored. In this case, it is necessary to adjust the parameters of the long-focus auxiliary camera to supplement the shooting of the moving blind area to obtain complete movement information. The supplemented moving blind area image will be updated to the field-of-view complementary area, so as to achieve comprehensive monitoring of the movement trajectory of the target person. Such a design can effectively avoid potential safety hazards caused by monitoring blind areas and improve the safety and efficiency of tunnel construction.

[0031] 106. The convolutional network based on multi-head attention is used to extract the behavior characteristics of the target person in the updated field-of-view complementary area.

[0032] Among them, multi-head attention is a feature extraction method that can effectively capture the behavior details of the target person. It deeply analyzes the target person's image in the updated field-of-view complementary area and extracts key behavior characteristics. In specific implementation, the multi-head attention mechanism will divide the target person's image into multiple regions, and each region will independently calculate the correlation weights with other regions. These weights reflect the degree of association between different regions, which helps the model to more accurately understand the behavior pattern of the target person. Through steps such as weighted summation and feature splicing, the multi-head attention mechanism can generate a feature vector containing rich behavior information, providing strong support for subsequent behavior anomaly recognition. The application of this method enables more accurate identification of the abnormal behavior of the target person and improves the accuracy and reliability of monitoring.

[0033] 107. The behavior characteristics are input into a pre-trained behavior prediction classification model to identify whether the behavior of the target person is abnormal.

[0034] Among them, the behavior prediction classification model is a model constructed based on deep learning algorithms. It can accurately classify the input behavior characteristics to judge whether the behavior of the target person is abnormal. During the training phase of the model, it learns through a large number of historical behavior characteristic samples and their corresponding behavior anomaly categories, and continuously optimizes the network parameters to improve the accuracy and generalization ability of classification. In actual application, the extracted behavior characteristics of the target person are input into the model, and the model will quickly output a judgment result indicating whether the behavior of the target person belongs to abnormal behavior. If an abnormal behavior is identified, the system can immediately trigger an alarm mechanism to notify relevant personnel for timely handling, thereby effectively preventing the occurrence of safety accidents and ensuring the safe and smooth progress of tunnel construction.

[0035] It should be noted that, by leveraging the respective advantages of the wide-angle main camera and the telephoto auxiliary camera, this application constructs a wide-angle field of view and a telephoto field of view from the wide-angle video image and the telephoto video image captured by them respectively, and performs field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary area, thereby forming a complete and continuous field-of-view complementary area. At the same time, target detection is performed on the target personnel in the telephoto video image, and the target personnel detected by the target detection are tracked to obtain the movement area of the target personnel, and it is judged whether the movement area completely falls within the field-of-view complementary area. If the movement area does not completely fall within the field-of-view complementary area, it indicates that there is a movement blind area, and the telephoto auxiliary camera is used to supplement the shooting of the movement blind area, and the supplemented movement blind area is updated to the field-of-view complementary area. Additionally, the behavior features of the target personnel in the updated field-of-view complementary area are extracted based on the convolutional network with multi-head attention, and the behavior features are input into the pre-trained behavior prediction classification model to identify whether the behavior of the target personnel is abnormal, thus effectively solving the blind area vision problem, improving the accuracy of real-time monitoring of tunnel construction videos, and effectively avoiding target missing detection.

[0036] In some embodiments, determining the wide-angle field of view and the telephoto field of view based on the wide-angle video image and the telephoto video image, and performing field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary area, includes: 201. Based on the wide-angle video image, with the optical center as the origin, establish a wide-angle field of view in the polar coordinate system.

[0037] Among them, with the optical center (H1, 0) as the origin, the wide-angle field of view of the sector in the polar coordinate system is: In the formula, is the wide-angle field of view, is the radius, is the angle, is the initial angle.

[0038] 202. Based on the telephoto video image, with the optical center as the origin, establish a telephoto field of view in the rectangular coordinate system.

[0039] Among them, with the optical center (X c , Y c ) as the origin, the rectangular telephoto field of view established in the rectangular coordinate system is: In the formula, is the telephoto field of view, are the horizontal and vertical coordinates of the telephoto field of view, is the lower limit of the abscissa, , are the lower and upper limits of the ordinate respectively.

[0040] 203. Determine the overlapping area and the non-overlapping area according to the wide-angle field of view and the telephoto field of view.

[0041] Among them, the overlapping area is the part jointly covered by the wide-angle field of view and the telephoto field of view, while the non-overlapping area is the unique field of view of each. The non-overlapping area includes the wide-angle exclusive area and the telephoto exclusive area. By accurately dividing the overlapping area and the non-overlapping area, it is possible to more clearly understand the monitoring ranges of the wide-angle main camera and the telephoto auxiliary camera respectively and their complementary relationship.

[0042] Specifically, step 203 includes: Step 2031: Convert the updated wide-angle field of view to a rectangular coordinate system, and fuse the overlapping area of the wide-angle field of view and the telephoto field of view in the rectangular coordinate system through weighted averaging to obtain the fused overlapping area.

[0043] Among them, the updated wide-angle field of view is converted from a polar coordinate system to a rectangular coordinate system, so as to facilitate fusion with the telephoto field of view in the same coordinate system. After obtaining the wide-angle field of view in the rectangular coordinate system, it is compared with the telephoto field of view to determine their overlapping area. The overlapping area is the part jointly covered by the wide-angle field of view and the telephoto field of view. This part of the area can enjoy both the broad field of view of the wide-angle field of view and the fine details of the telephoto field of view, so it has extremely high monitoring value.

[0044] For the overlapping area, the present invention adopts the method of weighted average fusion for processing. Specifically, it is to perform weighted averaging according to the pixel values of the wide-angle field of view and the telephoto field of view in the overlapping area according to a certain weight to obtain the fused overlapping area, so as to make full use of the information of the two fields of view, and enable the fused overlapping area to have better performance in terms of field of view broadness and detail clarity.

[0045] Step 2032: Map each pixel point in the wide-angle exclusive area to the rectangular coordinate system through the initial mapping matrix to obtain the first pixel point; and map each pixel point in the telephoto exclusive area back to the polar coordinate through the initial mapping matrix to obtain the second pixel point.

[0046] Among them, let each pixel point in the wide-angle exclusive area be I p ( )The pixel points of the telephoto exclusive area in the rectangular coordinate system are I c (x c ,y c ), and the non-overlapping area is defined as: Wide-angle exclusive area:

[0047] Telephoto exclusive area:

[0048] By introducing the initial mapping matrix, each pixel point in the wide-angle exclusive area is mapped to the rectangular coordinate system, and the obtained first pixel point is: In the formula, is the first pixel point, is the initial mapping matrix, and T is the matrix transpose.

[0049] By introducing the initial mapping matrix, each pixel point in the long - focal exclusive area is reflected onto the polar coordinates, and the second pixel point is obtained as: In the formula, is the second pixel point.

[0050] Among them, by introducing the scale adjustment parameter ∈(0, 1], the initial mapping matrix is constructed as: In the formula, is the scale - scaling coefficient, is the rotation - angle adjustment coefficient, 、 are the translation parameters, is the scale adjustment parameter, and is determined by the spatial position of the non - overlapping area.

[0051] Among them, In the formula, 、 are the initial scaling coefficient and the target scaling coefficient respectively.

[0052] In the formula, is the initial angle, is the angle adjustment amount.

[0053] Step 2033: According to the first pixel point and the second pixel point, weighted fusion is performed through the initial fusion weight to obtain the non - overlapping area after initial fusion.

[0054] Among them, the initial fusion weight is calculated by combining gradient and distance information, and the initial fusion weight is: In the formula, is the wide - angle field - of - view fusion weight, is the long - focal field - of - view fusion weight, 、 are the gradient - threshold adjustment parameters, 、 are the image gradients, representing the texture complexity, is the distance from the polar - coordinate point to the boundary of the overlapping area, is the distance - attenuation coefficient, is the distance from the rectangular - coordinate point to the boundary of the overlapping area.

[0055] Weighted fusion is performed using the initial fusion weights to obtain the non-overlapping region after initial fusion. The set of pixel points in is as follows: is the pixel point image of the long-focus exclusive region reflected onto the polar coordinates, is the pixel point image of the long-focus exclusive region in the rectangular coordinate system, is the pixel point image of the wide-angle exclusive region mapped onto the rectangular coordinate system, is the pixel point image of the wide-angle exclusive region on the polar coordinates.

[0056] Step 2034: Construct a global energy constraint function based on the non-overlapping region after initial fusion.

[0057] Among them, the global energy constraint function is used to constrain the smoothness and consistency in the image fusion process.

[0058] The global energy constraint function is:

[0059] In the formula, is the global energy, is the regularization coefficient, which controls the gradient smoothness.

[0060] Step 2035: Through optimizing and solving the global energy constraint function, obtain the fusion weights and mapping matrix under the minimization of the global energy constraint function.

[0061] Among them, an iterative optimization algorithm, such as the gradient descent method or the conjugate gradient method, etc., is used to optimize and solve the global energy constraint function. In each iteration process, according to the current fusion weights and mapping matrix, calculate the global energy, and update the fusion weights and mapping matrix through the gradient information to gradually reduce the global energy until the convergence condition or the preset number of iterations is reached. Finally, obtain the optimal fusion weights and mapping matrix under the minimization of the global energy constraint function. These optimal parameters can ensure the smoothness and consistency in the image fusion process, so that the fused image can have better performance in terms of wide field of view and detail clarity.

[0062] Step 2036: Optimize the non-overlapping region after initial fusion according to the fusion weights and mapping matrix to obtain the optimized non-overlapping region.

[0063] Among them, according to the finally determined fusion weights and mapping matrix, the initially fused non-overlapping regions are further optimized. The purpose of this step is to ensure that the fusion effect of the non-overlapping regions is more natural and accurate, avoiding obvious splicing traces or information loss. By fine-tuning the fusion weights, the information contributions of the wide-angle field of view and the telephoto field of view in the non-overlapping regions can be balanced, making the fused image more consistent and coherent in the overall visual effect. At the same time, using the mapping matrix to accurately map pixel points can further reduce the errors introduced by coordinate conversion and improve the accuracy of image fusion. After the optimization process, the non-overlapping regions will be able to better fuse with the overlapping regions to form a complete and high-quality field-of-view complementary region.

[0064] 204. Adjust the focal length parameter of the wide-angle main camera according to the area ratio of the overlapping region and the non-overlapping region, update the wide-angle field of view according to the adjusted focal length parameter, and stop when the area ratio is greater than the preset area ratio threshold to obtain the updated wide-angle field of view.

[0065] Among them, since the overlapping region represents the region jointly covered by the wide-angle main camera and the telephoto auxiliary camera, its area ratio can reflect the degree of field-of-view overlap between the two cameras. To obtain a better field-of-view complementary effect, the present invention dynamically adjusts the focal length parameter of the wide-angle main camera according to the area ratio of the overlapping region and the non-overlapping region.

[0066] Specifically, first calculate the area ratio of the overlapping region and the non-overlapping region, and then determine whether this ratio meets the preset area ratio threshold. If it does not meet (the area ratio threshold is small, indicating that the overlapping region is small, that is, the shooting of the wide-angle main camera needs to be optimized), then adjust the focal length parameter of the wide-angle main camera according to the difference in the area ratio to expand or shrink the range of the wide-angle field of view, thereby increasing or decreasing the area of the overlapping region. After adjusting the focal length parameter, update the wide-angle field of view and recalculate the area ratio until the area ratio is greater than the preset area ratio threshold. In this way, by continuously adjusting the focal length parameter of the wide-angle main camera, the area ratio of the overlapping region and the non-overlapping region can be optimized, thus ensuring the quality and monitoring effect of the field-of-view complementary region. After obtaining the updated wide-angle field of view, combine it with the telephoto field of view again to obtain the updated field-of-view complementary region, providing more accurate and reliable image information for subsequent target detection and abnormal behavior recognition.

[0067] 205. Redetermine the overlapping region and the non-overlapping region according to the updated wide-angle field of view, and integrate the overlapping region and the non-overlapping region into the field-of-view complementary region.

[0068] Among them, after obtaining the updated wide-angle field of view, it is necessary to re-determine the overlapping area and non-overlapping area. The purpose of this step is to ensure that after the focal length parameter is adjusted, the division of the overlapping area and non-overlapping area is still accurate, so as to provide a reliable image basis for subsequent target detection and abnormal behavior recognition. Specifically, according to the relative position relationship between the updated wide-angle field of view and the telephoto field of view, recalculate the overlapping part and non-overlapping part between them, and splice the overlapping area and non-overlapping area to obtain a complementary field of view area.

[0069] In some embodiments, target detection is performed on the target person in the telephoto video image, and the detected target person is tracked to obtain the movement area of the target person, including: Step S301: Use the YOLO target detection algorithm to perform target detection on the target person in the telephoto video image, and obtain the position information and bounding box of the target person.

[0070] Among them, the YOLO (You Only Look Once) target detection algorithm is a deep learning-based target detection algorithm. It can predict the position and category of the target simultaneously in a single forward propagation, and has the advantages of fast detection speed and high accuracy. In the present invention, using the YOLO target detection algorithm to perform target detection on the telephoto video image can quickly and accurately identify the target person in the telephoto video image and obtain its position information and bounding box.

[0071] Step S302: According to the position information and the bounding box, use the optical flow method to calculate the optical flow vector of the target person between adjacent frames in the telephoto video image.

[0072] Among them, the optical flow method is a method for describing the movement of pixel points in an image. By analyzing the movement trajectories of pixel points in an image sequence, it can calculate the movement speed and direction of the target person between adjacent frames. In the present invention, using the optical flow method to calculate the optical flow vector of the target person between adjacent frames in the telephoto video image can realize the tracking of the movement trajectory of the target person and provide basic data for subsequent abnormal behavior recognition. Specifically, first, according to the position information and the bounding box of the target person, determine the target area to be tracked, and then use the optical flow method to calculate the optical flow vector of the pixel points in this area, so as to obtain the movement trajectory of the target person. The optical flow vector includes the optical flow direction and the optical flow speed.

[0073] Step S303: According to the optical flow vector, track the displacement trajectory points of the bounding box of the detected target person, and determine the movement trajectory of the target person according to the bounding box displacement; among them, the movement trajectory includes multiple bounding box displacement trajectory points; Among them, after tracking the bounding box displacement trajectory points of the target person, these trajectory points need to be connected to form the movement trajectory of the target person. The movement trajectory is the continuous movement path of the target person in the video image, which reflects the behavior pattern and movement characteristics of the target person. By analyzing the movement trajectory, it is possible to further determine whether the target person has abnormal behaviors, such as sudden acceleration, deceleration, or change of movement direction.

[0074] Step S304: According to the movement trajectory, perform dilation calculation on each bounding box displacement trajectory point to obtain the dilation neighborhood of each bounding box displacement trajectory point.

[0075] Among them, dilation calculation is a morphological operation. By expanding the target area, it can fill the holes or small gaps inside the target area, making the target area more complete and continuous. In the present invention, performing dilation calculation on each bounding box displacement trajectory point can obtain its dilation neighborhood, thereby further expanding the detection range of the target person and improving the accuracy and robustness of target detection.

[0076] Specifically, taking each bounding box displacement trajectory point as the center, setting a dilation radius, and then considering all pixel points within this radius as part of the target person, so as to obtain the dilated target area. The dilation neighborhood is the area between the dilated target area and the original bounding box. By analyzing the dilation neighborhood, it is possible to further determine the behavior pattern and movement characteristics of the target person, such as whether the target person collides with other objects or enters a prohibited area.

[0077] Step S305: Perform connected component connection processing on the dilation neighborhood of each bounding box displacement trajectory point to obtain the dilated movement area, and screen the valid areas in the dilated movement area through a density threshold to obtain the movement area of the target person.

[0078] Among them, connected component connection processing is an image processing technology that can connect adjacent pixel points or regions with similar attributes in the image to form a larger connected area.

[0079] In the present invention, performing connected component connection processing on the dilation neighborhood of each bounding box displacement trajectory point can connect the movement trajectories of the target person in the video image to form a complete movement area. However, due to the possible presence of noise or interference in the image, some invalid or false areas are included in the dilated movement area.

[0080] Therefore, it is necessary to remove these invalid areas through density threshold screening to obtain the true movement area of the target person.

[0081] Specifically, first, calculate the density of the dilated motion area, that is, the number of pixels per unit area. Then, regard the area with density lower than the preset threshold as an invalid area and remove it, so as to obtain the true motion area of the target person, which can further improve the accuracy and robustness of target detection and provide more reliable image information for subsequent abnormal behavior recognition. After obtaining the motion area of the target person, it can be further analyzed and processed, such as calculating feature parameters such as motion speed, motion direction, and motion trajectory, so as to achieve a comprehensive description and recognition of the behavior pattern and motion characteristics of the target person. These feature parameters can be used as input data for subsequent abnormal behavior recognition to determine whether the target person has abnormal behavior.

[0082] In some embodiments, according to the motion trajectory, perform dilation calculation on each bounding box displacement trajectory point to obtain the dilation neighborhood of each bounding box displacement trajectory point, including: Step S3041: According to each bounding box displacement trajectory point, convert the bounding box displacement trajectory point to a spatio-temporal point in the spatio-temporal domain and determine the density of the spatio-temporal point.

[0083] Among them, let the motion trajectory of the target person be: In the formula, is the position in the rectangular coordinate, is the timestamp, and N is the number of trajectory points.

[0084] Calculate the influence range of the trajectory point in the spatio-temporal domain through the spatio-temporal joint Gaussian kernel function: In the formula, is the influence range in the spatio-temporal domain, , is the spatial dimension bandwidth parameter, which controls the attenuation speed of the density by the spatial distance, is the time dimension bandwidth parameter, is the abscissa difference between the i-th trajectory point and the (i - 1)-th trajectory point, is the ordinate difference between the i-th trajectory point and the (i - 1)-th trajectory point, is the time difference between the i-th trajectory point and the (i - 1)-th trajectory point.

[0085] Among them, this kernel function smooths and diffuses the spatio-temporal influence of the trajectory point through the Gaussian distribution. The influence intensity decays in the spatial dimension with as the scale, and different weights are given according to the time interval in the time dimension, ensuring that the trajectory points closer to the current moment contribute more to the density calculation, so as to accurately characterize the aggregation characteristics of the trajectory point in the spatio-temporal domain.

[0086] For any spatio-temporal point , its density estimate is: In the formula, is the spatio-temporal point density.

[0087] Step S3042: Determine the dilation amplitude of the bounding box displacement trajectory points according to the density of the spatio-temporal points.

[0088] Among them, the dilation amplitude refers to the size of the expanded area when performing a dilation operation on the bounding box displacement trajectory points. In the present invention, determining the dilation amplitude according to the density of the spatio-temporal points can achieve adaptive dilation processing for different density regions. Specifically, a region with a higher density indicates that the target person appears more frequently in this region, and there may be more complex motion patterns or behavioral characteristics. Therefore, a larger dilation amplitude is required to fully cover the motion range of the target person; while a region with a lower density indicates that the target person appears less frequently in this region, and the motion pattern is relatively simple. Therefore, a smaller dilation amplitude can be used to avoid inaccurate target detection caused by excessive expansion.

[0089] The dilation amplitude is calculated as: In the formula, is the dilation amplitude, is the time weight sharpening parameter, which controls the weight increase rate of the recent trajectory; is the latest timestamp in the trajectory, ensuring that the closer the trajectory, the greater the impact on dilation.

[0090] Step S3043: Determine the dilation radius of the bounding box displacement trajectory points according to the dilation amplitude and the velocity vector in the optical flow vector.

[0091] Among them, the dilation radius refers to the radius of the circular area expanded with this point as the center when performing a dilation operation on the bounding box displacement trajectory points. In the present invention, determining the dilation radius according to the dilation amplitude and the velocity vector in the optical flow vector can achieve an accurate description and expansion of the target person's motion trajectory. Specifically, the size of the dilation radius should match the motion speed and direction of the target person to ensure that the expanded motion area can accurately cover the actual motion range of the target person. At the same time, since the motion speed and direction of the target person may change over time, it is necessary to dynamically adjust the dilation radius according to the velocity vector in the optical flow vector to adapt to the motion changes of the target person.

[0092] Among them, the dilation radius of each trajectory point is jointly determined by the density and the historical speed: In the formula, is the trajectory point dilation radius, is the basic dilation radius (such as 1.5 times the default target average size), , are the density and velocity weight coefficients respectively, is the velocity vector in the optical flow vector, which is estimated by the aforementioned optical flow method.

[0093] Step S3044: Determine the dilation neighborhood corresponding to the bounding box displacement trajectory point according to the dilation radius; wherein, the dilation neighborhood is a circular area drawn with the bounding box displacement trajectory point as the center and the dilation radius.

[0094] Among them, since the person is in dynamic motion, the potential motion range of the target person in the video image is determined through the dilation neighborhood for subsequent abnormal behavior recognition.

[0095] In some embodiments, the behavior features of the target person in the updated field-of-view complementary region are extracted based on a convolutional network with multi-head attention, including: Step 601: Input the target person image in the updated field-of-view complementary region into the convolutional network model with multi-head attention; wherein, the convolutional network model with multi-head attention includes multiple self-attention heads including a convolutional network; Step 602: Calculate the correlation weights between different trajectory points in the target person image through each self-attention head to obtain a correlation weight matrix; Among them, the self-attention head generates a correlation weight matrix by calculating the correlation scores between any two positions in the input data.

[0096] In the present invention, each self-attention head calculates the correlation weights between different positions in the target person image respectively, so as to obtain a correlation weight matrix. These matrices reflect the degree of association between different positions in the target person image.

[0097] Specifically, each self-attention head calculates the dot product similarity between any two trajectory point positions according to the trajectory point representation of the target person image, and normalizes it to a correlation score through the softmax function. Then, these scores are used as weights, and a correlation weight matrix is constructed based on the weights corresponding to each trajectory point.

[0098] Step 603: Extract the features of the target person at each trajectory point from the target person image through the convolutional network, and perform weighted summation on all features in the target person image according to the correlation weight matrix to obtain a weighted feature map.

[0099] Among them, the convolutional network can be a 3D convolutional network (Two-Stream 3D CNN), including a spatial stream network and a temporal stream network. The spatial stream network extracts the static pose features of the target person at each trajectory point, and the temporal stream network extracts the action dynamic features of the target person at each trajectory point.

[0100] Among them, the weighted feature map reflects the importance of features at different positions in the target person's image. In the present invention, by performing weighted summation on the behavioral features in the target person's image according to the correlation weight matrix, key regions related to the behavioral features of the target person can be highlighted, and irrelevant background information can be suppressed, thereby improving the accuracy and robustness of behavioral feature extraction.

[0101] Specifically, for each position in the target person's image, according to its corresponding correlation weight, the feature value at that position is multiplied by the corresponding weight, and the feature values at all positions are summed to obtain the weighted feature map. The weighted feature map not only contains the features of the target person at each trajectory point but also incorporates the correlation information between the positions of the target person at different trajectory points, providing a richer and more accurate feature representation for subsequent behavioral anomaly recognition. By analyzing and processing the weighted feature map, the behavioral features of the target person, such as posture and movement, can be further extracted, thereby realizing a comprehensive description and recognition of the behavioral patterns of the target person.

[0102] Step 604: Stitch the weighted feature maps and process them through a fully connected layer to obtain the behavioral feature vector of the target person; wherein, the behavioral feature vector is used to characterize the behavioral pattern of the target person.

[0103] Among them, the fully connected layer is a neural network layer that can map the input feature map to an output vector of a fixed size. In the present invention, by stitching the weighted feature maps and processing them through a fully connected layer, the behavioral feature vector of the target person can be obtained. This vector contains the behavioral feature information of the target person, such as posture and movement, and can be used for subsequent behavioral anomaly recognition.

[0104] In some embodiments, the method further includes: Step S11: Collect historical behavioral feature samples of the workers in the tunnel construction video and the corresponding behavioral anomaly categories of the historical behavioral feature samples; Step S12: Construct a training set and a test set according to the historical behavioral feature samples and the corresponding behavioral anomaly categories of the historical behavioral feature samples; Step S13: Construct an initial behavioral prediction classification model, and the initial behavioral prediction classification model uses a deep learning algorithm; Step S14: Train the behavioral prediction classification model using the training set to obtain the behavioral prediction classification model; Step S15: Test the behavioral prediction classification model using the test set, and optimize the network parameters of the behavioral prediction classification model according to the test results to obtain the optimized behavioral prediction classification model.

[0105] Among them, the behavior prediction and classification model is a model used to classify and predict the behaviors of target personnel in tunnel construction videos. In the present invention, a deep learning algorithm is adopted to construct an initial behavior prediction and classification model because the deep learning algorithm can automatically learn feature representations from a large amount of data and achieve accurate classification and prediction of the behaviors of target personnel.

[0106] Specifically, the deep learning algorithm can learn from historical behavior feature samples in the training set, automatically extract feature information related to the behaviors of target personnel, and construct a classifier capable of distinguishing different behavior anomaly categories. During the training process, the deep learning algorithm will continuously adjust the parameters of the model to minimize the error between the prediction result and the actual result, thereby improving the classification and prediction performance of the model.

[0107] As Figure 2 shown, an embodiment of the present application provides a real-time abnormal behavior recognition system for tunnel construction videos. A wide-angle main camera and a telephoto auxiliary camera are set in any one of multiple target construction areas of the tunnel, including: An image acquisition module 100, configured to respectively collect wide-angle video images and telephoto video images in the target construction area through the wide-angle main camera and the telephoto auxiliary camera; A field of view combination module 200, configured to determine a wide-angle field of view and a telephoto field of view according to the wide-angle video image and the telephoto video image, and perform field of view combination on the wide-angle field of view and the telephoto field of view to obtain a field of view complementary area; A moving area determination module 300, configured to perform target detection on target personnel in the telephoto video image and track the target personnel detected by the target detection to obtain the moving area of the target personnel; An area judgment module 400, configured to judge whether the moving area completely falls within the field of view complementary area; A blind area supplementary shooting module 500, configured to, when the moving area does not completely fall within the field of view complementary area, perform supplementary shooting on the moving blind area through the telephoto auxiliary camera and update the supplemented moving blind area to the field of view complementary area; wherein, the moving blind area is a part of the moving area that does not fall within the field of view complementary area; A feature extraction module 600, configured to extract the behavior features of target personnel in the updated field of view complementary area based on a convolutional network with multi-head attention; A behavior anomaly recognition module 700, configured to input the behavior features into a pre-trained behavior prediction and classification model to identify whether the behavior of the target personnel is abnormal.

[0108] In some embodiments, the field of view combination module 200 is configured to: Establish a wide-angle field of view in the polar coordinate system with the optical center as the origin according to the wide-angle video image; Based on the long - focal - length video image, with the optical center as the origin, establish a long - focal - length field of view in a rectangular coordinate system; Determine the overlapping area and non - overlapping area according to the wide - angle field of view and the long - focal - length field of view; According to the area ratio of the overlapping area and the non - overlapping area, adjust the focal length parameter of the wide - angle main camera, update the wide - angle field of view according to the adjusted focal length parameter, and stop when the area ratio is greater than a preset area ratio threshold to obtain the updated wide - angle field of view; Redetermine the overlapping area and the non - overlapping area according to the updated wide - angle field of view, and integrate the overlapping area and the non - overlapping area into a field - of - view complementary area.

[0109] In some embodiments, the non - overlapping area includes a wide - angle exclusive area and a long - focal - length exclusive area; Determine the overlapping area and the non - overlapping area according to the wide - angle field of view and the long - focal - length field of view, including: Convert the updated wide - angle field of view to a rectangular coordinate system, and fuse the overlapping area of the wide - angle field of view and the long - focal - length field of view in the rectangular coordinate system by weighted average to obtain the fused overlapping area; Map each pixel point in the wide - angle exclusive area to the rectangular coordinate system through the initial mapping matrix to obtain the first pixel point; and reflect each pixel point in the long - focal - length exclusive area to the polar coordinate system through the initial mapping matrix to obtain the second pixel point; Perform weighted fusion on the first pixel point and the second pixel point through the initial fusion weight to obtain the initially fused non - overlapping area; Construct a global energy constraint function according to the initially fused non - overlapping area; wherein, the global energy constraint function is used to constrain the smoothness and consistency in the image fusion process; Through optimization and solution of the global energy constraint function, obtain the fusion weight and the mapping matrix under the minimization of the global energy constraint function; Optimize the initially fused non - overlapping area according to the fusion weight and the mapping matrix to obtain the optimized non - overlapping area.

[0110] In some embodiments, the motion area determination module 300 is used for: Adopt the YOLO object detection algorithm to perform object detection on the target personnel in the long - focal - length video image, and obtain the position information and bounding box of the target personnel; According to the position information and the bounding box, use the optical flow method to calculate the optical flow vector of the target personnel between adjacent frames in the long - focal - length video image; According to the optical flow vector, track the bounding box displacement trajectory points of the detected target personnel, and determine the motion trajectory of the target personnel according to the bounding box displacement; wherein, the motion trajectory includes multiple bounding box displacement trajectory points; According to the motion trajectory, perform dilation calculation on each bounding box displacement trajectory point to obtain the dilation neighborhood of each bounding box displacement trajectory point; Perform connected component connection processing on the dilation neighborhood of each bounding box displacement trajectory point to obtain the dilated motion region, and filter the valid regions in the dilated motion region through a density threshold to obtain the motion region of the target person.

[0111] In some embodiments, according to the motion trajectory, performing dilation calculation on each bounding box displacement trajectory point to obtain the dilation neighborhood of each bounding box displacement trajectory point includes: According to each bounding box displacement trajectory point, convert the bounding box displacement trajectory point to a spatio-temporal point in the spatio-temporal domain and determine the density of the spatio-temporal point; Determine the dilation amplitude of the bounding box displacement trajectory point according to the density of the spatio-temporal point; Determine the dilation radius of the bounding box displacement trajectory point according to the dilation amplitude and the velocity vector in the optical flow vector; Determine the dilation neighborhood corresponding to the bounding box displacement trajectory point according to the dilation radius; wherein, the dilation neighborhood is a circular region drawn with the bounding box displacement trajectory point as the center and the dilation radius.

[0112] In some embodiments, the feature extraction module 600 is used for: Input the target person image in the updated field-of-view complementary region into the convolutional network model based on multi-head attention; wherein, the convolutional network model based on multi-head attention includes multiple self-attention heads including a convolutional network; Calculate the correlation weights between different trajectory points in the target person image through each self-attention head respectively to obtain a correlation weight matrix; Extract the features of the target person at each trajectory point from the target person image through the convolutional network, and perform weighted summation on all the features in the target person image according to the correlation weight matrix to obtain a weighted feature map; Stitch the weighted feature maps and process them through a fully connected layer to obtain the behavior feature vector of the target person; wherein, the behavior feature vector is used to characterize the behavior pattern of the target person.

[0113] In some embodiments, the system further includes: a model training module, which is used for: Collect historical behavior feature samples of the workers in the tunnel construction video and the corresponding behavior anomaly categories of the historical behavior feature samples; Construct a training set and a test set according to the historical behavior feature samples and the corresponding behavior anomaly categories of the historical behavior feature samples; Construct an initial behavior prediction classification model, and the initial behavior prediction classification model adopts a deep learning algorithm; Train the behavior prediction classification model using the training set to obtain the behavior prediction classification model; Test the behavior prediction classification model using the test set, and optimize the network parameters of the behavior prediction classification model according to the test results to obtain the optimized behavior prediction classification model.

[0114] As Figure 3 shown, an embodiment of the present application provides an electronic device. The electronic device 10 includes a memory 20 and a processor 30. A computer program is stored in the memory 20. When the computer program is executed by the processor 30, the processor 30 is caused to execute the steps of the method for real-time abnormal behavior recognition of tunnel construction videos in the above embodiments.

[0115] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the method for real-time abnormal behavior recognition of tunnel construction videos in the above embodiments are implemented.

[0116] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, electronic devices, and computer storage media can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0117] It should be noted that the user information (including but not limited to user images, user portrait information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0118] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0119] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0120] In several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0121] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0123] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the methods described in the various embodiments of the present invention through a computer device (which may be a personal computer, a server, or a network device, etc.). The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs.

[0124] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A real-time abnormal behavior recognition method for tunnel construction videos, characterized in that, Set a wide-angle main camera and a telephoto auxiliary camera in any one of multiple target construction areas in the tunnel, including: Collect the wide-angle video image and the telephoto video image in the target construction area through the wide-angle main camera and the telephoto auxiliary camera respectively; Determine the wide-angle field of view and the telephoto field of view according to the wide-angle video image and the telephoto video image, and perform field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary area; Perform target detection on the target personnel in the telephoto video image, and track the target personnel detected by the target detection to obtain the movement area of the target personnel; Judge whether the movement area completely falls within the field-of-view complementary area; When the movement area does not completely fall within the field-of-view complementary area, the telephoto auxiliary camera is used to supplement the shooting of the movement blind area, and the supplemented movement blind area is updated to the field-of-view complementary area; wherein, the movement blind area is the partial movement area where the movement area does not fall within the field-of-view complementary area; Extract the behavior features of the target personnel in the updated field-of-view complementary area based on the convolutional network with multi-head attention; Input the behavior features into a pre-trained behavior prediction classification model to identify whether the behavior of the target personnel is abnormal.

2. The real-time abnormal behavior recognition method for tunnel construction videos according to claim 1, characterized in that The step of determining the wide-angle field of view and the telephoto field of view according to the wide-angle video image and the telephoto video image, and performing field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary area includes: According to the wide-angle video image, with the optical center as the origin, establish a wide-angle field of view in the polar coordinate system; According to the telephoto video image, with the optical center as the origin, establish a telephoto field of view in the rectangular coordinate system; Determine the overlapping area and the non-overlapping area according to the wide-angle field of view and the telephoto field of view; Adjust the focal length parameter of the wide-angle main camera according to the area ratio of the overlapping area and the non-overlapping area, update the wide-angle field of view according to the adjusted focal length parameter, and stop when the area ratio is greater than a preset area ratio threshold to obtain an updated wide-angle field of view; Redetermine the overlapping area and the non-overlapping area according to the updated wide-angle field of view, and integrate the overlapping area and the non-overlapping area into the field-of-view complementary area.

3. The real-time abnormal behavior recognition method for tunnel construction videos according to claim 2, wherein The non-overlapping area includes a wide-angle exclusive area and a telephoto exclusive area; The step of determining the overlapping area and the non-overlapping area according to the wide-angle field of view and the telephoto field of view includes: Convert the updated wide-angle field of view to the rectangular coordinate system, and fuse the overlapping area of the wide-angle field of view and the telephoto field of view in the rectangular coordinate system through weighted average to obtain a fused overlapping area; Map each pixel point in the wide-angle exclusive area to the rectangular coordinate system through the initial mapping matrix to obtain the first pixel point; and map each pixel point in the telephoto exclusive area to the polar coordinate system through the initial mapping matrix to obtain the second pixel point; Perform weighted fusion on the first pixel point and the second pixel point through the initial fusion weight to obtain an initially fused non-overlapping area; Construct a global energy constraint function according to the initially fused non-overlapping regions; wherein, the global energy constraint function is used to constrain the smoothness and consistency during the image fusion process; By optimizing and solving the global energy constraint function, obtain the fusion weights and mapping matrix under the minimization of the global energy constraint function; Optimize the initially fused non-overlapping regions according to the fusion weights and the mapping matrix to obtain the optimized non-overlapping regions.

4. The method for real-time abnormal behavior recognition of tunnel construction videos according to claim 1, wherein, The target detection of the target person in the long-focus video image and the tracking of the detected target person to obtain the motion region of the target person include: Use the YOLO target detection algorithm to perform target detection on the target person in the long-focus video image, and obtain the position information and bounding box of the target person; According to the position information and the bounding box, use the optical flow method to calculate the optical flow vector of the target person between adjacent frames in the long-focus video image; According to the optical flow vector, track the displacement trajectory points of the bounding box of the detected target person, and determine the motion trajectory of the target person according to the displacement of the bounding box; wherein, the motion trajectory includes multiple displacement trajectory points of the bounding box; According to the motion trajectory, perform dilation calculation on each displacement trajectory point of the bounding box to obtain the dilation neighborhood of each displacement trajectory point of the bounding box; Perform connected domain connection processing on the dilation neighborhood of each displacement trajectory point of the bounding box to obtain the dilated motion region, and screen the valid region in the dilated motion region through a density threshold to obtain the motion region of the target person.

5. The real-time abnormal behavior recognition method for tunnel construction videos according to claim 4, characterized in that The performing dilation calculation on each displacement trajectory point of the bounding box according to the motion trajectory to obtain the dilation neighborhood of each displacement trajectory point of the bounding box includes: According to each displacement trajectory point of the bounding box, convert the displacement trajectory point of the bounding box to a spatio-temporal point in the spatio-temporal domain, and determine the density of the spatio-temporal point; Determine the dilation amplitude of the displacement trajectory point of the bounding box according to the density of the spatio-temporal point; Determine the dilation radius of the displacement trajectory point of the bounding box according to the dilation amplitude and the velocity vector in the optical flow vector; Determine the dilation neighborhood corresponding to the displacement trajectory point of the bounding box according to the dilation radius; wherein, the dilation neighborhood is a circular region drawn with the displacement trajectory point of the bounding box as the center and the dilation radius.

6. The real-time abnormal behavior recognition method for tunnel construction videos according to claim 1, wherein The extracting the behavior features of the target person in the updated field-of-view complementary region by the convolutional network based on multi-head attention includes: Input the target person image in the updated field-of-view complementary region into the convolutional network model based on multi-head attention; wherein, the convolutional network model based on multi-head attention includes multiple self-attention heads including a convolutional network; Calculate the correlation weights between different trajectory points in the target person image through each self-attention head respectively to obtain a correlation weight matrix; Extract the features of the target person at each trajectory point from the target person image through the convolutional network, and perform weighted summation on all the features in the target person image according to the correlation weight matrix to obtain a weighted feature map; Stitch the weighted feature maps and process them through a fully-connected layer to obtain the behavioral feature vector of the target person; wherein, the behavioral feature vector is used to characterize the behavioral pattern of the target person.

7. The real-time abnormal behavior recognition method for tunnel construction videos according to claim 1, characterized in that It further includes: Collect historical behavioral feature samples of the operators in the tunnel construction video and the corresponding behavioral anomaly categories of the historical behavioral feature samples; Construct a training set and a test set according to the historical behavioral feature samples and the corresponding behavioral anomaly categories of the historical behavioral feature samples; Construct an initial behavioral prediction classification model, and the initial behavioral prediction classification model uses a deep learning algorithm; Use the training set to train the behavioral prediction classification model to obtain a behavioral prediction classification model; Use the test set to test the behavioral prediction classification model, and optimize the network parameters of the behavioral prediction classification model according to the test results to obtain an optimized behavioral prediction classification model.

8. A real-time abnormal behavior recognition system for tunnel construction videos, characterized in that, Set a wide-angle main camera and a telephoto auxiliary camera in any one of multiple target construction areas of the tunnel, including: An image acquisition module, configured to collect wide-angle video images and telephoto video images in the target construction area through the wide-angle main camera and the telephoto auxiliary camera respectively; A field-of-view combination module, configured to determine a wide-angle field of view and a telephoto field of view according to the wide-angle video image and the telephoto video image, and perform field-of-view combination on the wide-angle field of view and the telephoto field of view to obtain a field-of-view complementary area; A moving area determination module, configured to perform target detection on the target person in the telephoto video image and track the detected target person to obtain the moving area of the target person; An area judgment module, configured to judge whether the moving area completely falls within the field-of-view complementary area; A blind area supplementary shooting module, configured to, when the moving area does not completely fall within the field-of-view complementary area, perform supplementary shooting on the moving blind area through the telephoto auxiliary camera and update the supplemented moving blind area to the field-of-view complementary area; wherein, the moving blind area is the partial moving area where the moving area does not fall within the field-of-view complementary area; A feature extraction module, configured to extract the behavioral features of the target person in the updated field-of-view complementary area based on a convolutional network with multi-head attention; A behavioral anomaly recognition module, configured to input the behavioral features into a pre-trained behavioral prediction classification model to identify whether the behavior of the target person is abnormal.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the tunnel construction video real-time abnormal behavior recognition method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, the steps of the tunnel construction video real-time abnormal behavior recognition method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Tunnel operation safety incident detection system based on video recognition

    CN105338304A

  • Abnormity detection method and related equipment and device

    CN111372043A

  • Method for detecting staying behavior of personnel in public place

    CN118334743A

Cited By

  • Examination room abnormal behavior detection method and system based on image recognition

    CN121121655A