Task pushing method and device based on foundation pit and storage medium
By collecting image data in construction sites, identifying the relationship between workers and foundation pits, and using the operation classification network to identify the operation modes, the inspection tasks are solved, and the problems of high manual monitoring cost and prone to errors and omissions in foundation pit safety management in construction sites are achieved, and more efficient and accurate foundation pit safety management is achieved.
Patent Information
- Application Number
- CN202510006533.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In construction sites, the safety management of foundation pits increases the risk caused by the foundation pit due to the high cost of manual monitoring and the prone to fatigue.
By collecting multi-frame image data in construction sites, performing semantic segmentation and distance statistics, identifying the relationship between the operator and the foundation pit, loading the job classification network, identifying the operation mode, and building inspection tasks, which are automatically pushed to the management personnel for execution.
It reduces fatigue errors and omissions caused by long-term browsing of surveillance videos, improves the automation and accuracy of foundation pit safety management, promptly pushes inspection tasks, and reduces the risks brought by foundation pits.
Smart Images

Figure CN119992447A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of natural language processing, and in particular, relate to a task push method, device, and storage medium based on a foundation pit. Background Art
[0002] With the continuous improvement of economy and technology, higher requirements are placed on the quality of the entire engineering construction. Among them, foundation pits face certain risks, such as collapse, burial, falling from heights, mechanical injuries and electric shock, and their safety management is receiving more and more attention.
[0003] Management personnel are usually arranged at construction sites to view surveillance videos of the construction sites in real time. When workers are found approaching foundation pits, they will be warned on site.
[0004] This model has high labor costs, and managers become tired from browsing surveillance videos for a long time, which can easily lead to errors and omissions, thus increasing the risks posed by the foundation pit. Summary of the invention
[0005] In view of this, an embodiment of the present application provides a task push method, device and storage medium based on a foundation pit, so as to reduce the risks brought by the foundation pit.
[0006] A first aspect of an embodiment of the present application provides a task pushing method based on a foundation pit, comprising:
[0007] Collect multiple frames of image data at a construction site;
[0008] Performing semantic segmentation on the image data to obtain a foundation pit and a plurality of workers;
[0009] Counting the distance between each of the workers and the edge of the foundation pit in the image data;
[0010] The coordinates of the same operator in multiple frames of the image data are combined into a first operation sequence, and the multiple distances are combined into a second operation sequence;
[0011] For a plurality of the operators, a plurality of the first operation sequences are combined into a first operation matrix, and a plurality of the second operation sequences are combined into a second operation matrix;
[0012] Load the job classification network;
[0013] Inputting the first operation matrix and the second operation matrix into the operation classification network to identify the operation modes of a plurality of the operators;
[0014] An inspection task is constructed according to the operation mode, and the inspection task is pushed to the management personnel for execution.
[0015] Optionally, the job classification network includes a first encoder, a second encoder, a third encoder, a fourth encoder, a feature interaction module, a feature fusion module and a head structure;
[0016] The step of inputting the first operation matrix and the second operation matrix into the operation classification network to identify the operation modes of the plurality of operators comprises:
[0017] Inputting the first operation matrix into the first encoder for encoding to obtain a first operation feature;
[0018] Inputting the second operation matrix into the second encoder for encoding to obtain a second operation feature;
[0019] Inputting the first operation feature and the second operation feature into the interactive part of features in the feature interaction module to obtain a third operation feature and a fourth operation feature;
[0020] Inputting the third operation feature into the third encoder for encoding to obtain a fifth operation feature;
[0021] Inputting the fourth operation characteristic into the fourth encoder for encoding to obtain a sixth operation characteristic;
[0022] Inputting the fifth operation feature and the sixth operation feature into the feature fusion module and fusing them into a seventh operation feature;
[0023] The seventh operation feature is input into the head structure to classify the operation modes of the plurality of operators.
[0024] Optionally, the first encoder includes three first convolutional layers, and the second encoder includes three second convolutional layers;
[0025] The step of inputting the first operation matrix into the first encoder for encoding to obtain a first operation feature includes:
[0026] Inputting the first operation matrix into the three first convolutional layers in sequence to perform convolution operations to obtain first operation features;
[0027] The step of inputting the second operation matrix into the second encoder for encoding to obtain a second operation feature includes:
[0028] The second operation matrix is sequentially input into three second convolutional layers to perform convolution operations to obtain second operation features.
[0029] Optionally, the step of inputting the first operation feature and the second operation feature into the feature interaction module to interact with each other to obtain a third operation feature and a fourth operation feature comprises:
[0030] Performing an upsampling operation on the first job feature to obtain a first intermediate feature;
[0031] Adding the first intermediate feature and the second operation feature to obtain a third operation feature;
[0032] Performing a downsampling operation on the third operation feature to obtain a second intermediate feature;
[0033] The second intermediate feature is added to the first intermediate feature to obtain a fourth operation feature.
[0034] Optionally, the third encoder includes three third convolutional layers, and the fourth encoder includes three fourth convolutional layers;
[0035] The step of inputting the third operation feature into the third encoder for encoding to obtain the fifth operation feature comprises:
[0036] Inputting the third operation feature into the three third convolutional layers in sequence to perform convolution operations, thereby obtaining a fifth operation feature;
[0037] The step of inputting the fourth operation feature into the fourth encoder for encoding to obtain a sixth operation feature comprises:
[0038] The fourth operation feature is sequentially input into the three fourth convolutional layers to perform a convolution operation to obtain a sixth operation feature.
[0039] Optionally, the feature fusion module includes a fifth convolutional layer and a sixth convolutional layer;
[0040] The step of inputting the fifth operation feature and the sixth operation feature into the feature fusion module and fusing them into a seventh operation feature includes:
[0041] Inputting the fifth operation feature into the fifth convolutional layer to perform a convolutional layer operation to obtain a third intermediate feature;
[0042] splicing the third intermediate feature and the sixth operating feature into a fourth intermediate feature;
[0043] The fourth intermediate feature is input into the sixth convolutional layer to perform a convolutional layer operation to obtain a seventh operation feature.
[0044] Optionally, constructing an inspection task according to the operation mode includes:
[0045] Generate job description information using the first job matrix and the second job matrix;
[0046] Combining the operation mode and the operation description information into target operation information;
[0047] Searching for task record information similar to the target operation information in a preset task library; the task record information is information recorded when the inspection task was historically executed;
[0048] An inspection task is constructed based on the task record information and the target operation information.
[0049] Optionally, constructing the inspection task according to the task record information and the target operation information includes:
[0050] Extracting task key points from the task record information;
[0051] Writing the task key points and the target operation information into a preset template to obtain instructions;
[0052] The instructions are input into a large language model to construct an inspection task.
[0053] A second aspect of an embodiment of the present application provides a task pushing device based on a foundation pit, comprising:
[0054] An image data acquisition module, used for acquiring multiple frames of image data at a construction site;
[0055] A semantic segmentation module, used to perform semantic segmentation on the image data to obtain a foundation pit and a plurality of workers;
[0056] A distance detection module, used for counting the distance between each of the workers and the edge of the foundation pit in the image data;
[0057] A sequence composition module, used for composing the coordinates of the same operator in multiple frames of the image data into a first operation sequence, and composing multiple distances into a second operation sequence;
[0058] A matrix composition module, for forming a first operation matrix with a plurality of the first operation sequences and a second operation matrix with a plurality of the second operation sequences for a plurality of the operators;
[0059] A job classification network loading module is used to load the job classification network;
[0060] A work mode identification module, used for inputting the first work matrix and the second work matrix into the work classification network to identify work modes of a plurality of the workers;
[0061] The inspection task processing module is used to construct the inspection task according to the operation mode and push the inspection task to the management personnel for execution.
[0062] Optionally, the job classification network includes a first encoder, a second encoder, a third encoder, a fourth encoder, a feature interaction module, a feature fusion module and a head structure;
[0063] The operation mode recognition module is also used for:
[0064] Inputting the first operation matrix into the first encoder for encoding to obtain a first operation feature;
[0065] Inputting the second operation matrix into the second encoder for encoding to obtain a second operation feature;
[0066] Inputting the first operation feature and the second operation feature into the interactive part of features in the feature interaction module to obtain a third operation feature and a fourth operation feature;
[0067] Inputting the third operation feature into the third encoder for encoding to obtain a fifth operation feature;
[0068] Inputting the fourth operation characteristic into the fourth encoder for encoding to obtain a sixth operation characteristic;
[0069] Inputting the fifth operation feature and the sixth operation feature into the feature fusion module and fusing them into a seventh operation feature;
[0070] The seventh operation feature is input into the head structure to classify the operation modes of the plurality of operators.
[0071] Optionally, the first encoder includes three first convolutional layers, and the second encoder includes three second convolutional layers;
[0072] The operation mode recognition module is also used for:
[0073] Inputting the first operation matrix into the three first convolutional layers in sequence to perform convolution operations to obtain first operation features;
[0074] The operation mode recognition module is also used for:
[0075] The second operation matrix is sequentially input into three second convolutional layers to perform convolution operations to obtain second operation features.
[0076] Optionally, the operation mode recognition module is further used for:
[0077] Performing an upsampling operation on the first job feature to obtain a first intermediate feature;
[0078] Adding the first intermediate feature and the second operation feature to obtain a third operation feature;
[0079] Performing a downsampling operation on the third operation feature to obtain a second intermediate feature;
[0080] The second intermediate feature is added to the first intermediate feature to obtain a fourth operation feature.
[0081] In one embodiment of the present application, the third encoder includes three third convolutional layers, and the fourth encoder includes three fourth convolutional layers;
[0082] The operation mode recognition module is also used for:
[0083] Inputting the third operation feature into the three third convolutional layers in sequence to perform convolution operations, thereby obtaining a fifth operation feature;
[0084] The step of inputting the fourth operation feature into the fourth encoder for encoding to obtain a sixth operation feature comprises:
[0085] The fourth operation feature is sequentially input into the three fourth convolutional layers to perform a convolution operation to obtain a sixth operation feature.
[0086] Optionally, the feature fusion module includes a fifth convolutional layer and a sixth convolutional layer;
[0087] The operation mode recognition module is also used for:
[0088] Inputting the fifth operation feature into the fifth convolutional layer to perform a convolutional layer operation to obtain a third intermediate feature;
[0089] splicing the third intermediate feature and the sixth operating feature into a fourth intermediate feature;
[0090] The fourth intermediate feature is input into the sixth convolutional layer to perform a convolutional layer operation to obtain a seventh operation feature.
[0091] Optionally, the inspection task processing module is further used to:
[0092] Generate job description information using the first job matrix and the second job matrix;
[0093] Combining the operation mode and the operation description information into target operation information;
[0094] Searching for task record information similar to the target operation information in a preset task library; the task record information is information recorded when the inspection task was historically executed;
[0095] An inspection task is constructed based on the task record information and the target operation information.
[0096] Optionally, the inspection task processing module is further used to:
[0097] Extracting task key points from the task record information;
[0098] Writing the task key points and the target operation information into a preset template to obtain instructions;
[0099] The instructions are input into a large language model to construct an inspection task.
[0100] A third aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the task push method based on the foundation pit as described in the first aspect above is implemented.
[0101] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the task push method based on the foundation pit as described in the first aspect above.
[0102] A fifth aspect of an embodiment of the present application provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the task push method based on the foundation pit described in the first aspect.
[0103] In this embodiment, multiple frames of image data are collected at the construction site; semantic segmentation is performed on the image data to obtain the foundation pit and multiple workers; the distance between each worker and the edge of the foundation pit is counted in the image data; the coordinates of the same worker in the multiple frames of image data are combined into a first operation sequence, and multiple distances are combined into a second operation sequence; for multiple workers, multiple first operation sequences are combined into a first operation matrix, and multiple second operation sequences are combined into a second operation matrix; the operation classification network is loaded; the first operation matrix and the second operation matrix are input into the operation classification network to identify the operation mode of multiple workers; inspection tasks are constructed based on the operation mode, and the inspection tasks are pushed to the management personnel for execution. In this embodiment, the operation mode is identified based on the movement mode of the workers relative to the foundation pit in computer vision, and the inspection tasks are constructed based on the operation mode in natural language processing. It has a high degree of automation and high accuracy, effectively reduces the errors and omissions caused by fatigue caused by long-term browsing of monitoring videos, pushes the inspection tasks to the management personnel in real time, and conducts inspections on the foundation pit site in a timely manner, reduces the risks brought by the foundation pit, and ensures the safety of construction operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0105] Figure 1 It is a schematic diagram of a task pushing method based on a foundation pit provided in an embodiment of the present application;
[0106] Figure 2 is a schematic diagram of a job classification network provided in an embodiment of the present application;
[0107] Figure 3 It is a schematic diagram of a task pushing device based on a foundation pit provided in an embodiment of the present application;
[0108] Figure 4 It is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0109] In the following description, specific details such as specific system structures, technologies, etc. are proposed for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from hindering the description of the present application.
[0110] The technical solution of the present application is described below through specific embodiments.
[0111] Reference Figure 1 , shows a schematic diagram of a task pushing method based on a foundation pit provided in an embodiment of the present application, which may specifically include the following steps:
[0112] Step 101: Collect multiple frames of image data at a construction site.
[0113] In this embodiment, one or more cameras may be deployed on-site at the construction site, and each camera may collect video data facing the construction site, wherein the video data includes multiple frames of image data.
[0114] Step 102: semantically segment the image data to obtain a foundation pit and multiple workers.
[0115] In this embodiment, semantic segmentation networks can be pre-built and trained, such as FCN (Fully Convolution Networks), U-Net (U-net), SegNet (Segmentation Networks), DeepLab (Deep Convolutional Nets), etc.
[0116] The image data of each frame is input into the semantic segmentation network for semantic segmentation to obtain the foundation pit (pixel points) and multiple workers (pixel points).
[0117] Step 103: Count the distances between each worker and the edge of the foundation pit in the image data.
[0118] In each frame of image data, projection or other methods may be used to count the shortest distance (pixel distance) between each worker (pixel point) and the edge of the foundation pit (pixel point).
[0119] Step 104: The coordinates of the same operator in multiple frames of image data are combined into a first operation sequence, and multiple distances are combined into a second operation sequence.
[0120] In this embodiment, algorithms such as KCF (Kernelized Correlation Filter) can be used to track each operator (pixel point) in multiple frames of image data.
[0121] For the same operator (pixel point), on the one hand, its multiple coordinates in the multi-frame image data are organized into a first operation sequence in chronological order, and on the other hand, its multiple distances in the multi-frame image data are organized into a second operation sequence in chronological order.
[0122] The first operation sequence and the second operation sequence reflect the movement rules of a single operator during operation, and can reflect the characteristics of their operation mode to a certain extent.
[0123] Step 105 : for multiple operators, multiple first operation sequences are combined into a first operation matrix, and multiple second operation sequences are combined into a second operation matrix.
[0124] In this embodiment, a plurality of operators may be configured in an order, such as from left to right, from top to bottom, etc., and a plurality of first operation sequences may be organized into a first operation matrix according to the order of the operators, and a plurality of second operation sequences may be organized into a second operation matrix according to the order of the operators.
[0125] Both the first operation sequence and the second operation sequence belong to the sequence data of a single operator. Considering that construction sites basically involve the collaborative work of multiple operators, multiple first operation sequences are formed into a first operation matrix, and multiple second operation sequences are formed into a second operation matrix. This can improve the collaborative features of multiple operators in the first operation matrix and the second operation matrix and improve the accuracy of identifying operation modes.
[0126] Step 106: Load the job classification network.
[0127] In this embodiment, a job classification network can be constructed and trained in advance based on deep learning. The job classification network is a multi-classification neural network. The structure of the job classification network is not limited to artificially designed neural networks, but can also be a neural network optimized by a model quantization method, a neural network searched for characteristics of construction by a NAS (Neural Architecture Search) method, and so on. This embodiment does not impose any restrictions on this.
[0128] When training the job classification network, the historical first job matrix and the second job matrix are selected as samples, the job modes marked by the managers are used as labels, and the cross entropy is selected as the loss function. The job classification network is supervised and trained, so that the job classification network has the ability to classify the job modes of operators.
[0129] Among them, the operation mode is set for the foundation pit, for example, staying close to the foundation pit, constructing close to the foundation pit, staying far away from the foundation pit, constructing far away from the foundation pit, and so on.
[0130] Step 107: Input the first operation matrix and the second operation matrix into the operation classification network to identify the operation modes of multiple operators.
[0131] In practical applications, the first operation matrix and the second operation matrix are scaled to a specified size. If the scaling is completed, the first operation matrix and the second operation matrix are input into the operation classification network. The operation classification network performs multi-classification tasks and identifies the operation mode when multiple operators work together.
[0132] In one embodiment of the present invention, Figure 2 As shown, the job classification network includes the first encoder Encoder_1, the second encoder Encoder_2, the third encoder Encoder_3, the fourth encoder Encoder_4, the feature interactive module (Feature Interactive Module, FIM), the feature fusion module (Feature Fusion Module ion Module, FFM) and the head structure Head.
[0133] Then, in this embodiment, step 107 may include the following steps:
[0134] Step 1071: Input the first operation matrix into the first encoder for encoding to obtain the first operation feature.
[0135] In this embodiment, if Figure 2 As shown, the first operation matrix can be input into the first encoder Encoder_1 for preliminary encoding to obtain the first operation feature E1.
[0136] For example, Figure 2 As shown, the first encoder Encoder_1 includes three first convolutional layers Conv_1. Then, in this example, the first operation matrix can be sequentially input into the three first convolutional layers Conv_1 to perform convolution operations to obtain the first operation feature E1.
[0137] Step 1072: Input the second operation matrix into the second encoder for encoding to obtain the second operation feature.
[0138] In this embodiment, if Figure 2 As shown, the second operation matrix can be input into the second encoder Encoder_2 for preliminary encoding to obtain the second operation feature E2.
[0139] For example, Figure 2 As shown, the second encoder Encoder_2 includes three second convolutional layers Conv_2. Then, in this example, the second operation matrix can be sequentially input into the three second convolutional layers Conv_2 to perform convolution operations to obtain the second operation feature E2.
[0140] Step 1073: Input the first operation feature and the second operation feature into the interactive part feature of the feature interaction module to obtain the third operation feature and the fourth operation feature.
[0141] In this embodiment, if Figure 2 As shown, the first operation feature E1 and the second operation feature E2 can be input into the feature interaction module FIM to interact with each other to obtain the third operation feature E3 and the fourth operation feature E4.
[0142] Exemplarily, in the feature interaction module FIM, an upsampling operation UpSample is performed on the first job feature E1 to obtain a first intermediate feature, the first intermediate feature is added to the second job feature E2 Add to obtain a third job feature E3, a downsampling operation DownSample is performed on the third job feature E3 to obtain a second intermediate feature, the second intermediate feature is added to the first intermediate feature Add to obtain a fourth job feature E4.
[0143] Step 1074: input the third operation feature into the third encoder for encoding to obtain the fifth operation feature.
[0144] In this embodiment, if Figure 2 As shown, the third operation feature E3 can be input into the third encoder Encoder_3 for deep encoding to obtain the fifth operation feature E5.
[0145] For example, Figure 2As shown, the third encoder Encoder_3 includes three third convolutional layers Conv_3. Then, in the example, the third operation feature E3 is sequentially input into the three third convolutional layers Conv_3 to perform convolution operations to obtain the fifth operation feature E5.
[0146] Step 1075: Input the fourth operation feature into the fourth encoder for encoding to obtain the sixth operation feature.
[0147] In this embodiment, if Figure 2 As shown, the fourth operation feature can be input into the fourth encoder Encoder_4 for deep encoding to obtain the sixth operation feature.
[0148] For example, Figure 2 As shown, the fourth encoder Encoder_4 includes three fourth convolutional layers Conv_4. Then, in this example, the fourth operation feature E4 is sequentially input into the three fourth convolutional layers Conv_4 to perform convolution operations to obtain the sixth operation feature E6.
[0149] Step 1076: Input the fifth operation feature and the sixth operation feature into a feature fusion module to fuse them into a seventh operation feature.
[0150] In this embodiment, if Figure 2 As shown, the fifth operation feature E5 and the sixth operation feature E6 are input into the feature fusion module FFM and fused into the seventh operation feature E7.
[0151] For example, Figure 2 As shown, the feature fusion module includes a fifth convolutional layer Conv_5 and a sixth convolutional layer Conv_6. Then, in this example, the fifth operation feature E5 is input into the fifth convolutional layer Conv_5 to perform a convolutional layer operation to obtain a third intermediate feature, the third intermediate feature and the sixth operation feature E6 are concatenated into a fourth intermediate feature, and the fourth intermediate feature is input into the sixth convolutional layer Conv_6 to perform a convolutional layer operation to obtain a seventh operation feature E7.
[0152] In this embodiment, the first operating matrix and the second operating matrix are preliminarily encoded, and then some features are exchanged with each other, and then deep encoding is continued, and finally they are merged together. While retaining the original coordinate features and distance features, feature fusion is gradually performed to improve the quality of the final fused features.
[0153] Step 1077: Input the seventh operation feature into the head structure to classify and obtain the operation modes of multiple operators.
[0154] In this embodiment, the seventh operation feature E7 is input into the head structure Head and classified to obtain the operation modes of multiple operators.
[0155] For example, Figure 2 As shown, the head structure Head includes three fully connected layers (FC). Then, in this embodiment, the seventh operation feature E7 is sequentially input into the three fully connected layers FC and mapped into the eighth operation feature. The eighth operation feature is activated using activation functions such as Sigmoid, and the probabilities that multiple operators belong to multiple operation modes are obtained, and the operation mode with the highest probability is output.
[0156] Step 108: Construct an inspection task according to the operation mode, and push the inspection task to the management personnel for execution.
[0157] In this embodiment, NLP (Natural Language Processing) technology can be used to construct appropriate inspection tasks according to the operation mode.
[0158] In a specific implementation, the first job matrix and the second job matrix can be input into a text generator, such as LSTM (Long Short Term Memory), Transformer, etc., and the first job matrix and the second job matrix can be used to generate job description information, and the job mode and the job description information can be combined into target job information.
[0159] The target operation information is encoded into a vector, and the vector is used to search for task record information similar to the target operation information in a preset task library; wherein the task record information is information recorded when the inspection task is executed historically.
[0160] At this time, the inspection task can be constructed based on the task record information and target operation information.
[0161] Furthermore, the task record information is structured data, so the key points of the task can be extracted from the task record information based on the data structure, such as the residence time at a certain location, the warning method for the operator, etc.
[0162] The task key points and target operation information are written into a preset template to obtain a prompt, and the prompt is input into a large language model (Large Language Model) to construct an inspection task.
[0163] Inspection tasks are pushed to management personnel (accounts) through SMS and instant messaging. Management personnel perform inspection tasks, inspect foundation pits according to the requirements of inspection tasks, and warn relevant workers. When the inspection is completed, management personnel can modify the operation mode, obtain new samples and labels, and update the operation classification network through Fine-tuning and other modes to continuously improve the performance of the operation classification network, and generate new task record information and store it in the task library.
[0164] In this embodiment, multiple frames of image data are collected at the construction site; semantic segmentation is performed on the image data to obtain the foundation pit and multiple workers; the distance between each worker and the edge of the foundation pit is counted in the image data; the coordinates of the same worker in the multiple frames of image data are combined into a first operation sequence, and multiple distances are combined into a second operation sequence; for multiple workers, multiple first operation sequences are combined into a first operation matrix, and multiple second operation sequences are combined into a second operation matrix; the operation classification network is loaded; the first operation matrix and the second operation matrix are input into the operation classification network to identify the operation mode of multiple workers; inspection tasks are constructed based on the operation mode, and the inspection tasks are pushed to the management personnel for execution. In this embodiment, the operation mode is identified based on the movement mode of the workers relative to the foundation pit in computer vision, and the inspection tasks are constructed based on the operation mode in natural language processing. It has a high degree of automation and high accuracy, effectively reduces the errors and omissions caused by fatigue caused by long-term browsing of monitoring videos, pushes the inspection tasks to the management personnel in real time, and conducts inspections on the foundation pit site in a timely manner, reduces the risks brought by the foundation pit, and ensures the safety of construction operations.
[0165] It should be noted that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0166] Reference Figure 3 , shows a schematic diagram of a task pushing device based on a foundation pit provided in an embodiment of the present application, which may specifically include the following modules:
[0167] An image data acquisition module 301 is used to acquire multiple frames of image data at a construction site;
[0168] A semantic segmentation module 302, used to perform semantic segmentation on the image data to obtain a foundation pit and a plurality of workers;
[0169] A distance detection module 303 is used to count the distance between each of the workers and the edge of the foundation pit in the image data;
[0170] A sequence composition module 304, configured to compose the coordinates of the same operator in multiple frames of the image data into a first operation sequence, and compose multiple distances into a second operation sequence;
[0171] A matrix forming module 305 is used for forming a first operating matrix with a plurality of the first operating sequences and forming a second operating matrix with a plurality of the second operating sequences for a plurality of the operating personnel;
[0172] The job classification network loading module 306 is used to load the job classification network;
[0173] A work mode identification module 307, used for inputting the first work matrix and the second work matrix into the work classification network to identify work modes of a plurality of the workers;
[0174] The inspection task processing module 308 is used to construct an inspection task according to the operation mode and push the inspection task to the management personnel for execution.
[0175] In one embodiment of the present application, the job classification network comprises a first encoder, a second encoder, a third encoder, a fourth encoder, a feature interaction module, a feature fusion module and a head structure;
[0176] The operation mode recognition module 307 is also used for:
[0177] Inputting the first operation matrix into the first encoder for encoding to obtain a first operation feature;
[0178] Inputting the second operation matrix into the second encoder for encoding to obtain a second operation feature;
[0179] Inputting the first operation feature and the second operation feature into the interactive part of features in the feature interaction module to obtain a third operation feature and a fourth operation feature;
[0180] Inputting the third operation feature into the third encoder for encoding to obtain a fifth operation feature;
[0181] Inputting the fourth operation characteristic into the fourth encoder for encoding to obtain a sixth operation characteristic;
[0182] Inputting the fifth operation feature and the sixth operation feature into the feature fusion module and fusing them into a seventh operation feature;
[0183] The seventh operation feature is input into the head structure to classify the operation modes of the plurality of operators.
[0184] In one embodiment of the present application, the first encoder includes three first convolutional layers, and the second encoder includes three second convolutional layers;
[0185] The operation mode recognition module 307 is also used for:
[0186] Inputting the first operation matrix into the three first convolutional layers in sequence to perform convolution operations to obtain first operation features;
[0187] The operation mode recognition module 307 is also used for:
[0188] The second operation matrix is sequentially input into three second convolutional layers to perform convolution operations to obtain second operation features.
[0189] In one embodiment of the present application, the operation mode identification module 307 is further used to:
[0190] Performing an upsampling operation on the first job feature to obtain a first intermediate feature;
[0191] Adding the first intermediate feature and the second operation feature to obtain a third operation feature;
[0192] Performing a downsampling operation on the third operation feature to obtain a second intermediate feature;
[0193] The second intermediate feature is added to the first intermediate feature to obtain a fourth operation feature.
[0194] In one embodiment of the present application, the third encoder includes three third convolutional layers, and the fourth encoder includes three fourth convolutional layers;
[0195] The operation mode recognition module 307 is also used for:
[0196] Inputting the third operation feature into the three third convolutional layers in sequence to perform convolution operations, thereby obtaining a fifth operation feature;
[0197] The step of inputting the fourth operation feature into the fourth encoder for encoding to obtain a sixth operation feature comprises:
[0198] The fourth operation feature is sequentially input into the three fourth convolutional layers to perform a convolution operation to obtain a sixth operation feature.
[0199] In one embodiment of the present application, the feature fusion module includes a fifth convolutional layer and a sixth convolutional layer;
[0200] The operation mode recognition module 307 is also used for:
[0201] Inputting the fifth operation feature into the fifth convolutional layer to perform a convolutional layer operation to obtain a third intermediate feature;
[0202] splicing the third intermediate feature and the sixth operating feature into a fourth intermediate feature;
[0203] The fourth intermediate feature is input into the sixth convolutional layer to perform a convolutional layer operation to obtain a seventh operation feature.
[0204] In one embodiment of the present application, the inspection task processing module 308 is further used to:
[0205] Generate job description information using the first job matrix and the second job matrix;
[0206] Combining the operation mode and the operation description information into target operation information;
[0207] Searching for task record information similar to the target operation information in a preset task library; the task record information is information recorded when the inspection task was historically executed;
[0208] An inspection task is constructed based on the task record information and the target operation information.
[0209] In one embodiment of the present application, the inspection task processing module 308 is further used to:
[0210] Extracting task key points from the task record information;
[0211] Writing the task key points and the target operation information into a preset template to obtain instructions;
[0212] The instructions are input into a large language model to construct an inspection task.
[0213] An embodiment of the present application provides a task pushing device based on a foundation pit. By using this device, each step in the aforementioned method embodiments can be implemented.
[0214] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment part.
[0215] Reference Figure 4 , shows a schematic diagram of a terminal device provided by an embodiment of the present application. Figure 4 As shown, the terminal device 400 in the embodiment of the present application includes: a processor 410, a memory 420, and a computer program 421 stored in the memory 420 and executable on the processor 410. When the processor 410 executes the computer program 421, the steps in each embodiment of the above-mentioned task push method based on the foundation pit are implemented. Alternatively, when the processor 410 executes the computer program 421, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0216] Exemplarily, the computer program 421 may be divided into one or more modules / units, which are stored in the memory 420 and executed by the processor 410 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which may be used to describe the execution process of the computer program 421 in the terminal device 400.
[0217] The terminal device 400 may include, but is not limited to, a processor 410 and a memory 420. Those skilled in the art will appreciate that Figure 4 It is only an example of the terminal device 400 and does not constitute a limitation on the terminal device 400. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device 400 may also include input and output devices, network access devices, buses, etc.
[0218] The processor 410 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0219] The memory 420 may be an internal storage unit of the terminal device 400, such as a hard disk or memory of the terminal device 400. The memory 420 may also be an external storage device of the terminal device 400, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 400. Further, the memory 420 may also include both an internal storage unit of the terminal device 400 and an external storage device. The memory 420 is used to store the computer program 421 and other programs and data required by the terminal device 400. The memory 420 may also be used to temporarily store data that has been output or is to be output.
[0220] An embodiment of the present application also discloses a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the task push method based on the foundation pit as described in the aforementioned embodiments is implemented.
[0221] An embodiment of the present application further discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the task push method based on the foundation pit as described in the above-mentioned embodiments is implemented.
[0222] An embodiment of the present application further discloses a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the task push method based on the foundation pit described in the aforementioned embodiments.
[0223] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application is described in detail with reference to the above-mentioned embodiments, a person skilled in the art should understand that the technical solutions described in the above-mentioned embodiments can still be modified, or some of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A task pushing method based on foundation pit, characterized in that: include: Collect multiple frames of image data at a construction site; Performing semantic segmentation on the image data to obtain a foundation pit and a plurality of workers; Counting the distance between each of the workers and the edge of the foundation pit in the image data; The coordinates of the same operator in multiple frames of the image data are combined into a first operation sequence, and the multiple distances are combined into a second operation sequence; For a plurality of the operators, a plurality of the first operation sequences are combined into a first operation matrix, and a plurality of the second operation sequences are combined into a second operation matrix; Load the job classification network; Inputting the first operation matrix and the second operation matrix into the operation classification network to identify the operation modes of a plurality of the operators; An inspection task is constructed according to the operation mode, and the inspection task is pushed to the management personnel for execution.
2. The method according to claim 1, characterized in that The job classification network first encoder, second encoder, third encoder, fourth encoder, feature interaction module, feature fusion module and head structure; The step of inputting the first operation matrix and the second operation matrix into the operation classification network to identify the operation modes of the plurality of operators comprises: Inputting the first operation matrix into the first encoder for encoding to obtain a first operation feature; Inputting the second operation matrix into the second encoder for encoding to obtain a second operation feature; Inputting the first operation feature and the second operation feature into the interactive part of features in the feature interaction module to obtain a third operation feature and a fourth operation feature; Inputting the third operation feature into the third encoder for encoding to obtain a fifth operation feature; Inputting the fourth operation characteristic into the fourth encoder for encoding to obtain a sixth operation characteristic; Inputting the fifth operation feature and the sixth operation feature into the feature fusion module and fusing them into a seventh operation feature; The seventh operation feature is input into the head structure to classify the operation modes of the plurality of operators.
3. The method according to claim 2, characterized in that The first encoder includes three first convolutional layers, and the second encoder includes three second convolutional layers; The step of inputting the first operation matrix into the first encoder for encoding to obtain a first operation feature includes: Inputting the first operation matrix into the three first convolutional layers in sequence to perform convolution operations to obtain first operation features; The step of inputting the second operation matrix into the second encoder for encoding to obtain a second operation feature includes: The second operation matrix is sequentially input into three second convolutional layers to perform convolution operations to obtain second operation features.
4. The method according to claim 3, characterized in that The step of inputting the first operation feature and the second operation feature into the feature interaction module to interact with each other to obtain the third operation feature and the fourth operation feature comprises: Performing an upsampling operation on the first job feature to obtain a first intermediate feature; Adding the first intermediate feature and the second operation feature to obtain a third operation feature; Performing a downsampling operation on the third operation feature to obtain a second intermediate feature; The second intermediate feature is added to the first intermediate feature to obtain a fourth operation feature.
5. The method according to claim 4, characterized in that The third encoder includes three third convolutional layers, and the fourth encoder includes three fourth convolutional layers; The step of inputting the third operation feature into the third encoder for encoding to obtain the fifth operation feature comprises: Inputting the third operation feature into the three third convolutional layers in sequence to perform convolution operations, thereby obtaining a fifth operation feature; The step of inputting the fourth operation feature into the fourth encoder for encoding to obtain a sixth operation feature comprises: The fourth operation feature is sequentially input into the three fourth convolutional layers to perform a convolution operation to obtain a sixth operation feature.
6. The method according to claim 5, characterized in that The feature fusion module includes a fifth convolutional layer and a sixth convolutional layer; The step of inputting the fifth operation feature and the sixth operation feature into the feature fusion module and fusing them into a seventh operation feature includes: Inputting the fifth operation feature into the fifth convolutional layer to perform a convolutional layer operation to obtain a third intermediate feature; splicing the third intermediate feature and the sixth operating feature into a fourth intermediate feature; The fourth intermediate feature is input into the sixth convolutional layer to perform a convolutional layer operation to obtain a seventh operation feature.
7. The method according to any one of claims 1 to 6, characterized in that The construction of the inspection task according to the operation mode includes: Generate job description information using the first job matrix and the second job matrix; Combining the operation mode and the operation description information into target operation information; Searching for task record information similar to the target operation information in a preset task library; the task record information is information recorded when the inspection task was historically executed; An inspection task is constructed based on the task record information and the target operation information.
8. The method according to claim 7, characterized in that The constructing of the inspection task according to the task record information and the target operation information includes: Extracting task key points from the task record information; Writing the task key points and the target operation information into a preset template to obtain instructions; The instructions are input into a large language model to construct an inspection task.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the task pushing method based on the foundation pit is implemented as described in any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, the task pushing method based on the foundation pit is implemented as described in any one of claims 1-8.
Citation Information
Patent Citations
Personnel positioning system in foundation pit construction and risk assessment method
CN111307110A
Safety belt wearing identification and detection method for various high-altitude operation construction sites
CN112990232A
Image compression method and image compression device
CN113014927A
Construction method for deep and thick soft soil foundation pit
CN113186934A
Coal mine personnel behavior detection method and device, and storage medium
CN114038067A