Method, apparatus, and product for identifying objects at construction sites based on dual backbone fusion

The dual backbone fusion-based object identification method improves safety monitoring at construction sites by reducing labor costs and increasing accuracy through real-time detection of abnormal operations.

JP7843394B2Active Publication Date: 2026-04-09CHINA THREE GORGES CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional safety monitoring at construction sites requires significant labor costs and lacks real-time accuracy, leading to potential safety risks and low monitoring effectiveness.

Method used

An object identification method based on dual backbone fusion using a target detection model with a first backbone network, feature segmentation network, second backbone network, neck network, and detection head network to identify construction objects and determine abnormal operations in real-time.

Benefits of technology

This method reduces labor costs and enhances safety monitoring accuracy by enabling real-time detection of abnormal behaviors, thereby reducing safety accidents at construction sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843394000002
    Figure 0007843394000002
  • Figure 0007843394000003
    Figure 0007843394000003
  • Figure 0007843394000004
    Figure 0007843394000004
Patent Text Reader

Abstract

To provide an object identification method capable of monitoring an abnormal operation of a construction image in real time, saving labor costs required for safety monitoring at a construction site, improving an accuracy rate of the safety monitoring at the construction site, and reducing an occurrence rate of safety accidents.SOLUTION: The method includes: collecting a construction image; recognizing the construction image based on a target detection model to obtain a recognition result, the recognition result including a category of a construction object and location information of the construction object; and determining whether the construction image is associated with an abnormal action based on the recognition result. The ith feature segmentation module in the feature segmentation network performs convolution processing and segmentation processing on the ith type of first image feature, and outputs the jth type of sub-image feature to the jth feature fusion module connected to the ith feature segmentation module.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiments of this application relate to the technical field of image recognition, and more particularly to a method, apparatus, and product for identifying objects at construction sites based on dual backbone fusion. [Background technology]

[0002] With the continuous development of society, engineering industries such as hydroelectric power generation and civil engineering and construction are developing rapidly. However, in recent years, there is a possibility of safety accidents occurring due to factors such as people and vehicles getting too close to each other during construction.

[0003] To reduce safety accidents, conventional monitoring methods typically involve personnel patrols. Specifically, specialized patrol inspectors are stationed at construction sites, regularly inspecting the sites and monitoring the distance between construction workers and construction machinery. If this distance is shorter than the safe distance, an alarm is issued.

[0004] Traditional monitoring methods require a certain amount of labor costs. Furthermore, because patrol personnel cannot inspect construction sites in real time, traditional monitoring methods still have certain safety risks and the accuracy of safety monitoring is relatively low. [Overview of the project] [Problems that the invention aims to solve]

[0005] The embodiment of the present invention provides an object identification method at a construction site based on dual backbone fusion, which can save labor costs required for safety monitoring at construction sites, and the real-time monitoring of the embodiment of the present invention can improve the accuracy of safety monitoring at construction sites and further reduce the incidence of safety accidents.

[0006] Accordingly, the embodiments of the present invention further provide an object identification device, electronic equipment, and machine-readable media for construction sites based on dual backbone fusion, thereby ensuring the realization and application of the above method. [Means for solving the problem]

[0007] To solve the above problem, an embodiment of the present invention is a method for identifying objects at a construction site based on dual backbone fusion, wherein the method is Steps to collect construction images, A step of identifying the construction image and obtaining an identification result based on a target detection model, wherein the identification result includes the category of the construction object and location information of the construction object category in the construction image, and the category of the construction object includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category. The step includes determining whether the construction image is related to abnormal operation based on the identification result, The target detection model includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network includes n feature segmentation modules, the second backbone network includes n feature fusion modules, and at least some of the feature fusion modules are connected to channel-to-pixel modules, where n is a positive integer greater than 1. The step of identifying the construction image based on the target detection model is: The first backbone network includes the step of determining n types of first image features corresponding to the construction image, The i-th feature segmentation module in the feature segmentation network performs convolution and segmentation on a first image feature of kind i, the resulting segmentation process includes i sub-image features, and outputs the j-th sub-image features to the connected j-th feature fusion module, wherein i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the i sub-image features of kind i output by the same i-th feature segmentation module are sub-image features with different numbers of channels, one feature fusion module corresponds to one number of channels, and different feature fusion modules correspond to different numbers of channels. The i-th feature fusion module in the second backbone network performs a first fusion process on at least one sub-image feature to obtain a first fusion result, and the channel-to-pixel module in the second backbone network determines a second image feature based on the first fusion result and outputs the second image feature to the neck network. The neck network performs a second fusion process on the second image feature to obtain the result of the second fusion process, The detection head network discloses a method for identifying objects at a construction site based on dual backbone fusion, which includes the step of determining the identification result based on the second fusion processing result.

[0008] An embodiment of the present application is an object identification device, wherein the device is A collection module for collecting construction images, An object identification module for identifying a construction image and obtaining an identification result based on a target detection model, wherein the identification result includes the category of the construction object and the location information of the construction object category in the construction image, and the category of the construction object includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category. The module includes an abnormality determination module for determining whether the construction image is related to abnormal operation based on the identification result, The target detection model includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network includes n feature segmentation modules, the second backbone network includes n feature fusion modules, and at least some of the feature fusion modules are connected to channel-to-pixel modules, where n is a positive integer greater than 1. The object identification module is A first image feature determination module for determining n types of first image features corresponding to the construction image, using the first backbone network, A convolutional partitioning module that uses the i-th feature partitioning module in a feature partitioning network to perform convolution and partitioning on a first image feature of type i, the resulting partitioning process includes i-th type sub-image features, and outputs the j-th type sub-image features to a connected j-th feature fusion module, wherein i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the i-th type sub-image features output by the same i-th feature partitioning module are sub-image features with different numbers of channels, one feature fusion module corresponds to one type of channel count, and different feature fusion modules correspond to different numbers of channels, and A first fusion module that uses the i-th feature fusion module in the second backbone network to perform a first fusion process on at least one sub-image feature to obtain a first fusion result, and uses the channel-to-pixel module in the second backbone network to determine a second image feature based on the first fusion result, and outputs the second image feature to the neck network, A second fusion module for performing a second fusion process on the aforementioned second image features using a neck network and obtaining the result of the second fusion process, Further disclosure is an object identification device that includes an identification result determination module for determining the identification result based on the second fusion processing result using a detection head network.

[0009] Selectable, the categories for wearing safety protective equipment include the category for not wearing safety protective equipment. The aforementioned abnormality detection module is If the identification result includes the category of "not wearing safety protective equipment," then a first abnormality determination module determines that the construction image is related to abnormal operation, or The system includes a second abnormality determination module that determines distance information between the human category and the construction machine category based on the location information corresponding to the human category and the construction machine category included in the identification result, and determines whether the construction image is related to abnormal operation based on the distance information.

[0010] Selectively, the convolutional partitioning module is, A convolution module that uses the i-th feature partitioning module in the feature partitioning network to perform a convolution on the i-th type first image feature to obtain an image feature with c channels, A partitioning module for dividing an image feature with c channels into i types of sub-image features with different numbers of channels, depending on the channel dimension, wherein the number of channels in the j-th sub-image feature in the i types of sub-image features is t*2 (j-1) It includes a split module.

[0011] Optionally, the second backbone network further includes a first convolutional module connected prior to the feature fusion module. The first fusion module described above is A receiving module for receiving n-i+1 types of sub-image features and convolutional image features output by a first convolutional module, wherein the number of channels corresponding to the n-i+1 types of sub-image features is the first number of channels, and the number of channels corresponding to the convolutional image features is the second number of channels, An interpolation module for performing interpolation processing on n - i + 1 types of sub - image features by using the i - th feature fusion module, where the number of channels corresponding to the n - i + 1 types of sub - image features after interpolation processing is the second number of channels, the interpolation module, A fusion processing module for performing fusion processing on the n - i + 1 types of sub - image features after interpolation processing and the convolutional image features by using the i - th feature fusion module, and includes.

[0012] Optionally, the first convolutional module connected before the first feature fusion module is used to perform convolutional processing on the construction image and output the corresponding convolutional image features to the first feature fusion module.

[0013] Optionally, the first backbone network includes a second convolutional module and n - 1 processing units connected in sequence, and the processing unit includes a third convolutional module and a channel - to - pixel module. The second convolutional module is connected to the first feature segmentation module and is used to output the first type of first image features to the first feature segmentation module. The channel - to - pixel modules included in the n - 1 processing units are respectively connected to the corresponding n - 1 feature segmentation modules and are used to output the first image features to the corresponding feature segmentation modules.

[0014] Optionally, the training process of the target detection model is Inputting a construction image sample into the target detection model, and the target detection model outputs a prediction result corresponding to the construction image sample. The prediction result includes prediction box information corresponding to the category of the construction object, and the construction image sample corresponds to the measured box information, the step, Determining loss information corresponding to the measured box information and the prediction box information based on the minimum point distance function based on a horizontal rectangle, Updating the parameters of the target detection model based on the loss information, and includes.

[0015] An embodiment of the present application is an electronic device, including a processor and a memory in which executable code is stored. When the executable code is executed, the processor is caused to execute the method described in the embodiment of the present application, and further discloses an electronic device.

[0016] An embodiment of the present application further discloses a machine-readable medium in which executable code is stored, and when the executable code is executed, the processor is caused to execute the method described in the embodiment of the present application.

Advantages of the Invention

[0017] Embodiments of the present application include the following advantages.

[0018] In the technical solution of the embodiment of the present application, by means of image recognition, the abnormal operation of the construction image can be monitored in real time, and the labor cost required for safety monitoring at the construction site can be saved. Moreover, the above real-time monitoring can improve the accuracy rate of safety monitoring at the construction site, and further reduce the occurrence rate of safety accidents.

[0019] First, the n feature segmentation modules of the feature segmentation network in the embodiment of the present application can play a role in aggregating sub-image features at multiple levels. Based on this, the n feature fusion modules of the second backbone network can effectively fuse sub-image features at multiple levels. In the forward propagation stage of the target detection model, richer image features can be provided, and the representation ability of the second image features output by the second backbone network can be enhanced.

[0020] Furthermore, in the reverse propagation calculation process of the target detection model, the calculation of the deviation between aggregated multi-level sub-image feature information and the true value transmits more reliable gradient information, guiding parameter learning of the target detection model. This helps the first backbone network extract more important and accurate image features. Therefore, the embodiment of this application can further improve the accuracy of the identification results of construction images, and based on this, the accuracy of safety monitoring at construction sites can be improved. [Brief explanation of the drawing]

[0021] [Figure 1] This is a schematic diagram illustrating the application environment of an object identification method at a construction site based on dual backbone fusion, one embodiment of the present invention. [Figure 2] This is a schematic diagram of the step flow of an object identification method at a construction site based on dual backbone fusion, one embodiment of the present invention. [Figure 3] This is a schematic diagram of the structure of a target detection model according to one embodiment of the present invention. [Figure 4] This is a schematic diagram of the structure of the first backbone network 301, the feature segmentation network 302, and the second backbone network 303 of one embodiment of the present application. [Figure 5] This is a schematic diagram of the structure of a channel-to-pixel module according to one embodiment of the present invention. [Figure 6] This is a schematic diagram of the structure of a bottleneck module according to one embodiment of the present invention. [Figure 7] This is a schematic diagram of the structure of the second backbone network in one embodiment of the present invention. [Figure 8] This is a schematic diagram of the structure of the first backbone network of one embodiment of the present invention. [Figure 9] This is a schematic diagram of the structure of a neck network according to one embodiment of the present invention. [Figure 10] This is a schematic diagram of the structure of a spatial pyramid pooling module according to one embodiment of the present invention. [Figure 11]This is a schematic diagram of the structure of an object identification device for construction sites based on dual backbone fusion, according to one embodiment of the present invention. [Figure 12] This is a schematic diagram of the structure of the device provided by one embodiment of the present invention. [Modes for carrying out the invention]

[0022] To make the above-mentioned objectives, features, and advantages of this application clearer and easier to understand, the application will be described in further detail below, with reference to the drawings and specific embodiments.

[0023] The embodiment of this invention can be applied to engineering industries such as hydroelectric power generation and civil engineering and construction, and can be used to save labor costs while simultaneously enabling safety monitoring during the construction process, reducing the incidence of safety accidents during construction, and improving the accuracy of safety monitoring at construction sites.

[0024] Conventional technology typically involves deploying specialized patrol inspectors to construction sites. These inspectors periodically inspect the site, monitoring the distance between construction workers and machinery, and issuing an alarm if the distance falls below the safe distance. Conventional technology requires a certain level of labor costs. Furthermore, because patrol inspectors cannot inspect the construction site in real time, conventional monitoring methods still carry certain safety risks, resulting in a relatively low accuracy rate for safety monitoring.

[0025] In response to the technical challenges of the prior art, such as the consumption of labor costs and the relatively low accuracy of safety monitoring, an embodiment of the present invention provides an object identification method, which specifically includes the steps of: collecting construction images; identifying the construction images based on a target detection model to obtain an identification result, wherein the identification result specifically includes the category of the construction object and the location information of the construction object category in the construction image, wherein the category of the construction object specifically includes at least one of the categories of human, construction machinery, safety protective equipment, and safety sign; and determining, based on the identification result, whether the construction image is related to abnormal operation. The above target detection model specifically includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network specifically includes n feature segmentation modules, and the second backbone network specifically includes n feature fusion modules, with at least some feature fusion modules followed by channel-to-pixel modules, where n is a positive integer greater than 1. The process of identifying the above construction images based on the target detection model is, specifically, The above-mentioned first backbone network includes the step of determining n types of first image features corresponding to the above-mentioned construction image, The i-th feature segmentation module in the feature segmentation network performs convolution and segmentation on the i-th type first image feature, and the resulting segmentation process includes i-th type sub-image features, which are output to the connected j-th feature fusion module, wherein i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the i-th type sub-image features output by the same i-th feature segmentation module are sub-image features with different numbers of channels, one feature fusion module corresponds to one type of channel count, and different feature fusion modules correspond to different numbers of channels. The i-th feature fusion module in the second backbone network performs a first fusion process on at least one sub-image feature to obtain a first fusion result, and the channel-to-pixel module in the second backbone network determines a second image feature based on the first fusion result and outputs the second image feature to the neck network. The above neck network performs a second fusion process on the above second image features to obtain the result of the second fusion process, The above detection head network includes the step of determining the identification result based on the above second fusion processing result.

[0026] The image recognition method of the embodiment of this invention allows for real-time monitoring of abnormal behavior in construction images, thereby saving labor costs required for safety monitoring at construction sites. Furthermore, this real-time monitoring can improve the accuracy of safety monitoring and reduce the incidence of safety-related accidents.

[0027] The i-th feature segmentation module of the embodiment of the present application segmentes the first image feature to obtain i types of sub-image features with different numbers of channels, and outputs the j-th type of sub-image feature to the connected j-th feature fusion module. Since different i-th feature fusion modules can process first image features at different levels, the n feature segmentation modules of the embodiment of the present application can serve to aggregate multiple levels of sub-image features. For example, the n feature segmentation modules can provide the first feature fusion module with n levels of sub-image features, and the n feature segmentation modules can provide the second feature fusion module with (n-1) levels of sub-image features.

[0028] First, the n feature division modules of the feature division network in the embodiment of the present invention can aggregate sub-image features at multiple levels. Based on this, the n feature fusion modules of the second backbone network can effectively fuse sub-image features at multiple levels, providing richer image features during the forward propagation phase of the target detection model and enhancing the expressive power of the second image features output by the second backbone network.

[0029] Furthermore, in the reverse propagation calculation process of the target detection model, the calculation of the deviation between aggregated multi-level sub-image feature information and the true value transmits more reliable gradient information, guiding parameter learning of the target detection model. This helps the first backbone network extract more important and accurate image features. Therefore, the embodiment of this application can further improve the accuracy of the identification results of construction images, and based on this, the accuracy of safety monitoring at construction sites can be improved.

[0030] Referring to Figure 1, which shows a schematic diagram of the application environment for the object identification method at a construction site based on dual backbone fusion according to an embodiment of the present invention, the image acquisition terminal 101 and the server 102 can exchange data based on a wireless or wired network.

[0031] In actual applications, the image acquisition terminal 101 may be equipped with an image acquisition device having an image acquisition function, such as an image sensor. The image acquisition device can collect construction images, and the image acquisition terminal 101 can transmit the construction images to the server 102 according to a preset time cycle. Of course, the image acquisition device can also collect construction videos, and the image acquisition terminal 101 can transmit the construction videos to the server 102 according to a preset time cycle. In this case, the server 102 can analyze the construction images from the construction videos.

[0032] After receiving construction images transmitted by the image collection terminal 101, the server 102 can process the construction images using the method of the embodiment of the present invention and determine whether the construction images are related to abnormal operation. If so, it can output notification information. For example, the notification information can be sent to a pre-configured user's terminal. The notification information indicates the occurrence of abnormal operation and prompts the pre-configured user to take appropriate processing measures.

[0033] Examples of abnormal operation include situations where the distance between a person (construction worker) and construction machinery is shorter than the safe distance, where a person is not wearing safety protective equipment, or where safety signs are not installed in the construction environment corresponding to the construction image.

[0034] Server 102 can send information to a pre-configured user in the event of abnormal operation, prompting the pre-configured user to take appropriate action. The pre-configured user may be the manager of the construction environment. For example, if the distance between a person and construction machinery is shorter than the safety distance, an example of an action could be to play a first alarm sound using the speaker closest to the construction site corresponding to the distance construction image, prompting the construction worker in the construction image to move away from the construction machinery. Also, if a person is not wearing safety protective equipment, an example of an action could be to play a second alarm sound using the speaker closest to the construction site corresponding to the construction image, prompting the construction worker in the construction image to wear a safety cap. Furthermore, if no safety signs are installed in the construction environment corresponding to the construction image, the manager of the construction environment can install safety signs such as safety slogans in the construction environment. Those skilled in the art will understand that various action measures can be used depending on the requirements of the actual application, and that the embodiments of this application do not limit the specific action measures.

[0035] Example of the method Referring to Figure 2, which shows a schematic step flow of an object identification method at a construction site based on dual backbone fusion according to one embodiment of the present invention, the method specifically, Step 201 involves collecting construction images, Step 202, which identifies the construction image based on a target detection model and obtains an identification result, wherein the identification result specifically includes the category of the construction object and the location information of the construction object category in the construction image, and the category of the construction object specifically includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category, Step 203 includes determining whether the construction image is related to abnormal operation based on the above identification result, The above target detection model specifically includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network specifically includes n feature segmentation modules, and the second backbone network specifically includes n feature fusion modules, with channel-to-pixel modules connected after at least some of the feature fusion modules, where n may be a positive integer greater than 1. Based on the target detection model in step 202 above, the process of identifying the construction image is as follows: The first backbone network performs step 221, which determines n types of first image features corresponding to the above construction image, Step 222 is a feature segmentation module in a feature segmentation network that performs convolution and segmentation on a first image feature of kind i, the resulting segmentation process includes a sub-image feature of kind i, and outputs a sub-image feature of kind j to a connected j-th feature fusion module, where i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the sub-image features of kind i output by the same i-th feature segmentation module are sub-image features with different numbers of channels, one feature fusion module corresponds to one number of channels, and different feature fusion modules correspond to different numbers of channels. Step 223: The i-th feature fusion module in the second backbone network performs a first fusion process on at least one sub-image feature to obtain a first fusion result, and the channel-to-pixel module in the second backbone network determines a second image feature based on the first fusion result and outputs the second image feature to the neck network. Step 224 involves the neck network performing a second fusion process on the above second image features to obtain the result of the second fusion process. The detection head network includes step 225, which determines an identification result based on the second fusion processing result described above.

[0036] The steps included in the embodiment of the method shown in Figure 2 can be performed by a server, which can leverage its abundant computing resources to rapidly process construction images and monitor abnormal behavior in construction images in real time based on image recognition. The embodiments of this application do not limit the specific implementer of the embodiment of the method shown in Figure 2.

[0037] In the embodiment of the present invention, since the construction site can be monitored in real time based on construction images, real-time monitoring by human labor becomes unnecessary, and labor costs associated with construction monitoring can be saved.

[0038] In step 201, the server can receive construction images transmitted by the image collection terminal according to a predetermined time cycle. The image collection terminal is installed at the construction site and can be used to collect construction images of the construction site in real time and transmit the construction images to the server according to a predetermined time cycle.

[0039] In step 202, the construction image is identified based on the target detection model, and the identification result obtained specifically includes the category of the construction object and the location information of the construction object category in the construction image, and the category of the construction object specifically includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category.

[0040] Examples of the human category include construction workers, etc. Examples of the construction machinery category include backhoe excavators, tower cranes, dump trucks, truck cranes, loaders, hooks, pump trucks, smooth rollers, concrete mixer trucks, pile drivers, etc. The safety equipment wearing category specifically includes categories of people wearing safety equipment or people not wearing safety equipment. The safety sign category specifically includes safety slogans, etc.

[0041] Referring to Figure 3, which shows a schematic diagram of the structure of a target detection model of one embodiment of the present invention, it specifically includes a first backbone network 301, a feature segmentation network 302, a second backbone network 303, a neck network 304, and a detection head network 305 connected in order.

[0042] The first backbone network 301 is used to extract features from the input image and obtain n types of first image features. During the training phase, the input image may be a sample construction image. During the image recognition phase, the input image may be a real-time construction image.

[0043] The feature segmentation network 302 is used to perform convolution and segmentation on n types of first image features and output the multiple types of sub-image features obtained by segmentation to the second backbone network 303.

[0044] The second backbone network 303 is used to perform first fusion processing and channel-to-pixel processing on multiple types of sub-image features to obtain second image features, and to output these second image features to the neck network 304.

[0045] The neck network 304 is used to perform a second fusion process on the second image feature described above and to obtain the result of the second fusion process.

[0046] The detection head network 305 is used to determine the identification result based on the results of the second fusion process described above.

[0047] Referring to Figure 4, it shows a schematic diagram of the structure of a first backbone network 301, a feature segmentation network 302, and a second backbone network 303 of one embodiment of the present invention, where the first backbone network 301 outputs n types of first image features to n feature segmentation modules of the feature segmentation network 302.

[0048] The feature partitioning network 302 specifically includes n feature partitioning modules, which are denoted as the first feature partitioning module 321, the second feature partitioning module 322, ... and the nth feature partitioning module 32n, respectively.

[0049] The i-th feature partitioning module in the feature partitioning network performs convolution and partitioning on the i-th type first image feature, and the resulting partitioning process includes i-th type sub-image features, which are output to the connected j-th feature fusion module as j-th type sub-image features. i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and if i is greater than 1, then the i-th type sub-image features output by the same i-th feature partitioning module are sub-image features with different numbers of channels, one feature fusion module corresponds to one type of channel count, and different feature fusion modules correspond to different numbers of channels.

[0050] The second backbone network 303 specifically includes n feature fusion modules, which are denoted as the first feature fusion module 331, the second feature fusion module 332, ... and the nth feature fusion module 33n, respectively.

[0051] In a specific implementation, the process by which the i-th feature segmentation module in the feature segmentation network performs convolution and segmentation on the i-th type first image feature is, specifically, The i-th feature segmentation module in the feature segmentation network performs a convolution process on the first image feature of the i-th type to obtain an image feature with c channels in step A1, Step A2 is to divide the image feature with c channels into i types of sub-image features with different numbers of channels according to the channel dimension. The number of channels of the j-th type of sub-image feature in the i types of sub-image features is t*2 (j-1) where t can be determined by those skilled in the art according to the actual application needs. Examples of the value of t include 32, 64, etc. This includes step A2.

[0052] The number of channels of the i types of sub-image features is respectively t*2 0 、t*2 1 、t*2 [[ID=1,3]] 2 …、t*2 (i-1) That is. The sum of the number of channels corresponding to the i types of sub-image features may be c, that is, t*2 0 +t*2 1 +t*2 2 …+t*2 (i-1) =c.

[0053] When n is 5, the i-th feature segmentation module can output the j-th type of sub-image feature to the connected j-th feature fusion module. Specific examples include the following.

[0054] For example, the segmentation processing result of the first feature segmentation module 321 includes one type of m-channel first sub-image feature, and one type of m-channel first sub-image feature is output to the first feature fusion module 331.

[0055] Also, for example, the segmentation processing result of the second feature segmentation module 3,22 includes one type of m-channel second sub-image feature and one type of 2m-channel third sub-image feature. One type of m-channel second sub-image feature is output to the first feature fusion module 331. One type of 2m-channel third sub-image feature is output to the second feature fusion module 332.

[0056] Furthermore, for example, the segmentation result of the third feature segmentation module 323 (not shown) includes a fourth sub-image feature of one type m channel, a fifth sub-image feature of one type 2m channel, and a sixth sub-image feature of one type 4m channel. The fourth sub-image feature of one type m channel is output to the first feature fusion module 331. The fifth sub-image feature of one type 2m channel is output to the second feature fusion module 332. The sixth sub-image feature of one type 4m channel is output to the third feature fusion module 333.

[0057] Alternatively, the segmentation result of the fourth feature segmentation module 324 (not shown) includes a seventh sub-image feature of one type of m channel, an eighth sub-image feature of one type of 2m channel, a ninth sub-image feature of one type of 4m channel, and a tenth sub-image feature of one type of 8m channel. The seventh sub-image feature of one type of m channel is output to the first feature fusion module 331. The eighth sub-image feature of one type of 2m channel is output to the second feature fusion module 332. The ninth sub-image feature of one type of 4m channel is output to the third feature fusion module 333. The tenth sub-image feature of one type of 8m channel is output to the fourth feature fusion module 334.

[0058] Alternatively, the segmentation result of the fifth feature segmentation module 325 (not shown) includes a 11th sub-image feature of one type of m channel, a 12th sub-image feature of one type of 2m channel, a 13th sub-image feature of one type of 4m channel, a 14th sub-image feature of one type of 8m channel, and a 15th sub-image feature of one type of 16m channel. The 11th sub-image feature of one type of m channel is output to the first feature fusion module 331. The 12th sub-image feature of one type of 2m channel is output to the second feature fusion module 332. The 13th sub-image feature of one type of 4m channel is output to the third feature fusion module 333. The 14th sub-image feature of one type of 8m channel is output to the fourth feature fusion module 334. The 15th sub-image feature of one type of 16m channel is output to the fifth feature fusion module 335.

[0059] The i-th feature fusion module in the second backbone network performs a first fusion process on at least one sub-image feature to obtain a first fusion result. Based on this first fusion result, the channel-to-pixel module in the second backbone network determines a second image feature and outputs the second image feature to the neck network.

[0060] In the selectable implementations of the present invention, the second backbone network may further include a first convolutional module connected prior to the feature fusion module. In this case, the process by which the i-th feature fusion module in the second backbone network performs the first fusion process on at least one type of sub-image feature is, specifically, Step B1 involves the i-th feature fusion module receiving n-i+1 types of sub-image features and a convolutional image feature output by the first convolutional module, wherein the number of channels corresponding to the n-i+1 types of sub-image features is the first number of channels, and the number of channels corresponding to the convolutional image feature is the second number of channels. Step B2 is a step in which the i-th feature fusion module performs interpolation on n-i+1 types of sub-image features, and the number of channels corresponding to the n-i+1 types of sub-image features after interpolation is the second number of channels. The i-th feature fusion module includes step B3, which performs a fusion process on the n-i+1 types of sub-image features and convolutional image features after interpolation.

[0061] The first convolution module, the second convolution module, or the first convolution processing module in the embodiments of this application all fall within the scope of a convolutional structure. In one example, the convolutional structure specifically includes at least one convolutional layer, at least one batch normalization layer, and at least one activation function. Those skilled in the art will understand that a desired convolutional structure can be used depending on the needs of the actual application, and that the embodiments of this application do not limit the specific convolutional structure.

[0062] For example, the number of first channels corresponding to n types of sub-image features received by the first feature fusion module is t*2. 0 Those skilled in the art can use corresponding interpolation techniques according to the needs of their actual applications, and the embodiments of this application do not limit the specific interpolation techniques.

[0063] The process by which the i-th feature fusion module performs fusion on the n-i+1 types of sub-image features and convolutional image features after interpolation specifically includes the i-th feature fusion module adding values ​​corresponding to the n-i+1 types of sub-image features and convolutional image features after interpolation. The sub-image features and convolutional image features can correspond to multidimensional matrices, and the above addition process may be the addition of element values ​​of the multidimensional matrices.

[0064] In another selectable implementation of the present invention, a first convolution module connected before the first feature fusion module is used to perform convolution on the construction image and output the corresponding convolutional image features to the first feature fusion module.

[0065] A channel-to-pixel module is connected after at least some feature fusion modules. Referring to Figure 5, which shows a schematic diagram of the structure of a channel-to-pixel module of one embodiment of the present invention, the channel-to-pixel module specifically includes a first convolution module 501, a splitting module 502, M bottleneck modules 503, a connection module 504, and a second convolution module 505, where M may be a positive integer greater than 1.

[0066] The first convolution processing module 501 is used to perform a first convolution process on the first fusion result output by the feature fusion module to obtain a first convolution processing result.

[0067] The splitting module 502 is used to split the result of the first convolution into two parts with the same number of channels. These two parts may include the features of the first part and the features of the second part. Assuming that the number of channels before splitting is the third number of channels, the number of channels after splitting may be the fourth number of channels, and the third number of channels may be twice the fourth number of channels.

[0068] After the features of the first part are processed by M bottleneck modules 503, the resulting bottleneck processing features are entered into the connection module 504. The first convolution result and the features of the second part are also entered into the connection module 504. The connection module 504 is used to perform connection operations on the first convolution result, the features of the second part, and the bottleneck processing features to obtain connection features.

[0069] The connection features are entered into the second convolution module 505, which restores the number of channels for the connection features, for example, restoring the number of channels from the fourth channel to the third channel.

[0070] Referring to Figure 6, which shows a schematic diagram of the structure of a bottleneck module of one embodiment of the present invention, the bottleneck module specifically includes a third convolution module 601 and a fourth convolution module 602.

[0071] The third convolution module 601 is used to reduce the number of channels in the input features to half of the original number and obtain the result of the second convolution.

[0072] The fourth convolution module 603 is used to double the number of channels in the second convolution result to obtain the third convolution result. The third convolution result has the same number of channels as the input features. The third convolution result and the input features are fused to obtain the output features.

[0073] Referring to Figure 7, which shows a schematic diagram of the structure of a second backbone network of one embodiment of the present application, the second backbone network specifically includes a first convolutional module A701, a first feature fusion module 702, a first convolutional module B703, a second feature fusion module 704, a first channel-to-pixel module 705, a first convolutional module C706, a third feature fusion module 707, a second channel-to-pixel module 708, a first convolutional module D709, a fourth feature fusion module 710, a third channel-to-pixel module 711, a first convolutional module E712, a fifth feature fusion module 713, and a fourth channel-to-pixel module 714.

[0074] The first convolution module A701 is used to perform convolution on the construction image and output the corresponding convolutional image feature A to the first feature fusion module.

[0075] The first feature fusion module 702 receives five types of sub-image features A from the first feature segmentation module and convolutional image features A from the first convolution module A701, performs interpolation on the five types of sub-image features A, and assumes that the number of channels corresponding to the five types of sub-image features A after interpolation is the second number of channels, and that a fusion process is performed on the five types of sub-image features A after interpolation and the convolutional image features A to obtain the first fusion result A.

[0076] The first convolution module B703 is used to perform a convolution operation on the first fusion result A to obtain the convolutional image feature B.

[0077] The second feature fusion module 704 receives four types of sub-image features B from the second feature segmentation module and convolutional image features B from the first convolution module B 703. It is used to assume that the number of channels corresponding to the four types of sub-image features B after interpolation is the second channel number, and that a fusion process is performed on the four types of sub-image features B after interpolation and the convolutional image features B to obtain the first fusion result B.

[0078] The first channel two-pixel module 705 is used to determine intermediate image features based on the first fusion result B described above.

[0079] The first convolution module C706 is used to perform convolution on intermediate image features to obtain convolutional image features C.

[0080] The third feature fusion module 707 receives three types of sub-image features C from the third feature segmentation module and a convolutional image feature C from the first convolution module C706. It is assumed that interpolation is performed on the three types of sub-image features C, that the number of channels corresponding to the three types of sub-image features C after interpolation is the second channel number, and that fusion is performed on the three types of sub-image features C and the convolutional image feature C to obtain the first fusion result C.

[0081] The second channel two-pixel module 708 is used to determine the second image feature A based on the first fusion result C and to output the second image feature A to the neck network.

[0082] The first convolution module D709 is used to perform a convolution operation on the second image feature A to obtain the convolutional image feature D.

[0083] The fourth feature fusion module 710 receives two types of sub-image features D from the fourth feature segmentation module, receives convolutional image features D from the first convolution module D709, performs interpolation on the three types of sub-image features D, assumes that the number of channels corresponding to the two types of sub-image features D after interpolation is the second channel number, and assumes that a fusion process is performed on the two types of sub-image features D and the convolutional image features D after interpolation to obtain the first fusion result D.

[0084] The third channel two-pixel module 711 is used to determine the second image feature B based on the first fusion result D and to output the second image feature B to the neck network.

[0085] The first convolution module E712 is used to perform a convolution operation on the second image feature B to obtain the convolutional image feature E.

[0086] The fifth feature fusion module 713 receives one type of sub-image feature E from the fifth feature segmentation module, receives a convolutional image feature E from the first convolution module E706, performs interpolation on the one type of sub-image feature E, assumes that the number of channels corresponding to the one type of sub-image feature E after interpolation is the second channel number, and assumes that a fusion process is performed on the one type of sub-image feature E after interpolation and the convolutional image feature E to obtain a first fusion result E.

[0087] The fourth channel two-pixel module 714 is used to determine the second image feature C based on the first fusion result E and to output the second image feature C to the neck network.

[0088] In a specific implementation, the first backbone network may specifically include a second convolutional module and n-1 processing units connected in sequence, and each processing unit may specifically include a third convolutional module and a channel-to-pixel module. The second convolution module is connected to the first feature segmentation module and used to output a first image feature of the first kind to the first feature segmentation module. The channel-to-pixel modules included in the n-1 processing units are each connected to the corresponding n-1 feature segmentation modules and used to output a first image feature to the corresponding feature segmentation modules.

[0089] Referring to Figure 8, which shows a schematic diagram of the structure of a first backbone network in one embodiment of the present invention, the first backbone network specifically includes, in order, a second convolutional module 801, a third convolutional module A802, a fifth channel-to-pixel module 803, a third convolutional module B804, a sixth channel-to-pixel module 805, a third convolutional module C806, a seventh channel-to-pixel module 807, a third convolutional module D808, and an eighth channel-to-pixel module 809. The structure of the fifth channel-to-pixel module 803 is shown in Figure 5 and will not be described in detail here.

[0090] The second convolution module 801 is used to perform convolution on the construction image to obtain the first image feature A, and to transmit the first image feature A to the first feature segmentation module.

[0091] The third convolution module A802 is used to perform a convolution operation on the first image feature A to obtain the first convolution result.

[0092] The fifth channel two-pixel module 803 is used to determine the first image feature B based on the first convolution result and to transmit the first image feature B to the second feature segmentation module.

[0093] The third convolution module B804 is used to perform convolution on the first image feature B to obtain the second convolution result.

[0094] The sixth channel two-pixel module 805 is used to determine the first image feature C based on the second convolution result and to transmit the first image feature C to the third feature segmentation module.

[0095] The third convolution module C806 is used to perform a convolution operation on the first image feature C to obtain the third convolution result.

[0096] The seventh channel two-pixel module 807 is used to determine the first image feature D based on the third convolution result and to transmit the first image feature D to the fourth feature segmentation module.

[0097] The third convolution module D808 is used to perform convolution on the first image feature D to obtain the fourth convolution result.

[0098] The seventh channel two-pixel module 807 is used to determine the first image feature E based on the fourth convolution result and to transmit the first image feature E to the fifth feature segmentation module.

[0099] In step 224, the neck network is used to perform a second fusion process on the second image feature described above to obtain the result of the second fusion process.

[0100] Referring to Figure 9, which shows a schematic diagram of the structure of a neck network of one embodiment of the present invention, the neck network specifically includes a spatial pyramid pooling module 901, a first upsampling module 902, a first connection module 903, a ninth channel-to-pixel module 904, a second upsampling module 905, a second connection module 906, a tenth channel-to-pixel module 907, a fourth convolution module 908, a third connection module 909, an eleventh channel-to-pixel module 910, a fifth convolution module 911, a fourth connection module 912, and a twelfth channel-to-pixel module 913.

[0101] The spatial pyramid pooling module 901 receives a second image feature C and is used to perform spatial pyramid pooling on the second image feature C. The spatial pyramid pooling process may include convolution and max pooling operations, enabling deep fusion of the second image feature C.

[0102] Referring to Figure 10, which shows a schematic diagram of the structure of a spatial pyramid pooling module of one embodiment of the present invention, the spatial pyramid pooling module specifically includes a sixth convolution module 1001, p maximum pooling modules 1002, a fifth connection module 1003, and a seventh convolution module 1004. p may be a positive integer greater than 1, and the value of p in the figure is 3.

[0103] The sixth convolution module 1001 is used to perform a convolution operation on the input second image feature C to obtain the fifth convolution result.

[0104] Each of the p maximum pooling modules 1002 is used to perform a maximum pooling operation on the input image features and to obtain the corresponding p maximum pooling results.

[0105] The result of the fifth convolution and the results of the p maximum pooling operations are input to the fifth connection module 1003, which then performs deep fusion on the fifth convolution result and the p maximum pooling operations to obtain the corresponding deep fusion result.

[0106] The seventh convolution module 1004 is used to perform a convolution operation on the deep fusion result to obtain the sixth convolution result. The sixth convolution result is provided to the first upsampling module 902.

[0107] The first upsampling module 902 is used to perform a first upsampling process on the sixth convolution result to obtain the first upsampling result.

[0108] The first connection module 903 is used to perform a connection process on the second image feature B and the first upsampling result to obtain the first connection result.

[0109] The ninth channel two-pixel module 904 is used to determine the first processing result based on the first connection result.

[0110] The second upsampling module 905 is used to perform a second upsampling process on the first processing result and obtain the second upsampling result.

[0111] The second connection module 906 is used to perform a connection process on the second upsampling result and the second image feature A to obtain a second connection result. The above connection process may also be a joining process of two types of image features.

[0112] The 10th channel two-pixel module 907 is used to determine the second fusion processing result A based on the second connection result. The second fusion processing result A is output to the detection head network.

[0113] The fourth convolution module 908 is used to perform a convolution operation on the second fusion result A to obtain the sixth convolution result.

[0114] The third connection module 909 is used to perform a connection process on the sixth convolution result and the first processing result to obtain the third connection result.

[0115] The 11th channel two-pixel module 910 is used to determine the second fusion processing result B based on the third connection result. The second fusion processing result B is output to the detection head network.

[0116] The fifth convolution module 911 is used to perform a convolution operation on the second fusion result B to obtain the seventh convolution result.

[0117] The fourth connection module 912 is used to perform a connection process on the sixth and seventh convolution results to obtain the fourth connection result.

[0118] The 12th channel two-pixel module 913 is used to determine the second fusion processing result C based on the fourth connection result. The second fusion processing result C is output to the detection head network.

[0119] In step 225, the detection head network can determine the identification result based on the results of the second fusion process described above.

[0120] The detection head network may include at least one detection module. Based on the second fusion processing result, the detection module can be used to perform classification and regression calculations using a convolutional module and a convolutional layer, respectively, to obtain the category of construction object in the construction image and the positional information of the category of construction object in the construction image.

[0121] In summary, in the embodiments of the present invention, the n feature division modules of the feature division network can perform the role of aggregating sub-image features at multiple levels, and based on this, the n feature fusion modules of the second backbone network can effectively fuse sub-image features at multiple levels, thus providing richer image features during the forward propagation phase of the target detection model and enhancing the expressive power of the second image features output by the second backbone network.

[0122] Furthermore, in the reverse propagation calculation process of the target detection model, the calculation of the deviation between aggregated multi-level sub-image feature information and the true value transmits more reliable gradient information, guiding parameter learning of the target detection model. This helps the first backbone network extract more important and accurate image features. Therefore, the embodiment of this application can further improve the accuracy of the identification results of construction images, and based on this, the accuracy of safety monitoring at construction sites can be improved.

[0123] The target detection model of this embodiment is experimentally analyzed using a target detection criterion dataset from a construction site. The 19,404 marked training set images from the construction site target detection criterion dataset are used as the training dataset, and the 4,000 validation set images are used as the test dataset. AP (Average Precision) 50 and AP75 are defined as evaluation metrics. AP50 represents the average precision when the cross-overunion threshold is 0.5. AP75 represents the average precision when the cross-overunion threshold is 0.75.

[0124] Experimental results showed that the detection accuracy of the embodiment of the present invention in a target detection criterion dataset for construction sites was superior to that of the prior art. Using the AP50 index as an example, the detection accuracy of the first, second, and third versions of the target detection model improved by 13.6%, 8.7%, and 4.3% respectively compared to the prior art, verifying the effectiveness of the method according to the present invention. The first, second, and third versions of the target detection model correspond to different parameter quantities.

[0125] In embodiments of the present invention, the training process for the target detection model may include forward propagation and backward propagation.

[0126] Forward propagation calculates sequentially from the embedding layer to the processing layer based on the parameters of the target detection model, ultimately obtaining predictive information for the identification result. This predictive information is used to determine the loss information.

[0127] Backward propagation allows for the sequential calculation and updating of the target detection model parameters from the output layer to the input layer, based on loss information. The target detection model typically uses a neural network structure, and its parameters may include parameters such as neural network weights. During the backward propagation process, gradient information for the target detection model parameters can be determined, and this gradient information can be used to update the target detection model parameters. For example, backward propagation can store gradient information for the target detection model parameters by sequential calculation from the processing layer to the embedding layer, based on the chain rule of calculus.

[0128] In this embodiment, the training process for the target detection model is as follows: Step C1 is a step in which a construction image sample is input to a target detection model, and the target detection model outputs a prediction result corresponding to the construction image sample, wherein the prediction result includes prediction box information corresponding to the category of the construction object, and the construction image sample corresponds to the measured box information, and Step C2 determines loss information corresponding to the measured box information and the predicted box information based on a minimum point distance function based on a horizontal rectangle, The process includes step C3, which updates the parameters of the target detection model based on the loss information.

[0129] The process for acquiring the above construction image samples is, specifically, Step D1 involves randomly selecting four original construction images from the construction image set, Step D2 involves performing random enhancement operations on four original construction images to obtain four enhanced construction images, Step D3 involves fusing four enhanced construction images into a single fused image, Step D4 includes marking measured box information on a single fused image to obtain one training sample.

[0130] The random enhancement operations described above specifically include at least one of the following operations: inversion, random scaling, random color transformation, and random perspective transformation.

[0131] The process of merging the four enhanced construction images in step D3 into a single fused image specifically includes the steps of: positioning the four enhanced construction images into a single intermediate image according to the offsets [0,0], [0,243], [320,320], and [320,0]; cropping the portion of the intermediate image that exceeds the size range and reducing the range of the measurement box so as not to cross the boundary, thereby finally obtaining a single fused image. The size range specifically includes the coordinate range corresponding to [0,0] to [320,320].

[0132] Referring to equation (1), it shows a process for determining loss information corresponding to the measured box information and the predicted box information based on a minimum point distance function based on a horizontal rectangle. JPEG0007843394000001.jpg140170

[0133] A minimum point distance function based on a horizontal rectangle helps ensure that the predicted box is geometrically close to the measured box, and performs metric optimization by calculating the distance between the top-left and bottom-right corners of the predicted and measured boxes, especially when the predicted and measured boxes have the same aspect ratio but different width and height values. Because the minimum point distance function is more sensitive to positional errors in the predicted box, it can be useful in improving the identification accuracy of target detection models.

[0134] In step 203, based on the above identification results, it can be determined whether or not the construction image is related to abnormal operation.

[0135] Examples of abnormal operation include situations where the distance between a person (construction worker) and construction machinery is shorter than the safe distance, where a person is not wearing safety protective equipment, or where safety signs are not installed in the construction environment corresponding to the construction image.

[0136] In the implementation of the present invention, the category of wearing safety protective equipment includes the category of not wearing safety protective equipment. The process for determining whether the above construction image is related to abnormal operation based on the above identification results is, specifically, If the identification result includes the category of "not wearing safety protective equipment", step C1 determines that the construction image is related to abnormal operation, or Step C2 includes determining distance information between the human category and the construction machine category based on the location information corresponding to the human category and the construction machine category included in the identification result, and determining whether the construction image is related to abnormal operation based on the distance information.

[0137] In one example, the process for determining distance information between the human category and the construction machinery category specifically includes first determining the image distance between the human category and the construction machinery category based on the location information corresponding to the human category and the construction machinery category included in the identification result, and then converting the image distance to an actual distance based on the scale factor of the construction image, the actual distance being usable as distance information between the human category and the construction machinery category.

[0138] If the distance value corresponding to the above distance information is shorter than the safe distance, it can be determined that the construction image is related to abnormal operation. If the above construction image is related to abnormal operation, the embodiment of the present invention can send corresponding presentation information to a pre-configured user so that the pre-configured user can process the abnormal operation.

[0139] In summary, the object identification method of the embodiment of the present invention uses an image recognition method to monitor abnormal behavior in construction images in real time, thereby saving labor costs required for safety monitoring at construction sites. Furthermore, this real-time monitoring can improve the accuracy of safety monitoring at construction sites and reduce the incidence of safety accidents.

[0140] First, the n feature division modules of the feature division network in the embodiment of the present invention can aggregate sub-image features at multiple levels. Based on this, the n feature fusion modules of the second backbone network can effectively fuse sub-image features at multiple levels. Thus, in the forward propagation stage of the target detection model, richer image features can be provided, and the expressive power of the second image features output by the second backbone network can be enhanced.

[0141] Furthermore, in the reverse propagation calculation process of the target detection model, the calculation of the deviation between aggregated multi-level sub-image feature information and the true value transmits more reliable gradient information, guiding parameter learning of the target detection model. This helps the first backbone network extract more important and accurate image features. Therefore, the embodiment of this application can further improve the accuracy of the identification results of construction images, and based on this, the accuracy of safety monitoring at construction sites can be improved.

[0142] For the sake of simplicity, the embodiments of the method are all expressed as a combination of a series of operations. However, those skilled in the art will know that, according to the embodiments of this application, certain steps can be performed in other orders or simultaneously, and therefore the embodiments of this application are not limited to the order of operations described. Furthermore, those skilled in the art should know that all embodiments described in the specification belong to preferred embodiments, and that the operations relating to them are not necessarily essential to the embodiments of this application.

[0143] Based on the above embodiment, this embodiment further provides an object identification device, which, with reference to Figure 11, may specifically include a collection module 1101, an object identification module 1102, and an anomaly determination module 1103.

[0144] The collection module 1101 is used to collect construction images. The object identification module 1102 is used to identify the construction image and obtain identification results based on the target detection model, the identification results include the category of the construction object and the location information of the construction object category in the construction image, the category of the construction object includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category. The abnormality detection module 1103 is used to determine whether the construction image is related to abnormal operation based on the identification result. The target detection model includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network includes n feature segmentation modules, the second backbone network includes n feature fusion modules, and at least some of the feature fusion modules are connected to channel-to-pixel modules, where n is a positive integer greater than 1. Specifically, the object identification module 1102 is: A first image feature determination module 1121 for determining n types of first image features corresponding to the construction image, using the first backbone network, A convolutional partitioning module 1122 that uses the i-th feature partitioning module in a feature partitioning network to perform convolution and partitioning on a first image feature of type i, the resulting partitioning process includes i-th type sub-image features, and outputs j-th type sub-image features to a connected j-th feature fusion module, wherein i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the i-th type sub-image features output by the same i-th feature partitioning module are sub-image features with different numbers of channels, one feature fusion module corresponds to one type of channel count, and different feature fusion modules correspond to different numbers of channels, A first fusion module 1123 in the second backbone network performs a first fusion process on at least one sub-image feature using the i-th feature fusion module in the second backbone network to obtain a first fusion result, and uses the channel-to-pixel module in the second backbone network to determine a second image feature based on the first fusion result and outputs the second image feature to the neck network. A second fusion module 1124 is provided to perform a second fusion process on the aforementioned second image features using a neck network and to obtain the second fusion processing result, The system includes an identification result determination module 1125 for determining the identification result based on the second fusion processing result using a detection head network.

[0145] Selectable, the categories for wearing safety protective equipment include the category for not wearing safety protective equipment. The aforementioned abnormality detection module is If the identification result includes the category of "not wearing safety protective equipment," then a first abnormality determination module determines that the construction image is related to abnormal operation, or The system includes a second abnormality determination module that determines distance information between the human category and the construction machine category based on the location information corresponding to the human category and the construction machine category included in the identification result, and determines whether the construction image is related to abnormal operation based on the distance information.

[0146] Selectively, the convolutional partitioning module is, A convolution module that uses the i-th feature partitioning module in the feature partitioning network to perform a convolution on the i-th type first image feature to obtain an image feature with c channels, A partitioning module for dividing an image feature with c channels into i types of sub-image features with different numbers of channels, depending on the channel dimension, wherein the number of channels in the j-th sub-image feature in the i types of sub-image features is t*2 (j-1 ) includes a split module and .

[0147] Optionally, the second backbone network further includes a first convolutional module connected prior to the feature fusion module. The first fusion module described above is A receiving module for receiving n-i+1 types of sub-image features and convolutional image features output by a first convolutional module, wherein the number of channels corresponding to the n-i+1 types of sub-image features is the first number of channels, and the number of channels corresponding to the convolutional image features is the second number of channels, An interpolation module for performing interpolation on n-i+1 types of sub-image features using the i-th feature fusion module, wherein the number of channels corresponding to the n-i+1 types of sub-image features after interpolation is the number of second channels, and It includes a fusion processing module that uses the i-th feature fusion module to perform fusion processing on the n-i+1 types of sub-image features and convolutional image features after interpolation processing.

[0148] Optionally, a first convolution module connected before the first feature fusion module is used to perform convolution on the construction image and output the corresponding convolutional image features to the first feature fusion module.

[0149] Selectively, the first backbone network includes a second convolutional module and n-1 processing units connected in sequence, the processing units including a third convolutional module and a channel-to-pixel module, The second convolution module is connected to the first feature segmentation module and used to output a first image feature of the first kind to the first feature segmentation module. The channel-to-pixel modules included in the n-1 processing units are each connected to the corresponding n-1 feature segmentation modules and used to output a first image feature to the corresponding feature segmentation modules.

[0150] Selectively, the training process for the target detection model is: A step comprising inputting a construction image sample into a target detection model, the target detection model outputting a prediction result corresponding to the construction image sample, wherein the prediction result includes prediction box information corresponding to the category of the construction object, and the construction image sample corresponds to the measured box information, A step of determining loss information corresponding to the measured box information and the predicted box information based on a minimum point distance function based on a horizontal rectangle, The process includes the step of updating the parameters of the target detection model based on the loss information.

[0151] In summary, the object identification device of the embodiment of the present invention can monitor abnormal behavior in construction images in real time using an image recognition method, thereby saving labor costs required for safety monitoring at construction sites. Furthermore, this real-time monitoring can improve the accuracy of safety monitoring at construction sites and reduce the incidence of safety accidents.

[0152] First, the n feature division modules of the feature division network in the embodiment of the present invention can aggregate sub-image features at multiple levels. Based on this, the n feature fusion modules of the second backbone network can effectively fuse sub-image features at multiple levels. Thus, in the forward propagation stage of the target detection model, richer image features can be provided, and the expressive power of the second image features output by the second backbone network can be enhanced.

[0153] Furthermore, in the reverse propagation calculation process of the target detection model, the calculation of the deviation between aggregated multi-level sub-image feature information and the true value transmits more reliable gradient information, guiding parameter learning of the target detection model. This helps the first backbone network extract more important and accurate image features. Therefore, the embodiment of this application can further improve the accuracy of the identification results of construction images, and based on this, the accuracy of safety monitoring at construction sites can be improved.

[0154] Embodiments of the present invention further provide a non-volatile readable storage medium having one or more modules (programs) stored therein, which, when applied to a device, can cause the device to execute the instructions for each step of the embodiment of the present invention.

[0155] Embodiments of the present application provide one or more machine-readable media on which instructions are stored, causing an electronic device to perform one or more methods of the above embodiments when executed by one or more processors. In embodiments of the present application, the electronic device includes various devices such as terminal devices and servers (clusters).

[0156] The embodiments of this disclosure can be implemented as a device that uses any suitable hardware, firmware, software, or any combination thereof to achieve a desired configuration, and such device may include electronic equipment such as terminal devices and servers (clusters). Figure 12 schematically shows an exemplary device 1300 that can be used to implement each embodiment described herein.

[0157] In one embodiment, Figure 12 shows an exemplary device 1300, which includes one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the (one or more) processors 1302, a memory 1306 coupled to the control module 1304, an NVM (non-volatile memory) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0158] The processor 1302 may include one or more single-core or multi-core processors, and the processor 1302 may include any combination of general-purpose processors or dedicated processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 1300 can be used as a terminal device, server (cluster), or other device as described in the embodiments of the present application.

[0159] In some embodiments, the apparatus 1300 may include one or more computer-readable media having instructions 1314 (e.g., memory 1306 or non-volatile memory / storage device 1308), and one or more processors 1302 configured to execute the instructions 1314 and implement a module in combination with the one or more computer-readable media to perform the operations described herein.

[0160] In one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the (one or more) processors 1302 and / or any suitable device or component communicating with the control module 1304.

[0161] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0162] Memory 1306 can be used, for example, to load and store data and / or instructions 1314 of device 1300. In one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM (Dynamic Random Access Memory). In some embodiments, memory 1306 may include DDR4 synchronous dynamic random access memory.

[0163] In one embodiment, the control module 1304 may include one or more input / output controllers to interface with the non-volatile memory / storage device 1308 and (one or more) input / output devices 1310.

[0164] For example, the non-volatile memory / storage device 1308 can be used to store data and / or instructions 1314. The non-volatile memory / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives, one or more optical disk drives, and / or one or more digital multipurpose optical disk drives).

[0165] The non-volatile memory / storage device 1308 may include storage resources that are a physical part of the device on which the device 1300 is installed, or it may be accessible by the device but not necessarily part of the device. For example, the non-volatile memory / storage device 1308 can be accessed via a network through one or more input / output devices 1310.

[0166] One or more input / output devices 1310 can interface to the device 1300 to communicate with any other suitable devices, and the input / output devices 1310 may include communication components, audio components, sensor components, etc. A network interface 1312 can interface to the device 1300 to communicate over one or more networks, and the device 1300 can wirelessly communicate with one or more components of a wireless network based on any standard and / or protocol in one or more wireless network standards and / or protocols, for example, by accessing and wirelessly communicating with wireless networks based on communication standards, such as WiFi (Wireless Fidelity), 2G (2nd Generation wireless telephone technology), 3G (3rd Generation wireless telephone technology), 4G (4th Generation wireless telephone technology), 5G (5th Generation wireless telephone technology), etc., or a combination thereof.

[0167] In one embodiment, at least one of the (one or more) processors 1302 can be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 1304. In one embodiment, at least one of the (one or more) processors 1302 can be packaged together with the logic of one or more controllers of the control module 1304 to form a system-in-package. In one embodiment, at least one of the (one or more) processors 1302 can be integrated onto the same die as the logic of one or more controllers of the control module 1304. In one embodiment, at least one of the (one or more) processors 1302 can be integrated onto the same die as the logic of one or more controllers of the control module 1304 to form a system-on-chip.

[0168] In each embodiment, the apparatus 1300 may include, but is not limited to, a server, a desktop computing device, or a terminal device such as a mobile computing device (e.g., a laptop computing device, a handheld computing device, a touchscreen device, a netbook, etc.). In each embodiment, the apparatus 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, the apparatus 1300 includes one or more video cameras, a keyboard, a liquid crystal display screen (touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit, and a speaker.

[0169] In the detection device, a master chip can be used as the processor or control module, sensor data, location information, etc. are stored in memory or a non-volatile memory / storage device, the sensor group can be used as input / output devices, and the communication interface may include a network interface.

[0170] In the case of the apparatus embodiment, since it is essentially the same as the method embodiment, the explanation is relatively simple, and relevant parts should be referred to the partial explanation of the method embodiment.

[0171] Each example in this specification is described step by step, with each example focusing on its differences from the others, and identical and similar parts between the examples should be referred to from one another.

[0172] Embodiments of the present application will be described with reference to flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products relating to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions are provided to the processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing terminal device, and the instructions executed by the processor of the computer or other programmable data processing terminal device can generate a machine that generates an apparatus for realizing the functions specified in one flow of a flowchart or one or more blocks of multiple flows and / or block diagrams.

[0173] These computer program instructions may also be stored in computer-readable memory that can operate a computer or other programmable data processing terminal device in a particular way, thereby generating a product that includes an instruction unit that implements a function specified for one or more flows in a flowchart and / or one or more blocks in a block diagram.

[0174] These computer program instructions may be loaded onto a computer or other programmable data processing terminal device, thereby executing a series of operational steps on the computer or other programmable terminal device to generate computer implementation processing, and the instructions executed on the computer or other programmable terminal device provide steps to realize the functions specified in one or more flows of a flowchart and / or one or more blocks of a block diagram.

[0175] While preferred embodiments of the embodiments of this application have been described, those skilled in the art, knowing the basic creative concepts, can make additional changes and modifications to these embodiments. Therefore, the appended claims are intended to be construed as including all changes and modifications that fall within the scope of the preferred embodiments and embodiments of this application.

[0176] Finally, it should be noted that, in this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply that such an actual relationship or order exists between these entities or operations. Furthermore, the terms "includes," "contains," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device containing a set of elements also includes other elements not expressly described, or elements specific to such a process, method, article, or terminal device. Unless further limited, an element defined by the phrase "contains" does not preclude the presence of other identical elements in a process, method, article, or terminal device containing that element.

[0177] The object identification method, apparatus, electronic device, and machine-readable medium according to the present invention have been described in detail above. While this specification has used specific examples to illustrate the principles and embodiments of the present application, the above description of the embodiments is merely intended to aid in understanding the method and core concept of the present application. Furthermore, those skilled in the art will recognize that there are modifications in specific embodiments and scope of application based on the concept of the present application, and therefore, the contents of this specification should not be understood as limiting the present application.

Claims

1. A method for identifying an object at a construction site based on dual backbone fusion, wherein the object identification method is: The collection module collects construction images, A step in which an object identification module identifies the construction image based on a target detection model and obtains an identification result, wherein the identification result includes the category of the construction object and the location information of the construction object category in the construction image, and the category of the construction object includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category. The abnormality detection module includes the step of determining whether the construction image is related to abnormal operation based on the identification result, The target detection model includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network includes n feature segmentation modules, the second backbone network includes n feature fusion modules, and at least some of the feature fusion modules are connected to channel-to-pixel modules, where n is a positive integer greater than 1. The step of identifying the construction image based on the target detection model is: The first backbone network includes the step of determining n types of first image features corresponding to the construction image, The i-th feature segmentation module in the feature segmentation network performs convolution and segmentation on a first image feature of type i, and the resulting segmentation process includes i sub-image features, which are output to a connected j-th feature fusion module, wherein i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the i sub-image features of type i output by the same i-th feature segmentation module are sub-image features with different numbers of channels, one feature fusion module corresponds to one number of channels, and different feature fusion modules correspond to different numbers of channels. The i-th feature fusion module in the second backbone network performs a first fusion process on at least one sub-image feature to obtain a first fusion result, and the channel-to-pixel module in the second backbone network determines at least one second image feature based on the first fusion result and outputs the second image feature to the neck network. The neck network performs a second fusion process on two or more of the second image features to obtain the result of the second fusion process. A method for identifying an object at a construction site based on dual backbone fusion, characterized in that the detection head network includes the step of determining the identification result based on the second fusion processing result.

2. The aforementioned categories of safety protective equipment wearing include categories of safety protective equipment not wearing, The step of determining whether the construction image is related to abnormal operation based on the identification result is: If the identification result includes the category of "not wearing safety protective equipment," the step of determining that the construction image is related to abnormal operation, or The method according to claim 1, characterized by comprising the steps of determining distance information between the human category and the construction machine category based on location information corresponding to the human category and the construction machine category included in the identification result, and determining whether or not the construction image is related to abnormal operation based on the distance information.

3. The step in which the i-th feature segmentation module in the feature segmentation network performs convolution and segmentation on the i-th type first image feature is: The i-th feature segmentation module in the feature segmentation network performs a convolution operation on the i-th type first image feature to obtain an image feature with c channels. The step of dividing an image feature with c channels into i types of sub-image features with different numbers of channels, according to the channel dimension, wherein the number of channels in the j-th sub-image feature in the i types of sub-image features is t*2 (j-1) The method according to claim 1, comprising the steps of:

4. The second backbone network further includes a first convolutional module connected in front of the feature fusion module, The step of the i-th feature fusion module in the second backbone network performing a first fusion process on at least one sub-image feature is: The i-th feature fusion module receives n-i+1 types of sub-image features and convolutional image features output by the first convolutional module, wherein the number of channels corresponding to the n-i+1 types of sub-image features is the first number of channels, and the number of channels corresponding to the convolutional image features is the second number of channels. The i-th feature fusion module is a step of performing interpolation on n-i+1 types of sub-image features, wherein the number of channels corresponding to the n-i+1 types of sub-image features after interpolation is the second number of channels. The method according to claim 1, characterized in that the i-th feature fusion module includes the step of performing a fusion process on n-i+1 types of sub-image features and convolutional image features after interpolation processing.

5. The method according to 4, characterized in that a first convolution module connected before the first feature fusion module is used to perform convolution on a construction image and output the corresponding convolutional image features to the first feature fusion module.

6. n feature partitioning modules include a first feature partitioning module, a second feature partitioning module, a third feature partitioning module, a fourth feature partitioning module, and a fifth feature partitioning module, and n feature fusion modules include a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, and a fifth feature fusion module, The segmentation result of the first feature segmentation module includes a first sub-image feature of one type m channel, and this first sub-image feature of one type m channel is output to the first feature fusion module. The splitting result of the second feature splitting module includes a second sub-image feature of one type m channel and a third sub-image feature of one type 2m channel. The second sub-image feature of one type m channel is output to the first feature fusion module, and the third sub-image feature of one type 2m channel is output to the second feature fusion module. The splitting result of the third feature splitting module includes a fourth sub-image feature of one type m channel, a fifth sub-image feature of one type 2m channel, and a sixth sub-image feature of one type 4m channel. The fourth sub-image feature of one type m channel is output to the first feature fusion module, the fifth sub-image feature of one type 2m channel is output to the second feature fusion module, and the sixth sub-image feature of one type 4m channel is output to the third feature fusion module. The segmentation result of the fourth feature segmentation module includes a seventh sub-image feature of one type m channel, an eighth sub-image feature of one type 2m channel, a ninth sub-image feature of one type 4m channel, and a tenth sub-image feature of one type 8m channel. The seventh sub-image feature of one type m channel is output to the first feature fusion module 331, the eighth sub-image feature of one type 2m channel is output to the second feature fusion module, the ninth sub-image feature of one type 4m channel is output to the third feature fusion module 333, and the tenth sub-image feature of one type 8m channel is output to the fourth feature fusion module 334. The method according to claim 1, wherein the division processing result of the fifth feature division module includes a 11th sub-image feature of one type m channel, a 12th sub-image feature of one type 2m channel, a 13th sub-image feature of one type 4m channel, a 14th sub-image feature of one type 8m channel, and a 15th sub-image feature of one type 16m channel, the 11th sub-image feature of one type m channel is output to the first feature fusion module, the 12th sub-image feature of one type 2m channel is output to the second feature fusion module, the 13th sub-image feature of one type 4m channel is output to the third feature fusion module, the 14th sub-image feature of one type 8m channel is output to the fourth feature fusion module, and the 15th sub-image feature of one type 16m channel is output to the fifth feature fusion module.

7. The method according to claim 1, characterized in that the second backbone network includes, in order, a first convolutional module A, a first feature fusion module, a first convolutional module B, a second feature fusion module, a first channel-to-pixel module, a first convolutional module C, a third feature fusion module, a second channel-to-pixel module, a first convolutional module D, a fourth feature fusion module, a third channel-to-pixel module, a first convolutional module E, a fifth feature fusion module, and a fourth channel-to-pixel module.

8. The first convolution module A is used to perform convolution on the construction image and output the corresponding convolutional image feature A to the first feature fusion module. The first feature fusion module receives five types of sub-image features A from the first feature segmentation module, receives convolutional image features A from the first convolution module A, performs interpolation on the five types of sub-image features A, assumes that the number of channels corresponding to the five types of sub-image features A after interpolation is the second number of channels, and assumes that a fusion process is performed on the five types of sub-image features A after interpolation and the convolutional image features A to obtain the first fusion result A. The first convolution module B is used to perform a convolution operation on the first fusion result A to obtain the convolutional image feature B. The second feature fusion module receives four types of sub-image features B from the second feature segmentation module, receives convolutional image features B from the first convolution module B, performs interpolation on the four types of sub-image features B, assumes that the number of channels corresponding to the four types of sub-image features B after interpolation is the second channel number, and assumes that a fusion process is performed on the four types of sub-image features B after interpolation and the convolutional image features B to obtain the first fusion result B. The first channel two-pixel module is used to determine intermediate image features based on the first fusion result B. The first convolution module C is used to perform a convolution operation on intermediate image features to obtain convolutional image features C. The third feature fusion module receives three types of sub-image features C from the third feature segmentation module, receives convolutional image features C from the first convolution module C, performs interpolation on the three types of sub-image features C, assumes that the number of channels corresponding to the three types of sub-image features C after interpolation is the second number of channels, and assumes that a fusion process is performed on the three types of sub-image features C and the convolutional image features C to obtain the first fusion result C. The second channel two-pixel module is used to determine the second image feature A based on the first fusion result C and to output the second image feature A to the neck network. The first convolution module D is used to perform a convolution operation on the second image feature A to obtain the convolutional image feature D. The fourth feature fusion module receives two types of sub-image features D from the fourth feature segmentation module, receives convolutional image features D from the first convolution module D, performs interpolation on the three types of sub-image features D, assumes that the number of channels corresponding to the two types of sub-image features D after interpolation is the second channel number, and assumes that fusion is performed on the two types of sub-image features D and the convolutional image features D after interpolation to obtain the first fusion result D. The third channel two-pixel module is used to determine the second image feature B based on the first fusion result D and to output the second image feature B to the neck network. The first convolution module E is used to perform a convolution operation on the second image feature B to obtain the convolutional image feature E. The fifth feature fusion module receives one type of sub-image feature E from the fifth feature segmentation module, receives a convolutional image feature E from the first convolution module E, performs interpolation on the one type of sub-image feature E, assumes that the number of channels corresponding to the one type of sub-image feature E after interpolation is the second number of channels, and assumes that a fusion process is performed on the one type of sub-image feature E after interpolation and the convolutional image feature E to obtain a first fusion result E. The method according to 7, characterized in that the fourth channel two-pixel module is used to determine a second image feature C based on the first fusion result E and to output the second image feature C to the neck network.

9. The first backbone network includes a second convolutional module and n-1 processing units connected in sequence, and each processing unit includes a third convolutional module and a channel-to-pixel module. The method according to claim 1, wherein the second convolution module is connected to the first feature segmentation module and used to output a first image feature of the first kind to the first feature segmentation module, and the channel-to-pixel modules included in the n-1 processing units are each connected to the corresponding n-1 feature segmentation modules and used to output a first image feature to the corresponding feature segmentation modules.

10. In the training process of the aforementioned target detection model, A construction image sample is input to a target detection model, the target detection model outputs a prediction result corresponding to the construction image sample, the prediction result includes prediction box information corresponding to the category of the construction object, and the construction image sample corresponds to the measured box information. Based on the minimum point distance function based on the horizontal rectangle, loss information corresponding to the measured box information and the predicted box information is determined. The method according to any one of 1 to 9, characterized in that the parameters of the target detection model are updated based on the loss information.

11. An object identification device, wherein the object identification device is A collection module for collecting construction images, An object identification module for identifying a construction image and obtaining an identification result based on a target detection model, wherein the identification result includes the category of the construction object and the location information of the construction object category in the construction image, and the category of the construction object includes at least one of the following: human category, construction machinery category, safety protective equipment wearing category, and safety sign category. The module includes an abnormality determination module for determining whether the construction image is related to abnormal operation based on the identification result, The target detection model includes a first backbone network, a feature segmentation network, a second backbone network, a neck network, and a detection head network, wherein the feature segmentation network includes n feature segmentation modules, the second backbone network includes n feature fusion modules, and at least some of the feature fusion modules are connected to channel-to-pixel modules, where n is a positive integer greater than 1. The object identification module is A first image feature determination module for determining n types of first image features corresponding to the construction image, using the first backbone network, A convolutional partitioning module that uses the i-th feature partitioning module in a feature partitioning network to perform convolution and partitioning on a first image feature of type i, the resulting partitioning process includes i-th type sub-image features, and outputs the j-th type sub-image features to a connected j-th feature fusion module, wherein i and j are positive integers, i is less than or equal to n, j is less than or equal to i, and when i is greater than 1, the i-th type sub-image features output by the same i-th feature partitioning module are sub-image features with different numbers of channels, one feature fusion module corresponds to one type of channel count, and different feature fusion modules correspond to different numbers of channels. A first fusion module that uses the i-th feature fusion module in the second backbone network to perform a first fusion process on at least one sub-image feature to obtain a first fusion result, and uses the channel-to-pixel module in the second backbone network to determine at least one second image feature based on the first fusion result, and outputs the second image feature to the neck network, A second fusion module for performing a second fusion process on two or more of the aforementioned second image features using a neck network and obtaining the result of the second fusion process, An object identification device characterized by including an identification result determination module for determining the identification result based on the second fusion processing result using a detection head network.

12. It is an electronic device, Processor and An electronic device comprising a memory in which executable code is stored, wherein when the executable code is executed, the processor is instructed to execute the method according to any one of claims 1 to 9.

13. A machine-readable medium that stores executable code and, when the executable code is executed, causes a processor to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Safety management support device, safety management support program, and storage medium

    JP2017033047A

  • Information process system, information processing device, server device, program, or method

    JP2021043932A

  • Joint perception model training, joint perception method, device, and medium

    JP2023131117A

  • Image detection method, apparatus, device, medium, and program

    JP2023518160A