Target identification method and device based on intelligent construction site, equipment and medium
By using the smart construction site-based target recognition method on construction sites, identifying and tracking target objects on construction sites, the problem of difficulty in full coverage of manual patrols is solved, and intelligent early warning and cost reduction for construction site safety is achieved.
Patent Information
- Application Number
- CN202510164795.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In construction sites, manual inspections and records are difficult to cover all areas in full, resulting in many challenges in safety management and high labor costs.
The target recognition method based on smart construction sites is adopted to obtain real-time monitoring videos, identify and mark target objects with activity capabilities, such as construction machinery and personnel, and form historical sequences to trace the position changes of the target objects and comprehensively determine whether safety warning information is generated.
The monitoring accuracy of key elements of the construction site has been improved, potential safety hazards are discovered in a timely manner, intelligent early warning of construction site safety has been achieved, and labor costs have been reduced.
Smart Images

Figure CN120126073A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a target recognition method, device, equipment and medium based on a smart construction site. Background Art
[0002] With the acceleration of urbanization, safety management of construction sites has become an increasingly important issue. Despite the development of information technology in recent years, manual methods are still widely used for safety management on many small and medium-sized construction sites. Although manual inspections and records can timely discover some safety hazards, due to limited manpower, it is impossible to fully cover all areas, resulting in many challenges in safety management.
[0003] At present, the commonly used target recognition method on construction sites is generally manual observation, that is, workers or managers personally patrol the site and identify target objects through visual inspection. This method ensures the safe management of the construction site to a certain extent, but in the complex and changeable construction site environment, in order to ensure the safe management of the construction site, a lot of manpower costs are required, resulting in a waste of manpower costs. Summary of the invention
[0004] In order to reduce the manpower cost in the target identification process on the construction site, the present application provides a target identification method, device, equipment and medium based on a smart construction site.
[0005] In the first aspect, the present application provides a target recognition method based on a smart construction site, which adopts the following technical solutions: A target recognition method based on a smart construction site, comprising: Obtaining a real-time monitoring video corresponding to the current construction site, and extracting a current image corresponding to the current moment from the real-time monitoring video; Identify target objects in the current image and mark each target object on the current image, wherein the target object is an object capable of moving on the current construction site; Obtain the historical annotated images corresponding to each historical moment, and sort the historical annotated images according to the shooting time to obtain a historical sequence; Determining a position change of each target object in the historical sequence; Based on the position change of each target object and the target object in the current image, it is determined whether to generate safety warning information.
[0006] By adopting the above technical solution, by obtaining the real-time monitoring video of the current construction site and extracting the current image, the target object in the current image is identified and annotated, especially for objects with the ability to move, such as construction machinery, personnel, etc., the monitoring accuracy of key elements of the construction site is improved, which helps to timely discover potential safety hazards. By obtaining historical annotated images and forming historical sequences, the historical position of the target object can be traced, providing a data basis for analyzing the movement trajectory of the object, which helps to grasp the overall trend of construction site activities. Then, the position change of each target object in the historical sequence is determined, and the dynamic behavior of the object can be accurately captured. Based on the position change of the target object and the target object in the current image, it is comprehensively judged whether it is necessary to generate safety warning information, realizing intelligent warning of construction site safety. In this process, no human participation is required, thereby reducing the labor cost of target identification in the construction site.
[0007] In a possible implementation manner, identifying the target object in the current image includes: Using an edge recognition algorithm to recognize the current image to obtain a current object in the current image; Calculating the similarity between each current object in the current image and a preset object; Comparing the similarity between each current object and the preset object with the corresponding similarity threshold respectively, to obtain a comparison result corresponding to each current object, wherein the comparison result is greater than, less than, or equal to; The current object whose comparison result is greater than is determined as the target object.
[0008] By adopting the above technical solution, the edge recognition algorithm is used to identify the current image. This step can efficiently separate the contour of the object in the image from the complex background, which not only improves the accuracy of recognition, but also greatly reduces the computational complexity. By calculating the similarity between each current object and the preset object in the current image, the quantitative comparison of the object features is realized, and the similarity between each current object and the preset object is compared with the corresponding similarity threshold to ensure that only sufficiently similar objects are identified as target objects, thereby avoiding misidentification. Finally, the current object with a comparison result greater than is determined as the target object, making the entire recognition process both efficient and accurate.
[0009] In a possible implementation manner, marking each target object on the current image includes: Determine the object category and object name corresponding to each target object, wherein the object type includes a strong object and a weak object; Based on the object category and the object name corresponding to each target object, a first label corresponding to each target object is generated, and each target object is labeled based on the first label corresponding to each target object.
[0010] By adopting the above technical solution, the object category and object name corresponding to each target object are determined, and the target object is accurately identified through detailed classification and naming. Classifying objects into categories such as strong objects (such as large construction machinery) and weak objects (such as construction personnel) will help to carry out different treatments and management according to the importance and risk level of the objects in the future. Based on the object category and object name corresponding to each target object, the first label corresponding to each target object is generated, and the key information of the target object is presented intuitively and concisely. Each target object is labeled based on the first label corresponding to each target object, and the key information of the target object is directly displayed on the image.
[0011] In a possible implementation manner, determining the position change of each target object in the historical sequence includes: Obtaining basic information of a target object having a first label in each historically annotated image, the basic information including object color and object texture; Based on the object color and object texture corresponding to each target object in each historically annotated image, generate a second label corresponding to each target object; Calculating the similarity between any two target objects based on the first label and the second label, and determining the associated object of each target object in any historical annotated image based on the similarity between the any two target objects, wherein the any two target objects are from two different historical annotated images; Based on the associated objects of each target object in any historical annotated image, a position change of each target object in the historical sequence is determined.
[0012] By adopting the above technical scheme, the basic information of the target object with the first label in each historical annotated image is obtained, and the second label corresponding to each target object is generated based on the object color and object texture corresponding to each target object in each historical annotated image. Then, the similarity between any two target objects is calculated based on the first label and the second label, and the associated object of each target object in any historical annotated image is determined. Therefore, by comprehensively considering the category and appearance characteristics of the target object, accurate matching of the target object between different historical annotated images is achieved, which is helpful to track the movement trajectory of the target object and provide reliable data support for subsequent position change analysis. Based on the associated objects of each target object in any historical annotated image, the position change of each target object in the historical sequence is determined. By comparing the position information of the target object at different historical moments, the position change of the target object is obtained, which not only helps to grasp the overall trend of construction site activities, but also provides a strong decision-making basis for safety warning and construction management.
[0013] In a possible implementation, based on an associated object of each target object in any historical annotated image, a position change of each target object in the historical sequence is determined, and the process of determining the position change of each target object includes: Acquire a historical position and a current position corresponding to the target object, wherein the historical position is a position of an associated object corresponding to the target object in the historical sequence, and the current position is a position of the target object in the current image; Put each historical position into the current image in sequence according to the order of each historical annotated image in the historical sequence, and generate a third label corresponding to each historical position according to the order of each historical annotated image in the historical sequence; Annotate the current image based on the third label corresponding to each historical position to obtain a target image; Based on each historical position and the current position in the target image, a position change of the target object is determined.
[0014] By adopting the above technical solution, the historical position and current position corresponding to the target object are obtained, and each historical position is placed in the current image in the order of each historical annotated image in the historical sequence, and a corresponding third label is generated. The historical position information of the target object can be intuitively presented on the current image through timeline backtracking, which facilitates the visualization analysis of position changes. Then, the current image is annotated based on the third label corresponding to each historical position to obtain the target image, and the position change of the target object is determined based on each historical position and the current position in the target image. By comparing the position information of the target object at different time points, the position change of the target object is obtained, which not only helps to grasp the key information such as the moving trajectory and speed of the target object, but also provides a strong decision-making basis for safety warning, construction management, etc.
[0015] In a possible implementation, determining whether to generate safety warning information based on the position change of each target object and the target object in the current image includes: Based on the position change of each target object, predict the motion trajectory of each target object in the current image, wherein the motion trajectory includes the motion speed and the motion route; Determine whether there are overlapping trajectory points in the motion trajectories of any two target objects; If there are overlapping trajectory points in the respective motion trajectories of the arbitrary two target objects, the arbitrary two target objects are respectively determined as a first object and a second object, first labels corresponding to the first object and the second object are obtained, and a first influence degree and a second influence degree are determined based on the first labels corresponding to the first object and the second object, and whether to generate safety warning information is determined based on the first influence degree and the second influence degree; If there are no overlapping trajectory points in the motion trajectories of any two target objects, it is determined that no safety warning information is generated.
[0016] By adopting the above technical solution, by predicting the motion trajectory of each target object in the current image, including the motion speed and the motion route, and by determining whether there are overlapping trajectory points in the respective motion trajectories of any two target objects, the potential conflict or collision between the target objects is predicted. For the target objects with overlapping trajectory points, by obtaining their respective corresponding first labels, and determining the first impact degree and the second impact degree based on the first label, the warning information is made more accurate and targeted by fully considering the category and characteristics of the target object. For example, for the potential collision between large construction machinery and construction personnel, the warning information will be more urgent and important, so that timely measures can be taken to protect the safety of construction personnel. Based on the first impact degree and the second impact degree, it is decided whether to generate safety warning information. If the impact degree is high, the warning information is generated to remind the relevant personnel to pay attention to safety; if the impact degree is low or there is no potential danger, the warning information is not generated to avoid unnecessary panic and interference, thereby ensuring the accuracy and effectiveness of the warning information and improving the intelligent level of construction site safety management.
[0017] In the second aspect, the present application provides a target recognition device based on a smart construction site, which adopts the following technical solutions: A target recognition device based on a smart construction site, comprising: An acquisition module is used to acquire a real-time monitoring video corresponding to the current construction site, and extract a current image corresponding to the current moment from the real-time monitoring video; A recognition module, used to recognize target objects in the current image and mark each target object on the current image, wherein the target object is an object capable of moving on the current construction site; A sorting module obtains the historical annotated images corresponding to each historical moment, and sorts the historical annotated images according to the shooting time to obtain a historical sequence; A first determination module is used to determine the position change of each target object in the historical sequence; The second determination module is used to determine whether to generate safety warning information based on the position change of each target object and the target object in the current image.
[0018] In a possible implementation, when identifying the target object in the current image, the recognition module is specifically configured to: Using an edge recognition algorithm to recognize the current image to obtain a current object in the current image; Calculating the similarity between each current object in the current image and a preset object; Comparing the similarity between each current object and the preset object with the corresponding similarity threshold respectively, to obtain a comparison result corresponding to each current object, wherein the comparison result is greater than, less than, or equal to; The current object whose comparison result is greater than is determined as the target object.
[0019] In a third aspect, the present application provides an electronic device, which adopts the following technical solution: An electronic device, comprising: at least one processor; Memory; At least one application, wherein at least one application is stored in a memory and configured to be executed by at least one processor, and the at least one application is configured to: execute the target recognition method based on the smart construction site described in the first aspect above.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium, comprising: storing a computer program that can be loaded by a processor and execute the target recognition method based on a smart construction site described in the first aspect above.
[0021] In summary, this application includes the following beneficial technical effects: By acquiring the real-time monitoring video of the current construction site and extracting the current image, the target objects in the current image are identified and labeled, especially for objects with the ability to move, such as construction machinery and personnel, etc., which improves the monitoring accuracy of key elements of the construction site and helps to discover potential safety hazards in a timely manner. By acquiring historical annotated images and forming historical sequences, the historical position of the target object can be traced, providing a data basis for analyzing the movement trajectory of the object, which helps to grasp the overall trend of construction site activities. Then, the position change of each target object in the historical sequence is determined, and the dynamic behavior of the object can be accurately captured. Based on the position change of the target object and the target object in the current image, it is comprehensively judged whether it is necessary to generate safety warning information, realizing intelligent warning of construction site safety. In this process, no human participation is required, thereby reducing the labor cost of target identification in the construction site. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of a target recognition method based on a smart construction site provided in an embodiment of the present application; Figure 2 It is a block diagram of a target recognition device based on a smart construction site provided in an embodiment of the present application; Figure 3 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following is combined with Figure 1 -Attached Figure 3 This application is described in further detail.
[0024] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0025] The present application embodiment provides a target recognition method based on a smart construction site, such as Figure 1 As shown, the method provided in the embodiment of the present application is performed by an electronic device, which can be a server or a terminal device, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application. The method includes steps S101 to S105, wherein: Step S101: obtain the real-time monitoring video corresponding to the current construction site, and extract the current image corresponding to the current moment from the real-time monitoring video.
[0026] The electronic device establishes a network connection with the camera installed on the construction site, and uses a video stream transmission protocol (such as RTSP, HTTP-FLV, etc.) to obtain a real-time monitoring video stream. Using a video processing library (such as OpenCV), the image at the current moment is extracted from the video stream at a preset frame rate (such as 1 frame per second) as the current image. For example, assuming that the current time is 10:00:05, the electronic device captures the image at that moment from the real-time monitoring video stream being received according to the set frame rate.
[0027] Step S102: Identify target objects in the current image and mark each target object on the current image.
[0028] The target objects are objects that have the ability to move on the current construction site, such as construction workers, transport vehicles, and transport manipulators.
[0029] Specifically, the current image can be processed using a pre-trained deep learning target detection model (such as YOLO, Faster-RCNN, etc.). More specifically, the current image is input into the pre-trained deep learning target detection model, and the processed current image output by the deep learning target detection model is obtained. Among them, the trained deep learning target detection model is trained with a large number of image samples and annotated image samples. The trained deep learning target detection model will perform feature extraction and classification judgment on each object in the current image, and identify target objects with activity capabilities (such as workers, vehicles, etc.). For each identified target object, the deep learning target detection model annotates the corresponding position information (such as bounding box coordinates) in the current image, and outputs the annotated current image. The electronic device obtains the annotated current image output by the deep learning target detection model.
[0030] Step S103: Obtain the historical annotated images corresponding to the historical moments, and sort the historical annotated images according to the shooting time to obtain a historical sequence.
[0031] The historical sequence is an ordered set of historically annotated images sorted according to the shooting time.
[0032] Specifically, the historical annotated images of the current construction site can be obtained from the database corresponding to the current construction site, and the obtained historical annotated images can be sorted according to the shooting time information recorded in the image metadata. More specifically, the sorting algorithm in the programming language (such as the sorted function in Python) can be used to arrange the images in order from the earliest to the latest shooting time to form a historical sequence. For example, three historical annotated images are obtained, and the shooting times are 10:00:00, 10:00:02, and 10:00:04, respectively. After sorting, a historical sequence arranged in chronological order is formed.
[0033] Step S104: Determine the position change of each target object in the historical sequence.
[0034] The position change includes the moving direction and the moving speed.
[0035] Specifically, for each target object in the historical sequence (the same target object is associated by category and relative position in different images), its position information in each historical annotated image (such as the center coordinates of the bounding box) is extracted. By calculating the difference between the position coordinates of the same target object in adjacent images, the displacement of the target object at adjacent moments is obtained. For example, in one image, the center coordinates of the bounding box of a worker are (x1, y1), and in the next moment image, the center coordinates of the bounding box of the same worker are (x2, y2), then the displacement is ((x2-x1), (y2-y1)).
[0036] Furthermore, the average speed can be obtained by calculating the average displacement value according to the displacement values at multiple adjacent moments, and the motion speed can be obtained. The motion direction can be determined by the direction of the displacement vector, thereby obtaining the position change of each target object in the historical sequence.
[0037] Step 105: Determine whether to generate safety warning information based on the position change of each target object and the target object in the current image.
[0038] Specifically, the preset safety rules are obtained from the database corresponding to the current construction site, and the position change of the target object and the state of the target object in the current image are judged according to the preset safety rules. For example, if the preset rule is "workers are prohibited from staying in the dangerous area for a long time", first determine whether there are workers in the pre-set dangerous area in the current image (by comparing the image coordinates with the coordinate range of the dangerous area), and at the same time combine the position change of the worker in the historical sequence. If it is found that the worker stays in the dangerous area for more than a certain threshold (such as 5 minutes), a safety warning information is generated. Among them, the warning information can include the category, location and description of the violation of the target object, and notify relevant personnel through SMS, APP push or on-site sound and light alarm.
[0039] The embodiment of the present application provides a target recognition method based on a smart construction site. By acquiring the real-time monitoring video of the current construction site and extracting the current image, the target object in the current image is identified and labeled, especially for objects with the ability to move, such as construction machinery, personnel, etc., the monitoring accuracy of key elements of the construction site is improved, which helps to timely discover potential safety hazards. By acquiring historical annotated images and forming historical sequences, the historical position of the target object can be traced, providing a data basis for analyzing the movement trajectory of the object, which helps to grasp the overall trend of construction site activities. Then, the position change of each target object in the historical sequence is determined, and the dynamic behavior of the object can be accurately captured. Based on the position change of the target object and the target object in the current image, it is comprehensively judged whether it is necessary to generate safety warning information, realizing intelligent warning of construction site safety. In this process, no human participation is required, thereby reducing the labor cost in the target recognition process on the construction site.
[0040] In a possible implementation of the embodiment of the present application, in the above step S102, identifying the target object in the current image includes: The current image is identified by using an edge recognition algorithm to obtain the current object in the current image; Calculate the similarity between each current object in the current image and the preset object; The similarity between each current object and the preset object is respectively compared with the corresponding similarity threshold to obtain a comparison result corresponding to each current object, and the comparison result is greater than, less than, or equal to; The current object whose comparison result is greater than is determined as the target object.
[0041] Specifically, after obtaining the current image, the acquired current image can be preprocessed, such as grayscale conversion, to convert the color image into a grayscale image, which can simplify the subsequent processing process. After that, the smoothed image is processed using an edge recognition algorithm (such as the Canny edge detection algorithm). The Canny algorithm will go through a series of steps to first calculate the gradient of each pixel in the image, find points with larger gradient values as possible edge points, then perform non-maximum suppression on these edge points, remove those points that are not true edges, and finally determine which edge points are true edges through double threshold processing, thereby outlining the edge contours of objects in the image. Based on these edge contours, each current object in the current image can be divided.
[0042] Furthermore, multiple preset objects and the similarity threshold corresponding to each preset object are obtained from the database corresponding to the current construction site, and each current object in the current image is compared with each preset object to obtain the similarity between each current object and each preset object. For each current object, the similarity value between it and the preset object and the corresponding similarity threshold are compared. If the similarity is greater than the threshold, it means that the current object is highly similar to the preset object; if the similarity is less than the threshold, it means that the similarity is low; if the similarity is equal to the threshold, it is in a critical state.
[0043] Furthermore, all the comparison results of the current object are traversed, and those current objects whose comparison results are "greater than" are screened out.
[0044] A possible implementation of the embodiment of the present application, in the above embodiment, marking each target object on the current image includes: Determine the object category and object name corresponding to each target object, and the object type includes strong objects and weak objects; Based on the object category and the object name corresponding to each target object, a first label corresponding to each target object is generated, and each target object is labeled based on the first label corresponding to each target object.
[0045] Among them, strong objects usually have larger size, stronger power or destructive power, such as large construction machinery; weak objects are relatively fragile and have weaker self-protection ability, such as workers.
[0046] Specifically, for each target object, a preset object corresponding to the target object is obtained, that is, a preset object corresponding to the target object whose similarity is greater than a corresponding similarity threshold is obtained, and an object name corresponding to the preset object is obtained as the object name corresponding to the target object. The object name corresponding to each preset object is stored in the database corresponding to the current construction site.
[0047] A strong object category table and a weak object category table are obtained from a database corresponding to the current construction site, and the strong object category table and the weak object category table contain various object names respectively. After the electronic device identifies the target object, an object corresponding to the object name of each target object is obtained from the strong object category table or the weak object category table to obtain the object type of each target object, that is, if the object name of the target object is the same as an object name in the strong object category table, then the target object is a strong object; if the object name of the target object is the same as an object name in the weak object category table, then the target object is a weak object.
[0048] Furthermore, a corresponding first label is generated according to the object category and object name of each target object. The first label can be in the form of a text string, containing information about the object category and object name, such as "strong object-construction vehicle" and "weak object-worker". For the generated first label, the position of the target object in the current image is obtained, and its range in the image is generally represented by the bounding box of the target object (such as a rectangular box). Then, the text display of the first label is added near (such as above or next to) the bounding box of the target object on the current image, using a color that has a significant contrast with the image background (such as white font displayed on a black background) to ensure that the label is clearly visible and complete the labeling of the target object.
[0049] A possible implementation of the embodiment of the present application, in the above embodiment, determining the position change of each target object in the historical sequence includes: Obtaining basic information of a target object having a first label in each historically annotated image, the basic information including object color and object texture; Based on the object color and object texture corresponding to each target object in each historically annotated image, generate a second label corresponding to each target object; Calculating the similarity between any two target objects based on the first label and the second label, and determining the associated object of each target object in any historical annotated image based on the similarity between the any two target objects, wherein the any two target objects come from two different historical annotated images; Based on the associated objects of each target object in any historical annotated image, the position change of each target object in the historical sequence is determined.
[0050] Specifically, each historical annotated image and the corresponding annotation information are obtained from the database corresponding to the current construction site, and the annotation information contains relevant content of the target object with the first label. For each target object with the first label, its area in the image is analyzed. More specifically, in terms of color, the color statistics of the pixels in the area can be calculated, such as the main color tone, the distribution ratio of the color, etc. For texture, texture analysis methods, such as grayscale co-occurrence matrix and other technologies, can be used to extract parameters reflecting the surface texture characteristics of the object, such as the roughness of the texture, so as to obtain the two basic information of the color and texture of the target object.
[0051] Furthermore, after obtaining the object color and object texture of each target object, the object color and object texture information of each target object can be integrated and encoded. Specifically, the color information and texture features can be quantified, for example, the colors can be classified and encoded according to a certain color gamut range, and the parameters of the texture features can be converted into a specific numerical combination. Then, these processed color and texture information are combined into a unique identifier, which is the second label. For example, a specific string format can be used to combine the color code and the texture code to form a second label.
[0052] For the first label, if the first labels of the two target objects are exactly the same, it means that they are consistent in object category and name, which will give a certain similarity score; if they are partially the same, the corresponding score will be given according to the proportion of the same part. For the second label, the color and texture encoding of the two target objects can be compared, and the similarity can be measured by calculating the distance between them (such as Euclidean distance). The closer the distance, the higher the similarity. The similarity scores of the first label and the second label are weighted and summed to obtain the comprehensive similarity of the two target objects. Then, for each target object, the target object with the highest similarity in other historical annotated images is found and determined as the associated object of the target object in the corresponding historical annotated image.
[0053] In this embodiment, based on the associated objects of each target object in any historical annotated image, the position change of each target object in the historical sequence is determined, and the process of determining the position change of each target object includes: obtaining the historical position and current position corresponding to the target object, the historical position is the position of the associated object corresponding to the target object in the historical sequence, and the current position is the position of the target object in the current image; placing each historical position in the current image in sequence according to the order of each historical annotated image in the historical sequence, and generating a third label corresponding to each historical position in the order of each historical annotated image in the historical sequence; annotating the current image based on the third label corresponding to each historical position to obtain the target image; and determining the position change of the target object based on each historical position and the current position in the target image.
[0054] After obtaining the associated objects of each target object in any historically annotated image, the position of each associated object can be extracted from the stored historically annotated image data and used as the historical position of the target object. At the same time, the position information of the target object in the current image is read.
[0055] Furthermore, each historical position is taken out in turn according to the shooting time sequence of the historical annotated images in the historical sequence. Since the historical annotated images and the current image may have different coordinate systems or resolutions, the coordinates of the historical positions can be transformed so that they can be accurately mapped to the coordinate system of the current image, and after the transformation is completed, these historical positions are placed at the corresponding positions of the current image.
[0056] Furthermore, for each placed historical position, a third label is generated for it in the order of the historical sequence. The third label can include the shooting time information of the historical annotated image corresponding to the historical position, for example, the label is generated in the form of "timestamp-historical position sequence number", such as "2025010110:00:00-1", where "2025010110:00:00" is the shooting time of the historical annotated image, and "1" represents the sequence number of the historical position in the historical position sequence of the current target object.
[0057] After placing the historical locations in the current image and generating the third label, you can use the image annotation tool to annotate each historical location on the current image. The annotation method can be to display the corresponding third label text near the historical location (such as above the bounding box), using a color that has a clear contrast with the image background (such as white font displayed on a black background) to ensure that the label is clearly visible. After the annotation is completed, an image containing all the historical location annotation information is obtained, which is defined as the target image.
[0058] In the target image, the adjacent historical positions and the current position are compared in sequence according to the time sequence information contained in the third tag corresponding to the historical position. For each two adjacent positions (between historical positions or between historical positions and the current position), the displacement between them is calculated. The displacement can be calculated by the coordinate difference, for example, the difference between the center coordinates of the bounding boxes of the two positions is calculated. According to the magnitude and direction of the displacement, the movement direction and movement distance of the target object between adjacent time points are determined. Furthermore, the average speed of the target object in different time periods is calculated in combination with the shooting time interval of the historical annotated images corresponding to the historical positions. For example, the displacement of two adjacent positions is divided by the corresponding time interval to obtain the average speed in the time period. By comprehensively analyzing the movement direction, movement distance and average speed corresponding to each target object, the position change of the target object is comprehensively determined.
[0059] A possible implementation of the embodiment of the present application, in the above embodiment, determining whether to generate safety warning information based on the position change of each target object and the target object in the current image includes: Based on the position change of each target object, predict the motion trajectory of each target object in the current image, the motion trajectory including the motion speed and the motion route; Determine whether there are overlapping trajectory points in the motion trajectories of any two target objects; If there are overlapping trajectory points in the motion trajectories of any two target objects, the any two target objects are respectively determined as the first object and the second object, and the first labels corresponding to the first object and the second object are obtained, and based on the first labels corresponding to the first object and the second object, a first influence degree and a second influence degree are determined, and based on the first influence degree and the second influence degree, it is determined whether to generate safety warning information; If there are no overlapping trajectory points in the motion trajectories of any two target objects, it is determined that no safety warning information is generated.
[0060] Based on the position data of the target object at the historical moment, the Kalman filter model can be used to predict its future position. For the speed of movement, the historical speed can be obtained by calculating the difference of the historical position change and dividing it by the time interval, and then the future speed of movement can be predicted by combining the speed change trend. For the movement route, the moving direction and turning conditions of the target object in the historical sequence can be analyzed to infer its possible future movement direction and thus depict the movement route. For example, if a target object has been moving at a constant speed along a straight line for a period of time in the past, it is predicted that it will most likely continue to move along the straight line at a similar speed in the future.
[0061] Furthermore, by analyzing the motion trajectory of each pair of target objects, the predicted motion trajectory is discretized into a series of trajectory points, each of which contains position coordinates and corresponding time information. Then the trajectory points of any two target objects are compared one by one to check whether there is a situation where the position coordinates are similar (within the preset error range) at the same time point. If such a trajectory point exists, it is considered that there are overlapping trajectory points in the motion trajectories of the two target objects; if not, it is considered that their motion trajectories do not overlap. For example, the motion trajectories of the two target objects in the next 10 seconds are divided into one trajectory point per second, and then the position coordinates of each time point are compared.
[0062] When it is determined that there are overlapping trajectory points in the motion trajectories of any two target objects, the two target objects are named the first object and the second object respectively, and then the first tags corresponding to the first object and the second object are read, and the first tags contain the category and name information of the object. Further, the first impact degree and the second impact degree are determined for different object categories and names. Specifically, if the first object is a large construction vehicle (strong object) and the second object is a worker (weak object), then when their motion trajectories overlap, the large construction vehicle may cause greater harm to the worker, while the manual labor may cause less harm to the large construction vehicle, so the impact degree of the first object (first impact degree) is a low impact value, and the consequences of the second object (worker) being injured are serious, and its impact degree (second impact degree) is also a high impact value. Among them, the first impact degree and the second impact degree respectively include a high impact value and a low impact value.
[0063] Further, when the first impact degree and the second impact degree are both low impact values, it is determined not to generate security warning information; when the first impact degree and / or the second impact degree are high impact values, it is determined to generate security warning information.
[0064] Furthermore, after the motion trajectories of all target objects are compared, if no overlapping trajectory points are found, it means that under the currently predicted motion trajectories, the target objects will not collide or interfere with each other. Based on this judgment, the electronic device determines that there is no safety risk at present and will not generate safety warning information.
[0065] The above-mentioned embodiment introduces a target recognition method based on a smart construction site from the perspective of method flow. The following embodiment introduces a target recognition device based on a smart construction site from the perspective of a virtual module or a virtual unit. For details, please refer to the following embodiment.
[0066] See also Figure 2The target recognition device 20 based on the smart construction site may specifically include: an acquisition module 201, an identification module 202, a sorting module 203, a first determination module 204 and a second determination module 205, wherein: A target recognition device 20 based on a smart construction site, comprising: The acquisition module 201 is used to acquire the real-time monitoring video corresponding to the current construction site, and extract the current image corresponding to the current moment from the real-time monitoring video; The recognition module 202 is used to recognize the target objects in the current image and mark each target object on the current image. The target object is an object that has the ability to move on the current construction site. The sorting module 203 obtains the historical annotated images corresponding to the historical moments, and sorts the historical annotated images according to the shooting time to obtain a historical sequence; A first determination module 204 is used to determine the position change of each target object in the historical sequence; The second determination module 205 is used to determine whether to generate safety warning information based on the position change of each target object and the target object in the current image.
[0067] In a possible implementation of the embodiment of the present application, when the recognition module 202 recognizes the target object in the current image, it is specifically used to: The current image is identified by using an edge recognition algorithm to obtain the current object in the current image; Calculate the similarity between each current object in the current image and the preset object; The similarity between each current object and the preset object is respectively compared with the corresponding similarity threshold to obtain a comparison result corresponding to each current object, and the comparison result is greater than, less than, or equal to; The current object whose comparison result is greater than is determined as the target object.
[0068] In a possible implementation of the embodiment of the present application, when the recognition module 202 marks each target object on the current image, it is specifically used to: Determine the object category and object name corresponding to each target object, and the object type includes strong objects and weak objects; Based on the object category and the object name corresponding to each target object, a first label corresponding to each target object is generated, and each target object is labeled based on the first label corresponding to each target object.
[0069] In a possible implementation of the embodiment of the present application, when determining the position change of each target object in the historical sequence, the first determination module 204 is specifically configured to: Obtaining basic information of a target object having a first label in each historically annotated image, the basic information including object color and object texture; Based on the object color and object texture corresponding to each target object in each historically annotated image, generate a second label corresponding to each target object; Calculating the similarity between any two target objects based on the first label and the second label, and determining the associated object of each target object in any historical annotated image based on the similarity between the any two target objects, wherein the any two target objects come from two different historical annotated images; Based on the associated objects of each target object in any historical annotated image, the position change of each target object in the historical sequence is determined.
[0070] In a possible implementation of the embodiment of the present application, the first determination module 204 determines the position change of each target object in the historical sequence based on the associated objects of each target object in any historical annotated image. The process of determining the position change of each target object is specifically used to: Obtain the historical position and current position corresponding to the target object, the historical position is the position of the associated object corresponding to the target object in the historical sequence, and the current position is the position of the target object in the current image; Put each historical position into the current image in sequence according to the order of each historical annotated image in the historical sequence, and generate a third label corresponding to each historical position according to the order of each historical annotated image in the historical sequence; Annotate the current image based on the third label corresponding to each historical position to obtain a target image; Based on each historical position and current position in the target image, the position change of the target object is determined.
[0071] In a possible implementation of the embodiment of the present application, when the second determination module 205 determines whether to generate safety warning information based on the position change of each target object and the target object in the current image, it is specifically used to: Based on the position change of each target object, predict the motion trajectory of each target object in the current image, the motion trajectory including the motion speed and the motion route; Determine whether there are overlapping trajectory points in the motion trajectories of any two target objects; If there are overlapping trajectory points in the motion trajectories of any two target objects, the any two target objects are respectively determined as the first object and the second object, and the first labels corresponding to the first object and the second object are obtained, and based on the first labels corresponding to the first object and the second object, a first influence degree and a second influence degree are determined, and based on the first influence degree and the second influence degree, it is determined whether to generate safety warning information; If there are no overlapping trajectory points in the motion trajectories of any two target objects, it is determined that no safety warning information is generated.
[0072] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0073] See also Figure 3 , the embodiment of the present application also introduces an electronic device from the perspective of a physical device, such as Figure 3 As shown, Figure 3 The electronic device 300 shown includes: a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, such as through a bus 302. Optionally, the electronic device 300 may also include a transceiver 304. It should be noted that in actual applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation on the embodiments of the present application.
[0074] The processor 301 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. The processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0075] The bus 302 may include a path to transmit information between the above components. The bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0076] The memory 303 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0077] The memory 303 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 301. The processor 301 is used to execute the application code stored in the memory 303 to implement the contents shown in the above method embodiment.
[0078] The electronic devices include but are not limited to: mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., and can also be servers, etc. Figure 3 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0079] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content in the aforementioned method embodiment.
[0080] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0081] The above are only some implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A target recognition method based on a smart construction site, characterized in that: include: Obtaining a real-time monitoring video corresponding to the current construction site, and extracting a current image corresponding to the current moment from the real-time monitoring video; Identify target objects in the current image and mark each target object on the current image, wherein the target object is an object capable of moving on the current construction site; Obtain the historical annotated images corresponding to each historical moment, and sort the historical annotated images according to the shooting time to obtain a historical sequence; Determining a position change of each target object in the historical sequence; Based on the position change of each target object and the target object in the current image, it is determined whether to generate safety warning information.
2. The target recognition method based on smart construction site according to claim 1 is characterized in that: The identifying the target object in the current image includes: Using an edge recognition algorithm to recognize the current image to obtain a current object in the current image; Calculating the similarity between each current object in the current image and a preset object; Comparing the similarity between each current object and the preset object with the corresponding similarity threshold respectively, to obtain a comparison result corresponding to each current object, wherein the comparison result is greater than, less than, or equal to; The current object whose comparison result is greater than is determined as the target object.
3. The target recognition method based on smart construction site according to claim 2 is characterized in that: The marking each target object on the current image includes: Determine the object category and object name corresponding to each target object, wherein the object type includes a strong object and a weak object; Based on the object category and the object name corresponding to each target object, a first label corresponding to each target object is generated, and each target object is labeled based on the first label corresponding to each target object.
4. The target recognition method based on smart construction site according to claim 3 is characterized in that: The determining of the position change of each target object in the historical sequence includes: Obtaining basic information of a target object having a first label in each historically annotated image, the basic information including object color and object texture; Based on the object color and object texture corresponding to each target object in each historically annotated image, generate a second label corresponding to each target object; Calculating the similarity between any two target objects based on the first label and the second label, and determining the associated object of each target object in any historical annotated image based on the similarity between the any two target objects, wherein the any two target objects are from two different historical annotated images; Based on the associated objects of each target object in any historical annotated image, a position change of each target object in the historical sequence is determined.
5. The target recognition method based on smart construction site according to claim 4 is characterized in that: The process of determining the position change of each target object in the historical sequence based on the associated objects of each target object in any historical annotated image includes: Acquire a historical position and a current position corresponding to the target object, wherein the historical position is a position of an associated object corresponding to the target object in the historical sequence, and the current position is a position of the target object in the current image; Put each historical position into the current image in sequence according to the order of each historical annotated image in the historical sequence, and generate a third label corresponding to each historical position according to the order of each historical annotated image in the historical sequence; Annotate the current image based on the third label corresponding to each historical position to obtain a target image; Based on each historical position and the current position in the target image, a position change of the target object is determined.
6. The target recognition method based on a smart construction site according to any one of claims 3 to 5, characterized in that: The determining whether to generate safety warning information based on the position change of each target object and the target object in the current image includes: Based on the position change of each target object, predict the motion trajectory of each target object in the current image, wherein the motion trajectory includes the motion speed and the motion route; Determine whether there are overlapping trajectory points in the motion trajectories of any two target objects; If there are overlapping trajectory points in the respective motion trajectories of the arbitrary two target objects, the arbitrary two target objects are respectively determined as a first object and a second object, first labels corresponding to the first object and the second object are obtained, and a first influence degree and a second influence degree are determined based on the first labels corresponding to the first object and the second object, and whether to generate safety warning information is determined based on the first influence degree and the second influence degree; If there are no overlapping trajectory points in the motion trajectories of any two target objects, it is determined that no safety warning information is generated.
7. A target recognition device based on a smart construction site, characterized in that: include: An acquisition module is used to acquire a real-time monitoring video corresponding to the current construction site, and extract a current image corresponding to the current moment from the real-time monitoring video; A recognition module, used to recognize target objects in the current image and mark each target object on the current image, wherein the target object is an object capable of moving on the current construction site; A sorting module obtains the historical annotated images corresponding to each historical moment, and sorts the historical annotated images according to the shooting time to obtain a historical sequence; A first determination module is used to determine the position change of each target object in the historical sequence; The second determination module is used to determine whether to generate safety warning information based on the position change of each target object and the target object in the current image.
8. The target recognition device based on the smart construction site according to claim 7 is characterized in that: When identifying the target object in the current image, the recognition module is specifically used to: Using an edge recognition algorithm to recognize the current image to obtain a current object in the current image; Calculating the similarity between each current object in the current image and a preset object; Comparing the similarity between each current object and the preset object with the corresponding similarity threshold respectively, to obtain a comparison result corresponding to each current object, wherein the comparison result is greater than, less than, or equal to; The current object whose comparison result is greater than is determined as the target object.
9. An electronic device, characterized in that: The electronic device includes: at least one processor; Memory; At least one application, wherein at least one application is stored in a memory and configured to be executed by at least one processor, and the at least one application is configured to: execute the target recognition method based on a smart construction site as described in any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed in a computer, the computer is caused to execute the target recognition method based on a smart construction site as described in any one of claims 1 to 6.
Citation Information
Patent Citations
A Method and System for Worker and Machinery Collision Avoidance Prediction Based on Computer Vision
CN114937240A
Safety monitoring system and method based on mine intelligent construction site
CN118172732A
Collision early warning method and apparatus, and head-mounted VR device, and storage medium
WO2023092641A1
Cited By
ESG-oriented coal mine equipment safe operation optimization method and system
CN121214353A