Intelligent inspection method and system for special operation physical examination room
Through intelligent inspection methods that integrate real-life navigation maps and image features, the inspection paths are dynamically planned, and dangerous areas are identified and visualized, which solves the problem that traditional manual inspections are difficult to fully cover in high-risk operating environments, and efficient and accurate examination room inspections are achieved, reducing safety hazards.
Patent Information
- Application Number
- CN202510425495.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Traditional manual inspections are difficult to fully cover the examination room area in high-risk operating environments, and there are blind spots in inspections, which reduces the accuracy of inspections and brings safety hazards.
An intelligent inspection method that integrates real-life navigation maps and image features is adopted. By obtaining examination room data and images in real time, patrol the inspection path dynamically, identifying dangerous areas and visualizing them, using drones or robots to conduct high-risk areas, and combining image features and environmental parameters for fusion analysis to determine the congestion index.
It improves the accuracy and coverage of inspections, reduces the subjectivity and error of human judgments, effectively reduces the risks and hidden dangers in the examination room, ensures the comprehensiveness and pertinence of inspections, and avoids the blind spots of manual inspections.
Smart Images

Figure CN120279488A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of path planning, and in particular, to an intelligent inspection method and system for a special operation physical examination room. Background Art
[0002] The inspection work in special operation examination rooms (such as high-risk industries like electricity, chemical industry, and construction) is of great significance for ensuring operation safety and preventing accidents. The traditional inspection method mainly relies on inspectors to regularly check the equipment, environment, etc. in the examination room to ensure that they meet safety standards.
[0003] Specifically, the inspectors need to check the equipment and environment in the examination room one by one along a fixed route. The accuracy of manual inspection highly depends on the experience and skills of the inspectors. However, with the expansion of the scale of the examination room, the increase in equipment complexity, and the improvement of safety requirements, especially in high-risk operation environments such as high-altitude operations and welding operations, in order to avoid potential safety hazards to personnel, manual inspection often cannot fully cover the examination room area, resulting in inspection blind spots. If there is a problem with a certain piece of equipment in the examination room, it will not only reduce the accuracy of manual inspection but also bring risk hazards to the personnel in the examination room.
[0004] Currently, there is no effective solution to the above problems. Summary of the Invention
[0005] The embodiments of the present application provide an intelligent inspection method and system for a special operation physical examination room, which are used to effectively improve the accuracy of inspecting the special operation physical examination room, thereby effectively reducing the risk hazards in the examination room.
[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, there is provided an intelligent inspection method for a special operation physical examination room, which is applied to a central control device. The method includes:
[0008] Obtain the examination room data and environmental parameters in real time, and establish a real-scene navigation map based on the examination room data, where the real-scene navigation map is used to be linked with the examination room data in real time;
[0009] Obtain the examination room images in real time through an inspection device, where the inspection device is wirelessly or wiredly connected to the central control device, and the inspection device is provided with a camera, and the camera is used to collect the examination room images;
[0010] Extract the image features of the examination room images, and divide the inspection path into n regions based on the image features, where n is a positive integer;
[0011] Based on the real-scene navigation map and the image features, execute a dynamic path planning strategy to generate the inspection path of the inspection device;
[0012] Fuse and analyze environmental parameters and image features to determine the congestion index on the inspection path;
[0013] Based on the congestion index, determine dangerous areas among n areas and visualize the dangerous areas.
[0014] In a possible implementation manner of the first aspect, the examination room data includes three-dimensional space data, and a real-scene navigation map is established based on the examination room data, including:
[0015] Convert the three-dimensional space data into a grid map;
[0016] Use a pre-built neural network model to perform semantic segmentation on the three-dimensional space data to obtain a semantic segmentation result;
[0017] Fuse the semantic segmentation result with the grid map to obtain a real-scene navigation map including semantic information.
[0018] In another possible implementation manner of the first aspect, the dynamic path planning strategy includes:
[0019] Obtain historical inspection data;
[0020] Calculate the global optimal path based on the improved A* algorithm, where the improved A* algorithm is improved based on a heuristic function;
[0021] Real-time monitor the pedestrian flow density and obstacle distribution on the inspection path;
[0022] Determine whether there are path blocking areas and / or dangerous areas in the global optimal path based on the pedestrian flow density and obstacle distribution;
[0023] In the case where there are path blockages and / or dangerous areas in the global optimal path, execute a local path replanning strategy and generate a final path;
[0024] Among them, the local path replanning strategy includes:
[0025] Identify the boundary coordinates of the path blocking area and / or dangerous area;
[0026] In the real-scene navigation map, assign a high passage cost value to the grid nodes of the path blocking area and / or dangerous area;
[0027] Taking the current position as the starting point and the target inspection point as the ending point, apply the D*Lite algorithm to calculate the local optimal path, where the target inspection point is the grid node closest to the grid nodes of the path blocking area and / or dangerous area outside the boundary coordinates;
[0028] When multiple consecutive path blocking areas and / or dangerous areas are detected, a segmented planning strategy is adopted to decompose the path into multiple sub-paths;
[0029] Apply the ant colony optimization algorithm to each sub-path to generate a locally optimal sub-path;
[0030] Connect all the locally optimal sub-paths to obtain a locally re-planned path.
[0031] In another possible implementation of the first aspect, based on the real-scene navigation map and image features, a dynamic path planning strategy is executed to generate the inspection path of the inspection device, including:
[0032] Divide the real-scene navigation map into multiple grid nodes and calculate the passing cost value of each grid node;
[0033] In response to the user's setting instruction for the real-scene navigation map, determine the inspection points in the real-scene navigation map and the corresponding weight coefficients for each inspection point;
[0034] Use the improved A* algorithm to calculate the initial global path;
[0035] Take the real-time acquired environmental parameters and image features as the input parameters of the pre-constructed multi-feature fusion model to dynamically update the passing cost value of each grid node;
[0036] When the change amount of the passing cost value exceeds the preset threshold, adopt the preset path re-planning algorithm to generate a new inspection path;
[0037] When the change amount of the passing cost value does not exceed the preset threshold, take the initial global path as the inspection path of the inspection device.
[0038] In another possible implementation of the first aspect, divide the inspection path into n regions based on image features, including:
[0039] Based on the K-means clustering algorithm, cluster the image features to obtain at least one clustering region;
[0040] Calculate the feature vectors of each clustering region and calculate the similarity between the feature vectors;
[0041] According to the similarity between the feature vectors, divide the inspection path into n consecutive regions, where each region is assigned a unique identifier.
[0042] In another possible implementation of the first aspect, fuse and analyze the environmental parameters and image features to determine the congestion index on the inspection path, including:
[0043] The personnel density, equipment distribution, and operation behaviors are extracted from the image features through the YOLOv8 model;
[0044] The congestion index is output through a multi-layer perceptron neural network model. Among them, the input layer of the multi-layer perceptron neural network model includes environmental parameters, personnel density, equipment distribution, and operation behaviors, and the output layer is the congestion index.
[0045] In another possible implementation manner of the first aspect, based on the congestion index, dangerous areas are determined among n areas, including:
[0046] When the congestion index of an area is lower than the preset first congestion index threshold T1, the area is marked as a safe area;
[0047] When the congestion index of an area is between the first congestion index threshold T1 and the second congestion index threshold T2, the area is marked as a warning area;
[0048] When the congestion index of an area is higher than the second congestion index threshold T2, the area is marked as a dangerous area, where 0 < T1 < T2 < 100.
[0049] In another possible implementation manner of the first aspect, the camera is also used to collect the examination room video. The examination room video includes multiple video frames. The method further includes:
[0050] Input the video frames into a pre-built YOLOv8 model to identify predefined types of violation behaviors;
[0051] For each type of violation behavior, record the start time and end time of the violation behavior;
[0052] Within a preset time interval, extract key video frames at a frequency of m frames per second, and add a timestamp, a violation type label, and a violation area bounding box to each key video frame;
[0053] Sort the key video frames in chronological order, generate a violation behavior evidence package, store the violation behavior evidence package in the security database of the central control device, and construct a violation behavior index for the security database.
[0054] In another possible implementation manner of the first aspect, the patrol device transmits the examination room image to the central control device, including:
[0055] The patrol device performs H.264 encoding and compression on the examination room image, establishes an RTMP streaming channel, and pushes the encoded data stream to the central control device;
[0056] Perform AES-256 encryption on the encoded data stream to obtain an encrypted data stream;
[0057] When the patrol device detects a network interruption, it stores the encrypted data stream in the local cache, and after detecting the network recovery, it transmits the encrypted data stream to the central control device.
[0058] In a second aspect, the present application provides a patrol system, including:
[0059] A central control device for executing the above-mentioned intelligent patrol method for the physical examination room of special operations;
[0060] A patrol device wirelessly or wiredly connected to the central control device.
[0061] Through the above technical solutions, the central control device can obtain the examination room data and environmental parameters in real time, and establish a real-scene navigation map based on the examination room data. The real-scene navigation map is linked with the examination room data in real time, providing accurate spatial positioning and navigation support for patrol. Secondly, the patrol device can obtain the examination room images in real time. The patrol device is wirelessly or wiredly connected to the central control device, and the patrol device is equipped with a camera for collecting the examination room images, enabling the patrol device to quickly cover all areas of the examination room, avoiding the time-consuming problem of manual patrol. At the same time, by extracting the image features of the examination room images and dividing the patrol path into n regions based on the image features, where n is a positive integer, the comprehensive coverage of the examination room area by the patrol is realized, further improving the patrol efficiency. In addition, based on the real-scene navigation map and image features, a dynamic path planning strategy is executed to generate the patrol path of the patrol device. This solution can adjust the patrol route in real time according to the actual situation of the examination room, ensuring the comprehensiveness and pertinence of the patrol. At the same time, by fusing and analyzing the environmental parameters and image features to determine the congestion index on the patrol path, the state of the examination room can be judged more accurately, reducing the subjectivity and error of human judgment. Finally, based on the congestion index, dangerous areas are determined among the n regions and visualized, enabling the management personnel to quickly understand the state of the examination room and formulate targeted improvement measures, thus effectively preventing the occurrence of accidents. For high-risk areas, this solution uses patrol devices (such as drones, robots, etc.) to be able to enter high-risk areas for patrol, avoiding the blind area problem of manual patrol, significantly expanding the coverage range of the examination room patrol. In addition, this solution uses image feature extraction technology to accurately identify changes in the examination room equipment and environment. By fusing and analyzing the image features and environmental parameters, the state of the examination room can be judged more accurately, reducing the subjectivity and error of human judgment, not only improving the patrol efficiency and accuracy, but also expanding the patrol coverage range, effectively reducing the risk hidden dangers of the examination room, thus effectively preventing the occurrence of accidents and ensuring the safety of the examination room personnel.
[0062] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1Schematic flowchart of an intelligent patrol inspection method for a special operation physical examination room provided by an embodiment of the present application;
[0064] Figure 2 Schematic structural diagram of a patrol inspection system provided by an embodiment of the present application;
[0065] Figure 3 Schematic structural diagram of a patrol inspection device provided by an embodiment of the present application;
[0066] Figure 4 Schematic diagram of the setting of patrol inspection points on a real - scene navigation map provided by an embodiment of the present application. Detailed implementation manners
[0067] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the specific implementation manners described herein are only for explaining and interpreting the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0068] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present application, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0069] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0070] The technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0071] Figure 1 Schematically shown is a flowchart of an intelligent inspection method for a special operation physical examination room according to an embodiment of the present application. As Figure 1 shown, an embodiment of the present application provides an intelligent inspection method for a special operation physical examination room, which is applied to a central control device. The method may include the following steps.
[0072] S110. Obtain the examination room data and environmental parameters in real time, and establish a real-scene navigation map based on the examination room data, wherein the real-scene navigation map is used to be linked with the examination room data in real time;
[0073] S120. Obtain the examination room images in real time through an inspection device, wherein the inspection device is wirelessly or wiredly connected to the central control device, and the inspection device is provided with a camera for collecting the examination room images;
[0074] S130. Extract the image features of the examination room images, and divide the inspection path into n regions based on the image features, where n is a positive integer;
[0075] S140. Based on the real-scene navigation map and the image features, execute a dynamic path planning strategy to generate the inspection path of the inspection device;
[0076] S150. Perform fusion analysis on the environmental parameters and the image features to determine the congestion index on the inspection path;
[0077] S160. Based on the congestion index, determine the dangerous regions among the n regions and visualize the dangerous regions.
[0078] Figure 2 Shown is a schematic structural diagram of an inspection system provided by an embodiment of the present application. Among them, the central control device and the inspection device communicate through wireless or wired connection. Hereinafter, wireless connection is taken as an example. As Figure 2 shown, the inspection system includes a central control device and an inspection device. The central control device and the inspection device communicate through wireless connection. The inspection device is equipped with a camera for collecting the examination room images. The central control device obtains the examination room data and environmental parameters and establishes a real-scene navigation map. The examination room images are transmitted to the central control device, and the image features are extracted and the inspection path regions are divided. Based on the real-scene navigation map and the image features, dynamic path planning is performed to generate the inspection path. The environmental parameters and the image features are subjected to fusion analysis to determine the congestion index. Based on the congestion index, the dangerous regions are determined and visualized.
[0079] In this embodiment, the central control device may be a device with a processor such as a tablet computer, desktop computer, laptop computer, handheld computer, wearable device, notebook computer, ultra-mobile personal computer (UMPC), netbook, etc. Of course, the central control device may also be a server. The specific form of the central control device is not particularly limited in the embodiments of the present application.
[0080] The central control device obtains the examination room data and environmental parameters in real time through a variety of sensors and data acquisition devices. The examination room data may include, but is not limited to: three-dimensional space data of the examination room, examination room layout information, equipment distribution information, personnel distribution information, etc. The environmental parameters may include, but are not limited to: temperature, humidity, noise, light intensity, gas concentration, etc.
[0081] In a possible implementation manner, the examination room data can be obtained in the following ways: 1. Obtain the three-dimensional point cloud data of the examination room by scanning with a lidar (LiDAR); 2. Obtain the real-time video stream of the examination room through a fixed camera; 3. Obtain the basic layout information of the examination room through a pre-established CAD model of the examination room; 4. Obtain the location information of personnel and equipment in the examination room through RFID tags or Bluetooth beacons.
[0082] The environmental parameters can be obtained through an environmental sensor network distributed throughout the examination room. These sensors are connected to the central control device by wired or wireless means and transmit data in real time.
[0083] Based on the obtained examination room data, the central control device establishes a real-scene navigation map. The real-scene navigation map is a high-precision map that integrates actual scene information and can reflect the real-time state of the examination room. The real-scene navigation map is linked to the examination room data in real time. When the state of the examination room changes, the real-scene navigation map will also be updated accordingly.
[0084] In a possible implementation manner, the process of establishing the real-scene navigation map includes: 1. Convert the three-dimensional point cloud data into a grid map, and each grid represents an area in the examination room; 2. Integrate the scene information extracted from the video stream with the grid map; 3. Overlay the static layout information in the CAD model on the map; 4. Dynamically update the location information of personnel and equipment on the map.
[0085] The real-scene navigation map can adopt a multi-layer structure, and different layers represent different types of information. For example: the basic layer represents the physical layout of the examination room, the dynamic layer represents the real-time locations of personnel and equipment, and the semantic layer represents the functional attributes of different areas, etc.
[0086] The inspection device obtains real-time images of the examination room. The inspection device can be a mobile robot, a drone, or other movable intelligent devices, equipped with a high-definition camera for collecting real-time images of the examination room. The inspection device is connected to the central control device through a wireless network, which can be Wi-Fi, 5G, or a dedicated wireless communication network.
[0087] Figure 3 The following shows a schematic structural diagram of an inspection device provided by an embodiment of the present application, as Figure 3 shown. In a possible implementation, the inspection device may include: 1. Mobile chassis: providing mobility, which can be wheeled, tracked, or other forms; 2. Imaging system: including a high-definition camera, a pan-tilt head, etc., for collecting images of the examination room; 3. Communication module: for wireless communication with the central control device; 4. Navigation system: for realizing autonomous navigation and positioning; 5. Power system: providing the energy requirements of the inspection device.
[0088] The camera can be fixed or pan-tilt type. The pan-tilt type camera can achieve a 360-degree horizontal rotation and a -90-degree to +90-degree vertical pitch to obtain a wider field of view. The resolution of the camera is not lower than 1080p, and the frame rate is not lower than 30fps to ensure image quality.
[0089] The images collected by the inspection device are transmitted to the central control device in real time through a wireless network. During the transmission process, video coding technologies such as H.264 or H.265 can be used for compression to reduce bandwidth occupancy.
[0090] The central control device processes the examination room images received from the inspection device, extracts image features, and divides the inspection path into n regions based on these features, where n is a positive integer.
[0091] The extraction of image features can adopt a variety of computer vision technologies, including but not limited to: 1. Edge detection: using operators such as Canny and Sobel to detect edges in the image; 2. Corner detection: using algorithms such as Harris and FAST to detect corners in the image; 3. Feature descriptors: using algorithms such as SIFT, SURF, and ORB to extract local features of the image; 4. Deep learning features: using a pre-trained convolutional neural network (CNN) to extract high-level semantic features of the image.
[0092] In a possible implementation, the process of image feature extraction includes: 1. Preprocessing the image, including denoising, illumination equalization, etc.; 2. Using a pre-trained deep neural network such as ResNet-50 to extract deep features of the image; 3. Using principal component analysis (PCA) to reduce the dimension of the features and retain the main information; 4. Associating the extracted features with geographical location information to form a feature map with spatial attributes.
[0093] Based on the extracted image features, the central control device divides the inspection path into n regions. The region division can be based on the following factors: 1. Similarity of image features: Regions with similar features can be grouped into one category; 2. Spatial continuity: Spatially continuous regions tend to be grouped into one category; 3. Functional attributes: Regions with the same function can be grouped into one category; 4. Risk level: Regions with similar risk levels can be grouped into one category.
[0094] In a possible implementation, the region division can use the K-means clustering algorithm to cluster regions with similar features into one category. The steps of clustering include: 1. Initialize K cluster centers; 2. Assign each region to the nearest cluster center; 3. Recalculate the center of each cluster; 4. Repeat steps 2 and 3 until the cluster centers no longer change or reach the maximum number of iterations.
[0095] Each divided region is assigned a unique identifier for subsequent path planning and risk assessment.
[0096] In one of the implementation manners, the central control device executes a dynamic path planning strategy based on the real-scene navigation map and the extracted image features to generate an optimal inspection path for the inspection device.
[0097] The dynamic path planning strategy considers multiple factors, including but not limited to: 1. Examination hall layout: Avoid impassable areas such as walls and large equipment; 2. Inspection coverage rate: Ensure that all important areas are inspected; 3. Path length: Try to shorten the path length on the premise of meeting the inspection requirements; 4. Risk factors: Avoid high-risk areas or give priority to inspecting high-risk areas; 5. Real-time obstacles: Adjust the path according to the real-time detected obstacles.
[0098] In a possible implementation, the process of dynamic path planning includes: 1. Convert the real-scene navigation map into a cost map, and each grid cell has a cost value representing the difficulty of passing through this area; 2. Dynamically update the cost map according to the image features and environmental parameters; 3. Use the improved A* algorithm to calculate the global optimal path; 4. During the inspection process, dynamically adjust the path according to the real-time obtained information.
[0099] The improved A* algorithm is an optimized version based on the traditional A* algorithm, and its heuristic function considers multiple factors such as distance, risk level, inspection priority, etc. The core formula of the algorithm is:
[0100] f(n) = g(n) + h(n) + r(n) + p(n);
[0101] Where: f(n) is the total evaluation value of node n; g(n) is the actual cost from the starting point to node n; h(n) is the estimated cost from node n to the target; r(n) is the risk factor of node n; p(n) is the inspection priority factor of node n.
[0102] When the examination room environment changes, such as the emergence of new obstacles or the gathering of people, the system will recalculate the paths in the affected areas to ensure that the inspection equipment can complete the inspection tasks safely and efficiently.
[0103] In a possible implementation, the central control device fuses and analyzes the environmental parameters and image features to calculate the congestion index of each area on the inspection path. The congestion index is a comprehensive indicator that measures the degree of regional congestion and potential risks.
[0104] There are various methods for the fusion analysis of environmental parameters and image features, including but not limited to: 1. Feature-level fusion: Fusing environmental parameters and image features at the feature level; 2. Decision-level fusion: Making decisions based on environmental parameters and image features respectively, and then fusing the decision results; 3. Model-level fusion: Constructing a fusion model that can simultaneously process environmental parameters and image features.
[0105] In a possible implementation, the process of fusion analysis includes: 1. Extracting information such as personnel density, equipment distribution, and operation behavior from image features; 2. Standardizing the environmental parameters to make them comparable to image features; 3. Constructing a multi-layer perceptron neural network model, where the input layer includes environmental parameters and image features, and the output layer is the congestion index; 4. Training the neural network model with historical data to enable it to accurately predict the congestion index.
[0106] The calculation formula of the congestion index can be expressed as:
[0107] CI = f(ED, PD, ED, OB, EP).
[0108] Where: CI is the congestion index; ED is the equipment distribution density; PD is the personnel distribution density; OB is the operation behavior risk score; EP is the comprehensive score of environmental parameters; f is the function represented by the neural network model.
[0109] The value range of the congestion index is from 0 to 100. The larger the value, the higher the degree of congestion and the greater the risk.
[0110] Based on the calculated congestion index, the central control device determines the dangerous areas among the n areas and visually displays the dangerous areas on the real-scene navigation map.
[0111] The determination of the dangerous area can be based on the congestion index threshold. For example, when the congestion index of an area is lower than the threshold T1, the area is marked as a safe area; when the congestion index of the area is between the thresholds T1 and T2, the area is marked as a warning area; when the congestion index of the area is higher than the threshold T2, the area is marked as a dangerous area.
[0112] Among them, 0 < T1 < T2 < 100, and the specific threshold can be adjusted according to the actual application scenario.
[0113] In a possible implementation, T1 can be set to 30 and T2 can be set to 70, that is, the area with a congestion index lower than 30 is a safe area, the area with a congestion index between 30 and 70 is a warning area, and the area with a congestion index higher than 70 is a dangerous area.
[0114] The visualization of the dangerous area can be achieved in the following ways: 1. Use different colors on the real - scene navigation map to mark areas with different risk levels. For example, green represents a safe area, yellow represents a warning area, and red represents a dangerous area; 2. Use a heat map to display the distribution of the congestion index, and the color from blue to red represents the congestion index from low to high; 3. Add a flashing or other visual effect to the dangerous area to attract the attention of the operator; 4. Overlay risk - type icons on the dangerous area, such as crowd gathering, equipment abnormality, etc.
[0115] The visualization result can be displayed in real - time on the display interface of the central control device, or can be remotely viewed through a mobile terminal or a Web interface. Based on the visualization result, the operator can timely discover potential risks and take corresponding intervention measures.
[0116] Through the above steps, the intelligent inspection method provided by the embodiments of this application can achieve comprehensive monitoring and risk assessment of the special - operation physical examination room, improve the efficiency and accuracy of the examination room safety management, and reduce the occurrence probability of safety accidents.
[0117] In practical applications, this method can be flexibly adjusted and optimized according to the specific examination room environment and requirements to adapt to different application scenarios.
[0118] In this embodiment, the central control device is used to obtain the examination room data and environmental parameters in real time, and a real-scene navigation map is established based on the examination room data. The real-scene navigation map is linked with the examination room data in real time, providing accurate spatial positioning and navigation support for the patrol. Secondly, the patrol device is used to obtain the examination room images in real time. The patrol device is connected to the central control device wirelessly or wiredly. The patrol device is equipped with a camera, which is used to collect the examination room images, enabling the patrol device to quickly cover all areas of the examination room, avoiding the time-consuming problem of manual patrol. At the same time, by extracting the image features of the examination room images and dividing the patrol path into n areas based on the image features, where n is a positive integer, the comprehensive coverage of the examination room area by the patrol is achieved, further improving the patrol efficiency. In addition, based on the real-scene navigation map and the image features, a dynamic path planning strategy is executed to generate the patrol path of the patrol device. This solution can adjust the patrol route in real time according to the actual situation of the examination room, ensuring the comprehensiveness and pertinence of the patrol. At the same time, the environmental parameters and the image features are fused and analyzed to determine the congestion index on the patrol path, enabling a more accurate judgment of the examination room status, reducing the subjectivity and error of human judgment. Finally, based on the congestion index, dangerous areas are determined among the n areas and visualized, enabling the management personnel to quickly understand the examination room status and formulate targeted improvement measures, thus effectively preventing the occurrence of accidents. For high-risk areas, this solution uses patrol devices (such as drones, robots, etc.) to be able to enter the high-risk areas for patrol, avoiding the blind area problem of manual patrol, significantly expanding the coverage range of the examination room patrol. In addition, this solution uses the image feature extraction technology to accurately identify the changes in the examination room equipment and environment. By fusing and analyzing the image features and the environmental parameters, the examination room status can be judged more accurately, reducing the subjectivity and error of human judgment. It not only improves the patrol efficiency and accuracy but also expands the patrol coverage range, effectively reducing the risk potential of the examination room, thus effectively preventing the occurrence of accidents and ensuring the safety of the examination room personnel.
[0119] In one implementation manner of this embodiment, the examination room data includes three-dimensional space data. Establishing a real-scene navigation map based on the examination room data includes the following steps:
[0120] S210. Convert the three-dimensional space data into a grid map;
[0121] S220. Perform semantic segmentation on the three-dimensional space data using a pre-constructed neural network model to obtain a semantic segmentation result;
[0122] S230. Fuse the semantic segmentation result with the grid map to obtain a real-scene navigation map including semantic information.
[0123] The examination room data includes three-dimensional spatial data, which are usually acquired by devices such as lidar (LiDAR), depth cameras, or 3D scanners. The three-dimensional spatial data contains information such as the geometric shape, size, and spatial position of the examination room environment, and exists in the form of point clouds, meshes, or voxels. These data can accurately reflect the physical layout of the examination room, including the positions and shapes of fixed facilities such as walls, doors, windows, equipment, desks, and chairs, providing basic data support for the subsequent construction of navigation maps. When collecting three-dimensional spatial data, factors such as sampling density, accuracy, and coverage need to be considered to ensure the integrity and accuracy of the data. In the physical examination room for special operations, the three-dimensional spatial data is usually collected by multi-site scanning, and the data from multiple sites are registered and fused to form a complete three-dimensional model of the examination room. These data are usually stored in point cloud file formats (such as.pcd,.ply, etc.) or 3D model formats (such as.obj,.stl, etc.), and are preprocessed through data processing software, including operations such as denoising, downsampling, and outlier filtering, to improve data quality and processing efficiency. High-quality three-dimensional spatial data is the key prerequisite for constructing an accurate real-scene navigation map, directly affecting the subsequent navigation and inspection effects.
[0124] The process of converting three-dimensional spatial data into a grid map is to discretize continuous three-dimensional spatial information into a regular grid structure, which is convenient for computer processing and the application of path planning algorithms. When specifically implementing, first, the resolution of the grid needs to be determined, that is, the physical size of each grid cell, which is usually determined according to the size of the inspection equipment, the navigation accuracy requirements, and the computing resource limitations. For example, the grid size can be set to 10 cm × 10 cm, which can ensure sufficient accuracy without causing excessive computational complexity. During the conversion process, the three-dimensional spatial data is projected onto a horizontal plane to form a two-dimensional occupancy grid map. For each grid cell, the occupancy probability value is calculated according to the distribution of the point cloud in its corresponding space. The specific calculation method can use the Bayesian update formula. When the point cloud density exceeds the threshold within a certain height range (usually between 10 cm and 200 cm above the ground), the grid is marked as occupied; when the point cloud density is lower than the threshold, the grid is marked as free; when there is not enough observation data, the grid is marked as unknown. In addition, according to the height information of the point cloud, a height value can be assigned to each grid cell to form a 2.5D elevation map, which better represents the height changes of the environment. The generation process of the grid map also needs to consider the filtering of dynamic obstacles. By comparing multiple scans, temporary obstacles are identified and removed, and the static environmental structure is retained. The converted grid map is stored in matrix form, and each element represents the state (occupied, free, or unknown) of the corresponding grid, providing a basis for subsequent path planning and navigation.
[0125] Using a pre - constructed neural network model to perform semantic segmentation on 3D spatial data, the obtained semantic segmentation results are for understanding the functional attributes of different objects and regions in the environment.
[0126] Semantic segmentation is a task in computer vision that aims to assign each pixel or point in an image or point cloud to a predefined semantic category.
[0127] In the special operation physical examination site scenario, common semantic categories include ground, wall, door, window, desks and chairs, equipment, personnel passage, dangerous area, etc. The pre - constructed neural network model usually adopts deep learning architectures such as PointNet++, SparseConvNet or MinkowskiNet, etc. These models are specifically designed to process 3D point cloud data. The training process of the model requires a large amount of 3D data with semantic annotations. The transfer learning method can be used. First, pre - train on large - scale public datasets (such as S3DIS, ScanNet), and then fine - tune on the specific examination site data.
[0128] In practical applications, first pre - process the 3D point cloud data into an input format acceptable to the model, including steps such as point cloud downsampling, normal vector calculation, and feature extraction. Then, input the processed point cloud into the neural network model. The model learns the local and global features of the point cloud through multi - layer feature extraction and non - linear transformation, and finally predicts semantic labels for each point. The output of the model is a vector of the same size as the input point cloud, and each element represents the probability distribution of the corresponding point belonging to each semantic category. By taking the category corresponding to the maximum probability as the final prediction result, the semantic segmentation task is completed.
[0129] To improve the segmentation accuracy, multi - view information can be combined to fuse and process RGB images and point cloud data, taking advantage of their complementary advantages. In addition, post - processing methods such as conditional random field (CRF) can be applied to optimize the spatial consistency of the segmentation results. The results of semantic segmentation are stored in the form of category labels for each point, providing semantic information for subsequent navigation map construction. Through semantic segmentation, the functional structure of the environment can be better understood, different types of regions and objects can be distinguished, providing a higher - level environmental understanding ability for intelligent inspection.
[0130] Fusing the semantic segmentation results with the grid map to obtain a real - scene navigation map including semantic information is a key step in constructing an advanced navigation map. The fusion process first requires establishing a spatial correspondence between the semantic segmentation results and the grid map, which is usually achieved through coordinate transformation. For each grid cell, count the number of points in each semantic category that fall into this grid, and use the Majority Voting method to determine the main semantic category of this grid. For example, if 60% of the points in a grid are classified as "equipment", 30% as "ground", and 10% as "wall", then the semantic label of this grid is determined to be "equipment". To handle the uncertainty in semantic segmentation, a semantic probability distribution can be maintained for each grid instead of just a single label, so that the semantic uncertainty can be considered in subsequent path planning. During the fusion process, the functional attributes of different semantic categories also need to be considered. For example, assign a traversability cost to each semantic category, which represents the difficulty or risk for inspection equipment to pass through this type of area. Usually, the traversability costs of the "ground" and "passage" categories are relatively low, while those of the "equipment" and "wall" categories are relatively high or impassable. In addition, different inspection priorities can be set according to semantic categories. For example, "hazardous areas" and "critical equipment" have higher inspection priorities. The fused real - scene navigation map is stored in a multi - layer structure, including an occupancy layer, a semantic layer, a cost layer, a priority layer, etc., with each layer representing different types of information. For the convenience of visualization and human - machine interaction, different colors can be assigned to different semantic categories and displayed intuitively in a 2D or 3D view. The real - scene navigation map also needs to support real - time updates, and be able to dynamically adjust the map content when the environment changes or new observation data is obtained. The real - scene navigation map integrating semantic information not only contains the geometric structure of the environment, but also the functional attributes of the environment, and can support more intelligent path planning and decision - making, improving the efficiency and safety of inspection.
[0131] In this embodiment, by converting three-dimensional spatial data into a grid map, using a pre-built neural network model for semantic segmentation, and fusing the semantic segmentation results with the grid map, a real-scene navigation map containing rich semantic information is finally constructed. This advanced navigation map not only represents the geometric structure of the examination room environment but also includes the functional attributes of different regions and objects in the environment, providing a solid foundation for intelligent patrol inspection. Based on this real-scene navigation map, the patrol inspection system can understand the semantic structure of the environment, distinguish different types of regions, such as dangerous areas, equipment areas, personnel passages, etc., so as to achieve more intelligent path planning and decision-making. For example, the system can prioritize the patrol inspection of dangerous areas, avoid impassable areas, reasonably arrange the patrol inspection path, and improve the patrol inspection efficiency. At the same time, the semantic information can also help the system better understand environmental changes and identify abnormal situations. For example, detecting an obstacle in an area that should originally be empty may indicate a potential safety hazard. In addition, the multi-layer structure design of the real-scene navigation map enables it to support multiple application scenarios and meet different levels of navigation and decision-making needs. Generally speaking, this real-scene navigation map integrating semantic information significantly improves the environmental understanding ability, path planning ability, and safety monitoring ability of the intelligent patrol inspection system for special operation physical examination rooms, providing strong support for ensuring the safety of the examination room.
[0132] In one implementation of this embodiment, the dynamic path planning strategy includes the following steps:
[0133] S310. Obtain historical patrol inspection data;
[0134] S320. Calculate the global optimal path based on the improved A* algorithm, where the improved A* algorithm is improved based on the heuristic function;
[0135] S330. Real-time monitor the crowd density and obstacle distribution on the patrol inspection path;
[0136] S340. Determine whether there are path blocking areas and / or dangerous areas in the global optimal path based on the crowd density and obstacle distribution;
[0137] S350. In the case where there are path blocking and / or dangerous areas in the global optimal path, execute the local path replanning strategy and generate the final path;
[0138] Among them, the local path replanning strategy includes:
[0139] S1. Identify the boundary coordinates of the path blocking area and / or dangerous area;
[0140] S2. In the real-scene navigation map, assign a high passage cost value to the grid nodes of the path blocking area and / or dangerous area;
[0141] S3. Starting from the current position and ending at the target inspection point, apply the D*Lite algorithm to calculate the locally optimal path, where the target inspection point is the grid node closest to the grid nodes in the distance path blockage area and / or dangerous area outside the boundary coordinates;
[0142] S4. When multiple consecutive path blockage areas and / or dangerous areas are detected, adopt a segmented planning strategy to decompose the path into multiple sub-paths;
[0143] S5. Apply the ant colony optimization algorithm to each sub-path to generate locally optimal sub-paths;
[0144] S6. Connect all the locally optimal sub-paths to obtain the locally re-planned path.
[0145] Historical inspection data usually includes multi-dimensional information such as inspection path records, inspection point access order, travel time of each section, variation law of pedestrian flow density, frequency of obstacle appearance and location distribution, etc. These data are stored in the form of time series, and each record contains fields such as timestamp, location coordinates, inspection status, and environmental parameters. The process of obtaining historical inspection data first requires extracting the original data from the system database, which is usually stored in a structured form in a relational database (such as MySQL) or a time series database (such as InfluxDB). Data extraction can be achieved through SQL query statements.
[0146] It should be noted that the acquisition of historical inspection data is not limited to the internal data of the system, but can also integrate external data sources, such as the pedestrian flow statistics data of the examination room monitoring system, the access records of the access control system, etc.
[0147] The improvement of the improved A* algorithm is mainly reflected in the design of the heuristic function. The heuristic functions used in the traditional A* algorithm are usually Euclidean distance or Manhattan distance, and these simple distance metrics cannot fully consider the complex environmental characteristics of the special operation physical examination room. The improved A* algorithm introduces a multi-factor weighted heuristic function, comprehensively considering factors such as distance, historical travel time, congestion probability, and safety risk. The specific heuristic function can be expressed as:
[0148] h(n) = ω1·d(n, target node) + ω2·t + ω3·p + ω4·r;
[0149] Where d(n, target node) represents the Euclidean distance from node n to the target node, t represents the historical average travel time of node n, p represents the historical congestion probability of node n, r represents the safety risk value of node n, and ω1, ω2, ω3, and ω4 are the corresponding weight coefficients, which can be dynamically adjusted according to actual needs. For example, during the exam period, the value of ω3 can be increased to pay more attention to avoiding congested areas; near dangerous operation areas, the value of ω4 can be increased to pay more attention to safety. In the implementation process of the improved A* algorithm, a priority queue (usually a binary heap) is used to maintain the nodes to be expanded, and each time the node with the smallest f(n) = g(n) + h(n) value is taken out from the queue for expansion, where g(n) represents the actual cost from the starting point to node n.
[0150] In another embodiment, to improve the algorithm efficiency, a bidirectional A* search strategy can also be adopted, starting the search from both the starting point and the end point simultaneously. When the two search directions meet, the optimal path is found. In addition, to handle large-scale grid maps, a hierarchical A* algorithm (HPA*) can be used. First, a rough path is planned at the abstract level, and then the local path is optimized at the fine-grained level. The calculated global optimal path is a series of continuous grid nodes, representing the complete path from the starting point to the end point. These nodes are arranged in the order of access to form the inspection path. The calculation result of the global optimal path directly affects the inspection efficiency and safety and is the basis of dynamic path planning.
[0151] In this embodiment, the real-time monitoring system continuously collects and analyzes the dynamic changes of the inspection environment through multi-sensor fusion technology. The sensor configuration usually includes a lidar (LiDAR), a depth camera, an RGB camera on the robot platform, and fixed surveillance cameras distributed throughout the examination room. The lidar can accurately measure the distance and contour of surrounding objects by emitting laser beams and receiving reflected signals, forming 360-degree point cloud data, which is suitable for detecting the position and shape of obstacles. The depth camera can generate depth images and provide three-dimensional structure information of the scene, helping to identify obstacles at different heights. The RGB camera captures color images and, combined with computer vision algorithms, can identify the types and states of people, equipment, and other objects. The fixed surveillance cameras provide a wider field of view, covering the blind spots of the robot's own sensors. The crowd density monitoring uses object detection and tracking algorithms based on deep learning, such as YOLOv8 to detect people, and combines multi-object tracking algorithms such as DeepSORT to track the movement trajectories of people.
[0152] The crowd density calculation formula is: ρ = N / A, where N is the number of people in the area and A is the area of the area. According to the density value, the crowd situation can be divided into four levels: sparse (ρ < 0.2 people / m²), normal (0.2 people / m² ≤ ρ < 0.5 people / m²), crowded (0.5 people / m² ≤ ρ < 1 people / m²), and extremely crowded (ρ ≥ 1 people / m²).
[0153] Obstacle detection uses point cloud segmentation and clustering algorithms to divide the point cloud data into different clusters, and each cluster represents a potential obstacle. For each detected obstacle, its position coordinates, size, moving speed, direction and other attributes are recorded. The real-time monitoring system updates the environmental state at a frequency of 10Hz or higher to ensure that environmental changes can be captured in a timely manner. The monitoring data is transmitted to the central processing unit through a wireless network, registered and fused with the real-scene navigation map to form a real-time environmental state representation, providing a basis for subsequent path planning decisions.
[0154] The determination of path blocked areas and dangerous areas adopts a method of multi-index fusion, comprehensively considering factors such as pedestrian flow density, obstacle distribution, passage space width and safety risks. For pedestrian flow density, when the pedestrian flow density in a certain area exceeds a preset threshold (usually 0.8 people / ㎡), this area is marked as a potential path blocked area. The pedestrian flow density threshold can be adjusted according to the functional characteristics of different areas. For example, a higher threshold can be set at the entrance of an examination room, while a lower threshold can be set in the equipment operation area. For obstacle distribution, first calculate the space range occupied by the obstacles, and then evaluate the width of the remaining passage space. When the width of the remaining passage space is less than the width of the inspection equipment plus the safety margin (usually 1.5 times the width of the equipment), this area is marked as a path blocked area. For example, if the width of the inspection robot is 60cm and the safety margin is 30cm, when the passage space width is less than 90cm, it is determined as a path blocked area. The determination of dangerous areas considers various risk factors, including the density of people, the moving speed of obstacles, special equipment in the area (such as high-voltage electrical appliances, dangerous chemicals, etc.) and historical safety event records.
[0155] Risk assessment can adopt the fuzzy logic method, map each risk factor to a risk value in the range of [0,1], and then calculate the comprehensive risk index through weighted summation. When the comprehensive risk index exceeds the safety threshold, this area is marked as a dangerous area. During the determination process, the time factor also needs to be considered to distinguish between permanent blockage / hazard and temporary blockage / hazard. For each grid node on the globally optimal path, check whether it is located in the determined path blocked area or dangerous area. If there are such nodes, mark the corresponding section in the global path, and record the range, degree and estimated duration of the blocked / hazardous area, providing a basis for subsequent local path replanning. This multi-dimensional path state assessment ensures the safety and feasibility of the inspection task.
[0156] When an impassable or high-risk area is detected on the global path, it is necessary to adjust the path in a timely manner to avoid these areas while maintaining the overall optimality of the path as much as possible. Local path replanning first needs to determine the scope of replanning, usually centered on the blocked / dangerous area and expanding a certain distance (such as 10 meters) outward as the replanning area. Within the replanning area, improved path planning algorithms such as the D*Lite algorithm or the RRT* algorithm are applied to generate local alternative paths. The D*Lite algorithm is suitable for path replanning in dynamic environments because it can efficiently handle environmental changes and only recalculate the affected part of the path. The RRT* algorithm is suitable for dealing with environments with complex geometric constraints and can quickly find a feasible path. Local path replanning also needs to consider time factors. For temporary blockages (such as short-term personnel gatherings), a waiting strategy can be selected. When the expected waiting time is shorter than the detour time, choose to wait in place until the blockage is eliminated; when the expected waiting time is longer than the detour time, choose the detour strategy.
[0157] In this embodiment, the detour strategy needs to consider the balance of energy consumption, time, and safety. Through a multi-objective optimization method, the detour path with the minimum comprehensive cost is calculated. After the local path replanning is completed, it is necessary to seamlessly connect the local alternative path with the unaffected part of the original global path to form a new complete path. During the connection process, it is necessary to ensure the smoothness and continuity of the path and avoid sharp turns or unnecessary back-and-forth movements.
[0158] In one implementation manner of this embodiment, methods such as Bezier curves or spline interpolation can be applied to smooth the path connection points. Through local path replanning, it is possible to flexibly respond to local environmental changes while maintaining the overall optimality of the global path, ensuring the continuity and safety of the inspection task.
[0159] The identification of boundary coordinates uses computer vision and point cloud processing technologies, combined with real-time monitoring data and a real-scene navigation map. First, spatial clustering is performed on the detected blocked area or dangerous area, and adjacent blocked / dangerous grids are merged into a continuous area. The clustering algorithm can use the DBSCAN algorithm to group based on the spatial adjacency relationship and state similarity of the grids.
[0160] For each clustering area, its outer contour points are extracted as the boundary point set. Boundary extraction can use contour tracking algorithms such as the Moore neighborhood tracking algorithm. Starting from a boundary point of the area, move along the area boundary according to specific rules (such as clockwise) until returning to the starting point, and record all the boundary points passed along the way.
[0161] For subsequent processing convenience, the convex hull algorithm can be used to calculate the convex hull of the area as a simplified representation of the area. For areas with complex shapes, geometric figures such as the minimum bounding rectangle or ellipse can be used for approximate representation to further simplify the boundary description. The boundary coordinates also need to be appended with timestamp and confidence information to indicate the timeliness and reliability of the boundary.
[0162] In this embodiment, the passage cost value reflects the difficulty or risk level of the inspection device passing through a certain grid and is an important input for the path planning algorithm. The cost allocation adopts a multi-level cost model, and different cost values are set according to the degree of blockage / hazard. For completely blocked areas (such as the space occupied by fixed obstacles), an infinite cost value (in actual implementation, it can be a sufficiently large value, such as 10000) is set to ensure that the path planning algorithm absolutely avoids these areas.
[0163] For partially blocked areas (such as areas with a high density of people but still passable), gradient cost values are set according to the degree of blockage. The following formula can be used:
[0164] Passage cost value = Basic passage cost × (1 + α × Population density + β × Obstacle density);
[0165] Among them, the basic passage cost is usually 1, and α and β are weight coefficients used to adjust the influence degree of the population and obstacles on the cost.
[0166] It should be noted that for dangerous areas, the passage cost value not only considers the passage difficulty but also the safety risk. To achieve smooth path planning, the cost allocation is not limited to the grids within the blocked / hazardous areas, and a cost gradient also needs to be set around the area. The cost gradient can be implemented using a distance function. The closer the grid is to the boundary of the blocked / hazardous area, the higher the cost value; the farther away, the cost value gradually decreases to the basic cost. The following formula can be used for this gradient cost allocation:
[0167] d = Basic passage cost + (Maximum cost value - Basic passage cost) × exp(-d 2 / σ 2 );
[0168] Among them, d is the distance from the grid to the boundary of the blocked / hazardous area, the maximum cost value is the maximum cost value at the boundary, and σ is a parameter controlling the gradient decay rate. The cost allocation also needs to consider the time factor. For temporary blockage / hazard, the cost value will decay with time; for permanent blockage / hazard, the cost value remains unchanged. The update frequency of the cost value is consistent with the environmental monitoring frequency to ensure that the cost map can timely reflect environmental changes. By reasonably allocating the passage cost value, the path planning algorithm can naturally avoid high-cost areas and generate a safe and efficient inspection path.
[0169] Starting from the current position and ending at the target inspection point, the D*Lite algorithm is applied to calculate the locally optimal path. Here, the target inspection point is the grid node closest to the grid nodes in the distance path blocking area and / or dangerous area outside the boundary coordinates. The D*Lite algorithm is an incremental search algorithm suitable for path replanning in dynamic environments and can efficiently handle environmental changes without complete replanning.
[0170] In practical applications, the selection of the target inspection point is crucial. It is necessary to ensure that it is outside the boundary of the blocking / dangerous area and has the minimum deviation from the original planned path. The distance from the grid nodes outside the boundary of the blocking / dangerous area to the original planned path can be calculated, and the node with the minimum distance can be selected as the target inspection point. Through the D*Lite algorithm, the locally optimal path from the current position to the target inspection point can be quickly calculated, effectively avoiding the blocking / dangerous area while maintaining the overall optimality of the path.
[0171] When multiple consecutive path blocking areas and / or dangerous areas are detected, adopting a segmented planning strategy and decomposing the path into multiple sub-paths is an effective method for dealing with complex environments. Segmented planning first requires identifying all blocking / dangerous areas and analyzing their spatial distribution relationships. When the distance between two blocking / dangerous areas is less than a preset threshold (such as twice the length of the inspection device), they are regarded as continuous areas; when the distance is greater than the threshold, they are regarded as independent areas. For continuous blocking / dangerous areas, they can be merged into a large composite area for simplified processing; for independent blocking / dangerous areas, they can be processed separately. Path decomposition uses the key point extraction method, and a series of key points are selected on the original global path as the demarcation points of the sub-paths. The selection of key points considers various factors, including the position and shape of the blocking / dangerous areas, the topological structure of the environment, and the requirements of the inspection task, etc. Commonly used key points include: the entrance and exit points of the blocking / dangerous areas, the turning points of the path, the inspection points, and the feature points in the environment (such as doorways, corridor intersections, etc.). Key point extraction can use the curvature-based method to calculate the curvature of each point on the path and select the curvature peak points as candidate key points; it can also use the visibility-based method. Starting from a key point and moving along the path until a point where the line of sight is blocked is encountered, this point is used as the next key point. After determining the key points, the original path is decomposed into multiple sub-paths, and each sub-path connects two adjacent key points. The number of sub-paths depends on the environmental complexity and the distribution of the blocking / dangerous areas, usually 2 - 5. Segmented planning also needs to consider the coherence and smoothness between sub-paths to ensure that there are no sharp turns or unnecessary round-trip movements at the connection points of sub-paths. Path smoothing techniques such as Bezier curve interpolation or spline interpolation can be applied at the connection points to make the path transition more natural.
[0172] The advantage of segmented planning lies in decomposing the complex global path planning problem into multiple simple local path planning problems, reducing the computational complexity, improving the planning efficiency, and being able to respond more flexibly to local environmental changes.
[0173] The ant colony optimization algorithm is a swarm intelligence optimization algorithm inspired by the foraging behavior of ants and is suitable for solving combinatorial optimization problems such as path planning. When applying the ACO algorithm, the area between the starting point and the ending point of the sub-path needs to be discretized into grids first, and each grid node serves as a position that an ant can visit. The core of the algorithm is the pheromone update mechanism. The pheromone concentration represents the attractiveness of the path, and initially, the pheromone concentrations of all paths are equal. In each iteration, multiple "ants" (virtual search agents) start from the starting point and select the next visited node according to the pheromone concentration and heuristic information (such as the distance to the target point). After each ant completes the path search, the path quality is evaluated according to the path length and safety. The higher the quality of the path, the greater the pheromone increment it obtains.
[0174] In one implementation of this embodiment, to adapt to the characteristics of the special operation physical examination room, the ACO algorithm has been improved in many aspects: introducing a safety factor and incorporating the safety of the grid (distance from the blocked / hazardous area) into the heuristic information; adopting the elite ant colony strategy, and only the ants that find better paths can release pheromones; introducing a local search mechanism to locally optimize the found paths, such as path smoothing and redundant point deletion. The algorithm is iteratively executed until the maximum number of iterations is reached or the path quality does not improve significantly for several consecutive iterations. Finally, the path with the highest quality is selected as the local optimal sub-path. The advantage of the ACO algorithm is that it can effectively handle path planning problems in complex environments, especially in the case of multiple blocked / hazardous areas, and can find the optimal path that balances distance, safety, and smoothness.
[0175] When there are path blockages and / or dangerous areas in the global optimal path, it is necessary to execute a local path replanning strategy to generate a new safe path. The local path replanning strategy includes the following steps: First, identify the boundary coordinates of the path blockage area and / or dangerous area. For example, the boundary coordinates of a certain dangerous area are from (x1, y1) to (x2, y2). Then, in the real-time navigation map, assign high passage cost values to the grid nodes in these areas. For example, set the passage cost value of the dangerous area to 1000, which is much higher than that of the ordinary area. Next, starting from the current position and ending at the target inspection point, apply the D*Lite algorithm to calculate the local optimal path. The D*Lite algorithm is an incremental path planning algorithm that can quickly update the path when the environment changes. For example, in a certain local path replanning, the D*Lite algorithm generates a shortest path that bypasses the dangerous area. When multiple consecutive path blockage areas and / or dangerous areas are detected, adopt a segmented planning strategy, decompose the path into multiple sub-paths, and apply the ant colony optimization algorithm to each sub-path to generate local optimal sub-paths. For example, a certain path is decomposed into 3 sub-paths, and the ant colony optimization algorithm generates optimal solutions for each sub-path. Finally, connect all the local optimal sub-paths to obtain the local replanned path. For example, a certain local replanned path is composed of 3 sub-paths connected together, with a total length of 50 meters, which is 10 meters shorter than the original path.
[0176] It should be noted that connecting all the local optimal sub-paths to obtain the local replanned path is a dynamic path planning, and it is necessary to ensure that the connected path is smooth, continuous and feasible.
[0177] When connecting sub-paths, it is first necessary to check whether the connection points of adjacent sub-paths are consistent. If not, it is necessary to generate a transition path segment. The generation of the transition path segment can adopt the Bezier curve interpolation method, and the selection of control points considers the tangent direction of the path to ensure smooth transition.
[0178] The connected path may have redundant points and unnecessary twists and turns, and path optimization is required. Path optimization includes point simplification and smoothing. Point simplification can use the Douglas-Peucker algorithm to reduce the number of path points while maintaining the path shape; smoothing can use methods such as moving average or spline interpolation to eliminate sharp corners and mutations in the path.
[0179] Through the dynamic path planning strategy, this embodiment can effectively cope with the complex changes in the examination room environment, ensuring the efficient operation of the inspection equipment and the safety of the examination room. Obtaining historical inspection data provides a reliable reference for path planning, and the improved A* algorithm improves the efficiency and accuracy of global path planning. Real-time monitoring of the crowd density and obstacle distribution can promptly detect path-blocked areas and dangerous areas, providing a basis for local path replanning. The local path replanning strategy uses the D*Lite algorithm and the ant colony optimization algorithm to generate a safe and efficient new path, avoiding operation delays and safety hazards of the inspection equipment. It not only improves the inspection efficiency and accuracy but also expands the inspection coverage, effectively reducing the risk potential in the examination room and providing strong guarantee for the safety of the examination room personnel.
[0180] In one implementation manner of this embodiment, based on the real-scene navigation map and image features, a dynamic path planning strategy is executed to generate the inspection path of the inspection equipment, including the following steps:
[0181] S410. Divide the real-scene navigation map into multiple grid nodes and calculate the traversal cost value of each grid node;
[0182] S420. Respond to the user's setting instruction for the real-scene navigation map to determine the inspection points in the real-scene navigation map and the weight coefficient corresponding to each inspection point;
[0183] S430. Use the improved A* algorithm to calculate the initial global path;
[0184] S440. Take the environmentally acquired parameters and image features obtained in real time as input parameters of a pre-constructed multi-feature fusion model to dynamically update the traversal cost value of each grid node;
[0185] S450. When the change amount of the traversal cost value exceeds a preset threshold, use a preset path replanning algorithm to generate a new inspection path;
[0186] S460. When the change amount of the traversal cost value does not exceed the preset threshold, take the initial global path as the inspection path of the inspection equipment.
[0187] In this embodiment, first, the real - scene navigation map needs to be divided into multiple grid nodes, and the traversal cost value of each grid node is calculated. Specifically, the real - scene navigation map is usually based on high - precision map data of the actual environment, including information such as the geometric structure of the environment, the distribution of obstacles, and terrain features. Dividing the real - scene navigation map into grids is to discretize the continuous space, facilitating subsequent path - planning algorithms. The grid division uses the uniform grid method, that is, the entire map area is divided into regular square grids according to a preset grid size (such as 0.5 meters × 0.5 meters). Each grid node represents a position point in the actual environment and has a unique coordinate identifier (i, j), where i represents the row index and j represents the column index. The choice of grid size needs to balance the computational complexity and planning accuracy. Too large a grid will lead to insufficient path - planning accuracy, while too small a grid will increase the computational burden.
[0188] For each grid node, its traversal cost value needs to be calculated, which reflects the difficulty for the inspection device to pass through this node. The calculation of the traversal cost value comprehensively considers various factors, including: terrain factors (such as slope, flatness), obstacle factors (such as the distance to static obstacles, the predicted position of dynamic obstacles), environmental factors (such as lighting conditions, weather conditions), and task - specific factors (such as device power requirements, signal coverage intensity).
[0189] In response to the user's setting instruction for the real - scene navigation map, the inspection points in the real - scene navigation map and the weight coefficient corresponding to each inspection point are determined. An inspection point refers to a key position that the inspection device must reach during the inspection task, usually equipment, facilities, or areas that need to be inspected with emphasis. The user can directly mark inspection points on the real - scene navigation map through an interactive interface or batch - set them by importing a predefined list of inspection points. Each inspection point has a clear coordinate position (x, y) on the map and is mapped to the nearest grid node. In practical applications, the setting of inspection points usually considers factors such as equipment distribution, inspection frequency requirements, and safety - hazard risks to ensure that all key areas are covered by the inspection.
[0190] Figure 4 shows a schematic diagram of the setting of inspection points on the real - scene navigation map provided by the embodiment of the present application, as Figure 4 shown, which shows the overall system architecture of the setting of inspection points on the real - scene navigation map, including: the user marks inspection points or imports a predefined list through an interactive interface, the storage and processing process of inspection - point data, the process of mapping coordinates to grid nodes, the key attributes of inspection points (coordinates, weight coefficients, etc.), and the components of the real - scene navigation map.
[0191] In a specific implementation, for each inspection point, the user can also set a corresponding weight coefficient, which reflects the importance or priority of the inspection point. The value range of the weight coefficient is usually between 0 and 1, and the larger the value, the more important the inspection point. The determination of the weight coefficient can be based on various factors, such as equipment failure history, equipment importance, inspection urgency, etc. For example, for critical equipment in high-risk areas, a relatively high weight coefficient (such as 0.8 - 1.0) can be set; for ordinary equipment in routine inspections, a medium weight coefficient (such as 0.4 - 0.7) can be set; for auxiliary inspection points, a relatively low weight coefficient (such as 0.1 - 0.3) can be set.
[0192] For example, if a certain device has abnormal data recently, the system will automatically increase the weight coefficient of the inspection point corresponding to this device.
[0193] For the improved A* algorithm, please refer to the above description, and it will not be elaborated here in this application. Through the improved A* algorithm, the initial global path can be calculated.
[0194] The multi-feature fusion model is a deep learning model that can comprehensively process multi-source heterogeneous data, mainly composed of a feature extraction module, a feature fusion module, and a cost prediction module. The feature extraction module is responsible for extracting effective features from data from different sources, including environmental parameters such as temperature, humidity, and light obtained from environmental sensors, as well as image data obtained from cameras. For image data, a pre-trained convolutional neural network (such as ResNet-50 or EfficientNet) is used to extract high-level semantic features, which can identify key information such as obstacles, terrain changes, and human activities.
[0195] In this embodiment, the feature fusion module uses the attention mechanism and cross-modal fusion technology to effectively integrate features from different sources. Specifically, first, various features are normalized, and then the relevance between different features is calculated through the multi-head attention mechanism. The formula is as follows:
[0196]
[0197] Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively, which are transformed from different features, and d k is the feature dimension. Through the attention mechanism, the model can automatically learn the importance weights of different features and perform weighted fusion.
[0198] The cost prediction module then predicts the update amount of the passage cost value of each grid node based on the fused features. This module adopts a multi-layer perceptron structure, and the ReLU activation function is used in the last layer to ensure that the output cost value is non-negative. The prediction formula can be:
[0199] ΔC(i,j) = MLP(FF(i,j));
[0200] Where ΔC(i,j) is the updated value of the traversal cost of grid node (i,j), and FF(i,j) is the fused feature corresponding to this node.
[0201] In practical applications, the multi-feature fusion model triggers an update calculation at fixed time intervals (such as 1 second) or when a significant environmental change is detected. The updated traversal cost value is:
[0202] C new (i,j) = C old (i,j) + ΔCost(i,j);
[0203] In this way, environmental changes can be sensed in real time, such as detecting temporary obstacles (such as pedestrians, vehicles), identifying adverse terrain conditions (such as slippery ground, temporary water accumulation), or sensing changes in environmental conditions (such as insufficient light, bad weather), and the traversal cost value can be adjusted accordingly to ensure the safety and feasibility of the inspection path.
[0204] When the change amount of the traversal cost value exceeds a preset threshold, a preset path replanning algorithm is adopted to generate a new inspection path. First, evaluate whether the change amount of the traversal cost value exceeds the preset threshold.
[0205] When the change amount exceeds the preset threshold, the path replanning algorithm is triggered. To balance computational efficiency and path quality, a hierarchical replanning strategy is adopted: First, evaluate the impact degree of the changed area on the current path. If the changed area has no intersection with the current path or has little impact, only the local path is replanned; if the changed area has a significant impact on the current path, a global path replanning is performed. The local path replanning uses the D*Lite algorithm, which is a variant of the A* algorithm and is designed for path replanning in dynamic environments and can efficiently update the affected path segments. The global path replanning re-executes the improved A* algorithm but uses the previously calculated result as the initial solution to accelerate the convergence process.
[0206] Through the threshold-based path replanning mechanism, the inspection path can be adjusted in time when the environment changes significantly, ensuring the safety and effectiveness of the inspection task, while avoiding unnecessary waste of computing resources.
[0207] When the change amount of the traversal cost value does not exceed the preset threshold, the initial global path is used as the inspection path of the inspection device. At this time, the initial global path is maintained, but the following several optimization processes are carried out: First, perform path validity verification to ensure that there are no new impassable areas on the current path. The verification method is to check whether the traversal cost values of all grid nodes on the path become infinite. If so, local path adjustment is triggered.
[0208] In this embodiment, by dividing the real - scene navigation map into grid nodes and calculating the cost of passage, combining the inspection points and weight coefficients set by the user, using the improved A* algorithm to generate an initial global path, dynamically updating the cost of passage based on the multi - feature fusion model, and deciding whether to re - plan the path according to the change amount, intelligent and adaptive inspection path planning is realized, so that it can effectively cope with complex and changeable environments, balance inspection efficiency and safety, reduce waste of computing resources, and improve the completion quality of inspection tasks.
[0209] In one implementation of this embodiment, based on image features, the inspection path is divided into n regions, including the following steps:
[0210] S510. Cluster the image features based on the K - means clustering algorithm to obtain at least one clustering region;
[0211] S520. Calculate the feature vectors of each clustering region and calculate the similarity between the feature vectors;
[0212] S530. Divide the inspection path into n continuous regions according to the similarity between the feature vectors, where each region is assigned a unique identifier.
[0213] In the actual implementation process, image features are usually high - dimensional vectors (such as 2048 - dimensional), which can effectively capture the semantic information and visual features of images. To reduce the computational complexity, these high - dimensional features are usually dimension - reduced, such as using principal component analysis (PCA) to reduce the feature dimension to 128 or 256 dimensions while retaining most of the effective information.
[0214] After obtaining the image features, the K - means clustering algorithm is applied to cluster these features. The K - means algorithm is an iterative clustering method, and its core idea is to divide n data points into k clusters so that each data point belongs to the cluster center closest to it. The specific steps of the algorithm include: first, randomly select k points as the initial cluster centers; then iteratively execute two steps: (1) assign each data point to the cluster represented by the closest cluster center, (2) recalculate the center point of each cluster (i.e., the average value of all points within the cluster); when the cluster center points no longer change significantly or reach the maximum number of iterations, the algorithm terminates.
[0215] In practical applications, the value of K can be set according to the requirements of the patrol inspection task and the environmental complexity. For example, in a complex and variable industrial environment, more clustering regions (such as K = 8 - 12) may be required; while in a relatively simple environment, fewer clustering regions (such as K = 3 - 5) may be sufficient. After clustering, each clustering region represents a set of images with similar visual features, and these regions may be continuous in physical space (such as the same type of corridor) or scattered (such as equipment regions distributed at different locations but with similar features).
[0216] In the previous step, the image features have been clustered into multiple regions by the K-means algorithm, and each region contains a set of image points with similar features. To more accurately characterize the characteristics of each clustering region, a representative feature vector needs to be calculated for each region. The feature vector comprehensively considers the feature distribution of all image points within the region. Specifically, the calculation of the feature vector can adopt the method of weighted average, that is, calculate the weighted average of all feature vectors to reduce the influence of outliers and improve the representativeness of the feature vector.
[0217] After obtaining the feature vectors of each clustering region, it is necessary to calculate the similarity between the feature vectors to evaluate the degree of association between different regions. The calculation of similarity can adopt the calculation method of cosine similarity to measure the similarity of the directions of two vectors.
[0218] In the previous two steps, the clustering regions have been obtained through K-means clustering, and the feature similarities between regions have been calculated. However, these clustering regions may be scattered in space, while the patrol inspection path needs to be a continuous linear structure. Therefore, it is necessary to map the clustering results onto the patrol inspection path to form a continuous path region division. This process uses the dynamic programming algorithm, regarding the patrol inspection path as a one-dimensional sequence, and finding the optimal segmentation points to make the internal feature similarity of the divided regions high and the feature differences between regions significant.
[0219] The specific implementation method is as follows. First, represent the patrol inspection path as a series of ordered position point sequences P = {p1, p2,..., p m}, and each position point is associated with the label of the clustering region it belongs to. Then, define an objective function to evaluate the quality of the path division:
[0220]
[0221] where S = {S1, S2,..., S n} is a division scheme of the path, S n represents the nth region, and I(S i ) measures the similarity within the region S i (the higher the better), and I(Si ,S i+1 ) Measure the similarity between adjacent regions (the lower the better), where λ1 and λ2 are weight coefficients.
[0222] Use the dynamic programming algorithm to solve the optimal partitioning scheme: Define DP[i][j] to represent the optimal score for partitioning the first i points of the path into j regions. The state transition equation is:
[0223]
[0224] Among them, Score(k + 1, i) represents the score for taking the position points p k+1 to p i as a region. By filling the DP table, the optimal partitioning scheme can finally be obtained.
[0225] After determining n consecutive regions, assign a unique identifier to each region. Specifically, digital coding (such as 1, 2, ……, n) can be used as the basic identifier; secondly, combined with the main feature types of the regions, such as "A1" representing a device-dense area, "B2" representing a corridor area, "C3" representing an open space, etc.; in addition, the location information of the regions can also be encoded in the identifier, such as "N1" representing the northern region 1, "SE2" representing the southeast region 2, etc.
[0226] Finally, each region not only has a unique identifier but also is associated with a series of metadata, such as the main feature description of the region, a list of key devices, the recommended inspection speed, and precautions, etc. This information can be stored in the region attribute database for the inspection system to refer to when performing tasks. In this way, the inspection path is divided into continuous regions with clear boundaries and rich semantic information, providing a structured spatial framework for subsequent inspection strategy formulation and execution.
[0227] In this embodiment, the K-means clustering algorithm is used to cluster the image features, calculate the feature vectors and their similarities of the clustering regions, and accordingly divide the inspection path into continuous regions with unique identifiers, realizing the intelligent zoning management of the inspection path in a complex environment, achieving the environmental understanding based on visual features, enabling the inspection system to identify and distinguish different types of regions (such as device areas, corridors, open spaces, etc.), and improving the environmental perception ability; secondly, optimizing the inspection strategy through region division, formulating differentiated inspection parameters according to different region characteristics, and improving the inspection efficiency; in addition, simplifying the task planning and anomaly location, enabling operators to accurately refer to specific regions, reducing communication costs, and improving the fault response speed; the region division provides a spatial index framework for data management, facilitating the organization and retrieval of inspection data by region and supporting more refined data analysis.
[0228] In one implementation of this embodiment, the environmental parameters and image features are fused and analyzed to determine the congestion index on the inspection path, including the following steps:
[0229] S610. Extract the personnel density, equipment distribution, and operation behavior from the image features through the YOLOv8 model;
[0230] S620. Output the congestion index through a multi-layer perceptron neural network model. Among them, the input layer of the multi-layer perceptron neural network model includes environmental parameters, personnel density, equipment distribution, and operation behavior, and the output layer is the congestion index.
[0231] In this embodiment, YOLOv8 is a real-time object detection algorithm. Its core principle is to divide the image into grids and simultaneously predict the target bounding boxes, class probabilities, and confidence levels that may exist in each grid cell during a single forward propagation. Compared with traditional two-stage detectors, YOLOv8 adopts a single-stage detection architecture, which greatly improves the detection speed and maintains a high detection accuracy through multiple technological innovations.
[0232] The network architecture of YOLOv8 mainly consists of three parts: the backbone network, the neck network, and the detection head. The backbone network is responsible for extracting multi-scale features of the image. It adopts the CSPDarknet structure, which reduces the computational amount while maintaining the feature extraction ability through techniques such as depthwise separable convolution and cross-stage partial connection; the neck network adopts the PANet (Path Aggregation Network) structure to achieve the fusion of features at different scales and enhance the detection ability for targets of different sizes; the detection head is responsible for the final target localization and classification, and improves the detection accuracy through multi-scale prediction boxes and dynamic assignment strategies.
[0233] In practical applications, the YOLOv8 model needs to be specifically trained for the inspection scenario. The training dataset contains a large number of labeled industrial environment images, covering various personnel, equipment, and operation behavior categories. Specifically, the personnel categories can be subdivided into workers, managers, visitors, etc.; the equipment categories can include fixed equipment (such as production lines, control cabinets) and mobile equipment (such as forklifts, transport vehicles); the operation behaviors can be divided into routine operations (such as equipment inspection, material handling) and special operations (such as maintenance, debugging). During the training process, a transfer learning strategy is adopted, based on the pre-trained YOLOv8 model, and it is adapted to the specific scenario through fine-tuning. At the same time, data augmentation techniques (such as rotation, scaling, color jitter, etc.) are applied to improve the generalization ability of the model.
[0234] After the model training is completed, during the inspection process, the real-time collected images will be sent to the YOLOv8 model for processing. The model outputs include the bounding box coordinates, class labels, and confidence scores of various detected objects. Based on these raw outputs, three key metrics are further calculated:
[0235] 1. Personnel density: It is calculated by counting the number of detected personnel within a unit area. The specific calculation formula is:
[0236]
[0237] where n is the total number of detected objects, w i is the weight of the i-th object (which can be determined based on the confidence level and object size), I(·) is the indicator function, which takes the value of 1 when the i-th object belongs to the set C of personnel categories 人员 and 0 otherwise, and A is the actual area (square meters) covered by the image.
[0238] 2. Equipment distribution: The kernel density estimation method is used to generate a heat map of equipment distribution and extract distribution features such as aggregation degree and uniformity. The equipment distribution metric can be expressed as:
[0239] Equipment distribution = {equipment density, aggregation degree, uniformity, proportion of main equipment types}
[0240] 3. Operational behavior: Analyze the interaction relationship between the detected personnel and equipment to identify different types of operational behaviors. Through time series analysis and pose estimation, static operations (such as observing and recording) and dynamic operations (such as handling and debugging) are distinguished. The operational behavior metric can be expressed as:
[0241] Operational behavior = {operation type distribution, operation intensity, operation duration, proportion of abnormal operations}
[0242] To improve the stability of feature extraction, the analysis results of consecutive multiple frames of images are usually subjected to temporal smoothing to reduce the influence of fluctuations in single-frame detection. At the same time, combined with the position and viewing angle information of the camera, spatial correction is performed on the detection results to ensure the comparability of data obtained at different positions.
[0243] These three types of features (personnel density, equipment distribution, and operational behavior) extracted by the YOLOv8 model provide key inputs for subsequent calculation of the congestion index, and they jointly reflect the activity status and potential congestion risks in the industrial environment. This vision-based feature extraction method has the advantages of wide coverage and rich information compared with traditional sensor monitoring, and can comprehensively capture various factors affecting the traffic conditions of the inspection path.
[0244] Fuse and analyze the environmental parameters with the image features extracted in the previous step, and finally generate the congestion index of the inspection path. The Multi-Layer Perceptron (MLP) is a feedforward neural network consisting of an input layer, one or more hidden layers, and an output layer. The layers are connected by fully connected connections and can learn complex non-linear mapping relationships between the input and output. In this application, the architecture of the MLP model is carefully designed to achieve the effective fusion of environmental parameters and image features.
[0245] The input layer of the MLP model receives various types of feature data. First are the environmental parameters, including physical environmental indicators such as temperature, humidity, noise level, light intensity, air flow velocity, etc. These parameters are usually collected in real time by various sensors distributed on the inspection path. Second are the three types of image features extracted by the YOLOv8 model in the previous step: personnel density, equipment distribution, and operation behavior. To enable the effective fusion of features of different types and dimensions, first use Z-score normalization to standardize all input features.
[0246] The hidden layer of the MLP model adopts a multi-layer structure. A typical configuration includes 3 - 5 hidden layers, with each layer containing 64 - 256 neurons. The hidden layer uses the ReLU (Rectified Linear Unit) activation function, and its mathematical expression is:
[0247] f(x) = max((0, x);
[0248] The ReLU activation function has advantages such as simple calculation and stable gradients, which helps to solve the problem of gradient disappearance in the training of deep networks. Between the hidden layers, batch normalization technology can also be applied to accelerate network training and improve model stability by standardizing the input distribution of each layer. To prevent overfitting, a Dropout layer is added between the hidden layers, randomly discarding a certain proportion (usually 0.2 - 0.5) of neurons to enhance the generalization ability of the model.
[0249] The output layer of the MLP model contains one neuron, and the Sigmoid activation function is used to map the output to the interval [0, 1], representing the normalized value of the congestion index. The mathematical expression of the Sigmoid function is:
[0250]
[0251] The closer the value of the congestion index is to 1, the more congested the path is; the closer it is to 0, the smoother the path is. To improve the interpretability of the congestion index, the interval [0, 1] can be divided into multiple levels, such as 0 - 0.2 for "smooth", 0.2 - 0.4 for "slightly congested", 0.4 - 0.6 for "moderately congested", 0.6 - 0.8 for "severely congested", and 0.8 - 1.0 for "extremely congested".
[0252] The training of the MLP model adopts the supervised learning method and requires constructing a training data set containing input features and congestion index labels. The label data can be obtained through multiple methods: one is expert annotation, where domain experts evaluate the congestion degree based on scene images and environmental data; the second is rule-based automatic annotation, which calculates the preliminary congestion index according to predefined rules (such as personnel density threshold, travel time, etc.); the third is semi-supervised learning, which combines a small amount of labeled data and a large amount of unlabeled data for training. During the training process, the mean squared error (MSE) or mean absolute error (MAE) is used as the loss function, the Adam optimizer is used for parameter update, the learning rate is initially set to 0.001, and a learning rate decay strategy is used.
[0253] After the model training is completed, in actual applications, the real-time collected environmental parameters and image features are fed into the MLP model, and the congestion index of the current inspection path is output. This index not only reflects the immediate congestion state of the path but also can be used to predict the congestion trend in the short term, providing a decision-making basis for inspection path planning and adjustment. Through regular retraining and online learning, the model can continuously adapt to environmental changes and new congestion patterns and maintain prediction accuracy.
[0254] The technical solution of this embodiment for fusing environmental parameters and image features to determine the congestion index on the inspection path extracts the personnel density, equipment distribution, and operation behavior in the image through the YOLOv8 model, and combines the multi-layer perceptron neural network model to fuse and analyze these features and environmental parameters, realizing the accurate assessment of the congestion status of the inspection path in the industrial environment. This solution first realizes the intelligent fusion of multi-source data, organically combines traditional environmental sensing data with the results of advanced computer vision analysis, creates a more comprehensive and accurate scene understanding ability, and effectively improves the congestion prediction accuracy; secondly, it provides a real-time congestion index, which not only reflects the current state but also can predict the future congestion trend, making the inspection task planning more forward-looking; in addition, through fine-grained congestion analysis, different types of congestion causes (such as personnel gathering, equipment blockage, special operations, etc.) can be identified, and it supports adaptive inspection path optimization, and the inspection route and time arrangement can be dynamically adjusted according to the real-time congestion index, reducing the inspection delay time.
[0255] In one implementation of this embodiment, determining a dangerous area among n areas based on the congestion index includes the following steps:
[0256] S710. When the congestion index of an area is lower than the preset first congestion index threshold T1, mark the area as a safe area;
[0257] S720. When the congestion index of an area is between the first congestion index threshold T1 and the second congestion index threshold T2, mark the area as a warning area;
[0258] S730. When the congestion index of an area is higher than the second congestion index threshold T2, mark the area as a dangerous area, where 0 < T1 < T2 < 100.
[0259] When the congestion index of an area is lower than the preset first congestion index threshold T1, mark the area as a safe area. The first congestion index threshold T1 is a scientifically analyzed and practically verified safe upper limit value, representing the maximum congestion level at which the area can maintain normal operation without posing potential risks to personnel safety and equipment operation.
[0260] During the implementation process, the system continuously monitors the congestion index of each area. When it detects that the congestion index of a certain area is lower than T1, it immediately marks the area as a safe area. After being marked as a safe area, the area is usually represented in green on the visualization interface, and the relevant safety status information is recorded in the database for subsequent statistical analysis and report generation.
[0261] The marking of the safe area has multiple practical significances: First, it indicates that the current personnel and equipment density in this area is within a controllable range, and normal production and operation activities can proceed smoothly; Second, the safe area can be used as a priority evacuation route or a temporary assembly point in case of emergency; Third, the proportion and distribution of the safe area can reflect the safety status of the overall environment and provide a basis for management decisions.
[0262] It should be noted that the determination of the safe area is a dynamic process. As the personnel flow and equipment operation status in the area change, the safety status of the area may change. Therefore, the system needs to update the congestion index calculation results at an appropriate frequency (usually every 5 - 10 seconds) to ensure the real-time and accuracy of the safety status assessment. In addition, to avoid frequent state switching caused by short-term fluctuations, a time smoothing mechanism can be introduced, that is, only when the congestion index is lower than T1 at multiple consecutive time points (such as 3 consecutive detections), the area is marked as a safe area.
[0263] When the congestion index of a region is between the first congestion index threshold T1 and the second congestion index threshold T2, the region is marked as a warning area. The warning area represents a transitional state between safety and danger, indicating that the congestion level within the region has exceeded the ideal range but has not reached the dangerous level that requires immediate intervention.
[0264] The interval width between the first congestion index threshold T1 and the second congestion index threshold T2 directly affects the sensitivity of the warning mechanism. If the interval is too narrow, it may cause the system to be overly sensitive and trigger warnings frequently; if the interval is too wide, the warning may lose its timeliness. Therefore, the difference between T1 and T2 is usually set at 15 to 30 percentage points. This range can provide enough buffer time for preventive intervention without overly expanding the coverage of the warning state.
[0265] In practical applications, the marking of the warning area is usually accompanied by the initiation of a series of preventive measures. First, on the visualization interface, the warning area will be represented by yellow or orange, forming an intuitive visual cue; second, the system will send a warning notification to relevant management personnel, and the notification methods can include interface pop-ups, text messages, emails, or push messages from a dedicated application; third, for specific types of regions, preset response strategies may be triggered, such as adjusting the personnel flow in adjacent regions, temporarily restricting new personnel from entering, increasing the monitoring frequency, etc.
[0266] To improve the accuracy of the warning mechanism, the change rate of the congestion index can be introduced as an auxiliary judgment factor. Even if the congestion index has not reached T1, but if its growth rate is extremely fast (for example, it increases by more than 10 percentage points in a short period), it may also trigger an early warning.
[0267] When the congestion index of a region is higher than the second congestion index threshold T2, the region is marked as a dangerous area. This step is the highest warning level in the regional safety status assessment, indicating that the congestion level within the region has reached the critical level that may threaten personnel safety, equipment operation, or production order, and immediate intervention measures are required. The second congestion index threshold T2 is a critical value determined based on safety risk assessment and emergency management requirements, and is usually set between 60 and 80, and the specific value depends on the characteristics and safety requirements of the region.
[0268] In practical applications, when a certain area is marked as a dangerous area, a series of emergency response measures will be triggered. First, on the visual interface, the dangerous area will be represented by red or flashing red, and may be accompanied by audible and visual alarm signals; second, the system will immediately send high-priority alarms to security managers, area leaders, and relevant emergency teams, ensuring the timely transmission of information through multiple channels; third, according to the preset emergency plan, corresponding intervention measures will be initiated, such as evacuating non-essential personnel in the area, suspending or adjusting the examination room activities, starting emergency equipment (such as exhaust systems, fire-fighting equipment), etc.; fourth, detailed event logs will be recorded, including the start time, duration, maximum congestion index value of the dangerous state, and the response measures taken, providing a basis for subsequent analysis and improvement.
[0269] In addition, the determination of the dangerous area can also be combined with other safety monitoring indicators to form a multi-dimensional risk assessment. For example, by combining environmental parameters such as temperature, gas concentration, and noise level, or information such as equipment operating status and abnormal behavior detection results, a more comprehensive dangerous state assessment model can be constructed:
[0270] Degree of danger = w1·Congestion index + w2·Environmental risk index + w3·Behavior risk index +...;
[0271] Among them, w1, w2, w3, etc. are the weight coefficients of each index, which can be adjusted according to the characteristics and safety priorities of different areas. When the comprehensive degree of danger exceeds the preset threshold, the dangerous area is marked.
[0272] The timely identification and handling of dangerous areas are of crucial significance for preventing safety accidents, ensuring personnel safety, and maintaining normal production order. By establishing clear determination criteria and response mechanisms, effective intervention can be carried out before the risk evolves into actual harm, minimizing potential losses to the greatest extent.
[0273] In this embodiment, by setting two key thresholds, a three-level safety state assessment system is established, realizing the precise classification and timely warning of the regional safety status.
[0274] In one implementation manner of this embodiment, the camera is also used to collect examination room videos, and the examination room videos include multiple video frames. The following steps are further included:
[0275] S810: Input the video frames into a pre-constructed YOLOv8 model to identify predefined types of violation behaviors;
[0276] S820: For each type of violation behavior, record the start time and end time of the violation behavior;
[0277] S830. Extract key video frames at a frequency of m frames per second within a preset time interval, and add a timestamp, a violation type label, and a bounding box of the violation area to each key video frame;
[0278] S840. Sort the key video frames in chronological order, generate an evidence package of violation behaviors, store the evidence package of violation behaviors in the security database of the central control device, and construct an index of violation behaviors in the security database.
[0279] Input the video frames into a pre - constructed YOLOv8 model to identify predefined violation behavior types. For the specific content of the YOLOv8 model, please refer to the above description, and it will not be elaborated in this embodiment of the present application. In practical applications, the examination room videos captured by the camera will be decomposed into a sequence of consecutive video frames. Each frame is a static image, and these images will be sent to a pre - trained YOLOv8 model for processing and analysis.
[0280] In the model input stage, each video image frame will undergo pre - processing, including size adjustment (usually adjusted to 640×640 or 416×416 pixels), normalization (scaling pixel values to the range of 0 - 1), and channel arrangement adjustment, etc., to meet the input requirements of the YOLOv8 model. The processed image data will be sent to the input layer of the model, and then through the multi - layer convolutional neural network of the model for feature extraction and analysis.
[0281] The working principle of the YOLOv8 model is to divide the input image into a grid of S×S, and each grid is responsible for predicting the objects contained therein. For each grid cell, the model will predict B bounding boxes and their confidence scores, as well as the probability distribution of C categories.
[0282] In the application of examination room monitoring, the predefined violation behavior types usually include but are not limited to: whispering to each other (communication between two or more examinees), passing items, using unauthorized electronic devices (such as mobile phones, smart watches), looking at others' test papers, using hidden reference materials, leaving the seat, etc. The YOLOv8 model will analyze each frame of the image and output the detected violation behavior types, locations (represented by bounding boxes), and confidence scores.
[0283] Compared with traditional manual monitoring, the automated recognition system can monitor multiple examination rooms simultaneously, and will not miss violation behaviors due to fatigue or distraction, greatly improving the invigilation efficiency and fairness. In addition, the recognition method based on deep learning can also adapt to different environmental conditions and behavior changes. Through continuous learning and updating, the recognition accuracy can be continuously improved.
[0284] For each type of violation, record the start time and end time of the violation. This step is a key link in tracking and recording the violation event in the time dimension after the YOLOv8 model successfully identifies the violation. In the actual implementation process, the time recording of violations adopts a continuous tracking mechanism, and the specific process is as follows: When the YOLOv8 model first detects a specific type of violation (such as using a mobile phone) in a certain frame and the confidence level exceeds the preset threshold (usually set to 0.75 or 0.8), the system will record the current timestamp as the start time of the violation. This timestamp usually contains information such as year, month, day, hour, minute, second, and millisecond to ensure the accuracy of time recording. For example, if it is detected that candidate A uses a mobile phone at 09:45:23.456 on June 15, 2023, this time point will be recorded as the start time of this violation.
[0285] After that, continuously track the presence of this violation in subsequent frames. Since the video is a continuous sequence of frames, violations are usually detected in multiple consecutive frames. To avoid wrongly believing that the violation has ended due to temporary detection failures (such as brief undetected due to occlusion, light change, etc.), a time tolerance window (usually 0.5 - 2 seconds) can be set. Only when the violation is not detected in multiple consecutive frames or exceeds the preset time window will it be considered that the violation has ended, and the timestamp of the last detection of this behavior will be recorded as the end time.
[0286] For multiple possible violations, the system will maintain time records separately for each behavior. For example, if a candidate first whispers to each other (09:30:15 - 09:30:45) and then uses a mobile phone (09:35:20 - 09:35:40), the start and end times of these two different types of violations will be recorded separately. The data structure of time records usually includes the following fields: violation ID (unique identifier), violation type (such as "using a mobile phone", "whispering to each other", etc.), start time (timestamp accurate to milliseconds), end time (timestamp accurate to milliseconds), duration (end time minus start time), violator (such as candidate ID or seat number), confidence score (the degree of certainty of the model's judgment on this violation), etc.
[0287] Within the preset time interval, extract key video frames at a frequency of m frames per second, and add timestamps, violation type labels, and violation area bounding boxes to each key video frame to ensure the integrity of evidence while avoiding storing and processing too much redundant data and improving the efficiency and availability of the system.
[0288] The preset time interval generally refers to a time range from a few seconds before the start time of the violation to a few seconds after the end time. For example, if the system records a whispering behavior occurring from 09:45:30 to 09:45:50, the preset time interval may be set from 09:45:25 to 09:45:55, that is, expanding 5 seconds before and after the violation to capture the complete behavior process and context.
[0289] The extraction frequency parameter m of the key video frames is an adjustable value, usually determined according to the type, duration of the violation and system resource limitations. For rapidly changing violations (such as passing items), a higher frame rate may be required (such as m = 10, that is, 10 frames are extracted per second) to ensure capturing the key moments of the behavior; while for violations with a long duration and slow changes (such as viewing non-permitted materials for a long time), a lower frame rate (such as m = 2) can be used to save storage space. In practical applications, the m value is usually set between 2 and 15 to balance evidence integrity and system resource consumption.
[0290] The extraction of key video frames can be carried out in two ways: uniform sampling or intelligent sampling. Uniform sampling means that within the preset time interval, one frame is extracted every 1 / m seconds; while intelligent sampling dynamically adjusts the sampling frequency based on factors such as the degree of change of the frame content and the confidence score of the violation behavior, extracting more frames at critical moments (such as at the beginning or near the end of the violation behavior) and fewer frames in relatively static stages.
[0291] For each extracted key video frame, the following annotation information can be added: timestamp, violation type label, and violation area bounding box.
[0292] Specifically, the timestamp can accurately record the acquisition time of the frame, usually including date and time information. The timestamp can be added by directly superimposing it on the image (usually in the upper left or upper right corner of the image), or stored together with the image as metadata. The violation type label can clearly mark the type of violation behavior detected in the frame, such as "using mobile phone", "whispering", etc. When multiple violation behaviors are detected in one frame, all detected types will be listed. The label is usually superimposed on the image in text form, close to the corresponding violation area or placed uniformly at a specific position of the image. The violation area bounding box can accurately indicate the location where the violation behavior occurs by drawing a rectangular box on the image. The bounding box is usually drawn in eye-catching colors (such as red, yellow), and the line width and style can be adjusted as needed to improve visibility. For different types of violation behaviors, different colored bounding boxes can be used for distinction. The coordinate information (x, y coordinates of the upper left and lower right corners) of the bounding box will also be saved as metadata for subsequent analysis and processing.
[0293] In addition, other auxiliary information can be added to the key frames, such as the confidence score of the violation behavior, the examination room number, the examinee information (such as seat number), etc., to enhance the integrity and traceability of the evidence.
[0294] Finally, sort the key video frames in chronological order, generate a violation evidence package, store the violation evidence package in the security database of the central control device, and construct a violation index for the security database. Specifically, first, the chronological sorting of the key video frames is the basis for constructing a coherent evidence chain. Although the timestamp of each frame has been recorded during the extraction process, due to factors such as parallel processing or network transmission delays, the received frames may be out of order. Therefore, the system will perform strict chronological sorting based on the timestamp information of each frame to ensure the temporal continuity of the evidence. The sorting algorithm usually adopts a preset sorting method to ensure that frames with the same timestamp maintain their original relative order. The sorted frame sequence can clearly show the complete development process of the violation behavior, from the state before the behavior starts, to the occurrence, duration, and end of the violation behavior, forming a complete timeline evidence chain.
[0295] Next, organize the sorted key video frames into a structured violation evidence package. The evidence package is a logically organized data set that contains all the evidence materials and metadata related to a specific violation event. A typical violation evidence package includes the following components: 1. Evidence package metadata: including basic information such as a unique identifier (UUID), violation type, violation time range (start time and end time), violation location (examination room number), information about the involved persons (such as examinee ID or seat number), evidence generation time, evidence package version number, etc. 2. Set of key video frames: all the annotated key video frames arranged in chronological order, with each frame containing information such as timestamp, violation type label, and violation area bounding box. 3. Violation behavior summary: a brief description of the violation behavior, including information such as behavior type, duration, severity assessment, etc., to facilitate a quick understanding of the violation situation. 4. Technical verification information: including digital signatures, hash values, etc., used to verify the integrity and authenticity of the evidence and prevent the evidence from being tampered with. 5. Access control information: defines which roles or users have access to the evidence package and the operations that can be performed at different permission levels (such as viewing, exporting, deleting, etc.).
[0296] The evidence package is usually organized in a standardized data format (such as JSON, XML, or a proprietary binary format) and may be compressed and encrypted to reduce storage space requirements and protect sensitive information.
[0297] The generated evidence packages of violations are securely stored in the security database of the central control device. The security database is a database system specifically designed for storing sensitive data and has advanced security features such as access control, data encryption, audit logs, etc. During the storage process, the following operations are performed: 1. Data verification: Check the integrity and format correctness of the evidence package to ensure that all necessary fields are filled and conform to the expected format. 2. Data encryption: Encrypt the evidence package, usually using a high-strength encryption algorithm such as AES-256, to ensure that even if the database is accessed without authorization, the evidence content will not be leaked. 3. Access control settings: Set the corresponding database access permissions according to the permission control information in the evidence package to ensure that only authorized users can access specific evidence packages. 4. Storage confirmation: After completing the storage operation, generate storage confirmation information, including storage time, storage location, storage status, etc., and record it in the system log.
[0298] Finally, build an index of violations in the security database, which is a key data structure for improving the efficiency of evidence retrieval and query. The index of violations can be set based on the time dimension, that is, an index is established based on the occurrence time of violations, supporting queries by time range (such as finding all violations within a specific date or time period).
[0299] The implementation of the index usually uses efficient data structures such as B+ trees, hash tables, or inverted indexes to ensure fast queries even on large-scale data sets.
[0300] This embodiment uses a pre-built YOLOv8 deep learning model to perform real-time analysis on video frames, which can accurately identify various predefined violations, such as whispering to each other, using mobile phones, etc. The marked key frames are organized into evidence packages of violations in chronological order, stored in the security database, and a multi-dimensional index is built to ensure the integrity, security, and retrievability of the evidence. Thus, it significantly improves the efficiency and accuracy of the examination room monitoring, reduces the workload of invigilators; provides objective and detailed evidence of violations, reduces the possibility of disputes and appeals; through data analysis, reveals the patterns and trends of violations, and provides a basis for optimizing examination management.
[0301] In one implementation of this embodiment, the inspection device transmits the examination room images to the central control device, including the following steps:
[0302] S910. The inspection device performs H.264 encoding and compression on the examination room images, establishes an RTMP streaming channel, and pushes the encoded data stream to the central control device;
[0303] S920. Encrypt the encoded data stream with AES-256 to obtain an encrypted data stream;
[0304] S930. When the inspection device detects a network interruption, it stores the encrypted data stream in the local cache and, after detecting network recovery, transmits the encrypted data stream to the central control device.
[0305] During the process of H.264 encoding and compressing the examination room images, the inspection device first collects the original image data of the examination room. These data are usually in uncompressed YUV or RGB format and have a large data volume. To improve the transmission efficiency, the H.264 encoding algorithm is used to compress the images. The H.264 encoding process includes macroblock partitioning, intra-frame prediction, inter-frame prediction, transform quantization, and entropy coding, etc. Specifically, the image is first divided into macroblocks of 16×16 pixels, and then each macroblock is subjected to predictive coding, including I-frames (intra-frame prediction), P-frames (forward prediction), and B-frames (bidirectional prediction). Intra-frame prediction uses the already encoded adjacent pixels in the current frame for prediction, while inter-frame prediction uses the similar regions in the reference frames for motion estimation and compensation. The predicted residual data undergoes discrete cosine transform (DCT) and quantization processing to further reduce data redundancy.
[0306] Finally, entropy coding methods such as CAVLC or CABAC are used to perform lossless compression on the quantization coefficients. H.264 encoding can reduce the data volume to 1 / 50 to 1 / 100 of the original data while maintaining relatively high image quality. After encoding is completed, the inspection device establishes an RTMP (Real-Time Messaging Protocol) push stream channel. This protocol is based on TCP and can provide stable streaming media transmission.
[0307] The RTMP channel establishment process includes three stages: handshake, connection, and push stream. Through this channel, the encoded data stream is pushed to the central control device to achieve the transmission of real-time monitoring images.
[0308] Encrypt the encoded data stream to ensure the security of the examination room monitoring data. Specifically, the AES-256 encryption algorithm is used to encrypt the data stream after H.264 encoding. AES (Advanced Encryption Standard) is a symmetric encryption algorithm, and its 256-bit key length provides extremely high security. The encryption process first requires generating a 256-bit random key, which can be shared between the patrol device and the central control device through a secure key exchange protocol. During encryption, the data stream is divided into 128-bit (16-byte) data blocks, and each data block undergoes four transformation operations: SubBytes (byte substitution), ShiftRows (row shift), MixColumns (column mixing), and AddRoundKey (round key addition), for a total of 14 rounds of transformation. The encryption mode adopts the CBC (Cipher Block Chaining) or GCM (Galois / Counter Mode) mode to enhance the encryption strength and prevent replay attacks. Even if the encrypted data is intercepted, it cannot be decrypted without the correct key, thus ensuring the confidentiality and integrity of the examination room monitoring data during transmission. The encrypted data stream maintains the same format structure as the original encoded data, but the content has been randomized to ensure that only the authorized central control device can correctly decrypt and restore the monitoring screen.
[0309] The network interruption handling mechanism is a key technology to ensure the integrity of the monitoring data. When the patrol device detects a network interruption, it will start the local caching mechanism to ensure that the monitoring data will not be lost due to network problems. The network interruption detection adopts multiple mechanisms, including TCP connection status monitoring, ICMP probe packet (ping) detection, and heartbeat packet mechanism. When the detection fails continuously for multiple times (usually 3 - 5 times) or no response is received after the heartbeat packet times out, a network interruption is determined. After the network interruption, the patrol device redirects the encrypted data stream to the local storage device, and a circular buffer structure can be used for storage. This structure can maximize the preservation of the latest monitoring data within a limited storage space. During the storage process, the timestamp and sequence number of each data packet are recorded for subsequent retransmission in chronological order. The local storage uses high-speed flash memory or SSD to ensure that the write speed meets the storage requirements of real-time monitoring data. At the same time, a data integrity verification mechanism, such as CRC32 or MD5 checksum, is implemented to prevent data corruption during storage. When the network resumes, after confirming the stable connection through the aforementioned network detection mechanism, the data retransmission process will be started. The retransmission adopts a priority strategy, first transmitting real-time data, and then transmitting cached data in chronological order to ensure the continuity of monitoring. Flow control is implemented during the retransmission process to avoid network congestion or interruption again due to sudden large traffic. Through this mechanism, even in an environment with unstable network, the complete recording and transmission of the examination room monitoring data can be ensured.
[0310] Under the premise of ensuring image quality, this embodiment significantly reduces the bandwidth requirement. The encryption mechanism ensures that the monitoring data is not accessed or tampered with without authorization during transmission, guaranteeing the security and fairness of the examination process. The network interruption handling mechanism solves the problem of network instability in practical applications, ensuring the continuity and integrity of the monitoring data, and preventing the loss of key monitoring images due to network fluctuations. This makes the examination hall monitoring system efficient, secure, and reliable, capable of meeting the strict monitoring requirements of various examination scenarios and providing a solid technical guarantee for examination fairness.
[0311] As Figure 2 shown, the embodiment of the present application also provides an inspection system, including:
[0312] A central control device for executing the intelligent inspection method for the physical examination hall of special operations as described above;
[0313] An inspection device wirelessly or wiredly connected to the central control device.
[0314] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0315] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0316] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0317] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0318] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0319] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0320] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0321] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0322] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An intelligent inspection method for a physical examination room for special operations, characterized in that, Applied to a central control device, the method includes: Obtaining examination room data and environmental parameters in real time, and establishing a real - scene navigation map based on the examination room data, where the real - scene navigation map is used to be linked with the examination room data in real time; Obtaining examination room images in real time through a patrol device, where the patrol device is wirelessly or wiredly connected to the central control device, and the patrol device is equipped with a camera for collecting examination room images; Extracting the image features of the examination room images, and dividing the patrol path into n regions based on the image features, where n is a positive integer; Based on the real - scene navigation map and the image features, implementing a dynamic path planning strategy to generate the patrol path of the patrol device; Fusing and analyzing the environmental parameters and the image features to determine the congestion index on the patrol path; Based on the congestion index, determining dangerous regions among the n regions and visualizing the dangerous regions.
2. The method according to claim 1, characterized in that, The examination room data includes three - dimensional space data. Establishing a real - scene navigation map based on the examination room data includes: Converting the three - dimensional space data into a grid map; Performing semantic segmentation on the three - dimensional space data using a pre - constructed neural network model to obtain a semantic segmentation result; Fusing the semantic segmentation result with the grid map to obtain a real - scene navigation map including semantic information.
3. The method according to claim 1, wherein The dynamic path planning strategy includes: Obtaining historical patrol data; Calculating the global optimal path based on an improved A* algorithm, where the improved A* algorithm is improved based on a heuristic function; Real - time monitoring of the crowd density and obstacle distribution on the patrol path; Determining whether there are path - blocked regions and / or dangerous regions in the global optimal path based on the crowd density and obstacle distribution; In the case where there are path - blocked regions and / or dangerous regions in the global optimal path, implementing a local path re - planning strategy and generating a final path; Among them, the local path re - planning strategy includes: Identifying the boundary coordinates of the path - blocked regions and / or dangerous regions; In the real - scene navigation map, assigning a high passage cost value to the grid nodes of the path - blocked regions and / or dangerous regions; Taking the current position as the starting point and the target patrol point as the ending point, applying the D*Lite algorithm to calculate the local optimal path, where the target patrol point is the grid node closest to the grid nodes of the path - blocked regions and / or dangerous regions outside the boundary coordinates; When detecting multiple consecutive path - blocked regions and / or dangerous regions, adopting a segmented planning strategy to decompose the path into multiple sub - paths; Applying the ant colony optimization algorithm to each sub - path to generate local optimal sub - paths; Connecting all the local optimal sub - paths to obtain a local re - planned path.
4. The method according to claim 3, wherein Based on the real - scene navigation map and the image features, implementing a dynamic path planning strategy to generate the patrol path of the patrol device, including: Dividing the real - scene navigation map into multiple grid nodes and calculating the passage cost value of each grid node; Responding to the user's setting instruction for the real - scene navigation map, determining the patrol points in the real - scene navigation map and the weight coefficients corresponding to each patrol point; Calculating the initial global path using the improved A* algorithm; Taking the real - time obtained environmental parameters and image features as input parameters of a pre - constructed multi - feature fusion model to dynamically update the passage cost value of each grid node; When the change amount of the passage cost value exceeds a preset threshold, a preset path re-planning algorithm is adopted to generate a new inspection path; When the change amount of the passage cost value does not exceed the preset threshold, the initial global path is used as the inspection path of the inspection device.
5. The method according to claim 1, wherein Based on the image features, the inspection path is divided into n regions, including: Based on the K-means clustering algorithm, the image features are clustered to obtain at least one clustering region; Calculate the feature vectors of each clustering region and calculate the similarity between the feature vectors; According to the similarity between the feature vectors, the inspection path is divided into n continuous regions, where each region is assigned a unique identifier.
6. The method according to claim 1, characterized in that, Fuse and analyze the environmental parameters and image features to determine the congestion index on the inspection path, including: Extract the personnel density, equipment distribution and operation behavior from the image features through the YOLOv8 model; Output the congestion index through a multi-layer perceptron neural network model, where the input layer of the multi-layer perceptron neural network model includes environmental parameters, personnel density, equipment distribution and operation behavior, and the output layer is the congestion index.
7. The method according to claim 6, wherein Based on the congestion index, determine the dangerous regions among the n regions, including: When the congestion index of a region is lower than the preset first congestion index threshold T1, the region is marked as a safe region; When the congestion index of a region is between the first congestion index threshold T1 and the second congestion index threshold T2, the region is marked as a warning region; When the congestion index of a region is higher than the second congestion index threshold T2, the region is marked as a dangerous region, where 0 < T1 < T2 < 100.
8. The method according to claim 1, wherein The camera is also used to collect the examination room video, and the examination room video includes multiple video frames. The method further includes: Input the video frames into a pre-constructed YOLOv8 model to identify the predefined types of violation behaviors; For each type of violation behavior, record the start time and end time of the violation behavior; Within a preset time interval, extract key video frames at a frequency of m frames per second, and add a timestamp, a violation type label and a violation region bounding box to each key video frame; Sort the key video frames in chronological order, generate an evidence package of violation behaviors, store the evidence package of violation behaviors in the security database of the central control device, and construct an index of violation behaviors in the security database.
9. The method according to claim 1, wherein The inspection device transmits the examination room image to the central control device, including: The inspection device performs H.264 encoding and compression on the examination room image, establishes an RTMP push stream channel, and pushes the encoded data stream to the central control device; Perform AES-256 encryption on the encoded data stream to obtain an encrypted data stream; When the inspection device detects a network interruption, store the encrypted data stream in the local cache, and when the network recovery is detected, transmit the encrypted data stream to the central control device.
10. An inspection system, characterized in that, Including: A central control device for executing the intelligent inspection method for a special operation physical examination room according to any one of claims 1 to 9; An inspection device, wirelessly or wiredly connected to the central control device.
Citation Information
Patent Citations
Illegal operation analysis method based on inspection process
CN116579609A
Campus inspection robot navigation method based on large model fusion environment and biological multi-modal information
CN118329044A
Path planning method and device for 5G nursing robot
CN119336038A
Public place security inspection management system based on Internet of Things
CN119723702A
System, method, and device to proactively detect in real time one or more threats in crowded areas
US12243306B1
Cited By
Intelligent inspection method and system based on low-altitude economy
CN120523226A
Mobile robot control device giving consideration to field navigation patrol and transportation load
CN120821235A
High-precision autonomous navigation robot and autonomous navigation method thereof
CN121274997A
A high-precision autonomous navigation robot and an autonomous navigation method thereof
CN121274997B
Fire-fighting equipment intelligent inspection and state monitoring method and electronic equipment
CN121436903A