Adaptive Quantization Inference Method for Dynamic Accuracy of Vehicle Perception
By constructing edge-side adaptive computing power constraint boundary parameters and dynamic precision adaptive quantization matching strategies, the problem of insufficient energy and precision caused by fixed precision in the perception scheme of inspection robots is solved. This achieves high energy efficiency and high precision adaptation of the perception model under different conditions, and extends the operation time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DECK SMART TECH CO LTD
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-31
AI Technical Summary
In existing inspection robot perception solutions, fixed-precision neural network models cannot dynamically adjust energy consumption and perception accuracy based on real-time battery life and environmental complexity. This results in the system maintaining high energy efficiency and shortening operation time when power is limited, and it is difficult to provide sufficient recognition sensitivity in complex environments.
By constructing edge-side adaptive computing power constraint boundary parameters based on endurance status and inspection field of view activity distribution characteristics, a dynamic accuracy adaptive layer-by-layer quantization matching strategy is established to generate a vehicle perception dynamic accuracy adaptive quantization inference mapping map, thereby realizing real-time adjustment of the perception model's accuracy and computing power resource allocation.
It achieves dynamic adaptation of the perception model to high efficiency under different operating conditions, extends the robot's working time, and maintains high-precision perception quality in complex environments, meeting the needs of intelligent inspection in modern parks.
Smart Images

Figure CN122491480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and autonomous driving technology, and more specifically, to an adaptive quantization reasoning method for vehicle-mounted perception dynamic accuracy. Background Technology
[0002] With the rapid development of autonomous driving and artificial intelligence technologies, inspection robots have been widely used in areas such as park security and security patrols. Inspection robots are typically equipped with multiple sensors and use edge-based perception models to monitor and identify crowd gatherings, uncivilized behavior, and safety hazards in the environment in real time, thereby achieving unmanned patrol and early warning in security scenarios.
[0003] In existing inspection robot perception solutions, fixed-precision neural network models are typically used for inference. This solution first trains and quantizes a fixed-width perception model offline and deploys it in the robot's edge computing module. Then, the robot captures real-time video streams through a camera and directly inputs them into the fixed-width model. Finally, the model outputs fixed target recognition results and classification information.
[0004] However, this fixed-precision perception scheme has significant technical drawbacks. During actual operation, the remaining battery power of the inspection robot and the complexity of the external environment (such as pedestrian density) dynamically change. The fixed-precision inference model cannot adjust energy consumption based on real-time battery status, resulting in the system maintaining high energy efficiency even when battery power is limited, thus shortening the robot's operating time. Furthermore, when environmental interference is significant or the scene becomes complex, the fixed low-bit-width quantization model struggles to provide sufficient recognition sensitivity and lacks a mechanism to adaptively adjust the perception depth based on environmental characteristics. This leads to low adaptability between perception accuracy and computing resources, failing to meet the needs of adaptive inspection in complex environments. Summary of the Invention
[0005] This application provides an adaptive quantization inference method for vehicle-mounted perception dynamic accuracy to at least alleviate the above-mentioned technical problems.
[0006] An adaptive quantization inference method for vehicle-mounted perception dynamic accuracy includes: Step 1: Based on the current endurance status representation parameters of the inspection robot and the activity distribution characteristics of the inspection field of view in the corresponding public scene, construct adaptive computing power constraint boundary parameters for the inspection robot in its current operating state at the edge. Step 2: Based on the pre-built basic multi-source perception diagnostic mapping map for the inspection scenario and the edge-side adaptive computing power constraint boundary parameters, establish a dynamic accuracy adaptive layer-by-layer quantization matching strategy for real-time behavior detection. Step 3: Based on the dynamic precision adaptive layer-by-layer quantization matching strategy, generate the vehicle perception dynamic precision adaptive quantization inference mapping map. Step 4: Obtain the real-time visual monitoring sequence of the inspection robot, and determine the intelligent safety early warning decision level for the current scenario by mapping the real-time visual monitoring sequence to the vehicle perception dynamic accuracy adaptive quantization inference mapping map.
[0007] Optionally, step 1 includes: Obtain the real-time topological profile of the LiDAR environment from the LiDAR sensors deployed on the chassis of the inspection robot; The inspection robot acquires a public scene inspection video payload synchronously captured by the vision sensor mounted on its chassis. Align the lidar environment topology with the public scene inspection video payload to construct a multi-source sensing synchronous data array; Based on the multi-source sensing synchronous data array, the correlation features between dynamic target aggregation and abnormal activity trajectories in public scenes are determined, thereby generating the activity distribution features of the inspection field of view.
[0008] Optionally, step 1 further includes: Obtain the current battery life status parameters fed back by the battery monitoring system built into the chassis of the inspection robot; By integrating the current endurance status representation parameters with the inspection field activity distribution characteristics, an edge-side adaptive computing power constraint boundary parameter is constructed for the current operating state.
[0009] Optionally, step 2 includes: Using the edge-side adaptive computing power constraint boundary parameters as the basis for inference accuracy adjustment, the association of computational mapping nodes inside the basic multi-source perception diagnostic mapping map is decoupled to determine the layer-level mutation candidate mapping level of the basic multi-source perception diagnostic mapping map under the current computing power constraint; Based on the layer-level mutation candidate mapping hierarchy, a dynamic accuracy-adaptive layer-by-layer quantization matching strategy for real-time behavior detection is established.
[0010] Optionally, based on the layer-level mutation candidate mapping hierarchy, a dynamic accuracy-adaptive layer-by-layer quantization matching strategy for real-time behavior detection is established, including: The quantization bit width response mapping range of the layer-level mutation candidate mapping level is detected, and then the quantization accuracy compression mapping node parameters of each layer-level mutation candidate mapping level under the baseline operating conditions are determined. The quantization precision compression mapping node parameters are matched with the edge-side adaptive computing power constraint boundary parameters to establish a dynamic precision adaptive layer-by-layer quantization matching strategy for real-time behavior detection.
[0011] Optionally, step 3: Based on the dynamic precision adaptive layer-by-layer quantization matching strategy, generate an on-board perception dynamic precision adaptive quantization inference mapping map, including: The configuration parameters of the dynamic precision adaptive layer-by-layer quantization matching strategy are analyzed to obtain the quantization floating-point conversion step size calibration parameter; Extract the underlying structure sequence of the basic multi-source sensing diagnostic mapping map to obtain the initial weight parameter distribution sequence; The quantization floating-point conversion step size calibration parameter is injected into the initial weight parameter distribution sequence to generate a step size fusion weight distribution sequence; Determine the activation response threshold of the step-size fusion weight distribution sequence to generate a dynamic precision-trimmed weight distribution sequence; Based on the dynamic precision pruning weight distribution sequence, an adaptive quantization inference mapping map for vehicle perception dynamic precision is generated.
[0012] Optionally, based on the dynamic precision pruning weight distribution sequence, an adaptive quantization inference mapping map for vehicle perception dynamic precision is generated, including: Extract the signal transmission mechanism data within the basic multi-source sensing diagnostic mapping map to obtain the feature transmission information payload; The activation distribution of the feature transfer information payload is adjusted using the quantization floating-point conversion step size calibration parameter to generate a quantization scaling adapted activation feature payload. The network computation graph is reconstructed based on the dynamic precision pruning weight distribution sequence and the quantization scaling adaptation activation feature load to generate an on-board perception dynamic precision adaptive quantization inference mapping graph.
[0013] Optionally, step 4 includes: During the cycle in which the inspection robot chassis operates according to the preset inspection path rules, the visual stream data continuously collected by the vision sensor is converted to obtain a real-time visual monitoring sequence.
[0014] Optionally, step 4 includes: The real-time visual monitoring sequence is mapped into the vehicle perception dynamic accuracy adaptive quantization inference mapping map to generate the visual monitoring sequence to be inferred. Based on the preset calculation rules of the vehicle perception dynamic accuracy adaptive quantization inference map, the adaptive bit width forward diagnostic result of the visual monitoring sequence to be inferred is inferred, so as to output the abnormal behavior diagnostic result for the current scene. Assess the potential hazard level of the abnormal behavior diagnosis results, and then determine the intelligent safety early warning decision level for the current scenario.
[0015] Optionally, step 4 further includes: triggering the inspection intervention scheduling strategy based on the intelligent security early warning decision level to generate an inspection anomaly intervention scheduling notification; and transmitting the inspection anomaly intervention scheduling notification to a preset remote control center.
[0016] Technical advantages of the technical solution provided in this application This application presents a vehicle-mounted perception dynamic accuracy adaptive quantization reasoning method. Addressing the technical shortcomings of traditional perception schemes, such as the inability to adjust computing power consumption based on real-time conditions and low adaptability between perception accuracy and resources, this method constructs adaptive computing power constraint boundary parameters on the edge side based on parameters representing the current battery endurance and the activity distribution characteristics of the inspection field of view. This solves the problem of the disconnect between perception logic and physical operating conditions (such as remaining energy and environmental pressure) in traditional schemes. Compared to traditional fixed-precision single reasoning methods, this application integrates battery endurance data and field-of-view activity characteristics, enabling edge-side computing power resources to collaboratively model based on the robot's real-time "physical strength" and the "pressure" of scene recognition. This achieves a scientific setting of computing power constraint boundaries and provides accurate physical reference for subsequent dynamic accuracy adjustment.
[0017] An adaptive layer-by-layer quantization matching strategy is established based on computational power constraint boundary parameters and a basic multi-source sensing diagnostic mapping map. This strategy solves the problems of fixed quantization accuracy and difficulty in flexibly adapting to environmental changes in traditional schemes. Traditional schemes can only operate at a single precision, while this application decouples the node associations within the diagnostic mapping map and probes the response range of the candidate mapping layers for layer-level mutations. This allows the model to find a better bit width combination based on the current constraint boundary. Compared with the traditional fixed bit width scheme, this strategy realizes the transformation from static bit width to dynamic layer-by-layer matching, enabling the sensing model to achieve a higher performance ratio under different operating conditions.
[0018] This application addresses the lack of dynamic precision adjustment mechanisms in traditional schemes by generating an adaptive quantization inference mapping graph through parsing the matching strategy and reconstructing the network computation graph. By injecting the quantization floating-point conversion step size calibration parameter into the parameter sequence and reconstructing the computation graph, online reconfiguration based on the aforementioned generated strategy is possible. Compared to traditional schemes that require model redeployment, this application achieves online, real-time model precision adjustment through scaling and adapting the feature transfer payload and pruning the weight distribution, thus improving the dynamic adaptability of the perception model to complex inspection environments.
[0019] Finally, by mapping real-time monitoring sequences to inference maps to determine early warning decision levels and trigger scheduling, a closed loop of "state perception - policy generation - model reconstruction - adaptive inference" is formed. Compared with traditional solutions, this application, while achieving high recognition sensitivity, can dynamically reduce or increase computing power overhead as battery life fluctuates and scene complexity changes. This allows the robot to extend its working time by appropriately reducing redundant accuracy when power is limited, while ensuring perception quality by increasing accuracy in complex core areas. This provides stronger support for inspection tasks and better meets the long-term stable operation requirements of intelligent inspection in modern parks. Attached Figure Description
[0020] Figure 1 This is a flowchart of an adaptive quantization reasoning method for vehicle-mounted perception dynamic accuracy according to an embodiment of this application.
[0021] Figure 2 This is a schematic diagram of a vehicle-mounted perception dynamic accuracy adaptive quantization inference device based on edge computing power constraints, according to an embodiment of this application.
[0022] Figure 3 This is a schematic diagram of an electronic device structure according to an embodiment of this application. Detailed Implementation
[0023] like Figure 1 As shown, this embodiment of the present application provides a vehicle-mounted perception dynamic accuracy adaptive quantization inference method, which includes: Step 1: Based on the current endurance status representation parameters of the inspection robot and the activity distribution characteristics of the inspection field of view in the corresponding public scene, construct adaptive computing power constraint boundary parameters for the inspection robot in its current operating state at the edge. Step 2: Based on the pre-built basic multi-source perception diagnostic mapping map for the inspection scenario and the edge-side adaptive computing power constraint boundary parameters, establish a dynamic accuracy adaptive layer-by-layer quantization matching strategy for real-time behavior detection. Step 3: Based on the dynamic precision adaptive layer-by-layer quantization matching strategy, generate the vehicle perception dynamic precision adaptive quantization inference mapping map. Step 4: Obtain the real-time visual monitoring sequence of the inspection robot, and determine the intelligent safety early warning decision level for the current scenario by mapping the real-time visual monitoring sequence to the vehicle perception dynamic accuracy adaptive quantization inference mapping map.
[0024] Optionally, step 1 includes: Obtain the real-time topological profile of the LiDAR environment from the LiDAR sensors deployed on the chassis of the inspection robot; The inspection robot acquires a public scene inspection video payload synchronously captured by the vision sensor mounted on its chassis. Align the lidar environment topology with the public scene inspection video payload to construct a multi-source sensing synchronous data array; Based on the multi-source sensing synchronous data array, the correlation features between dynamic target aggregation and abnormal activity trajectories in public scenes are determined, thereby generating the activity distribution features of the inspection field of view.
[0025] Preferably, the specific implementation process of obtaining the real-time LiDAR environmental topology contour collected by the LiDAR sensor deployed on the chassis of the inspection robot in step 1 is as follows: During the inspection robot's movement along a preset inspection path, the LiDAR sensor deployed on its chassis in the forward or circumferential direction emits laser pulse beams at a fixed scanning frequency (e.g., 10 Hz to 20 Hz) to the surrounding public scene and receives echo pulse signals reflected back from the surfaces of environmental objects. The signal processing circuit inside the LiDAR sensor calculates the three-dimensional spatial coordinates of each laser footpoint in a local spherical or Cartesian coordinate system with the sensor itself as the origin, based on the time difference of flight and the emission angle information of the laser pulse beam, thereby forming a raw point cloud frame composed of a massive number of discrete spatial points. Each data point in this raw point cloud frame not only carries the distance information of the scene object surface relative to the LiDAR sensor, but also implicitly contains the geometric distribution characteristics of the object surface in the horizontal orientation and vertical height dimensions. Subsequently, inter-frame motion compensation and noise filtering were performed on multiple consecutively acquired raw point cloud frames to eliminate the offset of point cloud data in the spatial reference frame caused by the movement of the inspection robot's chassis at different times, as well as outlier noise introduced by environmental interference factors such as dust and flying insects. The compensated and filtered point cloud data was further subjected to ground point cloud segmentation and non-ground point cloud clustering operations to identify and extract the contour information representing permanent or semi-permanent structures within the scene, such as building walls, road edges, fixed public facility supports, and the ground. The extracted contour information of these structures collectively constitutes a lidar environmental topology contour corresponding to the current moment, reflecting the macroscopic geometric layout of the environment surrounding the inspection robot. This lidar environmental topology contour is essentially a sparse three-dimensional geometric structure descriptor, which clearly defines the boundary of the traversable area where the inspection robot is located at the current moment and the spatial occupancy of the main static obstacles in the surrounding area.
[0026] Preferably, the specific implementation process of obtaining the public scene inspection video payload synchronously captured by the vision sensor mounted on the inspection robot chassis in step 1 is as follows: A vision sensor, such as a color wide-angle camera or depth camera, synchronously controlled by the same time reference source as the LiDAR sensor, continuously acquires optical images of the same public scene while the LiDAR sensor acquires each frame of the LiDAR environment topology contour, generating a video data stream strictly aligned with the LiDAR point cloud data in timestamps. Each frame of image data in this video data stream is called a frame of the public scene inspection video payload, which contains the red, green, and blue three-channel pixel color information or depth value information of the public scene in the visible light band. The internal image signal processor of the vision sensor performs a series of preprocessing operations on the raw electrical signal output by the photosensitive element, including automatic white balance correction, automatic exposure control, gamma correction, and lens distortion correction, to ensure that each frame of the output public scene inspection video payload has stable brightness, contrast, and color reproduction, and eliminates edge barrel distortion or pincushion distortion caused by the wide-angle lens. The preprocessed public scene inspection video payload is cached in the local memory buffer of the inspection robot in the form of a time-series image frame sequence, awaiting subsequent data-level fusion with the synchronously acquired LiDAR environmental topology contour. The dense pixel-level texture information and color semantic information carried by each frame of the public scene inspection video payload precisely compensate for the lack of surface appearance details in the LiDAR environmental topology contour, which only contains sparse geometric structures. This provides a rich information foundation for the subsequent accurate identification of the type, quantity, and accurate pixel-level spatial location of dynamic targets in the scene.
[0027] Preferably, the specific implementation process of aligning the LiDAR environmental topology contour and the public scene inspection video payload in step 1 to construct a multi-source sensing synchronous data array is as follows. Based on the frame pairs of the LiDAR environmental topology contour and the public scene inspection video payload that have already completed timestamp synchronization, the spatial coordinate system alignment operation needs to be performed first. This spatial alignment process utilizes a pre-calibrated extrinsic parameter calibration matrix between the LiDAR sensor and the vision sensor. This extrinsic parameter calibration matrix describes the relative installation position and attitude relationship of the two sensors on the inspection robot chassis, specifically including a 3x3 rotation matrix and a 3x1 translation vector. By performing matrix multiplication and vector addition operations on each three-dimensional spatial point constituting the LiDAR environmental topology contour, the coordinates of the point are transformed from its original LiDAR local coordinate system to the camera coordinate system of the vision sensor. Subsequently, using the known intrinsic parameter matrix of the vision sensor, i.e., the projection matrix containing parameters such as focal length and principal point coordinates, the three-dimensional point in the camera coordinate system is further projected onto the two-dimensional image plane to obtain the pixel coordinate position that accurately corresponds to the LiDAR point. By traversing all valid 3D points in the lidar environment topology contour and performing the aforementioned coordinate transformation and projection operations, the geometrically corresponding pixel region can be determined on each frame of the public scene inspection video payload. The lidar environment topology contour frame and the public scene inspection video payload frame, aligned with their timestamps, along with the point-to-point or point-to-pixel region mapping index established through the aforementioned spatial alignment operations, are uniformly encapsulated into a data structure. This data structure is a basic building block of the multi-source sensing synchronous data array. Multiple consecutive frames of this data structure, arranged in temporal order, constitute the complete multi-source sensing synchronous data array. Each unit in this multi-source sensing synchronous data array contains a fused representation of geometric structure information and visual texture information within the same scene at the same time, providing a unified and standardized data foundation for subsequent data-driven mining of the activity patterns of dynamic targets.
[0028] Preferably, the specific implementation process of determining the association features of dynamic target aggregation and abnormal activity trajectories in public scenes based on the multi-source sensing synchronous data array in step 1 is as follows. After obtaining the multi-source sensing synchronous data array, this application first performs dynamic target detection and tracking processing on each frame of the public scene inspection video payload in the array. This processing uses a visual target detector based on a deep convolutional neural network to infer the image frames. The visual target detector contains a feature extraction backbone network for extracting multi-scale visual features and a detection head network for predicting target bounding boxes and categories. Through this visual target detector, multiple dynamic targets, such as pedestrians, moving vehicles, or other robots, can be identified from each frame of the public scene inspection video payload, and the bounding box coordinates and semantic category labels of each dynamic target can be obtained. At the same time, a multi-target tracker based on Kalman filtering and Hungarian algorithm matching is used to associate and maintain the identity of the same dynamic target appearing in multiple consecutive frames of the public scene inspection video payload, thereby generating the motion trajectory of each dynamic target in the pixel coordinate system in a continuous time series. Subsequently, this application maps the pixel region corresponding to each dynamic target within its bounding box back to the lidar environmental topology contour part of the multi-source sensing synchronous data array at the same time stamp, so as to index out the spatially corresponding 3D point cloud subset of the dynamic target from this part. By performing geometric center calculation or bounding box fitting on the 3D point cloud subset, the true position coordinates and approximate volume scale information of the dynamic target in 3D physical space can be obtained. Based on the position coordinates of all dynamic targets in 3D space and their motion trajectory evolving over time, this application further calculates the density distribution heatmap characterizing the spatial distribution aggregation degree of dynamic targets, as well as statistics such as trajectory direction entropy and velocity mutation rate characterizing the degree of abnormality of their motion patterns. These density distribution heatmaps and motion pattern statistics together constitute the correlation features between dynamic target aggregation and abnormal activity trajectories. Among them, the dynamic target aggregation feature indicates which areas in the scene have a high density of people or objects, while the abnormal activity trajectory feature indicates which dynamic targets' motion patterns (e.g., sudden acceleration, reverse movement, long-term loitering) deviate from the normal behavior pattern.
[0029] Preferably, the specific implementation process of generating the inspection field of view activity distribution characteristics based on the correlation features between dynamic target aggregation and abnormal activity trajectories in step 1 is as follows. Based on the correlation features between dynamic target aggregation and abnormal activity trajectories calculated in the aforementioned steps, this application performs spatial rasterization and time window statistical aggregation processing to generate an inspection field of view activity distribution characteristic map that can quantitatively describe the complexity of activities and potential perceptual pressure within the current field of view of the inspection robot. Specifically, this application first projects the three-dimensional perception space in front of the inspection robot onto a horizontal ground as a two-dimensional planar raster map, which consists of multiple raster units of a preset size (e.g., 0.5 meters by 0.5 meters). Subsequently, for each raster unit, this application statistically analyzes the frequency of occurrence, average movement speed, and trajectory anomaly score of dynamic targets falling within the raster unit within a selected time window (e.g., within the last five or ten seconds). The frequency of dynamic targets directly reflects the activity density of the grid cell area; the average movement speed reflects the behavior state of dynamic targets in the area—whether they are stationary, moving slowly, or moving rapidly; the trajectory anomaly score is a quantitative indicator that measures the degree to which the movement pattern of dynamic targets in the area deviates from the norm, and it can be calculated from the cumulative value of the angle of sudden change in movement direction or the variance of the rate of change of velocity. The statistics of these three dimensions—frequency of occurrence, average movement speed, and trajectory anomaly score—are weighted and fused to form a scalar value representing the activity pressure level of the grid cell. The scalar values of the activity pressure levels of all grid cells together constitute a two-dimensional numerical matrix corresponding to the current time window. This two-dimensional numerical matrix is the activity distribution characteristic of the inspection field of view. Areas with higher values in the activity distribution characteristic of the inspection field of view represent dense dynamic targets and complex and varied movement patterns in the spatial area, which means higher recognition difficulty and computational power consumption for the inspection robot's perception model; conversely, areas with lower values represent relatively open scenes or simple and regular dynamic target movement patterns, with less perception processing pressure. The distribution characteristics of the inspection field of view activity, as the final output of step 1, provide a quantitative basis for the "perceived pressure" on the scene side for the subsequent construction of edge-side adaptive computing power constraint boundary parameters.
[0030] Preferably, in the overall technical logic of step 1, the technical processing step of acquiring the topological contour of the LiDAR environment and the inspection video payload of the public scene and constructing a multi-source perception synchronous data array lays the foundation for multimodal data fusion for accurately extracting the correlation features of dynamic target aggregation and abnormal activity trajectories from the data source. A series of technical processes, including dynamic target detection, tracking, 3D spatial mapping, and behavioral pattern statistical analysis, are performed on the multi-source perception synchronous data array. The purpose is to transform the raw, unstructured sensor data stream into inspection field-of-view activity distribution features with clear physical meaning and task orientation. The final generated inspection field-of-view activity distribution features, containing spatial distribution information on activity density and motion complexity, accurately characterize the scene-side perception difficulty distribution faced by the inspection robot during task execution. Unlike traditional solutions that only focus on the sensor data itself while ignoring the task scene pressure inherent in the data, this application, through the aforementioned series of technical processing actions, generates inspection field-of-view activity distribution features that provide a quantifiable and computable structured information carrier for the subsequent collaborative modeling of scene "perception pressure" and the robot's own "endurance capacity." It is precisely because the generation process of the inspection field activity distribution characteristics fully integrates the accurate three-dimensional spatial geometric constraints provided by the lidar and the dense semantic dynamic information provided by the visual sensor that the edge-side adaptive computing power constraint boundary parameters established based on this feature can truly and dynamically reflect the real-time computing resource requirements and objective environment perception difficulties faced by the inspection robot in the current specific inspection scenario.
[0031] In the process of reasoning on each frame of a public scene inspection video payload using a visual object detector based on a deep convolutional neural network, the visual object detector first takes the current frame image of the public scene inspection video payload as input data and feeds it into its internal feature extraction backbone network. The feature extraction backbone network consists of a series of cascaded convolutional computation layers, activation function layers, and downsampling layers. Its technical role is to perform multi-level visual feature extraction from low-level edge textures to high-level semantic concepts on the input single-frame image data. Specifically, the first few layers of the feature extraction backbone network are responsible for capturing local details in the image, such as the edge contours, corners, and local color change patterns of objects. As the network layers deepen, the middle and later layers of the feature extraction backbone network gradually expand the receptive field, aggregating these local details into semantic information that can represent a larger image region, such as the shape of object parts, overall contours, and relative spatial relationships with the surrounding environment. Finally, the feature extraction backbone network outputs a set of multi-scale visual feature maps with different spatial resolutions and different levels of semantic abstraction in parallel at its multiple outputs at different depths. In these multi-scale visual feature maps, the high-resolution feature maps output from the shallow layer retain rich spatial location information, which helps to locate small or clearly defined dynamic targets; while the low-resolution feature maps output from the deep layer contain highly abstract semantic category information, which helps to distinguish dynamic targets with similar appearances but belonging to different semantic categories, such as distinguishing pedestrians from poles.
[0032] After the feature extraction backbone network extracts multi-scale visual feature maps, these maps are then passed to the detection head network in parallel or serially. The detection head network's role is to use the multi-scale visual feature maps output by the feature extraction backbone network as processing objects. For each spatial location point or predefined anchor box on each feature map, it performs two parallel sub-task predictions: one is the regression prediction of the target bounding box coordinates, and the other is the classification prediction of the target's semantic category label. The detection head network typically contains several shared convolutional transformation layers and two independent output branch convolutional layers, corresponding to the bounding box regression branch and the category classification branch, respectively. In the bounding box regression branch, the detection head network outputs a set of continuous values for each preset location or anchor box. These continuous values represent the offset of the predicted bounding box relative to the center point coordinates of the preset location or anchor box, as well as the scaling factors for width and height. In the category classification branch, the detection head network outputs a probability distribution vector for each preset location or anchor box. Each dimension of this probability distribution vector corresponds to a predefined dynamic target semantic category (e.g., pedestrian, vehicle, cyclist, etc.), and its value represents the confidence score that the target contained within the predicted bounding box belongs to that category.
[0033] The technical collaboration between the feature extraction backbone network and the detection head network manifests as a feedforward collaborative mechanism of "general feature supply and task-oriented prediction." The feature extraction backbone network does not directly depend on specific detection task objectives; it focuses on extracting general visual representations related to scene understanding from the pixel array of the public scene inspection video payload and provides them to the detection head network in the form of multi-scale visual feature maps. The detection head network, on the other hand, no longer directly processes the raw image pixels but directly utilizes the multi-scale visual feature maps, which are already abstracted by the feature extraction backbone network and rich in discriminative information. This allows it to focus on completing the two specific and well-defined detection sub-tasks of bounding box coordinate regression and category classification with a relatively lightweight network structure. Because the feature extraction backbone network undertakes the computationally intensive and parameter-intensive feature extraction work, the detection head network can quickly generate candidate detection results at each spatial location of the multi-scale visual feature map with relatively low computational cost. Subsequently, the visual object detector performs non-maximum suppression processing on all candidate detection results to filter out redundant bounding boxes repeatedly predicted for the same dynamic object, ultimately outputting the bounding box coordinates and semantic category labels of the multiple dynamic objects contained in each frame of the public scene inspection video payload. By leveraging the collaborative operation of the feature extraction backbone network and the detection head network in the aforementioned feedforward computation process, this application achieves the technical objective of simultaneously identifying multiple dynamic targets from a single image frame of a public scene inspection video payload, providing an accurate target instantiation information basis for the subsequent determination of the correlation features between dynamic target aggregation and abnormal activity trajectories.
[0034] Optionally, step 1 further includes: Obtain the current battery life status parameters fed back by the battery monitoring system built into the chassis of the inspection robot; By integrating the current endurance status representation parameters with the inspection field activity distribution characteristics, an edge-side adaptive computing power constraint boundary parameter is constructed for the current operating state.
[0035] Preferably, the specific implementation process for obtaining the current range status characterization parameter from the power monitoring system built into the inspection robot chassis in step 1 is as follows: The power monitoring system, deployed on the power supply bus inside the inspection robot chassis, typically consists of a coulomb counter chip, a voltage detection circuit, and a temperature sensor. This power monitoring system continuously monitors the real-time output current, terminal voltage, and cell temperature of the energy storage unit (e.g., a lithium battery pack) at a preset sampling period (e.g., once every second or every two seconds). The microcontroller unit inside the power monitoring system performs time integration on the real-time output current and combines it with compensation corrections for the terminal voltage and cell temperature to calculate the remaining usable charge of the energy storage unit at the current moment. This remaining usable charge is typically measured in ampere-hours (AHs) or watt-hours (WHs). Subsequently, the power monitoring system calculates the ratio of the calculated remaining usable charge to the factory-calibrated total usable charge of the energy storage unit at full charge, obtaining a dimensionless percentage value ranging from zero to one hundred percent. This dimensionless percentage value is the current range status characterization parameter. The current battery life parameter directly reflects the potential working time margin that the inspection robot can continue to perform inspection tasks without needing to return to its charging station midway. Unlike traditional solutions that treat the remaining battery power as an isolated alarm threshold trigger condition, this application uses the current battery life parameter as a core input variable for constructing the edge-side adaptive computing power constraint boundary parameter. Its value will directly affect the degree to which the subsequent computing power constraint boundary is tightened or relaxed.
[0036] Preferably, the specific implementation process of fusing the current endurance status representation parameter and the inspection field of view activity distribution feature in step 1 to construct the edge-side adaptive computing power constraint boundary parameter for the current operating state is as follows. After obtaining the current endurance status representation parameter representing the remaining energy margin of the inspection robot and the inspection field of view activity distribution feature representing the scene-side perceived pressure, this application normalizes and co-maps these two types of parameters that differ in physical dimensions and numerical distribution ranges to generate a unified overall constraint index that can guide the edge-side computing power allocation. First, for the inspection field of view activity distribution feature, this feature is originally a two-dimensional numerical matrix, where each element corresponds to the scalar value of the activity pressure level of a spatial grid cell. To facilitate fusion with the current endurance status representation parameter, this application performs a global pooling operation on the two-dimensional numerical matrix corresponding to the inspection field of view activity distribution feature, specifically max pooling or average pooling. If max pooling is used, the highest-valued activity stress level scalar value in the two-dimensional numerical matrix is extracted as the global activity stress index, reflecting the perception difficulty of the most severe area within the current field of view. If average pooling is used, the arithmetic mean of all activity stress level scalar values in the two-dimensional numerical matrix is calculated as the global activity stress index, reflecting the overall perception difficulty within the current field of view. This global activity stress index is a dimensionless scalar; a higher value indicates a dense concentration of dynamic targets and complex behaviors in the current scene, indicating a more urgent need for computing power.
[0037] Preferably, after obtaining the global activity pressure index, this application further constructs a two-dimensional computing power constraint mapping function. This computing power constraint mapping function uses the current endurance status representation parameter and the global activity pressure index as two independent input variables, and outputs an edge-side adaptive computing power constraint boundary parameter that is also normalized to a preset interval (e.g., between zero and one). The specific form of this computing power constraint mapping function can be designed as a weighted fusion function with adjustable weight coefficients, or as a nonlinear mapping relationship implemented through a pre-calibrated two-dimensional lookup table. Taking the weighted fusion function as an example, its operation process is as follows: multiply the current endurance status representation parameter by the first weight coefficient, multiply the global activity pressure index by the second weight coefficient, and then add the two product results. The sum is then clipped to the preset interval by a limiting function, thus forming the edge-side adaptive computing power constraint boundary parameter. The first weighting coefficient is a negative or negatively correlated mapping factor, which makes the contribution of the current battery life representation parameter (i.e., less battery power) to the edge-side adaptive computing power constraint boundary parameter tend to shrink as the current battery life representation parameter is lower (i.e., less battery power), thus guiding the subsequent inference process to reduce computing power overhead. The second weighting coefficient is a positive or positively correlated mapping factor, which makes the contribution of the global activity pressure index (i.e., more complex scenario) to the edge-side adaptive computing power constraint boundary parameter tend to expand as the global activity pressure index is higher, thus guiding the subsequent inference process to maintain the necessary perception accuracy. Through this weighted fusion function, a dynamic balance and collaborative constraint relationship is formed between the current battery life representation parameter and the global activity pressure index.
[0038] Preferably, the magnitude of the edge-side adaptive computing power constraint boundary parameter directly determines the decoupling depth and search range when searching for layer-level mutation candidate mapping levels for the basic multi-source perception diagnostic mapping map in the subsequent step 2. Specifically, when the value of the edge-side adaptive computing power constraint boundary parameter is low (e.g., close to 0.2 or 0.3), it indicates that the inspection robot is currently operating in a low-battery and simple scenario. At this time, the computing power constraint boundary is relatively strict. The subsequent step 2 will tend to decouple deeper-level computational mapping node associations from the basic multi-source perception diagnostic mapping map and select layer-level mutation candidate mapping levels with lower quantization bit widths (e.g., using four-bit or eight-bit integer quantization) to compress the computational load and memory usage of the model as much as possible. Conversely, when the value of the edge-side adaptive computing power constraint boundary parameter is high (e.g., approaching 0.8 or 0.9), it indicates that the inspection robot currently has sufficient power and is in a core inspection area with dense crowds and complex behaviors. In this case, the computing power constraint boundary is relatively loose, and subsequent step 2 will tend to maintain a shallow decoupling depth and select layer-level mutation candidate mapping levels with higher quantization bit widths (e.g., using 16-bit or 32-bit floating-point) to fully utilize the perception capabilities of edge-side computing resources. In this way, the edge-side adaptive computing power constraint boundary parameter provides a unified quantization constraint benchmark with clear physical meaning for the generation of subsequent dynamic precision adjustment strategies.
[0039] Preferably, unlike traditional solutions that treat power monitoring and perception tasks as two separate subsystems, this application utilizes a technique to construct edge-side adaptive computing power constraint boundary parameters by fusing current battery life representation parameters with the activity distribution characteristics of the inspection field of view. This achieves deep coupling between the robot's physical energy state and the objective perception requirements of the scene at the computing power decision-making level. In traditional solutions, the perception model runs continuously with fixed precision. When the robot's battery is insufficient, it can only passively respond by abruptly cutting off power or reducing frequency at the system level, which can easily lead to missed or false detections in critical inspection areas due to a sudden drop in computing power. However, the edge-side adaptive computing power constraint boundary parameters generated by the above-mentioned fusion process in this application enable the computing power allocation strategy to anticipate the changing trends of battery margin and scene pressure, thereby driving the subsequent quantitative inference process to make forward-looking precision adjustments. For example, as the battery level drops from 50% to 40%, even if the scene activity pressure index remains unchanged, the edge-side adaptive computing power constraint boundary parameters will be lowered accordingly. This prompts the subsequent steps to select a more conservative combination of quantization bit widths for the dynamic precision adaptive layer-by-layer quantization matching strategy, thereby reducing computing power consumption in a smooth and gradual manner. This avoids a precipitous drop in computing power caused by a critical battery alarm, and improves the robustness of the inspection robot's task execution during long-cycle operations.
[0040] In summary, the edge-side adaptive computing power constraint boundary parameter, constructed collaboratively based on the current battery life representation parameters and the activity distribution characteristics of the inspection field of view, lays a dynamically variable decision-making foundation for the entire vehicle-mounted perception dynamic accuracy adaptive quantization inference method. This edge-side adaptive computing power constraint boundary parameter is not a statically preset fixed threshold, but a dynamic variable that evolves in real time with the consumption of the inspection robot's remaining battery power and fluctuations in the pedestrian density of the inspected area. It is precisely because this edge-side adaptive computing power constraint boundary parameter can objectively reflect the dynamic game relationship between the inspection robot's "physical strength" and the "pressure" of the scene at the current moment that the dynamic accuracy adaptive layer-by-layer quantization matching strategy established in step 2 and the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping map generated in step 3 possess the ability to adjust perception accuracy and computing power consumption in real time according to changes in operating conditions. This technology fundamentally solves the technical defect of the disconnect between perception logic and physical operation in traditional perception solutions. It enables inspection robots to achieve fine-grained management and on-demand allocation of limited computing resources on the edge side while maintaining recognition sensitivity, in the face of complex and ever-changing inspection scenarios in the park, such as the empty square in the early morning to the crowded canteen entrance at noon. This extends the inspection coverage and continuous operation time within a single battery cycle.
[0041] Optionally, step 2 includes: Using the edge-side adaptive computing power constraint boundary parameters as the basis for inference accuracy adjustment, the association of computational mapping nodes inside the basic multi-source perception diagnostic mapping map is decoupled to determine the layer-level mutation candidate mapping level of the basic multi-source perception diagnostic mapping map under the current computing power constraint; Based on the layer-level mutation candidate mapping hierarchy, a dynamic accuracy-adaptive layer-by-layer quantization matching strategy for real-time behavior detection is established.
[0042] Preferably, in step 2, the edge-side adaptive computing power constraint boundary parameter is used as the basis for adjusting inference accuracy. The specific implementation process for decoupling the computational mapping node associations within the basic multi-source perception diagnostic mapping graph is as follows. The basic multi-source perception diagnostic mapping graph is a directed acyclic computational graph representation of a deep convolutional neural network model pre-constructed for visual behavior recognition tasks of typical dynamic targets (e.g., pedestrians, non-motorized vehicles, motorized vehicles, etc.) in inspection scenarios. This basic multi-source perception diagnostic mapping graph consists of multiple sequentially cascaded or cross-layer connected computational mapping nodes. Each computational mapping node corresponds to an operator layer in the deep convolutional neural network model, such as a convolutional computation layer, batch normalization layer, activation function layer, pooling layer, or fully connected layer. After obtaining the edge-side adaptive computing power constraint boundary parameter output in step 1, this application uses the edge-side adaptive computing power constraint boundary parameter as a threshold variable to control the depth of the decoupling operation. The decoupling operation first traverses all computational mapping nodes in the basic multi-source perception diagnostic mapping graph and extracts the statistical characteristics of the weight tensor stored inside each computational mapping node, such as the numerical distribution range, standard deviation, and information entropy of the weight tensor. Subsequently, for each computational mapping node, this application calculates a node redundancy score based on the information entropy of its weight tensor. A computational mapping node with lower information entropy indicates that its weight values are more concentrated, its response pattern to input features is relatively simple, and there is significant room for numerical precision compression. Conversely, a computational mapping node with higher information entropy indicates that its weight values are more dispersed, its ability to distinguish different input patterns is stronger, and numerical precision compression has a greater negative impact on overall perception performance. The node redundancy score of each computational mapping node is compared with the edge-side adaptive computing power constraint boundary parameters. When the node redundancy score is higher than the redundancy tolerance threshold obtained from the edge-side adaptive computing power constraint boundary parameter mapping, the computational mapping node is marked as a decoupled node, and its fixed computational mapping node association with subsequent computational mapping nodes is temporarily released to allow subsequent processes to configure a lower numerical quantization bit width for this computational mapping node independently. By traversing and comparing all computational mapping nodes in the basic multi-source perception diagnostic mapping graph, a set containing all nodes marked as decoupled is finally obtained. This set, along with the topological relationship between each computational mapping node in the set and its neighboring computational mapping nodes after decoupling, together constitute a decoupling state description of the computational mapping node association.
[0043] Preferably, in the specific implementation process of the association of computational mapping nodes within the above-mentioned decoupled basic multi-source sensing diagnostic mapping map, the various computational mapping nodes constituting the basic multi-source sensing diagnostic mapping map undertake different technical roles and cooperate in coordination through signal transmission and transformation during the forward propagation process. The technical roles and mutual cooperation relationships of the convolutional computation layer, batch normalization layer, activation function layer, pooling layer and fully connected layer are described in detail below.
[0044] As the core feature extraction and computational mapping node in the basic multi-source perception diagnostic mapping map, the convolutional computation layer internally stores a set of learnable convolutional kernel weight tensors. Each convolutional kernel weight tensor is a multi-dimensional array, typically a four-dimensional structure consisting of kernel height, kernel width, number of input channels, and number of output channels. When the input feature map from the previous computational mapping node or the visual monitoring sequence to be inferred enters the convolutional computation layer, the layer generates a two-dimensional output response map by sliding a window through each convolutional kernel weight tensor along the height and width directions of the input feature map, and performing element-wise multiplication and summation operations between the convolutional kernel weight tensor and the local receptive field region of the input feature map at each sliding window position. The output response maps of all output channels are stacked along the channel dimension to form the output feature map of the convolutional computation layer. The technical role of the convolutional computation layer is to automatically learn and extract local spatial patterns in the input data, such as edges, corners, textures, and higher-level semantic component features, through the parameterized form of the convolutional kernel weight tensor. In the decoupling operation, the statistical characteristics of the weight tensors extracted by the convolutional computation layer, especially the information entropy of the weight tensors, directly determine the node redundancy score of the corresponding convolutional computation layer. If the information entropy of the weight tensors of a certain convolutional computation layer is low, it indicates that its multiple convolutional kernels have learned similar or redundant local feature extraction patterns. Under the constraint of computing power, there is room for this convolutional computation layer to be marked as a decoupling node and quantized with a lower numerical quantization bit width.
[0045] Batch normalization layers typically follow convolutional layers and precede activation function layers. Their function is to independently normalize the output feature map of the convolutional layers along each channel dimension. Specifically, the batch normalization layer first calculates the mean and variance of all activation values in each channel within the current mini-batch training data. Then, it uses this mean and variance to standardize each activation value in that channel by subtracting the mean and dividing by the standard deviation, adjusting its distribution to near a standard normal distribution with a mean of zero and a variance of one. Next, the batch normalization layer introduces a pair of learnable scaling and translation parameters to perform an affine transformation on the standardized activation values, restoring the network's nonlinear expressive power that might have been weakened by the standardization operation. During inference, the batch normalization layer uses the global mean and global variance obtained through moving average statistics during training instead of the mini-batch statistics. The complementary role of the batch normalization layer in signal transmission is reflected in its effective mitigation of the internal covariate shift problem during deep neural network training, ensuring a relatively stable input signal distribution for subsequent activation function layers, thereby accelerating model convergence and improving final inference accuracy. In the decoupling operation, the batch normalization layer itself contains a small number of scaling and translation parameters. Its information entropy evaluation is usually combined with the preceding convolutional calculation layer as the basis for evaluating the overall redundancy of the calculation mapping node group.
[0046] The activation function layer follows the batch normalization layer or convolutional computation layer. Its technical role is to apply a pre-defined nonlinear mapping function to each element in the input feature map, thereby introducing nonlinear transformation capabilities into the deep neural network model. Commonly used activation functions include the modified linear unit (MRU) and its variants, such as the leaky MRU or the parameterized MRU. Taking the MRU as an example, its mathematical definition is that the output equals the input value when the input value is greater than zero, and the output is set to zero when the input value is less than or equal to zero. This nonlinear mapping enables the deep neural network model to fit arbitrarily complex nonlinear decision boundaries, not just combinations of linear transformations. The output signal of the activation function layer is the activation feature map after nonlinear activation. At the signal coordination level, the activation function layer, through its sparse activation characteristics (e.g., the MRU sets negative values to zero), keeps some neurons in an inactive state. This not only increases the sparsity and generalization ability of the model but also reduces the number of activation values participating in actual computation during forward propagation, indirectly reducing the computational load of subsequent mapping nodes. In the decoupling operation, the activation function layer itself does not contain learnable weight parameters, so it is not used as an independent node redundancy scoring object. However, its sparse activation degree will affect the information entropy distribution of the input activation tensor of the subsequent calculation of the mapping node.
[0047] Pooling layers are typically inserted periodically after several consecutive convolutional and activation function layers. Their function is to downsample the input feature map in terms of spatial dimension, reducing its spatial resolution and thus decreasing the computational load of subsequent mapping nodes while expanding the receptive field of subsequent convolutional layers. Common pooling operations include max pooling and average pooling. Taking max pooling as an example, the pooling layer divides the input feature map into several non-overlapping or partially overlapping rectangular sliding windows, and selects the maximum value among all activation values within each window as the representative output value for that window region. After processing by the pooling layer, the height and width of the output feature map are reduced to a fraction of the input feature map (e.g., half). The contribution of the pooling layer to signal processing lies in providing a degree of local translation invariance. Even if the target in the input image undergoes a slight translation, the max-value selection mechanism of the pooling layer can still ensure that downstream mapping nodes receive relatively stable feature responses. Meanwhile, the downsampling effect of the pooling layer compresses the spatial dimension of the feature map, resulting in a progressively decreasing amount of data required for subsequent computational mapping nodes. This is crucial for achieving efficient inference under resource-constrained conditions at the edge. In the decoupling operation, the pooling layer also does not contain learnable weight parameters, but its downsampling rate affects the computational graph topology of the entire basic multi-source perception diagnostic mapping map.
[0048] Fully connected layers are typically located at the end of the basic multi-source perceptual diagnostic mapping map, near the output computation mapping node. Their technical function is to flatten the high-level semantic feature maps extracted by the preceding convolutional and pooling layers into a one-dimensional feature vector. This feature vector is then multiplied by a two-dimensional weight matrix stored within the fully connected layer, linearly mapping the feature vector to an output dimension equal to the number of preset abnormal behavior categories. A fully connected layer can be viewed as a feature aggregation operation of the global receptive field; each output neuron establishes a weighted connection with all elements in the input feature vector. Therefore, the weight parameters in a fully connected layer typically account for a large proportion of the total parameters in the entire deep neural network model. After the output of the fully connected layer is normalized using a flexible maximum function, the confidence score of the visual monitoring sequence to be inferred belonging to each preset abnormal behavior category can be obtained. In the decoupling operation, the weight tensor of the fully connected layer usually has a high information entropy, indicating that its weight values are relatively dispersed and have a strong ability to distinguish between different categories. Therefore, its node redundancy score is often lower than the redundancy tolerance threshold. Under the boundary parameter of adaptive computing power constraint on the edge side, it is generally not marked as a decoupling node, so as to maintain the basic guarantee of the accuracy of the judgment of abnormal behavior diagnosis results.
[0049] The aforementioned convolutional computation layers, batch normalization layers, activation function layers, pooling layers, and fully connected layers, connected by pre-defined directed edges within the basic multi-source perception diagnostic mapping map, constitute a progressively deeper, hierarchical visual feature extraction and abstract signal processing system. The visual monitoring sequence to be inferred first passes through shallow convolutional computation layers, batch normalization layers, activation function layers, and pooling layers to extract low-level geometric features such as edges and corners. As the network depth increases, the middle-layer convolutional computation layers combine the low-level geometric features to form component-level semantic features, such as a pedestrian's head, torso, or a vehicle's wheel hub and window. The deeper convolutional computation layers and fully connected layers further aggregate the component-level semantic features into a global scene-level semantic representation, ultimately outputting the abnormal behavior diagnostic results. In the decoupling operation of step 2, this application utilizes the differentiated statistical characteristics of the information entropy of the weight tensor of computational mapping nodes of different depths and types. By comparing with the boundary parameters of adaptive computing power constraints on the edge side, it accurately identifies convolutional computation layers with low information entropy and high redundancy in contributing to the final perception accuracy, and marks them as decoupling nodes. This provides a node-level fine-grained decoupling basis for the subsequent generation of a dynamic precision adaptive layer-by-layer quantization matching strategy with layer-by-layer heterogeneous quantization bit width. This mechanism enables deep convolutional computation layers and fully connected layers that play a key role in feature extraction in the basic multi-source perception diagnostic mapping map to retain high numerical accuracy, while some shallow or mid-level convolutional computation layers with high redundancy in contributing to feature extraction can run with lower numerical accuracy under computing power constraints, realizing fine-grained dynamic control of the trade-off between model accuracy and inference computing power.
[0050] Preferably, after completing the decoupling state description of the computational mapping nodes, this application further determines the specific process of the layer-level mutation candidate mapping level of the basic multi-source sensing diagnostic mapping map under the current computing power constraints as follows. The layer-level mutation candidate mapping level refers to an alternative computational level with different quantization bit width configuration options generated for each computational mapping node based on the original hierarchical structure of the basic multi-source sensing diagnostic mapping map. For each decoupling node identified in the aforementioned decoupling state description, this application generates a set of mutation candidate mapping levels with different numerical quantization bit widths. For example, for an original computational mapping node that uses 32-bit floating-point numbers for weight storage and calculation, its corresponding layer-level mutation candidate mapping level may include a layer variant using 16-bit floating-point quantization, a layer variant using 8-bit integer quantization, and a layer variant using 4-bit integer quantization. Each layer-level mutation candidate mapping level retains the topological connection position of the original computational mapping node in the basic multi-source sensing diagnostic mapping graph. However, the weight tensors and activation tensors used when performing convolution or matrix multiplication operations are constrained to the set of discrete values that can be represented by the selected quantization bit width. Meanwhile, for computational mapping nodes whose redundancy scores are below the redundancy tolerance threshold and are not marked as decoupling nodes, this application treats them as precision-sensitive nodes. Their corresponding layer-level mutation candidate mapping levels are uniquely locked as the original unquantized or high-precision quantized (e.g., 16-bit floating-point) level variants, and low-bit quantization mutations are not allowed. Through the above processing, this application transforms the basic multi-source sensing diagnostic mapping graph from a static computational graph with fixed precision for all computational mapping nodes into a dynamic candidate computational graph structure with multiple parallel quantization bit width branches at some computational mapping node positions. Ultimately, the union of all candidate layer-level mutation mapping levels at the locations of all computational mapping nodes constitutes the set of candidate layer-level mutation mapping levels for the basic multi-source sensing diagnostic mapping map under the current computing power constraints. This set provides a complete search space definition for subsequent searches for better layer-by-layer quantization bit width combinations.
[0051] Preferably, the specific implementation process of establishing a dynamic accuracy adaptive layer-by-layer quantization matching strategy for real-time behavior detection based on the layer-level mutation candidate mapping hierarchy in step 2 is as follows. Based on the already obtained set of layer-level mutation candidate mapping hierarchies, this application transforms the strategy establishment problem into an optimization problem of searching for a better quantization configuration scheme that satisfies the edge-side adaptive computing power constraint boundary parameters within a discrete and finite combination space. First, this application defines an inference path from the input node to the output node of the basic multi-source perception diagnostic mapping map. Each computational mapping node on this inference path must select a specific variant from one or more corresponding layer-level mutation candidate mapping hierarchies. For each possible combination scheme of layer-level mutation candidate mapping hierarchies, this application calculates two core evaluation indicators for the combination scheme: one is the expected inference computing power consumption estimate, and the other is the expected perception accuracy decay estimate. The expected inference computing power consumption estimate can be obtained by weighting and estimating the weight bit width and activation bit width corresponding to each layer-level mutation candidate mapping hierarchy, combined with the theoretical floating-point operation count of the computational mapping node in the basic multi-source perception diagnostic mapping map. The expected perception accuracy decay estimate can be calculated using a preset empirical model of accuracy loss, based on the node redundancy score of each computational mapping node and the quantization bit width of the selected layer-level mutation candidate mapping level. Lower quantization bit width and lower node redundancy score result in a smaller accuracy decay estimate. The search algorithm searches the combinatorial space for layer-level mutation candidate mapping level combinations that minimize the expected inference computational power consumption estimate compared to the target computational power budget upper limit obtained from the edge-side adaptive computational power constraint boundary parameter mapping, and minimize the expected perception accuracy decay estimate. The search process can employ greedy algorithms, genetic algorithms, or reinforcement learning-based search strategies to quickly converge to an approximate optimal solution while meeting real-time requirements. When the search terminates and a satisfactory layer-level mutation candidate mapping level combination is determined, the binding relationship between each computational mapping node and its selected layer-level mutation candidate mapping level, as well as the sequential dependency relationship between computational mapping nodes arranged according to the original topological order of the basic multi-source perception diagnostic mapping map, are recorded as a dynamic accuracy adaptive layer-by-layer quantization matching strategy.
[0052] Preferably, in establishing a dynamic precision adaptive layer-by-layer quantization matching strategy, this application also involves probing the quantization bit width response mapping interval of the layer-level variant candidate mapping level to determine the quantization precision compression mapping node parameters of each layer-level variant candidate mapping level under baseline operating conditions. The quantization bit width response mapping interval refers to the numerical change curve formed by the mean cosine similarity or peak signal-to-noise ratio of the output feature map of a layer-level variant candidate mapping level relative to the unquantized original feature map when different quantization bit widths are used for that layer-level variant candidate mapping level. This application performs a forward propagation calibration once for each layer-level variant candidate mapping level variant of each computational mapping node under baseline operating conditions through offline pre-calibration. During calibration, a set of calibration image datasets covering typical inspection scenarios is input to the basic multi-source perception diagnostic mapping map, and the original output feature map of the computational mapping node in the unquantized state and the quantized output feature map under different quantization bit width configurations are recorded respectively. By calculating the statistical difference between the two output feature maps, a quantization precision compression mapping node parameter is generated for each quantization bit width. This parameter is a dimensionless scalar value; a larger value indicates greater information loss and more significant precision decay due to quantization. The quantization precision compression mapping node parameter is stored in the attribute descriptor of the layer-level mutation candidate mapping level. When subsequent search algorithms evaluate the expected perceived precision decay estimate for each layer-level mutation candidate mapping level combination scheme, they directly read and accumulate the quantization precision compression mapping node parameters of each selected layer-level mutation candidate mapping level to efficiently estimate the overall scheme's precision loss. By pre-probing the quantization bit width response mapping interval and calibrating the precision compression mapping node parameter, this application transforms the precision evaluation process, which originally required complex calculations at runtime, into a fast lookup and accumulation operation of pre-stored calibration values, thereby improving the real-time performance and feasibility of establishing a dynamic precision adaptive layer-by-layer quantization matching strategy.
[0053] Preferably, the established dynamic precision adaptive layer-by-layer quantization matching strategy is essentially a dynamically executable configuration list. It clarifies which quantization bit width layer-level mutation candidate mapping level each computational mapping node in the basic multi-source perception diagnostic mapping map should be specifically instantiated into, given the edge-side adaptive computing power constraint boundary parameters. There is a close causal relationship between the establishment process of this dynamic precision adaptive layer-by-layer quantization matching strategy and the edge-side adaptive computing power constraint boundary parameters generated in step 1. When the edge-side adaptive computing power constraint boundary parameters are lowered due to a decrease in the inspection robot's battery power or a reduction in scene activity pressure, the decoupling operation will mark more computational mapping nodes with higher node redundancy scores as decouplingable nodes, thereby generating more low-bit-width layer-level mutation candidate mapping level variants for these computational mapping nodes. Simultaneously, the upper limit of the target computing power budget followed by the search algorithm is tightened, forcing the search process to tend to select combinations containing more four-bit or eight-bit integer quantization variants. Conversely, when the edge-side adaptive computing power constraint boundary parameters are increased, the search space shrinks, and the search algorithm tends to retain more high-precision level variants. This linkage mechanism enables the dynamic precision adaptive layer-by-layer quantization matching strategy to adjust the quantization precision of each computational mapping node in the basic multi-source perception diagnostic mapping map in a continuous or quasi-continuous manner as the real-time status and scene pressure of the inspection robot change. Thus, while achieving the recognition sensitivity required for real-time behavior detection, it realizes dynamic on-demand trimming or on-demand enhancement of computing power consumption.
[0054] Preferably, the establishment of the dynamic precision adaptive layer-by-layer quantization matching strategy provides direct and specific construction instructions for generating the vehicle-mounted perception dynamic precision adaptive quantization inference mapping map in subsequent step 3. Specifically, step 3 will extract the corresponding underlying structure sequence and signal transmission mechanism data from the basic multi-source perception diagnostic mapping map based on the binding relationship between the computational mapping nodes and the layer-level mutation candidate mapping levels recorded in the dynamic precision adaptive layer-by-layer quantization matching strategy. Then, the weight distribution sequence and feature transmission information payload will be reparameterized and rescaled using the quantization floating-point conversion step size calibration parameter, ultimately reconstructing a vehicle-mounted perception dynamic precision adaptive quantization inference mapping map that can be directly deployed and run on the edge computing unit of the inspection robot. Unlike traditional solutions where the internal numerical precision of the perception model is permanently fixed once deployed, this application, through the dynamic precision adaptive layer-by-layer quantization matching strategy established in step 2, endows the subsequent reconstruction stage with the ability to dynamically adjust the numerical representation precision of each layer within the model according to real-time operating conditions. This enables the vehicle-mounted perception dynamic precision adaptive quantization inference mapping map to switch online from low-power, low-precision inference mode to high-power, high-precision inference mode within seconds or sub-second time granularity when the inspection robot moves from an empty underground parking lot into a crowded shopping mall atrium. This achieves a dynamic balance between the two mutually restrictive goals of extending overall battery life and improving perception quality in key areas, providing more reliable and efficient perception technology support for long-term unattended inspection operations in complex and dynamic public scenarios.
[0055] Optionally, based on the layer-level mutation candidate mapping hierarchy, a dynamic accuracy-adaptive layer-by-layer quantization matching strategy for real-time behavior detection is established, including: The quantization bit width response mapping range of the layer-level mutation candidate mapping level is detected, and then the quantization accuracy compression mapping node parameters of each layer-level mutation candidate mapping level under the baseline operating conditions are determined. The quantization precision compression mapping node parameters are matched with the edge-side adaptive computing power constraint boundary parameters to establish a dynamic precision adaptive layer-by-layer quantization matching strategy for real-time behavior detection.
[0056] Preferably, the specific implementation process of detecting the quantization bit-width response mapping interval of the candidate layer-level mutation map in step 2 is as follows. The quantization bit-width response mapping interval refers to the numerical mapping relationship curve of the similarity between the output response and the unquantized original output response of a candidate layer-level mutation map derived from a certain computational mapping node in the basic multi-source perception diagnostic map, when a series of discrete numerical quantization bit-width configurations are used. In order to obtain this quantization bit-width response mapping interval, this application adopts an offline pre-calibration method to perform forward propagation calibration on each candidate layer-level mutation map under the baseline operating conditions. Specifically, a set of calibration image datasets covering common operating scenarios of inspection robots is pre-collected. This calibration image dataset contains typical image samples under different lighting conditions, different dynamic target densities, and different scene structural complexities. The calibration image dataset is input frame by frame into the basic multi-source perception diagnostic map. For each computational mapping node in the basic multi-source perception diagnostic map, in its original state without any quantization processing, the original output feature map output by the computational mapping node to the next computational mapping node is recorded. The original output feature map is a multidimensional array, where each element is a 32-bit floating-point number, representing the response strength of the computational mapping node to the input features at the corresponding spatial location and channel dimension.
[0057] Preferably, after recording the original output feature map, this application applies different quantization bit width configurations to the weight tensor and input activation tensor of the computational mapping node to simulate the computational behavior of the computational mapping node when instantiated into different layer-level mutation candidate mapping levels. The applied quantization bit width configuration may include, but is not limited to, 16-bit floating-point quantization, 8-bit integer quantization, and 4-bit integer quantization. For each quantization bit width configuration, when performing the quantization operation, this application first calculates the absolute maximum value of the values in the weight tensor, and uses this absolute maximum value as a scaling factor to linearly map each floating-point value in the weight tensor and round it into the set of discrete integers that can be represented by the target quantization bit width; subsequently, in the forward computation process of the computational mapping node, the quantized integer weight values are dequantized back to the floating-point domain by multiplying by the scaling factor to perform approximate convolution operations or matrix multiplication operations. After the above quantization forward propagation process, a quantized output feature map can be obtained at the output end of the computational mapping node. A corresponding quantized output feature map is generated for each quantization bit width configuration. Subsequently, this application calculates a statistical similarity metric between each quantized output feature map and the aforementioned recorded original output feature map. For example, it calculates the mean cosine similarity of the two feature maps in high-dimensional space, or the peak signal-to-noise ratio (PSNR) of the two. Taking the mean cosine similarity as an example, its value ranges from zero to one. The closer the value is to one, the more consistent the geometric direction of the quantized output feature map and the original output feature map, and the smaller the information loss caused by quantization. By using different quantization bit width configurations as the horizontal axis variable and the calculated mean cosine similarity or PNR as the vertical axis variable, a curve reflecting the quantization response characteristics of the layer-level mutation candidate mapping layer can be plotted. This curve and its corresponding numerical variation pattern constitute the quantization bit width response mapping interval of the layer-level mutation candidate mapping layer.
[0058] Preferably, based on the detected quantization bit width response mapping range of the layer-level mutation candidate mapping levels, this application further determines the quantization precision compression mapping node parameters of each layer-level mutation candidate mapping level under baseline operating conditions. The quantization precision compression mapping node parameter is a dimensionless scalar value used to quantitatively characterize the degree of perceptual information loss introduced by instantiating a computational mapping node into a layer-level mutation candidate mapping level of a specified bit width, relative to the unquantized baseline state. Specifically, for each layer-level mutation candidate mapping level of each computational mapping node, this application extracts a similarity metric value corresponding to the adopted quantization bit width configuration from its corresponding quantization bit width response mapping range, for example, extracting the average cosine similarity value when using four-bit integer quantization. Subsequently, the similarity metric value is converted into a quantization precision compression mapping node parameter using a preset mapping function. For example, the quantization precision compression mapping node parameter can be defined as one minus the average cosine similarity value, making the value of this parameter vary between zero and one. The lower the value, the smaller the information loss caused by quantization precision compression, and the more suitable the layer-level mutation candidate mapping level is to be selected under computationally limited conditions. The generated quantization precision compression mapping node parameter, serving as an attribute descriptor for the layer-level mutation candidate mapping level, is persistently stored in the data structure associated with that layer-level mutation candidate mapping level. Through the aforementioned offline calibration process, all available layer-level mutation candidate mapping levels in the basic multi-source perception diagnostic mapping map are assigned corresponding quantization precision compression mapping node parameters. This lays the data foundation for quickly and accurately evaluating the accuracy loss of different quantization combination schemes during runtime.
[0059] Preferably, the specific implementation process of the dynamic precision adaptive layer-by-layer quantization matching strategy for real-time behavior detection in step 2, which matches the quantization precision compression mapping node parameters with the edge-side adaptive computing power constraint boundary parameters, is as follows. During the online operation phase of the inspection robot, after step 1 has outputted the edge-side adaptive computing power constraint boundary parameters reflecting the comprehensive constraint strength of the current operating state, this application uses these edge-side adaptive computing power constraint boundary parameters as the overall constraint condition for real-time strategy search. First, this application determines a target computing power budget upper limit based on the current value of the edge-side adaptive computing power constraint boundary parameters. This mapping relationship can be predefined using a monotonically decreasing function or a lookup table, such that the lower the value of the edge-side adaptive computing power constraint boundary parameters (representing a more scarce power supply or a simpler scenario), the lower the target computing power budget upper limit is set, and the more stringent the selection of the quantization bit width. Subsequently, this application initiates a search process executed within the combined space composed of all layer-level mutation candidate mapping levels. The goal of this search process is to find a complete inference path from the input to the output of the basic multi-source perception diagnostic mapping. Each computational mapping node on this inference path selects a specific layer-level mutation candidate mapping level. The combined scheme must simultaneously satisfy two constraints: first, the estimated expected inference computing power consumption of the combined scheme does not exceed the upper limit of the target computing power budget; second, among all combined schemes that satisfy the computing power constraints, the estimated expected perception accuracy decay of the combined scheme is minimized.
[0060] Preferably, when evaluating any candidate layer-level mutation candidate mapping layer combination scheme during the search process, the estimated inference computing power consumption is calculated as follows: For each layer-level mutation candidate mapping layer in the combination scheme, based on its weight bit width and activation bit width, combined with the theoretical floating-point operation count of the computing mapping node in the basic multi-source perception diagnostic mapping map, a preset weighted formula is used to estimate the computing resource consumption of the layer-level mutation candidate mapping layer in the actual inference process. The computing resource consumption of all selected layer-level mutation candidate mapping layers is accumulated to obtain the estimated inference computing power consumption of the combination scheme. At the same time, the estimated perception accuracy decay directly depends on the aforementioned offline calibrated quantization precision compression mapping node parameters. For each layer-level mutation candidate mapping layer in the combination scheme, the corresponding quantization precision compression mapping node parameters are read from its attribute descriptor. The quantization precision compression mapping node parameters of all selected layer-level mutation candidate mapping layers are weighted and accumulated or summed. The accumulated result is the estimated perception accuracy decay of the combination scheme. During weighted accumulation, different weight coefficients can be assigned based on the hierarchical depth of different computational mapping nodes in the basic multi-source perception diagnostic mapping map or their contribution to the final perception decision. For example, computational mapping nodes closer to the output end have a greater impact on accuracy, and their corresponding weight coefficients are set higher. In this way, the search process does not need to perform actual quantization inference to evaluate accuracy loss at runtime. Instead, it efficiently estimates the expected performance of each candidate solution through fast lookup of pre-stored calibration values and linear combination operations.
[0061] Preferably, the core of the matching process lies in using the target computing power budget upper limit obtained by mapping the edge-side adaptive computing power constraint boundary parameters as a hard constraint, and using the quantization precision compression mapping node parameters of each layer-level mutation candidate mapping level as a cost term in the optimization objective, thereby selecting the quantization configuration scheme with the minimum precision loss from the candidate solution set that satisfies the computing power constraint. When the search process converges and determines the optimal combination scheme of layer-level mutation candidate mapping levels, the correspondence between each computing mapping node in this combination scheme and the selected layer-level mutation candidate mapping level, along with the execution order of each computing mapping node arranged along the topological order of the basic multi-source perception diagnostic mapping map, are encapsulated and recorded as a dynamic precision adaptive layer-by-layer quantization matching strategy. The generation frequency of this dynamic precision adaptive layer-by-layer quantization matching strategy is synchronized with the update frequency of the edge-side adaptive computing power constraint boundary parameters, for example, triggering a re-search and strategy update every few seconds or when the inspection robot enters a new scene area. Through this mechanism, this application realizes the dynamic matching of quantization precision compression mapping node parameters and edge-side adaptive computing power constraint boundary parameters, so that the final generated dynamic precision adaptive layer-by-layer quantization matching strategy can accurately reflect the inspection robot's trade-off between computing power and precision under the current specific operating conditions.
[0062] In summary, the offline calibration step, which detects the quantization bit-width response mapping range of candidate layer-level mutation maps and determines the quantization precision compression mapping node parameters, and the online search step, which matches the edge-side adaptive computing power constraint boundary parameters during runtime, constitute an efficient collaborative system. The offline calibration step transforms the complex, data-dependent precision loss assessment process into the calibration of independent and reusable quantization precision compression mapping node parameters for each candidate layer-level mutation map. This technique significantly reduces the computational complexity of the online search, making it feasible to generate dynamic precision adaptive layer-by-layer quantization matching strategies in real-time on the resource-constrained edge computing units of inspection robots. The online search step uses the real-time changing edge-side adaptive computing power constraint boundary parameters as dynamic constraints to drive the search algorithm to find the optimal solution that satisfies the current working conditions in the feature space composed of quantization precision compression mapping node parameters. Unlike traditional solutions where the quantization strategy is static and cannot be changed once determined, this application establishes a dynamic precision adaptive layer-by-layer quantization matching strategy through the above matching process. This enables the basic multi-source perception diagnostic mapping map to adapt to various complex situations in park inspection scenarios, ranging from low to high activity levels, in a "one map, multiple uses, adaptable to different situations" form. This achieves an adaptive dynamic balance between extending the robot's endurance and maintaining high-reliability perception in key areas.
[0063] Optionally, step 3: Based on the dynamic precision adaptive layer-by-layer quantization matching strategy, generate an on-board perception dynamic precision adaptive quantization inference mapping map, including: The configuration parameters of the dynamic precision adaptive layer-by-layer quantization matching strategy are analyzed to obtain the quantization floating-point conversion step size calibration parameter; Extract the underlying structure sequence of the basic multi-source sensing diagnostic mapping map to obtain the initial weight parameter distribution sequence; The quantization floating-point conversion step size calibration parameter is injected into the initial weight parameter distribution sequence to generate a step size fusion weight distribution sequence; Determine the activation response threshold of the step-size fusion weight distribution sequence to generate a dynamic precision-trimmed weight distribution sequence; Based on the dynamic precision pruning weight distribution sequence, an adaptive quantization inference mapping map for vehicle perception dynamic precision is generated.
[0064] Preferably, the specific implementation process of parsing the configuration parameters of the dynamic precision adaptive layer-by-layer quantization matching strategy to obtain the quantization floating-point conversion step size calibration parameter in step 3 is as follows. After the establishment of the dynamic precision adaptive layer-by-layer quantization matching strategy is completed in step 2, the dynamic precision adaptive layer-by-layer quantization matching strategy is stored in the local memory of the inspection robot in the form of structured data. It records the identifier of the selected layer-level mutation candidate mapping level for each computational mapping node in the basic multi-source perception diagnostic mapping map, as well as the execution order of each computational mapping node arranged along the original topological order of the basic multi-source perception diagnostic mapping map. This application first parses and traverses the dynamic precision adaptive layer-by-layer quantization matching strategy, and reads the layer-level mutation candidate mapping level identifier corresponding to each entry. For each layer-level mutation candidate mapping level identifier, this application extracts the configured weight bit width value and activation bit width value from the attribute descriptor associated with the layer-level mutation candidate mapping level. For example, the weight bit width value is an eight-bit integer and the activation bit width value is an eight-bit integer, or the weight bit width value is a four-bit integer and the activation bit width value is an eight-bit integer, etc. Based on the extracted weight bit width and activation bit width values, this application calculates a quantization floating-point conversion step size calibration parameter to guide subsequent quantization parameter remapping. The quantization floating-point conversion step size calibration parameter is a floating-point scalar value whose magnitude determines the quantization step size used when mapping the original floating-point weight tensor and floating-point activation tensor to the target discrete integer numerical space. Specifically, for the weight tensor, the quantization floating-point conversion step size calibration parameter can be obtained by dividing the absolute maximum value of the weight tensor by half the target integer representation range (e.g., for 8-bit integer quantization, the representation range is -128 to +127, with a maximum absolute value of 127); for the activation tensor, it can be obtained by calculating the maximum moving average value of the activation tensor under baseline operating conditions and combining it with the target integer representation range. The calibration parameters for the weight-side quantization floating-point conversion step size and the calibration parameters for the activation-side quantization floating-point conversion step size corresponding to each computational mapping node, obtained through analytical calculation, are collected into a parameter set. This parameter set provides a calibration benchmark for numerical scaling and mapping of the initial weight parameter distribution sequence and feature transfer information payload in subsequent steps.
[0065] Preferably, the specific implementation process of extracting the underlying structure sequence of the basic multi-source sensing diagnostic mapping graph to obtain the initial weight parameter distribution sequence in step 3 is as follows. The basic multi-source sensing diagnostic mapping graph is represented in the storage medium as a directed acyclic graph data structure, where each computational mapping node is associated with a weight tensor containing the learnable parameters of that computational mapping node. This application performs a depth-first or breadth-first topological traversal of the basic multi-source sensing diagnostic mapping graph, visiting each computational mapping node sequentially according to the execution order of the computational mapping node in the inference process. For each traversed computational mapping node, this application reads the original 32-bit floating-point format weight tensor of the computational mapping node from its associated storage area without any quantization processing. For each weight tensor, this application expands its continuous storage sequence in memory into a one-dimensional weight parameter distribution sequence, where each element of the weight parameter distribution sequence is the original floating-point weight value. By concatenating the weight parameter distribution sequences of all computational mapping nodes according to the order of topological traversal, the initial weight parameter distribution sequence of the basic multi-source sensing diagnostic mapping graph is formed. This initial weight parameter distribution sequence comprehensively reflects the numerical magnitude and distribution range of all learnable parameters in the basic multi-source sensing diagnostic mapping before quantization adjustment, and serves as the processing object for subsequent quantization parameter injection and pruning operations. Unlike traditional solutions that directly use the initial weight parameter distribution sequence for inference deployment, this application provides an operable parameter basis for online dynamic quantization reconstruction by extracting the initial weight parameter distribution sequence.
[0066] Preferably, the specific implementation process of injecting the quantization floating-point conversion step size calibration parameter into the initial weight parameter distribution sequence to generate the step-size fused weight distribution sequence in step 3 is as follows. After obtaining the initial weight parameter distribution sequence and the quantization floating-point conversion step size calibration parameter corresponding to each computational mapping node, this application logically divides the initial weight parameter distribution sequence into multiple subsequences according to the boundaries of the computational mapping nodes, with each subsequence corresponding to the weight parameter distribution sequence of a computational mapping node. For each subsequence corresponding to a computational mapping node and its corresponding quantization floating-point conversion step size calibration parameter, this application performs element-wise quantization and dequantization operations to integrate the effect of the quantization floating-point conversion step size calibration parameter into the numerical representation of the weight parameters. The specific operation process is as follows: For each floating-point weight value in the subsequence, firstly, divide the floating-point weight value by the quantization floating-point conversion step size calibration parameter on the weight side to obtain a quotient value; then, perform a rounding operation on the quotient value (e.g., round to the nearest integer), mapping the continuous floating-point quotient values to discrete integer values; next, multiply the discrete integer value by the quantization floating-point conversion step size calibration parameter, remapping it back to the floating-point value domain. After the above division-rounding-multiplication operation, the original floating-point weight value is replaced with a new floating-point value calibrated by the quantization step size. This new floating-point value represents the closest value that the weight can represent under the target quantization bit width constraint. Re-attach all the subsequences after the above processing to form a new parameter sequence, which is the step-size fused weight distribution sequence. Although each element in the step-size fused weight distribution sequence is still a floating-point number in terms of data type, the effective precision of its value has been limited to discrete grid points determined by the quantization floating-point conversion step size calibration parameter, thereby achieving controllable compression of the precision of the initial weight parameter distribution sequence.
[0067] Preferably, the specific implementation process of determining the activation response threshold of the step-size fusion weight distribution sequence to generate the dynamic precision pruning weight distribution sequence in step 3 is as follows. Although the step-size fusion weight distribution sequence introduces quantization step-size constraints, it still retains some redundant weights with extremely small values. These redundant weights contribute little to the feature response during forward propagation but still consume computational and storage resources. To further prune redundancy, this application determines an activation response threshold for each subsequence (i.e., the weight part corresponding to each computational mapping node) in the step-size fusion weight distribution sequence, based on the layer type of the computational mapping node in the basic multi-source perception diagnostic mapping map and its position in the dynamic precision adaptive layer-by-layer quantization matching strategy. The activation response threshold can be determined based on the statistical distribution of the absolute values of the weights in the subsequence, for example, taking a certain quantile (such as the thirtieth quantile) of the absolute value of the weights in the subsequence as the activation response threshold. Alternatively, the activation response threshold can also be preset according to the quantization bit width of the layer-level mutation candidate mapping level corresponding to the computational mapping node. The lower the quantization bit width, the higher the activation response threshold is set to prune more low-value weights. After determining the activation response threshold, this application performs a threshold comparison operation on each weight value of the subsequence in the step-size fusion weight distribution sequence: if the absolute value of the weight value is less than the activation response threshold, the weight value is set to zero; if the absolute value of the weight value is greater than or equal to the activation response threshold, the weight value is retained. After performing the above threshold comparison and zeroing operation on all subsequences, the resulting new parameter sequence is the dynamic precision pruning weight distribution sequence. Based on the step-size fusion weight distribution sequence, the dynamic precision pruning weight distribution sequence further introduces weight sparsity, causing a large number of weights close to zero to be explicitly set to zero. This not only reduces the effective number of model parameters but also allows for acceleration during subsequent inference using sparse matrix computation libraries.
[0068] Preferably, the specific implementation process of generating the vehicle perception dynamic precision adaptive quantization inference mapping map based on the dynamic precision pruning weight distribution sequence in step 3 is as follows. After obtaining the dynamic precision pruning weight distribution sequence, this application needs to remap the one-dimensional parameter sequence back to the original computation graph topology of the basic multi-source perception diagnostic mapping map to form a deployable and executable vehicle perception dynamic precision adaptive quantization inference mapping map. This application first creates a new computation graph object with the same computation mapping node topology connection relationship as the basic multi-source perception diagnostic mapping map. Then, according to the topology traversal order used when extracting the initial weight parameter distribution sequence, each computation mapping node in the new computation graph object is visited sequentially. For the currently visited computation mapping node, a subsequence matching the shape of the weight tensor of the computation mapping node is extracted sequentially from the dynamic precision pruning weight distribution sequence. This subsequence is reshaped into the multi-dimensional weight tensor shape expected by the computation mapping node, and the reshaped weight tensor is assigned to the computation mapping node as the weight parameter used during its inference. Meanwhile, this application also adds the activation-side quantization floating-point conversion step size calibration parameter corresponding to the computational mapping node, obtained during the parsing of the dynamic precision adaptive layer-by-layer quantization matching strategy, as an inference runtime parameter to the attributes of the computational mapping node. During subsequent actual inference execution, the forward computation logic of this computational mapping node is configured as follows: first, the input feature transfer information payload is quantized and dequantized online using the activation-side quantization floating-point conversion step size calibration parameter; then, it is convolved or multiplied with the assigned quantization weight tensor. After traversing and configuring all computational mapping nodes in the new computational graph object, the resulting new computational graph object containing quantization parameters and quantization computation logic is the vehicle perception dynamic precision adaptive quantization inference mapping graph. This vehicle perception dynamic precision adaptive quantization inference mapping graph fully inherits the layer-by-layer mixed precision characteristics specified by the dynamic precision adaptive layer-by-layer quantization matching strategy, and its internal parameters have been compressed and sparsified through step size fusion and threshold pruning.
[0069] Preferably, the process of generating the vehicle-mounted perception dynamic precision adaptive quantization inference map concretizes the dynamic precision adaptive layer-by-layer quantization matching strategy described as an abstract strategy in step 2 into a model instance that can be directly loaded and executed for forward inference on the edge computing unit of the inspection robot. This vehicle-mounted perception dynamic precision adaptive quantization inference map differs from traditional static quantization models in that its internal weight distribution is not fixed during model training, but is reconstructed in real-time according to the dynamic precision adaptive layer-by-layer quantization matching strategy during the operation of the inspection robot. When the boundary parameters of the edge-side adaptive computing power constraint change and trigger step 2 to generate a new dynamic precision adaptive layer-by-layer quantization matching strategy, step 3 will be re-executed, undergoing the complete process of "strategy parsing - parameter extraction - step size injection - threshold pruning - computation graph reconstruction" again online, thereby generating a vehicle-mounted perception dynamic precision adaptive quantization inference map adapted to the new operating conditions within seconds or sub-seconds. This mechanism enables the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping map to quickly switch from a high compression ratio, low accuracy inference mode to a low compression ratio, high accuracy inference mode when the inspection robot moves from a low-dynamic-activity parking area to a high-dynamic-activity plaza area in a park inspection scenario. This achieves real-time dynamic matching between edge computing resources and perception accuracy requirements. This technology fundamentally differs from the limitation of traditional solutions where the model cannot be dynamically adjusted once quantized and deployed. It provides a flexible and efficient online reconstruction capability of the perception model for long-term adaptive inspection in complex dynamic public scenarios.
[0070] Optionally, based on the dynamic precision pruning weight distribution sequence, an adaptive quantization inference mapping map for vehicle perception dynamic precision is generated, including: Extract the signal transmission mechanism data within the basic multi-source sensing diagnostic mapping map to obtain the feature transmission information payload; The activation distribution of the feature transfer information payload is adjusted using the quantization floating-point conversion step size calibration parameter to generate a quantization scaling adapted activation feature payload. The network computation graph is reconstructed based on the dynamic precision pruning weight distribution sequence and the quantization scaling adaptation activation feature load to generate an on-board perception dynamic precision adaptive quantization inference mapping graph.
[0071] Preferably, the specific implementation process of extracting the signal transmission mechanism data within the basic multi-source perception diagnostic mapping map to obtain the feature transmission information payload in step 3 is as follows. The basic multi-source perception diagnostic mapping map, as a directed acyclic computation graph, contains not only the weight tensors of each computation mapping node itself, but also defines a series of data paths for transmitting intermediate computation results between computation mapping nodes. These data paths and the data flowing along them together constitute the signal transmission mechanism data. During inference, when a frame of public scene inspection video payload or a real-time visual monitoring sequence is sent to the input computation mapping node of the basic multi-source perception diagnostic mapping map, the input computation mapping node performs a convolution operation on the input data to generate a multi-dimensional intermediate feature map. This intermediate feature map is then transmitted to the next downstream computation mapping node as input along the directed edges defined by the signal transmission mechanism data. To accurately simulate this data flow behavior between nodes during the reconstruction stage, this application first performs a simulated forward traversal of the topology of the basic multi-source perception diagnostic mapping map before generating the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping map. During this traversal, this application does not perform actual data computation, but instead extracts the tensor shape and data type information of the data carried by each directed edge. The data shape information includes the size of the feature map in the batch processing dimension, channel dimension, height dimension, and width dimension, while the data type information indicates that the feature map was originally stored in a 32-bit floating-point format. By recording the data shape and data type information corresponding to each directed edge between all computational mapping nodes according to the topological connection relationship, a structural description of the feature transfer information payload is obtained. The feature transfer information payload is essentially an abstract specification of the intermediate activation tensors flowing between computational mapping nodes within the basic multi-source sensing diagnostic mapping graph. It clarifies the shape and initial data precision of the data on each edge during subsequent reconstruction of the computational graph, thus providing an operational object definition for the next step of numerically adjusting these intermediate activation tensors using quantization floating-point conversion step size calibration parameters.
[0072] Preferably, the specific implementation process of adjusting the activation distribution of the feature transfer information payload using the quantization floating-point conversion step size calibration parameter in step 3 to generate the quantization scaling adapted activation feature payload is as follows. After obtaining the structural description of the feature transfer information payload and the activation-side quantization floating-point conversion step size calibration parameter corresponding to each computational mapping node calculated in the preceding steps of step 3, this application begins to actively adjust the numerical precision of the actual flowing intermediate activation tensor. Specifically, when reconstructing the computation graph according to the dynamic precision adaptive layer-by-layer quantization matching strategy, whenever a computational mapping node completes its forward computation and generates an original intermediate activation tensor, the internal elements of the original intermediate activation tensor are all 32-bit floating-point numbers. Before the original intermediate activation tensor is passed to the downstream computational mapping node, this application performs online quantization and dequantization operations on the original intermediate activation tensor according to the activation-side quantization floating-point conversion step size calibration parameter bound to the current computational mapping node in the dynamic precision adaptive layer-by-layer quantization matching strategy. The operation process is as follows: First, each floating-point element value in the original intermediate activation tensor is divided by the activation-side quantization floating-point conversion step size calibration parameter to obtain a quotient value. Then, the quotient value is rounded to the nearest integer, mapping it to the set of discrete integers that can be represented by the target activation bit width (e.g., an eight-bit integer). Next, this discrete integer is multiplied by the same activation-side quantization floating-point conversion step size calibration parameter to restore it to a floating-point format. After the above processing, the original intermediate activation tensor is converted into a new activation tensor, which is the quantization-scaling adapted activation feature payload. The quantization-scaling adapted activation feature payload is still a floating-point number in terms of data type, but its effective precision is limited to a discrete grid determined by the activation-side quantization floating-point conversion step size calibration parameter, thereby achieving precise compression of the activation distribution of the feature transfer information payload. Since different computational mapping nodes in the basic multi-source sensing diagnostic mapping graph may be assigned different activation bit widths and activation-side quantization floating-point conversion step size calibration parameters, the numerical precision of the generated quantization scaling adaptation activation feature load at the output of different computational mapping nodes is heterogeneous and adapted layer by layer. This is completely consistent with the layer-by-layer mixed precision design intent of the dynamic precision adaptive layer-by-layer quantization matching strategy.
[0073] Preferably, the specific implementation process of reconstructing the network computation graph based on the dynamic precision pruning weight distribution sequence and the quantization scaling adaptation activation feature payload in step 3 to generate the vehicle perception dynamic precision adaptive quantization inference map is as follows. After obtaining the dynamic precision pruning weight distribution sequence containing layer-by-layer compressed weight parameters and the quantization scaling adaptation activation feature payload generation mechanism that defines the layer-by-layer activation precision adjustment rules, this application recombines these two types of quantized components into a complete computation graph object that can directly perform inference. First, this application instantiates a new computation graph container in memory that has the same number of computation mapping nodes and topological connection relationship as the basic multi-source perception diagnostic map. Then, according to the topological traversal order of the basic multi-source perception diagnostic map, each computation mapping node in the new computation graph container is constructed sequentially. For the currently constructed computation mapping node, this application reads the subsequence that matches the shape of the weight tensor of the computation mapping node sequentially from the dynamic precision pruning weight distribution sequence, and reshapes the subsequence into a multi-dimensional weight tensor with the same shape as the original weight tensor, and assigns it to the new computation mapping node as its static weight parameter. Simultaneously, this application binds the activation-side quantization floating-point conversion step size calibration parameter associated with the computational mapping node as a runtime attribute during inference to the output port of the computational mapping node. During subsequent actual inference, when the computational mapping node generates the original intermediate activation tensor, the activation-side quantization floating-point conversion step size calibration parameter bound to the output port will be automatically invoked to perform the aforementioned division-rounding-multiplication operations on the original intermediate activation tensor, thereby dynamically generating the quantization scaling adaptation activation feature payload at runtime and passing it to the downstream computational mapping node. After traversing and sequentially constructing all computational mapping nodes and their connections, the obtained result includes quantization weight parameters, quantization activation generation logic, and [other parameters related to the original [mechanism]]. Figure 1 The new computational graph object for the topology is the final generated vehicle perception dynamic accuracy adaptive quantization inference mapping graph.
[0074] Preferably, during the reconstruction of the network computation graph, the collaborative operation mechanism of the dynamic precision pruning weight distribution sequence and the quantization scaling adaptation activation feature payload is reflected in every computational step of the forward inference. When a real-time visual monitoring sequence is fed into the input computational mapping node of the vehicle perception dynamic precision adaptive quantization inference mapping graph, the input computational mapping node first uses its quantized weight tensor, which has undergone stride fusion and threshold pruning, to perform a convolution operation on the input data, generating the original intermediate activation tensor. Then, the activation-side quantization floating-point conversion stride calibration parameter bound to this computational mapping node is activated, converting the original intermediate activation tensor into a quantization scaling adaptation activation feature payload, and passing this quantization scaling adaptation activation feature payload to the second computational mapping node. The second computational mapping node also uses its own quantization weight tensor to perform a convolution operation on the received quantization scaling adaptation activation feature payload, generating a new original intermediate activation tensor, and again uses its bound activation-side quantization floating-point conversion stride calibration parameter for precision compression before passing it to the next computational mapping node. The above process propagates layer by layer forward along the topological order of the vehicle perception dynamic precision adaptive quantization inference mapping map until the output computational mapping node generates the final perception diagnostic result. Since the weight parameters of each computational mapping node have been compressed in precision through dynamic precision pruning of the weight distribution sequence, and the intermediate activations generated by each computational mapping node have been adapted in precision through quantization scaling to adapt the activation feature payload, the computational intensity and memory consumption of the entire forward inference process are effectively limited within the range allowed by the adaptive computing power constraint boundary parameters on the edge side. At the same time, for shallow or deep computational mapping nodes with high information entropy and sensitivity to precision, the dynamic precision adaptive layer-by-layer quantization matching strategy preserves a high weight bit width and activation bit width, thereby maintaining the ability to discriminate key visual features.
[0075] Preferably, the generation of the vehicle perception dynamic precision adaptive quantization inference map marks the completion of the transformation process from abstract strategy to concrete executor. The vehicle perception dynamic precision adaptive quantization inference map generated in this application has a layer-by-layer heterogeneous precision characteristic. Specifically, some computational mapping nodes in the vehicle perception dynamic precision adaptive quantization inference map may use four-bit integer quantization weights combined with eight-bit integer quantization activation to compress the data flow to a greater extent when computational resources are limited; while other computational mapping nodes that have a decisive impact on the final abnormal behavior diagnosis results may be retained as a higher precision representation of sixteen-bit floating-point or even thirty-two-bit floating-point to ensure the ability to identify subtle behavioral differences. This heterogeneous precision configuration is the result of the dynamic precision pruning weight distribution sequence and the quantization scaling adaptation activation feature load being accurately set in the reconstruction stage according to the dynamic precision adaptive layer-by-layer quantization matching strategy. By unifying the static compression on the weight side (reflected in the dynamic precision pruning weight distribution sequence) and the dynamic compression on the activation side (reflected in the generation rules of quantization scaling adaptation activation feature load) within the same computation graph framework, this application achieves fine-grained, layer-by-layer control over model accuracy and computational power consumption, providing key perception model support for inspection robots to achieve long-endurance, high-reliability autonomous inspection in complex and ever-changing park environments.
[0076] In summary, this application combines dynamic precision pruning weight distribution sequences with quantization scaling adaptation activation feature loads to reconstruct the network computation graph, enabling the vehicle-mounted perception dynamic precision adaptive quantization inference map to be regenerated online as operating conditions change. When the edge-side adaptive computing power constraint boundary parameters are updated due to scene or power changes, driving step 2 to generate a new dynamic precision adaptive layer-by-layer quantization matching strategy, the reconstruction process in step 3 will be retried. The newly generated vehicle-mounted perception dynamic precision adaptive quantization inference map will immediately replace the old map and take over the subsequent real-time visual monitoring sequence inference task. This online reconstruction mechanism makes the accuracy level of the perception model no longer a fixed attribute, but a dynamic variable that fluctuates in real time with the "physical strength" of the inspection robot and the "pressure" of the scene. For example, when the inspection robot enters a sparsely populated corridor area, the newly generated vehicle-mounted perception dynamic accuracy adaptive quantization inference map can significantly reduce computing power consumption by lowering the bit width of most computing mapping nodes, thereby saving energy and extending battery life. When it enters the densely populated core area of a plaza, it can quickly switch to a high-precision version to maintain a high detection rate for abnormal behaviors (such as running or leaving items behind). This dynamic adaptive capability fundamentally overcomes the dilemma of accuracy versus power consumption faced by traditional static quantization models when deployed on resource-constrained edge devices, providing a more flexible, efficient, and sustainable perception computing solution for intelligent inspection of modern parks.
[0077] Optionally, step 4 includes: During the cycle in which the inspection robot chassis operates according to the preset inspection path rules, the visual stream data continuously collected by the vision sensor is converted to obtain a real-time visual monitoring sequence.
[0078] Preferably, the specific implementation process of acquiring the visual stream data continuously collected by the vision sensor during the cycle of the inspection robot chassis operating according to the preset inspection path rules in step 4 is as follows. When the inspection robot completes self-check and initialization and begins to execute the inspection task, its chassis motion control system loads the preset inspection path rules generated in advance through teaching or synchronous positioning and mapping technology. The preset inspection path rules define a series of path point coordinates arranged sequentially in the public scene map and the target travel speed of the inspection robot between every two adjacent path points. When the inspection robot chassis runs according to the preset inspection path rules, the vision sensor mounted on the chassis is configured to continuously collect optical images of the public scene within its field of view at a fixed acquisition frame rate. The complementary metal-oxide-semiconductor image sensor array inside the vision sensor converts the received ambient light signal into an analog voltage signal, and then quantizes it into discrete digital pixel values through an analog-to-digital converter to form a continuous sequence of raw image frames. Each raw image frame is accompanied by a timestamp information generated by the system clock, which accurately records the time when the image frame was exposed. This series of raw image frames with timestamps constitutes the visual stream data. Visual stream data is written at high speed to a circular buffer in the local memory of the inspection robot in the form of a raw pixel array. This circular buffer can store consecutive image frames from the most recent few seconds (e.g., five to ten seconds), ensuring that subsequent processing steps can obtain temporally continuous image data closely related to the current inference time. Unlike traditional methods that only extract single frames for isolated analysis, the continuous acquisition and caching of visual stream data provides the necessary temporal dimension information for extracting the motion features and temporal behavior patterns of dynamic targets.
[0079] Preferably, the specific technical processing steps in step 4 for converting the continuously acquired visual stream data from the visual sensor to obtain a real-time visual monitoring sequence include frame filtering and format standardization of the visual stream data. Since the original acquisition frame rate of the visual stream data may be higher than the inference frame rate required by the vehicle-mounted perception dynamic accuracy adaptive quantization inference map or the inference frame rate that the edge computing unit of the inspection robot can handle, directly processing all the original image frames would cause unnecessary computational consumption. Therefore, after acquiring the visual stream data, this application first extracts the target image frames corresponding to the current time window sequentially from the circular buffer according to a preset sampling step size (e.g., sampling one frame every two or three frames), forming a downsampled image frame sequence. Subsequently, for each target image frame in the downsampled image frame sequence, this application performs format standardization processing. Format standardization includes converting the color space of the original image frames from the default format of the acquisition device (e.g., "red-green-blue" color space) to the color space required by the input layer of the vehicle perception dynamic precision adaptive quantization inference map (e.g., "blue-green-red" color space), and normalizing the integer representation of image pixel values from zero to 255 to a floating-point representation with a mean of 0.5 and a standard deviation of 0.5, or directly normalizing it to a floating-point interval of zero to one. Furthermore, format standardization also includes scaling the image frame size to the fixed input size specified by the input computation mapping node of the vehicle perception dynamic precision adaptive quantization inference map using bilinear interpolation or nearest-neighbor interpolation algorithms (e.g., width 640 pixels, height 480 pixels). After the above frame selection and format standardization processes, the resulting multidimensional array, arranged in chronological order and with uniform size and normalized values, constitutes the real-time visual monitoring sequence. Each element in the real-time visual monitoring sequence, i.e., the standardized single-frame image data, constitutes the input sample for a single forward inference of the vehicle perception dynamic precision adaptive quantization inference map.
[0080] Preferably, during the process of generating a real-time visual monitoring sequence, this application also simultaneously performs image data enhancement and denoising preprocessing to improve the quality stability of the real-time visual monitoring sequence under different lighting and weather conditions. For the target image frame extracted from the visual stream data, this application first evaluates its overall average brightness and contrast distribution. When the overall average brightness is lower than a preset underexposure threshold (e.g., the average pixel value is lower than 60), this application automatically performs adaptive histogram equalization processing on the target image frame to stretch the dynamic range of the dark areas of the image and enhance the texture details hidden in the shadows; when the overall average brightness is higher than a preset overexposure threshold (e.g., the average pixel value is higher than 200), this application performs gamma correction on the target image frame, compressing the brightness of the highlight areas through nonlinear mapping to restore the contour information of the locally overexposed areas. At the same time, for the Gaussian noise and salt-and-pepper noise that visual sensors are prone to generate in low-light environments, this application uses a designed lightweight bilateral filtering algorithm to perform spatial denoising on the target image frame. This lightweight bilateral filtering algorithm considers both the spatial Euclidean distance weight between neighboring pixels and the center pixel, as well as the pixel value difference weight, when calculating the filtered output for each pixel. This allows it to smooth noise in flat areas while preserving the sharpness of dynamic target edges and contours. The target image frame, after the aforementioned enhancement and denoising preprocessing, is then fed into a format normalization process. This enables the final real-time visual monitoring sequence to overcome the interference of adverse imaging conditions such as backlighting, shadows, and insufficient nighttime lighting commonly encountered in park inspections on the accuracy of subsequent visual perception inference.
[0081] Preferably, there is a close temporal synchronization and spatial correlation between the technical processing of acquiring the real-time visual monitoring sequence in step 4 and the motion state of the inspection robot chassis. During the operation of the inspection robot chassis according to the preset inspection path rules, the odometer or inertial measurement unit continuously feeds back the real-time pose information of the chassis at the current moment, including its horizontal coordinate, vertical coordinate, and yaw angle in the global map coordinate system. This application aligns the timestamp of each target image frame in the visual stream data with the timestamp of the chassis's real-time pose information, thereby associating each frame of image data in the real-time visual monitoring sequence with a corresponding inspection robot pose label at the acquisition time. This pose label is attached to the attribute field of each frame of the real-time visual monitoring sequence in the form of metadata. In this way, the real-time visual monitoring sequence not only contains the visual representation information of the dynamic target but also implicitly includes the inspection robot's own observation perspective and spatial position context in the scene. When subsequent steps map the real-time visual monitoring sequence onto the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping map, this associated pose label can be used to assist in inferring the physical velocity and direction of motion of the dynamic target. For example, by comparing the pixel position changes of the same dynamic target in two adjacent frames and combining this with motion compensation based on the robot's own pose changes, the true motion vector of the dynamic target in three-dimensional physical space can be obtained. Unlike traditional solutions where visual perception and robot localization are independent and the perception results lack a spatial reference, this application achieves synchronous binding of visual data and pose data during the generation of the real-time visual monitoring sequence, providing a fusion perception foundation for subsequent abnormal behavior diagnosis with spatial consistency.
[0082] Preferably, the generation rate and data length of the real-time visual monitoring sequence are dynamically adjusted based on the current inference mode of the vehicle-mounted perception dynamic precision adaptive quantization inference map and the changes in the edge-side adaptive computing power constraint boundary parameters. When the edge-side adaptive computing power constraint boundary parameters constructed in step 1 indicate that the current operating condition is low power and simple scenario, the vehicle-mounted perception dynamic precision adaptive quantization inference map generated in step 3 is in a high compression ratio, low precision inference mode, and its single-frame inference time is short. Under this condition, this application correspondingly increases the sampling step size of the visual stream data, that is, reduces the generation frame rate of the real-time visual monitoring sequence (e.g., from ten frames per second to three frames per second), to further match the low computing power consumption strategy and extend the inspection time of a single driving cycle. Conversely, when the edge-side adaptive computing power constraint boundary parameters indicate a high-battery and complex operating condition, the vehicle-mounted perception dynamic precision adaptive quantization inference map switches to a low-compression, high-precision inference form. Although the single-frame inference time increases, this application correspondingly reduces the sampling step size to obtain real-time visual monitoring sequences at a higher generation frame rate (e.g., 15 frames per second), ensuring the ability to capture fast-moving or abnormal behaviors in dense crowds. Furthermore, before being input into the vehicle-mounted perception dynamic precision adaptive quantization inference map, the real-time visual monitoring sequence can be organized into a short video segment (e.g., a four-dimensional tensor composed of sixteen consecutive frames) according to a preset temporal window length. This supports the temporal convolutional layers or three-dimensional convolutional layers in the vehicle-mounted perception dynamic precision adaptive quantization inference map in modeling the behavioral coherence of dynamic targets.
[0083] In summary, the process of converting visual stream data into real-time visual monitoring sequences within the cycle of the inspection robot chassis operating according to the preset inspection path rules constitutes a crucial data bridge connecting the dynamic visual information of the physical world with the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping. This process, through a series of orderly technical processing actions such as frame filtering, format standardization, image enhancement and denoising, and pose synchronization binding, transforms the raw visual stream data, which contains a large amount of redundant information and noise, into a structured real-time visual monitoring sequence with consistent size, standardized values, enhanced quality, and spatial context labels. As the direct processing object of subsequent operations in step 4, the data quality, temporal density, and adaptability to the inspection robot's own state directly determine the accuracy and real-time performance of the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping for abnormal behavior diagnosis. Unlike traditional solutions where visual data acquisition and processing are fixed and rigid, unable to be adjusted according to task context, the generation parameters of the real-time visual monitoring sequence in this application are controllable by the real-time state of the edge-side adaptive computing power constraint boundary parameters and the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping map. This achieves a dynamic balance between resource consumption at the data acquisition end and accuracy requirements at the model inference end. This technology enables inspection robots to provide efficient and reliable visual data input adapted to the current working conditions in park inspection applications, facing complex and ever-changing scenarios from empty roads in the early morning to crowded squares at midday, thus strongly supporting the accurate generation of subsequent intelligent safety early warning decisions.
[0084] Optionally, step 4 includes: The real-time visual monitoring sequence is mapped into the vehicle perception dynamic accuracy adaptive quantization inference mapping map to generate the visual monitoring sequence to be inferred. Based on the preset calculation rules of the vehicle perception dynamic accuracy adaptive quantization inference map, the adaptive bit width forward diagnostic result of the visual monitoring sequence to be inferred is inferred, so as to output the abnormal behavior diagnostic result for the current scene. Assess the potential hazard level of the abnormal behavior diagnosis results, and then determine the intelligent safety early warning decision level for the current scenario.
[0085] Preferably, the specific implementation process of mapping the real-time visual monitoring sequence to the vehicle-mounted perception dynamic precision adaptive quantization inference map in step 4 to generate the visual monitoring sequence to be inferred is as follows. In the preceding steps of step 4, a real-time visual monitoring sequence arranged in chronological order and after format standardization and quality enhancement has been obtained. Each frame of image data in the real-time visual monitoring sequence is stored in the local memory of the inspection robot in the form of a normalized multi-dimensional floating-point array. This application first expands the batch dimension and transforms the data layout of the real-time visual monitoring sequence according to the input tensor shape requirements specified by the input computation mapping node of the vehicle-mounted perception dynamic precision adaptive quantization inference map. Specifically, if the input computation mapping node of the vehicle-mounted perception dynamic precision adaptive quantization inference map expects to receive a four-dimensional tensor, whose dimension order is batch size, number of channels, image height, and image width, then this application reshapes the current target image frame in the real-time visual monitoring sequence from a three-dimensional array (number of channels, image height, image width) into a four-dimensional array with the batch dimension added to the first dimension, and adjusts the arrangement order of the channel dimensions to be consistent with the format during the pre-training of the input computation mapping node. The four-dimensional array, after batch dimensionality expansion and data layout transformation, is fed into the input port of the vehicle perception dynamic accuracy adaptive quantization inference map. This process maps the real-time visual monitoring sequence into the vehicle perception dynamic accuracy adaptive quantization inference map. After entering the vehicle perception dynamic accuracy adaptive quantization inference map, this four-dimensional array serves as the input data for the first forward propagation and is referred to as the visual monitoring sequence to be inferred. The generation of the visual monitoring sequence to be inferred signifies that the visual observation information of the external physical world has been completely transformed into a model-understandable numerical representation, providing a directly computable data carrier for subsequent adaptive bit-width forward diagnostics.
[0086] Preferably, the specific implementation process of inferring the adaptive bit-width forward diagnostic result of the visual monitoring sequence to be inferred based on the preset calculation rules of the vehicle perception dynamic precision adaptive quantization inference map in step 4 is as follows. The preset calculation rules inside the vehicle perception dynamic precision adaptive quantization inference map are determined in step 3 when reconstructing the network computation graph. Its core feature is that the weight parameters of each computation map node have been replaced by the dynamic precision pruning weight distribution sequence, and the output port of each computation map node is bound to the activation-side quantization floating-point conversion step size calibration parameter. When the visual monitoring sequence to be inferred is sent to the first computation map node of the vehicle perception dynamic precision adaptive quantization inference map, the first computation map node uses its quantized weight tensor, which has been fused by step size and pruned by threshold, to perform a convolution operation on the visual monitoring sequence to be inferred, generating the first original intermediate activation tensor. Immediately afterwards, the activation-side quantization floating-point conversion step size calibration parameter bound to the first computation map node is triggered, and a division-rounding-multiplication operation is performed on the first original intermediate activation tensor to generate the first quantization scaling adaptation activation feature payload, and the first quantization scaling adaptation activation feature payload is passed to the second computation map node. The second computational mapping node repeats the above process, convolving its own quantization weight tensor with the first quantization scaling adapted activation feature payload to generate a second original intermediate activation tensor. This second quantization scaling adapted activation feature payload is then generated through its bound activation-side quantization floating-point conversion step size calibration parameter and passed to the third computational mapping node. This process propagates layer by layer forward along the topological connections of the vehicle perception dynamic precision adaptive quantization inference mapping graph. At each computational mapping node, the data flow undergoes a precision-controlled weighting operation and a precision-controlled activation compression. Since the quantization bit widths used by different computational mapping nodes may be heterogeneous—some nodes use four-bit integer weights and eight-bit integer activation, while others use sixteen-bit floating-point weights and sixteen-bit floating-point activation—the numerical precision experienced by the data during the entire forward propagation process is dynamically changing and adapting layer by layer. This is precisely the manifestation of adaptive bit width forward diagnostics. Finally, the output computation mapping node generates a multidimensional output tensor. Each element of this multidimensional output tensor represents the confidence score of the dynamic target contained in the visual monitoring sequence to be inferred, indicating that it belongs to a certain preset abnormal behavior category (e.g., running, fighting, people falling, leaving behind items, area intrusion, etc.), or represents the position regression value of the dynamic target's bounding box. This multidimensional output tensor is the adaptive bit-width forward diagnostic result.
[0087] Preferably, the specific implementation process of outputting the abnormal behavior diagnosis result for the current scene based on the adaptive bit-width forward diagnostic result in step 4 is as follows. The adaptive bit-width forward diagnostic result is an unprocessed raw numerical tensor. The numerical information it directly contains needs to be decoded and thresholded to be transformed into an abnormal behavior diagnosis result with clear semantics. This application first performs an output decoding operation on the adaptive bit-width forward diagnostic result. If the output calculation mapping node of the vehicle perception dynamic accuracy adaptive quantization inference mapping map contains a classification branch and a regression branch, then a flexible maximum function is applied to the output tensor of the classification branch to normalize the original confidence score to a category probability distribution between zero and one; for the output tensor of the regression branch, coordinate decoding is performed in combination with the preset anchor box parameters to obtain the bounding box coordinates of the dynamic target in the corresponding image frame of the visual monitoring sequence to be inferred. Subsequently, this application sets a confidence screening threshold (e.g., 0.5), retains all detection results with category probabilities higher than the confidence screening threshold, and uses a non-maximum suppression algorithm to filter out highly overlapping bounding boxes repeatedly predicted for the same dynamic target, resulting in a set of simplified detection target instances. Each detected target instance includes a semantic category label, bounding box coordinates, and a confidence score. For these detected target instances, this application further performs temporal association and behavior determination. If adaptive bit-width forward diagnostic results from several historical frames have been cached before the current frame, this application uses a multi-target tracking algorithm to associate the detected target instance in the current frame with the trajectory of the same dynamic target in the historical frames, thereby forming the motion trajectory of the dynamic target within a continuous time window. By calculating the kinematic parameters of this motion trajectory, such as motion speed, rate of change of motion direction, and dwell time, and comparing them with a preset normal behavior template, it is determined whether the behavior of the dynamic target constitutes an anomaly. For example, if a dynamic target dwells in a preset restricted area for more than a time threshold (e.g., ten seconds), it is determined to be an area intrusion anomaly; if the instantaneous motion speed of a dynamic target exceeds a speed threshold (e.g., eight meters per second) and the duration exceeds a duration threshold, it is determined to be a running anomaly. Information on all dynamic targets in the current frame identified as exhibiting abnormal behavior, along with their abnormal behavior category, location, and confidence score, is compiled into a structured abnormal behavior description record. This record constitutes the abnormal behavior diagnosis result for the current scenario. The abnormal behavior diagnosis result is output in a machine-readable data structure format, providing factual basis for subsequent hazard classification and early warning decisions.
[0088] Preferably, the specific implementation process of evaluating the hazard classification of the abnormal behavior diagnosis results and then determining the intelligent security early warning decision level for the current scenario in step 4 is as follows. Although the abnormal behavior diagnosis results have clearly indicated the types of abnormal behaviors present in the current scenario and their locations, the actual threat levels of different abnormal behaviors to the safety and order of the public scenario vary. Therefore, it is necessary to map the abnormal behavior diagnosis results to the corresponding level of early warning decision according to the preset hazard classification rules. This application presets a set of hazard classification rule bases, which defines the correspondence between various abnormal behavior categories and hazard levels. For example, "running" behavior is usually assigned a lower level of attention (e.g., level one warning), "person falling to the ground" behavior is assigned a higher level of emergency (e.g., level three warning), and "fighting" or "crowd scattering and fleeing" behavior is assigned the highest level of criticality (e.g., level four warning). In addition, the hazard classification also comprehensively considers the spatial area attributes of the abnormal behavior. If the abnormal behavior diagnosis results indicate that the abnormal behavior occurs in a preset core security area (e.g., the entrance of a data center, the area around a hazardous materials warehouse), its hazard level is automatically increased by one level. This application matches each abnormal behavior description record in the abnormal behavior diagnosis results with a hazard classification rule base and extracts the corresponding hazard level value for that record. If the abnormal behavior diagnosis results of the current frame contain multiple abnormal behavior description records, the highest hazard level value among them is taken as the comprehensive hazard level of the current scenario. Based on this comprehensive hazard level, this application further determines the intelligent safety early warning decision level. The intelligent safety early warning decision level defines the type and urgency of the response action that the inspection robot should take. For example, when the overall hazard level is Level 1, the intelligent safety early warning decision level can be set to "logging," meaning that only the abnormal behavior diagnosis results are stored in the local storage medium without triggering external alarms. When the overall hazard level is Level 2, the intelligent safety early warning decision level is set to "local audio-visual alert," meaning that the inspection robot's own warning lights flash and a prompt voice is played. When the overall hazard level is Level 3, the intelligent safety early warning decision level is set to "remote notification," meaning that the abnormal behavior diagnosis results and on-site image snapshots are sent to a preset remote control center via a wireless communication module. When the overall hazard level is Level 4, the intelligent safety early warning decision level is set to "emergency alarm," which, in addition to remote notification, triggers a real-time audio-visual call request with the remote control center and uploads continuous video clips. Through the above-mentioned graded evaluation mechanism, the intelligent safety early warning decision level achieves differentiated and refined responses to abnormal events, avoiding the problems of high-frequency false alarms or missed critical events caused by a single alarm threshold in traditional solutions.
[0089] Preferably, in the process of generating abnormal behavior diagnostic results and determining the intelligent safety early warning decision level, there is a dynamic matching relationship between the adaptive bit width characteristics of the vehicle perception dynamic precision adaptive quantization inference map and the scene complexity. Since the vehicle perception dynamic precision adaptive quantization inference map is reconstructed online in step 3 according to the dynamic precision adaptive layer-by-layer quantization matching strategy, the quantization bit width of each computational mapping node within it has been optimized according to the edge-side adaptive computing power constraint boundary parameters. When the inspection robot is in a scene with sparse crowds and simple behavior, the edge-side adaptive computing power constraint boundary parameters are low, and the vehicle perception dynamic precision adaptive quantization inference map is in a high compression ratio, low precision inference form. Under this form, the distribution of category confidence scores in the adaptive bit width forward diagnostic results may be relatively smooth, but since the probability of abnormal events in the scene itself is low, the lower inference precision is still sufficient to distinguish between normal behavior and obvious abnormal behavior, and the lower computing power consumption helps to extend the battery life. Conversely, when the inspection robot enters a densely populated and behaviorally complex core area, the edge-side adaptive computing power constraint boundary parameters are increased, and the vehicle-mounted perception dynamic accuracy adaptive quantization inference mapping map switches to a low-compression-ratio, high-precision inference mode. In this mode, the model's ability to discriminate subtle behavioral differences is enhanced, enabling it to identify abnormal actions such as falling, running, or pushing from individuals in crowded areas. The recall and accuracy of abnormal behavior diagnosis results are both improved.
[0090] In summary, the process of mapping real-time visual monitoring sequences to the vehicle-mounted perception dynamic precision adaptive quantization inference mapping map and ultimately determining the intelligent safety warning decision level constitutes the final execution link in the perception, decision-making, and action chain of the vehicle-mounted perception dynamic precision adaptive quantization inference method of this application. This process organically connects the edge-side adaptive computing power constraint boundary parameters constructed in step 1, the dynamic precision adaptive layer-by-layer quantization matching strategy established in step 2, the vehicle-mounted perception dynamic precision adaptive quantization inference mapping map generated in step 3, and the real-time visual monitoring sequences obtained in the preceding part of step 4, forming a complete technical path from physical world state perception to model computing power and precision adaptive adjustment, and then to specific safety event diagnosis and graded response. Unlike the loose architecture of traditional solutions with fixed perception model precision, single alarm strategy, and independent operation of each module, this application achieves an adaptive linkage mechanism of "the more complex the scene, the higher the model precision and the stronger the alarm sensitivity; the simpler the scene, the lower the model power consumption and the longer the battery life" through the above series of orderly technical processing actions. When inspection robots perform long-term inspection tasks in modern parks, this mechanism ensures that they can autonomously allocate limited onboard edge computing resources to the spatiotemporal areas that require the most precise perception, regardless of the range of scenarios, from empty parking lots in the early morning to crowded restaurant entrances at noon. They can also intervene in a way that matches the severity of the event, thereby greatly improving the overall operational efficiency, security reliability, and sustainability of long-term unattended operation of the park's intelligent inspection system.
[0091] Optionally, step 4 further includes: triggering the inspection intervention scheduling strategy based on the intelligent security early warning decision level to generate an inspection anomaly intervention scheduling notification; and transmitting the inspection anomaly intervention scheduling notification to a preset remote control center.
[0092] Preferably, the specific implementation process of triggering the inspection intervention scheduling strategy based on the intelligent safety early warning decision level to generate the inspection anomaly intervention scheduling notification in step 4 is as follows. In the preceding steps of step 4, the intelligent safety early warning decision level for the current scenario has been determined. This intelligent safety early warning decision level is stored in the decision status register of the inspection robot's local memory in the form of an enumerated value or code. This application first matches the intelligent safety early warning decision level with the preset inspection intervention scheduling strategy. The inspection intervention scheduling strategy is a set of event response rules pre-configured in the inspection robot system, where each rule defines the triggering condition, response action type, and action execution parameters. The triggering condition corresponds one-to-one with the intelligent safety early warning decision level. For example, when the intelligent safety early warning decision level is "logging," the corresponding triggering condition is a low-level event; when the intelligent safety early warning decision level is "emergency alarm," the corresponding triggering condition is a high-level event. After matching the triggering condition corresponding to the current intelligent safety early warning decision level, this application schedules the corresponding functional modules of the inspection robot to execute the preset intervention action according to the response action type associated with the triggering condition. Specifically, if the response action type is a local record, this application serializes the abnormal behavior diagnosis result generated in step 4, along with its corresponding timestamp, the current position coordinates of the inspection robot, and the target image frame identifier in the associated real-time visual monitoring sequence, into a structured inspection abnormal event record, and appends this inspection abnormal event record to the event log file in the inspection robot's local non-volatile storage medium. If the response action type involves remote notification or emergency alarm, this application, based on the above local record, further calls the wireless communication module to extract the abnormal behavior diagnosis result, abnormal behavior description record, and on-site image snapshots or short video clips from the local memory, and packages and encapsulates the above data according to a preset communication protocol (e.g., message queue telemetry transmission protocol or hypertext transfer protocol). The encapsulated data packet is the inspection abnormality intervention scheduling notification, which contains an event type identifier, an event occurrence timestamp, inspection robot identification, geographical location information, abnormal behavior category label, confidence score, and supporting image or video data payload. In this way, the generation process of inspection anomaly intervention scheduling notification completes the transformation from abstract decision level to specific transmittable data structure.
[0093] Preferably, the specific implementation process of transmitting the inspection anomaly intervention scheduling notification to the preset remote control center in step 4 is as follows. After generating the inspection anomaly intervention scheduling notification, this application calls the wireless communication module built into the inspection robot chassis. This wireless communication module supports establishing a network connection with the preset remote control center located in a remote monitoring room or cloud platform through the park's wireless local area network or cellular mobile communication network. This application first queries the current network connection status of the wireless communication module. If the network connection is disconnected, the inspection anomaly intervention scheduling notification is temporarily stored in the local transmission queue, and a reconnection and retransmission mechanism is initiated. After the network is restored, the inspection anomaly intervention scheduling notifications in the queue are sent sequentially in a first-in-first-out order. If the network connection is normal, this application directly uses the inspection anomaly intervention scheduling notification as the payload and sends it to the network address and port of the preset remote control center through the established secure transmission layer encrypted channel (e.g., a transmission control protocol connection encrypted with a transmission layer security protocol). Upon receiving an abnormal inspection intervention dispatch notification, the pre-set remote control center's internal monitoring software parses, verifies, and visualizes the notification data packet. For example, it highlights the inspection robot icon exhibiting abnormal behavior on an electronic map interface and displays a pop-up window showing the abnormal behavior category, on-site images, and suggested handling measures. Simultaneously, if the response action triggered by the intelligent security early warning decision level includes on-site audio-visual intervention, this application simultaneously drives the audio-visual alarm device on the inspection robot chassis while sending the abnormal inspection intervention dispatch notification. This controls the warning lights to emit visible light signals at a preset flashing frequency and controls the speaker to play a preset warning voice corresponding to the abnormal behavior category (e.g., "Please note, there is abnormal behavior in the area; please stop immediately"). Through the aforementioned transmission and local linkage intervention actions, this application achieves real-time and reliable transmission of the inspection robot's perception, diagnosis, and decision-making results to the pre-set remote control center monitored by human security personnel, while simultaneously issuing warnings to on-site personnel. This constitutes a complete automated security response chain from visual data acquisition, adaptive precision inference, abnormal behavior diagnosis, hazard level assessment, to remote alarms and on-site intervention. This mechanism differs from the traditional approach where inspection robots merely serve as mobile image acquisition terminals and security decisions rely entirely on manual analysis at the back end, significantly reducing the time delay from the occurrence of an abnormal event to the intervention and response of security forces, and improving the efficiency of modern intelligent park inspection systems in handling sudden security incidents.
[0094] like Figure 2 As shown, this is an embodiment of an onboard perception dynamic accuracy adaptive quantization inference device, which includes: The first program unit is used to construct adaptive computing power constraint boundary parameters for the inspection robot in its current operating state based on the current endurance status representation parameters of the inspection robot and the corresponding public scene inspection field activity distribution characteristics. The second program unit is used to establish a dynamic accuracy adaptive layer-by-layer quantization matching strategy for real-time behavior detection based on the basic multi-source perception diagnostic mapping map pre-built for the inspection scenario and the edge-side adaptive computing power constraint boundary parameters. The third program unit is used to generate a vehicle perception dynamic precision adaptive quantization inference mapping map based on the dynamic precision adaptive layer-by-layer quantization matching strategy. The fourth program unit is used to acquire the real-time visual monitoring sequence of the inspection robot, and to determine the intelligent safety early warning decision level for the current scenario by mapping the real-time visual monitoring sequence to the vehicle perception dynamic accuracy adaptive quantization inference mapping map.
[0095] like Figure 3 As shown, an electronic device according to an embodiment of this application includes a memory and a processor. The memory stores a computer program, and the processor is used to execute the computer program to implement the method of any one of the embodiments of this application.
[0096] Figures 2-3 For an exemplary description, please refer to the above. Figure 1 This will not be elaborated upon here.
Claims
1. An in-vehicle perception dynamic precision adaptive quantization inference method, characterized in that, include: Step 1: Based on the current endurance status representation parameters of the inspection robot and the activity distribution characteristics of the inspection field of view in the corresponding public scene, construct adaptive computing power constraint boundary parameters for the inspection robot in its current operating state at the edge. Step 2: Based on the pre-built basic multi-source perception diagnostic mapping map for the inspection scenario and the edge-side adaptive computing power constraint boundary parameters, establish a dynamic accuracy adaptive layer-by-layer quantization matching strategy for real-time behavior detection. Step 3: Based on the dynamic precision adaptive layer-by-layer quantization matching strategy, generate the vehicle perception dynamic precision adaptive quantization inference mapping map. Step 4: Obtain the real-time visual monitoring sequence of the inspection robot, and determine the intelligent safety early warning decision level for the current scenario by mapping the real-time visual monitoring sequence to the vehicle perception dynamic accuracy adaptive quantization inference mapping map.
2. The method of claim 1, wherein, Step 1 includes: Obtain the real-time topological profile of the LiDAR environment from the LiDAR sensors deployed on the chassis of the inspection robot; The inspection robot acquires a public scene inspection video payload synchronously captured by the vision sensor mounted on its chassis. Align the lidar environment topology with the public scene inspection video payload to construct a multi-source sensing synchronous data array; Based on the multi-source sensing synchronous data array, the correlation features between dynamic target aggregation and abnormal activity trajectories in public scenes are determined, thereby generating the activity distribution features of the inspection field of view.
3. The method of claim 2, wherein, Step 1 also includes: Obtain the current battery life status parameters fed back by the battery monitoring system built into the chassis of the inspection robot; By integrating the current endurance status representation parameters with the inspection field activity distribution characteristics, an edge-side adaptive computing power constraint boundary parameter is constructed for the current operating state.
4. The method of claim 1, wherein, Step 2 includes: Using the edge-side adaptive computing power constraint boundary parameters as the basis for inference accuracy adjustment, the association of computational mapping nodes inside the basic multi-source perception diagnostic mapping map is decoupled to determine the layer-level mutation candidate mapping level of the basic multi-source perception diagnostic mapping map under the current computing power constraint; Based on the layer-level mutation candidate mapping hierarchy, a dynamic accuracy-adaptive layer-by-layer quantization matching strategy for real-time behavior detection is established.
5. The method of claim 4, wherein, Based on the layer-level mutation candidate mapping hierarchy, a dynamic accuracy-adaptive layer-by-layer quantization matching strategy for real-time behavior detection is established, including: The quantization bit width response mapping range of the layer-level mutation candidate mapping level is detected, and then the quantization accuracy compression mapping node parameters of each layer-level mutation candidate mapping level under the baseline operating conditions are determined. The quantization precision compression mapping node parameters are matched with the edge-side adaptive computing power constraint boundary parameters to establish a dynamic precision adaptive layer-by-layer quantization matching strategy for real-time behavior detection.
6. The method according to claim 1, characterized in that, Step 3 includes: The configuration parameters of the dynamic precision adaptive layer-by-layer quantization matching strategy are analyzed to obtain the quantization floating-point conversion step size calibration parameter; Extract the underlying structure sequence of the basic multi-source sensing diagnostic mapping map to obtain the initial weight parameter distribution sequence; The quantization floating-point conversion step size calibration parameter is injected into the initial weight parameter distribution sequence to generate a step size fusion weight distribution sequence; Determine the activation response threshold of the step-size fusion weight distribution sequence to generate a dynamic precision-trimmed weight distribution sequence; Based on the dynamic precision pruning weight distribution sequence, an adaptive quantization inference mapping map for vehicle perception dynamic precision is generated.
7. The method according to claim 6, characterized in that, Based on the dynamic precision pruning weight distribution sequence, an adaptive quantization inference mapping map for vehicle perception dynamic precision is generated, including: Extract the signal transmission mechanism data within the basic multi-source sensing diagnostic mapping map to obtain the feature transmission information payload; The activation distribution of the feature transfer information payload is adjusted using the quantization floating-point conversion step size calibration parameter to generate a quantization scaling adapted activation feature payload. The network computation graph is reconstructed based on the dynamic precision pruning weight distribution sequence and the quantization scaling adaptation activation feature load to generate an on-board perception dynamic precision adaptive quantization inference mapping graph.
8. The method according to claim 1, characterized in that, Step 4 includes: During the cycle in which the inspection robot chassis operates according to the preset inspection path rules, the visual stream data continuously collected by the vision sensor is converted to obtain a real-time visual monitoring sequence.
9. The method according to claim 1, characterized in that, Step 4 includes: The real-time visual monitoring sequence is mapped into the vehicle perception dynamic accuracy adaptive quantization inference mapping map to generate the visual monitoring sequence to be inferred. Based on the preset calculation rules of the vehicle perception dynamic accuracy adaptive quantization inference map, the adaptive bit width forward diagnostic result of the visual monitoring sequence to be inferred is inferred, so as to output the abnormal behavior diagnostic result for the current scene. Assess the potential hazard level of the abnormal behavior diagnosis results, and then determine the intelligent safety early warning decision level for the current scenario.
10. The method according to claim 9, characterized in that, Step 4 further includes: triggering the inspection intervention scheduling strategy based on the intelligent security early warning decision level to generate an inspection anomaly intervention scheduling notification; and transmitting the inspection anomaly intervention scheduling notification to a preset remote control center.