Inspection model training method, inspection method, inspection device and electronic equipment

By using a robot dog to collect and annotate video data in complex environments, and combining environmental and operational status information to train a detection model, the problem of insufficient generalization ability of visual detection models in non-standard perspectives and complex dynamic environments is solved, thereby improving detection accuracy and adaptability.

CN121999313APending Publication Date: 2026-05-08INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
Filing Date
2025-12-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing visual inspection models lack generalization ability when facing non-standard perspectives and complex dynamic environments, resulting in missed detections and false detections during park inspections.

Method used

By driving a robot dog to perform mobile inspections under multiple different inspection tasks, sample video data from the target perspective is collected and metadata is recorded. Target detection labels are annotated and dynamically enhanced. The detection model is trained by combining environmental information and operating status information to generate an inspection model.

Benefits of technology

It significantly improves the detection accuracy and generalization ability of the inspection model in edge scenarios such as large changes in lighting and harsh environments, and reduces the false alarm rate and missed detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999313A_ABST
    Figure CN121999313A_ABST
Patent Text Reader

Abstract

The invention provides an inspection model training method, an inspection method, an inspection device and electronic equipment, and the method comprises the steps: driving a robot dog to carry out the mobile inspection of a sample park under a plurality of different inspection tasks, and obtaining the sample video data of a target visual angle and the metadata with synchronous time sequence; the metadata comprises running state information and environment information of the robot dog; performing target detection label labeling on sample images in the sample video data to obtain a first data set; according to the operation state information, performing dynamic enhancement on each sample image in the first data set to obtain a second data set; and according to the environment information, obtaining a training weight corresponding to each sample image in the second data set, and according to the training weight and a target detection loss value corresponding to each sample image in the second data set output by the initialization detection model, training the initialization detection model to obtain an inspection model. According to the invention, the detection precision and adaptability in a low-view-angle and complex dynamic environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring technology, and in particular to a patrol model training method, patrol method, device and electronic equipment. Background Technology

[0002] With the continuous improvement of the level of intelligent management in smart parks, how to apply visual inspection models to conduct intelligent inspections of parks is an important issue that the industry urgently needs to address.

[0003] Existing visual detection models are typically built upon general-purpose deep learning object detection algorithms. To obtain a detection model, current technologies usually directly collect video streams from fixed security surveillance cameras within the park as training samples, or utilize publicly available general-purpose image datasets for training, aiming to obtain a detection model capable of recognizing various targets within the park.

[0004] However, since fixed surveillance cameras are usually located at high places and are stationary, the images they capture have a single perspective, stable background, and high target integrity, making it difficult for detection models to fully learn feature representations in complex environments. This results in insufficient generalization ability of the detection models when facing non-standard perspectives and complex dynamic environments, making it difficult to effectively output accurate detection results. Consequently, problems such as missed detections and false detections occur in actual park inspections. Summary of the Invention

[0005] This invention provides a method for training an inspection model, an inspection method, an apparatus, and an electronic device to address the shortcomings of existing detection models in terms of insufficient generalization ability when facing non-standard perspectives and complex dynamic environments, making it difficult to effectively output accurate detection results. This invention aims to improve the detection accuracy and adaptability under non-standard perspectives and complex dynamic environments.

[0006] This invention provides a method for training an inspection model, comprising: The robot dog is driven to perform mobile inspections of the sample area under multiple different inspection tasks, obtaining sample video data from the target perspective, as well as metadata synchronized with the sample video data in time. The metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective. The sample images in the sample video data are labeled with target detection tags to obtain the first dataset; Based on the running status information, the sample images in the first dataset are dynamically enhanced to obtain the second dataset; Based on the environmental information, the training weights corresponding to each sample image in the second dataset are obtained. Based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model, the initial detection model is trained to obtain the inspection model.

[0007] According to the inspection model training method provided by the present invention, the step of obtaining the training weights corresponding to each sample image in the second dataset based on the environmental information includes: Based on the environmental information, obtain the scene complexity and / or lighting conditions of each sample image in the second dataset; Based on the scene complexity and / or the lighting conditions, obtain the training weights corresponding to each sample image in the second dataset.

[0008] According to the present invention, a method for training an inspection model includes dynamically enhancing each sample image in the first dataset based on the running state information to obtain a second dataset, comprising: Based on the running status information, determine the affine transformation parameters and motion blur kernel parameters of each sample image in the first dataset; Based on the affine transformation parameters and the motion blur kernel parameters, data augmentation is performed on each sample image in the first dataset to obtain the second dataset.

[0009] According to a method for training an inspection model provided by the present invention, the first dataset is obtained by performing target detection label annotation on sample images in the sample video data, comprising: The sample video data is preprocessed to obtain a keyframe sequence; the preprocessing includes segmentation and archiving, resolution normalization, and deduplication. The keyframe sequence is labeled with target detection tags based on the target detection model to obtain simulated tags for each sample image in the keyframe sequence; The simulated tags are sent to each client. Upon receiving a tag correction instruction from any of the clients for the simulated tag, the simulated tag is corrected according to the tag correction instruction; The corrected simulated tag is sent back to each of the clients again, and the tag verification is performed iteratively until a verification pass instruction is received from all the clients. The corrected simulated label corresponding to the verification pass instruction is determined as the real label of each sample image in the keyframe sequence; The first dataset is constructed based on each sample image in the keyframe sequence and the real labels.

[0010] According to the inspection model training method provided by the present invention, the metadata also includes the location information of the robot dog; The steps for obtaining the target detection loss value include: Each sample image in the second dataset is input into the feature extraction network of the initial detection model to obtain multiple feature maps of different scales for each sample image in the second dataset. Based on the location information, determine the weight coefficients of the feature maps at each scale for each sample image in the second dataset; Based on the weighting coefficients, feature maps of multiple scales of each sample image in the second dataset are fused to obtain a fused feature map. The fused feature map is input into the detection network of the initial detection model to obtain the target detection results of each sample image in the second dataset; Based on the target detection results and the true labels of each sample image in the second dataset, obtain the target detection loss value corresponding to each sample image in the second dataset.

[0011] According to the inspection model training method provided by the present invention, the multiple different inspection tasks include multiple inspection tasks covering different areas, multiple inspection tasks covering different inspection time periods, and multiple inspection tasks covering different inspection environments.

[0012] The present invention also provides an inspection method, comprising: Acquire the image to be inspected from the robot dog's current inspection perspective; The image to be detected is input into the inspection model to obtain the target detection result of the image to be detected; When the target detection result indicates the presence of an abnormal target, an alarm is issued according to the type of the abnormal target; The inspection model is trained based on any of the inspection model training methods described above.

[0013] The present invention also provides an inspection model training device, comprising: The first data acquisition module is used to drive the robot dog to perform mobile inspections of the sample area under multiple different inspection tasks, obtain sample video data from the target perspective, and metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective. The first data management module is used to perform target detection labeling on sample images in the sample video data to obtain the first dataset; The second data management module is used to dynamically enhance each sample image in the first dataset according to the running status information to obtain the second dataset; The optimization module is used to obtain the training weights corresponding to each sample image in the second dataset based on the environmental information, and to train the initial detection model based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model to obtain the inspection model.

[0014] The present invention also provides an inspection device, comprising: The second data acquisition module is used to acquire the image to be inspected from the robot dog's current inspection perspective. The detection module is used to input the image to be detected into the inspection model to obtain the target detection result of the image to be detected; The alarm module is used to issue an alarm based on the type of the abnormal target when the target detection result indicates the presence of an abnormal target; The inspection model is trained based on any of the inspection model training methods described above.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the inspection model training method or inspection method as described above.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the inspection model training method or inspection method as described above.

[0017] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the inspection model training method or inspection method as described above.

[0018] The inspection model training method, inspection method, device, and electronic equipment provided by this invention solve the problem that existing general models cannot adapt to the unique low-angle and dynamic motion characteristics of robot dogs by driving a robot dog to collect low-view sample video data in real inspection scenarios and simultaneously recording metadata. Furthermore, by using operational status information to guide dynamic data enhancement, the visual interference experienced by the robot dog during movement can be accurately simulated, improving the model's anti-shake and anti-distortion capabilities. Simultaneously, by dynamically adjusting the training weights of samples using environmental information, targeted optimization training for difficult samples in complex environments is achieved, significantly improving the detection accuracy and generalization ability of the final inspection model in edge scenarios such as large changes in lighting and harsh environments, thereby effectively reducing the false negative and false positive rates in park inspections. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the inspection model training method provided by the present invention.

[0021] Figure 2 This is a flowchart illustrating the inspection method provided by the present invention.

[0022] Figure 3 This is a schematic diagram of the inspection model training device provided by the present invention.

[0023] Figure 4 This is a schematic diagram of the inspection device provided by the present invention.

[0024] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] With the continuous improvement of intelligent management in smart parks, mobile inspection robots (such as robot dogs) are widely used in daily security inspections, violation identification, and environmental monitoring. Compared to traditional cameras in fixed locations, robot dogs use cameras with lower viewing angles during inspections and are in a dynamic, moving state. Therefore, the video data they collect often has characteristics such as large changes in the proportion of target objects, numerous occlusions, and complex lighting conditions. Existing target detection models are mainly trained on data from fixed cameras, making it difficult to effectively output accurate detection results in complex dynamic environments and non-standard viewing angles in low-angle scenes. This leads to problems such as missed detections and false detections in actual park inspections, failing to fully utilize the automated recognition capabilities of mobile inspection robots.

[0027] To address this deficiency, this application provides a method for training an inspection model.

[0028] It should be noted that the execution subject of this method can be an inspection model training device, which can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, super mobile personal computers, netbooks, or personal digital assistants, etc., while non-mobile electronic devices can be servers, network attached storage devices, personal computers, etc. This invention does not impose specific limitations.

[0029] Figure 1 This is a flowchart illustrating the inspection model training method provided by the present invention. Figure 1 As shown, the method includes steps 110, 120, 130 and 140.

[0030] Step 110: Drive the robot dog to perform mobile inspections of the sample area under multiple different inspection tasks to obtain sample video data from the target perspective, as well as metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective.

[0031] Optionally, in this embodiment, the robot dog refers to a biomimetic quadruped robot with autonomous or remote-controlled mobility, capable of moving on complex terrain (such as steps, grass, and unpaved surfaces) within the park. The sample park refers to the physical location used to collect training data, which can be a real industrial park, logistics and warehousing area, office park, etc., or a specially constructed simulated test field.

[0032] To ensure the diversity and generalization ability of the training data, the robot dog needs to perform multiple different inspection tasks during the sample data collection process. This allows the robot dog to collect video data in multiple time periods and from all directions according to the set route and different inspection tasks, ensuring comprehensive coverage of the sample data.

[0033] The multiple inspection tasks here can be understood as tasks with differentiated combinations in terms of space, time, or task objectives. Since different regions provide rich negative samples of background textures, which helps reduce overfitting of the model to specific backgrounds, different inspection time periods ensure the model's adaptability to changes in lighting at different times, and different inspection environments ensure the model's adaptability to changes in different environments. Therefore, in one possible implementation, multiple inspection tasks include multiple inspection tasks covering different areas (such as office areas, warehouse areas, and perimeter walls), multiple inspection tasks covering different inspection time periods, and multiple inspection tasks covering different inspection environments (such as daytime, nighttime, rain / snow, occlusion, etc.), to ensure that the sample data has sufficient diversity and representativeness.

[0034] During the inspection, the robot dog uses one or more onboard data acquisition devices to capture real-time video of the sample area, thereby obtaining sample video data from the target perspective. It should be noted that the target perspective in this application specifically refers to the robot dog's own perspective (also known as a low-angle view). Unlike traditional fixed high-position surveillance cameras, the robot dog's camera is typically installed at a lower height (e.g., between 30cm and 80cm above the ground), and its height dynamically changes with the robot dog's movement. Therefore, the sample video data collected in this way is characterized by low shooting angles, cluttered backgrounds, and frequent ground obstructions.

[0035] In addition, while collecting sample video data, metadata synchronized with the sample video data in time is also recorded. Time synchronization means that each frame or segment of video data is precisely associated with metadata at a specific moment through a timestamp, so as to ensure that subsequent processing can know in what operating state and environment the image was captured under.

[0036] The metadata refers to data describing the attributes of the sample video data. In this embodiment, the metadata includes at least operational status information and environmental information. The operational status information characterizes the physical motion state of the robot dog during filming, and may include, but is not limited to, the robot dog's movement speed, acceleration, attitude angles (such as pitch, roll, and yaw angles), and joint gait information. This information reflects the degree of camera shake and viewing angle tilt of the camera mounted on the robot dog. The environmental information characterizes the objective conditions of the external environment in which the robot dog is located, and may include, but is not limited to, lighting information, weather conditions, the inspection route it is on, and location information.

[0037] Step 120: Target detection labels are added to the sample images in the sample video data to obtain the first dataset.

[0038] Optionally, after data acquisition is completed, the acquired sample video data can be uniformly stored and formatted for management.

[0039] Specifically, the sample video data collected by the robot dog during inspections is typically a continuous, raw video stream containing a large amount of redundant information and noise. Therefore, preprocessing of the sample images in the video data is necessary. This preprocessing includes, but is not limited to, segmentation and archiving, data cleaning, resolution normalization, and frame extraction to ensure the consistency and controllability of the input during subsequent training. Data cleaning includes, but is not limited to, removing blurry or unrecognizable images, removing completely duplicate static images, and eliminating invalid data where the camera lens is completely obscured, in order to improve data quality.

[0040] Subsequently, target detection labels are applied to each frame of the preprocessed sample images to obtain the true labels for each sample image. The labeled target objects are abnormal targets of interest in the inspection task, such as one or more of the following: flames, smoke, personnel not wearing safety helmets, and illegally placed objects. The labeled content may include the category of the target object and its position coordinates in the image, which is not specifically limited in this embodiment. The labeling method can be manual or a semi-automatic labeling method using pre-trained models for automatic labeling and manual review, which is not specifically limited in this embodiment.

[0041] The preprocessed and labeled sample images and their corresponding ground truth labels together constitute the first dataset.

[0042] Step 130: Based on the running status information, dynamically enhance each sample image in the first dataset to obtain the second dataset.

[0043] Optionally, since the robot dog is in a dynamic state during actual inspections, the captured images often suffer from motion blur and perspective distortion. To enable the trained inspection model to adapt to these dynamic features, the running status information recorded synchronously during data acquisition can be used to perform targeted dynamic enhancement processing on the first dataset.

[0044] Specifically, by utilizing the robot dog's operational status information when capturing each frame of sample images in the first dataset, specific distortion features of each frame are simulated or enhanced to generate training samples that more closely resemble harsh working conditions. For example, for any frame of sample image in the first dataset, if the operational status information corresponding to that frame indicates the robot dog is running at high speed or vibrating at high frequency, motion blur or Gaussian noise processing can be applied to that sample image to simulate image shake during actual inspections. Furthermore, if the operational status information corresponding to that frame indicates the robot dog is climbing a slope or tilting (i.e., a large pitch or roll angle), affine transformation techniques can be used to rotate, crop, or perform perspective transformations on that frame to simulate drastic changes in perspective.

[0045] This data augmentation, driven by real-world physical states, generates a second dataset containing more challenging samples that are difficult to capture during actual data acquisition. The sample images in this second dataset not only include the original target features but also incorporate motion-related disturbance features unique to the robot dog, thereby improving the model's robustness in dynamic scenes.

[0046] Step 140: Based on the environmental information, obtain the training weights corresponding to each sample image in the second dataset, and train the initial detection model based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model to obtain the inspection model.

[0047] Optionally, in order to address the problem of low recognition rate in complex environments such as nighttime, backlighting, and inclement weather, this embodiment introduces a training weight adjustment mechanism based on environmental information to train the model, so as to obtain an inspection model that can accurately identify target objects in non-standard perspectives and complex dynamic environments.

[0048] Specifically, an initial detection model is constructed; this initial detection model refers to a pre-constructed deep learning network model that has not yet undergone targeted training or pre-training on a general dataset, such as the YOLO series, Faster R-CNN, and other architectures.

[0049] Furthermore, based on environmental information, training weights are obtained for each sample image in the second dataset. These training weights refer to the contribution of the loss value generated by the sample image to the model parameter updates during the model's loss function calculation. Obtaining training weights is essentially a process of reweighting the importance of samples based on environmental difficulty. Specifically, the weights of each sample image are dynamically set according to the environmental information corresponding to it. For example, sample images with low scene complexity or good lighting conditions can be assigned lower training weights; while sample images with high scene complexity or poor lighting conditions can be assigned higher training weights. This forces the model to pay more attention to those difficult-to-identify harsh environmental samples during training, thus preventing the model from performing well only in simple scenes and failing in complex scenes.

[0050] During iterative training, sample images from the second dataset are input into the initialization detection model, which then outputs object detection results for each sample image in the second dataset. Based on the difference between the object detection results and the true labels for each sample image in the second dataset, the object detection loss value for each sample image in the second dataset can be calculated.

[0051] Next, the target detection loss value corresponding to each sample image in the second dataset is weighted using the training weights of each sample image in the second dataset. Based on the weighted loss value, the network parameters of the initial detection model are updated through multiple rounds of iterative training using the backpropagation algorithm and gradient descent optimizer until the trained initial detection model converges or reaches the preset accuracy index. After the model training is completed, the trained model should be tested for multiple rounds of performance using offline validation data covering different scenarios and conditions. If the test fails, the trained model is used as a new initial detection model and returned to step 110 for training until the test is passed. The trained model that passes the test is finally determined as the inspection model to ensure that it can meet the stability requirements of actual inspection.

[0052] It should be noted that the training process of deep learning models involves a massive amount of parameter calculation and high-dimensional matrix operations, resulting in an extremely high computational load. Therefore, in executing the training process of this embodiment, it is necessary to allocate computing resources reasonably. Specifically, the server or computing cluster executing the training task should make full use of hardware acceleration capabilities such as Graphics Processing Units (GPUs), Tensor Processing Units (TPUs), or dedicated intelligent acceleration chips. By adopting strategies such as data parallelism or model parallelism, training efficiency can be maximized while ensuring training accuracy, thereby shortening the model development cycle and meeting the needs of park inspection operations for rapid model iteration and updates.

[0053] The method provided in this embodiment solves the problem that existing general models cannot adapt to the unique low-angle and dynamic motion characteristics of robot dogs by driving a robot dog to collect low-angle sample video data in real inspection scenarios and simultaneously recording metadata. Furthermore, by using operational status information to guide dynamic data enhancement, it can accurately simulate visual interference during robot dog movement, improving the model's anti-shake and anti-distortion capabilities. Simultaneously, by dynamically adjusting the training weights of samples using environmental information, it achieves targeted optimization training for difficult samples in complex environments, significantly improving the detection accuracy and generalization ability of the final inspection model in edge scenarios such as large changes in lighting and harsh environments, thereby effectively reducing the false negative and false positive rates in park inspections.

[0054] In some embodiments, step 140 specifically includes: Based on the environmental information, obtain the scene complexity or lighting conditions of each sample image in the second dataset; Based on the scene complexity or the lighting conditions, obtain the training weights corresponding to each sample image in the second dataset.

[0055] Optionally, in practical robot dog inspection scenarios, environmental factors mainly interfere with target detection in terms of visual perception clarity and background interference. To enable the trained inspection model to better cope with these challenges, this embodiment provides a weight allocation strategy based on fine-grained environmental features.

[0056] Specifically, based on the environmental information of each sample image in the first dataset, the scene complexity or lighting conditions of each sample image in the second dataset are dynamically extracted. Lighting conditions are an indicator reflecting the brightness and quality of the light distribution in the image; they can be obtained through parsing from the environmental information or calculated. Scene complexity is an indicator reflecting the clutter of the image background and the interference with target recognition; it can be obtained by calculating the scene complexity from the parameters in the environmental information or by directly parsing from the environmental information, etc. This embodiment does not specifically limit this.

[0057] After quantifying the scene complexity or lighting conditions of each sample image in the second dataset, the training weights corresponding to each sample image in the second dataset can be obtained by associating the relationships between different scene complexities and different training weights, as well as the relationships between different lighting conditions and different training weights.

[0058] It should be noted that when establishing the mapping relationship between indicators (i.e., lighting conditions or scene complexity) and training weights, the core logic is that the greater the difficulty of sample recognition, the higher the training weight is assigned, so as to force the model to pay more attention to the error generated by the sample during backpropagation.

[0059] For example, when assigning weights based on lighting conditions, samples with extremely poor lighting conditions (such as being too dark or too exposed) are difficult to identify. Therefore, a baseline lighting range can be set. The training weights corresponding to sample images falling within this range are set as the baseline values. For sample images deviating from the baseline range, the weights increase linearly or non-linearly as the degree of deviation increases. For example, when the lighting intensity of a sample image is L < 50 Lux, the weight W = 1.0 + α × (50 - L), where α is a preset coefficient.

[0060] When assigning weights based on scene complexity, the more complex the scene, the higher the probability of false detection. Therefore, the scene complexity of the sample image can be divided into corresponding scene levels, and the weights set for the corresponding scene levels can be determined as the training weights for that sample image.

[0061] When assigning weights based on a combination of scene complexity and lighting conditions, the weights corresponding to the scene complexity and lighting conditions of the sample image can be weighted and summed, or the maximum value can be taken to determine the training weights of the sample image.

[0062] The method provided in this embodiment implements an adaptive sample reweighting training strategy. This strategy avoids the problem of simple samples dominating the gradient direction in traditional training, enabling the model to automatically focus on long-tail scenarios in robot dog inspections, such as low-light conditions and highly cluttered backgrounds. This not only improves the feature extraction capability of the inspection model under harsh working conditions but also effectively balances the detection performance under different scenarios, thereby significantly enhancing the model's generalization ability and reliability in real-world complex park environments.

[0063] In some embodiments, step 130 specifically includes: Based on the running status information, determine the affine transformation parameters and motion blur kernel parameters of each sample image in the first dataset; Based on the affine transformation parameters and the motion blur kernel parameters, data augmentation is performed on each sample image in the first dataset to obtain the second dataset.

[0064] Optionally, during the actual inspection process of the robot dog, its camera is rigidly or semi-rigidly connected to the robot dog body. Unlike wheeled robots or fixed monitoring systems, the robot dog's movement is accompanied by periodic bumps and tilting on unstructured surfaces. To enable the trained inspection model to be anti-shake and viewpoint robust, this embodiment proposes a data augmentation method based on physical state-driven principles.

[0065] Specifically, based on the operational status information, the affine transformation parameters and motion blur kernel parameters of each sample image in the first dataset are first obtained. The operational status information includes the robot dog's posture and motion data at the time of image acquisition, such as velocity and acceleration.

[0066] Affine transformation parameters describe the geometric deformation of an image and typically include at least one of rotation matrices, translation vectors, scaling factors, and shearing factors. During the determination of affine transformation parameters, runtime status information synchronized with the current sample image can be read from the metadata. Based on the robot dog's motion state displayed in the runtime status information, the affine transformation parameters are dynamically determined. Different motion states correspond to different calculation methods for the affine transformation parameters. For example, if the runtime status information shows the robot dog is climbing a slope (i.e., with a large pitch angle), a corresponding perspective transformation matrix or a vertical shearing parameter is calculated to simulate perspective distortion caused by the camera looking up or down. If the runtime status information shows the robot dog is tilting (i.e., with a large roll angle), a corresponding rotation matrix is ​​calculated to simulate rotation correction of the sample image. In this way, the generated affine transformation parameters can realistically reflect the viewing angle deviations that may occur during the robot dog's movement.

[0067] The motion blur kernel parameters mainly include the kernel size and blur direction. These parameters determine the intensity and shape of the simulated blur effect. During the determination of the motion blur kernel parameters, runtime status information synchronized with the current sample image can be read from the metadata. Based on the robot dog's motion state displayed in the runtime status information, the motion blur kernel parameters are dynamically determined. Different motion states correspond to different motion blur kernel parameters. For example, the faster the robot dog runs, the larger the blur kernel size, and the blur direction is consistent with the robot dog's motion vector direction.

[0068] After determining the above parameters, the determined affine transformation parameters can be used to perform spatial transformation on the image pixel matrix of each sample image in the first dataset to simulate the change of viewpoint; the determined motion blur kernel parameters can be used to perform convolution operation with each sample image in the first dataset to simulate motion blur.

[0069] The second dataset, generated after data augmentation, not only contains the original clear samples but also augmented samples that simulate various real motion disturbances.

[0070] The method provided in this embodiment abandons the blindness of traditional random data augmentation and instead adopts a targeted augmentation strategy for running status information. By mapping the robot dog's real running status information to image augmentation parameters, the second dataset can highly restore the visual characteristics of the robot dog when walking in complex terrain, such as severe image shaking and tilted horizon. As a result, the inspection model trained on this dataset can effectively learn the target invariance under these interference features, thereby significantly improving the recognition stability during dynamic inspection.

[0071] In some embodiments, step 120 specifically includes: The sample video data is preprocessed to obtain a keyframe sequence; the preprocessing includes segmentation and archiving, resolution normalization, and deduplication. The keyframe sequence is labeled with target detection tags based on the target detection model to obtain simulated tags for each sample image in the keyframe sequence; The simulated tags are sent to each client. Upon receiving a tag correction instruction from any of the clients for the simulated tag, the simulated tag is corrected according to the tag correction instruction; The corrected simulated tag is sent back to each of the clients again, and the tag verification is performed iteratively until a verification pass instruction is received from all the clients. The corrected simulated label corresponding to the verification pass instruction is determined as the real label of each sample image in the keyframe sequence; The first dataset is constructed based on each sample image in the keyframe sequence and the real labels.

[0072] It should be noted that, considering the large amount of low-view video data and the complex background, purely manual annotation is inefficient and purely automatic annotation is inaccurate. This embodiment provides an iterative data annotation scheme that couples human and machine.

[0073] Specifically, the sample video data is first processed sequentially through segmentation and archiving, resolution normalization, and deduplication to obtain a keyframe sequence. Segmentation and archiving divides long video streams into shorter segments based on time periods or task identifiers, facilitating distributed processing and storage. Resolution normalization adjusts all video frames to a uniform resolution to meet the subsequent model input requirements, given the different types of acquisition devices. Deduplication addresses the issue of numerous similar frames generated when the robot dog stops or moves slowly. By calculating the structural similarity or histogram difference between adjacent frames, images with a repetition rate exceeding a preset threshold are removed, retaining only representative keyframes.

[0074] Subsequently, in order to reduce manual workload, a pre-trained general object detection model was used to pre-label each sample image in the keyframe sequence. That is, it automatically identifies the potential target objects in each sample image in the keyframe sequence and generates simulated labels containing the category and location bounding box of the target object, i.e., the pre-labeling results.

[0075] To ensure label accuracy, after pre-labeling each sample image in the keyframe sequence, multiple rounds of manual review are introduced to guarantee the accuracy and consistency of the labeling. This results in a complete and accurately labeled low-view sample dataset, the first dataset, providing a solid data foundation for training the inspection model. Specifically, image task packages with simulated labels are first sent to each client (i.e., the annotator's terminal) so that the annotator can view the pre-labeling results. If errors are found in the simulated labels, such as missed detections, false detections, or inaccurate bounding boxes, a label correction command is generated on the client and returned to the inspection model training device. The inspection model training device updates the simulated labels according to the label correction command. To ensure quality, a multi-round verification mechanism is set up for iterative verification until the inspection model training device receives verification pass commands from all relevant clients for the corrected simulated labels of that batch of sample images. The corrected simulated labels corresponding to the verification pass commands are then determined as the true labels for each sample image in the keyframe sequence. Subsequently, each sample image in the keyframe sequence and the true labels are stored together to construct a high-quality first dataset.

[0076] The method provided in this embodiment, which combines automated preprocessing and multi-round manual review, solves the problems of massive invalid frames and inaccurate labeling in low-view video data. It not only ensures the extremely high accuracy of the labels in the first dataset, but also significantly reduces the cost and time of manual labeling, providing a solid data foundation for the training of the inspection model.

[0077] In some embodiments, the metadata further includes the location information of the robot dog; the step of obtaining the target detection loss value includes: Each sample image in the second dataset is input into the feature extraction network of the initial detection model to obtain multiple feature maps of different scales for each sample image in the second dataset. Based on the location information, determine the weight coefficients of the feature maps at each scale for each sample image in the second dataset; Based on the weighting coefficients, feature maps of multiple scales of each sample image in the second dataset are fused to obtain a fused feature map. The fused feature map is input into the detection network of the initial detection model to obtain the target detection results of each sample image in the second dataset; Based on the target detection results and the true labels of each sample image in the second dataset, obtain the target detection loss value corresponding to each sample image in the second dataset.

[0078] Optionally, since the scale of target objects varies greatly during low-angle inspection by the robot dog, for example, nearby objects occupy most of the frame while distant objects occupy only a few pixels. In addition, different regions or locations have different dependencies on the observation scale. Therefore, this embodiment utilizes this characteristic to optimize the loss calculation for each sample image.

[0079] Specifically, for each sample image in the second dataset, the sample image is input into the feature extraction network of the initialization detection model. The feature extraction network then extracts feature maps at multiple scales for that sample image, resulting in multiple feature maps of different scales. After obtaining the feature maps at different scales, the weight coefficients of the feature maps at each scale can be obtained based on the correlation between the robot dog's location information associated with the sample image and the weight coefficients of the feature maps at different scales. For example, if the location information indicates that the robot dog is in a narrow space, where the viewing distance is limited and there are many large targets, the weight coefficient of the deep feature map (i.e., the small-scale feature map) will be increased. If the location information indicates that the robot dog is in an open space, where the viewing distance is far and small hazards at a distance need to be considered, the weight coefficient of the shallow feature map (i.e., the large-scale feature map) will be increased.

[0080] After obtaining the weight coefficients of the feature maps at each scale of the sample image, these weight coefficients can be used to perform weighted fusion of the feature maps at each scale to obtain a fused feature map that enhances information at specific scales. The fused feature map is then input into the detection network of the initial detection model to obtain the predicted target detection result for the sample image.

[0081] After obtaining the predicted target detection result for the sample image, the difference between the predicted target detection result and the true label can be calculated to obtain the target detection loss value corresponding to the sample image. Since the feature fusion stage has already incorporated location-based scale preference, the target detection loss value calculated at this time can guide the model to pay more attention to the target scale most likely to occur in the current geographical environment.

[0082] The method provided in this embodiment addresses the problem of drastic target scale changes in low-view scenes by dynamically adjusting the fusion weights of multi-scale feature maps using location information. This context-aware feature fusion strategy enables the inspection model to automatically adjust its detection attention for different regions based on its environment, significantly improving detection accuracy across all scenes.

[0083] In some embodiments, this application also provides an inspection method for practical application based on the model obtained by the above training method. The execution subject of this method is an inspection device.

[0084] Figure 2 This is a flowchart illustrating the inspection method provided by the present invention; as shown below. Figure 2 As shown, the inspection method includes steps 210, 220 and 230.

[0085] Step 210: Obtain the image to be inspected from the current inspection perspective of the robot dog.

[0086] Optionally, when the robot dog is performing routine inspection tasks, it can acquire video streams captured by data acquisition devices mounted on the robot dog in real time and extract images to be detected from them, such as by extracting images to be detected from the video stream according to preset sampling frequency or key frame extraction algorithms; the images to be detected record the robot dog's perspective view at the current moment and current position.

[0087] In addition, since the inspection device provides a standardized Representational State Transfer API (RESTful API) interface to connect the robot dog with the existing video stream access and analysis platform in the park, the interface design realizes the decoupling and seamless integration of the end-side hardware and the cloud algorithm.

[0088] Therefore, after acquiring the image to be detected, the robot dog can automatically call the RESTful API interface via the wireless network, upload the image to be detected as a request parameter to the inspection device, thereby triggering the backend capability call of the inspection device to detect the image.

[0089] Step 220: Input the image to be detected into the inspection model to obtain the target detection result of the image to be detected.

[0090] Step 230: When the target detection result indicates the presence of an abnormal target, an alarm is issued according to the type of the abnormal target.

[0091] Optionally, before performing step 220, it is necessary to first follow the instructions. Figure 1 The training steps shown train the inspection model, and then encapsulate the trained inspection model (i.e., an AI model with special perspective capabilities) into a service to ensure real-time inference capabilities in actual inspection scenarios. Specifically, the model file and its dependent environment are packaged to generate an inference service component that can run on the park's video analytics platform.

[0092] Therefore, after receiving a video frame, the inspection model loaded inside the inference service component can be used to perform forward inference calculations to obtain the target detection result of the image to be detected. The target detection result includes the category of the identified target object, such as fireworks, illegally parked vehicles, etc., as well as the confidence score and the bounding box of the target object in the image to be detected.

[0093] After reasoning is completed, the inspection device can encapsulate the target detection results into structured data and return them to the business end, thereby enabling real-time display of the recognition results.

[0094] Simultaneously, the inspection device can also perform logical judgments on the target detection results. If a preset abnormal target is detected, such as a flame with a confidence level greater than a threshold, an alarm process is triggered. The specific method of issuing the alarm can be classified according to the type of abnormal target. For example, for high-risk anomalies such as flames or smoke, an audible and visual alarm is immediately triggered, and the highest-level pop-up alarm and on-site video clips are sent to the central control center. For low-risk anomalies such as improper placement, the location is recorded in the log and marked, and the inspection report prompts management personnel to handle the matter, thereby realizing real-time display of the identification results and alarm linkage.

[0095] Furthermore, during the aforementioned inference process, the inspection device simultaneously operates a service status monitoring and anomaly logging mechanism. This means it monitors the processor utilization, response latency, and throughput of the inference service component in real time; simultaneously, it logs anomalies such as inference failures, timeouts, or model errors. This provides data support to ensure the stability and maintainability of the model during long-term use.

[0096] It should be noted that the method provided in this embodiment has the following advantages: First, it can improve the accuracy of low-view scene detection. By specifically collecting low-view video data obtained by the robot dog in a real inspection environment, the artificial intelligence model is trained and optimized in a targeted manner, significantly enhancing the recognition accuracy of targets such as fireworks under complex low-view conditions. Second, it reduces missed detections and false detections, solving the problem of poor adaptability of traditional fixed camera models in dynamic low-view scenes, and effectively reducing the frequency of missed and false alarms during the inspection process. Third, it enhances the reliability of intelligent inspection. By deploying specially optimized artificial intelligence capabilities, it ensures the safety and stability of inspection tasks such as fireworks recognition in the park, and promotes the automation level of the intelligent management system. Fourth, it achieves scenario-based adaptation of technology, breaking through the dependence of existing artificial intelligence models on fixed-view data, constructing a dedicated recognition capability suitable for mobile robots in low-view dynamic scenes, and expanding the boundaries of technology application.

[0097] The method provided in this embodiment solves the problem that existing general models cannot adapt to the unique low-angle and dynamic motion characteristics of robot dogs by driving a robot dog to collect low-angle sample video data in real inspection scenarios and simultaneously recording metadata. Furthermore, by using operational status information to guide dynamic data enhancement, it can accurately simulate visual interference during robot dog movement, improving the model's anti-shake and anti-distortion capabilities. Simultaneously, by dynamically adjusting the training weights of samples using environmental information, it achieves targeted optimization training for difficult samples in complex environments, significantly improving the detection accuracy and generalization ability of the final inspection model in edge scenarios such as large changes in lighting and harsh environments, thereby effectively reducing the false negative and false positive rates in park inspections.

[0098] The inspection model training device provided by the present invention is described below. The inspection model training device described below and the inspection model training method described above can be referred to in correspondence.

[0099] Figure 3 This is a schematic diagram of the inspection model training device provided by the present invention; as shown. Figure 3 As shown, the device includes: The first data acquisition module 310 is used to drive the robot dog to perform mobile inspections of the sample area under multiple different inspection tasks, and obtain sample video data from the target perspective, as well as metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective. The first data management module 320 is used to perform target detection labeling on sample images in the sample video data to obtain a first dataset; The second data management module 330 is used to dynamically enhance each sample image in the first dataset according to the running status information to obtain the second dataset; The optimization module 340 is used to obtain the training weights corresponding to each sample image in the second dataset according to the environmental information, and to train the initial detection model according to the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model to obtain the inspection model.

[0100] The device provided in this embodiment solves the problem that existing general models cannot adapt to the unique low-angle and dynamic motion characteristics of robot dogs by driving a robot dog to collect low-angle sample video data in real inspection scenarios and simultaneously recording metadata. Furthermore, by using operational status information to guide dynamic data enhancement, it can accurately simulate visual interference during robot dog movement, improving the model's anti-shake and anti-distortion capabilities. Simultaneously, by dynamically adjusting the training weights of samples using environmental information, it achieves targeted optimization training for difficult samples in complex environments, significantly improving the detection accuracy and generalization ability of the final inspection model in edge scenarios such as large changes in lighting and harsh environments, thereby effectively reducing the false negative and false positive rates in park inspections.

[0101] Figure 4 This is a schematic diagram of the inspection device provided by the present invention; as shown. Figure 4 As shown, the device includes: The second data acquisition module 410 is used to acquire the image to be detected from the current inspection perspective of the robot dog; The detection module 420 is used to input the image to be detected into the inspection model to obtain the target detection result of the image to be detected; The alarm module 430 is used to issue an alarm according to the type of the abnormal target when the target detection result indicates the presence of an abnormal target; The inspection model is trained based on the inspection model training method provided in the above embodiments.

[0102] The device provided in this embodiment solves the problem that existing general models cannot adapt to the unique low-angle and dynamic motion characteristics of robot dogs by driving a robot dog to collect low-angle sample video data in real inspection scenarios and simultaneously recording metadata. Furthermore, by using operational status information to guide dynamic data enhancement, it can accurately simulate visual interference during robot dog movement, improving the model's anti-shake and anti-distortion capabilities. Simultaneously, by dynamically adjusting the training weights of samples using environmental information, it achieves targeted optimization training for difficult samples in complex environments, significantly improving the detection accuracy and generalization ability of the final inspection model in edge scenarios such as large changes in lighting and harsh environments, thereby effectively reducing the false negative and false positive rates in park inspections.

[0103] The apparatus provided by the present invention is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.

[0104] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can call logic instructions in the memory 530 to execute an inspection model training method. This method includes: driving a robot dog to perform mobile inspections of a sample area under multiple different inspection tasks, obtaining sample video data from the target perspective, and metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective; labeling sample images in the sample video data with target detection tags to obtain a first dataset; dynamically enhancing each sample image in the first dataset according to the operating status information to obtain a second dataset; obtaining training weights corresponding to each sample image in the second dataset according to the environmental information, and training the initialization detection model according to the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initialization detection model to obtain an inspection model; or executing an inspection method, which includes: obtaining the image to be detected from the robot dog's current inspection perspective; inputting the image to be detected into the inspection model to obtain the target detection result of the image to be detected; and issuing an alarm according to the type of the abnormal target when the target detection result indicates the presence of an abnormal target.

[0105] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0106] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the inspection model training method provided by the above methods. The method includes: driving a robot dog to perform mobile inspection of a sample area under multiple different inspection tasks to obtain sample video data from the target perspective, and metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective; and performing target detection labeling on the sample images in the sample video data to obtain a first dataset. Based on the running status information, dynamic enhancement is performed on each sample image in the first dataset to obtain a second dataset; based on the environment information, training weights corresponding to each sample image in the second dataset are obtained, and the initial detection model is trained based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model to obtain an inspection model, or an inspection method is executed, the method including: obtaining the image to be detected from the current inspection perspective of the robot dog; inputting the image to be detected into the inspection model to obtain the target detection result of the image to be detected; when the target detection result indicates the presence of an abnormal target, issuing an alarm according to the type of the abnormal target.

[0107] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program is implemented to perform the inspection model training method provided by the above methods. The method includes: driving a robot dog to perform mobile inspections of a sample area under multiple different inspection tasks, obtaining sample video data from the target perspective, and metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective; performing target detection labeling on sample images in the sample video data to obtain a first dataset; and, based on the operating status information, training the second... The images of each sample in a dataset are dynamically enhanced to obtain a second dataset. Based on the environmental information, the training weights corresponding to each sample image in the second dataset are obtained. Based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model, the initial detection model is trained to obtain an inspection model, or an inspection method is executed. The method includes: obtaining the image to be detected from the current inspection perspective of the robot dog; inputting the image to be detected into the inspection model to obtain the target detection result of the image to be detected; when the target detection result indicates the presence of an abnormal target, issuing an alarm according to the type of the abnormal target.

[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for training an inspection model, characterized in that, include: The robot dog is driven to perform mobile inspections of the sample area under multiple different inspection tasks, obtaining sample video data from the target perspective, as well as metadata synchronized with the sample video data in time. The metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective. The sample images in the sample video data are labeled with target detection tags to obtain the first dataset; Based on the running status information, the sample images in the first dataset are dynamically enhanced to obtain the second dataset; Based on the environmental information, the training weights corresponding to each sample image in the second dataset are obtained. Based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model, the initial detection model is trained to obtain the inspection model.

2. The inspection model training method according to claim 1, characterized in that, The step of obtaining the training weights corresponding to each sample image in the second dataset based on the environmental information includes: Based on the environmental information, obtain the scene complexity and / or lighting conditions of each sample image in the second dataset; Based on the scene complexity and / or the lighting conditions, obtain the training weights corresponding to each sample image in the second dataset.

3. The inspection model training method according to claim 1, characterized in that, The step of dynamically enhancing each sample image in the first dataset based on the running status information to obtain the second dataset includes: Based on the running status information, determine the affine transformation parameters and motion blur kernel parameters of each sample image in the first dataset; Based on the affine transformation parameters and the motion blur kernel parameters, data augmentation is performed on each sample image in the first dataset to obtain the second dataset.

4. The inspection model training method according to any one of claims 1-3, characterized in that, The first dataset is obtained by performing target detection and labeling on sample images in the sample video data, including: The sample video data is preprocessed to obtain a keyframe sequence; the preprocessing includes segmentation and archiving, resolution normalization, and deduplication. The keyframe sequence is labeled with target detection tags based on the target detection model to obtain simulated tags for each sample image in the keyframe sequence; The simulated tags are sent to each client. Upon receiving a tag correction instruction from any of the clients for the simulated tag, the simulated tag is corrected according to the tag correction instruction; The corrected simulated tag is sent back to each of the clients again, and the tag verification is performed iteratively until a verification pass instruction is received from all the clients. The corrected simulated label corresponding to the verification pass instruction is determined as the real label of each sample image in the keyframe sequence; The first dataset is constructed based on each sample image in the keyframe sequence and the real labels.

5. The inspection model training method according to any one of claims 1-3, characterized in that, The metadata also includes the location information of the robot dog; The steps for obtaining the target detection loss value include: Each sample image in the second dataset is input into the feature extraction network of the initial detection model to obtain multiple feature maps of different scales for each sample image in the second dataset. Based on the location information, determine the weight coefficients of the feature maps at each scale for each sample image in the second dataset; Based on the weighting coefficients, feature maps of multiple scales of each sample image in the second dataset are fused to obtain a fused feature map. The fused feature map is input into the detection network of the initial detection model to obtain the target detection results of each sample image in the second dataset; Based on the target detection results and the true labels of each sample image in the second dataset, obtain the target detection loss value corresponding to each sample image in the second dataset.

6. The inspection model training method according to any one of claims 1-3, characterized in that, The multiple inspection tasks include multiple inspection tasks covering different areas, multiple inspection tasks covering different inspection time periods, and multiple inspection tasks covering different inspection environments.

7. An inspection method, characterized in that, include: Acquire the image to be inspected from the robot dog's current inspection perspective; The image to be detected is input into the inspection model to obtain the target detection result of the image to be detected; When the target detection result indicates the presence of an abnormal target, an alarm is issued according to the type of the abnormal target; The inspection model is trained based on the inspection model training method described in any one of claims 1-6.

8. A training device for an inspection model, characterized in that, include: The first data acquisition module is used to drive the robot dog to perform mobile inspections of the sample area under multiple different inspection tasks, obtain sample video data from the target perspective, and metadata synchronized with the sample video data in time; the metadata includes the robot dog's operating status information and environmental information, and the target perspective is the robot dog's own perspective. The first data management module is used to perform target detection labeling on sample images in the sample video data to obtain the first dataset; The second data management module is used to dynamically enhance each sample image in the first dataset according to the running status information to obtain the second dataset; The optimization module is used to obtain the training weights corresponding to each sample image in the second dataset based on the environmental information, and to train the initial detection model based on the training weights and the target detection loss values ​​corresponding to each sample image in the second dataset output by the initial detection model to obtain the inspection model.

9. An inspection device, characterized in that, include: The second data acquisition module is used to acquire the image to be inspected from the robot dog's current inspection perspective. The detection module is used to input the image to be detected into the inspection model to obtain the target detection result of the image to be detected; The alarm module is used to issue an alarm based on the type of the abnormal target when the target detection result indicates the presence of an abnormal target; The inspection model is trained based on the inspection model training method described in any one of claims 1-6.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the inspection model training method as described in any one of claims 1 to 6, or the inspection method as described in claim 7.