A visual inspection method and device for a target scene

By randomly selecting target inspection locations and matching keyframe images using a visual inspection method, the problem of poor reliability and accuracy in traditional inspection methods is solved, achieving both reliability and accuracy in intelligent inspection, which is suitable for refined management and control in complex indoor environments.

CN121661726BActive Publication Date: 2026-04-21ZHE JIANG SHEN XIANG ZHI NENG KE JI YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHE JIANG SHEN XIANG ZHI NENG KE JI YOU XIAN GONG SI
Filing Date
2026-02-03
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing inspection methods rely on manual inspection, which has problems such as strong subjectivity, low information collection efficiency, difficulty in effectively collecting and analyzing data, high cost, and limited coverage frequency, making it difficult to achieve reliability and accuracy.

Method used

A visual inspection method is adopted. By randomly selecting target inspection locations, key frame images are obtained and matched with a pre-stored set of reference images to determine the target inspection locations. Inspection tasks are then performed based on the matching degree to ensure the reliability and accuracy of the inspection process.

Benefits of technology

It improves the reliability and accuracy of inspection tasks, avoids invalid inspection operations, enhances the credibility of inspection data, and adapts to refined management in complex indoor environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661726B_ABST
    Figure CN121661726B_ABST
Patent Text Reader

Abstract

This application discloses a visual inspection method and apparatus for a target scene. The method includes: acquiring keyframe images selected from inspection video data; matching the key feature data of the keyframe images with the reference feature data of each group of reference images in a pre-stored target scene reference image set to determine a matching degree dataset between the key feature data and the reference feature data of each group of reference images; determining whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is a target inspection location; if so, determining whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements; if so, executing the inspection of the target inspection object according to the generated inspection task execution instruction based on the target inspection object randomly determined from the target inspection location. This improves the reliability and accuracy of the inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent inspection technology, specifically to a visual inspection method and apparatus for a target scene. This application also relates to a visual inspection system for a target scene, as well as a computer storage medium and electronic equipment. Background Technology

[0002] With the accelerated advancement of digital transformation, various industries are increasingly demanding operational safety, compliance management, quality control, and service standardization. To promptly identify potential risks, ensure process standardization, and reduce the risk of human error, inspections have become an indispensable part of daily enterprise operations management. Traditionally, inspections rely on manual on-site checks, paper records, or simple electronic spreadsheets, and are widely used in various fields such as catering, retail, manufacturing, energy, transportation, healthcare, and property management.

[0003] Traditional inspections primarily rely on manual on-site checks and recording. This model has several inherent drawbacks: First, the inspection process is highly subjective, with different personnel having varying understandings and application standards, leading to a lack of objectivity and comparability in the results. Second, information collection efficiency is low, the problem feedback chain is long, and it is difficult to achieve real-time reporting and closed-loop rectification of problems. Third, a large amount of unstructured data (such as photos and text descriptions) is difficult to effectively collect and analyze, failing to support enterprise-level risk warning and decision optimization. Fourth, in cross-regional and multi-store scenarios, manual inspections are costly, have limited coverage frequency, and are prone to regulatory blind spots.

[0004] In recent years, with the development of artificial intelligence, computer vision, Internet of Things (IoT) and mobile Internet technologies, intelligent inspection methods have been introduced to send collected data to corresponding remote inspection systems for intelligent identification in order to check for abnormal risks and other problems. Summary of the Invention

[0005] This application provides a visual inspection method for a target scene to solve the problems of poor reliability and authenticity in the inspection process in the prior art.

[0006] This application provides a visual inspection method for a target scene, including:

[0007] Based on the inspection task start command generated by randomly selecting the target inspection location in the target scene, the key frame image selected from the inspection video data is obtained.

[0008] The key feature data of the key frame image is matched with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene;

[0009] Determine whether the inspection location corresponding to the maximum matching value selected from the matching dataset is the target inspection location;

[0010] If so, determine the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether it meets the target inspection location confirmation requirements;

[0011] If so, the inspection of the target inspection object is executed according to the inspection task execution instruction generated from the target inspection location.

[0012] In some embodiments, the step of executing an inspection of the target inspection object based on an inspection task execution instruction generated from the target inspection location includes:

[0013] Randomly select an object to be inspected from the objects to be inspected corresponding to the target inspection location as the target inspection object;

[0014] Obtain the inspection object collection video data of the target inspection object collected according to the inspection task execution instruction; determine whether the target inspection object has disappeared from the inspection object collection video data by tracking the target inspection object detected in the inspection object collection video data;

[0015] If not, determine whether the video data collected from the detected inspection object meets the requirements for generating the target inspection video data;

[0016] When the video data collected by the inspection object meets the requirements for generating the target inspection video data, the generated target inspection video data is inspected.

[0017] In some embodiments, it also includes:

[0018] If the target inspection object disappears from the target inspection video data, then record the target inspection object as abnormal, and / or reset the inspection status of the target inspection object to the initial state.

[0019] In some embodiments, it also includes:

[0020] After completing the inspection task of the target inspection object according to the inspection task execution instruction, a random number for confirming duplicate locations is generated.

[0021] Determine whether the random number at the repeated location is less than a preset random repetition threshold;

[0022] If so, return to the step of determining whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the requirements for confirming the target inspection location.

[0023] In some embodiments, it also includes:

[0024] The execution time of the step of monitoring whether the matching degree between the image feature data returned to the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the requirements for confirming the target inspection location is determined.

[0025] Determine whether the execution duration is greater than a duration threshold;

[0026] If so, then record that there is an anomaly at the target inspection location.

[0027] In some embodiments, the step of obtaining keyframe images selected from the inspection video data based on the inspection task start instruction generated according to the target inspection location randomly selected in the target scene includes:

[0028] The locations to be inspected in the target scenario are randomly sorted.

[0029] The target inspection location is determined by randomly selecting from the randomly sorted locations to be inspected.

[0030] The inspection task start command is generated based on the target inspection location;

[0031] According to the inspection task start command, key frame images selected from the inspection video data are obtained.

[0032] In some embodiments, the step of matching the key feature data of the keyframe image with the reference feature data of each group of reference images in a pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images includes:

[0033] Collect reference images of each inspection location in the target scene;

[0034] The reference image set for the target scene is generated by taking the reference image of each inspection location as a set of reference images;

[0035] Extract the key feature data of the keyframe images and the reference feature data group of each set of reference images;

[0036] Based on the similarity values ​​calculated between the key feature data and each reference feature data in the reference feature data group, a similarity array between the key frame image and each group of reference images is determined.

[0037] The matching score dataset is determined based on the mean of each similarity array.

[0038] In some embodiments, determining whether the inspection location corresponding to the maximum matching value selected from the matching dataset is the target inspection location includes:

[0039] Determine whether the identification information of the inspection location corresponding to the maximum similarity in the matching dataset is the same as the identification information of the target inspection location;

[0040] If so, the inspection location will be designated as the target inspection location.

[0041] In some embodiments, determining whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set satisfies the target inspection location confirmation requirement includes:

[0042] Calculate the similarity between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set;

[0043] Count the number of similarities greater than or equal to the similarity threshold;

[0044] Determine whether the quantity meets the preset ratio for confirming the target inspection location.

[0045] In some embodiments, it also includes:

[0046] If the requirement for confirming the target inspection location is not met, then the inspection location is recorded as abnormal.

[0047] This application also provides a visual inspection device for a target scene, comprising:

[0048] The acquisition unit is used to acquire key frame images selected from the inspection video data based on the inspection task start command generated by the randomly selected target inspection location in the target scene.

[0049] The first determining unit is used to match the key feature data of the key frame image with the reference feature data of each group of reference images in the pre-stored target scene reference image set, and determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene;

[0050] The second determining unit is used to determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location.

[0051] The third determining unit is used to determine, if yes, the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether the target inspection location confirmation requirement is met.

[0052] An execution unit is configured to, if so, execute an inspection of the target inspection object based on an inspection task execution instruction generated from the target inspection location.

[0053] This application also provides a visual inspection system for a target scene, including: a server and an edge;

[0054] The server is used to generate an inspection task start command based on a randomly selected target inspection location in the target scene and send it to the edge terminal; acquire inspection video data collected by the edge terminal and select key images from the inspection video data; match the key feature data of the key frame images with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location; if so, determine whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements; if so, execute the inspection of the target inspection object based on the inspection task execution command generated from the target inspection object randomly determined from the target inspection location.

[0055] The edge device collects inspection video data according to the inspection task start instruction obtained, and collects inspection video data for the target inspection object according to the inspection task execution instruction received from the server.

[0056] or,

[0057] The server is used to generate an inspection task start command based on a randomly selected target inspection location in the target scene, and send the inspection task start command to the edge terminal; and to store the target scene reference image set;

[0058] The edge device is used to collect inspection video data according to the inspection task start command, and select key frame images from the inspection video data; match the key feature data of the key frame images with the reference feature data of each group of reference images in the target scene reference image set pre-stored on the server, and determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location; if so, determine whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements;

[0059] When the server determines that the target inspection location confirmation requirement is met based on the result from the edge terminal, it generates an inspection task execution instruction to perform the inspection of the target inspection object based on the target inspection object randomly determined from the target inspection location, and sends the instruction to the edge terminal.

[0060] This application also provides a computer storage medium, including: a computer program, which, when the computer program is run, executes the visual inspection method for the target scenario described above.

[0061] This application also provides an electronic device, including:

[0062] processor;

[0063] The memory is used to store programs that process data generated by electronic devices. When the program is read and executed by the processor, it performs a visual inspection method for the target scenario described above.

[0064] Compared with the prior art, this application has the following advantages:

[0065] This application provides a visual inspection method for target scenes that randomly selects target inspection locations from the target scene. It then matches the key feature data of the target inspection location with a pre-stored set of reference images to obtain the matching degree. Based on the inspection location corresponding to the maximum matching degree, it determines whether it is a target inspection location. Further confirmation of the target inspection location is then performed. If the confirmation result meets the requirements, the inspection task is executed on the randomly determined target inspection object. This method ensures the reliability and accuracy of subsequent inspection tasks by repeatedly detecting and tracking the target inspection location in the initial stage of the inspection task. It avoids abnormal operations such as pre-preparing inspection locations or replacing locations during inspection, which could lead to subsequent inspections not being at the actual target inspection location, resulting in low reliability and invalid inspections. Attached Figure Description

[0066] Figure 1 This is a flowchart of a visual inspection method for a target scene provided in this application.

[0067] Figure 2 This is a flowchart illustrating an embodiment of a visual inspection method for a target scene provided in this application, specifically regarding the determination of the target inspection location based on a matching degree dataset.

[0068] Figure 3 This is a schematic diagram illustrating the determination of target inspection objects in the catering industry using a visual inspection method for a target scene provided in this application.

[0069] Figure 4 This is a flowchart illustrating an embodiment of a visual inspection method for a target scene provided in this application, specifically regarding the identification of the target inspection object.

[0070] Figure 5 This is a flowchart of an embodiment of the visual inspection method for a target scene provided in this application, specifically regarding the confirmation of recurring locations.

[0071] Figure 6 This is a structural schematic diagram of a visual inspection device for a target scene provided in this application.

[0072] Figure 7 This is a structural diagram of a visual inspection system for a target scene provided in this application.

[0073] Figure 8 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation

[0074] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0075] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The descriptive terms used in this application and the appended claims, such as "a," "first," and "second," are not intended to limit quantity or sequence, but rather to distinguish information of the same type from one another.

[0076] Based on the aforementioned background technology, it is clear that existing inspection methods are transitioning from traditional manual inspection to intelligent inspection. In particular, with the development of vision and multimodal technologies, industries such as catering and supermarkets are increasingly deploying algorithmic models to intelligently identify camera footage from remote inspection systems or photos and videos uploaded from stores, thereby detecting potential anomalies through intelligent inspection. For example, some inspection systems have introduced auxiliary means such as QR code scanning, NFC tag recognition, or location-based services (LBS) to verify whether inspection personnel have reached preset locations. However, in complex indoor environments, these technologies still face challenges such as insufficient positioning accuracy, signal obstruction, and location confusion. Especially in scenarios such as large supermarkets, where similar areas are densely distributed, location misjudgment is easily caused; and in enclosed spaces such as restaurant kitchens, GPS signals are limited, and static tag recognition can only reflect "approach" behavior, unable to determine whether a specific operation actually occurred. Furthermore, existing solutions generally lack the ability to comprehensively analyze inspection behavior, resulting in low reliability of inspection data and further making it difficult to effectively evaluate the quality of inspection execution.

[0077] Therefore, there is an urgent need for an intelligent inspection method that integrates multimodal perception, behavioral logic modeling, and trusted verification mechanisms. This method should be adaptable to highly dynamic and ever-changing business scenarios such as catering and supermarkets, enabling refined control of the inspection process and improved data credibility. This would provide reliable technical support for operational compliance, risk prevention and control, and service standardization.

[0078] Based on this, this application provides a visual inspection method for a target scene, such as... Figure 1 As shown, the method includes:

[0079] Step S101: Based on the inspection task start command generated according to the randomly selected target inspection location in the target scene, obtain the key frame image selected from the inspection video data;

[0080] Step S102: Match the key feature data of the key frame image with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene;

[0081] Step S103: Determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location;

[0082] Step S104: If yes, determine the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether the target inspection location confirmation requirement is met.

[0083] Step S105: If yes, then generate and issue an inspection task execution instruction based on the target inspection object randomly determined according to the target inspection location.

[0084] Before describing steps S101 to S105 in detail, a brief explanation of the concept of visual inspection and the system architecture of its application scenarios will be given.

[0085] Visual inspection refers to the process of automatically or assistedly inspecting and analyzing the appearance, state, location, markings, and environment of a target object using visual acquisition devices such as cameras and image sensors, combined with computer vision and artificial intelligence technologies (such as deep learning, image recognition, and object detection). It is one of the key technologies in an intelligent inspection system. Visual inspection can be applied in the catering industry to inspect and determine whether employee attire, hygiene, ingredients, and operational procedures meet requirements; it can be applied in large supermarkets to check the neatness of fresh produce displays, the presence of near-expiry labels, the closure of refrigerated cabinet doors, and the unobstructed nature of fire exits; it can be applied in warehousing and logistics to verify pallet labels, stacking height, warehouse occupancy status, and forklift operation compliance; and it can be applied in industrial manufacturing to detect product surface defects (scratches, cracks), assembly integrity, and label placement.

[0086] The visual inspection method for the target scene provided in this application can be applied to, for example... Figure 7The system shown illustrates the interaction during the inspection task process. The system includes a server (701) and an edge terminal (702). The server can issue inspection task instructions, configure and store relevant inspection tasks and requirements, analyze inspection data, and provide inspection prompts. The edge terminal can execute the inspection task according to the instructions. The edge terminal can be an electronic device, such as a computer, mobile phone, smart glasses, or smartwatch, used to collect inspection data and provide inspection information. Interaction between the server and the edge terminal can be achieved through inspection application software. This application software can log in using different methods, such as locally downloaded software programs, mini-programs, QR code login, identifier login, or link login.

[0087] The system will be described in detail later; this is only a brief overview. Steps S101 to S104 will be described in detail below.

[0088] Regarding step S101: Based on the inspection task start instruction generated according to the randomly selected target inspection location in the target scene, obtain the key frame image selected from the inspection video data.

[0089] The purpose of step S101 is to acquire keyframe images from the inspection video data.

[0090] In the entire inspection chain, from inspection initiation to inspection execution and inspection termination, inspection initiation is the foundation for subsequent inspection task execution and inspection termination phases. To avoid pre-emptive risks, this embodiment needs to establish the reliability of the inspection chain initiation phase. The target inspection locations in the target scenario are determined through random selection to avoid relying on pre-prepared data for inspections. The specific implementation process may include:

[0091] Step S101-1: Randomly sort the locations to be inspected in the target scene;

[0092] Step S101-2: Randomly select the target inspection location based on the randomly sorted locations to be inspected;

[0093] Step S101-3: Generate the inspection task start command based on the target inspection location;

[0094] Step S101-4: According to the inspection task start instruction, acquire the key frame image selected from the inspection video data.

[0095] In this embodiment, inspection tasks for the target scenario can be pre-stored on the server, and different inspection tasks can be configured for different inspection requirements. The configuration of inspection tasks can be dynamically adjusted at any time before the inspection begins, rather than being a fixed pattern. For example, it can configure the objects to be inspected at the inspection location, as well as the inspection task requirements for those objects. Alternatively, it can first configure the inspection items, and then group the objects to be inspected within those items according to their corresponding inspection locations. The configuration of inspection tasks can be tailored to the inspection requirements of different target scenarios and is not limited to the above configuration methods.

[0096] Taking the catering industry as an example, the inspection location in this embodiment can refer to the location or area where the inspection items are located, such as: the front desk reception area, the general customer area, a private room, a passageway, the kitchen processing room, cooking room, warehouse, cold storage, serving room, fruit room, dishwashing room, garbage room, etc. To avoid the inspection location being prepared in advance, the inspection location can be randomly selected, as described in steps S101-1 and S101-4: the inspection locations are first randomly sorted, and then randomly selected based on the randomly sorted locations, and the randomly selected location is determined as the target inspection location; then, the inspection task start instruction is generated according to the target inspection location, and the inspection task start instruction is sent to the edge terminal, which can then collect inspection video data through the acquisition device according to the received inspection task start instruction. The server selects keyframe images from the inspection video data obtained by the edge terminal. Therefore, step S101 can, on the one hand, determine the target inspection location by randomly selecting the inspection location, thereby ensuring the reliability of subsequent inspection tasks; on the other hand, it can reduce the amount of computation and improve processing efficiency by obtaining key frame images in the inspection video data and filtering out invalid images.

[0097] Regarding step S101-4: According to the inspection task start instruction, key frame images selected from the inspection video data are obtained. The specific implementation process may include at least three methods:

[0098] Method 1:

[0099] Step S101-4-11: According to the configured selection time, keyframe images are periodically captured from the inspection video data. For example, the interval for capturing images is set to T1 seconds, and a keyframe image is captured from the inspection video data every T1 seconds.

[0100] Method 2:

[0101] Step S101-4-21: Randomly select video frame images from the inspection video data according to the set selection time interval;

[0102] Step S101-4-22: Perform feature point detection on the video frame image to determine the number of feature points;

[0103] Step S101-4-23: Determine whether the number of feature points meets the quantity requirement;

[0104] Step S101-4-23: If yes, then the video frame image corresponding to the number of feature points is determined as the keyframe image.

[0105] Method 3:

[0106] Step S101-4-31: Obtain the acceleration value of the inertial sensor of the acquisition device during the inspection video data acquisition process;

[0107] Step S101-4-32: Determine whether the acceleration value is less than a set threshold;

[0108] Step S101-4-33: If yes, select the keyframe image.

[0109] During the execution of step S101, the collected inspection video data should meet the target inspection location specified in the inspection task start instruction. Therefore, it is necessary to determine whether the key frame images in the collected inspection video data correspond to the target inspection location, and then execute steps S102 and S103.

[0110] Regarding step S102: Match the key feature data of the key frame image with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene.

[0111] The purpose of step S102 is to determine whether the key feature data in the keyframe image matches the reference feature data of the reference images in the pre-stored standard reference image set. Therefore, it can be understood that this embodiment needs to pre-store the standard images corresponding to the inspection locations involved in the target scene as reference images, thereby improving the accuracy of the inspection locations after the inspection is started. The reference images can be pre-stored on the server as part of the inspection list, which also includes relevant information such as inspection locations and inspection objects, facilitating random selection. The specific implementation process of step S102 is as follows: Figure 2 As shown, it includes:

[0112] Step S102-11: Collect reference images of each inspection location in the target scene; wherein, the inspection locations can be determined according to inspection requirements, for example: inspection locations can be different inspection rooms or different locations within the same room. Therefore, the determination of inspection locations can be configured according to the inspection task requirements of the target scene. The configuration can be dynamically changed in real time, periodically fixed, or configured according to the attributes of the target scene, such as seasonal configuration, time-based configuration, etc. In this embodiment, the collected reference images are reference data prepared in advance for location determination during subsequent inspections. They can be collected by edge devices, and the reference images that meet the requirements are sent to the server for storage. The reference images stored on the server have identification information corresponding to the inspection locations.

[0113] Steps S102-12: The reference images of each inspection location are used as a set of reference images to obtain the target scene reference image set. To ensure the accuracy of subsequent target location determination and avoid discrepancies between the target and actual inspection locations due to pre-prepared inspection locations, which could lead to an incorrect inspection process from the initial stage, and considering the limitations of the viewing angle at the acquisition end (edge ​​end), multiple images can be acquired for each inspection location to form a set of reference images. All reference image sets from all inspection locations are then used as the target scene reference image set.

[0114] Step S102-13: Extract the key feature data of the key frame image and the reference feature data group of each group of reference images; different inspection locations have different characteristics, therefore, the key feature data of the key frame image and the reference feature data of each group of reference images can be extracted separately.

[0115] Step S102-14: Based on the similarity values ​​calculated between the key feature data and each reference feature data in the reference feature data group, determine the similarity array between the key frame image and each group of reference images;

[0116] Step S102-15: Determine the matching score dataset based on the mean of each similarity array.

[0117] Taking a catering application scenario as an example, the inspection locations in the target scenario may include: front desk-1, front hall-2, vegetable washing room-3, kitchen-4, and dishwashing room-5. Image samples, i.e., reference images, are collected from these 5 inspection locations. For example, if 10 images are collected from each inspection location, and features are extracted from each image, 5 sets of reference feature data can be obtained, as follows:

[0118] The reference feature data group for the front end is: [1-0,1-1,1-2,1-3,1-4,1-5,1-6,1-7,1-8,1-9];

[0119] The reference feature data set for the lobby is: [2-0,2-1,2-2,2-3,2-4,2-5,2-6,2-7,2-8-2-9];

[0120] The reference feature data set for the vegetable washing room is: [3-0,3-1,3-2,3-3,3-4,3-5,3-6,3-7,3-8,3-9];

[0121] The reference feature data set for the kitchen is: [4-0,4-1,4-4,4-3,4-4,4-5,4-6,4-7,4-8,4-9];

[0122] The reference feature data set for the dishwashing area is: [5-0, 5-1, 5-2, 5-3, 5-4, 5-5, 5-6, 5-7, 5-8, 5-9].

[0123] Assuming the target inspection location randomly selected in step S101 is kitchen-4, and 6 keyframe images are selected from the inspection video data, the key feature data group extracted from the 6 keyframe images is: [a,b,c,d,e,f]. The similarity between [a,b,c,d,e,f] and each feature in the reference feature data group of the front desk, lobby, vegetable washing room, kitchen, and dishwashing room is calculated as follows:

[0124] Front-end-1: The similarity between each key feature in the key feature data group [a,b,c,d,e,f] and each reference feature in the front-end reference feature data group [1-0,1-1,1-2,1-3,1-4,1-5,1-6,1-7,1-8,1-9] is calculated to obtain 60 similarity scores, and the mean S_1 of the 60 similarity scores is calculated, assuming S_1=0.12;

[0125] Front Hall-2: The similarity between each key feature in the key feature data set [a,b,c,d,e,f] and each reference feature in the front hall reference feature data set [2-0,2-1,2-2,2-3,2-4,2-5,2-6,2-7,2-8-2-9] is calculated to obtain 60 similarity scores, and the mean S_2 of the 60 similarity scores is calculated, assuming S_2=0.34;

[0126] Vegetable Washing Room-3: The similarity between each key feature in the key feature data set [a,b,c,d,e,f] and each reference feature in the reference feature data set [3-0,3-1,3-2,3-3,3-4,3-5,3-6,3-7,3-8,3-9] of the vegetable washing room is calculated to obtain 60 similarity scores, and the mean S_3 of the 60 similarity scores is calculated, assuming S_3=0.78;

[0127] Kitchen-4: The similarity between each key feature in the key feature data set [a,b,c,d,e,f] and each reference feature in the kitchen's reference feature data set [4-0,4-1,4-4,4-3,4-4,4-5,4-6,4-7,4-8,4-9] is calculated to obtain 60 similarity scores, and the mean S_4 of the 60 similarity scores is calculated, assuming S_4=0.52;

[0128] Dishwashing Room-5: The similarity between each key feature in the key feature data set [a,b,c,d,e,f] and each reference feature in the reference feature data set [5-0,5-1,5-2,5-3,5-4,5-5,5-6,5-7,5-8,5-9] of the dishwashing room is calculated, resulting in 60 similarity scores. The mean S_5 of the 60 similarity scores is then calculated, assuming S_5 = 0.44.

[0129] Use S=[0.12, 0.34, 0.78, 0.52, 0.44] as the matching score dataset.

[0130] The above is merely an example and is not intended to limit the scenarios in which the technical solution can be used. In this embodiment, the extraction of image feature data can be implemented using open-source models such as NetVLAD and EigenPlaces.

[0131] Regarding step S103: Determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location.

[0132] The purpose of step S103 is to determine whether the information displayed in the current inspection video data corresponds to the location of the target inspection site. The specific implementation process may include:

[0133] Step S103-1: Determine whether the identification information of the inspection location corresponding to the maximum similarity value in the matching dataset is the same as the identification information of the target inspection location;

[0134] Step S103-2: If yes, then the inspection location is determined as the target inspection location;

[0135] Step S103-3: If not, record that the inspection location is abnormal, that is, based on the information in the inspection video data, the current inspection location is incorrect, or the inspection video data was collected incorrectly.

[0136] Continuing with the above example, when the matching score dataset S=[0.12, 0.34, 0.78, 0.52, 0.44], the maximum value of 0.78 corresponds to the vegetable washing room-3, meaning the inspection location is the vegetable washing room-3. However, the assumed randomly selected target inspection location is kitchen-4. Therefore, the inspection location is different from the target inspection location, and this is recorded as an anomaly. If the matching score dataset S=[0.12, 0.34, 0.28, 0.52, 0.44], the maximum value of 0.52 corresponds to kitchen-4, which is the same as the assumed randomly selected target inspection location. In this case, step S104 is executed.

[0137] like Figure 4 As shown, regarding step S104: determine the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether it meets the target inspection location confirmation requirements.

[0138] To further ensure the reliability of the inspection location at the beginning of the inspection process, when the result of step S103 indicates that the inspection location is the target inspection location, step S104 is needed to confirm the target inspection location. This prevents errors in the inspection location due to external interference. Therefore, the specific implementation process of step S104 includes:

[0139] Step S104-1: Calculate the similarity between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set;

[0140] Step S104-2: Count the number of similarities greater than or equal to the similarity threshold;

[0141] Step S104-3: Determine whether the quantity meets the preset ratio for confirming the target inspection location.

[0142] Continuing with the previous example, suppose the inspection video data includes 10 images corresponding to 10 image feature data [a,b,c,d,e,f,g,h,i,j]. The image feature data [a,b,c,d,e,f,g,h,i,j] are compared with the five sets of reference feature data mentioned above, and the similarity is calculated as follows:

[0143] Each feature in the image feature data set [a,b,c,d,e,f,g,h,i,j] is compared with each reference feature in the front-end reference feature data set [1-0,1-1,1-2,1-3,1-4,1-5,1-6,1-7,1-8,1-9] to calculate the similarity, resulting in 100 similarity values.

[0144] Each feature in the image feature data set [a,b,c,d,e,f,g,h,i,j] is compared with each reference feature in the reference feature data set [2-0,2-1,2-2,2-3,2-4,2-5,2-6,2-7,2-8-2-9] of the front hall to obtain 100 similarity values;

[0145] The similarity between each feature in the image feature data set [a,b,c,d,e,f,g,h,i,j] and each reference feature in the reference feature data set [3-0,3-1,3-2,3-3,3-4,3-5,3-6,3-7,3-8,3-9] of the vegetable washing room is calculated to obtain 100 similarity values.

[0146] Each feature in the image feature data set [a,b,c,d,e,f,g,h,i,j] is compared with each reference feature in the kitchen reference feature data set [4-0,4-1,4-4,4-3,4-4,4-5,4-6,4-7,4-8,4-9] to calculate the similarity, resulting in 100 similarity values.

[0147] The similarity between each feature in the image feature data set [a,b,c,d,e,f,g,h,i,j] and each reference feature in the reference feature data set [5-0,5-1,5-2,5-3,5-4,5-5,5-6,5-7,5-8,5-9] of the dishwashing room is calculated to obtain 100 similarity values.

[0148] The similarity array can be represented as [0.56, 0.11, 0.21, 0.87, 0.54, ..., 0.nn]. 100 similarity values ​​are compared with a similarity threshold, and the similarity values ​​greater than or equal to the threshold are counted. This process can be performed by sorting the 100 similarity values ​​and then comparing them with the similarity threshold. Assuming the similarity threshold is set to 0.6, the number of similarity values ​​greater than or equal to 0.6 is counted. If the proportion of this number in the total similarity values ​​exceeds a preset ratio (e.g., 0.5), the target inspection location meets the confirmation requirements and can proceed with subsequent inspection tasks.

[0149] Understandably, the similarity threshold and preset ratio mentioned above can be set according to the requirements of different inspection scenarios, or according to historical inspection data. For example, if there are many abnormal risks in the historical inspection data, the threshold setting threshold can be increased; conversely, if there are few abnormal risks, the threshold setting threshold can be decreased.

[0150] When there are similar inspection locations, and the result of the target inspection location confirmation requirement in step S104 is yes, a second confirmation of the target inspection location can be performed by comparing location information or inspection location identifiers. This can also include:

[0151] Step S10a: Obtain the positioning information of the acquisition device that collects the inspection video data;

[0152] Step S10b: Compare the positioning information with the pre-stored location information of the inspection locations in the target scene;

[0153] Step S10c: Based on the comparison results, determine whether the secondary confirmation requirements for the target inspection location are met.

[0154] This will further ensure the consistency between video inspection data and the target inspection location when there are many similar inspection locations, and improve the reliability of subsequent inspection tasks.

[0155] Based on the above, when the inspection location corresponds to the target inspection location and is confirmed once or multiple times, step S105 is executed.

[0156] Regarding step S105: If the determination result of step S104 is yes, then the inspection of the target inspection object is executed according to the inspection task execution instruction generated from the target inspection location.

[0157] The purpose of step S105 is to generate inspection task execution instructions for subsequent execution of the inspection tasks. A target inspection location may include one or more inspection tasks, each corresponding to a different inspection object. For example... Figure 4 As shown, the specific implementation process of this step may include:

[0158] Step S105-1: Randomly select an object to be inspected from the objects to be inspected corresponding to the target inspection location as the target inspection object; each inspection location corresponds to an inspection object, which can be stored in advance on the server. After the target inspection location is determined, a random selection can be made based on the objects to be inspected corresponding to the target inspection location, and the randomly selected object to be inspected is taken as the target inspection object.

[0159] Step S105-2: According to the inspection task execution instruction, acquire inspection object collection video data including the target inspection object; the inspection task execution instruction may include: prompt information for performing collection operations on the target inspection object. The collection device (edge ​​device) performs video collection on the target inspection object. During the collection process, a visual algorithm model can be used to detect the object to be inspected in the target inspection location, and compare the detected object to be inspected with the target inspection object to determine whether the object to be inspected is the target inspection object. This process will generate collection operation prompt information for collecting the target inspection object, so as to accurately collect the target inspection object. The collection operation prompt information can be output in one or more ways, such as visual, text, and voice. In some implementation scenarios, the target inspection object can include different inspection objects or the same or similar inspection objects. For example, in the catering industry, if the target inspection location is the kitchen, and the randomly selected target inspection object is the cold dish area, the cold dish area may include multiple identical dishes (here, identical dishes refer to dishes with the same name, preparation, and raw materials) or different dishes. Therefore, random selection can be performed again within the cold dish area. Figure 3 As shown, the detected dishes are used as the target inspection object set. Each object to be inspected in the target inspection object set will display a detection box. Randomly select a dish from the objects to be inspected in the target inspection object set as the target inspection object, and display it in a form that can be distinguished from the detection boxes of other objects to be inspected. In this way, the inspectors can know the specific location of the target inspection object. The prompts for performing data collection on the target inspection object can guide inspection personnel to collect data on the target inspection object according to the collection requirements. For example, for situations requiring observation of small items, detailed features of item appearance (such as identifying discoloration, rot, insect eggs, oil stains, and damage), or identification of text (such as expiration dates, self-service food labels, and takeout labels), close-up photography is required, and the prompts include shooting distance requirements. For situations requiring identification of combined items (such as whether there is hand sanitizer, disinfectant alcohol, or a hand dryer around the sink, or whether tableware in the disinfection cabinet is arranged holistically), the prompts include the requirement to include these combined items or the outermost container in the frame and fill the field of view as much as possible. For situations requiring measurement of the length and thickness of processed food, the prompts include the requirement to place the item in the correct orientation next to the measuring tool on the worktable. For situations requiring measurement of food weight, the prompts include the requirement to first tare the food using the same container on an electronic scale, and then place the item on the electronic scale to weigh it, and so on. The list will not be provided here. Different prompts can be used for different target inspection objects. The purpose is to improve the efficiency of target inspection video data collection and the accuracy of the collection results.

[0160] Step S105-3: Based on the tracking of the target inspection object detected in the video data collected from the inspection object, determine whether the target inspection object has disappeared from the video data collected from the inspection object. The purpose of this step is to ensure that the target inspection object is always in the video data collected from the inspection object, avoiding abnormal risks such as replacement. In this embodiment, the continuous detection and tracking of the target inspection object can adopt a common multi-target tracking scheme such as tracking-by-detection. The core idea is to decompose the detection and tracking task into two stages: first, target detection is performed independently in the frame image to generate detection boxes, and then these detection boxes are connected into a continuous trajectory through a temporal association strategy. When the tracking target detection box is not displayed in the video data collected from the inspection object, it means that the target inspection object has disappeared from the video data collected from the inspection object.

[0161] Step S105-4: If not, determine whether the detected video data collected from the inspection object meets the requirements for generating target inspection video data; during this process, prompt information corresponding to the inspection collection requirements can be continuously generated; the purpose of this step is to determine whether the image information or video frame of the video data collected from the inspection object is clear and complete and will not affect subsequent inspections. This process can be automatically and completely judged by an algorithm model. For example, for cases where the detailed features of an item's appearance need to be observed, the size of the target detection box for item detection and tracking can be used for automatic judgment. When the size of the target detection box reaches the preset size, the detection step for the detailed features of the appearance is completed; for cases where text needs to be recognized, the text detection and recognition results of the OCR algorithm can be used for automatic judgment. The detection process is completed when the OCR text can be recognized, the confidence level of the text line output by the OCR model reaches a preset threshold, and the content of the text line string conforms to a preset template format (e.g., the expiration date field should conform to a preset date format). For measuring dimensions, video keyframes and text prompts can be simultaneously input into a universal multimodal language model to directly determine whether the item is aligned with the ruler according to the inspection requirements and whether the ruler's markings are clear. For measuring weight, video or keyframe sequences and text prompts can be input into the multimodal language model. After the measurement process is complete, the multimodal language model will return an automatically determined result based on the text prompts, such as whether the scale reading is clear. During the process of detecting whether the video data collected from the inspection object meets the requirements for generating the target inspection video data, prompts corresponding to the current display of the video data collected from the inspection object can be continuously generated to allow inspection personnel to adjust their collection operations in a timely manner and / or understand the relevant requirements for generating the target inspection video data.

[0162] Step S105-5: When the video data collected from the inspection object meets the requirements for generating the target inspection video data, the generated target inspection video data is inspected. The inspection results determine whether the target inspection object in the target inspection video data meets the inspection requirements. The purpose of this step is to inspect the target inspection object in the target inspection video data to obtain the inspection results. This can be considered a post-inspection stage. In the pre-inspection and acquisition stages, random location selection, target location determination, and target location confirmation effectively avoid unreliable or incorrect location issues in the pre-inspection stage. It also effectively avoids the risk of the target inspection object being replaced before generating the target inspection video data, providing a reliable data foundation for subsequent inspection of the target inspection object and determining whether it meets the inspection requirements. Specifically, the inspection of the target inspection object can be accomplished using a multimodal large language model. For example, the target inspection video data and text prompts (or voice prompts) can be used as input data for the multimodal large language model to obtain the inspection results for the target inspection object. Multimodal large language models can be used via API calls to determine whether a target inspection object meets inspection requirements. These models are trained on sample data provided by the target scenario. To integrate with inspection tasks, the parameters of the multimodal large language model can be configured to employ a greedy decoding strategy in its output data. For example:

[0163] Setting the temperature parameter to 0 or close to 0 (e.g., 0.01) ensures the determinism, accuracy, and reproducibility of the output results. The temperature parameter (T) is a hyperparameter that controls the randomness and diversity of the text generated by the language model, and it affects the probability distribution of the softmax output.

[0164] Setting the sampling parameter to False disables random sampling to ensure the determinism of the output results.

[0165] The random number parameter is set to a fixed integer, such as the random seed: an integer used to initialize the pseudo-random number generator. Setting a fixed seed ensures that, with the same input and parameters, the exact same output sequence can be obtained each time.

[0166] The above describes the situation where the target object has not disappeared. If the judgment result of step S105-3 is yes, that is, the target inspection object has disappeared from the inspection video data, the abnormal risk of the target inspection object can be recorded. Furthermore, the inspection status of the target inspection object can be reset to the initial state.

[0167] To improve the reliability of inspections during the execution of inspection task instructions targeting a specific inspection location, and to prevent errors in subsequent inspection tasks due to the replacement of the target inspection location after completing one inspection task, the target inspection location can be reconfirmed after executing one or more inspection tasks. Figure 5 As shown, the specific implementation process may include:

[0168] Step S106: After completing the inspection task of the target inspection object according to the inspection task execution instruction, generate a random number for confirming duplicate locations;

[0169] Step S107: Determine whether the random number of the repeated locations is less than a preset random repetition threshold;

[0170] Step S108: If yes, return to the step of determining whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the requirements for confirming the target inspection location; if no, return to step S105 to execute the inspection task for the next target inspection object.

[0171] The specific implementation process of step S106 includes:

[0172] Step S106-1: Determine whether there are other target inspection tasks at the target inspection location after the target inspection object's inspection task is completed;

[0173] Step S106-2: If yes, then after the inspection task of the current target inspection object is completed, generate a random number for confirming the repeated location; the random number can be a random number uniformly distributed between 0 and 1.

[0174] The random repetition threshold mentioned in step S107 can be R1, for example, 0.1. The weight of R1 can be set for different inspection locations to flexibly determine the repetition of target inspection locations. For example, for some similar target inspection locations, the weight of R1 can be set larger, and for dissimilar locations, a smaller weight can be set.

[0175] The process of performing step S108 also includes:

[0176] Step S109: Monitor the matching degree between the image feature data returned to the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether the execution time of the step of confirming the target inspection location is met;

[0177] Step S1010: Determine whether the execution duration is greater than the duration threshold; the duration threshold can be T2, for example: 15 seconds. Of course, this is just an example. The actual duration threshold can be dynamically adjusted and is not a fixed value.

[0178] Step S1011: If yes, record that the target inspection location is abnormal; if the execution time exceeds the time threshold, it indicates that the determination of the repeated target inspection location has failed, therefore, the target inspection location is abnormal. Then, you can return to step S101 to start the next round of inspection, or you can interrupt the inspection. If no, execute the next inspection task, i.e., return to step S105.

[0179] Understandably, when the above-mentioned abnormal situations occur, not only can they be recorded, but prompts can also be generated so that inspection personnel can know why the current inspection task cannot be performed and make corresponding adjustments.

[0180] The above describes an embodiment of a visual inspection method for a target scene provided in this application. This method can randomly select target inspection locations from the target scene, and obtain the matching degree between the key feature data of the target inspection location and a pre-stored reference image set. Then, it determines whether the inspection location corresponding to the maximum value in the matching degree is the target inspection location. Based on this, it is necessary to further confirm the target inspection location. If the confirmation result meets the confirmation requirements, the inspection task is executed on the randomly determined target inspection object. In this way, the reliability and accuracy of the subsequent inspection task execution can be guaranteed by multiple detection and tracking confirmation of the target inspection location in the initial stage of the inspection task. This avoids the operation of preparing inspection locations in advance for inspection, which may result in low reliability and invalid inspection results due to the subsequent inspection process not being the actual target inspection location.

[0181] In addition, during the inspection of the target inspection object, the random selection of the target inspection object and the continuous detection and tracking of the target inspection object in the inspection video data are also embedded, thereby ensuring the reliability of the target inspection object during the execution of the inspection task.

[0182] Furthermore, by integrating multimodal perception and intelligent verification mechanisms, a closed-loop anti-counterfeiting system is constructed throughout the entire inspection lifecycle, ensuring the authenticity, integrity, and traceability (risk record) of inspection locations. Through proactive risk prevention, fraudulent initiation is blocked at the source; through enhanced process feasibility, the authenticity of inspected objects and the standardization of operations are ensured during the inspection process. Multimodal perception and intelligent verification reduce the cost of manual review and enable automated information output. Therefore, the visual inspection method for the target scenario provided in this application not only ensures the rigidity of the inspection system's execution but also upgrades inspection from "passive compliance" to a core element of "proactive risk control," laying a solid foundation for enterprise safety production and digital governance.

[0183] The above is a detailed description of an embodiment of a visual inspection method for a target scene provided in this application. Corresponding to the aforementioned embodiment of a visual inspection method for a target scene, this application also discloses an embodiment of a visual inspection device for a target scene. Please refer to [link / reference]. Figure 6 Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For relevant details, please refer to the description of the method embodiments. The device embodiments described below are merely illustrative.

[0184] The visual inspection device for the target scene includes:

[0185] The acquisition unit 601 is used to acquire key frame images selected from the inspection video data based on the inspection task start instruction generated by the randomly selected target inspection location in the target scene.

[0186] The first determining unit 602 is used to match the key feature data of the key frame image with the reference feature data of each group of reference images in the pre-stored target scene reference image set, and determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene;

[0187] The second determining unit 603 is used to determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location.

[0188] The third determining unit 604 is used to determine, when the determining result of the second determining unit 603 is yes, the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether the target inspection location confirmation requirement is met.

[0189] The execution unit 605 is configured to perform an inspection of the target inspection object based on the inspection task execution instruction generated from the target inspection object randomly determined from the target inspection location when the determination result of the third determining unit 704 is yes.

[0190] The acquisition unit 601 includes: a sorting subunit, a determining subunit, a generating subunit, and an acquisition subunit;

[0191] The sorting subunit is used to randomly sort the locations to be inspected in the target scene;

[0192] The determining subunit is used to randomly select from the randomly sorted locations to be inspected to determine the target inspection location.

[0193] The generation subunit is used to generate the inspection task start command according to the target inspection location;

[0194] The acquisition subunit is used to acquire key frame images selected from the inspection video data according to the inspection task start instruction.

[0195] For details on the specific implementation process of the acquisition unit 601, please refer to the above description of step 101, which will not be elaborated here.

[0196] The first determining unit 602 includes: a collection subunit, a generation subunit, an extraction subunit, a calculation subunit, and a determining subunit;

[0197] The acquisition subunit is used to acquire reference images of each inspection location in the target scene;

[0198] The generation subunit is used to generate the target scene reference image set by taking the reference image of each inspection location as a set of reference images;

[0199] The extraction subunit is used to extract key feature data of the keyframe image and reference feature data group of each set of reference images;

[0200] The calculation subunit is used to determine the similarity array between the key frame image and each group of reference images based on the similarity value calculated between the key feature data and each reference feature data in the reference feature data group.

[0201] The determining subunit is used to determine the matching degree dataset based on the mean of each similarity array.

[0202] For details on the specific implementation process of the first determining unit 602, please refer to the above description of step 102, which will not be elaborated here.

[0203] The second determining unit 603 includes: a first determining subunit and a second determining subunit;

[0204] The first determining subunit is used to determine whether the identification information of the inspection location corresponding to the maximum similarity in the matching degree dataset is the same as the identification information of the target inspection location;

[0205] The second determining subunit is used to determine the inspection location as the target inspection location if the determination result of the first determining subunit is yes.

[0206] For details on the specific implementation process of the second determining unit 603, please refer to the above description of step 103, which will not be elaborated here.

[0207] The third determining unit 604 includes: a calculation subunit, a statistics subunit, and a determining subunit;

[0208] The calculation subunit is used to calculate the similarity between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set;

[0209] The statistical subunit is used to count the number of similarities greater than or equal to the similarity threshold;

[0210] The determining subunit is used to determine whether the quantity meets the preset ratio for confirming the target inspection location.

[0211] The third determining unit 604 may further include: a recording unit, used to record an abnormality at the inspection location if the requirement for confirming the target inspection location is not met.

[0212] The specific implementation process of the third determining unit 604 can be referred to the above description of step S104, and will not be detailed here.

[0213] The specific implementation process of the execution unit 605 may include: selecting a subunit, acquiring a subunit, detecting and tracking a subunit, determining a first subunit, and determining a second subunit;

[0214] The selection subunit is used to randomly select an object to be inspected from the objects to be inspected corresponding to the target inspection location as the target inspection object;

[0215] The acquisition subunit is used to acquire video data of the target inspection object collected according to the inspection task execution instruction; the detection and tracking subunit is used to determine whether the target inspection object has disappeared from the video data of the inspection object by tracking the target inspection object detected in the video data of the inspection object.

[0216] The first determining subunit is used to determine whether the video data collected from the detected inspection object meets the requirements for generating target inspection video data when the tracking result of the detection and tracking subunit is negative.

[0217] The second determining subunit is used to inspect the generated target inspection video data when the determination result of the first determining subunit is yes.

[0218] It also includes a recording subunit, which is used to record an anomaly of the target inspection object if the tracking result of the detection and tracking subunit is yes, and / or to reset the inspection status of the target inspection object to the initial state.

[0219] Also includes:

[0220] The repeating unit is used to generate a random number for confirming the repeating location after completing the inspection task of the target inspection object according to the inspection task execution instruction.

[0221] The judgment unit is used to determine whether the random number of the repeated locations is less than a preset random repetition threshold;

[0222] The return unit is used to return to the third determining unit 604 to perform the following steps when the judgment result of the judgment unit is yes: determining whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the requirements for confirming the target inspection location.

[0223] For details on the implementation process of the visual inspection device for the target scene, please refer to the above-described embodiment of the visual inspection method for the target scene, which will not be described in detail here.

[0224] Based on the above, this application also provides a visual inspection system for a target scene, such as... Figure 7 As shown, the system includes a server 701 and an edge terminal 702, and the specific implementation process can include two embodiments.

[0225] Example 1:

[0226] The server 701 is used to generate an inspection task start instruction based on a randomly selected target inspection location in the target scene and send it to the edge terminal; acquire inspection video data collected by the edge terminal and select key images from the inspection video data; match the key feature data of the key frame images with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location; if so, determine whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements; if so, execute the inspection of the target inspection object based on the inspection task execution instruction generated from the target inspection object randomly determined from the target inspection location.

[0227] The edge device 702 collects inspection video data according to the obtained inspection task start instruction, and collects inspection video data for the target inspection object according to the inspection task execution instruction received from the server.

[0228] Example 2:

[0229] The server 701 is used to generate an inspection task start command based on a randomly selected target inspection location in the target scene, and to send the inspection task start command to the edge terminal.

[0230] The edge terminal 702 is used to collect inspection video data according to the inspection task start command, and select key frame images from the inspection video data; match the key feature data of the key frame images with the reference feature data of each group of reference images in the target scene reference image set pre-stored on the server, and determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location; if so, determine whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements;

[0231] When the server determines that the target inspection location confirmation requirement is met based on the result from the edge terminal, it generates an inspection task execution instruction to perform the inspection of the target inspection object based on the target inspection object randomly determined from the target inspection location, and sends the instruction to the edge terminal.

[0232] It should be noted that in this embodiment, the edge terminal 702 needs to collect inspection data through acquisition devices, such as inspection video data and inspection object video data, and output prompt information during the execution of the inspection task. The server terminal 701 needs to process random data during the execution of the inspection task. Therefore, the execution of tasks by the server terminal 701 and the edge terminal 702 can be interchanged. For example, when the performance of the edge terminal 702 meets the corresponding data processing requirements, it can be executed according to Embodiment 2; otherwise, it can be executed according to Embodiment 1. This allows the system to dynamically adjust the execution mode based on the performance of the local edge terminal when executing inspection tasks, realizing diversified system deployment for inspection task execution.

[0233] Based on the above, this application also provides a computer storage medium, including: a computer program, which, when the computer program is run, executes the relevant content of the visual inspection method for the target scenario described above.

[0234] Based on the above, this application also provides an electronic device, such as... Figure 8 As shown, the electronic device includes:

[0235] Processor 801;

[0236] The memory 802 is used to store a program for processing data generated by the electronic device. When the program is read and executed by the processor, it performs the relevant content of the visual inspection method for the target scene described above.

[0237] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0238] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0239] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0240] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0241] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0242] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

Claims

1. A visual inspection method for a target scene, characterized in that, include: Based on the inspection task start command generated by randomly selecting the target inspection location in the target scene, the key frame image selected from the inspection video data is obtained. The key feature data of the key frame image is matched with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; Determine whether the inspection location corresponding to the maximum matching value selected from the matching dataset is the target inspection location; If so, determine the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether it meets the target inspection location confirmation requirements; If so, the inspection of the target inspection object is executed according to the inspection task execution instruction generated from the target inspection location.

2. The method according to claim 1, characterized in that, The step of executing an inspection task execution instruction based on a target inspection object randomly determined from the target inspection location to inspect the target inspection object includes: Randomly select an object to be inspected from the objects to be inspected corresponding to the target inspection location as the target inspection object; Obtain the inspection object collection video data of the target inspection object collected according to the inspection task execution instruction; determine whether the target inspection object has disappeared from the inspection object collection video data by tracking the target inspection object detected in the inspection object collection video data; If not, determine whether the video data collected from the detected inspection object meets the requirements for generating the target inspection video data; When the video data collected by the inspection object meets the requirements for generating the target inspection video data, the generated target inspection video data is inspected.

3. The method according to claim 2, characterized in that, Also includes: If the target inspection object disappears from the target inspection video data, then record the target inspection object as abnormal, and / or reset the inspection status of the target inspection object to the initial state.

4. The method according to claim 1 or 2, characterized in that, Also includes: After completing the inspection task of the target inspection object according to the inspection task execution instruction, a random number for confirming duplicate locations is generated. Determine whether the random number at the repeated location is less than a preset random repetition threshold; If so, return to the step of determining whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the requirements for confirming the target inspection location.

5. The method according to claim 4, characterized in that, Also includes: The execution time of the step of monitoring whether the matching degree between the image feature data returned to the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the requirements for confirming the target inspection location is determined. Determine whether the execution duration is greater than a duration threshold; If so, then record that there is an anomaly at the target inspection location.

6. The method according to claim 1, characterized in that, The step of generating an inspection task start command based on a randomly selected target inspection location in the target scene, and acquiring keyframe images selected from the inspection video data, includes: The locations to be inspected in the target scenario are randomly sorted. The target inspection location is determined by randomly selecting from the randomly sorted locations to be inspected. The inspection task start command is generated based on the target inspection location; According to the inspection task start command, key frame images selected from the inspection video data are obtained.

7. The method according to claim 1, characterized in that, The step of matching the key feature data of the keyframe image with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images includes: Collect reference images of each inspection location in the target scene; The reference image set for the target scene is generated by taking the reference image of each inspection location as a set of reference images; Extract the key feature data of the keyframe images and the reference feature data group of each set of reference images; Based on the similarity values ​​calculated between the key feature data and each reference feature data in the reference feature data group, a similarity array between the key frame image and each group of reference images is determined. The matching score dataset is determined based on the mean of each similarity array.

8. The method according to claim 5, characterized in that, Determining whether the inspection location corresponding to the maximum matching value selected from the matching dataset is the target inspection location includes: Determine whether the identification information of the inspection location corresponding to the maximum similarity in the matching dataset is the same as the identification information of the target inspection location; If so, the inspection location will be designated as the target inspection location.

9. The method according to claim 1, characterized in that, The determination of whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements includes: Calculate the similarity between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set; Count the number of similarities greater than or equal to the similarity threshold; Determine whether the quantity meets the preset ratio for confirming the target inspection location.

10. The method according to claim 1, characterized in that, Also includes: If the requirement for confirming the target inspection location is not met, then the inspection location is recorded as abnormal.

11. A visual inspection device for a target scene, characterized in that, include: The acquisition unit is used to acquire key frame images selected from the inspection video data based on the inspection task start command generated by the randomly selected target inspection location in the target scene. The first determining unit is used to match the key feature data of the key frame image with the reference feature data of each group of reference images in the pre-stored target scene reference image set, and determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; The second determining unit is used to determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location. The third determining unit is used to determine, if yes, the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set, and whether the target inspection location confirmation requirement is met. An execution unit is configured to, if so, execute an inspection of the target inspection object based on an inspection task execution instruction generated from the target inspection location.

12. A visual inspection system for a target scene, characterized in that, include: Server-side and edge-side; The server is used to generate an inspection task start command based on a randomly selected target inspection location in the target scene and send it to the edge terminal; acquire inspection video data collected by the edge terminal and select key images from the inspection video data; match the key feature data of the key frame images with the reference feature data of each group of reference images in the pre-stored target scene reference image set to determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location; if so, determine whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements; if so, execute the inspection of the target inspection object based on the inspection task execution command generated from the target inspection object randomly determined from the target inspection location. The edge device collects inspection video data according to the inspection task start instruction obtained, and collects inspection video data for the target inspection object according to the inspection task execution instruction received from the server. or, The server is used to generate an inspection task start command based on a randomly selected target inspection location in the target scene, and send the inspection task start command to the edge terminal; and to store the target scene reference image set; The edge device is used to collect inspection video data according to the inspection task start command, and select key frame images from the inspection video data; match the key feature data of the key frame images with the reference feature data of each group of reference images in the target scene reference image set pre-stored on the server, and determine the matching degree dataset between the key feature data and the reference feature data of each group of reference images; wherein, the target scene reference image set is the reference image set of all inspection locations in the target scene; determine whether the inspection location corresponding to the maximum matching value selected from the matching degree dataset is the target inspection location; if so, determine whether the matching degree between the image feature data of all images in the inspection video data and the reference feature data of all reference images of the target inspection location in the reference image set meets the target inspection location confirmation requirements; When the server determines that the target inspection location confirmation requirement is met based on the result from the edge terminal, it generates an inspection task execution instruction to perform the inspection of the target inspection object based on the target inspection object randomly determined from the target inspection location, and sends the instruction to the edge terminal.

13. A computer storage medium, characterized in that, include: A computer program that, when executed, performs the method as described in any one of claims 1 to 10.

14. An electronic device, characterized in that, include: processor; A memory for storing a program for processing data generated by an electronic device, wherein the program, when read and executed by the processor, performs the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Unmanned aerial vehicle inspection photo processing method and device, computer equipment, readable storage medium and program product

    CN119474444A

  • Data center inspection method and device

    WO2017076328A1