Store automation inspection method, model training method, and related device

By training chain stores with a multimodal inspection model for domain adaptation, the problems of high cost, low efficiency and inconsistent standards of manual inspection have been solved. This has enabled automated inspections at all times and high frequency, improving inspection efficiency and accuracy and forming a complete inspection closed loop.

CN122290035APending Publication Date: 2026-06-26ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, the inspection of chain stores relies on manual methods, which leads to high costs, low efficiency, inconsistent standards, difficulty in achieving full-coverage inspections at all times and high frequencies, and problems of missed inspections and misjudgments.

Method used

A multimodal inspection model is used for domain adaptation training. By acquiring image samples and inspection rule texts of the target domain, and combining a visual encoder and a text encoder, the model can perform qualification checks on image frames, output single-frame inspection results, and summarize multi-frame results to determine the inspection results.

Benefits of technology

It reduces inspection costs, improves inspection efficiency and accuracy, ensures automatic inspections at all times and high frequency, forms a complete inspection closed loop, and meets the needs of enterprises for refined operation and management of stores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290035A_ABST
    Figure CN122290035A_ABST
Patent Text Reader

Abstract

This specification provides an automated store inspection method, model training method, and related equipment. The automated store inspection method includes: acquiring information associated with the inspection task, including the store identifier, inspection items, and inspection time period of the store to be inspected; acquiring image frames of the store to be inspected; acquiring inspection rule text corresponding to the inspection items, including the pass / fail criteria and judgment basis for each inspection item, the judgment basis describing the scene features, object features, and / or behavioral features related to the inspection item; calling a multimodal inspection model to perform pass / fail checks on the image frames based on the inspection rule text, and outputting single-frame inspection results; the multimodal inspection model is obtained by fine-tuning a general multimodal model using training samples from the domain of the store to be inspected; and summarizing the single-frame inspection results corresponding to all image frames within the inspection time period to determine the inspection result of the store to be inspected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to an automated store inspection method, a multimodal inspection model training method, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the rapid development of the chain store industry, the number of stores has surged and their geographical distribution has become increasingly dispersed. As a result, enterprises have an ever-increasing demand for refined management of store operation standards and service quality.

[0003] Currently, store inspections still primarily rely on manual methods, namely, regular on-site visits and inspections by headquarters quality inspectors or regional supervisors. This traditional inspection method has several shortcomings: First, the cost of inspection is high. Stores located in scattered areas need to invest a lot of manpower, travel and time costs, which is difficult to maintain in the long term. Secondly, the inspection efficiency is low, the frequency of manual inspections is limited, and it is impossible to achieve full-coverage inspections at all times and with high frequency. Third, the lack of standardized inspection standards and the strong subjectivity of manual judgment make it easy to miss or misjudge inspections. This makes it difficult to detect violations and non-compliance issues in store operations in a timely manner, and it is impossible to form an effective inspection loop, which in turn affects the standardization of enterprise operations and user experience. Summary of the Invention

[0004] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, an automated store inspection method is proposed, comprising: Obtain information associated with the triggered inspection task, including the store identifier pointing to the store to be inspected, inspection items related to store operation and management, and preset inspection time period; Based on the store identifier, obtain the image frames captured by the camera of the store to be inspected during the inspection period; Obtain the inspection rule text corresponding to the inspection item. The inspection rule text includes the pass / fail judgment criteria and judgment basis for the inspection item. The judgment basis is used to describe the scene features, object features and / or behavioral features related to the inspection item. A multimodal inspection model is invoked to perform a qualification check on the image frame based on the inspection rule text and output a single-frame inspection result; wherein, the multimodal inspection model is obtained by performing domain adaptation fine-tuning on a general multimodal model using training samples from the domain to which the store to be inspected belongs; Summarize the single-frame inspection results corresponding to all image frames within the inspection time period, and determine the inspection results of the store to be inspected for the inspection item within the inspection time period.

[0005] According to a second aspect of one or more embodiments of this specification, a multimodal inspection model training method is proposed, comprising: Acquire training samples in the target domain, which include: image samples of stores in the target domain as input to the model, inspection rule texts corresponding to preset inspection items, and qualification judgment results as supervision labels. The image samples and the inspection rule text are input into a general multimodal model to obtain the qualification prediction results; With the goal of minimizing the error between the qualification determination result and the qualification prediction result, the parameters of the general multimodal model are adjusted and optimized to obtain a trained multimodal inspection model, which is then applied to the store automated inspection method as described in the first aspect.

[0006] According to a third aspect of the embodiments of this specification, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect or the second aspect.

[0007] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first or second aspect.

[0008] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in the first or second aspect.

[0009] As can be seen from the above embodiments, this specification can achieve the following technical effects: (1) Reduce inspection costs, improve inspection efficiency and coverage. Automatic inspection can be carried out at all times and at high frequency according to preset time periods to avoid manual omissions.

[0010] (2) Unify inspection standards, clarify the qualification judgment standards and judgment basis through the inspection rule text, and combine the multimodal inspection model adapted to the store field to reduce subjective misjudgment and improve inspection accuracy.

[0011] (3) To achieve systematic summarization and overall judgment of inspection results, the inspection results of stores in the inspection items can be quickly clarified.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0013] Figure 1 This is a flowchart of a multimodal inspection model training method provided in an exemplary embodiment.

[0014] Figure 2 This is an exemplary embodiment of the architecture diagram of a general multimodal model.

[0015] Figure 3 This is an exemplary embodiment of an inspection service system architecture diagram.

[0016] Figure 4 This is a flowchart of an exemplary embodiment of a store automated inspection method.

[0017] Figure 5 This is a schematic diagram of acquiring an image frame provided in an exemplary embodiment.

[0018] Figure 6 This is a schematic diagram of the structure of a device provided in an exemplary embodiment. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0020] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0021] With the rapid development of the chain store model, the number of stores has exploded, and their geographical distribution has become increasingly dispersed. This has placed higher demands on the refined management of store operations, service quality, and safety compliance across various formats, including gyms and retail stores. Store inspections, as a core component of ensuring operational compliance, are a crucial means for businesses to control service quality and mitigate operational risks. Currently, chain store inspections still primarily rely on traditional manual inspection methods, which have numerous shortcomings that negatively impact operational compliance and user experience.

[0022] To address the pain points of manual inspections, the inventors considered using a general multimodal model for inspections. However, they discovered that the general multimodal model has significant limitations in adaptability: it is designed for general scenarios and does not take into account the specific needs of store inspections, such as specific inspection items, scenario characteristics, and judgment criteria. It cannot accurately identify special scenarios in store inspections, such as gym equipment returning to its place, mirror reflection counting, and items to be cleaned in the store. Furthermore, it is difficult to deeply integrate with store inspection rules, resulting in low inspection accuracy and high misjudgment rate, which fails to meet the actual needs of refined store inspections.

[0023] To address the aforementioned shortcomings, this specification provides a multimodal inspection model training method adapted to store inspection scenarios, and an automated store inspection method implemented based on the multimodal inspection model obtained by this training method. By performing domain-specific adaptation training on a general multimodal model, a multimodal inspection model that can accurately adapt to store inspection scenarios is trained, thereby achieving automation, standardization, and precision in store inspections. This solves the drawbacks of traditional manual inspections and the problem of insufficient adaptability of general multimodal models, reduces inspection costs, improves inspection efficiency and accuracy, forms a complete closed loop for store inspections, and meets the needs of enterprises for refined store operation management.

[0024] Please see Figure 1 This paper demonstrates a multimodal inspection model training method that can be executed by electronic devices, including but not limited to servers, industrial control computers, embedded devices, and edge computing devices. These electronic devices possess the hardware resources for data storage, data processing, and model training. They can independently execute training sample acquisition and subsequent model training processes, or they can be implemented through multi-device collaboration, such as having terminal devices handle sample collection and servers handle sample processing and model training, adapting to deployment requirements in different scenarios. The multimodal inspection model training method includes: In S100, training samples for the target domain are obtained. The training samples include: image samples of stores in the target domain as input to the model, inspection rule texts corresponding to preset inspection items, and qualification judgment results as supervision labels.

[0025] The target domain refers to the specific domain to which the store to be inspected belongs, including but not limited to the domain of gyms, retail stores, and restaurants. The store scene characteristics and inspection items differ for different target domains. Therefore, it is necessary to obtain exclusive training samples for that domain to avoid cross-domain sample interference that could lead to a decrease in model training accuracy.

[0026] Regarding the acquisition of image samples, the optional implementation methods include, but are not limited to, the following two: First, by using cameras deployed in the target area stores to collect image data related to the inspection items in the stores according to a preset acquisition frequency; Second, by retrieving the pre-collected and stored image data of the target area stores from the local storage module of the electronic device or the associated cloud database. The image data needs to be pre-processed, and the pre-processing steps include image denoising, size normalization, and grayscale correction to eliminate interference caused by light and equipment precision during the image acquisition process and improve the uniformity of the image samples.

[0027] The image samples cover both normal and abnormal scenarios of store inspections in the target field. For example, in the gym field, there are normal scene images of equipment being put back in place and abnormal scene images of equipment not being put back in place. This ensures that the model can learn the feature differences in different scenarios and improve the accuracy of subsequent recognition.

[0028] The inspection rule text consists of standardized text data that corresponds one-to-one with preset inspection items. It is pre-defined by technical personnel based on the operational specifications and inspection standards of stores in the target area. The text format conforms to the input requirements of a multimodal model, clearly recording the pass / fail criteria and judgment basis for each inspection item. Setting the inspection rule text as model input allows the model to clearly define the core principles of inspection judgment, avoiding reliance on subjective judgments based solely on image features, reducing the probability of misjudgment, and ensuring that the model's output judgment results meet the standardized requirements of store inspections.

[0029] Supervisory labels represent the pass / fail assessment results. These standardized labels are manually annotated by labelers based on the assessment criteria in the inspection rule text. Each set of image samples is matched with the corresponding combination of the inspection rule text, and the label types include: pass / fail, explanations of the reasons, and the location of the abnormal region corresponding to the failure. The purpose of supervisory labels is to provide an error comparison benchmark for model training.

[0030] In S102, the image samples and inspection rule text are input into the general multimodal model to obtain the qualification prediction results.

[0031] For example, please refer to Figure 2 The general multimodal model includes a visual encoder, a text encoder, a cross-modal fusion layer, and a detection layer.

[0032] The visual encoder employs a pre-trained visual Transformer model, such as the CLIP-ViT model based on the Vision Transformer architecture. CLIP (Contrastive Language-Image Pretraining) is a method for pre-training images and text through contrastive learning. ViT (Vision Transformer) is a computer vision model based on the Transformer architecture. The visual encoder receives image samples of stores in the target domain as input. Through multi-layer self-attention mechanisms and convolutional operations, it extracts deep visual features from the images, outputting a high-dimensional feature vector or feature map. This output encodes information about objects in the image related to store inspection items, such as fitness equipment, items to be cleaned, store staff, mirrors, shelves, the layout of the area to be inspected, object positions, and states. These are collectively referred to as image features, providing the core visual basis for subsequent qualification judgments based on inspection rule text.

[0033] The text encoder employs a pre-trained language model, such as a Transformer-based model or the BERT (Bidirectional Encoder Representations from Transformers) model. The text encoder receives inspection rule texts that correspond one-to-one with the store's preset inspection items as input. Through a multi-layer Transformer encoder, it extracts deep semantic and contextual features from the inspection rule text, outputting a high-dimensional feature vector. This output encodes key information in the inspection rules, such as the pass / fail criteria, judgment basis, and core requirements of the inspection items, collectively referred to as text features. This provides semantic-level criterion support for subsequent fusion with image features and the achievement of standardized inspection judgments.

[0034] The cross-modal fusion layer effectively integrates features from different modalities (images and text) to achieve a deep association between visual features and semantic features of inspection rules in store inspection scenarios. For example, to suit store inspection scenarios, a cross-attention mechanism can be used. This layer's processing can include: using text feature vectors as query vectors and image feature vectors as both key and value vectors. Through attention calculation, the similarity between the query vector (text features, i.e., inspection rule requirements) and the key vector (image features, i.e., the actual store scene) is first calculated and normalized to obtain an attention weight distribution, allowing the model to focus on regions and objects in the image related to the inspection rules. Subsequently, the value vector (image features) is weighted and summed according to this weight distribution to generate a visual context vector modulated by text features. Similarly, reverse interactive calculation can be performed, using image features as the query vector and text features as both key and value vectors. Through cross-attention computation, this layer can output a unified shared feature representation that deeply integrates the visual appearance information of the actual store scene with the textual semantic information of the inspection rules. It accurately associates the actual store status with the inspection judgment criteria, thereby providing a complementary and closely related joint feature foundation for the subsequent detection layer to output the qualification prediction results, ensuring that the prediction results meet the actual needs of store inspection.

[0035] The detection layer is used to receive the unified shared feature representation output by the cross-modal fusion layer. Through the preset classification reasoning logic, it outputs the corresponding input qualification prediction result, thereby mapping the fused cross-modal features into a qualification judgment conclusion that meets the store inspection requirements. Optionally, it can also output relevant auxiliary information for judgment at the same time, providing a more comprehensive reference for subsequent inspection and handling and model parameter adjustment.

[0036] For example, the detection layer can adopt a fully connected neural network structure, including an input layer, a hidden layer, and an output layer. The input layer receives a high-dimensional shared feature vector output from the cross-modal fusion layer. The hidden layer performs a non-linear transformation on the feature vector using activation functions (such as ReLU or Sigmoid) to further explore the correlation between features and the pass / fail judgment result, the reason for the judgment, and abnormal regions. The output layer can adopt a multi-output structure. On one hand, it outputs a binary classification judgment result (e.g., pass / fail) and the corresponding prediction confidence. The prediction confidence is used to characterize the reliability of the model's output prediction result, providing a reference for subsequent model parameter adjustments. On the other hand, if the judgment result is unqualified, the reason for the unqualified result can be output simultaneously. The location information of the abnormal area is derived from the judgment criteria in the inspection rule text and image feature analysis. For example, the fitness equipment is not placed in the preset return position, there are objects to be cleaned in the area to be inspected, or the image acquisition equipment is obstructed. The location information of the abnormal area can be represented by normalized center point coordinates (x,y), width (w), and height (h), where the values ​​of x, y, w, and h are all in the range of [0,1]. It can be calculated by the ratio of image pixel coordinates to image size, which can accurately locate the specific area in the image that caused the non-compliance, so that relevant personnel can quickly locate the abnormality and carry out the handling work. If the judgment result is qualified, the core basis for the qualification can also be output to match the relevant requirements of the inspection rule text.

[0037] In S104, the parameters of the general multimodal model are adjusted and optimized with the goal of minimizing the error between the qualification judgment result and the qualification prediction result, so as to obtain the trained multimodal inspection model.

[0038] First, the electronic device calculates the error between the conformity judgment result and the conformity prediction result. This error is used to quantify the degree of deviation between the model prediction result and the actual judgment result, providing a quantitative basis for subsequent parameter adjustments.

[0039] By combining the multi-dimensional results output by the model, such as binary classification results, reasons for non-compliance, and location information of abnormal areas, a multi-task loss function can be used to calculate the comprehensive error, avoiding the limitation of single error calculation in taking into account multiple output dimensions. Specifically, the error for binary classification results can be calculated using the cross-entropy loss function, which effectively measures the difference between the predicted probability and the true label in binary classification scenarios. The error for abnormal area location information can be calculated using the smoothed L1 loss function, which effectively suppresses the impact of outliers on error calculation and improves the accuracy of location prediction. The prediction error for reasons for non-compliance can be calculated using the cosine similarity loss function, which measures the deviation between the semantic features of the non-compliance reasons output by the model and the semantic features of the true reasons. Finally, the errors of each dimension are integrated into a comprehensive loss value through weighted summation, which serves as the basis for adjusting model parameters. The weight coefficients can be preset based on the actual needs of store inspections; this embodiment does not impose any restrictions on this.

[0040] Secondly, the electronic device iteratively adjusts and optimizes the parameters of the general multimodal model with the goal of minimizing the aforementioned comprehensive loss value. The model parameters include all trainable parameters of the visual encoder, text encoder, cross-modal fusion layer, and detection layer. Parameter adjustment employs gradient descent-type optimization algorithms, with optional implementations including stochastic gradient descent, adaptive moment estimation, and root mean square propagation. Through backpropagation, the comprehensive loss value is propagated back to each layer of the model, calculating the gradient of each trainable parameter with respect to the comprehensive loss value. Then, according to the update rules of the optimization algorithm, each parameter is iteratively updated, gradually reducing the comprehensive loss value, so that the model's prediction results gradually approach the actual pass / fail judgment results.

[0041] Furthermore, iteration termination conditions can be set to ensure sufficient model training and avoid overfitting. The iteration termination conditions can be any combination of the following: First, when the overall loss value is less than a preset loss threshold; second, when the number of model iterations reaches a preset upper limit; third, when the decrease in the overall loss value is less than a preset magnitude threshold during a preset number of consecutive iterations. When the iteration termination conditions are met, the parameter adjustment process ends, and the trained multimodal inspection model is obtained.

[0042] It should be noted that for various inspection items in store inspection scenarios, such as store hygiene inspection, staff on-duty status inspection, camera working status inspection, and fitness equipment return detection, a dedicated training sample set can be prepared for each inspection item. Each dedicated training sample set follows the sample composition requirements of S100. During training, the dedicated training sample sets corresponding to all inspection items are aggregated into a model training set, or the training sample sets for each inspection item are input into a general multimodal model in batches for training. Through feature extraction and fusion in S102 and parameter iterative adjustment in S104, the model gradually converges, ultimately possessing the ability to accurately inspect different inspection items. It can receive corresponding inputs for any preset inspection item and output the pass / fail result, the reason for failure, and the location information of abnormal areas that meet the judgment criteria of that inspection item. This adapts to the multi-dimensional and full-scenario inspection needs of stores, eliminating the need to train models separately for different inspection items, effectively reducing model training costs and improving the model's versatility and practicality.

[0043] In some embodiments, the aforementioned trained multimodal inspection model can be directly applied to the automated store inspection method provided in the embodiments of this specification, realizing a closed-loop connection between model training and inspection application.

[0044] Based on the above multimodal inspection model, the automated store inspection method can be implemented in various ways to adapt to the deployment needs of enterprises of different sizes and different store scenarios, as follows: In one possible implementation, this automated store inspection method can be executed independently by a cloud server within the inspection service system. It requires no modification to existing store hardware, relying solely on pre-deployed cameras in the store. Through a pre-defined communication protocol, it retrieves video streams and acquires the image frames needed for store inspection. This fully reuses existing hardware resources, avoids new hardware investment, effectively reduces implementation costs and technical barriers for enterprises, eliminates the need for on-site hardware debugging by professionals, and is more suitable for large-scale deployment scenarios in chain stores. It enables rapid unified inspection management across multiple stores and regions.

[0045] Please see Figure 3 In addition to the cloud server 11, the inspection service system also includes several terminals, such as PC (Personal Computer) 12 and mobile phone 13. Of course, users can obviously also use the following terminals ( Figure 3 (Not shown): Tablet devices, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smartwatches, etc.), etc., and one or more embodiments in this specification are not intended to limit this.

[0046] During operation, the terminal can run the inspection service program, acting as a client for the inspection service. This client-side program can be a native application installed on the terminal, or it can be a mini-program, quick app, or other similar form. Alternatively, when using web technologies such as HTML5, the relevant functions can be implemented through a browser-displayed page. This browser can be a standalone browser application or a browser module embedded in some applications. Users can configure inspection tasks, view inspection results, and receive inspection alarm notifications through the terminal.

[0047] In another possible implementation, in addition to cloud servers and terminals, edge devices, such as edge computing boxes, can be deployed at each store to be inspected, adopting a "cloud-edge collaboration" architecture.

[0048] For example, cloud devices can send inspection service images, trained multimodal inspection models, and related configuration parameters, such as inspection item settings, inspection frequency, and alarm thresholds, to edge devices deployed in each store through secure communication channels such as HTTPS / WebSocket. This supports dynamic updates of inspection strategies and remote upgrades of models without the need for manual on-site operation, thus improving the flexibility and convenience of inspection management.

[0049] Edge devices deployed locally in stores receive video stream data transmitted from store cameras in real time, call the locally loaded multimodal inspection model to perform real-time inference calculations, and quickly generate inspection results, such as pass / fail judgment results, reasons for non-compliance, abnormal area location information and abnormal alarm information, achieving millisecond-level inspection response, avoiding delays caused by data uploading to the cloud, and ensuring that inspection anomalies can be detected in a timely manner.

[0050] Edge devices can report preliminary inspection results generated by local inference to cloud devices in real time. The cloud devices then process the reported results in different levels: high-confidence inspection results with a prediction confidence level that reaches a preset threshold are directly stored in the cloud database to form an inspection record; low-confidence inspection results with a prediction confidence level that is lower than the preset threshold are triggered to perform a second review by the multimodal inspection model deployed in the cloud. Combining the historical inspection data and inspection rule text stored in the cloud, the final inspection conclusion is output, effectively avoiding possible misjudgments that may occur in local inference at the edge device, and balancing inspection efficiency and judgment accuracy.

[0051] In some embodiments, please refer to Figure 4 The flowchart shown below illustrates the automated store inspection method. The following explanation uses a cloud server as an example to illustrate the method, which includes: In S400, information associated with the triggered inspection task is obtained, including the store identifier pointing to the store to be inspected, inspection items related to store operation and management, and the preset inspection time period.

[0052] The store identifier is a unique identifier for the store to be inspected, including but not limited to the store number, store code, and unique IP address. This identifier is used by the cloud server to accurately locate the store and ensure that the image data collected by the store's cameras can be accurately obtained subsequently. The inspection items are various inspection items preset based on the store's operational specifications. These may include store hygiene inspection items, camera working status inspection items, equipment return inspection items, and staff on-duty status inspection items. Different stores can flexibly configure the corresponding inspection items according to their own needs. The preset inspection time period is the execution time range of the inspection task. It can be flexibly set by the user according to the store's operational scenarios (such as store opening hours, peak hours, before and after closing, etc.) to ensure that the inspection can cover the key nodes of store operation.

[0053] For example, users can configure inspection tasks through the terminal in the aforementioned inspection service system, including but not limited to: ① Select preset inspection items, which can be selected individually or in multiple ways to meet multi-dimensional inspection needs.

[0054] ② Bind stores to be inspected. Accurately associate stores through store identifiers. Multiple stores can be bound at the same time, supporting batch inspections.

[0055] ③ Bind to a specific camera in the store to be inspected (optional). The camera is precisely associated through its identifier. Users can accurately select the camera in the corresponding area to be inspected according to the inspection requirements, avoiding data interference from irrelevant cameras and improving inspection efficiency.

[0056] ④ Set inspection time periods, which can be set for single inspection periods or periodic inspection periods.

[0057] ⑤ Based on historical images collected by cameras in stores to be inspected, mark the areas to be inspected or excluded corresponding to a certain inspection item (optional) to provide a reference for subsequent inspection judgment.

[0058] After the user completes the configuration, the terminal will synchronize the configured inspection task information to the cloud server, and the cloud server will then trigger and execute the subsequent tasks.

[0059] For example, the triggering methods for inspection tasks include, but are not limited to, the following two: one is timed triggering, where the cloud server automatically triggers the inspection task at the start of the user-preset inspection time period and begins the inspection process; the other is manual triggering, where the user manually triggers the inspection task for a specific store or specific inspection item through the terminal's interactive interface, suitable for scenarios such as temporary inspections and special inspections. Regardless of the triggering method used, the cloud server will immediately execute the S400 step after the task is triggered to obtain all information associated with the inspection task, laying the foundation for subsequent inspection operations.

[0060] In S402, based on the store identification, image frames captured by the camera of the store to be inspected during the inspection period are obtained.

[0061] Based on the existing hardware deployment of the stores, this step provides two optional methods for acquiring image frames, which can be adapted to different hardware configuration scenarios of different stores without the need for a large amount of new hardware investment, thus reducing the implementation threshold for enterprises.

[0062] In one possible implementation, the cloud server can directly establish a communication connection with the cameras in the store to be inspected. After the communication connection is established, the cloud server sends a data acquisition command to the camera based on the inspection time period obtained by the S400. The camera responds to the command and transmits the video stream data it has collected during the inspection time period to the cloud server in real time. After receiving the video stream data, the cloud server extracts image frames for inspection judgment according to a preset frame extraction frequency, thus completing the acquisition of image frames.

[0063] In another possible implementation, please refer to Figure 5 Data relay can be achieved through an NVR (Network Video Recorder). An NVR is a video storage and management device specifically designed for network video surveillance systems. Its core function is to connect all cameras within a store via a network, receiving and storing the video signals captured by the cameras, while also providing video management, data query, and data transmission capabilities. Cameras in the store to be inspected are pre-connected to the NVR, transmitting the captured video streams to the NVR for storage in real time. A cloud server establishes a secure communication connection with the store's NVR using pre-defined standardized communication protocols, such as GB / T28181-2016, RTSP, and ONVIF. After the connection is established, the cloud server sends a video stream acquisition command to the NVR based on the store identifier and the inspection time period. The NVR responds to this command, transmitting its stored video stream data for the corresponding inspection time period to the cloud server. The cloud server then performs frame extraction processing to obtain the image frames to be inspected. This implementation method fully reuses existing surveillance storage equipment in stores, adapts to the hardware status of most stores, and is highly versatile with low deployment costs.

[0064] In some embodiments, the frame sampling frequency of image frames can be dynamically adjusted. The cloud server first presets a basic frame sampling frequency as a benchmark, and then collects store operation scenario parameters in real time, such as the start and end times of group classes, closing times, and peak customer traffic. Combined with historical inspection anomaly data, the algorithm accurately identifies high-incidence periods and normal periods of inspection anomalies, and automatically adjusts the frame sampling frequency dynamically accordingly. For high-incidence periods of anomalies, such as equipment return checks after group classes or hygiene checks before store closure, the frame sampling frequency is automatically increased and the inspection duration of the corresponding period is extended to ensure that no abnormal scenarios are missed and the judgment is more accurate. For normal periods with no anomalies or extremely low probability of anomalies, the frame sampling frequency is automatically reduced to reduce the collection and processing of redundant image data, achieving a balance between inspection accuracy and system resource consumption.

[0065] To further improve the accuracy of subsequent inspection and judgment, the cloud server can perform preprocessing operations on the extracted single image frames. Specifically, this includes filtering image frames containing invalid images, including but not limited to solid color images, distorted images, mosaic images, blurry images, and images with unidentifiable valid content. For the valid image frames selected after preprocessing, the cloud server can perform unified standardization processing (such as size normalization, grayscale correction, and noise reduction) to ensure that the format and quality of all image frames are consistent, meeting the input requirements of the multimodal inspection model and providing reliable image data support for subsequent qualification judgment.

[0066] In S404, the inspection rule text corresponding to the inspection item is obtained. The inspection rule text includes the pass / fail criteria and judgment basis for the inspection item. The judgment basis is used to describe the scene features, object features and / or behavioral features related to the inspection item.

[0067] The inspection rule text is pre-formulated based on the operational norms and inspection standards of the field to which the store to be inspected belongs, and stored in the local database of the cloud server or the associated cloud database. It corresponds one-to-one with the inspection items, that is, one inspection item corresponds to one exclusive inspection rule text, ensuring the pertinence of the judgment standards.

[0068] The pass / fail criteria clearly stipulate the specific conditions that must be met for each inspection item to pass. For example, for the store hygiene inspection item, the pass / fail criteria can be set as "the area to be inspected has no obvious stains, no debris accumulation, and no standing water on the ground"; for the fitness equipment return inspection item, the pass / fail criteria can be set as "all fitness equipment is placed in the preset return area, and the equipment is neatly arranged, without tilting or missing parts".

[0069] The judgment criteria are supplementary explanations to the pass / fail criteria, used to accurately describe various characteristics related to the inspection item, providing clear recognition guidance for the multimodal inspection model, helping the model quickly locate the core areas, objects, or behaviors related to the inspection item in the image, and improving the accuracy of the judgment. These criteria may include, but are not limited to, the following three categories: First, scene characteristics, i.e., the overall characteristics of the scene to be inspected corresponding to the inspection item, such as "the area to be inspected is the store entrance reception area, and the scene characteristics include a reception desk, green plants, and signs, with no obstructions"; second, object characteristics, i.e., the specific characteristics of the core objects involved in the inspection item, such as for the equipment placement inspection item, the object characteristics can be described as "the fitness equipment is a stationary bike, the main color is black, the placement area is the fitness area on the north side of the store, and the coordinate range is (x1, y1) to (x2, y2)"; third, behavioral characteristics, i.e. the personnel behavioral norms involved in the inspection item, such as for the employee on-duty status inspection item, the behavioral characteristics can be described as "employees must wear uniforms, stand behind the reception desk, and not engage in long-term absences from their posts, playing on their mobile phones, or other violations."

[0070] Furthermore, based on the inspection item information obtained from S400, the cloud server can accurately retrieve the inspection rule text uniquely associated with each inspection item from either the local or cloud database. If the user has customized the judgment criteria for the inspection item when configuring the inspection task on S400, the cloud server will prioritize retrieving the adjusted inspection rule text to ensure that the inspection rules are consistent with the user's configured inspection requirements. After retrieval, the cloud server performs structured processing on the inspection rule text, converting it into a text format that meets the input requirements of the multimodal inspection model.

[0071] In S406, the multimodal inspection model is invoked so that it performs a qualification check on the image frame based on the inspection rule text and outputs the single-frame inspection result. The multimodal inspection model is obtained by performing domain adaptation fine-tuning on the general multimodal model using training samples from the domain to which the store to be inspected belongs.

[0072] The cloud server synchronously inputs a single image frame and the inspection rule text into the multimodal inspection model. It should be noted that if the inspection task is configured with multiple inspection items, the cloud server will call the model separately for each inspection item and perform a pass / fail check on the image frame in sequence to avoid inference confusion caused by mixed judgment of multiple inspection items.

[0073] For example, based on the model's output capabilities, the single-frame inspection result includes, but is not limited to: 1. A single-frame pass / fail judgment conclusion, i.e., the status of the inspection item corresponding to the image frame is "pass" or "fail," and may also include a judgment confidence level to characterize the reliability of the judgment result; 2. If the judgment conclusion is "fail," the reason for failure and the location information of the abnormal area are output simultaneously. The reason for failure is derived by combining the pass / fail judgment criteria in the inspection rule text with the visual feature analysis of the image frame. The location information of the abnormal area is represented by normalized center point coordinates (x, y), width (w), and height (h) to accurately locate the specific area in the image that causes failure; 3. If the judgment conclusion is "pass," the pass / fail judgment basis is output, clarifying which specific judgment criteria in the inspection rule text the image frame meets, ensuring the traceability and interpretability of the judgment result; 4. Other output information, such as the group class attendance statistics.

[0074] In S408, the single-frame inspection results corresponding to all image frames within the inspection period are summarized to determine the inspection results of the store to be inspected for the inspection items within the inspection period.

[0075] For example, the cloud server summarizes the single-frame inspection results corresponding to all image frames output by S406 within the inspection period, establishes a single-frame inspection result summary table, and records the acquisition time, pass / fail judgment conclusion, judgment confidence level, reason for failure (if any), and abnormal area location information (if any) for each frame image, ensuring that all single-frame results are traceable and queryable; then, the summarized single-frame inspection results are subjected to anomaly screening to remove low-confidence misjudgment results (such as single-frame results with judgment confidence level lower than a preset threshold) to avoid such results affecting the accuracy of the final judgment; finally, based on the filtered valid single-frame inspection results, a preset comprehensive judgment rule is used to determine the inspection result of the store to be inspected for the inspection item within the inspection period.

[0076] The comprehensive judgment rules can be flexibly set according to the store inspection needs. Exemplary judgment rules include, but are not limited to, the following: First, the full pass judgment rule, which means that the inspection item is judged to be passable only when the single frame inspection results corresponding to all valid image frames within the inspection period are "passable" and the judgment confidence is higher than the preset threshold. This is suitable for scenarios with extremely high inspection accuracy requirements, such as store safety and compliance inspection items. Second, the proportional pass judgment rule, which means that a pass ratio threshold is set, such as 85% or 90%. If the proportion of passable single frame results within the inspection period is higher than the threshold, the inspection item is judged to be passable. This is suitable for scenarios with high tolerance for instantaneous anomalies, such as store hygiene inspection items. Third, the anomaly threshold judgment rule, which means that an anomaly number threshold is set. If the number of unqualified single frame results within the inspection period does not exceed the threshold, the inspection item is judged to be passable. This is suitable for regular inspection scenarios.

[0077] Furthermore, after completing the comprehensive judgment, the cloud server can record the final judgment result, and at the same time associate the summarized single-frame inspection results, inspection rule text, image frames and other related data to form a complete inspection record, which is stored in the cloud database for easy subsequent query, review and traceability.

[0078] In one possible implementation, the single-frame inspection result includes a judgment on whether it is qualified, and an abnormal region in the corresponding image frame when it is determined to be unqualified. For example, the abnormal region information is represented by normalized center point coordinates (x, y), width (w), and height (h), which can accurately locate the abnormal position and provide clear guidance for subsequent processing.

[0079] If the inspection results of the store to be inspected are unqualified for the inspection items during the inspection period, the cloud server will trigger a real-time alarm process to generate standardized alarm information. The alarm information shall include at least the store identifier of the store to be inspected, the name of the inspection item, the inspection period, and the inspection result, but is not limited to these. It may also include the reason for non-compliance. The reason for non-compliance may be derived by combining the results of single frame inspection, the location information of abnormal area and / or the corresponding abnormal image frame thumbnail, etc. After generating the alarm information, the cloud server pushes the alarm information to the preset terminal through the preset communication method.

[0080] The preset terminals are terminal devices that users pre-configure on the inspection service platform, including but not limited to terminals for store managers, regional supervisors, and headquarters quality inspectors. Based on the importance of the inspection items, the alarms can be pushed to the corresponding level of terminals to ensure that relevant personnel can receive alarms in a timely manner, know the details of the anomalies, and quickly carry out anomaly handling. At the same time, the cloud server will record the alarm push time and reception status to form an alarm handling ledger, which will facilitate the subsequent tracking of the handling progress.

[0081] Otherwise, if the inspection results of the stores to be inspected are satisfactory within the inspection period, the cloud server can execute the daily inspection report generation process when it detects that the daily inspection report generation task has been triggered. There are two ways to trigger the daily inspection report generation task: one is timed triggering, where the cloud server automatically triggers the generation task according to a preset daily report generation time (such as a fixed time period each day); the other is manual triggering, where users manually trigger the daily inspection report generation task for a specific store and a specific inspection period through the interactive interface of the inspection service platform.

[0082] When generating a daily inspection report, the cloud server will summarize the single-frame inspection results of all image frames corresponding to the store to be inspected within the inspection time period, and at the same time associate the inspection rule text corresponding to the inspection item to generate a standardized daily store inspection report. The daily store inspection report includes the pass / fail conclusion of the inspection item, and the pass / fail conclusion is supported by the pass / fail judgment criteria in the inspection rule text.

[0083] For example, the daily inspection report can clearly record the pass / fail criteria for the inspection item, and summarize the pass / fail status of all single-frame inspection results within the inspection period. This explains how the actual state of the store meets the pass / fail criteria in the inspection rules text, ensuring that the pass / fail conclusions are interpretable and persuasive. In addition, the daily store inspection report can also include basic inspection information (store identification, inspection period, inspection item, inspection method), and summary statistics of single-frame inspection results (number of pass frames, pass rate, average confidence level). After generation, the cloud server stores the daily inspection report in the cloud database and can also push it to preset terminals simultaneously, providing reliable written evidence for store operation management and inspection result review, further improving the closed-loop management of the entire inspection process.

[0084] 1. Inspection item: Statistics on the number of participants in group classes.

[0085] In some embodiments, the store to be inspected belongs to the service store sector, such as the gym sector, and the store to be inspected is an offline store of a chain gym; the inspection item is the group class attendance statistics item, which is used to check whether the actual number of attendees during the group class matches the number of attendees, and to avoid operational risks caused by abnormal group class attendance (such as overcrowding, missing registration, etc.); the inspection time period is consistent with the duration of the group class, that is, the inspection time period of the inspection task is set to the start and end time of this group class.

[0086] For example, a user can configure an inspection task for the group class attendance statistics item at a specific store. Specific configuration details may include: selecting "Group Class Attendance Statistics" as the inspection item; binding the store to be inspected (precisely associated via store identifier); binding the imaging device deployed in the group class classroom (i.e., the high-definition camera inside the classroom, used to capture video streams during the group class); setting the inspection time period to the duration of the current group class; and marking the mirror area location and the instructor's face and clothing features based on historical image frames from the classroom (simultaneously recording the inspection rule text). After the user completes the configuration, the inspection service platform synchronizes the inspection task information to the cloud server. The cloud server automatically triggers the inspection task at the start time of the group class (19:00) and obtains all information associated with the inspection task.

[0087] In this embodiment, the inspection rule text corresponding to the group class attendance statistics item is pre-formulated based on the store's group class operation specifications, stored in the cloud database of the cloud server, and uniquely associated with the group class attendance statistics item.

[0088] The pass / fail criteria in the inspection rules text are described as follows: When counting the number of participants in a group class, the double counting of participants caused by mirror reflection and the number of instructors must be excluded. Only the actual number of participants in the group class should be counted to ensure the accuracy of the count and avoid statistical deviations caused by mirror reflection and the inclusion of instructors.

[0089] The judgment criteria in the inspection rule text provide clear recognition guidance for the multimodal inspection model, helping the model accurately distinguish between students, instructors, and mirror reflection areas. Specifically, it includes two parts: First, the location information of the mirror area in the image frame pre-annotated by the user. For example, this location information is represented by normalized center point coordinates (x,y), width (w), and height (h) (the values ​​of x, y, w, and h are all in the range of [0,1]). When configuring the inspection task through the terminal's interactive interface, the user manually annotates the area where the mirror is located based on the historical image frames of the group class classroom of the store to be inspected and saves it to the inspection rule text. This is used by the model to identify and exclude the reflected human figures in the area. Second, the facial feature information and clothing feature information related to the instructor. The facial feature information is the facial feature vector of the group class instructor, for example, pre-collected and entered through a facial recognition device. The clothing feature information is the uniform features of the group class instructor, such as uniform color, logo, and style. This is used by the model to accurately identify the instructors at the group class site and then exclude them during the headcount.

[0090] During the inspection process, the cloud server acquires image frames captured by the imaging device in the group class classroom of the store to be inspected within the inspection period, based on the store identifier. The cloud server synchronously inputs the image frames and corresponding inspection rule text into the multimodal inspection model. The multimodal inspection model used in this embodiment is obtained by fine-tuning a general multimodal model using domain-specific training samples from the service store domain, such as the gym domain. Its training samples include image samples from gym group class scenarios, inspection rule text related to group class attendance statistics, and qualification judgment labels, which can accurately adapt to the attendance statistics requirements of gym group class scenarios.

[0091] The multimodal inspection model performs the following reasoning process: First, it extracts human features and mirror area features from the image frame using a visual encoder, and combines this with the mirror area location information in the inspection rule text to identify and eliminate duplicate human figures caused by mirror reflections. Second, it combines the instructor's facial features and clothing features from the inspection rule text to identify and exclude the instructor from the image frame. Finally, it counts the remaining number of students and outputs the single-frame inspection result. In this embodiment, the single-frame inspection result includes the group class attendance statistics, i.e., the actual number of students in the image frame after excluding mirror reflections and instructors.

[0092] The cloud server summarizes the single-frame inspection results for all image frames during the duration of the group class. The maximum value among the number of people statistics for all image frames during the duration of the group class is determined as the actual number of participants in this group class. This is to take into account that during the group class, students may get up, move around, or temporarily leave the screen, which may cause deviations in the single-frame number of people statistics. Taking the maximum value is closest to the actual highest number of students during the group class, ensuring the objectivity of the statistics.

[0093] The cloud server establishes a communication connection with the course management system of the service store through a preset interface, and synchronizes the number of sign-in students for this group class, that is, the number of students who made an appointment to sign in through the course management system. The actual number of students attending the class is compared with the number of sign-in students. If the actual number of students attending the class exceeds the number of sign-in students, the inspection result of the store to be inspected for the group class attendance statistics during the duration of the group class is deemed unqualified.

[0094] Based on the above inspection results, the cloud server immediately triggers a real-time alarm process: based on the unqualified inspection results and the summarized single-frame inspection results, standardized alarm information is generated. This alarm information includes the store identifier to be inspected, group class attendance statistics, inspection time period, reason for non-compliance (e.g., the actual number of attendees is 25, exceeding the number of sign-in attendees by 20), attendance statistics details (number of valid statistical frames, maximum number of attendees, number of sign-in attendees), and the corresponding abnormal image frame (e.g., the image frame with an attendance count of 25 people) thumbnail. Subsequently, the cloud server pushes the alarm information to preset terminals, such as the store manager terminal and the headquarters operations supervisor terminal.

[0095] If the actual number of participants in this group class does not exceed the number of attendees, for example, if the actual number of participants is 18 and the number of attendees is 20, the inspection result is considered qualified, and the cloud server will not trigger an alarm process. When the store inspection daily report generation task is triggered, the store inspection daily report is generated based on the single-frame inspection results within the duration of this group class, in conjunction with the inspection rule text of the group class participant statistics item. The daily report clearly records the qualified conclusion of the group class participant statistics item, and uses the qualified judgment criteria in the inspection rule text (excluding mirror reflection duplicate counts and the number of instructors) as evidence to show that the group class participant statistics process meets the judgment criteria and the actual number of participants (18 people) does not exceed the number of attendees (20 people). At the same time, the single-frame participant statistics results, the number of frames sampled, the average judgment confidence level, and other information are summarized to provide a reliable basis for store operation management.

[0096] In this embodiment, the multimodal inspection model can accurately exclude duplicate counts of mirror reflections and the number of instructors. This is mainly based on the judgment criteria in the inspection rule text and the fine-tuning and optimization of the model based on training samples specific to the service store domain. It solves the problem of headcount deviation caused by mirror reflections and difficulty in distinguishing instructors and students in group class scenarios in general multimodal models, ensuring the accuracy and reliability of inspection results. This reflects its adaptability and practicality in the headcount statistics inspection scenario of group classes in the service store domain.

[0097] 2. Inspection Items: Check for occupancy of shelving units.

[0098] In some embodiments, the store to be inspected belongs to the service store sector, such as the gym sector, and the store to be inspected is an offline store of a chain gym; the inspection item is the shelf occupancy check item, which is used to check the actual occupancy of the shelves in the service store, to avoid problems such as clutter, safety hazards and decreased user experience caused by excessive shelf occupancy, and to ensure the standardization of the operating environment of the service store; for example, the inspection time period is consistent with the peak business hours of the service store to ensure that the inspection can fully cover the time period when the shelves are used most frequently and accurately reflect the actual occupancy status of the shelves.

[0099] Users configure an inspection task on the terminal for the "Shelf Occupancy Inspection" item in this service-type store. Specific configuration includes, but is not limited to: selecting "Shelf Occupancy Inspection" as the inspection item; binding the store to be inspected; binding the camera deployed in the area where the shelf is located within the service-type store; setting the inspection time period to the service-type store's peak business hours (e.g., 18:00-21:00); marking the shelf area location based on historical image frames of the service-type store's shelf (synchronously entering the inspection rule text); and setting a preset threshold for the actual shelf occupancy rate. After the user completes the configuration, the terminal synchronizes the inspection task information to the cloud server. The cloud server automatically triggers the inspection task at the start of the inspection time period and obtains all information associated with the inspection task.

[0100] In this embodiment, the inspection rule text corresponding to the shelf occupancy inspection item is pre-formulated based on the service store operation specifications, stored in the cloud database of the cloud server, and uniquely associated with the shelf occupancy inspection item.

[0101] The pass / fail criteria in the inspection rules text describe that the actual occupancy rate of the shelves does not exceed the preset percentage threshold. The preset percentage threshold is set by the user in advance according to the purpose, size and operational needs of the shelves and service stores to ensure that the shelves have enough free space to meet the user's temporary storage needs, while keeping the shelves clean and orderly.

[0102] The judgment criteria in the inspection rule text include the location information of the shelves in the image frames pre-annotated by the user, which is used to provide clear recognition guidance for the multimodal inspection model, help the model accurately locate the shelf area and calculate the actual occupancy ratio. Users can manually annotate the complete area of ​​the shelf based on historical image frames of the area where the shelves are located in the store to be inspected and save it to the inspection rule text.

[0103] During the inspection process, the cloud server acquires image frames from the imaging device in the area where the shelves in the store to be inspected are located within the inspection period, based on the store identifier. The cloud server preprocesses the acquired image frames, filtering out invalid images such as those with distorted images, blurry images, abnormal lighting, obstructions blocking the shelves, or images where the shelves are not present. For valid image frames, it performs size normalization and noise reduction to ensure the image frame quality meets the input requirements of the multimodal inspection model. Subsequently, the preprocessed image frames and corresponding inspection rule text are synchronously input into the multimodal inspection model. The multimodal inspection model used in this embodiment is obtained by fine-tuning a general multimodal model using domain-specific training samples from the service store domain. Its training samples include image samples of service store shelf scenarios, inspection rule text related to shelf occupancy checks, and qualification judgment labels, accurately adapting to the needs of service store shelf occupancy detection.

[0104] The multimodal inspection model executes the following reasoning process: First, it extracts visual features from the image frame using a visual encoder, combines this with the shelf location information in the inspection rule text, accurately locates the complete shelf area in the image frame, and calculates the total area of ​​the shelf. Second, it identifies items within the shelf area, such as user backpacks, water bottles, and fitness accessories, and calculates the area occupied by these items. Finally, it calculates the actual occupancy ratio of the shelf by using the ratio of the occupied area to the total area, compares this actual occupancy ratio with a preset ratio threshold, and outputs the single-frame inspection result. In this embodiment, the single-frame inspection result includes a judgment (yes or no) on whether the actual occupancy ratio of the shelf exceeds the preset ratio threshold, along with the actual occupancy ratio value and the judgment confidence level, ensuring the traceability of the judgment result.

[0105] For example, the cloud server summarizes the single-frame inspection results corresponding to all image frames within the inspection period, and checks the single-frame inspection results corresponding to each image frame one by one. If the single-frame inspection results corresponding to the image frames within the inspection period all indicate that the actual occupancy ratio of the shelves exceeds the preset ratio threshold, it is determined that the inspection result of the store to be inspected for the shelf occupancy inspection item within the inspection period is unqualified.

[0106] Based on the above inspection results, the cloud server immediately triggers a real-time alarm process: based on the unqualified inspection results and the summarized single-frame inspection results, standardized alarm information is generated. This alarm information includes the store identification to be inspected, the inspection item (shelf occupancy inspection item), the inspection time period (18:00-21:00), the reason for non-compliance (the actual occupancy rate of the built-in shelves exceeds the preset threshold of 70% throughout the inspection), occupancy rate details (average actual occupancy rate, highest occupancy rate), and the corresponding abnormal image frame (such as the image frame with the highest occupancy rate) thumbnail; subsequently, the cloud server pushes the alarm information to preset terminals, such as the store manager's terminal and the store cleaning staff's terminal.

[0107] If at least one single-frame inspection result indicates that the actual occupancy rate of the shelving unit does not exceed the preset threshold during the current inspection period, the inspection result is deemed qualified. When the cloud server detects that the store inspection daily report generation task has been triggered, it generates a store inspection daily report based on the single-frame inspection results during the current inspection period and the inspection rule text for the shelving unit occupancy inspection item. The daily report clearly records the qualified conclusion of the shelving unit occupancy inspection item, and uses the qualified judgment criteria in the inspection rule text as evidence to show that the current inspection process meets the judgment criteria and is not over-occupied throughout the entire process. At the same time, it summarizes information such as the average occupancy rate, the number of valid inspection frames, and the average judgment confidence level in the single-frame inspection results, providing a reliable basis for store operation management and daily maintenance of shelving units.

[0108] In this embodiment, the multimodal inspection model can accurately calculate the actual occupancy ratio of the shelves. This is mainly based on the shelf location information in the inspection rule text and the fine-tuning and optimization of the model based on training samples specific to the service store domain. This solves the problem of occupancy ratio calculation deviation caused by the inability of general multimodal models to accurately locate shelf areas and distinguish between items and empty areas in service store shelf scenarios. This ensures the accuracy and reliability of the inspection results and reflects the adaptability and practicality of shelf occupancy inspection in service store scenarios.

[0109] 3. Inspection items: Inspection items for the return of movable equipment to their proper place.

[0110] In some embodiments, the store to be inspected belongs to the service industry, such as the gym industry, and the store to be inspected is an offline store of a chain gym; the inspection item is the movable equipment placement check item, which is used to check whether the movable equipment in the store is placed in the preset placement position according to the specifications, to avoid operational risks such as messiness, safety hazards and user inconvenience caused by the random stacking of movable equipment, and to ensure the standardization and safety of fitness areas such as gyms; the inspection time period can be consistent with the time after the service industry stores close, to ensure that the inspection can accurately check the placement of movable equipment after the closing time, and avoid interference caused by temporary non-placement due to user use of equipment during the business hours.

[0111] For example, a user configures an inspection task for the "Mountainous Equipment Placement Inspection" item at a service store. Specific configuration details may include: selecting the inspection item as "Mountainous Equipment Placement Inspection"; binding the store to be inspected; binding the cameras deployed at the service store; setting the inspection time period to the period after the service store closes for business; and marking the preset placement locations of various movable equipment based on historical image frames from the service store (while simultaneously entering the inspection rule text). After the user completes the configuration, the inspection service platform synchronizes the inspection task information to the cloud server. The cloud server automatically triggers the inspection task at the start of the inspection time period and obtains all information associated with the inspection task.

[0112] In this embodiment, the inspection rule text corresponding to the movable equipment return inspection item is pre-formulated based on the service store operation specifications, stored in the cloud database of the cloud server, and uniquely associated with the movable equipment return inspection item.

[0113] The pass / fail criteria in the inspection rules text are described as follows: Movable devices must not be placed in non-movable device storage areas. Ensure that all movable devices are placed in the designated storage areas in accordance with regulations. Maintain the cleanliness and orderliness of service stores, avoid safety hazards such as user bumps and inconvenience caused by disorderly stacking of devices, and facilitate quick access to devices for subsequent users.

[0114] The judgment criteria in the inspection rule text are used to provide clear recognition guidance for the multimodal inspection model, helping the model to accurately locate the preset return area of ​​movable equipment and determine whether the equipment placement is compliant. Specifically, it is the preset return position information of movable equipment in the image frame pre-annotated by the user. For example, this position information is represented by normalized center point coordinates (x,y), width (w), and height (h). When the user configures the inspection task through the terminal's interactive interface, based on the historical image frames of service stores, the preset return area of ​​each type of movable equipment, such as exercise bikes, dumbbell racks, yoga mats, etc., is manually annotated and saved to the inspection rule text. This is used by the model to accurately identify the preset return area, laying the foundation for subsequent judgment on whether movable equipment is placed in a non-return position.

[0115] During the inspection process, the cloud server acquires image frames from the imaging devices of the stores to be inspected within the inspection period, based on the store identifier. The cloud server preprocesses the acquired image frames, filtering out invalid images such as those with distorted images, blurriness, abnormal lighting, or obstructions. For valid image frames, it performs size normalization and noise reduction to ensure the image frame quality meets the input requirements of the multimodal inspection model. Subsequently, the preprocessed image frames and corresponding inspection rule text are synchronously input into the multimodal inspection model. The multimodal inspection model used in this embodiment is obtained by fine-tuning a general multimodal model using domain-specific training samples from the service store domain. Its training samples include image samples of movable equipment scenarios, inspection rule text related to movable equipment placement checks, and pass / fail labels, accurately adapting to the needs of movable equipment placement detection.

[0116] The multimodal inspection model executes the following reasoning process: First, it extracts features of movable devices and preset return areas from the image frame using a visual encoder, and combines this with preset return location information from the inspection rule text to accurately locate the preset return areas of various movable devices in the image frame. Second, it identifies all movable devices in the image frame and locates the actual placement position of each device. Finally, it compares the actual placement position of each movable device with its corresponding preset return area to determine whether any movable devices are placed in non-preset return locations, and outputs the single-frame inspection result. In this embodiment, the single-frame inspection result includes a judgment conclusion (yes or no) on whether a movable device is placed in a non-preset return location, along with the type of illegally placed movable device, its actual placement location information, and the judgment confidence level, ensuring the traceability of the judgment result.

[0117] The cloud server summarizes the single-frame inspection results corresponding to all image frames within the inspection period. If the single-frame inspection results corresponding to all image frames within the inspection period indicate that the movable device is placed in a non-preset return position, the inspection result of the store to be inspected for the return of movable devices within the inspection period is deemed unqualified.

[0118] When the inspection results are unqualified, the cloud server triggers an alarm, generating alarm information including store identification, inspection items, time period, reason for non-compliance, and abnormal images, and pushes it to the preset terminal; conversely, if the inspection results are qualified, no alarm is triggered. Instead, when the daily report generation task is triggered, a daily report is generated in combination with the inspection rules, clarifying the qualified conclusion and summarizing relevant information to provide a basis for operation and management.

[0119] In this embodiment, the multimodal inspection model can accurately determine whether movable equipment is placed in a non-preset return position. This is mainly based on the support of the preset return position information of movable equipment in the inspection rule text, as well as the fine-tuning and optimization of the model based on training samples specific to the service store domain. This solves the judgment deviation problem caused by the inability of general multimodal models to accurately locate the preset return area and distinguish between the actual placement position of the equipment and the return area in the movable equipment return inspection scenario of service stores. This ensures the accuracy and reliability of the inspection results, reflecting its adaptability and practicality in the movable equipment return inspection scenario of service stores.

[0120] 4. Inspection items: Check the on-duty status of store staff.

[0121] In some embodiments, the inspection item is a staff on-duty status check item. This inspection item is used to check whether the staff in designated positions in the store are on duty during the inspection period, to avoid operational risks such as no response to customer inquiries and no handling of safety hazards caused by staff being absent from their posts, and to ensure the standardization and continuity of store operation services. The inspection period is consistent with the store's business hours to ensure that the inspection can fully cover the key time periods when staff are on duty and accurately reflect the actual on-duty status of the staff.

[0122] For example, a user can configure an inspection task for the "Employee On-Duty Status Check" item in a store. Specific configuration details may include: selecting the inspection item as "Employee On-Duty Status Check"; binding the store to be inspected; setting the inspection time period to the store's current business hours; and labeling facial and clothing features related to the store employees based on their information (while simultaneously entering the inspection rule text). After the user completes the configuration, the inspection service platform synchronizes the inspection task information to the cloud server. The cloud server automatically triggers the inspection task at the start of the inspection time period and retrieves all information associated with the inspection task.

[0123] In this embodiment, the inspection rule text corresponding to the employee on-duty status check item is pre-formulated based on the store operation specifications, stored in the cloud database of the cloud server, and uniquely associated with the employee on-duty status check item.

[0124] The pass / fail criteria in the inspection rules text are described as follows: The designated position in the store to be inspected must have a staff member on duty during the inspection period to ensure that the designated position can respond to user inquiries and handle emergencies in a timely manner, ensure the orderly operation of store services, and avoid affecting user experience or causing safety hazards due to staff leaving their posts.

[0125] The judgment criteria in the inspection rule text are used to provide clear recognition guidance for the multimodal inspection model, helping the model to accurately identify store employees in designated positions and determine whether the employees are on duty. Specifically, these criteria include: facial feature information and / or clothing feature information related to the store employees that are pre-labeled by the user. For example, facial feature information is the facial feature vector of all on-duty store employees in the store (pre-collected and entered through a facial recognition device); clothing feature information is the uniform features of the store employees (such as uniform color, logo, style, etc.). When the user configures the inspection task through the terminal's interactive interface, the above feature information is labeled and saved to the inspection rule text, which is used by the model to accurately identify store employees in designated positions, laying the foundation for subsequent judgments on whether store employees are on duty.

[0126] During the inspection process, the cloud server acquires image frames from the cameras of the stores to be inspected within the inspection period, based on the store identifier. The cloud server preprocesses the acquired image frames, filtering out invalid images such as those with distorted images, blurry images, abnormal lighting, or images without people in the frame. For valid image frames, it performs size normalization and noise reduction to ensure the image frame quality meets the input requirements of the multimodal inspection model. Subsequently, the preprocessed image frames and corresponding inspection rule text are synchronously input into the multimodal inspection model. The multimodal inspection model used in this embodiment is obtained by fine-tuning a general multimodal model using store-specific training samples. Its training samples include image samples of specific store job scenarios, inspection rule text related to employee on-duty status checks, and pass / fail labels, accurately adapting to the needs of store employee on-duty status detection.

[0127] The multimodal inspection model executes the following reasoning process: First, it extracts facial and clothing features from the image frame using a visual encoder, and combines this with the facial and / or clothing feature information of the store clerks in the inspection rule text to accurately identify whether a store clerk meets the criteria in the image frame. Second, it locates the area of ​​a designated post in the image frame and determines whether the identified store clerk is within that designated post area. Finally, it comprehensively determines whether a store clerk is on duty at the designated post and outputs the single-frame inspection result. In this embodiment, the single-frame inspection result includes a judgment on whether the store clerk is on duty (yes or no), along with a judgment confidence level. If no store clerk is identified or the store clerk is not in the designated post area, it is judged as "not on duty," ensuring the traceability of the judgment result.

[0128] The cloud server summarizes the single-frame inspection results corresponding to all image frames within the inspection period. If the single-frame inspection results corresponding to all image frames within the inspection period indicate that the store staff is not on duty, the inspection result of the store to be inspected for the staff on duty status inspection item within the inspection period is deemed unqualified.

[0129] When the inspection results are unqualified, the cloud server triggers an alarm, generates an alarm message containing the store identifier, inspection item, time period, reason for non-compliance, and abnormal image, and pushes it to the preset terminal; when the inspection results are qualified, the daily report generation task is triggered, and a daily report is generated in combination with the inspection rules, clarifying the qualified conclusion and summarizing relevant information to provide a basis for operation management.

[0130] In this embodiment, the multimodal inspection model can accurately determine whether a store employee is on duty at a designated position. This is mainly based on the support of the store employee's facial feature information and / or clothing feature information in the inspection rule text, as well as the fine-tuning and optimization of the model based on store-specific training samples. This solves the judgment bias problem caused by the inability of general multimodal models to accurately identify store employees and determine whether personnel are in designated positions in store employee on-duty detection scenarios, ensuring the accuracy and reliability of inspection results. This demonstrates the adaptability and practicality of the model in store-specific employee on-duty status inspection scenarios.

[0131] 5. Inspection items: Store hygiene inspection items.

[0132] In some embodiments, the inspection item is a store hygiene inspection item, which is used to check whether there are any items to be cleaned in the designated inspection area of ​​the store during the inspection period. This avoids operational risks such as decreased user experience and hygiene hazards caused by substandard hygiene, and ensures the cleanliness and standardization of the store's operating environment. The inspection period is consistent with the time between store business hours and after business hours, ensuring that the inspection can accurately check the key time periods for hygiene and cleaning, and accurately reflect the actual hygiene status of the store.

[0133] For example, a user configures an inspection task for a store's hygiene inspection items. Specific configuration details may include: selecting the inspection item as "Store Hygiene Inspection Item"; binding the store to be inspected; binding cameras deployed in the store's inspection area (such as the front desk area, around gym equipment, and restroom entrance); setting the inspection time period to the store's business breaks and after closing time; and based on historical image frames of the store's inspection area, annotating the location information and characteristics of the items to be cleaned in the inspection area (simultaneously entering the inspection rule text). After the user completes the configuration, the inspection service platform synchronizes the inspection task information to the cloud server. The cloud server automatically triggers the inspection task at the start of the inspection time period, obtaining all information associated with the inspection task, including store identification, store hygiene inspection items, inspection time period, and bound imaging device information.

[0134] In this embodiment, the inspection rule text corresponding to the store hygiene inspection items is pre-formulated based on the store operation specifications, stored in the cloud database of the cloud server, and uniquely associated with the store hygiene inspection items.

[0135] The pass / fail criteria in the inspection rules text are described as follows: the area to be inspected has no items to be cleaned during the inspection period, ensuring that the area to be inspected is clean and tidy, avoiding hygiene hazards caused by the accumulation of debris, stains and other items to be cleaned, improving the user fitness experience, and ensuring that the store's operating environment meets hygiene standards.

[0136] The judgment criteria in the inspection rule text provide clear recognition guidance for the multimodal inspection model, helping the model accurately locate the area to be detected and determine whether there are objects to be cleaned within the area. These criteria include: the location information of the area to be detected in the image frame pre-annotated by the user, and the feature information of the objects to be cleaned. The location information of the area to be detected is represented by normalized center point coordinates (x, y), width (w), and height (h). The feature information of the objects to be cleaned includes their shape, color, and size, such as tissues, beverage bottles, dust stains, and discarded fitness equipment. When the user configures the inspection task through the terminal's interactive interface, the above location and feature information are annotated and saved to the inspection rule text. This information is used by the model to accurately locate the area to be detected and identify the objects to be cleaned, laying the foundation for subsequent judgments on the presence of objects to be cleaned within the area.

[0137] During the inspection process, the cloud server acquires image frames from the imaging device in the inspected area of ​​the store within the inspection time period, based on the store identifier. The cloud server preprocesses the acquired image frames, filtering out invalid images such as those with distorted images, blurry images, abnormal lighting, or obstructions blocking the inspected area. For valid image frames, it performs size normalization and noise reduction to ensure the image frame quality meets the input requirements of the multimodal inspection model. Subsequently, the preprocessed image frames and corresponding inspection rule text are synchronously input into the multimodal inspection model. The multimodal inspection model used in this embodiment is obtained by fine-tuning a general multimodal model using store-specific training samples. Its training samples include image samples of the store's inspected area, inspection rule text related to store hygiene inspection, and pass / fail labels, accurately adapting to the needs of store hygiene inspection.

[0138] The multimodal inspection model executes the following reasoning process: First, it extracts region and object features from the image frame using a visual encoder, and combines this with the location information of the region to be detected in the inspection rule text to accurately locate the region to be detected in the image frame. Second, it combines the feature information of the object to be cleaned in the inspection rule text to identify whether there is an object to be cleaned that meets the conditions within the region to be detected. Finally, it comprehensively judges whether there is an object to be cleaned in the region to be detected and outputs a single-frame inspection result. In this embodiment, the single-frame inspection result includes a judgment conclusion (yes or no) on whether there is an object to be cleaned in the region to be detected, along with a judgment confidence level. If an object to be cleaned is identified, it is judged as "there is an object to be cleaned", ensuring the traceability of the judgment result.

[0139] The cloud server summarizes the single-frame inspection results corresponding to all image frames within the inspection period. If the single-frame inspection results corresponding to all image frames within the inspection period indicate the presence of items to be cleaned, the inspection result of the store to be inspected for the store hygiene inspection items within that inspection period is deemed unqualified.

[0140] When the inspection results are unqualified, the cloud server triggers an alarm, generates an alarm message containing the store identifier, inspection item, time period, reason for non-compliance, and abnormal image, and pushes it to the preset terminal; when the inspection results are qualified, the daily report generation task is triggered, and a daily report is generated in combination with the inspection rules, clarifying the qualified conclusion and summarizing relevant information to provide a basis for operation management.

[0141] In this embodiment, the multimodal inspection model can accurately determine whether there are items to be cleaned in the area to be inspected. This is mainly based on the location information of the area to be inspected and the feature information of the items to be cleaned in the inspection rule text, as well as the fine-tuning and optimization of the model based on training samples specific to the store domain. This solves the problem of judgment bias caused by the inability of general multimodal models to accurately locate the area to be inspected and to distinguish between items to be cleaned and normal items in the store hygiene inspection scenario. This ensures the accuracy and reliability of the inspection results and reflects the adaptability and practicality of the store hygiene inspection scenario in the store domain.

[0142] 6. Inspection Items: Camera working status inspection items.

[0143] In some embodiments, the inspection item is a camera working status check item. This inspection item is used to check whether all inspection cameras in the store maintain normal working status during the inspection period, avoid operational risks such as the inability to collect inspection images or the failure of inspection tasks due to camera obstruction, and ensure the smooth implementation of the store's automated inspection process.

[0144] For example, a user can configure an inspection task for the camera's operational status check item at a store. Specific configuration details may include: selecting the inspection item as "Camera Operational Status Check Item"; binding the store to be inspected; and setting the inspection period to all day. After the user completes the configuration, the inspection service platform will synchronize the inspection task information to the cloud server. The cloud server will automatically trigger the inspection task at the start of the inspection period and obtain all information associated with the inspection task.

[0145] In this embodiment, the inspection rule text corresponding to the camera working status inspection item is pre-formulated based on the store inspection specifications, stored in the cloud database of the cloud server, and uniquely associated with the camera working status inspection item.

[0146] The pass / fail criteria in the inspection rules text are described as follows: The camera must maintain normal working condition during the inspection period to ensure that the camera can continuously collect clear and effective inspection images, provide reliable image data support for various inspection tasks in the store, and avoid problems such as missed inspections and misjudgments caused by camera obstruction.

[0147] The judgment criteria in the inspection rule text provide clear recognition guidance for the multimodal inspection model, helping the model accurately determine whether the camera is obstructed and whether it is in normal working condition. These criteria include: pre-set reference feature information corresponding to the obstruction scenario. The reference feature information for the obstruction scenario includes the shape, color, and obstruction form of various obstructions (such as curtains, debris, fingers, dust, etc.). Optionally, it may also include baseline image feature information of the camera's unobstructed acquisition range, pre-labeled by the user.

[0148] During the inspection process, the cloud server obtains image frames from all inspection cameras in the store to be inspected within the inspection period, based on the store identifier.

[0149] The multimodal inspection model used in this embodiment is obtained by fine-tuning the general multimodal model using domain-specific training samples from the store domain. The training samples include image samples of normal camera operation and occlusion scenarios, inspection rule texts related to camera operation status checks, and qualification judgment labels, which can accurately adapt to the needs of store camera operation status detection.

[0150] The multimodal inspection model performs the following reasoning process: First, it extracts scene features and object features from image frames using a visual encoder, and combines this with occlusion scene reference feature information from the inspection rule text to identify whether features matching the occlusion scene exist in the image frame. Second, it comprehensively judges whether the image captured by the camera is occluded and whether the occlusion affects the effectiveness of image acquisition. Finally, it outputs the single-frame inspection result. In this embodiment, the single-frame inspection result includes a judgment conclusion on whether the camera is occluded ("yes" or "no"), along with a judgment confidence level. If an image matching the occlusion scene reference feature is identified, it is determined that "the camera is occluded," ensuring the traceability of the judgment result.

[0151] The cloud server summarizes the single-frame inspection results corresponding to all image frames within the inspection period. If the single-frame inspection results corresponding to all image frames within the inspection period indicate that the camera is blocked, the inspection result of the store to be inspected for the camera working status inspection item within the inspection period is determined to be unqualified.

[0152] When the inspection results are unqualified, the cloud server triggers an alarm, generates an alarm message containing the store identifier, inspection item, time period, reason for non-compliance, and abnormal image, and pushes it to the preset terminal; when the inspection results are qualified, the daily report generation task is triggered, and a daily report is generated in combination with the inspection rules, clarifying the qualified conclusion and summarizing relevant information to provide a basis for operation management.

[0153] In some embodiments, the inspection rule text corresponding to each inspection item can be adaptively optimized, in addition to being adjusted based on the configuration information entered by the user during the inspection task configuration process. For example, the cloud server automatically collects and analyzes all historical inspection data, including single-frame inspection results, final qualification judgment results, and misjudgment / omission cases reported by manual review. Through data mining algorithms, it identifies unreasonable aspects of the judgment criteria in the inspection rule text, such as vague descriptions of the features of the object to be cleaned, deviations in the coordinate marking of the mirror area, and inaccurate definition of the instrument return position range. It then automatically generates rule optimization suggestions, which are then dynamically iterated upon by the user after confirmation.

[0154] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0155] Figure 6 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 6 As shown, device 600 mainly consists of a communication interface 602, a user interface 604, a processor 606, and a data storage 608. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 610. The communication interface 602 enables device 600 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 602 may include an antenna and related processing devices for wireless communication with a radio access network or access point. Furthermore, the communication interface 602 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 602 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 602 may also include multiple physical communication interfaces, such as Wi-Fi interfaces, Bluetooth interfaces, and wide-area wireless interfaces.

[0156] User interface 604 includes receiving user input and providing output to the user. Therefore, user interface 604 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 604 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 604 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 600 may support remote access from other devices via communication interface 602 or another physical interface (not shown). User interface 604 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 604 may also be configured as a display device for rendering or displaying text fragments.

[0157] Processor 606 may contain one or more general-purpose processors and / or special-purpose processors.

[0158] Data storage 608 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 606. Data storage 608 may include removable and non-removable components.

[0159] Processor 606 is capable of executing program instructions 618 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 608 to perform the various functions described herein. Data storage 608 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 600, enable device 600 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 618 by processor 606 may result in processor 606 using data 612.

[0160] For example, program instructions 618 may include an operating system 622 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 600 and one or more applications 620 (e.g., a browser, social application, or game application). Similarly, data 612 may include operating system data 616 and application data 614. Operating system data 616 is primarily accessible to the operating system 622, while application data 614 is primarily accessible to one or more applications 620. Application data 614 may reside in a file system visible or hidden from the user of device 600.

[0161] Application 620 can communicate with operating system 622 through one or more application programming interfaces (APIs). These APIs help application 620 read and / or write application data 614, transmit or receive information via communication interface 602, receive or display information on user interface 604, etc.

[0162] In some terminology, application 620 may be simply referred to as "app". Furthermore, application 620 can be downloaded to device 600 through one or more online app stores or app markets. However, applications can also be installed on device 600 in other ways, such as through a web browser or a physical interface on device 600 (e.g., a USB port).

[0163] In some embodiments, the automated store inspection device can be applied to, for example... Figure 6 The device shown is used to implement the technical solution described in this specification. The automated store inspection device may include: The acquisition module is used to acquire information associated with the triggered inspection task. This information includes the store identifier pointing to the store to be inspected, inspection items related to store operation and management, and the preset inspection time period.

[0164] The acquisition module is also used to acquire image frames collected by the cameras of the stores to be inspected during the inspection period, based on the store identifier.

[0165] The acquisition module is also used to acquire the inspection rule text corresponding to the inspection item. The inspection rule text includes the qualification judgment criteria and judgment basis for the inspection item. The judgment basis is used to describe the scene characteristics, object characteristics and / or behavioral characteristics related to the inspection item.

[0166] The inspection module is used to call the multimodal inspection model so that the multimodal inspection model can perform qualification checks on image frames based on inspection rule text and output single-frame inspection results. The multimodal inspection model is obtained by performing domain adaptation fine-tuning on a general multimodal model using training samples from the domain to which the store to be inspected belongs.

[0167] The inspection module is also used to summarize the single-frame inspection results corresponding to all image frames within the inspection period and determine the inspection results of the store to be inspected for the inspection items within the inspection period.

[0168] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0169] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.

[0170] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0171] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0172] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0173] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.

[0174] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0175] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0176] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0177] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0178] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.

Claims

1. An automated store inspection method, comprising: Obtain information associated with the triggered inspection task, including the store identifier pointing to the store to be inspected, inspection items related to store operation and management, and preset inspection time period; Based on the store identifier, obtain the image frames captured by the camera of the store to be inspected during the inspection period; Obtain the inspection rule text corresponding to the inspection item. The inspection rule text includes the pass / fail judgment criteria and judgment basis for the inspection item. The judgment basis is used to describe the scene features, object features and / or behavioral features related to the inspection item. A multimodal inspection model is invoked to perform a qualification check on the image frame based on the inspection rule text and output a single-frame inspection result; wherein, the multimodal inspection model is obtained by performing domain adaptation fine-tuning on a general multimodal model using training samples from the domain to which the store to be inspected belongs; Summarize the single-frame inspection results corresponding to all image frames within the inspection time period, and determine the inspection results of the store to be inspected for the inspection item within the inspection time period.

2. The method according to claim 1, wherein the domain of the store to be inspected includes service stores; the inspection item includes group class participant statistics; and the inspection time period includes the duration of the group class. The pass / fail criteria in the inspection rules text are used to describe: when counting the number of people in a group class, the duplicate counting of people caused by mirror reflection and the number of instructors must be excluded; The judgment basis in the inspection rule text comprises: The location information of the mirror area in the image frame pre-annotated by the user, as well as the facial feature information and / or clothing feature information related to the coach; The single-frame inspection results include the group class participant statistics; The process of summarizing the single-frame inspection results corresponding to all image frames within the inspection time period and determining the inspection results of the store to be inspected for the inspection item within the inspection time period includes: The maximum value among the number of people in all image frames within the duration of the group class is determined as the actual number of people attending the class. The attendance figures for this group class are synchronized from the course management system, and the actual number of attendees is compared with the attendance figures. If the actual number of attendees exceeds the number of sign-in attendees, the inspection result for the group class attendance statistics of the store to be inspected during the duration of the group class is deemed unqualified.

3. The method according to claim 1, wherein the inspection items include a shelf occupancy inspection item; The pass / fail criteria in the inspection rule text are used to describe that the actual occupancy rate of the shelving does not exceed a preset percentage threshold. The judgment criteria in the inspection rule text include the position information of the shelf in the image frame pre-marked by the user; The single-frame inspection result includes a determination of whether the actual occupancy ratio of the shelf exceeds the preset ratio threshold. The process of summarizing the single-frame inspection results corresponding to all image frames within the inspection time period and determining the inspection results of the store to be inspected for the inspection item within the inspection time period includes: If the single-frame inspection results corresponding to all image frames within the inspection period indicate that the actual occupancy ratio of the shelving exceeds the preset ratio threshold, it is determined that the inspection result of the store to be inspected for the shelving occupancy inspection item is unqualified within the inspection period.

4. The method according to claim 1, wherein the inspection items include a movable equipment return-to-place inspection item; The pass / fail criteria in the inspection rules text are used to describe that movable instruments must not be placed in non-positioned locations. The judgment criteria in the inspection rule text include: the preset return position information of movable instruments in the image frames pre-marked by the user; The single-frame inspection results include: a determination of whether the movable device is placed in a non-preset return position; The process of summarizing the single-frame inspection results corresponding to all image frames within the inspection time period and determining the inspection results of the store to be inspected for the inspection item within the inspection time period includes: If the single-frame inspection results of all image frames within the inspection period indicate that the movable device is placed in a non-preset return position, the inspection result of the store to be inspected for the return of movable devices within the inspection period is deemed unqualified.

5. The method according to claim 1, wherein the inspection items include checks on the on-duty status of store employees; The pass / fail criteria in the inspection rules text are used to describe that: the designated positions in the stores to be inspected must have staff on duty during the inspection period. The judgment basis in the inspection rule text comprises: User-pre-labeled facial and / or clothing features related to store staff; The single-frame inspection result includes a determination of whether the store clerk is on duty; The process of summarizing the single-frame inspection results corresponding to all image frames within the inspection time period and determining the inspection results of the store to be inspected for the inspection item within the inspection time period includes: If the single-frame inspection results for all image frames within the inspection period indicate that the store clerk is not on duty, the inspection result for the store to be inspected during the inspection period is deemed unqualified.

6. The method according to claim 1, wherein the inspection items include store hygiene inspection items; The pass / fail criteria in the inspection rules text are used to describe that the area to be inspected has no items to be cleaned during the inspection period. The judgment criteria in the inspection rule text include the location information of the area to be detected in the image frame pre-annotated by the user, and the feature information of the object to be cleaned; The single-frame inspection result includes a determination of whether there are objects to be cleaned in the area to be detected. The process of summarizing the single-frame inspection results corresponding to all image frames within the inspection time period and determining the inspection results of the store to be inspected for the inspection item within the inspection time period includes: If the single-frame inspection results corresponding to all image frames within the inspection period indicate the presence of items to be cleaned, the inspection results for the store hygiene inspection items within the inspection period are deemed unqualified.

7. The method according to claim 1, wherein the inspection items include camera working status inspection items; The pass / fail criteria in the inspection rules text are used to describe that the camera must maintain normal working status during the inspection period. The judgment criteria in the inspection rule text include reference feature information corresponding to the occlusion scene; The single-frame inspection result includes a determination of whether the camera is obstructed; The process of summarizing the single-frame inspection results corresponding to all image frames within the inspection time period and determining the inspection results of the store to be inspected for the inspection item within the inspection time period includes: If the single-frame inspection results for all image frames within the inspection period indicate that the camera is blocked, the inspection result for the camera working status of the store to be inspected within the inspection period is deemed unqualified.

8. The method according to any one of claims 1 to 7, wherein the single-frame inspection result includes a judgment on whether it is qualified, and an abnormal region in the image frame corresponding to the judgment of being unqualified; The method further includes: If the inspection result of the store to be inspected for the inspection item is unqualified during the inspection period, an alarm message is generated and pushed to a preset terminal. The alarm message includes at least the store identifier, the inspection item, the inspection period, and the inspection result. Otherwise, when the store inspection daily report generation task is triggered, the store inspection daily report is generated based on the single-frame inspection results corresponding to all image frames of the store to be inspected within the inspection time period and associated with the inspection rule text; wherein, the store inspection daily report contains the inspection item qualification conclusion, and the qualification conclusion is supported by the qualification judgment criteria in the inspection rule text.

9. A method for training a multimodal inspection model, comprising: Acquire training samples in the target domain, which include: image samples of stores in the target domain as input to the model, inspection rule texts corresponding to preset inspection items, and qualification judgment results as supervision labels. The image samples and the inspection rule text are input into a general multimodal model to obtain the qualification prediction results; With the goal of minimizing the error between the qualification determination result and the qualification prediction result, the parameters of the general multimodal model are adjusted and optimized to obtain a trained multimodal inspection model. The multimodal inspection model is applied to the store automated inspection method as described in any one of claims 1 to 8.

10. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-9 by executing the executable instructions.

11. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-9.

12. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-9.