A multimodal generative security detection method and system applied to the logistics field

By using a multimodal generative security detection method, effective images are generated using scene text description information. The dataset is fused and the feature flow is decoupled to perform environmental risk situation classification and dynamic detection. This solves the problems of sample scarcity and insufficient robustness in logistics and warehousing security detection, and achieves efficient security early warning and intelligent decision-making.

CN122135091APending Publication Date: 2026-06-02SHAANXI SHUTUHANG INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI SHUTUHANG INFORMATION TECH CO LTD
Filing Date
2026-02-13
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing logistics and warehousing safety detection technologies suffer from problems such as sample scarcity and imbalance, limited data augmentation methods, single perception modality, insufficient generalization of feature extraction, separation of detection and risk identification processes, lack of dynamic self-learning ability, limited system intelligence level, and insufficient robustness in complex and dynamic scenarios.

Method used

A multimodal generative security detection method is adopted. By acquiring scene text description information and feedback information from real logistics scene images, effective images are generated, the dataset is fused and the feature flow is decoupled, environmental risk situation classification is performed, dynamic detection parameter thresholds are calculated to guide local feature flow detection, and security early warning decisions and uncertainty assessments are performed.

Benefits of technology

It improves the accuracy and flexibility of target detection, enhances the system's adaptability and generalization ability in complex dynamic environments, realizes end-to-end intelligent decision-making closed loop, and improves the real-time performance and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135091A_ABST
    Figure CN122135091A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal generative safety detection method and system applied in the logistics field, comprising: acquiring scene text description information and feedback information of real logistics scene images to obtain effective generated images; combining all effective generated images and real logistics scene images into a fusion dataset; for each sample image, decoupling features into local and global feature flows; classifying the global feature flows according to environmental risk situations to obtain the probability that the sample image belongs to K risk scene categories; calculating a set of dynamic detection parameter thresholds based on each probability; using the K sets of dynamic detection parameter thresholds to guide local feature flow detection to obtain K sets of detection results; making a safety warning decision based on the K sets of detection results and the probabilities of the K risk scene categories, and performing uncertainty assessment; generating feedback information when the uncertainty condition is met. This invention solves the problems of lack of physical constraints on samples and the inability of detection thresholds to dynamically adapt to the scene in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of logistics, specifically relating to a multimodal generative safety detection method and system applied in the logistics field. Background Technology

[0002] With the rapid development of intelligent logistics and smart warehousing, automated equipment, unmanned transport vehicles (AGVs), robotic arm sorting systems, and human collaborative operations are widely used in warehousing environments.

[0003] In this dynamic scenario of multi-entity collaboration, safety detection and risk warning have become key links in ensuring operational efficiency and personnel safety.

[0004] Existing logistics and warehousing security detection technologies primarily rely on visual monitoring and traditional target detection algorithms to identify the locations of personnel, vehicles, and goods through video images, thereby assessing potential intrusion, collision risks, or violations. These technologies employ deep learning models (such as Faster R-CNN, YOLO, and SSD) to detect surveillance video frames or camera images, identifying targets such as people, vehicles, goods, pallets, and forklifts, and determining the presence of risks such as collisions, boundary crossings, or violations. These methods typically use two-dimensional images as input, extract features using convolutional neural networks (CNNs), and then output target locations and category labels through region proposal networks or dense prediction structures.

[0005] However, existing technologies still have the following shortcomings in complex and dynamic logistics and warehousing scenarios: 1. High sample scarcity and data imbalance: Safety accidents and abnormal behaviors (such as collisions and accidental entry) in logistics and warehousing are characterized by low frequency and sporadic occurrence, resulting in a scarcity of real and effective abnormal samples that are difficult to collect and reproduce. This makes traditional detection models that rely on a large amount of labeled data face a serious sample imbalance problem, prone to overfitting during training, with poor generalization ability, and unable to cover diverse potential risk scenarios.

[0006] 2. Limited data augmentation techniques and lack of semantic realism: Traditional data augmentation methods (such as image transformation) struggle to generate anomalous samples with complex semantic information. Although generative models (such as GANs and diffusion models) are used for data synthesis, the generated samples differ from real-world scenes in style, physical plausibility, and semantics. The lack of effective semantic consistency verification and alignment mechanisms means that direct use may introduce noise and affect model performance.

[0007] 3. Limited perceptual modality and incomplete scene understanding: Existing technologies mainly rely on single visual signals (RGB images or videos) and fail to effectively integrate voice and text (work instructions, alarm logs). This single-modal perception method lacks a comprehensive understanding of the work context, behavioral intent, and risk context, making it difficult for the system to distinguish between "normal work" and "dangerous behavior".

[0008] 4. Insufficient feature extraction generalization and lack of unified semantic representation: Mainstream detection models extract features based on supervised convolutional neural networks. Their capabilities are highly dependent on labeled data for specific tasks and are sensitive to interference factors such as complex lighting changes, viewpoint occlusion, object stacking, and dynamic backgrounds in warehouse environments. The lack of a shared, semantically unified feature representation space across different scenarios or modalities results in weak transfer and generalization capabilities of the model across scenarios and devices.

[0009] 5. The detection and risk identification processes are fragmented, and the decision-making logic is dispersed: Current systems typically process "target detection" (outputting category and location) and "risk semantic understanding" (determining risk type, level, and behavioral compliance) as two independent modules processed sequentially. There is a lack of effective information exchange and collaborative decision-making mechanisms between the two, failing to achieve an end-to-end intelligent decision-making loop from "perception" to "cognition" to "early warning," thus affecting the accuracy and real-time performance of the overall system judgment.

[0010] 6. Lack of dynamic self-learning and closed-loop optimization capabilities: The warehousing environment, operating modes, and equipment layout are constantly changing. Existing systems are mostly "trained once, deployed statically," unable to utilize new data generated during runtime (especially new anomalous samples) for online adaptive optimization and iteration of the model. This leads to a gradual degradation of detection performance over time, and maintenance and updates require high costs for manual retraining.

[0011] 7. Limited system intelligence and lack of semantic understanding: Existing technologies mostly remain at the low-level perception stage of "target appearance," lacking the ability to understand complex events, behavioral logic, and safety procedures at a high level. The system struggles to automatically associate multi-target, multi-step behavioral sequences and cannot generate interpretable risk assessment criteria and early warning information, thus limiting its application value in proactive safety control and intelligent decision support.

[0012] 8. Insufficient robustness to complex dynamic environments: In summary, existing methods still have significant deficiencies in overall robustness, stability and adaptability in logistics warehousing, a practical application with high complexity (diverse scenarios, dense targets, dynamic interactions) and uncertainty (lighting changes, temporary occlusion). It is difficult to maintain continuous and reliable monitoring and early warning performance under various real and non-ideal operating scenarios. Summary of the Invention

[0013] To address the aforementioned problems in the existing technology, this invention provides a multimodal generative safety detection method and system applied in the logistics field. The technical problem to be solved by this invention is achieved through the following technical solution: In a first aspect, embodiments of the present invention provide a multimodal generative safety detection method applied in the logistics field, the method comprising: Obtain scene text description information of real logistics scene images and feedback information from the previous detection. Based on the image generation method, obtain the effective generated images corresponding to the scene text description information. Merge all the effective generated images and real logistics scene images to obtain a fusion dataset. For each sample image in the fusion dataset, its features are decoupled into local feature streams and global feature streams; environmental risk situation classification is performed on the global feature streams to obtain the probability that the sample image belongs to each of K risk scene categories; based on the probability of each risk scene category, a set of dynamic detection parameter thresholds corresponding to the sample image is calculated; the obtained K sets of dynamic detection parameter thresholds are used to guide the detection of the local feature streams to obtain K sets of detection results for the sample image, each set of detection results including target location, target category, and target detection confidence; Based on the K sets of detection results and the probabilities of the K risk scenario categories, a safety warning decision is made; and an uncertainty assessment is performed on the K sets of detection results. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scenario, and feedback information is generated for the next detection.

[0014] In one embodiment of the present invention, obtaining the effective generated image corresponding to the scene text description information based on the image generation method includes: Using the feedback information, the scene text description information is mapped into a structured scene cue vector; wherein, the feedback information contains at least one specific scene feature identifier, used to indicate under what scene feature conditions the corresponding detection process produces unreliable detection results or fails; the scene cue vector is composed of multiple scene feature parameters, each scene feature parameter corresponding to a controllable scene feature and associated with a corresponding weight value, the weight value being used to characterize the degree of emphasis of the corresponding scene feature in the image generation process, and the weight value being adjusted according to the feedback information; Using a preset generation model, generate an image corresponding to the scene prompt vector; The generated image is verified, and if the verification is successful, it is determined to be a valid generated image.

[0015] In one embodiment of the present invention, the scene features in the scene cue vector include one or more of the following: scene lighting features, occlusion degree features, target motion state features, spatial relationship features between people and equipment, and the number or density of targets in the scene.

[0016] In one embodiment of the present invention, the generated image is verified, and if the verification passes, it is determined to be a valid generated image, including: The generated image undergoes semantic consistency verification and physical logic verification. If both verifications pass, the generated image is determined to be a valid generated image. The semantic consistency verification uses the CLIP model to calculate the cosine similarity between the generated image and the corresponding scene cue vector to filter out generated images whose cosine similarity does not meet the requirements. The physical logic verification uses DINOv3 to extract the spatial geometric features of the generated image and performs preset physical rule checks to filter out generated images whose physical rule checks do not meet the requirements.

[0017] In one embodiment of the present invention, the preset physical rules include: Gravity constraint rules are used to detect whether cargo is suspended. The proportional constraint rule is used to detect whether the size ratio of personnel to goods stacked is within a preset reasonable threshold range.

[0018] In one embodiment of the present invention, a set of dynamic detection parameter thresholds is calculated based on the probability of each risk scenario category, including: Based on the probability of each risk scenario category, a dynamic confidence threshold and a dynamic non-maximum suppression threshold are calculated using a preset risk-sensitivity mapping relationship. These thresholds constitute a corresponding set of dynamic detection parameter thresholds. The calculation formula used is as follows: ; in, The threshold for the generated dynamic detection parameters is denoted as , and the threshold for the dynamic confidence level is denoted as . or dynamic NMS threshold ; The basic confidence threshold; This is the adjustment coefficient; For activation functions; For the first The probability of each risk scenario category, Among them, the introduction The generated dynamic detection parameter thresholds include a dynamic confidence threshold and a dynamic non-maximum suppression threshold. Different base confidence thresholds, adjustment coefficients, or activation functions are used to calculate the dynamic confidence threshold and the dynamic non-maximum suppression threshold, respectively. In one embodiment of the present invention, a safety warning decision is made based on the K sets of detection results and the probabilities of the K risk scenario categories, including: When the probabilities of the K groups of detection results and the K risk scene categories each satisfy at least one warning trigger condition, the sample image is determined to have a safety hazard, and a warning signal is output; wherein, the at least one warning trigger condition includes: In the K sets of detection results, at least one candidate target belongs to a predefined high-risk category; In the K sets of detection results, the target detection confidence of at least one candidate target exceeds the preset security criterion threshold; Among the probabilities of the K risk scenario categories, there exists a predefined high-risk global scenario category.

[0019] In one embodiment of the present invention, satisfying the uncertainty condition includes: In the detection results corresponding to the same sample image, the detection results of the candidate target are unstable; or, When multiple frames of images corresponding to the sample images are available, the detection results at the same spatial location in adjacent images frequently appear or disappear, resulting in inconsistent detection results. In the detection results corresponding to the same sample image, the detection results of the candidate target are unstable, which is achieved by satisfying at least one of the following conditions: In the K groups of detection results corresponding to the same sample image, the target detection confidence of at least one candidate target is within a preset fuzzy range; In the K groups of detection results corresponding to the same sample image, the detection confidence distribution of at least one candidate target exhibits a high entropy state, indicating that the category determination of the candidate target is unclear.

[0020] In one embodiment of the present invention, when the uncertainty condition is met, the sample image is determined to be a difficult example or a cognitive blind spot scene, and feedback information is generated, including: When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scene. Based on the detection intermediate results corresponding to the sample image, feature information describing the reasons for unreliable detection or detection failure is extracted to form feedback parameters. The feedback parameters together with the sample image constitute feedback information. The feedback parameters include: A global scene feature description used to describe the overall environmental state of a sample image, wherein the global scene feature description is used to characterize scene attributes related to unreliable detection or detection failure; Key semantic feature identifiers are used to identify candidate targets or semantic regions that cause unstable detection results, wherein the key semantic feature identifiers are used to indicate semantic objects that have class confusion, confidence fluctuations or unclear judgments during the detection process.

[0021] Secondly, embodiments of the present invention provide a multimodal generative safety detection system applied in the logistics field, the system comprising: The sample enhancement module is used to obtain scene text description information of real logistics scene images and feedback information from the previous detection, obtain valid generated images corresponding to the scene text description information based on the image generation method, and merge all valid generated images and real logistics scene images to obtain a fused dataset. An adaptive perception module is used to decouple the features of each sample image in the fusion dataset into a local feature stream and a global feature stream; classify the global feature stream according to environmental risk situation to obtain the probability of the sample image belonging to K risk scene categories; calculate a set of dynamic detection parameter thresholds corresponding to the sample image based on the probability of each risk scene category; and use all the obtained sets of dynamic detection parameter thresholds to guide the detection of the local feature stream to obtain K sets of detection results for the sample image, each set of detection results including target location, target category, and target detection confidence. The decision feedback module is used to make a safety warning decision based on the K sets of detection results and the probabilities of the K risk scenario categories; and to evaluate the uncertainty of the K sets of detection results. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scenario, and feedback information is generated for the next detection.

[0022] The beneficial effects of this invention are: The multimodal generative safety detection method and system for logistics provided in this invention first acquires scene text description information of real logistics scene images and feedback information from the previous detection. Based on the image generation method, effective generated images corresponding to the scene text description information are obtained. All effective generated images and real logistics scene images are merged to obtain a fusion dataset, achieving sample augmentation. Second, for each sample image in the fusion dataset, its features are decoupled into local feature streams and global feature streams. The global feature streams are classified according to environmental risk status to obtain the probability of the sample image belonging to K risk scene categories. Based on the probability of each risk scene category, a set of dynamic detection parameter thresholds corresponding to the sample image is calculated. The obtained K sets of dynamic detection parameter thresholds are used to guide the detection of the local feature streams to obtain K sets of detection results for the sample image. Finally, based on the K sets of detection results and the probabilities of the K risk scene categories, a safety warning decision is made. The uncertainty of the K sets of detection results is evaluated. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scene, and feedback information is generated for the next detection. This invention constructs a dual closed loop of "generation-detection" to generate effective new sample images using image and text multimodal methods, solving the problem of lack of physical constraints in the samples in the prior art. Furthermore, for each sample image, a set of dynamic detection parameter thresholds is specifically calculated to guide local target detection, solving the problem that the detection threshold cannot dynamically adapt to the scene in the prior art. This improves the accuracy and flexibility of target detection and enhances the system's generalization ability to quickly switch to adapt to different scenes. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating a multimodal generative safety detection method applied in the logistics field, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the principle of the multimodal generative safety detection method applied to the logistics field provided in this embodiment of the invention. Figure 3 This is a schematic diagram illustrating the structural principle of a multimodal generative safety detection system applied in the logistics field, provided by an embodiment of the present invention. Detailed Implementation

[0024] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0025] Firstly, embodiments of the present invention provide a multimodal generative security detection method applied in the logistics field. Please refer to [link to relevant documentation]. Figure 1 and Figure 2 Understanding this, the method may include the following steps: S1, obtain the scene text description information of the real logistics scene image and the feedback information of the previous detection, obtain the effective generated image corresponding to the scene text description information based on the image generation method, and merge all the effective generated images and the real logistics scene image to obtain the fusion dataset. The real logistics scene images in this embodiment of the invention can be real scene images of logistics warehousing scenarios, such as images of warehouses or logistics parks. These scenarios often contain goods, as well as warehousing equipment such as shelves and handling tools, and may also include personnel. This embodiment of the invention can pre-collect real images of various specific logistics scenarios to form a real image set. Furthermore, semantic recognition and other technologies are used to describe the scene with textual description information for each real logistics scene image in the real image set. The scene textual description information uses language to describe the content of the corresponding real logistics scene image. For example, it could be: in a warehouse, a worker wearing a safety helmet is facing a shelf and moving goods; there are multiple layers of shelves in front of and behind him, and goods are piled on the ground waiting to be placed.

[0026] Simultaneously, if feedback information was output during the previous use of this multimodal generative security detection method applied in the logistics field, it must also be obtained now. The feedback information includes at least one specific scenario feature identifier, used to indicate under what scenario feature conditions the corresponding detection process (i.e., the previous detection) produced an unreliable detection result or failed; the generation process and details of the feedback information will be explained in detail later.

[0027] In this embodiment of the invention, after obtaining the scene text description information of a real logistics scene image and the feedback information from the previous detection, a valid generated image corresponding to the scene text description information is obtained based on the image generation method, including the following steps: 1) Using the feedback information, the scene text description information is mapped into a structured scene prompt vector; The mapping method can employ existing related technologies. The resulting scene cue vector consists of multiple scene feature parameters. Each scene feature parameter corresponds to a controllable scene feature and is associated with a corresponding weight value. The weight value is used to characterize the degree of emphasis of the corresponding scene feature in the image generation process, and the weight value is adjusted according to the feedback information. The weight value will serve as the input control parameter for the subsequent generation model.

[0028] The scene features in the scene cue vector include one or more of the following: scene lighting features, occlusion features, target motion state features, spatial relationship features between people and equipment, and the number or density of targets in the scene. Targets can be people, equipment, etc.

[0029] Therefore, it can be understood that each scene text description, combined with available feedback information, generates a structured scene cue vector. Of course, if there is no feedback information, a structured scene cue vector can be generated solely using the scene text description information. Specifically, in the initial state of system execution or when feedback information is empty, the weight values ​​adopt a preset balanced distribution (or default weights), directly generating a structured scene cue vector based on the scene text description information, and subsequently producing the generated image.

[0030] 2) Using a preset generation model, generate the image corresponding to the scene prompt vector; This step aims to generate a virtual image with specific working condition characteristics from each real logistics scene image, based on its scene description information and the feedback information obtained.

[0031] This can be achieved using any existing generative model. In one optional implementation, the preset generative model is based on a Diffusion architecture, such as the StableDiffusion generative model, and is loaded with LoRA (Low-Rank Adaptation) weights fine-tuned based on real logistics scene images (LoRA weights are model parameters obtained by fine-tuning the generative model based on real logistics scene images during the training phase; they are internal model parameters). This gives the generative model "logistics domain knowledge," enabling it to understand the shapes of specific objects such as employees, shelves, pallets, and forklifts, rather than generating deformed objects using a general model.

[0032] 3) Verify the generated image, and determine it as a valid generated image if the verification is successful.

[0033] To ensure the quality of the generated image, this embodiment of the invention employs a dual verification method of physical logic and semantics for parallel verification.

[0034] Specifically, the steps include: The generated image undergoes semantic consistency and physical logic verification. If both verifications pass, it is determined to be a valid generated image. The semantic consistency check uses the CLIP model to calculate the cosine similarity between the generated image and the corresponding scene cue vector, and filters out generated images whose cosine similarity does not meet the requirements. The physical logic verification uses DINOv3 to extract the spatial geometric features of the generated image and performs preset physical rule checks to filter out generated images that do not meet the physical rule check requirements.

[0035] The preset physical rules include: Gravity constraint rules are used to detect whether cargo is suspended; that is, whether the bottom of the cargo's bounding box (i.e., the lower edge of the part of the cargo identified by the bounding box) is in contact with the supporting surface. Specifically, it is determined whether the cargo is suspended by calculating the Euclidean distance between the bottom edge of the target bounding box and the feature of the supporting surface.

[0036] The proportional constraint rule is used to detect whether the size ratio of personnel to goods stacking is within a preset reasonable threshold range, such as between 1:0.3 and 1:3.

[0037] Only generated images that pass both physical logic and semantic verification are considered valid generated images. All valid generated images are merged with real logistics scene images to obtain a fused dataset for subsequent training.

[0038] S2, for each sample image in the fusion dataset, its features are decoupled into local feature streams and global feature streams; environmental risk situation classification is performed on the global feature streams to obtain the probability that the sample image belongs to each of K risk scene categories; based on the probability of each risk scene category, a set of dynamic detection parameter thresholds corresponding to the sample image is calculated; the obtained K sets of dynamic detection parameter thresholds are used to guide the detection of the local feature streams to obtain K sets of detection results for the sample image, each set of detection results including target location, target category, and target detection confidence; Please see Figure 2 In this embodiment of the invention, feature decoupling can be implemented using a feature decoupling network. For example, in one optional implementation, the feature decoupling network can use DINOv3 as the backbone network, which decouples the features of the input sample image into two data streams. The decoupling process is described in detail in the relevant technical descriptions. The local feature stream retains high-resolution spatial location information and is used to characterize the fine-grained features of local targets in the sample image. The CLS Token output by DINOv3 is used as the global feature stream. The global feature stream is formed by pooling the features to form a low-dimensional global feature representation, which is used to characterize the overall environmental semantic information of the image, such as light intensity and target density. The global feature stream is used to characterize the environmental risk situation.

[0039] K risk scenario categories can be pre-defined, such as scenarios prone to missed detection due to high-density stacking and insufficient lighting. For example, these could include categories of items and people in dimly lit areas of a warehouse, objects that are pressed to the bottom when goods are placed and stacked, high-safety-risk scenarios that are close to equipment, complex operation scenarios with a large number of targets or severe obstruction, and dynamic scenarios where targets are in rapid movement or employees are changing postures.

[0040] Environmental risk situation classification of the global feature stream can be achieved using an environmental risk situation classification model. For example, this model could be composed of a multilayer perceptron (MLP). After inputting the global feature stream, the probability of each sample image belonging to one of K risk scene categories can be obtained as the global scene recognition result. Here, the probability of the Kth risk scene category is... The probability corresponding to each risk scenario category is: , The environmental risk situation classification model can be trained using several global feature stream samples representing labels. The labels are the probabilities of each of the labeled global feature stream samples belonging to one of K risk scenario categories. For a detailed understanding of the specific training process, please refer to the training process of a conventional neural network; it will not be elaborated upon here.

[0041] Based on the probability of each risk scenario category, a set of dynamic detection parameter thresholds corresponding to the sample image is calculated. This process transforms the global scene cognition results into control instructions for the local feature flow detection process. This process specifically includes: Based on the probability of each risk scenario category, a dynamic confidence threshold and a dynamic non-maximum suppression threshold are calculated using a preset risk-sensitivity mapping relationship. These thresholds constitute a corresponding set of dynamic detection parameter thresholds. The calculation formula used is as follows: ; in, The threshold for the generated dynamic detection parameters is denoted as , and the threshold for the dynamic confidence level is denoted as . or dynamic NMS threshold ; The basic confidence threshold; This is the adjustment coefficient; For activation functions; For the first The probability of each risk scenario category, Among them, the introduction The generated dynamic detection parameter thresholds include a dynamic confidence threshold and a dynamic non-maximum suppression threshold. Different base confidence thresholds, adjustment coefficients, or activation functions are used to calculate the dynamic confidence threshold and the dynamic non-maximum suppression threshold, respectively. It should be noted that the "risk-sensitivity mapping relationship" in this embodiment refers to the correspondence in which the detection sensitivity parameters used in the target detection process are adjusted according to the probability that the sample image belongs to different risk scene categories. Detection sensitivity parameters include, but are not limited to, the target detection confidence threshold and the non-maximum suppression threshold. Through this mapping relationship, detection sensitivity is increased in high-risk scenes, and maintained or decreased in low-risk scenes. The above calculation formula is only one optional implementation of the risk-sensitivity mapping relationship, used to illustrate how to continuously or non-linearly adjust the detection parameters according to the probability of the risk scene category, and does not constitute a limitation on the risk-sensitivity mapping relationship.

[0042] In other words, in the embodiments of this invention, the probability of each risk scenario category A dynamic confidence threshold will be obtained. A dynamic NMS threshold , and All of these are calculated using the aforementioned general formula, although they are all based on the probability of the same risk scenario category. The calculations are performed, but the risk-sensitivity mapping relationships for the two are based on different baseline confidence thresholds. and adjustment coefficient This is to accommodate the functional differences of different detection parameters in the target detection process. Therefore, even for the same sample image and the probability of the same risk scene category, the calculated dynamic confidence threshold and dynamic non-maximum suppression threshold are different, thus ensuring the rationality and effectiveness of the detection parameter adjustment.

[0043] It is worth noting that the dynamic confidence threshold Determined based on the corresponding sample image, with the first The probability of each risk scenario category For example, if its value is higher than a certain threshold, the sample image is determined to be a high-risk or easily missed scenario. In this case, the dynamic confidence threshold calculated using the above formula is used. It will automatically decrease, for example, to 0.3, thereby forcing subsequent local object detection to improve recall, so that low-confidence but potentially real targets are retained during detection; conversely, if If the value is below the judgment threshold, the sample image is determined to be a complex background or a scene with a high risk of false alarms. In this case, the dynamic confidence threshold calculated using the above formula is used. It will automatically improve to eliminate suspected noisy targets and suppress false alarms during subsequent local target detection.

[0044] In this embodiment of the invention, DINOv3's Patch Tokens are used as local feature streams to preserve high-resolution spatial geometric features for target detection. Subsequently, the K sets of dynamic detection parameter thresholds obtained from the sample image are transmitted in the form of control instruction packets to guide the detection of subsequent local feature streams, i.e., to perform local target detection.

[0045] The process of local target detection includes: 1) Perform RoI feature alignment and extraction: The system receives a local feature stream and processes it using a Region Proposal Network (RPN) to generate a series of candidate boxes (Proposals). Each candidate box can be represented by the coordinates of its upper left and lower right corners, which correspond to the spatial scale of the original input image.

[0046] The Roi Align layer and local feature flow are used to extract corresponding feature map patches. The Roi Align layer employs bilinear interpolation to uniformly map candidate boxes of different sizes to feature vectors of a fixed size (e.g., 7). 7 256), to eliminate quantization errors and ensure that features of small objects (such as cargo tags at a distance) are preserved intact. The final output is the RoI feature vector. RoI stands for Region of Interest, referring to a local location. For specific processing procedures, please refer to relevant technical explanations, which will not be detailed here.

[0047] 2) Processing with a dual-head prediction network: The feature vectors of the RoI are received, passed through several fully connected layers, and then split into two parallel output branches: Classification Branch (Cls Head): Outputs the raw probability distribution P_raw (values ​​before or after Softmax) of each candidate box belonging to each target category (such as person, forklift, goods, background).

[0048] Regression Branch (Reg Head): Outputs the coordinate offset of each candidate box relative to the real object. , here The values ​​are the pixel coordinates of the top left, top right, bottom left, and bottom right corners of the candidate box, respectively, used to correct the bounding box position.

[0049] For details on the specific processing procedure, please refer to the relevant technical explanations; they will not be elaborated upon here.

[0050] Final output classification score And bounding box regression parameters. Those skilled in the art will understand that in object detection, each candidate box detected in an image will have a classification score, the numerical range of which is... In this embodiment of the invention, the classification score is output by the classification branch of the object detection network and is used to characterize the prediction result of whether the candidate box belongs to each object category; the object detection confidence is determined based on the classification score and is used to determine whether the candidate box constitutes a valid detection result, and is used as a criterion for comparison with the dynamic confidence threshold. In one embodiment, the object detection confidence can be directly obtained from the classification score; in other embodiments, the object detection confidence can also be obtained by further processing based on the classification score.

[0051] After obtaining the classification score and bounding box regression parameters, this invention does not directly use fixed detection parameters to filter candidate boxes. Instead, it introduces a dynamic detection parameter threshold based on the risk scene category probability corresponding to the sample image to post-process the local target detection results.

[0052] Specifically, in the candidate box filtering stage, the classification score output by the classification branch is compared with the dynamic confidence threshold, and only candidate targets whose classification scores meet the dynamic confidence threshold are retained; In the non-maximum suppression stage, a dynamic non-maximum suppression threshold, which is dynamically calculated based on the probability of risk scenario categories, is used to suppress the overlapping relationship between candidate targets.

[0053] In this way, the screening threshold can be lowered to improve target recall in high-risk or high-risk scenarios, while the screening threshold can be raised to suppress noisy targets in complex backgrounds or high false alarm risk scenarios, thereby achieving adaptive adjustment of detection sensitivity without changing the detection network structure.

[0054] 3) Dynamic parameter adaptation detection: This invention employs controlled dynamic post-processing logic. Compared to traditional detectors that use fixed configuration files in this step, this invention executes a real-time instruction response mechanism, using dynamic detection parameter thresholds of the corresponding sample image to guide local feature flow detection.

[0055] First, the received control command packet is parsed to obtain K sets of dynamic detection parameter thresholds corresponding to the sample image (each set contains...). +Dynamic NMS threshold ); Then, adaptive filtering is performed, specifically, the classification score of each candidate box is iterated through. Based on the dynamic confidence threshold obtained from the analysis Filter by category score Greater than the dynamic confidence threshold The candidate box.

[0056] Due to dynamic confidence threshold It was determined based on the sample image, therefore, targeted detection can be achieved compared to different sample images.

[0057] This invention employs controlled dynamic post-processing logic. Compared to traditional target detectors that use fixed configuration files in the post-processing stage, this invention uses a real-time command response mechanism to adaptively process local target detection results by utilizing dynamic detection parameter thresholds calculated for the current sample image.

[0058] Specifically, the received control instruction packet is first parsed to obtain K sets of dynamic detection parameter thresholds corresponding to the sample image, wherein each set of dynamic detection parameter thresholds includes at least a dynamic confidence threshold and a dynamic non-maximum suppression threshold.

[0059] Then, adaptive filtering is performed based on the dynamic confidence threshold: the candidate box set output by the local object detection network is traversed, and the corresponding classification score is obtained for each candidate box. and the classification score The classification score is compared with the corresponding dynamic confidence threshold and retained. Candidate boxes with a value greater than the dynamic confidence threshold are used to form a preliminary set of candidate targets.

[0060] Subsequently, for the candidate target set after preliminary screening, non-maximum suppression processing is performed using the dynamic non-maximum suppression threshold to remove candidate targets that are spatially highly overlapping and have low confidence, thereby obtaining the final detection result for the current sample image.

[0061] The final detection result includes the target location, target category, and corresponding classification score of the candidate target. This will be used as one output of the K sets of detection results for subsequent safety warning decisions and uncertainty assessment processes.

[0062] Since the dynamic detection parameter threshold is adaptively determined based on the risk scene category probability of the current sample image, the detection sensitivity can be adjusted in a targeted manner for different sample images, thereby improving the reliability of the detection results without changing the detection network structure.

[0063] Secondly, candidate box deduplication is performed, which uses a dynamic NMS threshold. Non-maximum suppression is performed to adapt to scenarios with varying target density.

[0064] For the candidate target set after preliminary screening, non-maximum suppression operation is performed using a dynamic non-maximum suppression threshold determined based on the current sample image. When the overlap between any two candidate targets is greater than the dynamic non-maximum suppression threshold, the candidate target with higher target detection confidence is retained, while the candidate target with lower target detection confidence is suppressed to eliminate duplicate detection results.

[0065] After the confidence threshold filtering and non-maximum suppression processing described above, K sets of detection results corresponding to the current sample image are output. Each set of detection results corresponds to a risk scene category, and each set of detection results contains information on multiple detection targets.

[0066] After these two steps, K sets of detection results for the current sample image are output. Each set of detection results includes the target location, target category, and target detection confidence score, represented as: ; in, Indicates the first Each group of test results corresponds to a risk scenario category; because there are multiple targets, It will contain multiple sets of detection information. ; Indicates the first The location of a target is usually represented by the coordinates of the bounding box containing the target, which is used to characterize the position of the target in the image; Indicates the first Each target category can be considered a category label; Indicates the first Target detection confidence. The specific process of local feature flow detection using confidence thresholds and NMS thresholds can be understood by referring to relevant technologies, and will not be detailed here. The key point of the embodiments of the present invention is... and Based on the corresponding sample images, the threshold dynamically changes during the detection process of multiple sample images.

[0067] S3. Based on the K sets of detection results and the probabilities of the K risk scenario categories, a safety warning decision is made; and the uncertainty of the K sets of detection results is evaluated. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scenario, and feedback information is generated for the next detection.

[0068] S3 does not trigger alarms solely based on the detection confidence of a single target, but rather considers multiple factors comprehensively. Specifically, it makes a safety warning decision based on the K sets of detection results and the probabilities of the K risk scenario categories, including: When the probabilities of the K groups of detection results and the K risk scene categories meet at least one warning triggering condition, the sample image is determined to have a safety hazard, and a warning signal is output. The at least one warning triggering condition includes: 1) In the K sets of detection results, at least one candidate target belongs to a predefined high-risk category; This invention allows for the predefinition of one or more high-risk categories. These high-risk categories can be predefined based on the safety requirements of logistics operations, and may include, but are not limited to: personnel targets within the work area; operating forklifts, automated guided vehicles, or other mobile equipment; goods located at heights or in unstable conditions; targets located less than a preset safe distance from the work equipment, etc.

[0069] 2) In the K groups of detection results, the target detection confidence of at least one candidate target exceeds the preset security criterion threshold; In this embodiment of the invention, the security criterion threshold can be set according to the detection accuracy and security requirements, for example, it can be set in the range of 0.6 to 0.8. When the target detection confidence of the candidate target is higher than the security criterion threshold, the candidate target is determined to have a high security risk.

[0070] 3) Among the probabilities of the K risk scenario categories, there exists a predefined high-risk global scenario category.

[0071] This section displays the probabilities of K risk scenario categories. Are there predefined high-risk global scenario categories?

[0072] In this embodiment of the invention, the high-risk global scenario category is used to characterize scenario modes where the overall working environment has a high security risk, and may include, but is not limited to: Scenarios where high-density goods are stacked and lighting is insufficient, making them prone to missed inspections; High-risk collaborative scenarios where personnel and equipment are in close proximity; Complex operational scenarios with a high density of targets and severe obstruction within the work area; Dynamic scenes where the target is in a state of rapid movement or frequent interaction.

[0073] When at least one of the above three warning triggering conditions is met, it is determined that there is a safety hazard in the current scene corresponding to the sample image. Then, a corresponding audible and visual alarm signal or equipment shutdown control command is output through the safety response interface as a warning signal. When outputting the warning signal, the following can be utilized: Identify the spatial area where the risk occurred, mark it on the sample image, and notify management and inspection personnel of the location and area of ​​the violation.

[0074] In S3, when the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scene, and feedback information is generated. This involves identifying cognitively unstable or uncertain sample images from K sets of detection results and generating feedback information for them, which is then used to generate the next image.

[0075] Among them, satisfying the uncertainty condition includes at least one of the following cases (1) or (2): (1) In the detection results corresponding to the same sample image, the detection results of the candidate target are unstable; at this time, it is impossible to accurately determine the detection results. This situation occurs when at least one of the following conditions is met: In the K groups of detection results corresponding to the same sample image, the target detection confidence of at least one candidate target falls within a preset fuzzy range; for example... ; In the K groups of detection results corresponding to the same sample image, the detection confidence distribution of at least one candidate target exhibits a high-entropy state, indicating that the category determination of the candidate target is unclear. High entropy means that the model cannot distinguish which category a candidate target belongs to; the predicted probabilities, i.e., the confidence levels, of each category are relatively close, and there is no obvious dominant category. Therefore, the category determination of this target has significant uncertainty.

[0076] Understandably, this situation is based on the uncertainty judgment of a single frame image (high entropy state). Specifically, when only a single frame sample image is obtained, if the predicted probabilities of each target category in the category prediction results of a candidate target are relatively close and no obvious dominant category is formed, it indicates that the detection model has difficulty in making a clear determination of the category of the candidate target. For example, when the target is partially occluded, at a distance, or in poor lighting conditions, the detection model may simultaneously assign similar prediction probabilities to categories such as "personnel," "forklift," and "goods," resulting in a large uncertainty in the category determination of the candidate target.

[0077] (2) When multiple frames of images corresponding to the sample images are available, the detection results at the same spatial location in adjacent images frequently appear or disappear, which is a result of inconsistent detection results. The above situation (1) can be determined with only one frame image, while situation (2) requires the current frame image to be determined together with the previous multiple frames image. If the detection results of the same spatial position in adjacent images frequently appear or disappear, which is manifested as detection flicker, then the detection results are inconsistent.

[0078] Understandably, this situation is based on uncertainty determination (inconsistent detection results) across multiple frames. Specifically, given a sequence of multiple frames corresponding to the sample images, if the detection result at the same spatial location frequently appears or disappears in adjacent images, exhibiting unstable or flickering detection results, then it can be determined that the detection result at that location has significant uncertainty. For example, when the target appears and disappears intermittently due to obstruction by mobile devices or goods, or in scenes with rapidly changing lighting conditions, the detection model may produce inconsistent detection results for the same target in consecutive frames.

[0079] When at least one of conditions (1) and (2) is met, the uncertainty condition is satisfied. When the uncertainty condition is satisfied, the sample image is determined to be a difficult sample or a cognitive blind spot scene. At this time, the original sample image is not directly used as feedback information. Instead, based on the detection intermediate result corresponding to the sample image, feature information describing the reasons for unreliable detection or detection failure is extracted to form feedback parameters. The feedback parameters together with the sample image constitute the feedback information.

[0080] The feedback parameters include: A global scene feature description used to describe the overall environmental state of a sample image, wherein the global scene feature description is used to characterize scene attributes related to unreliable detection or detection failure, such as illumination state, target density, etc. Key semantic feature identifiers are used to identify candidate targets or semantic regions that cause unstable detection results. These key semantic feature identifiers are used to indicate semantic objects that are subject to category confusion, confidence fluctuations, or unclear judgments during the detection process, such as black goods, strong backlighting, or severe occlusion.

[0081] The aforementioned feedback parameters are used in the next detection process to adjust the weight values ​​of the corresponding scene feature parameters in the scene cue vector, thereby triggering a new round of targeted sample (generated image) generation.

[0082] As mentioned above, the scene features in the scene cue vector include one or more of the following: scene lighting features, occlusion degree features, target motion state features, spatial relationship features between people and equipment, and the number or density of targets in the scene. Each scene feature in the scene cue vector has a corresponding weight value, which is used to characterize the degree of emphasis of the scene feature in the image generation process.

[0083] In this embodiment of the invention, the feedback parameter corresponds to at least a portion of the scene features contained in the scene prompt vector, that is, the features contained in the feedback parameter are a subset of the scene features in the scene prompt vector.

[0084] In the next detection or sample generation process, the weight values ​​of the scene feature parameters corresponding to the feedback parameters in the scene cue vector are adjusted according to the feedback parameters, specifically as follows: The weight values ​​of the scene feature parameters corresponding to the feedback parameters are relatively increased to enhance the expression of such scene features during image generation.

[0085] For example, when the feedback parameter indicates that the detection is unreliable or the detection failure mainly occurs in scenes with insufficient lighting or severe occlusion, the weight values ​​of the corresponding lighting feature parameters and occlusion feature parameters in the scene cue vector are increased when generating the next image, so that the generated image more centrally reflects the complex scene features such as low lighting and high occlusion. When the feedback parameters indicate that the detection instability mainly occurs in scenarios where personnel and equipment are close together or where targets are densely packed, the weight values ​​corresponding to the spatial relationship features between personnel and equipment or the density features of targets are increased to generate sample images that are closer to the actual high-risk working environment.

[0086] By employing the above methods, the generated images can more effectively cover typical scene features that lead to unreliable or failed detection, thereby avoiding the recurrence of detection blind spots under the same or similar scene conditions in subsequent detection processes and improving the robustness of the detection system to complex logistics scenarios.

[0087] The multimodal generative safety detection method for logistics provided in this invention first acquires scene text description information of real logistics scene images and feedback information from the previous detection. Based on the image generation method, it obtains valid generated images corresponding to the scene text description information. All valid generated images and real logistics scene images are merged to obtain a fusion dataset, achieving sample augmentation. Second, for each sample image in the fusion dataset, its features are decoupled into local feature streams and global feature streams. The global feature streams are classified according to environmental risk status to obtain the probability of the sample image belonging to K risk scene categories. Based on the probability of each risk scene category, a set of dynamic detection parameter thresholds corresponding to the sample image is calculated. The obtained K sets of dynamic detection parameter thresholds are used to guide the detection of the local feature streams to obtain K sets of detection results for the sample image. Finally, based on the K sets of detection results and the probabilities of the K risk scene categories, a safety warning decision is made. The uncertainty of the K sets of detection results is evaluated. When the uncertainty condition is met, the sample image is determined to be a difficult example sample or a cognitive blind spot scene, and feedback information is generated for the next detection. This invention constructs a dual closed loop of "generation-detection" to generate effective new sample images using image and text multimodal methods, solving the problem of lack of physical constraints in the samples in the prior art. Furthermore, for each sample image, a set of dynamic detection parameter thresholds is specifically calculated to guide local target detection, solving the problem that the detection threshold cannot dynamically adapt to the scene in the prior art, thus improving the accuracy and flexibility of target detection.

[0088] Meanwhile, the present invention adaptively calculates a set of dynamic detection parameter thresholds for different sample images based on their environmental risk characteristics, and uses these dynamic detection parameter thresholds to guide the local target detection process, thereby overcoming the shortcomings of the prior art where the detection threshold is fixed and cannot be dynamically adjusted according to scene changes.

[0089] Through the above technical solutions, the present invention can achieve flexible and reliable detection of high-risk targets and abnormal scenarios in complex and ever-changing logistics operation environments, effectively improve the recall capability of target detection in high-risk scenarios, and suppress false detections under complex background conditions, thereby improving the overall accuracy, stability and scenario adaptability of the safety detection system.

[0090] Secondly, corresponding to the above method embodiments, this invention also provides a multimodal generative safety detection system applied in the logistics field, such as... Figure 3 As shown, the system includes: The sample enhancement module is used to obtain scene text description information of real logistics scene images and feedback information from the previous detection, obtain valid generated images corresponding to the scene text description information based on the image generation method, and merge all valid generated images and real logistics scene images to obtain a fused dataset. An adaptive perception module is used to decouple the features of each sample image in the fusion dataset into a local feature stream and a global feature stream; classify the global feature stream according to environmental risk situation to obtain the probability of the sample image belonging to K risk scene categories; calculate a set of dynamic detection parameter thresholds corresponding to the sample image based on the probability of each risk scene category; and use all the obtained sets of dynamic detection parameter thresholds to guide the detection of the local feature stream to obtain K sets of detection results for the sample image, each set of detection results including target location, target category, and target detection confidence. The decision feedback module is used to make a safety warning decision based on the K sets of detection results and the probabilities of the K risk scenario categories; and to evaluate the uncertainty of the K sets of detection results. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scenario, and feedback information is generated for the next detection.

[0091] The sample augmentation module acts as a "decision-maker with common sense about physics," generating high-quality training data to address the system's detection weaknesses. The sample augmentation module includes a semantic parameter mapping unit, a knowledge-driven synthesis unit, and a logical-semantic dual checker.

[0092] The semantic parameter mapping unit has two input terminals. The first input terminal acquires the scene text description information of the real logistics scene image, and the second input terminal acquires the feedback information returned by the adaptive perception module regarding the existence of the previous detection. Using the feedback information, the scene text description information is mapped into a structured scene cue vector.

[0093] The knowledge-driven synthesis unit uses a preset generation model to generate an image corresponding to the scene prompt vector.

[0094] The logical-semantic dual verifier performs parallel verification of the generated image using both physical logic and semantic dual verification methods. Once the verification is passed, the image is determined to be a valid generated image.

[0095] Then, all the valid generated images and real logistics scene images are merged to obtain a fused dataset, which is then transmitted to the adaptive perception module.

[0096] The adaptive perception module is the core control center of this invention. It is not only responsible for detection, but also for "adjusting vision" according to the environment. Its input is each sample image in the fused dataset.

[0097] The adaptive perception module includes a feature decoupling extraction unit, a global scene cognition unit, a dynamic parameter adapter, and a local target detection unit.

[0098] The feature decoupling extraction unit decouples the features of the input sample image into a local feature flow and a global feature flow.

[0099] The global scene cognition unit performs environmental risk situation classification on the global feature stream to obtain the probability of the sample image belonging to each of the K risk scene categories.

[0100] The dynamic parameter adapter calculates a set of dynamic detection parameter thresholds corresponding to the sample image based on the probability of each risk scenario category.

[0101] The local target detection unit uses the obtained K sets of dynamic detection parameter thresholds to guide the detection of the local feature flow, and obtains K sets of detection results for the sample image. Each set of detection results includes the target location, target category and target detection confidence.

[0102] The decision feedback module includes a result fusion and early warning unit, as well as an uncertainty assessment unit.

[0103] The result fusion and early warning unit makes a safety early warning judgment based on the K sets of detection results and the probabilities of the K risk scene categories. When it is determined that the sample image has a safety hazard, it outputs an early warning signal.

[0104] The uncertainty assessment unit is used to determine the cognitive stability of the system in the current image. Specifically, it performs uncertainty assessment on the K groups of detection results. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scene, and feedback information is generated. The feedback information is transmitted to the semantic parameter mapping unit for the next detection.

[0105] For details on the specific processing procedures of each module of this system, please refer to the relevant content in the first section, which will not be elaborated here.

[0106] It should be noted that in the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0107] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0108] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0109] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A multimodal generative safety detection method applied in the logistics field, characterized in that, include: Obtain scene text description information of real logistics scene images and feedback information from the previous detection. Based on the image generation method, obtain the effective generated images corresponding to the scene text description information. Merge all the effective generated images and real logistics scene images to obtain a fusion dataset. For each sample image in the fusion dataset, its features are decoupled into local feature streams and global feature streams; environmental risk situation classification is performed on the global feature streams to obtain the probability that the sample image belongs to each of K risk scene categories; based on the probability of each risk scene category, a set of dynamic detection parameter thresholds corresponding to the sample image is calculated; the obtained K sets of dynamic detection parameter thresholds are used to guide the detection of the local feature streams to obtain K sets of detection results for the sample image, each set of detection results including target location, target category, and target detection confidence; Based on the K sets of detection results and the probabilities of the K risk scenario categories, a safety warning decision is made; and an uncertainty assessment is performed on the K sets of detection results. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scenario, and feedback information is generated for the next detection.

2. The multimodal generative safety detection method applied to the logistics field according to claim 1, characterized in that, The process of obtaining the effective generated image corresponding to the scene text description information based on the image generation method includes: Using the feedback information, the scene text description information is mapped into a structured scene cue vector; wherein, the feedback information contains at least one specific scene feature identifier, used to indicate under what scene feature conditions the corresponding detection process produces unreliable detection results or fails; the scene cue vector is composed of multiple scene feature parameters, each scene feature parameter corresponding to a controllable scene feature and associated with a corresponding weight value, the weight value being used to characterize the degree of emphasis of the corresponding scene feature in the image generation process, and the weight value being adjusted according to the feedback information; Using a preset generation model, generate an image corresponding to the scene prompt vector; The generated image is verified, and if the verification is successful, it is determined to be a valid generated image.

3. The multimodal generative safety detection method for application in the logistics field according to claim 2, characterized in that, The scene features in the scene prompt vector include one or more of the following: scene lighting features, occlusion features, target motion state features, spatial relationship features between people and equipment, and the number or density of targets in the scene.

4. The multimodal generative safety detection method applied to the logistics field according to claim 2, characterized in that, The generated image is validated, and if the validation passes, it is determined to be a valid generated image, including: The generated image undergoes semantic consistency verification and physical logic verification. If both verifications pass, the generated image is determined to be a valid generated image. The semantic consistency verification uses the CLIP model to calculate the cosine similarity between the generated image and the corresponding scene cue vector to filter out generated images whose cosine similarity does not meet the requirements. The physical logic verification uses DINOv3 to extract the spatial geometric features of the generated image and performs preset physical rule checks to filter out generated images whose physical rule checks do not meet the requirements.

5. The multimodal generative safety detection method for application in the logistics field according to claim 4, characterized in that, The preset physical rules include: Gravity constraint rules are used to detect whether cargo is suspended. The proportional constraint rule is used to detect whether the size ratio of personnel to goods stacked is within a preset reasonable threshold range.

6. The multimodal generative safety detection method for application in the logistics field according to claim 1, characterized in that, Based on the probability of each risk scenario category, a set of dynamic detection parameter thresholds is calculated, including: Based on the probability of each risk scenario category, a dynamic confidence threshold and a dynamic non-maximum suppression threshold are calculated using a preset risk-sensitivity mapping relationship. These thresholds constitute a corresponding set of dynamic detection parameter thresholds. The calculation formula used is as follows: ; in, The threshold for the generated dynamic detection parameters is denoted as , and the threshold for the dynamic confidence level is denoted as . or dynamic NMS threshold ; The basic confidence threshold; This is the adjustment coefficient; For activation functions; For the first The probability of each risk scenario category, Among them, the introduction The generated dynamic detection parameter thresholds include a dynamic confidence threshold and a dynamic non-maximum suppression (NMS) threshold; different base confidence thresholds are used for the dynamic confidence threshold and the dynamic non-maximum suppression threshold, respectively. Adjustment coefficient Or use activation functions for calculation .

7. The multimodal generative safety detection method for application in the logistics field according to claim 1, characterized in that, Based on the K sets of detection results and the probabilities of the K risk scenario categories, a security warning decision is made, including: When the probabilities of the K groups of detection results and the K risk scene categories each satisfy at least one warning trigger condition, the sample image is determined to have a safety hazard, and a warning signal is output; wherein, the at least one warning trigger condition includes: In the K sets of detection results, at least one candidate target belongs to a predefined high-risk category; In the K sets of detection results, the target detection confidence of at least one candidate target exceeds the preset security criterion threshold; Among the probabilities of the K risk scenario categories, there exists a predefined high-risk global scenario category.

8. The multimodal generative safety detection method applied to the logistics field according to claim 1, characterized in that, The conditions for satisfying uncertainty include: In the detection results corresponding to the same sample image, the detection results of the candidate target are unstable; or, When multiple frames of images corresponding to the sample images are available, the detection results at the same spatial location in adjacent images frequently appear or disappear, resulting in inconsistent detection results. In the detection results corresponding to the same sample image, the detection results of the candidate target are unstable, which is achieved by satisfying at least one of the following conditions: In the K groups of detection results corresponding to the same sample image, the target detection confidence of at least one candidate target is within a preset fuzzy range; In the K groups of detection results corresponding to the same sample image, the detection confidence distribution of at least one candidate target exhibits a high entropy state, indicating that the category determination of the candidate target is unclear.

9. The multimodal generative safety detection method for application in the logistics field according to claim 1 or 8, characterized in that, When the uncertainty condition is met, the sample image is identified as a difficult example or a cognitive blind spot scene, and feedback information is generated, including: When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scene. Based on the detection intermediate results corresponding to the sample image, feature information describing the reasons for unreliable detection or detection failure is extracted to form feedback parameters. The feedback parameters together with the sample image constitute feedback information. The feedback parameters include: A global scene feature description used to describe the overall environmental state of a sample image, wherein the global scene feature description is used to characterize scene attributes related to unreliable detection or detection failure; Key semantic feature identifiers are used to identify candidate targets or semantic regions that cause unstable detection results, wherein the key semantic feature identifiers are used to indicate semantic objects that have class confusion, confidence fluctuations or unclear judgments during the detection process.

10. A multimodal generative safety detection system applied in the logistics field, characterized in that, include: The sample enhancement module is used to obtain scene text description information of real logistics scene images and feedback information from the previous detection, obtain valid generated images corresponding to the scene text description information based on the image generation method, and merge all valid generated images and real logistics scene images to obtain a fused dataset. An adaptive perception module is used to decouple the features of each sample image in the fusion dataset into a local feature stream and a global feature stream; classify the global feature stream according to environmental risk situation to obtain the probability of the sample image belonging to K risk scene categories; calculate a set of dynamic detection parameter thresholds corresponding to the sample image based on the probability of each risk scene category; and use all the obtained sets of dynamic detection parameter thresholds to guide the detection of the local feature stream to obtain K sets of detection results for the sample image, each set of detection results including target location, target category, and target detection confidence. The decision feedback module is used to make a safety warning decision based on the K sets of detection results and the probabilities of the K risk scenario categories; and to evaluate the uncertainty of the K sets of detection results. When the uncertainty condition is met, the sample image is determined to be a difficult sample or a cognitive blind spot scenario, and feedback information is generated for the next detection.