An experimental sample intelligent access management system based on machine vision
By using RFID and an improved Florence-2 multimodal vision model for multimodal fusion verification, the problem of accurately perceiving sample identity and location information in high-density sample storage environments was solved, achieving high reliability and consistency in sample identification and avoiding mismatch and misplacement identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENYANG XINGYA CHUANGWEI TECH DEV CO LTD
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to achieve highly reliable fusion perception of experimental sample identity information and spatial location information in high-density sample storage environments. They lack unified association matching and consistency judgment of multi-source heterogeneous information and an effective verification mechanism, leading to frequent problems of sample mismatch and misidentification.
Using an RFID radio frequency antenna array and an improved Florence-2 multimodal vision model, the system generates candidate identities, cargo locations, and layer sets through an RFID scanning module. Visual data is acquired by an image acquisition module, and multimodal fusion verification is performed. The fusion verification module is used for association matching and consistency determination, and the separation and retrieval module is used for verification to ensure the accuracy and reliability of sample identification.
It achieves accuracy and stability in sample identification and spatial positioning in high-density sample storage environments, improves the decision reliability and consistency of the intelligent storage and retrieval process of experimental samples, and effectively avoids sample mismatch and misidentification problems.
Smart Images

Figure CN122491670A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent storage and retrieval technology for experimental samples, and in particular to an intelligent storage and retrieval management system for experimental samples based on machine vision. Background Technology
[0002] With the continuous improvement of automation levels in biomedical laboratories, the demand for intelligent management of experimental samples during storage, retrieval, and transfer is increasing. In scenarios where high-density sample storage equipment and automated mechanical actuators work in tandem, how to accurately perceive and reliably manage the identity and spatial location information of experimental samples has become a key issue in intelligent laboratory systems. Existing technical solutions mainly rely on RFID radio frequency identification technology or vision-based image recognition technology for the positioning and identification of experimental samples.
[0003] In RFID-based solutions, samples are identified using radio frequency tags, and the signal strength or phase information of the tags is obtained using an RFID antenna array to infer the sample's location. However, in high-density storage environments, the RFID positioning results are susceptible to multipath effects, signal obstruction, and interference between tags, leading to significant uncertainty and making it difficult to accurately distinguish samples in adjacent storage locations or layers.
[0004] In vision-based technical solutions, images are acquired using cameras, and deep learning models are used for target detection and recognition to obtain semantic information about the appearance and spatial location of samples. However, in practical applications, experimental samples often exhibit high similarity in appearance, occluded or blurred labels, and complex lighting conditions, leading to ambiguity in visual recognition results. Furthermore, purely visual methods struggle to directly obtain the unique identifier of a sample, and misidentification or mismatching can easily occur when multiple samples are densely arranged.
[0005] To overcome the limitations of single sensing methods, some existing technologies attempt to fuse RFID with visual information, but the following shortcomings still exist: existing fusion methods mostly adopt simple rule matching or fusion after independent decision-making, lacking structured association modeling for identity information, cargo location information, and layer information, making it difficult to achieve consistency verification of multi-dimensional information; when there are identification conflicts or information inconsistencies, there is a lack of effective backtracking and re-retrieval mechanisms, making it impossible to reliably handle complex situations such as object code separation and sample offset; after completing sample grabbing or storage operations, existing technologies usually lack a verification mechanism based on multi-modal information, making it difficult to detect mis-grabbing or placement errors in the operation process in a timely manner.
[0006] Existing technologies for intelligent storage and management of experimental samples generally suffer from the following technical defects: they cannot simultaneously achieve highly reliable fusion perception of sample identity information and precise spatial location information in complex storage environments; they lack an effective mechanism for unified association matching and consistency determination of multi-source heterogeneous information; they lack dynamic retrieval and correction capabilities in cases of inconsistency or anomalies; and they lack effective means of result verification after mechanical operation, making it difficult to guarantee the accuracy and reliability of the overall storage and retrieval process. Summary of the Invention
[0007] One objective of this invention is to propose an intelligent storage and management system for experimental samples based on machine vision. This invention effectively avoids the problems of sample mismatch and misidentification, and improves the reliability and consistency of decision-making in the intelligent storage and retrieval process of experimental samples.
[0008] An intelligent storage and management system for experimental samples based on machine vision according to an embodiment of the present invention includes: The radio frequency scanning module uses an RFID radio frequency antenna array to perform a global radio frequency scan on the target storage device, generating an RFID candidate identity set, an RFID candidate cargo location set, and an RFID candidate layer set. The image acquisition module receives access instructions, determines the candidate spatial area based on the RFID candidate cargo location set and the RFID candidate layer set, and acquires the corresponding image data. An improved Florence-2 module was developed by inputting image data into the improved Florence-2 multimodal vision model to obtain a set of visual semantic truth values. The fusion verification module associates and matches the RFID candidate identity set, the RFID candidate cargo location set, and the RFID candidate layer set with the visual semantic truth value set to generate a fusion verification result set. The determination module performs consistency determination based on the fusion verification result set, outputs the unique identity information and precise spatial location information of the target sample, and generates a consistency confirmation result. When the consistency condition is not met, the separation and retrieval module generates retrieval and positioning information with the help of the improved Florence-2 multimodal vision model, updates the RFID candidate identity set, RFID candidate cargo location set and RFID candidate layer set and then returns to the fusion verification module. When the retrieval and positioning fails or the current retrieval round count reaches the preset maximum number of retrievals, the item code separation and retrieval process ends. After obtaining the consistency confirmation result, the execution control module converts the precise spatial position information of the target sample into motion control parameters of the mechanical actuator, and controls the mechanical actuator to complete the target sample grabbing, transfer or storage operation; After the mechanical actuator completes its operation, the grasping module collects the grasped image data and inputs it into the improved Florence-2 multimodal vision model to generate a set of visual semantic truth values after grasping. It then performs a consistency check. If the consistency check fails, it triggers an exception handling process and returns to the step separation and retrieval module. If the consistency check passes, it completes the experimental sample storage and retrieval management.
[0009] Optionally, the radio frequency scanning module includes: The target area of the target storage device is read multiple times by the RFID radio frequency antenna array to obtain the radio frequency identity information, radio frequency signal strength, radio frequency phase and antenna spatial identification of each RFID tag, forming radio frequency spatiotemporal sensing data; Based on radio frequency spatiotemporal sensing data, the identity aggregation result corresponding to each RFID tag is determined, and an RFID candidate identity set is generated. Feature comparison is performed between the RFID candidate identity set and the device layer radio frequency reference model to determine the layer matching result and generate the RFID candidate layer set. Feature comparison is performed between the RFID candidate identity set and the device layer radio frequency reference model to determine the layer matching result and generate the RFID candidate layer set. Based on the identity aggregation results, hierarchical matching results, and cargo location matching results, corresponding radio frequency credibility is added to the RFID candidate identity set, RFID candidate hierarchical set, and RFID candidate cargo location set.
[0010] Optionally, the image acquisition module includes: Receive the experimental sample access instruction and parse out the target radio frequency identification information corresponding to the target experimental sample. Based on the target radio frequency identification information, extract the corresponding target candidate storage location subset and target candidate layer location subset from the RFID candidate storage location set and the RFID candidate layer location set, respectively. The height range of the target candidate layer location subset is used as the height constraint boundary of the candidate spatial region, and the horizontal and vertical range of the cargo location corresponding to the target candidate cargo location subset are used as the planar constraint boundary of the candidate spatial region. The candidate spatial region is formed under the joint constraint of the height constraint boundary and the planar constraint boundary. The coordinates of the center of the candidate spatial region, the horizontal span of the region, and the vertical span of the region are calculated based on the spatial boundary coordinates of the candidate spatial region to determine the number of images to be acquired for the candidate spatial region. The control vision acquisition device uses the regional center coordinates of the candidate spatial region as the acquisition center. During the acquisition process, it synchronously records the global physical coordinates of the acquisition center corresponding to each block image, and generates image data of the candidate spatial region containing the physical coordinate mapping relationship.
[0011] Optionally, the improved Florence-2 module includes: Image data of candidate spatial regions are input into the improved Florence-2 multimodal vision model in the form of block image data. Based on the global physical coordinates of the acquisition center, a layer constraint prompt sequence, a cargo location topology prompt sequence, and a location anchoring prompt sequence are constructed respectively. These are then spliced together in order to obtain a joint prompt sequence. Each block of image data and its corresponding joint cue sequence are input into the improved Florence-2 multimodal vision model, and slot-guided visual semantic joint decoding is performed to obtain the slot visual semantic response set corresponding to each block of image data. Based on the sample foreground mask, slot visual region and semantic integrity corresponding to the visual semantic response of each slot, the occupancy confidence corresponding to the visual semantic response of each slot is calculated, and the visual slot occupancy relationship mapping corresponding to each block image data is generated according to the occupancy confidence. For the visual semantic response of each slot occupied by the experimental sample, the centroid pixel coordinates of the sample foreground are calculated based on the sample foreground mask, and combined with the physical coordinate mapping relationship of the corresponding block image data, the spatial position information of the sample is obtained, and a visual semantic truth unit is constructed. Calculate the cross-block consistency association score for visual semantic truth value units, and merge visual semantic truth value units whose cross-block consistency association score is not lower than the preset merging threshold to obtain the visual semantic truth value set.
[0012] Optionally, the fusion verification module includes: Based on the sample spatial location information corresponding to each visual semantic truth unit in the visual semantic truth set, and perform layer association matching with each radio frequency layer information in the RFID candidate layer set, the layer association result is obtained. Based on the spatial location information of the sample corresponding to each visual semantic truth unit and the radio frequency location information in the RFID candidate location set, location association matching is performed to obtain the location association result. Identity association matching is performed based on the radio frequency identity information in the RFID candidate identity set and the sample text label description corresponding to each visual semantic truth unit. The identity association result is obtained by combining the sample appearance semantic description, sample text label description and sample color attribute formation corresponding to each visual semantic truth unit. Based on the identity association results of each visual semantic truth unit relative to each radio frequency identity information, the hierarchical association results of each visual semantic truth unit relative to each radio frequency layer information, and the cargo location association results of each visual semantic truth unit relative to each radio frequency cargo location information, joint association matching is performed to generate a fusion verification result set.
[0013] Optionally, the determination module includes: Based on the correspondence between the radio frequency identity information and visual semantic features corresponding to each fusion verification result, the identity consistency result corresponding to each fusion verification result is calculated, and the identity consistency judgment result corresponding to each fusion verification result is determined based on the identity consistency result. Based on the correspondence between the radio frequency layer information and the sample spatial location information corresponding to each fusion verification result, the layer consistency result corresponding to each fusion verification result is calculated, and the layer consistency judgment result corresponding to each fusion verification result is determined based on the layer consistency result. Based on the correspondence between the radio frequency location information and the sample spatial location information corresponding to each fusion verification result, the location consistency result corresponding to each fusion verification result is determined. The location consistency result is compared with the preset location consistency threshold to determine the location consistency judgment result corresponding to each fusion verification result. The following fusion verification results are retained: identity consistency judgment result is passed, layer consistency judgment result is passed, cargo location consistency judgment result is passed, the corresponding radio frequency identity information matches the target radio frequency identity information in the access instruction, and the consistency confirmation score is not lower than the preset consistency confirmation threshold, forming a consistency candidate result set; When the set of candidate consistency results contains only a unique fusion verification result, the radio frequency identification information and sample spatial location information corresponding to the fusion verification result are extracted and used as the unique identification information and precise spatial location information of the target sample, respectively, to generate a consistency confirmation result.
[0014] Optionally, the separation and retrieval module includes: When the consistency condition is not met, the item code separation and retrieval mechanism is triggered, the current retrieval round count is incremented by one, and a retrieval area and retrieval prompt sequence for open vocabulary semantic retrieval are constructed based on the target radio frequency identification information in the access instruction, the current pallet area, and the RFID candidate layer set. The retrieved image data corresponding to the retrieved area and the retrieved prompt sequence are input together into the Florence-2 multimodal vision model, which has been improved by the constraints of the intelligent storage and management scenario of experimental samples. Open vocabulary semantic retrieval for the object code separation scenario is performed to obtain a set of retrieved candidate results. Based on the semantic description of the appearance of the candidate sample, the text label description of the candidate sample, the color attribute of the candidate sample, the candidate foreground mask and the semantic retrieval score corresponding to each candidate result, the semantic retrieval matching result corresponding to each candidate result is calculated. Based on the physical coordinate mapping relationship between the candidate foreground mask and the corresponding image block corresponding to each candidate retrieval result, the candidate spatial location information corresponding to each candidate retrieval result is determined. Combined with the semantic retrieval matching result, the layer retrieval matching result, the cargo location retrieval matching result and the current pallet priority result, the retrieval positioning information corresponding to each candidate retrieval result is determined. The RFID candidate identity set, RFID candidate cargo location set, and RFID candidate layer set are updated based on the retrieval and positioning information. The system then determines whether to return to the fusion verification module based on the current retrieval round count, the preset maximum number of retrievals, and the retrieval and positioning score.
[0015] Optionally, the execution control module includes: Obtain the consistency confirmation results and experimental sample access instructions; determine the source operation position based on the precise spatial position information of the target sample in the consistency confirmation results; and determine the target operation position based on the operation type information in the experimental sample access instructions. Based on the source operating position, target operating position, preset approach height, preset departure height, preset safe obstacle avoidance height, and grasping compensation distance, calculate the corresponding grasping approach posture, grasping posture, safe transition posture, release approach posture, and release posture of the mechanical actuator; Based on the sample clamping specification parameters and spatial distances between each pose corresponding to the unique identity information of the target sample, calculate the clamping control parameters and kinematic timing control parameters of the mechanical actuator. The grasping approach pose, grasping pose, safe transition pose, release approach pose, and release pose are converted into joint control parameters corresponding to each joint of the mechanical actuator, and the motion control parameters of the mechanical actuator are generated by combining the clamping control parameters and kinematic timing control parameters. The mechanical actuator is controlled according to the motion control parameters of the mechanical actuator, so that the end effector sequentially completes the target sample grasping and approaching, clamping and grasping, safe pulling out to the transition position, target position transfer and release placement, thus completing the target sample grasping, transfer or storage operation.
[0016] Optionally, the capture module includes: The control end vision acquisition device acquires image data after the mechanical actuator completes the grasping, transfer or storage operation of the target sample, and determines the priority observation range of the target sample in the image data after grasping based on the clamping area of the end actuator corresponding to the current state of the mechanical actuator, thus forming the observation area after grasping; Based on the observed region after capture, a capture observation prompt sequence is constructed. The captured image data and the capture observation prompt sequence are input together into the Florence-2 multimodal vision model to generate a capture visual semantic ground value set. Consistency verification is performed based on the set of visual semantic truth values after capture and the unique identity information of the target sample. The consistency verification result corresponding to each set of visual semantic truth values after capture is calculated. Perform a consistency review decision based on the consistency review results of each crawled data, and generate a consistency review result; When the consistency verification result indicates that the consistency verification has failed, the control mechanical actuator performs an abnormal release operation to clear the end effector. The current captured image data, the captured visual semantic truth set, and the unique identity information of the target sample are used as input information for the abnormal handling process. After updating the actual identity status of the corresponding source operation position or setting a negative sample masking label based on the captured visual semantic truth set, the process returns to the separation and retrieval module. When the consistency verification result indicates that the consistency verification has passed, the target sample corresponding to the current experimental sample access instruction is identified as a valid sample that has been captured and verified, thus completing the intelligent access management of experimental samples.
[0017] The beneficial effects of this invention are: This invention constructs a multimodal heterogeneous data fusion verification algorithm for intelligent storage and management of experimental samples. It performs structured association modeling of RFID candidate identity set, RFID candidate cargo location set, RFID candidate layer set, and visual semantic truth set. Through multi-dimensional joint constraints of layer association matching, cargo location association matching, and identity association matching, a unified joint association score-driven fusion verification mechanism is formed. By introducing the synergistic optimization of spatial location constraints and semantic consistency constraints, radio frequency information and visual semantic information are coupled and matched in the same coordinate system. This effectively suppresses the matching ambiguity caused by radio frequency multipath interference and visual misidentification, significantly improves the accuracy and stability of experimental sample identity recognition and spatial positioning, and maintains reliable recognition performance even in high-density sample storage environments.
[0018] In the consistency determination process, this invention constructs a consistency confirmation mechanism with multi-condition collaborative constraints by weighting and fusing identity consistency results, hierarchical consistency results, and cargo location consistency results with a unified dimension. It also introduces a uniqueness screening strategy to achieve global consistency determination and conflict resolution of candidate fusion results. Through multi-source information joint constraints, a strong coupling relationship is formed between identity information and spatial location information. Even in the case of multiple candidate matching or local optimal conflict, a unique solution can still be determined, effectively avoiding sample mismatch and misplacement identification problems, and improving the decision reliability and consistency in the intelligent storage and retrieval process of experimental samples.
[0019] This invention takes the improved Florence-2 machine vision-based multimodal vision model of the intelligent storage and management system for experimental samples as its main body, which is constrained by the scenario of intelligent storage and management of experimental samples. At the processing logic and execution level, the model performs scene-constrained visual encoding and cue encoding on the input images and text, and completes feature interaction through the model's attention mechanism layer. Finally, it performs slot-guided visual semantic joint decoding. In terms of data transmission dimension and format, the bottom-level block image data and the joint cue sequence composed of sequentially spliced data (including the layer constraint cue sequence composed of layer field combination, the slot topology cue sequence formed by the arrangement of the slot space, and the position anchoring cue sequence transformed by global physical coordinates) are jointly input into the improved Florence-2 machine vision-based multimodal vision model of the intelligent storage and management system for experimental samples. The joint cue sequence, as the slot prior information in the physical space, directly intervenes in the model's attention mechanism layer for fusion. Within one forward inference cycle, the generated block joint semantic features are decoded and output as a multi-dimensional slot visual semantic response set. The mask is converted into spatial position information through physical coordinate mapping relationship and a visual semantic truth unit is constructed and output. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 The flowchart shows a machine vision-based intelligent storage and management system for experimental samples proposed in this invention. Figure 2 This is a structural block diagram of the improved Florence-2 multimodal vision model in an intelligent storage and management system for experimental samples based on machine vision proposed in this invention. Detailed Implementation
[0021] Example 1: Reference Figure 1 A machine vision-based intelligent storage and management system for experimental samples includes: The radio frequency scanning module uses an RFID radio frequency antenna array to perform a global radio frequency scan on the target storage device, generating an RFID candidate identity set, an RFID candidate cargo location set, and an RFID candidate layer set. In this embodiment, the radio frequency scanning module includes: The target area of the target storage device is read multiple times by the RFID radio frequency antenna array to obtain the radio frequency identity information, radio frequency signal strength, radio frequency phase and antenna spatial identification of each RFID tag, forming radio frequency spatiotemporal sensing data; Based on radio frequency spatiotemporal sensing data, the identity aggregation result corresponding to each RFID tag is determined, and an RFID candidate identity set is generated. In Example 1, the radio frequency signal strength coefficient and phase consistency coefficient of the RFID tag in the multi-round reading results are weighted and aggregated to obtain the identity aggregation result; the radio frequency identity information whose identity aggregation result is not lower than the preset identity screening threshold is determined as the RFID candidate identity set.
[0022] Feature comparison is performed between the RFID candidate identity set and the device layer radio frequency reference model to determine the layer matching result and generate the RFID candidate layer set. Feature comparison is performed between the RFID candidate identity set and the device layer radio frequency reference model to determine the layer matching result and generate the RFID candidate layer set. Based on the identity aggregation results, hierarchical matching results, and cargo location matching results, corresponding radio frequency credibility is added to the RFID candidate identity set, RFID candidate hierarchical set, and RFID candidate cargo location set.
[0023] In Example 1, based on the identity aggregation results corresponding to each radio frequency identity information in the RFID candidate identity set, the layer matching results corresponding to each radio frequency layer information in the RFID candidate layer set, and the cargo location matching results corresponding to each radio frequency cargo location information in the RFID candidate cargo location set, a corresponding radio frequency credibility is added to each candidate element in the RFID candidate identity set, the RFID candidate layer set, and the RFID candidate cargo location set.
[0024] The image acquisition module receives access instructions, determines the candidate spatial area based on the RFID candidate cargo location set and the RFID candidate layer set, and acquires the corresponding image data. In this embodiment, the image acquisition module includes: Receive the experimental sample access instruction and parse out the target radio frequency identification information corresponding to the target experimental sample. Based on the target radio frequency identification information, extract the corresponding target candidate storage location subset and target candidate layer location subset from the RFID candidate storage location set and the RFID candidate layer location set, respectively. The height range of the target candidate layer location subset is used as the height constraint boundary of the candidate spatial region, and the horizontal and vertical range of the cargo location corresponding to the target candidate cargo location subset are used as the planar constraint boundary of the candidate spatial region. The candidate spatial region is formed under the joint constraint of the height constraint boundary and the planar constraint boundary. In Example 1, the target candidate cargo location subset is mapped to the horizontal lower boundary coordinates, horizontal upper boundary coordinates, vertical lower boundary coordinates, and vertical upper boundary coordinates of each cargo location. The target candidate layer location subset is mapped to the height lower boundary coordinates and height upper boundary coordinates of each layer location. The minimum lower boundary coordinates and maximum upper boundary coordinates of the target candidate cargo location subset in the horizontal and vertical directions, as well as the minimum lower boundary coordinates and maximum upper boundary coordinates of the target candidate layer location subset in the height direction, are extracted respectively. Then, the minimum lower boundary coordinates in the horizontal, vertical, and height directions are expanded outward, and the maximum upper boundary coordinates in the horizontal, vertical, and height directions are expanded outward. The expanded horizontal boundary, vertical boundary, and height boundary together enclose the candidate space region.
[0025] The coordinates of the center of the candidate spatial region, the horizontal span of the region, and the vertical span of the region are calculated based on the spatial boundary coordinates of the candidate spatial region to determine the number of images to be acquired for the candidate spatial region. In Example 1, the center coordinates, horizontal span, and vertical span of the candidate spatial region are calculated based on the spatial boundary coordinates of the candidate spatial region. The acquisition height of the visual acquisition device is adjusted to the height plane corresponding to the center coordinates of the region. Under the premise of introducing a preset image overlap redundancy span, the horizontal span of the region is compared with the horizontal coverage range of a single frame image of the visual acquisition device under the height plane to determine the number of images acquired in the horizontal direction of the candidate spatial region. The vertical span of the region is compared with the vertical coverage range of a single frame image of the visual acquisition device under the height plane to determine the number of images acquired in the vertical direction of the candidate spatial region. Based on the number of images acquired in the horizontal direction and the number of images acquired in the vertical direction of the candidate spatial region, the number of images acquired corresponding to the candidate spatial region is determined.
[0026] The control vision acquisition device uses the regional center coordinates of the candidate spatial region as the acquisition center. During the acquisition process, it synchronously records the global physical coordinates of the acquisition center corresponding to each block image, and generates image data of the candidate spatial region containing the physical coordinate mapping relationship.
[0027] An improved Florence-2 module was developed by inputting image data into the improved Florence-2 multimodal vision model to obtain a set of visual semantic truth values. refer to Figure 2 In this embodiment, the Florence-2 module is improved, including: Image data of candidate spatial regions are input into the improved Florence-2 multimodal vision model in the form of block image data. Based on the global physical coordinates of the acquisition center, a layer constraint prompt sequence, a cargo location topology prompt sequence, and a location anchoring prompt sequence are constructed respectively. These are then spliced together in order to obtain a joint prompt sequence. In Example 1, each candidate layer in the target candidate layer subset is sorted according to the layer matching result corresponding to each radio frequency layer information. The layer number, layer height range, layer adjacency relationship and layer priority order are extracted and combined according to the preset field order to form a layer constraint prompt sequence.
[0028] Based on the horizontal and vertical ranges of the target candidate storage locations subset, the candidate distribution area is determined. The storage location number, row and column position, adjacency relationship, arrangement direction, and carrier structure of each candidate storage location within the area are extracted and organized according to the spatial arrangement order of the storage locations to form a storage location topology prompt sequence.
[0029] The global physical coordinates of the acquisition center corresponding to each block of image data are directly converted into horizontal position markers, vertical position markers, and height position markers, and then combined according to the preset position field order to form a position anchoring prompt sequence.
[0030] The layer constraint prompt sequence, the cargo location topology prompt sequence, and the location anchoring prompt sequence are then concatenated in a fixed order according to layer constraint, cargo location topology, and location anchoring to form a joint prompt sequence.
[0031] Each block of image data and its corresponding joint cue sequence are input into the improved Florence-2 multimodal vision model, and slot-guided visual semantic joint decoding is performed to obtain the slot visual semantic response set corresponding to each block of image data. In Example 1, slot-guided visual-semantic joint decoding refers to the process of transforming prior information about slots in the physical space into a sequence of feature cues, which is then used to improve the attention mechanism layer of the Florence-2 multimodal visual model. This forces the model to no longer blindly search globally during the decoding process, but instead use physical slots as anchor points to synchronously decouple and output the bounding box coordinates, text label semantics, and multidimensional features of appearance color within a forward inference cycle.
[0032] First, the image data of each block and the corresponding joint cue sequence are input into the improved Florence-2 multimodal vision model. The improved Florence-2 multimodal vision model performs visual encoding and cue encoding on each block of image data under the joint constraints of the layer constraint cue sequence, the cargo location topology cue sequence and the location anchoring cue sequence, so as to obtain the block joint semantic features corresponding to each block of image data.
[0033] Based on the block joint semantic features, according to the cargo location arrangement relationship corresponding to the cargo location topology prompt sequence, the slot candidate region in each block image data is decoded slot by slot to obtain the slot visual region corresponding to each block image data.
[0034] For each slot's visual region, the improved Florence-2 multimodal vision model is invoked to perform foreground extraction and semantic joint decoding within the slot, obtaining the sample foreground mask, sample appearance semantic description, sample text label description, and sample color attributes corresponding to each slot's visual region, as well as the model decoding confidence of each semantic attribute.
[0035] Based on the integrity of the sample foreground mask, the formation of the sample appearance semantic description, the formation of the sample text label description, and the formation of the sample color attribute within the visual region of each slot, post-processing weighted calculation is performed to determine the semantic integrity corresponding to the visual region of each slot. The visual region of each slot, the sample appearance semantic description, the sample text label description, the sample color attribute, the sample foreground mask, and the semantic integrity are combined to obtain the slot visual semantic response set corresponding to each block of image data.
[0036] Based on the sample foreground mask, slot visual region and semantic integrity corresponding to the visual semantic response of each slot, the occupancy confidence corresponding to the visual semantic response of each slot is calculated, and the visual slot occupancy relationship mapping corresponding to each block image data is generated according to the occupancy confidence. In Example 1, the mask occupancy factor is obtained by dividing the overlap area between the sample foreground mask and the slot visual area by the area of the slot visual area.
[0037] The semantic completeness coefficient is obtained by weighted normalization of whether a semantic description of the sample appearance is formed, whether a textual label description of the sample is formed, and whether a color attribute of the sample is formed.
[0038] The placeholder credibility is obtained by weighted fusion of mask placeholder coefficient, semantic completeness coefficient and semantic integrity.
[0039] The occupancy confidence level is compared with the occupancy screening threshold. When the occupancy confidence level is not lower than the occupancy screening threshold, the occupancy status of the corresponding slot visual area is determined to be occupied by the experimental sample. When the occupancy confidence level is lower than the occupancy screening threshold, the occupancy status of the corresponding slot visual area is determined to be unoccupied by the experimental sample.
[0040] The visual slot occupancy relationship mapping corresponding to each block of image data is formed by the visual area of each slot and its corresponding slot occupancy status.
[0041] For the visual semantic response of each slot occupied by the experimental sample, the centroid pixel coordinates of the sample foreground are calculated based on the sample foreground mask, and combined with the physical coordinate mapping relationship of the corresponding block image data, the spatial position information of the sample is obtained, and a visual semantic truth unit is constructed. In Example 1, the horizontal pixel coordinates of all foreground pixels within the sample foreground mask are summed and then divided by the sample foreground mask area, and the vertical pixel coordinates of all foreground pixels within the sample foreground mask are summed and then divided by the sample foreground mask area to obtain the centroid pixel coordinates of the sample foreground. By inputting the foreground centroid pixel coordinates of the sample into the physical coordinate mapping relationship of the corresponding block image data, the spatial location information of the sample is obtained. The visual semantic truth unit is composed of the semantic description of the sample appearance, the text label description of the sample, the color attribute of the sample, and the spatial location information of the sample.
[0042] Calculate the cross-block consistency association score for visual semantic truth value units, and merge visual semantic truth value units whose cross-block consistency association score is not lower than the preset merging threshold to obtain the visual semantic truth value set.
[0043] In Example 1, for visual semantic truth units with spatial overlap between different image data blocks, a weighted fusion is performed on the normalized semantic similarity between sample appearance semantic descriptions, the normalized text similarity between sample text label descriptions, the normalized attribute similarity between sample color attributes, and the spatial proximity between sample spatial location information to calculate a cross-block consistency association score. Visual semantic truth units with cross-block consistency association scores not lower than a preset merging threshold are merged. During merging, a weighted average processing based on occupancy confidence is performed on the sample spatial location information, and the results with the highest model decoding confidence are retained for sample appearance semantic descriptions, sample text label descriptions, and sample color attributes to obtain the visual semantic truth set corresponding to the candidate spatial region. At the same time, the visual slot occupancy relationship mappings corresponding to each image data block are summarized to obtain the visual slot occupancy relationship mappings corresponding to the candidate spatial region.
[0044] The spatial proximity between sample spatial location information is obtained by exponentially decaying the spatial distance between two visual semantic truth units.
[0045] Compared to the existing conventional machine vision-based intelligent storage and management system for experimental samples, the improved Florence-2 multimodal vision model for the Florence-2 system abandons the original model's mechanism of relying entirely on generalized natural language prompts for global blind search. Instead, it transforms the layer and storage location topology rules and physical location coordinates in the physical world into a sequence of feature prompts, forcibly intervening and constraining the attention mechanism layer. This forms a novel slot-guided visual-semantic joint decoding process, forcing the large model to stop blindly searching during inference and instead use the physical slot as an absolute anchor point. Within a forward inference cycle, it synchronously decouples and outputs the bounding boxes, text, and appearance features within the slot area. This not only eliminates background interference and overlapping false detections in high-density test tube array scenarios from the underlying principle but also opens up the mapping channel from two-dimensional pixel space to three-dimensional physical coordinates, improving the model's feature decoding accuracy and physical space perception capability in industrial-grade storage and retrieval scenarios.
[0046] The fusion verification module associates and matches the RFID candidate identity set, the RFID candidate cargo location set, and the RFID candidate layer set with the visual semantic truth value set to generate a fusion verification result set. In this embodiment, the fusion verification module includes: Based on the sample spatial location information corresponding to each visual semantic truth unit in the visual semantic truth set, and perform layer association matching with each radio frequency layer information in the RFID candidate layer set, the layer association result is obtained. In Example 1, based on the layer height range corresponding to each radio frequency layer information, and combined with the coordinate values of the sample spatial position information in the height direction corresponding to each visual semantic truth unit, the layer correlation result of each visual semantic truth unit relative to each radio frequency layer information is calculated.
[0047] When the coordinates of the sample's spatial location information in the height direction are within the range of the corresponding radio frequency (RF) layer height information, the layer correlation result is determined based on the degree of deviation between the coordinates of the sample's spatial location information in the height direction and the center height coordinates of the corresponding RF layer height information. The smaller the deviation, the higher the layer correlation result. When the coordinates of the sample's spatial location information in the height direction are not within the range of the corresponding RF layer height information, the layer correlation result is determined to be zero.
[0048] Based on the spatial location information of the sample corresponding to each visual semantic truth unit and the radio frequency location information in the RFID candidate location set, location association matching is performed to obtain the location association result. In Example 1, the center position and range of each radio frequency (RF) location are first determined based on the horizontal and vertical ranges of the location. Then, it is determined whether the spatial position information of the sample corresponding to each visual semantic truth unit falls within the location range of the corresponding RF location information in both the horizontal and vertical directions. When the sample spatial position information falls within the location range of the corresponding RF location information, it is normalized based on the deviation of the sample spatial position information from the center position of the location in both the horizontal and vertical directions. The horizontal normalization result and the vertical normalization result are then multiplied and fused to obtain the location association result. When the sample spatial position information does not fall within the location range of the corresponding RF location information, the location association result is set to zero.
[0049] Identity association matching is performed based on the radio frequency identity information in the RFID candidate identity set and the sample text label description corresponding to each visual semantic truth unit. The identity association result is obtained by combining the sample appearance semantic description, sample text label description and sample color attribute formation corresponding to each visual semantic truth unit. In Example 1, each radio frequency identification information is first converted into an identity character sequence for comparison with the sample text label description. The identity text matching coefficient is determined based on the degree of character matching between the identity character sequence corresponding to each radio frequency identification information and the sample text label description corresponding to each visual semantic truth unit. Then, the visual semantic completeness coefficient is determined based on the formation of the sample appearance semantic description, sample text label description, and sample color attribute corresponding to each visual semantic truth unit. The identity text matching coefficient and the visual semantic completeness coefficient are weighted and fused to obtain the identity association result of each visual semantic truth unit relative to each radio frequency identification information.
[0050] Based on the identity association results of each visual semantic truth unit relative to each radio frequency identity information, the hierarchical association results of each visual semantic truth unit relative to each radio frequency layer information, and the cargo location association results of each visual semantic truth unit relative to each radio frequency cargo location information, joint association matching is performed to generate a fusion verification result set.
[0051] In Example 1, a weighted fusion is performed based on the identity association results, hierarchical association results, and cargo location association results corresponding to each visual semantic truth unit to obtain a joint association score. For each visual semantic truth unit, the radio frequency identity information, radio frequency hierarchical information, and radio frequency cargo location information with the highest joint association score are selected and combined with the sample appearance semantic description, sample text label description, sample color attribute, and sample spatial location information corresponding to the visual semantic truth unit to obtain a fusion verification result. Each fusion verification result with a joint association score not lower than the result screening threshold is retained. If multiple visual semantic truth units in the retained fusion verification results correspond to the same radio frequency identity information or the same radio frequency cargo location information, causing association conflicts, a bipartite graph matching algorithm is introduced for global optimization allocation to ensure the uniqueness of the mapping between each radio frequency identity information, radio frequency cargo location information, and visual semantic truth unit, thereby forming a fusion verification result set. The fusion verification result set records the correspondence between radio frequency identity information, radio frequency cargo location information, radio frequency hierarchical information, and visual semantic features.
[0052] The determination module performs consistency determination based on the fusion verification result set, outputs the unique identity information and precise spatial location information of the target sample, and generates a consistency confirmation result. In this embodiment, the determination module includes: Based on the correspondence between the radio frequency identity information and visual semantic features corresponding to each fusion verification result, the identity consistency result corresponding to each fusion verification result is calculated, and the identity consistency judgment result corresponding to each fusion verification result is determined based on the identity consistency result. In Example 1, the identity association result in the heterogeneous data fusion verification stage is directly used as the identity consistency result, and the corresponding identity consistency judgment result is determined according to the preset identity consistency threshold. When the identity consistency result is not lower than the preset identity consistency threshold, the identity consistency judgment result is determined to be passed; when the identity consistency result is lower than the preset identity consistency threshold, the identity consistency judgment result is determined to be failed.
[0053] Based on the correspondence between the radio frequency layer information and the sample spatial location information corresponding to each fusion verification result, the layer consistency result corresponding to each fusion verification result is calculated, and the layer consistency judgment result corresponding to each fusion verification result is determined based on the layer consistency result. In Example 1, the hierarchical association result of the heterogeneous data fusion verification stage is directly used as the hierarchical consistency result, and the corresponding hierarchical consistency judgment result is determined according to the preset hierarchical consistency threshold. When the hierarchical consistency result is not lower than the preset hierarchical consistency threshold, the hierarchical consistency judgment result is determined to be passed; when the hierarchical consistency result is lower than the preset hierarchical consistency threshold, the hierarchical consistency judgment result is determined to be failed.
[0054] Based on the correspondence between the radio frequency location information and the sample spatial location information corresponding to each fusion verification result, the location consistency result corresponding to each fusion verification result is determined. The location consistency result is compared with the preset location consistency threshold to determine the location consistency judgment result corresponding to each fusion verification result. In Example 1, the location association result of the heterogeneous data fusion verification stage is directly used as the location consistency result, and the corresponding location consistency judgment result is determined according to the preset location consistency threshold. When the location consistency result is not lower than the preset location consistency threshold, the location consistency judgment result is determined to be passed; when the location consistency result is lower than the preset location consistency threshold, the location consistency judgment result is determined to be failed.
[0055] The following fusion verification results are retained: identity consistency judgment result is passed, layer consistency judgment result is passed, cargo location consistency judgment result is passed, the corresponding radio frequency identity information matches the target radio frequency identity information in the access instruction, and the consistency confirmation score is not lower than the preset consistency confirmation threshold, forming a consistency candidate result set; In Example 1, the consistency scores of identity, level, location, and joint association are weighted and fused to obtain the consistency confirmation score.
[0056] When the set of candidate consistency results contains only a unique fusion verification result, the radio frequency identification information and sample spatial location information corresponding to the fusion verification result are extracted and used as the unique identification information and precise spatial location information of the target sample, respectively, to generate a consistency confirmation result.
[0057] In Example 1, the number of fusion verification results in the consistency candidate result set is first counted. When the number of fusion verification results in the consistency candidate result set is equal to one, the radio frequency identification information of the corresponding fusion verification result is determined as the unique identification information of the target sample, the sample spatial location information of the corresponding fusion verification result is determined as the precise spatial location information of the target sample, and a consistency confirmation result is generated by combining the corresponding consistency confirmation score. When the number of fusion verification results in the consistency candidate result set is not equal to one, the consistency confirmation result is determined as an empty result.
[0058] When the consistency condition is not met, the separation and retrieval module generates retrieval and positioning information with the help of the improved Florence-2 multimodal vision model, updates the RFID candidate identity set, RFID candidate cargo location set and RFID candidate layer set and then returns to the fusion verification module. When the retrieval and positioning fails or the current retrieval round count reaches the preset maximum number of retrievals, the item code separation and retrieval process ends. In this embodiment, the separation retrieval module includes: When the consistency condition is not met, the item code separation and retrieval mechanism is triggered, the current retrieval round count is incremented by one, and a retrieval area and retrieval prompt sequence for open vocabulary semantic retrieval are constructed based on the target radio frequency identification information in the access instruction, the current pallet area, and the RFID candidate layer set. In Example 1, it is determined whether there is a fusion verification result in the heterogeneous data fusion verification stage. If there is, the fusion verification result with the highest joint association score is selected as the reference fusion verification result. Based on the visual semantic features and location information corresponding to the reference fusion verification result, the semantic anchoring information and slot drift constraint information are determined. Then, according to the layer order relationship corresponding to each radio frequency layer information in the RFID candidate layer set, the adjacent layer information adjacent to the radio frequency layer information corresponding to the reference fusion verification result is extracted. If there is no such information, the target appearance morphology prior is extracted only based on the target radio frequency identity information in the access instruction by calling the preset experimental sample information database. This prior is used as the semantic anchoring information for retrieval. The slot drift constraint information is set to the global search state. Then, according to the layer order relationship corresponding to each radio frequency layer information in the RFID candidate layer set, the adjacent layer information adjacent to the radio frequency layer information corresponding to the target radio frequency identity information is extracted. The current tray area is combined with the adjacent layer areas corresponding to each adjacent layer information to form the retrieval area.
[0059] Based on the target radio frequency identity information, retrieval semantic anchoring information, and slot drift constraint information in the access instruction, identity anchoring prompt fragment, semantic anchoring prompt fragment, and slot drift prompt fragment are constructed respectively, and spliced in a fixed order to form a retrieval prompt sequence.
[0060] The retrieved image data corresponding to the retrieved area and the retrieved prompt sequence are input together into the Florence-2 multimodal vision model, which has been improved by the constraints of the intelligent storage and management scenario of experimental samples. Open vocabulary semantic retrieval for the object code separation scenario is performed to obtain a set of retrieved candidate results. In Example 1, the search area is segmented into images to obtain search image data. This data, along with the search prompt sequence, is input into the Florence-2 multimodal vision model, which has been improved by constraints of the intelligent storage and management scenario for experimental samples. Under the joint constraints of identity anchoring prompt fragments, semantic anchoring prompt fragments, and slot drift prompt fragments, the Florence-2 multimodal vision model performs open-vocabulary semantic retrieval for the object code separation scenario in the current tray area and adjacent layer areas, obtaining a set of search candidate results. For each search candidate result, the candidate bounding box, the candidate sample appearance semantic description, the candidate sample text label description, the candidate sample color attribute, the candidate foreground mask, and the semantic retrieval score are output simultaneously.
[0061] Based on the semantic description of the appearance of the candidate sample, the text label description of the candidate sample, the color attribute of the candidate sample, the candidate foreground mask and the semantic retrieval score corresponding to each candidate result, the semantic retrieval matching result corresponding to each candidate result is calculated. In Example 1, the identity text retrieval matching result is calculated based on the degree of matching between the target radio frequency identity information in the access instruction and the candidate sample text label description corresponding to each retrieval candidate result.
[0062] Based on the semantic similarity between the semantic description of the appearance of the candidate sample corresponding to each back-finding candidate result and the semantic description of the appearance of the sample corresponding to the back-finding semantic anchoring information, and the attribute similarity between the color attribute of the candidate sample corresponding to each back-finding candidate result and the color attribute of the sample corresponding to the back-finding semantic anchoring information, the appearance color back-finding matching result is determined.
[0063] Then, based on the completeness of the formation of the candidate sample appearance semantic description, candidate sample text label description, and candidate sample color attribute corresponding to each candidate result, the semantically complete result of the retrieval is determined.
[0064] The identity text search matching results, appearance color search matching results, and semantic completeness search results are weighted and fused to obtain the semantic search matching results.
[0065] Based on the physical coordinate mapping relationship between the candidate foreground mask and the corresponding image block corresponding to each candidate retrieval result, the candidate spatial location information corresponding to each candidate retrieval result is determined. Combined with the semantic retrieval matching result, the layer retrieval matching result, the cargo location retrieval matching result and the current pallet priority result, the retrieval positioning information corresponding to each candidate retrieval result is determined. In Example 1, the centroid pixel coordinates of the candidate foreground are extracted based on the candidate foreground mask corresponding to each candidate search result, and the candidate spatial location information is obtained by combining the physical coordinate mapping relationship of the corresponding image block; the layer search matching result is determined according to the correspondence between the candidate spatial location information and the layer height range corresponding to each radio frequency layer information in the RFID candidate layer set; the storage location search matching result is determined according to the correspondence between the candidate spatial location information and the storage location range in the current pallet area and adjacent layer areas; and the priority result of the current pallet is determined according to whether the candidate spatial location information is located in the current pallet area.
[0066] The semantic search matching results, layer search matching results, cargo location search matching results, and current pallet priority results are weighted and fused to obtain a search and location score. The search candidate with the highest search and location score is selected, and inverse spatial mapping is performed on the physical grid of the current pallet area and adjacent layer areas based on the corresponding candidate spatial location information to obtain updated radio frequency cargo location information and updated radio frequency layer information. The target radio frequency identity information, updated radio frequency cargo location information, updated radio frequency layer information, candidate spatial location information, and search and location score in the access instruction are combined together as the search and location information.
[0067] The RFID candidate identity set, RFID candidate cargo location set, and RFID candidate layer set are updated based on the retrieval and positioning information. The system then determines whether to return to the fusion verification module based on the current retrieval round count, the preset maximum number of retrievals, and the retrieval and positioning score.
[0068] In Example 1, based on the target radio frequency identity information, updated radio frequency cargo location information, and updated radio frequency layer information in the retrieved location information, updated RFID candidate identity set, updated RFID candidate cargo location set, and updated RFID candidate layer set are generated respectively.
[0069] When the retrieval and positioning score is not lower than the preset retrieval and update threshold, the RFID candidate identity set, RFID candidate cargo location set, and RFID candidate layer set before retrieval are replaced with the updated RFID candidate identity set, updated RFID candidate cargo location set, and updated RFID candidate layer set, and then returned to the fusion verification module.
[0070] When the retrieval and positioning score is lower than the preset retrieval and update threshold, or when the current retrieval round count reaches the preset maximum number of retrievals, the current code separation retrieval mechanism is determined to be a retrieval failure, and the code will no longer be returned to the fusion verification module. The current retrieval round count will be cleared, and empty retrieval and positioning information will be output.
[0071] The preset maximum number of lookups indicates the maximum number of lookup rounds allowed for the code separation lookup mechanism to be executed for the same target RFID identity information. The current lookup round count is automatically initialized to zero when a new access instruction is received. Empty lookup location information indicates that no valid lookup location result has been obtained that can be used to update the RFID candidate identity set, RFID candidate cargo location set, and RFID candidate layer set.
[0072] After obtaining the consistency confirmation result, the execution control module converts the precise spatial position information of the target sample into motion control parameters of the mechanical actuator, and controls the mechanical actuator to complete the target sample grabbing, transfer or storage operation; In this embodiment, the execution control module includes: Obtain the consistency confirmation results and experimental sample access instructions; determine the source operation position based on the precise spatial position information of the target sample in the consistency confirmation results; and determine the target operation position based on the operation type information in the experimental sample access instructions. In Example 1, the precise spatial location information of the target sample is determined as the source operation position for the mechanical actuator to perform the grasping operation; when the operation type information is a take-out operation, the position corresponding to the preset sample exit station is determined as the target operation position; when the operation type information is a transfer operation or a storage operation, the target storage position carried in the experimental sample storage and retrieval instruction is determined as the target operation position.
[0073] Based on the source operating position, target operating position, preset approach height, preset departure height, preset safe obstacle avoidance height, and grasping compensation distance, calculate the corresponding grasping approach posture, grasping posture, safe transition posture, release approach posture, and release posture of the mechanical actuator; In Example 1, based on the precise spatial position information of the target sample and the target operation position, and combined with the preset attitude angle of the mechanical actuator performing the grasping operation on the experimental sample in a top-down manner, the grasping approach pose, grasping pose, safe transition pose, release approach pose, and release pose are generated respectively.
[0074] The approach pose is obtained by superimposing a preset approach height on the height direction of the source operating position. The grasp pose is obtained by superimposing a grasp compensation distance on the height direction of the source operating position. The safe transition pose is obtained by superimposing a preset safe obstacle avoidance height on the height direction of the source operating position. The release approach pose is obtained by superimposing a preset departure height on the height direction of the target operating position. The release pose is obtained by directly adopting the target operating position and maintaining a preset attitude angle.
[0075] Based on the sample clamping specification parameters and spatial distances between each pose corresponding to the unique identity information of the target sample, calculate the clamping control parameters and kinematic timing control parameters of the mechanical actuator. In Example 1, the preset experimental sample clamping specifications are queried based on the unique identity information of the target sample to obtain the standard clamping opening width, standard clamping closing width, and standard clamping force corresponding to the target sample. Combined with the safety opening compensation amount, the clamping opening width, clamping closing width, and clamping force corresponding to the mechanical actuator are determined. Then, based on the grasping approach posture, the spatial distance between the grasping posture and the release approach posture, and combined with the preset approach speed and preset transfer speed, the approach motion duration and transfer motion duration are determined, thereby generating the clamping control parameters and kinematic timing control parameters corresponding to the mechanical actuator.
[0076] The grasping approach pose, grasping pose, safe transition pose, release approach pose, and release pose are converted into joint control parameters corresponding to each joint of the mechanical actuator, and the motion control parameters of the mechanical actuator are generated by combining the clamping control parameters and kinematic timing control parameters. In Example 1, inverse kinematics solutions are performed on the grasping approach pose, grasping pose, safe transition pose, release approach pose, and release pose to obtain the corresponding joint control parameters. Then, the joint control parameters corresponding to the grasping approach pose, the grasping pose, the safe transition pose, the release approach pose, the release pose, the clamping opening width, the clamping closing width, the clamping force, the approach motion duration, and the transfer motion duration are organized in the execution order to form the motion control parameters of the mechanical actuator.
[0077] The mechanical actuator is controlled according to the motion control parameters of the mechanical actuator, so that the end effector sequentially completes the target sample grasping and approaching, clamping and grasping, safe pulling out to the transition position, target position transfer and release placement, thus completing the target sample grasping, transfer or storage operation.
[0078] After the mechanical actuator completes its operation, the grasping module collects the image data after grasping and inputs it into the improved Florence-2 multimodal vision model to generate a set of visual semantic truth values after grasping. It then performs a consistency check. If the consistency check fails, it triggers an exception handling process and returns to the separation and retrieval module. If the consistency check passes, it completes the storage and management of experimental samples.
[0079] In this embodiment, the capture module includes: The control end vision acquisition device acquires image data after the mechanical actuator completes the grasping, transfer or storage operation of the target sample, and determines the priority observation range of the target sample in the image data after grasping based on the clamping area of the end actuator corresponding to the current state of the mechanical actuator, thus forming the observation area after grasping; Based on the observed region after capture, a capture observation prompt sequence is constructed. The captured image data and the capture observation prompt sequence are input together into the Florence-2 multimodal vision model to generate a capture visual semantic ground value set. In Example 1, based on the post-capture visual semantic truth set, each post-capture visual semantic truth unit synchronously outputs the post-capture sample appearance semantic description, post-capture sample text label description, post-capture sample color attribute, post-capture sample foreground mask, and post-capture sample spatial location information.
[0080] Consistency verification is performed based on the set of visual semantic truth values after capture and the unique identity information of the target sample. The consistency verification result corresponding to each set of visual semantic truth values after capture is calculated. In Example 1, the identity verification matching result is first determined based on the degree of matching between the unique identity information of the target sample and the text label description of the captured sample corresponding to each captured visual semantic truth unit; then, the semantic completeness result is determined based on the formation of the semantic description of the appearance of the captured sample, the text label description of the captured sample, and the color attribute of the captured sample corresponding to each captured visual semantic truth unit; finally, the clamping area constraint result is determined based on the degree of overlap between the foreground mask of the captured sample and the clamping area of the end effector corresponding to each captured visual semantic truth unit.
[0081] The identity verification matching results, the semantic completeness results after crawling, and the clamping region constraint results are weighted and fused to obtain the consistency verification results after crawling.
[0082] Perform a consistency review decision based on the consistency review results of each crawled data, and generate a consistency review result; In Example 1, the post-crawl visual semantic truth units whose post-crawl consistency verification results are not lower than the preset post-crawl verification threshold are first screened to form a post-crawl verification candidate set.
[0083] The consistency review is considered successful when the candidate set for post-fetching verification contains only one post-fetching visual semantic truth unit; the consistency review is considered unsuccessful when the candidate set for post-fetching verification does not contain a post-fetching visual semantic truth unit or contains multiple post-fetching visual semantic truth units.
[0084] When the consistency verification result indicates that the consistency verification has failed, the control mechanical actuator performs an abnormal release operation to clear the end effector. The current captured image data, the captured visual semantic truth set, and the unique identity information of the target sample are used as input information for the abnormal handling process. After updating the actual identity status of the corresponding source operation position or setting a negative sample masking label based on the captured visual semantic truth set, the process returns to the separation and retrieval module. When the consistency verification result indicates that the consistency verification has passed, the target sample corresponding to the current experimental sample access instruction is identified as a valid sample that has been captured and verified, thus completing the intelligent access management of experimental samples.
[0085] In Example 1, controlling the mechanical actuator to perform an abnormal release operation includes controlling the mechanical actuator to return the clamped non-target sample to the original operation position, or to transfer the non-target sample to a preset abnormal item buffer area and release it.
[0086] Example 2: In a storage system for automated high-density sample retrieval, the target storage device is a multi-layer cryogenic storage structure. Each layer is equipped with multiple regularly arranged sample trays, and cryovials are densely packed in rows and columns within each tray. During an automated sampling task, the system receives a sample retrieval command requesting the removal of a sample with the target RFID information "RF-07-3C-1842" and its delivery to a preset sample output station. The target sample is a cryovial with an appearance highly similar to the surrounding samples. The tube is transparent, has a narrow label, a light blue identification ring on the cap, and the internal liquid is light yellow. Due to the long-term dense arrangement of the batch of samples, and the proximity of some samples to liquid reagent kits and metal rails, the system had previously experienced cross-reading between adjacent storage locations and sample misalignment.
[0087] The system initiates a global RFID antenna array scan of the target storage device. Within the current global RFID scan cycle, six rounds of scanning are performed, involving 12 groups of RFID antennas. A total of 1168 tags are read, of which 1104 are stably read and 64 are weakly read or have intermittent reads. For the target RFID identification information "RF-07-3C-1842", the system extracts 31 valid read events from the multiple rounds of reading results. The average RFID signal strength is in the top 18% of all samples, and the phase consistency is in the top 22%. RFID candidate identification information corresponding to the target is generated, and a RFID confidence score of 0.91 is added to the global candidate list. Based on the RFID signal strength and phase distribution, the system determines two candidate layers. The layer matching result for the current pallet layer is 0.87, and the layer matching result for the adjacent upper layer is 0.63. At the storage location level, the system obtains three main candidate storage locations, corresponding to three adjacent slots in the same column, with storage location matching results of 0.84, 0.69, and 0.55, respectively. The RFID candidate identity set contains the target's radio frequency identity information, the RFID candidate layer set converges to two layers, and the RFID candidate storage location set converges to three storage locations. The results indicate that relying solely on radio frequency sensing cannot directly and uniquely confirm the actual storage location of the target sample, as typical adjacent storage location interference exists.
[0088] The system determines the candidate spatial region based on the RFID candidate location set and the RFID candidate layer set. Using the three locations with the highest RFID matching results as the center, it expands outward by one slot both horizontally and vertically, simultaneously including the current layer and the adjacent upper layer. The resulting candidate spatial region covers two layers and eight location units. The visual acquisition device performs block image acquisition of the region, obtaining a total of 16 candidate spatial region image data. Each image records the global physical coordinates of the corresponding acquisition center. Based on the actual situation reflected in the images, there is a cryopreservation tube with a similar appearance in the target theoretical location of the current layer, but the label is partially obscured; in the corresponding column of the adjacent upper layer, there is also a cryopreservation tube with a light blue identification ring, the label orientation is clearer, and the liquid color is closer to the target sample.
[0089] The system inputs image data into the Florence-2 multimodal visual model, which has been improved based on the constraints of the intelligent storage and management scenario for experimental samples. Instead of performing an unconstrained search on the entire image, the model constructs a layer constraint prompt sequence based on the candidate layer set, a cargo location topology prompt sequence based on the candidate cargo location set, and a location anchoring prompt sequence based on the global physical coordinates of the acquisition center of the segmented images. After the combined prompt input, the model resolves 11 visual semantic ground truth units from 16 images and simultaneously outputs a visual slot occupancy relationship mapping. For the theoretical target cargo location in the current layer, the model only identifies "...1847" in the sample text label description, and the sample appearance semantic description is a transparent cryovial, a light blue identification ring, and a partially rolled-up label. The sample color attribute is a light yellowish liquid. For another sample in the adjacent upper layer, the model identifies "...1842" in the sample text label description, and the sample appearance semantic description is a transparent cryovial, a light blue identification ring, a flat label, and the sample color attribute is a light yellowish liquid. Furthermore, the corresponding slot occupancy status of the sample is clearly occupied. Through physical coordinate mapping, the model further provides sample spatial location information for two key visual semantic truth units: the current layer sample is located near the center of the theoretical storage location with a coordinate deviation of 2.8 mm; the adjacent upper layer sample is located near the center of the corresponding storage location with a coordinate deviation of 1.9 mm.
[0090] During the heterogeneous data fusion and verification phase, the system performs association matching one by one with the RFID candidate identity set, RFID candidate cargo location set, and RFID candidate layer set, and the visual semantic truth value set. For the current theoretical target cargo location sample, the identity association result is only 0.42, the layer association result is 0.87, the cargo location association result is 0.81, and the joint association score is 0.66. For the adjacent upper layer sample, the identity association result reaches 0.95, the layer association result is 0.63, the cargo location association result is 0.74, and the joint association score is 0.82. The system enters the consistency judgment phase. Since the identity consistency result of the current layer sample is lower than the preset identity consistency threshold, and the layer consistency result of the adjacent upper layer sample, although not low, conflicts with the current optimal layer, neither of them meets the requirement that the radio frequency identity information, radio frequency cargo location information, and radio frequency layer information simultaneously meet the preset consistency conditions with the visual semantic features. The consistency candidate result set is empty, and the system determines that the consistency conditions for this round are not met.
[0091] The system triggers the item code separation and retrieval mechanism. The current retrieval round count increases from 0 to 1. The system first selects the result with the highest joint association score from the fusion verification results as the reference fusion verification result, and extracts the sample appearance semantic description, sample text label description, and sample color attribute to form retrieval semantic anchoring information. At the same time, based on the current pallet area and the adjacent layer information in the RFID candidate layer set, a larger retrieval area is constructed. The retrieval area ultimately covers all 20 storage locations of the current pallet, as well as 20 storage locations in the same column and adjacent columns of the adjacent upper layer, for a total of 40 storage location units. The system generates a retrieval prompt sequence based on the target RFID identity information, retrieval semantic anchoring information, and slot drift constraint information, and again calls the Florence-2 multimodal vision model to perform open-vocabulary semantic retrieval. The second round of retrieval yields 7 retrieval candidate results, of which 3 candidate results have a significant textual correlation with "1842". The system further calculates the semantic back-finding matching results. For candidate samples located one storage location to the right of the adjacent upper layer position, the identity text back-finding matching result is 0.96, the appearance color back-finding matching result is 0.91, and the semantic completeness result is 1.00, resulting in a semantic back-finding matching result of 0.95. The system combines the candidate sample's layer back-finding matching result (0.89 in the height direction), storage location back-finding matching result (0.86 in the planar direction), and current pallet priority result (0.72) to obtain a weighted back-finding positioning score of 0.903. After inverse spatial mapping, the candidate spatial location information corresponding to the candidate sample is updated to the upper layer position and the first slot on the right in the same column. The RFID layer information is also updated to the upper layer position. Based on this, the system updates the RFID candidate identity set, RFID candidate storage location set, and RFID candidate layer set, and returns to the fusion verification stage to re-execute the association matching.
[0092] After the update, heterogeneous data fusion verification and consistency determination are performed again. At this point, the consistency score for the target candidate sample is 0.95, the stratum consistency score is 0.89, the location consistency score is 0.86, the joint association score is 0.93, and the consistency confirmation score is 0.914. The RF identity information corresponding to the result is consistent with the target RF identity information in the access instruction, and the consistency candidate result set contains only this fusion verification result. The system outputs the unique identity information of the target sample, "RF-07-3C-1842," and the precise spatial location information of the target sample, generating a consistency confirmation result.
[0093] During the mechanical execution phase, the system determines the source operation position based on the precise spatial location information of the target sample, and, based on the operation type information in the access command (which is a retrieval operation), determines the position corresponding to the preset sample removal station as the target operation position. The mechanical actuator calculates the grasping approach posture, grasping posture, safe transition posture, release approach posture, and release posture. The target sample is a standard cryopreservation tube. The system calls preset clamping specification parameters, determining the standard clamping opening width to be 14.0 mm, the safe opening compensation to be 2.0 mm, the standard clamping closing width to be 11.2 mm, and the standard clamping force to be 1.8 N. Combining the spatial distance between the source and target operation positions, the system calculates the approach motion duration to be 0.42 seconds and the transfer motion duration to be 1.36 seconds, and generates the motion control parameters for the mechanical actuator. The mechanical actuator sequentially completes the grasping approach, clamping grasp, safe extraction to the transition posture, target position transfer, and release placement according to the control parameters. The entire mechanical execution process takes 2.84 seconds, and no collision alarm occurs during the execution process.
[0094] After the mechanical actuator completes the sampling operation, the system immediately activates the end-effector vision acquisition device to collect post-grab image data, acquiring a total of 6 post-grab images. Based on the end-effector clamping area corresponding to the current state of the mechanical actuator, the system determines the post-grab observation area and constructs a post-grab observation prompt sequence. The post-grab image data and the post-grab observation prompt sequence are input together into the Florence-2 multimodal vision model. The model outputs 3 post-grab visual semantic ground truth units. One post-grab visual semantic ground truth unit located at the center of the clamping area provides the post-grab sample text label description as "RF-07-3C-1842," the post-grab sample appearance semantic description as a transparent cryopreservation tube, a light blue identification ring, a flat label, and the post-grab sample color attribute as a light yellow liquid. The system performs a post-grab consistency check. The identity verification matching result of the captured visual semantic truth unit is 0.98, the semantic completeness result is 1.00, and the clamping region constraint result is 0.94. The weighted fusion of these three results yields a post-capture consistency verification result of 0.973. Since only this one post-capture visual semantic truth unit has a post-capture consistency verification result higher than the preset post-capture verification threshold of 0.85, the system determines that the consistency verification has passed and identifies the target sample corresponding to the current experimental sample access command as a valid sample that has completed the post-capture verification. This completes the entire process of intelligent sample access management for this experiment.
[0095] In 120 consecutive simulated automatic sampling tasks of the same type, when using the traditional RFID standalone judgment method, there were 13 instances of misreading, 9 instances of misjudging the location, and 5 instances of inconsistency in the samples after grabbing, with a final first-time success rate of 77.5%. When using the traditional visual standalone recognition method, there were 11 instances of text misrecognition, 7 instances of mismatching similar appearances, and 4 instances of failure to verify after grabbing, with a first-time success rate of 81.7%. When using the method of this invention, inconsistencies were found in the first consistency judgment stage in 14 out of 120 tasks. Among these, 11 instances were successfully relocated and sampled through the object code separation and retrieval mechanism, 2 instances ended the retrieval process due to failure to locate and reaching the preset maximum number of retrieval attempts, and 1 instance was successfully retrieved after an anomaly was found in the post-grabbing verification stage and an anomaly was released. The method of this invention successfully completed 116 out of 120 tasks, achieving a first-pass sampling rate of 96.7%, an average identity recognition accuracy of 98.3%, an average cargo location positioning error of 2.1 mm, and a post-grab verification pass rate of 97.5%. Experiments demonstrate that this invention can achieve realistic, continuous, and effective automated intelligent storage and retrieval management of experimental samples in complex scenarios involving dense sample arrangement, radio frequency crosstalk, cargo location offset, and object-code separation, through specific calculation and decision-making results at each step.
[0096] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A machine vision-based intelligent sample access management system for an experiment, characterized in that, include: The radio frequency scanning module uses an RFID radio frequency antenna array to perform a global radio frequency scan on the target storage device, generating an RFID candidate identity set, an RFID candidate cargo location set, and an RFID candidate layer set. The image acquisition module receives access instructions, determines the candidate spatial area based on the RFID candidate cargo location set and the RFID candidate layer set, and acquires the corresponding image data. An improved Florence-2 module was developed by inputting image data into the improved Florence-2 multimodal vision model to obtain a set of visual semantic truth values. The fusion verification module associates and matches the RFID candidate identity set, the RFID candidate cargo location set, and the RFID candidate layer set with the visual semantic truth value set to generate a fusion verification result set. The determination module performs consistency determination based on the fusion verification result set, outputs the unique identity information and precise spatial location information of the target sample, and generates a consistency confirmation result. When the consistency condition is not met, the separation and retrieval module generates retrieval and positioning information with the help of the improved Florence-2 multimodal vision model, updates the RFID candidate identity set, RFID candidate cargo location set and RFID candidate layer set and then returns to the fusion verification module. When the retrieval and positioning fails or the current retrieval round count reaches the preset maximum number of retrievals, the item code separation and retrieval process ends. After obtaining the consistency confirmation result, the execution control module converts the precise spatial position information of the target sample into motion control parameters of the mechanical actuator, and controls the mechanical actuator to complete the target sample grabbing, transfer or storage operation; After the mechanical actuator completes its operation, the grasping module collects the grasped image data and inputs it into the improved Florence-2 multimodal vision model to generate a set of visual semantic truth values after grasping. It then performs a consistency check. If the consistency check fails, it triggers an exception handling process and returns to the step separation and retrieval module. If the consistency check passes, it completes the experimental sample storage and retrieval management. 2.The machine vision-based intelligent sample access management system according to claim 1, wherein, The radio frequency scanning module includes: The target area of the target storage device is read multiple times by the RFID radio frequency antenna array to obtain the radio frequency identity information, radio frequency signal strength, radio frequency phase and antenna spatial identification of each RFID tag, forming radio frequency spatiotemporal sensing data; Based on radio frequency spatiotemporal sensing data, the identity aggregation result corresponding to each RFID tag is determined, and an RFID candidate identity set is generated. Feature comparison is performed between the RFID candidate identity set and the device layer radio frequency reference model to determine the layer matching result and generate the RFID candidate layer set. Feature comparison is performed between the RFID candidate identity set and the device layer radio frequency reference model to determine the layer matching result and generate the RFID candidate layer set. Based on the identity aggregation results, hierarchical matching results, and cargo location matching results, corresponding radio frequency credibility is added to the RFID candidate identity set, RFID candidate hierarchical set, and RFID candidate cargo location set.
3. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The image acquisition module includes: Receive the experimental sample access instruction and parse out the target radio frequency identification information corresponding to the target experimental sample. Based on the target radio frequency identification information, extract the corresponding target candidate storage location subset and target candidate layer location subset from the RFID candidate storage location set and the RFID candidate layer location set, respectively. The height range of the target candidate layer location subset is used as the height constraint boundary of the candidate spatial region, and the horizontal and vertical range of the cargo location corresponding to the target candidate cargo location subset are used as the planar constraint boundary of the candidate spatial region. The candidate spatial region is formed under the joint constraint of the height constraint boundary and the planar constraint boundary. Based on the spatial boundary coordinates of the candidate spatial region, the center coordinates, horizontal span, and vertical span of the region are calculated to determine the number of images to be acquired for the candidate spatial region. The control vision acquisition device uses the regional center coordinates of the candidate spatial region as the acquisition center. During the acquisition process, it synchronously records the global physical coordinates of the acquisition center corresponding to each block image, and generates image data of the candidate spatial region containing the physical coordinate mapping relationship.
4. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The improved Florence-2 module includes: Image data of candidate spatial regions are input into the improved Florence-2 multimodal vision model in the form of block image data. Based on the global physical coordinates of the acquisition center, a layer constraint prompt sequence, a cargo location topology prompt sequence, and a location anchoring prompt sequence are constructed respectively. These are then spliced together in order to obtain a joint prompt sequence. Each block of image data and its corresponding joint cue sequence are input into the improved Florence-2 multimodal vision model, and slot-guided visual semantic joint decoding is performed to obtain the slot visual semantic response set corresponding to each block of image data. Based on the sample foreground mask, slot visual region and semantic integrity corresponding to the visual semantic response of each slot, the occupancy confidence corresponding to the visual semantic response of each slot is calculated, and the visual slot occupancy relationship mapping corresponding to each block image data is generated according to the occupancy confidence. For the visual semantic response of each slot occupied by the experimental sample, the centroid pixel coordinates of the sample foreground are calculated based on the sample foreground mask, and combined with the physical coordinate mapping relationship of the corresponding block image data, the spatial position information of the sample is obtained, and a visual semantic truth unit is constructed. Calculate the cross-block consistency association score for visual semantic truth value units, and merge visual semantic truth value units whose cross-block consistency association score is not lower than the preset merging threshold to obtain the visual semantic truth value set.
5. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The fusion verification module includes: Based on the sample spatial location information corresponding to each visual semantic truth unit in the visual semantic truth set, and perform layer association matching with each radio frequency layer information in the RFID candidate layer set, the layer association result is obtained. Based on the spatial location information of the sample corresponding to each visual semantic truth unit and the radio frequency location information in the RFID candidate location set, location association matching is performed to obtain the location association result. Identity association matching is performed based on the radio frequency identity information in the RFID candidate identity set and the sample text label description corresponding to each visual semantic truth unit. The identity association result is obtained by combining the sample appearance semantic description, sample text label description and sample color attribute formation corresponding to each visual semantic truth unit. Based on the identity association results of each visual semantic truth unit relative to each radio frequency identity information, the hierarchical association results of each visual semantic truth unit relative to each radio frequency layer information, and the cargo location association results of each visual semantic truth unit relative to each radio frequency cargo location information, joint association matching is performed to generate a fusion verification result set.
6. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The determination module includes: Based on the correspondence between the radio frequency identity information and visual semantic features corresponding to each fusion verification result, the identity consistency result corresponding to each fusion verification result is calculated, and the identity consistency judgment result corresponding to each fusion verification result is determined based on the identity consistency result. Based on the correspondence between the radio frequency layer information and the sample spatial location information corresponding to each fusion verification result, the layer consistency result corresponding to each fusion verification result is calculated, and the layer consistency judgment result corresponding to each fusion verification result is determined based on the layer consistency result. Based on the correspondence between the radio frequency location information and the sample spatial location information corresponding to each fusion verification result, the location consistency result corresponding to each fusion verification result is determined. The location consistency result is compared with the preset location consistency threshold to determine the location consistency judgment result corresponding to each fusion verification result. The following fusion verification results are retained: identity consistency judgment result is passed, layer consistency judgment result is passed, cargo location consistency judgment result is passed, the corresponding radio frequency identity information matches the target radio frequency identity information in the access instruction, and the consistency confirmation score is not lower than the preset consistency confirmation threshold, forming a consistency candidate result set; When the set of candidate consistency results contains only a unique fusion verification result, the radio frequency identification information and sample spatial location information corresponding to the fusion verification result are extracted and used as the unique identification information and precise spatial location information of the target sample, respectively, to generate a consistency confirmation result.
7. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The separation and retrieval module includes: When the consistency condition is not met, the item code separation and retrieval mechanism is triggered, the current retrieval round count is incremented by one, and a retrieval area and retrieval prompt sequence for open vocabulary semantic retrieval are constructed based on the target radio frequency identification information in the access instruction, the current pallet area, and the RFID candidate layer set. The retrieved image data corresponding to the retrieved area and the retrieved prompt sequence are input together into the Florence-2 multimodal vision model, which has been improved by the constraints of the intelligent storage and management scenario of experimental samples. Open vocabulary semantic retrieval for the object code separation scenario is performed to obtain a set of retrieved candidate results. Based on the semantic description of the appearance of the candidate sample, the text label description of the candidate sample, the color attribute of the candidate sample, the candidate foreground mask and the semantic retrieval score corresponding to each candidate result, the semantic retrieval matching result corresponding to each candidate result is calculated. Based on the physical coordinate mapping relationship between the candidate foreground mask and the corresponding image block corresponding to each candidate retrieval result, the candidate spatial location information corresponding to each candidate retrieval result is determined. Combined with the semantic retrieval matching result, the layer retrieval matching result, the cargo location retrieval matching result and the current pallet priority result, the retrieval positioning information corresponding to each candidate retrieval result is determined. The RFID candidate identity set, RFID candidate cargo location set, and RFID candidate layer set are updated based on the retrieved location information. The system then determines whether to return to the fusion verification module based on the current retrieve round count, the preset maximum retrieve count, and the retrieved location score.
8. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The execution control module includes: Obtain the consistency confirmation results and experimental sample access instructions; determine the source operation position based on the precise spatial position information of the target sample in the consistency confirmation results; and determine the target operation position based on the operation type information in the experimental sample access instructions. Based on the source operating position, target operating position, preset approach height, preset departure height, preset safe obstacle avoidance height, and grasping compensation distance, calculate the corresponding grasping approach posture, grasping posture, safe transition posture, release approach posture, and release posture of the mechanical actuator; Based on the sample clamping specification parameters and spatial distances between each pose corresponding to the unique identity information of the target sample, calculate the clamping control parameters and kinematic timing control parameters of the mechanical actuator. The grasping approach pose, grasping pose, safe transition pose, release approach pose, and release pose are converted into joint control parameters corresponding to each joint of the mechanical actuator, and combined with the clamping control parameters and kinematic timing control parameters to generate the motion control parameters of the mechanical actuator. The mechanical actuator is controlled according to the motion control parameters of the mechanical actuator, so that the end effector sequentially completes the target sample grasping and approaching, clamping and grasping, safe pulling out to the transition position, target position transfer and release placement, thus completing the target sample grasping, transfer or storage operation.
9. The machine vision-based intelligent storage and management system for experimental samples according to claim 1, characterized in that, The crawling module includes: The control end vision acquisition device acquires image data after the mechanical actuator completes the grasping, transfer or storage operation of the target sample, and determines the priority observation range of the target sample in the image data after grasping based on the clamping area of the end actuator corresponding to the current state of the mechanical actuator, thus forming the observation area after grasping; Based on the observed region after capture, a capture observation prompt sequence is constructed. The captured image data and the capture observation prompt sequence are input together into the Florence-2 multimodal vision model to generate a capture visual semantic ground value set. Consistency verification is performed based on the set of visual semantic truth values after capture and the unique identity information of the target sample. The consistency verification result corresponding to each set of visual semantic truth values after capture is calculated. Perform a consistency review decision based on the consistency review results of each crawled data, and generate a consistency review result; When the consistency verification result indicates that the consistency verification has failed, the control mechanical actuator performs an abnormal release operation to clear the end effector. The current captured image data, the captured visual semantic truth set, and the unique identity information of the target sample are used as input information for the abnormal handling process. After updating the actual identity status of the corresponding source operation position or setting a negative sample masking label based on the captured visual semantic truth set, the process returns to the separation and retrieval module. When the consistency verification result indicates that the consistency verification has passed, the target sample corresponding to the current experimental sample access instruction is identified as a valid sample that has been captured and verified, thus completing the intelligent access management of experimental samples.