Automatic guiding robot system for material transportation in traditional Chinese medicine efficacy substance screening laboratory
The automated guided robot system for material transportation in the screening laboratory of traditional Chinese medicine active substances, which integrates lidar and deep learning visual recognition, solves the problems of low efficiency, poor accuracy and contamination risk in material transportation in the screening laboratory of traditional Chinese medicine active substances, and realizes efficient and stable material transportation and management.
Patent Information
- Application Number
- CN202511332112.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-09
AI Technical Summary
Existing laboratories for screening active substances in traditional Chinese medicine suffer from problems such as low efficiency, poor accuracy, easy contamination, and lack of automation and intelligence in material transportation. In particular, in laboratory environments with high cleanliness requirements and complex spaces, traditional robot systems struggle to achieve accurate identification, stable grasping, and path planning.
An automated guided robot system for transporting materials in a laboratory for screening active substances of traditional Chinese medicine was constructed by integrating a lidar module, an environmental camera module, a biomimetic perception large model module, a large model ROS communication module, a ROS module, and an intelligent robotic arm module. The system utilizes lidar for precise mapping and navigation, and combines deep learning-based visual recognition and grasping algorithms to achieve accurate identification and stable grasping of experimental equipment such as cell plates.
It enables efficient, stable, and unmanned management of laboratory material transportation, reduces labor costs, avoids pollution risks, and is suitable for laboratory environments with high cleanliness and complex spaces, providing a full-process intelligent management solution.
Smart Images

Figure CN121083709A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of automation control technology, and in particular relates to an automatic guided robot system for material transportation in a laboratory for screening active substances of traditional Chinese medicine. Background Technology
[0002] As a crucial upstream link in the new drug development process, the screening laboratory for active ingredients in traditional Chinese medicine (TCM) is primarily responsible for tasks such as extraction, initial screening, and cell experiments. The experimental procedures are highly precise, demanding extremely high levels of cleanliness and standardization in the experimental environment. In actual operation, numerous experimental instruments, such as cell culture plates, reagent kits, and reaction vessels, need to be frequently transferred between multiple areas, including transfer windows, circular stacks, and workbenches. Currently, most laboratories still rely on manual material handling, which is not only inefficient and labor-intensive but also prone to inconsistencies, cross-contamination, and human error. This severely restricts the stability, reproducibility, and high-throughput processing capabilities of the experiments, failing to meet the growing demand for automation and standardized operations in modern TCM screening laboratories.
[0003] In recent years, with the rapid development of emerging technologies such as Automated Guided Vehicles (AGVs), LiDAR navigation, collaborative robotic arms, image recognition, deep learning algorithms, and Robot Operating System (ROS), laboratory material handling systems based on intelligent robot platforms have become a research hotspot. Although some laboratories have introduced robots for initial exploration, traditional robot systems still face several challenges in complex experimental environments. For example, limited perception range makes it difficult to accurately identify equipment positions and obstacles; unstable path planning makes it difficult to adapt to dynamically changing environments; insufficient grasping and positioning accuracy makes it impossible to reliably handle lightweight materials such as cell plates; and the lack of a unified communication framework among system modules leads to low integration, slow response speed, and poor scalability. These problems are particularly prominent in scenarios such as traditional Chinese medicine laboratories, which have limited space, dense equipment, and high cleanliness requirements, severely limiting the practical application effectiveness of traditional robot systems. Summary of the Invention
[0004] In view of this, the present invention aims to propose an automated guided robot system for material transportation in the laboratory for screening active substances of traditional Chinese medicine, so as to solve the problems of low efficiency, poor accuracy, easy contamination, and lack of automation and intelligence in the existing manual material handling in laboratories.
[0005] To achieve the above objectives, the technical solution of the present invention is implemented as follows: An automated guided robot system for material transport in a laboratory for screening active substances of traditional Chinese medicine includes a lidar module, an environmental camera module, a biomimetic sensing large model module, a large model ROS communication module, a ROS module, a drive chassis module, and an intelligent robotic arm module. The intelligent robotic arm module is used to collect image data of icon recognition points and grasp materials based on the image data. The intelligent robotic arm module includes a robotic arm unit, a target camera unit, a perception node, a grasping planning node, and a Moveit control node. The target camera unit is used to acquire image data of icon recognition points, use its built-in convolutional neural network model based on multi-dimensional feature fusion to locate the material storage location, and publish the location data message to the ROS module. The backbone of the convolutional neural network model consists of 4-stage hybrid dilated convolutional layers, and before the features are input into the backbone, attention weights are generated by a two-dimensional attention mechanism to process the original features. The lidar module is used to collect laboratory environmental data in real time, generate an obstacle grid map, and send distance data messages to the ROS module. The environmental camera module is used to acquire images of the laboratory environment in real time and send the image data messages to the biomimetic sensing large model module and the ROS module; The biomimetic perception large model module is used to perform visual analysis based on the image data messages to perceive the surrounding environment in real time and provide real-time descriptions of the surrounding environment. It obtains relevant environmental data and acquires operator instructions while responding to them. It is also used to send the relevant environmental data and the operator instructions to the large model ROS communication module. The large model ROS communication module is used to understand the relevant semantic information of the instructions given by the large model of different modalities, and after converting the understood semantic information into a standardized format, it aligns the semantic information in the standardized format with the ROS module interface according to the historical multi-round tasks, and sends the specific ROS system instructions to the ROS module. The ROS module is used to build relevant topics based on the location data message, distance data message, image data message, and specific ROS system instructions. The drive chassis module is used to subscribe to relevant topics in the ROS module and to drive the robot based on three-dimensional coordinates.
[0006] Furthermore, the first layer of the four-stage hybrid dilated convolutional layer in the backbone of the convolutional neural network model consists of a dilated convolutional layer with a dilation rate of 2 and a filter size of 3×3, the second layer is a standard convolutional layer with a filter size of 3×3, the third layer is a normalization layer, and the fourth layer is a ReLU activation function layer.
[0007] Furthermore, the two-dimensional attention mechanism generates attention weights to process the original features, specifically including the following steps: The original features are first subjected to average pooling to preserve local average features; Adjust the channel using a standard convolution with a filter size of 1×1, then first perform a convolution operation in the vertical direction using a convolution with a filter size of 1×3, then perform a convolution operation in the horizontal direction, and finally adjust the channel again using a standard convolution with a filter size of 1×1. The formula for vertical convolution is as follows: In the above formula, This represents the image data after adjustment by the convolutional layer. The shape of the weights in the convolutional layer; Secondly, the formula for horizontal convolution is as follows: In the above formula, This represents the image data after adjustment following the vertical convolution operation. The shape of the weights in the convolutional layer; An attention map is generated using the Sigmoid function, and the attention map is multiplied by the original features before being fed into the backbone to process the original features.
[0008] Furthermore, the biomimetic perception large model module includes a visual unit, an auditory unit, and a statement unit; The vision unit is used to perform real-time analysis of the image data messages based on the underlying vision big model. After encoding the image and extracting it into semantic vectors, it performs multimodal alignment to generate text language and obtain relevant environmental data. The auditory unit is used to understand operator instructions based on the underlying speech conversion model and convert operator speech instructions into text language. The statement unit is used to convert the text language generated by the visual unit into speech based on the underlying speech generation model, and to broadcast it in real time according to the operator's instructions.
[0009] Furthermore, the specific steps for the visual unit to process image data messages include: The formula for obtaining the laboratory environment image from the image data message, i.e., the input image, is as follows: In the above formula, Image height, Image width, This refers to the number of RGB channels. The input image is further segmented into indivual Each The size is The formula is as follows: ; Each Flattening it into a vector, the formula is as follows: In the above formula, ; Through linear embedding matrix E Mapped to d Dimensions, plus position encoding The formula is as follows: ; The resulting image token sequence is: ; Extract semantic vectors, Send in The multi-head attention mechanism layer processes data layer by layer, as shown in the following formula: ; In the above formula, , This indicates the attention of the bulls. Indicates a feedforward network; The final output visual feature sequence is: ; Introduction The query vectors are then subjected to multimodal alignment, as shown in the following formula: In the above formula, ; An attention mechanism is applied to each query vector and the visual feature sequence, as shown in the following formula: In the above formula, , , It is a visual feature sequence. The attention weights are calculated using the following formula: In the above formula, and Represents a learnable linear transformation matrix, where the dimension is 1. d × d ; get Language space cue vectors: ; Spatial cue vectors are concatenated to generate textual language, thus obtaining relevant environmental data.
[0010] Furthermore, the large model ROS communication module includes a context encapsulation unit, a model call interface unit, a format standardization unit, and a memory unit; The context encapsulation unit is used to transform the semantics of instructions given by large models of different modalities into a unified JSON semantic context; The model call interface unit generates a model call request based on the JSON semantic context and converts the data into an understandable input format; The format standardization unit is used to standardize the structural format of the output of large models with different modalities and to ensure the alignment of ROS module instruction interfaces. Memory units are used to retain long-term contextual information across multiple historical task iterations, building a long-term memory bank and continuously updating it.
[0011] Furthermore, the ROS module includes a topic building unit, a message transmission unit, and a visualization unit; The topic building unit is used to build related topics for multimodal data based on distance data messages, image data messages, positioning data messages, and specific ROS system commands; The message transmission unit is used to receive communication instructions and transmit specific data; The visualization unit is used to display the robot's location information in the laboratory in real time through an Rviz window.
[0012] Compared with existing technologies, the automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine described in this invention has the following advantages: (1) The automated guided robot system for material transportation in the screening laboratory of traditional Chinese medicine active substances described in this invention integrates a lidar module, an environmental camera module, a biomimetic perception large model module, a large model ROS communication module, a ROS module, a drive chassis module, and an intelligent robotic arm module to construct a highly integrated and functionally coordinated automated guided robot system specifically for material transfer in the screening laboratory of traditional Chinese medicine active substances. This system uses lidar for precise mapping and positioning navigation, achieves multi-directional flexible movement through Mecanum wheel drive, and combines deep learning-based visual recognition and grasping algorithms to achieve accurate identification, stable grasping, and fully automated handling of experimental equipment such as cell plates. The entire system requires no human intervention during operation, effectively solving the problems of low efficiency, large errors, and high pollution risk in traditional laboratory handling processes, significantly reducing labor costs and improving logistics scheduling efficiency.
[0013] (2) The automated guided robot system for material transport in the screening laboratory for medicinal substances of traditional Chinese medicine described in this invention is particularly suitable for laboratory environments with high cleanliness requirements, complex spatial structures, and strong reliance on automation, such as screening laboratories for medicinal substances of traditional Chinese medicine, biopharmaceutical workshops, and cell culture rooms. In these scenarios, laboratory operations require extremely high aseptic assurance and standardized operations. This system has functions such as non-contact grasping, automatic obstacle avoidance, and intelligent visual recognition, and can operate flexibly in confined spaces while avoiding contamination or process interruptions caused by human operation. In addition, this system can seamlessly connect with laboratory equipment such as pass-through windows, stacks, and conveyor belts to achieve intelligent management of the entire process from material receiving, test loading, result recovery to material classification and storage, providing a safe, efficient, and stable solution for laboratory automation construction, and has extremely high practical value and promotion prospects. Attached Figure Description
[0014] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a schematic diagram of the automated guided robot system for transporting materials in the laboratory for screening the active substances of traditional Chinese medicine, as described in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the screening laboratory for medicinal substances of traditional Chinese medicine as described in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the convolutional neural network model based on multi-dimensional feature fusion in the automated guided robot system for material transportation in the laboratory for screening active substances of traditional Chinese medicine as described in this embodiment of the invention. Figure 4 This is a schematic diagram of the two-dimensional attention mechanism in the automated guided robot system for material transport in the laboratory for screening the active substances of traditional Chinese medicine as described in this embodiment of the invention. Detailed Implementation
[0015] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0016] This invention provides an automated guided robot system for material transportation in a laboratory for screening active substances of traditional Chinese medicine. The system integrates a lidar module, an environmental camera module, a biomimetic sensing large model module, a large model ROS communication module, a ROS module, a drive chassis module, and an intelligent robotic arm module. Its core objective is to ensure automated material transportation in the laboratory for screening active substances of traditional Chinese medicine through technical means, thereby further improving the intelligent and continuous experimental operation of the laboratory.
[0017] In practical applications, the aforementioned robot system integrates a lidar module, a 4-axis intelligent robotic arm, a deep learning vision recognition unit, and a ROS modular communication architecture, possessing multiple functions such as environmental mapping, autonomous navigation, object recognition and grasping, and obstacle avoidance. The system can collect and process multi-source sensor data in real time, construct an environmental grid map, identify the location of experimental equipment, accurately plan grasping paths, and drive the robotic arm to collaboratively complete the automated transport of experimental materials such as cell plates. Compared to traditional methods, this invention significantly improves the efficiency, stability, and operational consistency of material transportation, significantly reduces human intervention and errors, and lays a technological foundation for achieving fully automated management in traditional Chinese medicine efficacy screening laboratories.
[0018] Specifically, an automated guided robot system for material transport in a laboratory for screening active substances of traditional Chinese medicine includes a lidar module, an environmental camera module, a biomimetic sensing large model module, a large model ROS communication module, a ROS module, a drive chassis module, and an intelligent robotic arm module.
[0019] The intelligent robotic arm module is used to collect image data of icon recognition points and grasp materials based on the image data. The intelligent robotic arm module includes a robotic arm unit, a target camera unit, a perception node, a grasping planning node, and a Moveit control node. The target camera unit is used to acquire image data of icon recognition points, use its built-in convolutional neural network model based on multi-dimensional feature fusion to locate the material storage location, and publish the location data message to the ROS module. The backbone of the convolutional neural network model consists of 4-stage hybrid dilated convolutional layers, and before the features are input into the backbone, attention weights are generated by a two-dimensional attention mechanism to process the original features.
[0020] The lidar module is used to collect laboratory environmental data in real time, generate an obstacle grid map, and send distance data messages to the ROS module.
[0021] The environmental camera module is used to acquire images of the laboratory environment in real time and send the image data messages to the biomimetic sensing large model module and the ROS module.
[0022] The biomimetic perception large model module is used to perform visual analysis based on the image data messages to perceive the surrounding environment in real time and provide real-time descriptions of the surrounding environment. It obtains relevant environmental data and acquires operator instructions while responding to them. It is also used to send the relevant environmental data and the operator instructions to the large model ROS communication module.
[0023] The large model ROS communication module is used to understand the semantic information of the instructions given by the large model in different modalities, and after converting the understood semantic information into a standardized format, it aligns the semantic information in the standardized format with the ROS module interface according to the historical multi-round tasks, and sends the specific ROS system instructions to the ROS module.
[0024] The ROS module is used to build relevant topics based on the location data message, distance data message, image data message, and specific ROS system instructions.
[0025] The drive chassis module is used to subscribe to relevant topics in the ROS module and to drive the robot based on three-dimensional coordinates.
[0026] Optionally, the LiDAR module specifically includes a laser ranging node and a map generation unit. The laser ranging node can sequentially emit and receive single-line lasers to collect distances between itself and laboratory equipment and obstacles in the traditional Chinese medicine active substance screening laboratory, and then send this distance to the map generation unit. The map generation unit can generate an environmental grid map based on the distance collected by the laser ranging node and publish the data message to the ROS module. The environmental grid map can be a 2D environmental grid map, and the obstacle occupancy value of this grid map is [value missing]. If the obstacle conditions in a grid are unknown, the grid value is -1, and the 2D grid map will be sent to the ROS module.
[0027] Specifically, a laser is emitted using the time-of-flight (TOF) method. After a certain period, the laser beam is reflected and received by a lidar system, and the interval is determined from the sensor message packets (sensor_msgs). The interval distance is calculated from the sensor message packets (sensor_msgs). ,constant Speed of light: The single-line laser, after rotating once by the radar, yields a global range array. A .in A The values contained within the raster matrix cells differ, and the map generation unit fills these values into the raster matrix cells in an orderly manner. The occupancy value of each raster matrix cell is [value to be filled in]. The higher the value, the more obvious the obstacle. The obstacle generation map in this environment is a raster map of the Traditional Chinese Medicine Efficacy Substance Screening Laboratory. This distance data and the raster map will then be sent to the ROS module.
[0028] In practical applications, the LiDAR module collects the distance between the laboratory material storage locations (pass-through windows, circular stacks, and desktop stacks) and the robot. It then generates a grid map from this distance data, which, along with the grid map, is sent to the ROS module. Specifically, the LiDAR module operates as follows: the laser ranging node uses a single-line laser to collect the distance between itself and the laboratory equipment and obstacles used for screening medicinal substances, and sends this distance data to the map generation unit. The map generation unit generates an environmental grid map based on the distance collected by the laser ranging node and publishes the data message to the ROS module.
[0029] Optionally, the environment camera module includes a depth camera unit, which can collect laboratory environmental data (such as the location of laboratory materials) in real time and send it to the biomimetic sensing large model module and the ROS module for real-time perception of the laboratory location.
[0030] In practical applications, the working steps of the environment camera module are as follows: the depth camera unit acquires environmental images in real time and sends them to the biomimetic perception large model module for real-time environmental perception through the topic building unit and message transmission unit.
[0031] Preferably, the biomimetic perception big model module includes a visual unit, an auditory unit, and a statement unit; the visual unit is used to perform real-time analysis of the image data messages based on the underlying visual big model, and after encoding the image and extracting it into semantic vectors, performs multimodal alignment to generate text language to obtain relevant environmental data; the auditory unit is used to understand the operator's instructions based on the underlying speech conversion big model, and convert the operator's speech instructions into text language; the statement unit is used to convert the text language generated by the visual unit into speech based on the underlying speech generation big model, and broadcast it in real time according to the operator's instructions.
[0032] Specifically, the biomimetic perception large-scale model module can perform visual analysis based on environmental images captured by the environmental camera module to perceive the surrounding environment in real time, provide real-time descriptions of the surrounding environment, and transmit relevant environmental data to the large-scale model ROS communication module. The underlying large-scale model of the vision unit analyzes the real-time images from the environmental camera module; the underlying large-scale model of the auditory unit parses the operator's commands; the description unit describes the image analysis from the vision unit, responds to the instructions from the auditory unit, and sends the large-scale model's analysis and instructions to the large-scale model ROS communication module.
[0033] In practical applications, the biomimetic sensing big model module receives data from the environmental camera module and analyzes and describes it by calling different underlying big models. It explores the environment in real time and describes the situation. The analysis results (the robot's own position, task point, distance difference between the position and the task point, etc. At this time, the analysis results are only the self-perception results of the biomimetic sensing big model. The analysis results need to be understood and decomposed by the big model ROS communication module before subsequent operations can be performed) will be connected to the ROS module through the big model ROS communication module to call the robot to perform specific operations.
[0034] The working steps of the biomimetic perception large model module are as follows: (1) Receive real-time images sent by the ROS module, and perform intelligent analysis through the vision unit to determine the real-time environmental position of the vehicle, and send the position data to the statement unit. For example, the underlying big model of the vision unit is the locally deployed open-source Tongyi Qianwen 2.5VL vision big model. This unit recognizes the real-time laboratory environment images processed by the ROS module, encodes the images, extracts them into semantic vectors, performs multimodal alignment, and then generates text language.
[0035] (2) The auditory unit can perceive operator commands, parse the commands, and send the parsed data to the presentation unit. For example, the underlying big model of the auditory unit is a locally deployed open-source STT big model, which converts operator instructions into text language for understanding. In actual use, if there are operator language instructions, the auditory unit calls the underlying STT model to extract the operator's acoustic features, decodes and generates text, and then returns the text to the visual big model for instruction parsing.
[0036] (3) The statement unit verbally describes the location data and broadcasts the current location in real time. If there are operator commands, it responds in real time. For example, the underlying model of the statement unit is a locally deployed open-source TTS model, which converts the text language of the vision unit into spoken language and broadcasts the laboratory environment in real time for the operator to understand. In actual use, the statement unit can annotate the parsed text of the vision unit and output speech features to generate real-time speech, so as to achieve the purpose of real-time exploration and broadcasting for the vehicle. It should be noted that the role of the statement unit is description. Depending on the deployment of different parameters of the large model, it can achieve the purpose of inference after describing the environment, and complete the human-computer interaction more smoothly.
[0037] Preferably, the specific steps for the visual unit to process image data messages include: The formula for obtaining the laboratory environment image from the image data message, i.e., the input image, is as follows: In the above formula, Image height, Image width, This refers to the number of RGB channels. The input image is further segmented into indivual Each The size is The formula is as follows: ; Each Flattening it into a vector, the formula is as follows: In the above formula, ; Through linear embedding matrix E Mapped to d Dimensions, plus position encoding The formula is as follows: ; The resulting image token sequence is: ; Extract semantic vectors, Send in The multi-head attention mechanism layer processes data layer by layer, as shown in the following formula: ; In the above formula, , This indicates the attention of the bulls. Indicates a feedforward network; The final output visual feature sequence is: ; Introduction The query vectors are then subjected to multimodal alignment, as shown in the following formula: In the above formula, ; An attention mechanism is applied to each query vector and the visual feature sequence, as shown in the following formula: In the above formula, , , It is a visual feature sequence. The attention weights are calculated using the following formula: In the above formula, and Represents a learnable linear transformation matrix, where the dimension is 1. d × d ; get Language space cue vectors: ; Spatial cue vectors are concatenated to generate textual language, thus obtaining relevant environmental data.
[0038] Preferably, the large-scale model ROS communication module includes a context encapsulation unit, a model call interface unit, a format standardization unit, and a memory unit. The context encapsulation unit is used to convert the semantics of the instructions given by different modal large models into a unified JSON semantic context. The model call interface unit generates a model call request based on the JSON semantic context and converts the data into an understandable input format. The format standardization unit is used to standardize the structural format of the output of different modal large models and ensure the alignment of ROS module instruction interfaces. The memory unit is used to retain long-term context information in historical multi-round tasks, establish a long-term memory bank, and continuously update it.
[0039] In practical applications, the large-scale model ROS communication module first understands the semantics of instructions given by different modal large models (such as task objectives, sensor data, and environmental states), and integrates the task objectives, sensor data, and environmental states generated in the system into a unified semantic context. Secondly, this module uniformly converts the understood information into JSON format (this structure is necessary for sub-unit understanding). Through custom design of a unified calling protocol or data interface, smooth collaboration between modules can be achieved. Then, based on historical multi-round tasks, this module aligns the information in JSON format with the ROS module interface, providing relevant instructions that the ROS module can understand. By standardizing the structure and format of different model outputs, ROS module instruction alignment can be ensured. It should be noted that the memory unit retains long-term context information throughout multi-round tasks or dialogues for flexible implementation in the next round of tasks.
[0040] The working steps of the large-scale ROS communication module are as follows: (1) The context encapsulation unit concatenates information such as task objectives, perception data, and environmental states generated in the system into unified model context data. For example, this unit first understands the surrounding information: in Image perception ( ) Text instructions (T ) wait.
[0041] Feature extraction is then performed. Finally, a unified JSON semantic structure is output: Example of JSON semantic structure: { "task": "Environmental recognition", "history": ["Previous recognition failed", "Current image to be passed to the window"], "inputs": { "image": "base64-encoded", "params": {"robot_id": "A1", "env_status": "normal"} } }
[0042] (2) The model call interface unit generates a model call request from the context and converts the data into an understandable input format. This unit first generates instructions and then calls the model: Generate instructions: Calling the model: The model then undergoes instruction conversion and adaptation: Finally, the model is encapsulated after execution. (3) Standardized format unit specifications and standard fields are converted into intermediate language based on the commonly used instruction knowledge base of the robot operating system in the screening laboratory for active substances of traditional Chinese medicine: Named entity recognition and linking: Where y represents context encapsulation information. This is a knowledge base of commonly used instructions for a custom robot operating system used in a screening laboratory for active substances in traditional Chinese medicine.
[0043] Field mapping: in intermediate language field set Finally, complete the rules (add missing fields): The model interface is unified through named entity recognition and linking, field mapping and rule completion.
[0044] (4) The memory unit establishes a long-term memory bank and continuously updates it (splicing operation).
[0045] Establish a long-term memory bank: Each It can include information from the last call, historical exception events, etc.
[0046] Memory update: .
[0047] Preferably, the ROS module includes a topic building unit, a message transmission unit, and a visualization unit; the topic building unit is used to build relevant topics for multimodal data based on distance data messages, image data messages, positioning data messages, and specific ROS system commands; the message transmission unit is used to receive communication commands and transmit specific data (distance, images); the visualization unit is used to display the robot's location information in the laboratory in real time through the Rviz window.
[0048] Specifically, the ROS module can build relevant topics, create Rviz visualization windows, and create subscription windows to connect to the driver chassis module based on sensor messages and specific ROS commands from the depth camera unit, LiDAR module, and camera unit. The topic building unit builds topics for multimodal data from the depth camera unit, LiDAR module, and camera unit.
[0049] Optionally, the drive chassis module includes drive nodes and motors (with Mecanum wheels at the ends). This module can subscribe to map topics and drive the robot based on 3D coordinates. In actual use, the drive nodes subscribe to the radar topics of the ROS module, calculate the linear movement and rotation angle of the 3D coordinates based on the obstacle distance, and drive the motors (with Mecanum wheels at the ends).
[0050] In practical applications, the drive node subscribes to the radar topic in the ROS module and uses the radar to measure the distance in real time and the laboratory pass-through window position, desktop stack position and ring stack position of the 2D grid map, and drive the motor (with Mecanum at the end) wheel to reach the designated position.
[0051] Preferably, the intelligent robotic arm module includes a 4-axis robotic arm (with a gripper at the end effector), a camera unit, a sensing node, a grasping planning node, and a Moveit control node. In actual use, the camera unit uses a built-in deep learning algorithm to accurately locate the material storage position (i.e., the transfer window icon, desktop stack icon, and circular stack icon) and publishes data messages to the ROS module; the sensing node subscribes to and processes image data (collected by the camera unit) in the ROS module, and receives image data messages by subscribing to camera topics in the ROS module to perceive the surrounding environment of the robotic arm; the grasping planning node calculates the grasping posture based on the image messages and sends the grasping posture to the Moveit node; the Moveit node performs path planning based on the robotic arm's grasping posture and calls the controller to realize the coordinated grasping of materials by the 4-axis robotic arm (with a gripper at the end effector).
[0052] Specifically, when the camera unit accurately locates the cell plate storage position using a built-in deep learning algorithm, the recognition point is the image. The built-in deep learning algorithm adopts a convolutional neural network algorithm based on multi-dimensional feature fusion, which can further improve the accuracy of different recognition points.
[0053] Preferably, the backbone of the algorithm model consists of a novel four-stage hybrid dilated convolutional layer. To avoid excessive feature loss during downsampling after convolution operations, and to further improve image feature capture, the first layer of the four-stage hybrid dilated convolutional layer backbone of the convolutional neural network model consists of a dilated convolutional layer with a dilation rate of 2 and a filter size of 3×3. The second layer is a standard convolutional layer with a filter size of 3×3. The third layer is a normalization layer, and the fourth layer is a ReLU activation function layer. In practical applications, this algorithm model improves model recognition performance by stacking dilated convolutional and ordinary convolutional layers. In practical applications, before features are fed into the backbone, attention weights need to be generated by a two-dimensional attention mechanism to process the original features. The specific steps include the following: The original features are first subjected to average pooling to reduce some redundant parameters in order to preserve local average features; The system first adjusts the number of channels using a standard convolution with a filter size of 1×1. Then, it performs a convolution operation in the vertical direction using a 1×3 filter size, followed by a convolution operation in the horizontal direction. Finally, it adjusts the number of channels again using a standard convolution with a filter size of 1×1. By extracting features from the data in two different dimensions, the extracted features are richer and the accuracy is higher. The 1×1 standard convolution can be used to adjust the number of data channels, allowing the data to be further fed into subsequent backbone convolutional layers.
[0054] The formula for vertical convolution is as follows: In the above formula, This represents the image data after adjustment by the convolutional layer. The shape of the weights in the convolutional layer; Secondly, the formula for horizontal convolution is as follows: In the above formula, This represents the image data after adjustment following the vertical convolution operation. The shape of the weights in the convolutional layer; An attention map is generated using the Sigmoid function, and the attention map is multiplied by the original features before being fed into the backbone to process the original features.
[0055] This robotic system can be operated by following these steps: Step S1: Use the lidar module to collect distance data between the laboratory material storage location and the robot, generate a grid map from the distance data, and send the distance data and grid map to the ROS module.
[0056] Step S2: Use the environmental camera module to collect the location of laboratory materials, generate image data and send it to the biomimetic sensing large model module and ROS module for real-time analysis.
[0057] Step S3: Receive image data, distance data, grid map and operator instructions using the biomimetic perception large model module, analyze and present them, and connect to the ROS module through the large model ROS communication module.
[0058] Specifically, step S3 includes the following steps: Step S31: The vision unit calls the underlying large vision model to recognize the real-time laboratory environment image processed by the environment camera module, segments the image, and splices it to generate text.
[0059] Step S32: If there is an operator's language instruction, the auditory unit calls the underlying STT model to extract the operator's acoustic features, decodes and generates text, and then returns the text to the visual big model for instruction parsing.
[0060] In step S33, the statement unit annotates the parsed text of the visual unit and outputs speech features to generate real-time speech, so as to achieve the purpose of real-time exploration and broadcasting of the car.
[0061] Step S34: The context encapsulation unit will transform heterogeneous information such as perception, instructions, and environment into a unified JSON semantic structure. The model call interface unit will generate a model call request from the context and convert the instruction data into an understandable input format.
[0062] Step S35: Standardize the format of the unit and standardize the fields, converting them into intermediate language based on the commonly used instruction knowledge base of the robot operating system in the screening laboratory for active substances of traditional Chinese medicine.
[0063] Step S36: The memory unit establishes a long-term memory bank and continuously updates it.
[0064] Step S4: Subscribe to the radar topic of the ROS module using the drive chassis module, and realize the robot drive to reach the transfer window based on the three-dimensional coordinates.
[0065] Specifically, step S4 includes the following steps: Step S41: The topic building unit generates a radar topic and packages the distance measured by the lidar module with the grid map into a message.
[0066] Step S42: The drive node subscribes to the radar topic and receives distance messages measured by the lidar module. Then, the drive node establishes a three-dimensional coordinate system with the robot itself and invokes the motors (with Mecanum wheels at the ends) to unlock linear motion in the x-axis direction, linear motion in the y-axis direction, and rotational motion in the z-axis direction.
[0067] During movement, the drive node adjusts its linear motion speed along the x-axis and y-axis, and its rotation angle along the z-axis (assuming an angular acceleration of ω=0.3rad / s, considering the robot's width) based on distance information measured by the lidar module. This moderate rotation speed effectively avoids collisions, achieving flexible transport. The robot's programmed sequence is to sequentially reach the transfer window, the desktop stack, and the circular stack.
[0068] Step S5: The intelligent robotic arm module collects the icon recognition points of the transfer window and generates image data. Based on the image data, the module successively performs perception, grasping planning, and call control to complete the precise positioning and grasping of the cell plate to be tested. The cell plate is stored in its own robotic arm storage location; then, the drive chassis module drives the robot to the desktop stack.
[0069] Specifically, step S5 includes the following steps: Step S51: The camera unit uses a built-in deep learning algorithm to accurately locate the cell plate storage location, i.e., the transfer window icon.
[0070] Step S52: The topic building unit generates camera topics and packages the image data recognized by the camera unit into messages. The camera topics subscribed to by the perception node receive data message packets from the ROS module and perform 6D pose estimation, i.e., position (3D coordinate system x, y, z) and orientation (roll, pitch, yaw) prediction. This pose estimate will be sent to the grasping planning node.
[0071] Step S53: The grasping planning node receives the pose estimate and, combined with the object image data, generates a graspable pose using the GraspNet deep learning model. The graspable pose is then sent to the Moveit control node.
[0072] Step S54: The Moveit control node receives the graspable posture, calls the controller plugin, and publishes the specific gripper control action message to the ROS module.
[0073] Step S55: The topic building unit generates a controller topic and packages the object recognition gripper control actions published by the Moveit control node into messages. The 4-axis robotic arm (with a gripper at the end) subscribes to the controller topic, grips the physical object (cell board), and stores it in the robotic arm's own storage location.
[0074] Step S56: The drive node continuously receives distance messages measured by the LiDAR module 1. Then, the drive node establishes a three-dimensional coordinate system using the robot itself and calls the motor (with a Mecanum wheel at the end) to move to the desktop stack recognition point.
[0075] Step S6: Once the robot reaches the desktop stack, it recognizes the desktop stack icon and replaces its stored cell plates for experiments with those already used in the experiment. Then, the drive chassis module propels the robot to the circular stack.
[0076] Specifically, step S6 includes the following steps: In step S61, camera unit 22 uses a built-in deep learning algorithm to accurately locate the cell plate storage position (desktop stack icon). The recognition result (image data message) will be sent to the ROS module.
[0077] Step S62: The topic building unit generates camera topics and packages the image data recognized by the camera unit into messages. The camera topics subscribed to by the perception node receive data message packets from the ROS module and perform 6D pose estimation, i.e., position (3D coordinate system x, y, z) and orientation (roll, pitch, yaw) prediction. This pose estimate will be sent to the grasping planning node 24. Step S63: The grasping planning node receives the pose estimate and, combined with the object image data, generates a graspable pose using the GraspNet deep learning model. The graspable pose is then sent to the Moveit control node.
[0078] Step S64: The Moveit control node receives the graspable posture, calls the controller plugin, and publishes the specific gripper control action message to the ROS module.
[0079] Step S65: The topic building unit generates a controller topic and packages the object recognition gripper control actions published by the Moveit control node into messages. The 4-axis robotic arm (with a gripper at its end) subscribes to the controller topic, grips the physical object (the cell plate that has completed the experiment), stores it in the robotic arm's own storage position, and replaces the cell plate that has not completed the experiment, which is transported by the transfer window.
[0080] Step S66: The drive node continuously receives distance messages measured by the LiDAR module. Then, the drive node establishes a three-dimensional coordinate system using the robot itself and uses a motor (with a Mecanum wheel at the end) to move to the circular stack identification point.
[0081] Step S7: The intelligent robotic arm module collects the icon recognition points of the circular stack and generates image data. Based on the image data, the module successively performs perception, grasping planning, and call control to send the experimental cell plates into the circular stack conveyor belt for storage. Afterwards, the drive chassis module drives the robot to the transfer window for continuous transport.
[0082] Specifically, step S7 includes the following steps: In step S71, camera unit 22 uses a built-in deep learning algorithm to accurately locate the cell plate storage position (desktop stack icon). The recognition result (image data message) will be sent to the ROS module.
[0083] Step S72: The topic building unit generates camera topics and packages the image data recognized by the camera unit into messages. The camera topics subscribed to by the perception node receive data message packets from the ROS module and perform 6D pose estimation, i.e., position (3D coordinate system x, y, z) and orientation (roll, pitch, yaw) prediction. This pose estimate will be sent to the grasping planning node.
[0084] Step S73: The grasping planning node receives the pose estimate and, combined with the object image data, generates a graspable pose using the GraspNet deep learning model. The graspable pose will then be sent to the Moveit control node.
[0085] Step S74: The Moveit control node receives the graspable posture, calls the controller plugin, and publishes the specific gripper control action message to the ROS module.
[0086] Step S75: The topic building unit generates a controller topic and packages the body recognition gripper control actions published by the Moveit control node into messages. The 4-axis robotic arm (with a gripper at its end) subscribes to the controller topic, grips the physical object (the cell plate that has completed the experiment), stores it in the robotic arm's own storage position, and replaces the cell plate that has not completed the experiment, which is transported by the transfer window.
[0087] Step S76: The drive node continuously receives distance messages measured by the lidar module. Then, the drive node establishes a three-dimensional coordinate system based on the robot itself and calls the motors (with Mecanum wheels at their ends) to the transfer window for continuous transport.
[0088] The automated guided robot system for transporting materials in the laboratory for screening the active ingredients of traditional Chinese medicine described in this invention solves the problems of low efficiency, poor accuracy, easy contamination, and lack of automation and intelligence in existing manual material handling in laboratories by using robots to transport and replace materials (i.e., cell plates) in the laboratory for screening the active ingredients of traditional Chinese medicine. It realizes the high efficiency, standardization and intelligence of material handling operations in the process of screening the active ingredients of traditional Chinese medicine.
[0089] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. An automated guided robot system for material transport in a laboratory for screening active substances of traditional Chinese medicine, characterized in that: It includes a lidar module, an environmental camera module, a biomimetic sensing large model module, a large model ROS communication module, a ROS module, a drive chassis module, and an intelligent robotic arm module; The intelligent robotic arm module is used to collect image data of icon recognition points and grasp materials based on the image data. The intelligent robotic arm module includes a robotic arm unit, a target camera unit, a perception node, a grasping planning node, and a Moveit control node. The target camera unit is used to acquire image data of icon recognition points, use its built-in convolutional neural network model based on multi-dimensional feature fusion to locate the material storage location, and publish the location data message to the ROS module. The backbone of the convolutional neural network model consists of 4-stage hybrid dilated convolutional layers, and before the features are input into the backbone, attention weights are generated by a two-dimensional attention mechanism to process the original features. The lidar module is used to collect laboratory environmental data in real time, generate an obstacle grid map, and send distance data messages to the ROS module. The environmental camera module is used to acquire images of the laboratory environment in real time and send the image data messages to the biomimetic sensing large model module and the ROS module; The biomimetic perception large model module is used to perform visual analysis based on the image data messages to perceive the surrounding environment in real time and provide real-time descriptions of the surrounding environment. It obtains relevant environmental data and acquires operator instructions while responding to them. It is also used to send the relevant environmental data and the operator instructions to the large model ROS communication module. The large model ROS communication module is used to understand the relevant semantic information of the instructions given by the large model of different modalities, and after converting the understood semantic information into a standardized format, it aligns the semantic information in the standardized format with the ROS module interface according to the historical multi-round tasks, and sends the specific ROS system instructions to the ROS module. The ROS module is used to build relevant topics based on the location data message, distance data message, image data message, and specific ROS system instructions. The drive chassis module is used to subscribe to relevant topics in the ROS module and to drive the robot based on three-dimensional coordinates.
2. The automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine according to claim 1, characterized in that: The first layer of the four-stage hybrid dilated convolutional layer backbone of the convolutional neural network model consists of a dilated convolutional layer with a dilation rate of 2 and a filter size of 3×3. The second layer is a standard convolutional layer with a filter size of 3×3. The third layer is a normalization layer. The fourth layer is a ReLU activation function layer.
3. The automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine according to claim 1 or 2, characterized in that, The two-dimensional attention mechanism generates attention weights to process the original features, specifically including the following steps: The original features are first subjected to average pooling to preserve local average features; Adjust the channel using a standard convolution with a filter size of 1×1, then first perform a convolution operation in the vertical direction using a convolution with a filter size of 1×3, then perform a convolution operation in the horizontal direction, and finally adjust the channel again using a standard convolution with a filter size of 1×1. The formula for vertical convolution is as follows: In the above formula, This represents the image data after adjustment by the convolutional layer. The shape of the weights in the convolutional layer; Secondly, the formula for horizontal convolution is as follows: In the above formula, This represents the image data after adjustment following the vertical convolution operation. The shape of the weights in the convolutional layer; An attention map is generated using the Sigmoid function, and the attention map is multiplied by the original features before being fed into the backbone to process the original features.
4. The automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine according to claim 1, characterized in that: The biomimetic perception large model module includes a visual unit, an auditory unit, and a statement unit; The vision unit is used to perform real-time analysis of the image data messages based on the underlying vision big model. After encoding the image and extracting it into semantic vectors, it performs multimodal alignment to generate text language and obtain relevant environmental data. The auditory unit is used to understand operator instructions based on the underlying speech conversion model and convert operator speech instructions into text language. The statement unit is used to convert the text language generated by the visual unit into speech based on the underlying speech generation model, and to broadcast it in real time according to the operator's instructions.
5. The automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine according to claim 4, characterized in that, The specific steps for the visual unit to process image data messages include: The formula for obtaining the laboratory environment image from the image data message, i.e., the input image, is as follows: In the above formula, Image height, Image width, This refers to the number of RGB channels. The input image is further segmented into indivual each The size is The formula is as follows: ; Each Flattening it into a vector, the formula is as follows: In the above formula, ; Through linear embedding matrix E Mapped to d Dimensions, plus position encoding The formula is as follows: ; The resulting image token sequence is: ; Extract semantic vectors, Send in The multi-head attention mechanism layer processes data layer by layer, as shown in the following formula: ; In the above formula, , This indicates the attention of the bulls. Indicates a feedforward network; The final output visual feature sequence is: ; Introduction The query vectors are then subjected to multimodal alignment, as shown in the following formula: In the above formula, ; An attention mechanism is applied to each query vector and the visual feature sequence, as shown in the following formula: In the above formula, , , It is a visual feature sequence. The attention weights are calculated using the following formula: In the above formula, and Represents a learnable linear transformation matrix, where the dimension is 1. d × d ; get Language space cue vectors: ; Spatial cue vectors are concatenated to generate textual language, thus obtaining relevant environmental data.
6. The automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine according to claim 1, characterized in that: The large model ROS communication module includes a context encapsulation unit, a model call interface unit, a format standardization unit, and a memory unit. The context encapsulation unit is used to transform the semantics of instructions given by large models of different modalities into a unified JSON semantic context; The model call interface unit generates a model call request based on the JSON semantic context and converts the data into an understandable input format; The format standardization unit is used to standardize the structural format of the output of large models with different modalities and to ensure the alignment of ROS module instruction interfaces. Memory units are used to retain long-term contextual information across multiple historical task iterations, building a long-term memory bank and continuously updating it.
7. The automated guided robot system for transporting materials in the laboratory for screening active substances of traditional Chinese medicine according to claim 1, characterized in that: The ROS module includes a topic creation unit, a message transmission unit, and a visualization unit; The topic building unit is used to build related topics for multimodal data based on distance data messages, image data messages, positioning data messages, and specific ROS system commands; The message transmission unit is used to receive communication instructions and transmit specific data; The visualization unit is used to display the robot's location information in the laboratory in real time through an Rviz window.