A highly reliable method for drug identification based on the fusion of image, RFID and voice multi-source data

Through the fusion method of image, RFID and voice multi-data, high-accuracy recognition of drugs is achieved, the problem of insufficient recognition accuracy of traditional Chinese medicines is solved, the risk of error recognition is reduced, and the service robot's requirements for high-reliability recognition are met.

CN115081567BActive Publication Date: 2025-06-24SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210675263.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-06-24
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The accuracy rate of the prior art in drug identification cannot reach 99%, and the risk of giving users the wrong drugs is high, and it cannot meet the requirements of service robots for high reliability identification.

Method used

A high-reliability recognition method based on the fusion of image, RFID and voice multivariate data is adopted to initially identify target drugs through image processing, use RFID for the second recognition, and inform users of the drug name, usage and properties through voice synthesis, as the third guarantee.

Benefits of technology

It significantly improves the accuracy of drug identification, reduces the risk of error identification, and meets the requirements of service robots for high reliability identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115081567B_ABST
    Figure CN115081567B_ABST
Patent Text Reader

Abstract

The present invention discloses a highly reliable method for identifying drugs based on the fusion of image, RFID and voice multi-source data. The main content of this method is as follows: First, the user provides a voice command to the robot, and the robot obtains the drug name information through voice recognition and enters the image recognition stage. In the image recognition stage, after image enhancement and RGB extraction of the image in front of the robot, the position information of the corresponding drug is obtained, and then the position information is sent to the robot and enters the RFID recognition stage. In the RFID recognition stage, the robot grabs the drug according to the target drug position information and places the drug on the RFID reader for identification. After ensuring accuracy, it enters the drug delivery stage. In the drug delivery stage, while delivering the drug to the user, the robot uses voice synthesis to inform the user of the drug name, usage dosage and properties as a third layer of guarantee. Through the present invention, highly reliable automatic identification of drugs can be achieved, thereby helping the elderly and disabled with limited mobility to take their medicine on time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of service robots, and particularly relates to a highly reliable method for identifying drugs based on the fusion of image, RFID and voice multi-source data. Background Art

[0002] With the development of the concept of smart home, service robots have received increasing attention. Many elderly people and disabled people have problems with inconvenient legs and feet. Therefore, we hope that service robots can automatically identify drugs and help users get medicine. Drug identification is an indispensable technology in the process of service robots retrieving and placing drugs. Whether the robot can achieve highly accurate identification of drugs marks whether the service robot has reliable home service capabilities. The goal of the drug identification task is for the robot to automatically confirm whether a certain drug is the drug that the user currently needs to take. Specifically, it is to enable the robot to automatically distinguish drugs using a variety of information, accurately pick up the correct drug for the user and remind the user of the usage and dosage. The earlier solution in the field of drug identification for service robots was to use an image recognition algorithm based on deep learning to train a large number of drug pictures and then perform identification. In this method, the accuracy of the neural network in identifying objects basically cannot reach 99%. Assuming that the user takes medicine three times a day, the probability of not making a mistake within a month is about 40.5%. Obviously, this accuracy does not meet our requirements for the identification accuracy of service robots, and the risk of the user taking the wrong drug is unacceptable to us. In recent years, the development of deep learning technology has continuously improved the identification accuracy, but the identification accuracy of existing models still cannot meet our requirements for service robots. Summary of the Invention

[0003] Aiming at the above problems existing in the prior art, the present invention aims to provide a highly reliable method for identifying drugs based on the fusion of image, RFID and voice multi-source data. The method initially identifies the target drug through image processing and sends the position of the target drug to the robot. The robot grabs the drug according to the position information and places the drug in the RFID identification area for a second identification. When the results of the first two identifications are both accurate, the robot determines that the identification result is initially accurate, and uses voice synthesis to inform the user of the drug name, usage and dosage, and drug properties. The user makes a self-judgment based on the voice information as a third layer of guarantee.

[0004] To achieve the object of the present invention, the technical solution adopted by the present invention is as follows:

[0005] A highly reliable method for identifying drugs based on the fusion of image, RFID and voice multi-source data, the method comprising the following steps:

[0006] Step 1: Obtain the drug name through speech recognition. If the speech recognition fails, the robot will announce the request for the user to issue the command again.

[0007] Step 2: Obtain the position information of the target drug through image recognition. If the recognition fails, go to alternative step 1.

[0008] Step 3: Place the drug in the RFID recognition area for recognition. If no recognition result is obtained, go to alternative step 2. If the recognition fails, go to alternative step 1.

[0009] Step 4: The robot delivers the drug to the user.

[0010] Step 5: The robot announces the drug-related information and requests the user to judge whether the recognition is successful and provide a voice command. If the user judges that the recognition fails, the robot will put the drug back to its original place and return to step 2. Otherwise, the entire drug-taking process ends and the robot enters the standby state.

[0011] Alternative step 1: Turn on the active light source, use vision to obtain the positions of all drugs through feature tags. The robot grabs the medicine bottles one by one according to the drug positions until the RFID recognition result matches. Then turn off the active light source and go to step 4. If no result matches, it is determined that there is no target drug in the current environment. After making a voice announcement, the robot enters the standby state.

[0012] Alternative step 2: Announce that the RFID information cannot be obtained, ask the user to identify it in detail, and then go to step 4.

[0013] As an improvement of the present invention, in step 1, the speech recognition method is as follows:

[0014] In order to save computing resources and maximize the utilization of computing resources, the present invention uses an offline speech recognition method to work, and adopts a pre-trained model to perform speech feature matching to achieve the purpose of speech recognition.

[0015] In the training stage, the CPU is used to pre-collect the speech information of relevant drug names and instructions. After software noise reduction processing and feature extraction, the acoustic features of the corresponding speech information are stored in the Flash space. In the speech recognition stage, the speech information is collected in real time and subjected to noise reduction processing and feature extraction. Then the real-time speech-related features are compared with the pre-trained features to obtain the feature error coefficient. When the feature error coefficient is less than 254 for each character, it is considered that the relevant instruction or drug name has been recognized by speech.

[0016] Because the speech recognition part needs to cooperate with the image recognition link, in order to save storage space, in the pre-training part, the speech information is directly bound to the drug name and image features so that this link can communicate directly with the image recognition link.

[0017] As an improvement of the present invention, in the step 1, the method of speech synthesis and broadcast is as follows:

[0018] Initially, according to the drug information and relevant instruction information, the cloud computing platform is used to perform the calculation of speech synthesis, and the speech synthesized by the cloud computing platform is stored as a.wav format file for standby.

[0019] In the speech broadcast part, a separate sub-thread is used for work. The period of the sub-thread is 0.1 s. Once it is found in the sub-thread that the shared instruction variable changes, the relevant speech file is broadcast according to the instruction. To avoid conflicts between threads, timing waiting will be performed in the speech broadcast thread until the complete speech is played, and then the control quantity will be set to an empty instruction to allow other tasks to write new instructions.

[0020] At the same time, in other tasks, if the speech broadcast function needs to be used, the current shared instruction variable needs to be monitored first. When the instruction variable is in the state of no instruction, the instruction is written, otherwise it enters the waiting state until the instruction variable becomes in the state of no instruction and then the instruction is written.

[0021] As an improvement of the present invention, in the step 2, the method of image recognition is as follows:

[0022] First, the image is de-distorted, and then the Gray World algorithm is used to enhance the original image. In the processed image, the RGB values of each region are extracted by the method of spatio-temporal scale averaging and compared with the RGB feature range provided by the speech part. When more than 95% of the pixel points in a certain rectangular range are within the corresponding RGB feature range, it is determined that the target drug has been recognized.

[0023] As an improvement of the present invention, in the step 2, the method of position estimation is as follows:

[0024] To save hardware resources, the centroid and pixel number information of the target drug obtained by recognition in the de-distorted image by a monocular camera are used to confirm the direction information and distance information of the target drug relative to the robot according to the following formula:

[0025] sinθ aim =K θ (Col aim -Col0)

[0026]

[0027]

[0028] In the formula, θ aimθ is the horizontal angle information of the target drug relative to the direction directly in front of the robot, d is the linear distance information between the target drug and the robot, Col aim is the column number where the centroid of the detected target drug is located in the image, Col0 is the column number of the camera center pixel point, num pix is the number of pixels occupied by the drug, K θ and K d are the parameter conversion constants for the corresponding angle information and distance information respectively, and their specific values are determined by the actual sizes of the camera and the target drug.

[0029] The position information of the target drug relative to the robot is confirmed by the following formula according to the distance and direction information:

[0030] Pos x = d sinθ aim

[0031] Pos y = d cosθ aim

[0032] In the formula, Pos x and Pos y are the horizontal coordinate values of the drug position relative to the robot position, while the vertical coordinate value Pos z can be directly obtained from the relative height information between the robot and the drug.

[0033] Since the results of RFID identification and image recognition may be inconsistent, in the case of identification errors, the CPU will use alternative step 2 to obtain the coordinates of all drugs that the robot can reach. Using this method can avoid huge computational amounts in the case of accurate identification.

[0034] As an improvement of the present invention, in the said step 3, the method for the robot to perform RFID identification on drugs is:

[0035] First, the robot calculates the target angles of each joint from the coordinates of the target drug, grabs the drug and places it in the RFID identification area. The selected robot structure satisfies the Pieper criterion, and the following formula is used for inverse kinematics solution by the analytical method:

[0036] θ1 = θ aim

[0037] Trans x = l2 + l3sinθ2 + l4sin(θ2 + θ3) + l5sinθ tool

[0038] Trans y = l1 + l3cosθ2 + l4cos(θ2 + θ3) + l5cosθ tool

[0039] θ4 = θ tool -θ2 - θ3

[0040] Where θ1, θ2, θ3, θ4 are target angle values, and θ tool is the angle between the Z-axis direction of the end effector coordinate system we require and the world coordinate system. l1, l2, l3, l4, l5 are the lengths of the robot links, and Trans x and Trans y are respectively the horizontal linear distance and the vertical linear distance of the medicine relative to the robot, and are obtained using the following formula:

[0041]

[0042] Trans y = Pos z

[0043] When the medicine is in the RFID identification area, an independent RFID reader is used to communicate with the robot through a TTL serial port to obtain relevant information about the medicine (including name, usage dosage, and properties). The robot compares the target medicine information received by the CPU and the target medicine information identified by the RFID to confirm whether the currently grasped medicine is the same as the target medicine, and sends the recognition result to the CPU for further processing.

[0044] As an improvement of the present invention, in step 4, the method for the robot to take the medicine to the user's side is as follows:

[0045] Initially, a lidar and the gmapping algorithm are used in cooperation with a PC to build a map of the indoor environment, and several regular user positions such as the edge of the bed and the edge of the chair are marked on the map. At the same time, the numbers of each feature position are stored.

[0046] After the robot grabs the medicine, the user provides a voice command to the robot to confirm which feature position the medicine needs to be delivered to. Then the robot uses the lidar information and the map information to confirm its own pose, and then uses the A* algorithm for path planning according to the marked position, and controls its own movement according to the planned path. The robot periodically judges its current position during the movement to achieve the purpose of path closed-loop control. After moving to the target position, the robot holds the medicine above its head to ensure that the user can get the medicine.

[0047] The method for extracting feature points from the map information obtained by the lidar is:

[0048]

[0049]

[0050] Presult = |P R - P r |

[0051] In the above formula, if P result is greater than the set threshold P t then it is determined that this point is one of the feature points and can be used for feature matching.

[0052] As an improvement of the present invention, in step 5, the method for the robot to proofread the voice is as follows:

[0053] The robot determines whether some command variables in the voice broadcast part are empty. If not, it waits. After the waiting ends, it modifies the command variables to the current drug name, usage and dosage, and properties for broadcast, turns on the voice recognition to obtain the user's feedback, and determines whether the drug is correct. If the user feedbacks that the drug is incorrect, it returns to step 2. Otherwise, it determines that the task is completed and enters the standby state.

[0054] As an improvement of the present invention, in the alternative step 1, the method for obtaining the drug position through the feature label is as follows:

[0055] The present invention uses a holographic reflective material as the feature label of the drug. Under the condition that the robot actively illuminates, the RGB value of the feature label under the camera's view will be much larger than that of other objects. Using this as the threshold for binary segmentation can ensure that the robot can quickly and accurately find the target drug.

[0056] In the recognition task, first turn on the active light source at the camera position for illumination. The vision part performs binary image processing according to the lower limit of the RGB value of the feature label, then uses the DBSACN method to remove noise points and distinguish different pixel point clusters. Each cluster corresponds to a medicine bottle. Then, according to the image information of different medicine bottles, the following formula is used for position estimation:

[0057]

[0058]

[0059] In the formula, Position_medicine x is the coordinate of the estimated drug position relative to the X-axis of the robot coordinate system, Position_medicine y is the coordinate of the estimated drug position relative to the Y-axis of the robot coordinate system, K distance is the distance estimation parameter, which is determined by the size parameters of the camera and the holographic feature label, num pix_medicine is the number of pixel points in the cluster to be estimated, K rot is the angle estimation parameter, which is determined by the camera, Col medicineThe coordinates of the estimated cluster's centroid position, Col middle It is the coordinate value of the center column of the entire screen.

[0060] After the drug position estimation is completed, the estimated positions are sent to the robot one by one in a queue for grasping and RFID identification.

[0061] As an improvement of the present invention, in the alternative step 2, the method for obtaining the location of the drug through the feature tag is:

[0062] It is determined that the current drug does not have an RFID tag or the tag is damaged, and a layer of information has been lost during the identification process. The probability of identification error is relatively large, and voice notification is required to inform the user to carefully identify the drug.

[0063] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects: In view of the problem of drug identification, the present invention proposes a highly reliable drug identification method based on image, RFID and voice multivariate data fusion. The image recognition part uses a method based on a monocular camera to estimate the position of the target object based on size information and image centroid information, so that the distance estimation method requires very few computing resources and hardware resources. RFID identification uses an independent RFID identifier to communicate with the robot in the form of TTL serial communication, so that the accuracy of drug identification by this method is significantly improved compared with the simple image recognition method. In terms of robot movement and positioning, we use a PC-assisted mapping method to ensure that the map information obtained by the robot is very accurate, and at the same time use a 2D point cloud feature point extraction method to save the time required for robot posture perception and enhance the real-time and reliability of robot movement. In the voice information part, we use cloud computing to synthesize drug-related property information to assist users in autonomously identifying drugs, which as the third information ensures the accuracy of drug identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is a flow chart of a highly reliable drug identification method based on image, RFID and voice multivariate data fusion of the present invention;

[0065] Figure 2 It is a flow chart of the speech recognition method adopted by the present invention;

[0066] Figure 3 It is a flow chart of the voice broadcast method adopted by the present invention;

[0067] Figure 4 It is a flow chart of the method of the robot used in the present invention to bring medicines to the user;

[0068] Figure 5 It is a flow chart of the speech proofreading method adopted by the present invention;

[0069] Figure 6 It is the flowchart of the method for obtaining the location of drugs through feature tags adopted by the present invention; Specific embodiments

[0070] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0071] As Figure 1 shown, the present invention proposes a highly reliable drug identification method based on the fusion of image, RFID and voice multi-source data. The detailed steps of this method are as follows:

[0072] (1) Voice recognition:

[0073] As Figure 2 shown, in order to save computing resources and maximize the utilization of computing resources, the present invention uses an offline voice recognition method to work, and adopts a pre-trained model to perform voice feature matching to achieve the purpose of voice recognition.

[0074] In the training stage, the CPU is used to pre-collect the voice information of relevant drug names and instructions. After software noise reduction processing and feature extraction, the acoustic features of the corresponding voice information are stored in the Flash space; in the voice recognition stage, the voice information is collected in real time and subjected to noise reduction processing and feature extraction, and then the real-time voice-related features are compared with the pre-trained features to obtain the feature error coefficient. When the feature error coefficient is less than 254 for each word, it is considered that the relevant instruction or drug name has been recognized by voice.

[0075] Because the voice recognition part needs to cooperate with the image recognition link, in order to save storage space, in the pre-training part, the voice information is directly bound to the drug name and image features so that this link can communicate directly with the image recognition link.

[0076] If a voice recognition error occurs, the method of voice broadcast is used to prompt the user to resend the voice instruction and restart the voice recognition.

[0077] (2) Voice synthesis and voice broadcast:

[0078] As Figure 3 shown, initially, the voice synthesis calculation is performed using the cloud computing platform according to the drug information and relevant instruction information, and the voice synthesized by the cloud computing platform is stored as a.wav format file for later use.

[0079] The voice broadcast part works using a separate sub-thread. The period of the sub-thread is 0.1 s. Once it is found in the sub-thread that the shared command variable has changed, relevant voice files will be broadcast according to the instructions. To avoid conflicts between threads, timing waits will be performed in the voice broadcast thread until the complete voice is played, and then the control quantity will be set to an empty instruction, allowing other threads to write instructions.

[0080] If the voice broadcast function needs to be used in other steps, the current broadcast situation will be judged, and instructions can be written to call the voice broadcast function only after the broadcast state becomes an empty instruction.

[0081] (3) Image recognition and distance estimation:

[0082] First, the image is de-distorted, and then the Gray World algorithm is used to enhance the original image. In the processed image, the RGB values of each region are extracted using the spatio-temporal scale averaging method and compared with the RGB feature range provided by the voice part. When more than 95% of the pixel points within a certain rectangular range are within the corresponding RGB feature range, it is determined that the target drug has been recognized.

[0083] To save hardware resources, the centroid and pixel number information of the target drug obtained by recognition in the de-distorted image by a monocular camera are used to confirm the direction information and distance information of the target drug relative to the robot according to the following formula:

[0084] sinθ aim =K θ (Col aim -Col0)

[0085]

[0086]

[0087] In the formula, θ aim is the horizontal angle information of the target drug relative to the front direction of the robot, d is the linear distance information between the target drug and the robot, Col aim is the column number where the centroid of the detected target drug is located in the image, Col0 is the column number of the camera center pixel point, num pix is the number of pixel points occupied by the drug, K θ and K d are the parameter conversion constants for the corresponding angle information and distance information respectively, and their specific values are determined by the camera and the target drug.

[0088] According to the distance and direction information, the position information of the target drug relative to the robot is confirmed by the following formula:

[0089] Pos x = d sinθ aim

[0090] Pos y = d cosθ aim

[0091] Where Pos x and Pos y are the coordinate values on the X-axis and Y-axis of the drug position relative to the robot position, and the vertical coordinate value Pos z can be directly obtained from the relative height information between the robot and the drug.

[0092] Since the results of RFID identification and image recognition may be inconsistent, in the case of recognition errors, the CPU will use alternative step 2 to obtain the coordinates of all drugs that the robot can reach. Using this method can avoid huge computational amounts in the case of accurate recognition.

[0093] Affected by the lighting conditions, image recognition failures may occur. Limited by cost, providing an independent light source for the drug position will not only increase costs but also cause problems such as light pollution.

[0094] As Figure 6 shown, for this problem, the present invention uses a holographic reflective material as the characteristic label of the drug. Under the condition of the robot actively lighting, the RGB value of the characteristic label under the camera's view will be much larger than that of other objects. Using this as a threshold for binary segmentation can ensure that the robot can quickly and accurately find the target drug.

[0095] In the recognition task, first turn on the active light source at the camera position for lighting. The vision part performs image binary processing according to the lower limit of the RGB value of the characteristic label, then uses the DBSACN method to remove noise points and distinguish different pixel point clusters. Each cluster corresponds to a medicine bottle, and then the following formula is used for position estimation according to the picture information of different medicine bottles:

[0096]

[0097]

[0098] Where Position_medicine x is the coordinate of the estimated drug position relative to the X-axis of the robot coordinate system, Position_medicine y is the coordinate of the estimated drug position relative to the Y-axis of the robot coordinate system, K distance is the distance estimation parameter, which is determined by the size parameters of the camera and the holographic characteristic label, num pix_medicineK is the number of pixels in the cluster to be estimated. rot Col is the angle estimation parameter, which is determined by the camera. medicine Col is the column coordinate value of the centroid position of the cluster to be estimated. middle is the coordinate value of the center column of the entire picture.

[0099] After completing the estimation of the drug position, the estimated positions are sent to the robot one by one in the form of a queue for grasping and RFID identification.

[0100] (4) The robot performs RFID identification on the drug:

[0101] In the case of obtaining the target drug through image recognition, this step is entered for RFID identification. First, the robot calculates the target angles of each joint from the coordinates of the target drug, grasps the drug and places it in the RFID identification area. The selected robot structure satisfies the Pieper criterion, and the following formula is used for inverse kinematics solution by the analytical method:

[0102] θ1 = θ aim

[0103] Trans x = l2 + l3sinθ2 + l4sin(θ2 + θ3) + l5sinθ tool

[0104] Trans y = l1 + l3cosθ2 + l4cos(θ2 + θ3) + l5cosθ tool

[0105] θ4 = θ tool - θ2 - θ3

[0106] In the formula, θ1, θ2, θ3, θ4 are the target angle values, θ tool is the angle between the Z-axis direction of the end effector coordinate system and the world coordinate system that we require, l1, l2, l3, l4, l5 are the lengths of the robot links, Trans x and Trans y are respectively the horizontal linear distance and the vertical linear distance of the drug relative to the robot, and are obtained using the following formula:

[0107]

[0108] Trans y = Pos z

[0109] When the drug is in the RFID recognition area, an independent RFID reader is used to communicate with the robot through the TTL serial port to obtain drug-related information (including name, dosage method, and properties). The robot compares the target drug information received by the CPU with the target drug information recognized by the RFID to confirm whether the currently grasped drug is the same as the target drug, and sends the recognition result to the CPU for further processing.

[0110] If the RFID fails to recognize the drug, it is determined that the drug RFID is damaged or lost. One piece of information has been lost during the recognition process, and the probability of recognition error is relatively large. The user needs to be informed by voice to distinguish carefully.

[0111] If the recognized RFID information is inconsistent with the target RFID recognition information, it is considered that there is an error in the result of the image processing part. As Figure 6 shown, first, turn on the active light source at the camera position for lighting. After the visual part performs binary image processing based on the lower limit of the RGB value of the feature label, the DBSACN method is used to remove noise points and distinguish different pixel point clusters. Each cluster corresponds to a medicine bottle. Then, the position estimation is performed using the following formula based on the image information of different medicine bottles:

[0112]

[0113]

[0114] In the formula, Position_medicine x is the coordinate of the estimated drug position relative to the X-axis of the robot coordinate system, Position_medicine y is the coordinate of the estimated drug position relative to the Y-axis of the robot coordinate system, K distance is the distance estimation parameter, which is determined by the size parameters of the camera and the holographic feature label, num pix_medicine is the number of pixel points in the estimated cluster, K rot is the angle estimation parameter, which is determined by the camera, Col medicine is the column coordinate value of the centroid position of the estimated cluster, Col middle is the coordinate value of the center column of the entire picture.

[0115] After completing the drug position estimation, the estimated positions are sent to the robot one by one in the form of a queue for grasping and RFID recognition.

[0116] (5) The robot takes the drug to the user's side:

[0117] Initially, a lidar and the gmapping algorithm are used in conjunction with a PC to construct a map of the indoor environment, and several regular user positions, such as beside the bed and beside the chair, are marked on the map as characteristic positions. Meanwhile, the numbers of each characteristic position are stored.

[0118] After the robot grabs the medicine, the user provides a voice command to the robot to confirm which characteristic position the medicine needs to be delivered to. Then the robot uses the lidar information and the map information to confirm its own pose, and then uses the A* algorithm for path planning based on the marked positions, and controls its own movement according to the planned path. During the movement, the robot periodically judges its current position to achieve the purpose of path closed-loop control. After moving to the target position, the robot holds the medicine above its head to ensure that the user can get the medicine.

[0119] The method for extracting characteristic points from the map information obtained by the lidar is as follows:

[0120]

[0121]

[0122] P result =|P R -P r |

[0123] In the above formula, if P result is greater than the set threshold P t then it is determined that this point is one of the characteristic points and can be used for feature matching.

[0124] (6) Perform voice calibration:

[0125] As Figure 5 shown, the robot determines whether the command variables in the voice broadcast part are empty. If not, it waits. After the waiting ends, the command variables are modified to the current medicine name, usage and dosage, and properties broadcast, and voice recognition is enabled to obtain the user's feedback. If the user determines that the recognition is incorrect, it returns to step 2. Otherwise, it determines that the task is completed and enters the standby state.

[0126] It should be noted that the above embodiments are only preferred embodiments of the present invention and do not limit the protection scope of the present invention. Any equivalent replacement or substitution made on the basis of the above technical solutions belongs to the protection scope of the present invention.

Claims

1. A highly reliable drug identification method based on the fusion of image, RFID, and voice multi-source data, characterized in that, The method includes the following steps: Step 1: Obtain the drug name through speech recognition. If the speech recognition fails, the system will announce a voice prompt asking the user to issue the command again. Step 2: Obtain the position information of the target drug through image recognition. If the recognition fails, enter alternative step 1. Step 3: Place the drug in the RFID recognition area for recognition. If no recognition result is obtained, enter alternative step 2. If the recognition fails, enter alternative step 1. Step 4: The robot delivers the drug to the user. Step 5: The robot announces the drug-related information through voice, and requests the user to judge whether the recognition is successful and provide a voice command. If the user judges that the recognition fails, the robot will put the drug back to its original place and return to step 2. Otherwise, the entire drug-taking process ends and the robot enters the standby state. Alternative step 1: Turn on the active light source, use vision to obtain the positions of all drugs through feature tags. The robot grabs the medicine bottles one by one according to the drug positions until the RFID recognition result matches. Then turn off the active light source and enter step 4. If no match is found, it is determined that there is no target drug in the current environment. After a voice announcement, enter the standby state. Alternative step 2: Announce through voice that the RFID information cannot be obtained, and ask the user to identify it in detail, then enter step 4. Among them, in step 2, the method for position estimation is as follows: Based on the centroid and pixel quantity information of the target drug in the undistorted image obtained by recognition, the direction information and distance information of the target drug relative to the robot are confirmed according to the following formula: sinθ aim = K θ (Col aim - Col0) where θ aim is the horizontal angle information of the target drug relative to the direction directly in front of the robot, d is the linear distance information between the target drug and the robot, Col aim is the column number where the centroid of the detected target drug is located in the image, Col0 is the column number of the camera center pixel point, num pix is the number of pixels occupied by the drug, K θ and K d are respectively the parameter conversion constants corresponding to the angle information and the distance information, and their specific values are determined by the actual sizes of the camera and the target drug; Based on the distance and direction information, the position information of the target drug relative to the robot is confirmed according to the following formula: Pos x = d sinθ aim Pos y = d cosθ aim where Pos x and Pos y are the horizontal coordinate values of the drug position relative to the robot position, and the vertical coordinate value Pos z can be directly obtained from the relative height information between the robot and the drug; In step 3, the method for the robot to perform RFID recognition on the drug is as follows: First, the robot calculates the target angles of each joint from the coordinates of the target drug, grabs the drug and places it in the RFID recognition area. The selected robot structure satisfies the Pieper criterion, and the following formula is used to solve the inverse kinematics analytically to obtain the target angles of each joint: θ1 = θ aim Trans x = l2 + l3sinθ2 + l4sin(θ2 + θ3) + l5sinθ tool Trans y = l1 + l3cosθ2 + l4cos(θ2 + θ3) + l5cosθ tool θ4 = θ tool -θ2 - θ3 where θ1, θ2, θ3, θ4 are target angle values, and θ tool is the angle between the Z-axis direction of the end effector coordinate system we require and the world coordinate system. l1, l2, l3, l4, l5 are the lengths of the robot links. Trans x and Trans y are the horizontal and vertical linear distances of the medicine relative to the robot, respectively, and are obtained using the following formula: Trans y = Pos z When the drug is in the RFID recognition area, use an independent RFID reader to communicate with the robot through a TTL serial port to obtain drug-related information, including name, usage amount, and properties. The robot compares the target drug information received by the CPU with the target drug information recognized by the RFID to confirm whether the currently grabbed drug is the same as the target drug, and sends the recognition result to the CPU for the next step of processing. In alternative step 1, the method for obtaining the drug position through image recognition of feature tags is as follows: A holographic reflective material is used as the feature tag of the drug. Under the condition that the robot actively lights up, the RGB value of the feature tag in the camera view will be much larger than that of other objects. Using this as a threshold for binary segmentation can ensure that the robot can quickly and accurately find the target drug. In the recognition task, first turn on the active light source at the camera position for lighting. The vision part performs binary image processing on the image according to the lower limit of the RGB value of the feature tag, and then uses the DBSACN method to remove noise points and distinguish different pixel clusters. Each cluster corresponds to a medicine bottle. Then, based on the image information of different medicine bottles, the following formula is used for position estimation: Where Position_medicine x is the coordinate of the estimated medicine position relative to the X-axis of the robot coordinate system, Position_medicine y is the coordinate of the estimated medicine position relative to the Y-axis of the robot coordinate system, K distance is the distance estimation parameter, which is determined by the size parameters of the camera and the holographic feature label, num pix_medicine is the number of pixel points in the estimated cluster, K rot is the angle estimation parameter, which is determined by the camera, Col medicine is the column coordinate value of the centroid position of the estimated cluster, Col middle is the coordinate value of the center column of the entire picture; After estimating the position of the medicine, the estimated positions are sent to the robot one by one in the form of a queue for grasping and RFID identification.

2. The highly reliable drug identification method based on the fusion of image, RFID and voice multi-source data according to claim 1, wherein: In step 1, the method of speech recognition is as follows: In the training stage, the CPU is used to pre-collect the speech information of relevant medicine names and instructions. After software noise reduction processing and feature extraction, the acoustic features of the corresponding speech information are stored in the Flash space. In the speech recognition stage, the speech information is collected in real time and subjected to noise reduction processing and feature extraction. Then, the real-time speech-related features are compared with the pre-trained features to obtain the feature error coefficient. When the feature error coefficient is less than 254 for each character, it is considered that the relevant instruction or medicine name has been recognized by speech.

3. A highly reliable drug identification method based on the fusion of image, RFID, and voice multi-source data according to claim 1, characterized in that: In step 1, the method of speech synthesis and broadcasting is as follows: Initially, according to the medicine information and relevant instruction information, the cloud computing platform is used to calculate the speech synthesis and the synthesized speech of the cloud computing platform is stored as a.wav format file for later use. In the speech broadcasting part, a separate sub-thread is used to work. The period of the sub-thread is 0.1 s. Once it is found in the sub-thread that the shared instruction variable changes, the relevant speech file is broadcast according to the instruction. After the corresponding speech file is broadcast, the instruction variable is set to the no-instruction state, allowing other tasks to write new instructions. At the same time, in other tasks, if the speech broadcasting function needs to be used, the current shared instruction variable needs to be monitored first. When the instruction variable is in the no-instruction state, the instruction is written. Otherwise, it enters the waiting state until the instruction variable becomes in the no-instruction state and then the instruction is written.

4. A highly reliable drug identification method based on the fusion of image, RFID and voice multi-source data according to claim 1, characterized in that: In step 2, the method of image recognition is as follows: First, the image is de-distorted, and then the Gray World algorithm is used to enhance the original image. In the processed image, the RGB values of each region are extracted by the method of spatio-temporal scale averaging and compared with the RGB feature range provided by the speech part. When more than 95% of the pixel points in a certain rectangular range are within the corresponding RGB feature range, it is determined that the target medicine has been recognized.

5. A highly reliable drug identification method based on the fusion of image, RFID and voice multi-source data according to claim 1, characterized in that: In step 4, the method for the robot to take the medicine to the user's side is as follows: Initially, the lidar and the gmapping algorithm are used in cooperation with the PC to build a map of the indoor environment, and several regular user positions such as the edge of the bed and the edge of the chair are marked in the map. At the same time, the numbers of each feature position are stored. After the robot grabs the medicine, the user provides a voice instruction to the robot to confirm which feature position the medicine needs to be delivered to. Then the robot uses the lidar information and the map information to confirm its own pose, and then uses the A* algorithm for path planning according to the marked position, and controls its own movement according to the planned path. During the movement, the robot periodically judges its current position to achieve the purpose of path closed-loop control. After moving to the target position, the robot holds the medicine above its head to ensure that the user can get the medicine. The method for extracting the feature points of the map information obtained by the lidar is as follows: P result = |P R -P r | In the above formula, if P result is greater than the set threshold P t then it is determined that this point is one of the feature points for feature matching.

6. A highly reliable drug identification method based on the fusion of image, RFID and voice multi-source data according to claim 1, characterized in that: In step 5, the method for the robot to proofread the speech is as follows: The robot determines whether the command variables of the voice broadcast part are empty. If not, it waits. After waiting, it modifies the command variables to the current drug name, usage, dosage and properties. It turns on voice recognition to obtain user feedback and judge whether the drug is correct. If the user feedback is incorrect, it returns to step 2. Otherwise, it determines that the task is over and enters the standby state.

7. A highly reliable drug identification method based on the fusion of image, RFID and voice multi-source data according to claim 1, characterized in that: In the alternative step 2, the method for providing feedback on the information loss situation is: It is determined that the current drug does not have an RFID tag or the tag is damaged, and a layer of information has been lost during the identification process. The probability of identification error is relatively large, and voice notification is required to inform the user to carefully identify the drug.

Citation Information

Patent Citations

  • Service robot control platform system and multimode intelligent interaction and intelligent behavior realizing method thereof

    CN102323817A

  • Intelligent voice type guide blind robot control method and control system thereof

    CN110109457A