Hospital intelligent consumable management cabinet system based on multi-modal visual large model

The intelligent consumables management cabinet system based on multimodal visual large model solves the problem of inaccurate identification of consumable types and quantities in traditional manual management, achieving efficient and accurate consumables management and improving the hospital's management efficiency.

CN120998449AActive Publication Date: 2025-11-21SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

Patent Information

Application Number
CN202511517288.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Traditional medical consumables management methods rely on manual verification, which leads to problems of misidentification of type and quantity. In particular, the high similarity of the outer packaging of different medical consumables results in insufficient accuracy of image recognition technology.

Method used

The hospital intelligent consumables management cabinet system, based on a multimodal visual large model, acquires video data and weight change values ​​through a data acquisition module, extracts frames and optical flow information through an image extraction module, divides the object's motion trajectory, and combines the type and quantity recognition modules for identification. The management module then updates the inventory.

Benefits of technology

It improves the accuracy of consumable management and the efficiency of information updates, and can accurately identify the types and quantities of items put in and taken out, which greatly improves the efficiency of hospital consumable management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998449A_ABST
    Figure CN120998449A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of consumable management, in particular to a hospital intelligent consumable management cabinet system based on a multi-modal visual large model, which collects video data and weight change values of a management cabinet from door opening to door closing and screens out invalid video data through the duration of the video data. Therefore, the recognition accuracy in the application is improved. Besides, the optical flow information in the video data is extracted, and the optical flow information is utilized to identify the motion trail of the object in the video, so that the whole data is decomposed into a putting-in stage image frame when the hand of the user enters and a taking-out stage image frame when the hand of the user is extracted; and identifying the image frames of the two or more stages to obtain the type of the put-in article and the type of the taken-out article. In combination with the motion trail and the weight change value, the types and the number of the articles put in by the user and the types and the number of the articles taken out by the user can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of consumables management technology, specifically a hospital intelligent consumables management cabinet system based on a multimodal visual large model. Background Technology

[0002] Currently, medical consumables are widely used in medical institutions, and their management is an important part of hospital supplies management. Traditional medical consumable management mainly relies on manual verification. For example, medical staff or managers record the receipt, issuance, inventory quantity, and usage of consumables using paper registers or spreadsheets (such as Excel).

[0003] Although manual verification is still used in some primary care or small medical institutions, with the rapid increase in the types and usage of medical consumables, and the increasing requirements for medical quality and safety management, this management method has gradually revealed many drawbacks. Therefore, existing technologies utilize image recognition technology to identify consumables in stock. However, for medical consumables, the outer packaging of different consumables often has a high degree of similarity, so simply using image recognition patterns can easily lead to misidentification of type and quantity. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a hospital intelligent consumables management cabinet system based on a multimodal visual large model, so as to solve the problems in the background art.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] The intelligent hospital consumables management cabinet system based on a multimodal visual large model of the present invention includes:

[0007] The data acquisition module is used to start timing, collect the initial weight of multiple storage units and video when the cabinet door of the consumable management cabinet is opened, and stop timing, video acquisition and the final weight of multiple storage units when the cabinet door of the consumable management cabinet is closed, so as to obtain the duration of a consumable collection cycle, video data and weight change values ​​of multiple storage units; and to take the video data of consumable collection cycles with a duration greater than a preset duration threshold as valid video data.

[0008] The image extraction module is used to extract frames from the effective video data to obtain multiple keyframes;

[0009] A segmentation module is used to extract optical flow information from multiple key frames and construct the motion trajectory of an object within the key frames using the optical flow information; and to divide the multiple key frames into insertion stage image frames and extraction stage image frames based on the object motion trajectory.

[0010] The category recognition module is used to recognize the image frames of the insertion stage and the image frames of the removal stage respectively, and obtain the category of the inserted item, the initial quantity of the inserted item, the category of the removed item, and the initial quantity of the removed item.

[0011] The quantity recognition module is used to verify the initial quantity of the put-in items and the initial quantity of the taken-out items based on the weight change value of the multiple storage units, the motion trajectory of the object in the key frame, the type of put-in items, and the type of taken-out items, so as to obtain the quantity of put-in items and the quantity of taken-out items.

[0012] The management module is used to construct cabinet inventory change information based on the type of items put in, the quantity of items put in, the type of items taken out, and the quantity of items taken out, and to manage consumables based on the cabinet inventory change information.

[0013] In one embodiment of this application, frames are extracted from the valid video data to obtain multiple keyframes, including:

[0014] The valid video data is subjected to interval frame extraction to obtain multiple initial frames;

[0015] Multiple initial frames are converted to grayscale to obtain multiple grayscale image frames;

[0016] Perform a difference operation on adjacent grayscale image frames to obtain a difference image; then perform binarization on the difference image to obtain a binarized image;

[0017] Extract the motion region from the binarized image, and based on the proportion of the motion region, when the proportion of the motion region is greater than a set threshold, determine that the next frame in the adjacent grayscale image frame is a dynamic frame, and use the dynamic frame as a key frame.

[0018] In one embodiment of this application, optical flow information from multiple keyframes is extracted, and the optical flow information is used to construct the motion trajectory of an object within the keyframes, including:

[0019] Extract key feature points from the keyframe, wherein the key feature points include hand joint feature points and corner points;

[0020] The optical flow field in the key frame is calculated based on the key feature points to obtain the motion vector of each key feature point. ,in, This represents the displacement of the key feature point along the x-axis. This indicates the displacement of the key feature point along the y-axis;

[0021] Based on motion vectors Density clustering is performed on the key feature points of each keyframe to obtain one or more key feature point clusters;

[0022] Based on all motion vectors in the key feature point cluster Calculate the average value to obtain the average vector. And calculate the average coordinates of the key feature points in the key feature point cluster. ,in, This represents the average displacement along the x-axis. This represents the average displacement along the y-axis. The average coordinates of the x-axis. This represents the average coordinate of the y-axis;

[0023] The average vector is converted from a biaxial displacement representation to a representation of direction angle and uniaxial displacement, and then combined with the average coordinates. Obtain the target vector ,in:

[0024]

[0025]

[0026] In the formula, It is the direction angle. It is the displacement;

[0027] Target vector based on multiple keyframes Construct the motion trajectories of one or more objects.

[0028] In one embodiment of this application, the target vector is based on multiple keyframes. Constructing the motion trajectories of one or more objects, including:

[0029] Obtain a background image with the same viewpoint as the video data;

[0030] The target vector of the multiple key frames The motion feature image is obtained by mapping it onto the background image;

[0031] In the motion feature image, with each target vector starting position Define a radius of for the origin. The matching range, where, , This is an empirical proportional parameter;

[0032] Based on the matching range, any two adjacent target vectors are matched, wherein when the end of one of the target vectors falls within the matching range of the adjacent target vector, the two adjacent target vectors are determined to be matched.

[0033] By fitting the starting points of all matching target vectors, the trajectory of the object is obtained.

[0034] In one embodiment of this application, the plurality of keyframes are divided into placement stage image frames and removal stage image frames based on the object's motion trajectory, including:

[0035] Calculate the distance between the starting points of any two adjacent target vectors in the trajectory of the object. And based on the time interval between adjacent keyframes Calculate the velocity of an object moving along a multi-segment trajectory , ;

[0036] Based on the orientation angle of each keyframe The speed of an object and collection time point Construct the feature vector of the stage to which each keyframe belongs. ;

[0037] Construct a timeline and combine the feature vectors of the stages to which multiple keyframes belong. Mapped to the time axis;

[0038] Based on a pre-built sliding window, the sliding window slides along the time axis, and at each slide, the average direction angle and average speed within the sliding window are calculated. When the average direction angle falls within a preset first angle range and the average speed is greater than or equal to a preset speed threshold, the segment within the sliding window is determined to be an insertion segment. When the average direction angle falls within a preset second angle range and the average speed is greater than or equal to a preset speed threshold, the segment within the sliding window is determined to be an extraction segment. When the average speed is less than a preset speed threshold, the segment within the sliding window is determined to be an operation segment.

[0039] The consecutive insertion and removal segments are merged to obtain the time interval of the insertion stage and the time interval of the removal stage. The insertion stage image frames are extracted from the time interval of the insertion stage and the removal stage image frames are extracted from the time interval of the removal stage.

[0040] In one embodiment of this application, the image frames of the insertion stage and the image frames of the removal stage are identified respectively to obtain the type of item inserted, the initial quantity of items inserted, the type of item removed, and the initial quantity of items removed, including:

[0041] Remove the non-moving background regions from the input and output image frames to obtain the input image;

[0042] The input image is input into a pre-built recognition model to obtain the type of items put in, the initial quantity of items put in, the type of items taken out, and the initial quantity of items taken out.

[0043] In one embodiment of this application, the method for constructing the recognition model includes:

[0044] Acquire a sample image and perform expansion processing on the sample image to obtain an expanded sample. The expansion processing includes image scaling, rotation, cropping, low-light enhancement, ghosting removal, and dynamic de-glare processing.

[0045] The sample images are then labeled to obtain training samples;

[0046] The artificial neural network is trained by backpropagation based on the training samples to obtain the recognition model.

[0047] In one embodiment of this application, the initial quantity of items placed and the initial quantity of items taken out are verified based on the weight change values ​​of the plurality of storage units, the object motion trajectory within the keyframe, the type of items placed, and the type of items taken out, to obtain the quantity of items placed and the quantity of items taken out, including:

[0048] Obtain a pre-built consumable weight data table, wherein the consumable weight data table includes weight references for various consumables;

[0049] The target storage unit that produces the weight change is determined based on the object's motion trajectory within the keyframe.

[0050] Extract the weight change value of the target storage unit from the weight change values ​​of the plurality of storage units. ;

[0051] Obtain the reference weight of the type of item put in or taken out from the consumable weight data table. When the reference weight is less than the minimum weight threshold, use the initial quantity of the put-in items and the initial quantity of the taken-out items as the put-in items quantity and the taken-out items quantity, respectively. When the reference weight is greater than or equal to the minimum weight threshold, calculate a reference change value based on the initial quantity of the put-in items and the reference weight of the put-in item type. Alternatively, a reference change value can be calculated based on the initial quantity of items taken out and the reference weight of the types of items taken out. ;

[0052] Calculate the weight change value Compared with reference change value deviation rate , ;

[0053] The deviation rate When the deviation rate is less than or equal to a preset deviation rate threshold, the initial quantity of items put in and the initial quantity of items taken out are taken as the quantity of items put in and the quantity of items taken out, respectively; in the deviation rate When the deviation rate exceeds a preset threshold, the weight change value is calculated. The ratio of the reference weight to the target item type The target item type is either the type of item to be placed or the type of item to be taken out.

[0054] The ratio If the non-integer part is less than the target value, then the ratio will be... As an estimated quantity, the estimated quantity and the valid video data are sent to the target object. After receiving verification information from the target object, the quantity of items put in and the quantity of items taken out are obtained; in the ratio When the non-integer part is greater than or equal to the target value, an error message and the valid video data are sent to the target object. After receiving the verification information from the target object, the number of items put in and the number of items taken out are obtained.

[0055] In one embodiment of this application, it further includes:

[0056] The identity verification module is used to verify the identity information received from the user, and to unlock the consumables management cabinet when the verification is successful.

[0057] The beneficial effects of this invention are as follows: The hospital intelligent consumables management cabinet system based on a multimodal visual large model of this invention improves the recognition accuracy by collecting video data and weight change values ​​from the opening to closing of the cabinet, and filtering out invalid video data based on the duration of the video data. Furthermore, this application extracts optical flow information from the video data and uses this information to identify the motion trajectory of objects within the video, thereby decomposing the entire data into image frames representing the insertion stage when the user's hand enters and the removal stage when the hand exits. By recognizing image frames from two or more stages, the types of items inserted and removed are obtained. Combining the motion trajectory and weight change values, this application can accurately identify the types and quantities of items inserted and removed by the user. This application has the advantages of accurate recognition and high information update efficiency, greatly improving the management efficiency of hospital consumables. Attached Figure Description

[0058] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0059] Figure 1 This is an application scenario diagram of a hospital intelligent consumables management cabinet system based on a multimodal visual large model, as shown in one embodiment of this application;

[0060] Figure 2 This is a diagram of the actual hardware structure of a hospital intelligent consumables management cabinet system based on a multimodal visual large model, as shown in one embodiment of this application.

[0061] Figure 3 This is a structural diagram of a hospital intelligent consumables management cabinet system based on a multimodal visual large model, as shown in one embodiment of this application;

[0062] Figure 4 This is a schematic diagram of hand movement features in one embodiment of this application;

[0063] Figure 5 This is a schematic diagram of a motion feature image in one embodiment of this application;

[0064] Figure 6 This is a flowchart illustrating a hospital intelligent consumables management method based on a multimodal visual large model, as shown in one embodiment of this application. Detailed Implementation

[0065] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0066] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the layers related to the present invention and are not drawn according to the actual number, shape and size of the layers in the actual implementation. In the actual implementation, the form, number and proportion of each layer can be arbitrarily changed, and the layer layout may also be more complex.

[0067] Numerous details are explored in the following description to provide a more thorough explanation of embodiments of the invention; however, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details.

[0068] Figure 1 This is an application scenario diagram of a hospital intelligent consumables management cabinet system based on a multimodal visual large model, as shown in one embodiment of this application. Figure 1As shown, the management cabinet of this application is divided into multiple storage units by multiple partitions. Each storage unit is equipped with a gravity sensing module at its bottom. A photoelectric sensor is installed at the junction of the cabinet door 110 and the cabinet body to sense the opening and closing status of the cabinet door. A camera 120 with a downward-facing camera is installed on the top of the cabinet body. An RFID sensing module 130 is installed on the front of the cabinet body. The cabinet body and the cabinet door 110 are locked by an electromagnetic lock 140. When a user needs to open the cabinet door 110 to retrieve the required items, they can verify their identity information by swiping their hospital employee card through the RFID sensing module 130. When the RFID sensing module 130 senses the hospital employee card, the system will automatically read the card information and verify the employee's identity. After the identity is verified, the relevant equipment will automatically unlock the cabinet door.

[0069] Figure 2 This is a diagram illustrating the actual hardware structure of a hospital intelligent consumables management cabinet system based on a multimodal visual large model, as shown in one embodiment of this application. Figure 2 As shown, from a hardware structure perspective, this application also includes a microcontroller and an analysis device. When cabinet door 110 is opened, the photoelectric sensor transmits the opening information to the microcontroller, which then controls camera 120 to begin recording. When cabinet door 110 is closed, the photoelectric sensor transmits the closing information to the microcontroller, which then controls camera 120 to stop recording. The video data captured by camera 130 is transmitted to the analysis device for filtering, processing, and analysis to obtain consumable change information. Furthermore, the electromagnetic door lock of the management cabinet is also controlled by the microcontroller. When the RFID sensing module detects pre-registered radio frequency identification information, the microcontroller uses a related signal amplification circuit to control the electromagnetic door lock to unlock and open the door.

[0070] Figure 3 This is a structural diagram of a hospital intelligent consumables management cabinet system based on a multimodal visual large model, as shown in one embodiment of this application. Figure 3 As shown: This embodiment of the hospital intelligent consumables management cabinet system based on a multimodal visual large model includes:

[0071] The identity verification module 310 is used to verify the identity information received from the user and unlock the consumables management cabinet when the verification is successful.

[0072] First, if a user needs to access consumables, they need to verify their identity information using their employee card. Once the verification is successful, the cabinet door will unlock automatically.

[0073] The data acquisition module 320 is used to start timing, acquire the initial weight of multiple storage units and video when the cabinet door of the consumable management cabinet is opened, and stop timing, video acquisition and acquisition of the final weight of multiple storage units when the cabinet door of the consumable management cabinet is closed, so as to obtain the duration of a consumable collection cycle, video data and weight change values ​​of multiple storage units; and to take the video data of consumable collection cycles with a duration greater than a preset duration threshold as valid video data.

[0074] In its normal state, the cabinet door is kept closed by a spring hinge. The photoelectric sensor, consisting of a transmitter and a receiver, is typically integrated into a single device. It works by emitting light (usually infrared) from the transmitter. When the cabinet door is closed, the light is reflected back and received by the receiver. Once the door is opened, the light reflection path is interrupted, and the receiver cannot receive enough reflected light, thus triggering a change in the sensor's state. This allows for automatic detection of changes in the cabinet door's status.

[0075] When the cabinet door switches from closed to open, the microcontroller automatically controls the camera to start video recording and simultaneously collects the weight values ​​of multiple storage units at the start time. When the cabinet door is closed, the microcontroller automatically controls the camera to stop recording video, and simultaneously collects the weight values ​​of multiple storage units at the end time. At this time, the first Weight change value of each storage unit .

[0076] To avoid unnecessary processing of video data generated when cabinet doors are accidentally opened (rapidly opened and closed), videos shorter than 2 seconds are invalidated (it is generally believed that access cannot be completed normally within 2 seconds). Videos longer than 2 seconds are considered valid video data. Only valid video data undergoes further processing.

[0077] Image extraction module 330 is used to extract frames from the effective video data to obtain multiple keyframes;

[0078] In this application, the effective video data has a frame rate of 60, meaning there are 60 frames of video data per second. Therefore, it is necessary to extract keyframes from this data to save subsequent analysis and computation time and improve analysis efficiency. The frame extraction process includes:

[0079] (1) Perform interval frame extraction on the effective video data to obtain multiple initial frames;

[0080] First, by extracting frames at fixed intervals (such as extracting one frame every 10 frames), the total amount of video data is significantly reduced, thereby reducing storage requirements and computational load.

[0081] (2) Perform grayscale conversion on multiple initial frames to obtain multiple grayscale image frames;

[0082] Then, the color image is converted to a grayscale image (keeping only the brightness information) to reduce the computational overhead of the color channels.

[0083] (3) Perform a difference operation on adjacent grayscale image frames to obtain a difference image; and perform binarization processing on the difference image to obtain a binarized image;

[0084] By calculating the absolute difference between adjacent grayscale frames (inter-frame difference method), pixel changes between frames are detected, thereby locating moving regions. The difference image is then converted into a binary image (black and white image) to highlight the difference between the foreground (moving region) and the background.

[0085] (4) Extract the motion region in the binarized image, and based on the proportion of the motion region, when the proportion of the motion region is greater than a set threshold, determine that the next frame in the adjacent grayscale image frame is a dynamic frame, and use the dynamic frame as a key frame.

[0086] By statistically analyzing the pixel proportion of moving regions in the binarized image, it is determined whether the current frame contains significant motion. When the proportion of moving regions exceeds a set threshold, the next frame is determined to be a dynamic frame and is retained as a keyframe for subsequent analysis or storage.

[0087] The segmentation module 340 is used to extract optical flow information from multiple key frames and construct the motion trajectory of an object within the key frames using the optical flow information; and to divide the multiple key frames into insertion stage image frames and extraction stage image frames based on the object motion trajectory.

[0088] After extracting keyframes from the captured images of moving objects, this application analyzes the object's motion direction and trajectory using optical flow information, thereby dividing the keyframes into different motion stages. Object recognition is then performed on image frames at different motion stages. The process of extracting optical flow information from multiple keyframes and constructing the object's motion trajectory within each keyframe using this information includes:

[0089] (1) Extract key feature points from the key frame, wherein the key feature points include hand joint feature points and corner points;

[0090] By extracting hand joint feature points (such as fingers and wrist) and image corner points (such as edge intersections), the focus is placed on regions in the video that may contain dynamic information. Processing only key feature points rather than the entire image pixels significantly reduces subsequent computational complexity.

[0091] Corner detection is achieved using response functions (such as Harris). Hand joint feature points can be extracted using models such as OpenPose.

[0092] (2) Calculate the optical flow field in the key frame based on the key feature points to obtain the motion vector of each key feature point. ,in, This represents the displacement of the key feature point along the x-axis. This indicates the displacement of the key feature point along the y-axis;

[0093] This application is based on sparse optical flow. It uses optical flow algorithms (such as Lucas-Kanade) to calculate the displacement vector of each key feature point between adjacent frames, describing its motion trajectory and obtaining the optical flow field. The optical flow vector directly reflects the spatial motion of the feature points, providing a data foundation for subsequent clustering and trajectory analysis.

[0094] (3) Based on motion vectors Density clustering is performed on the key feature points of each keyframe to obtain one or more key feature point clusters;

[0095] If multiple moving objects exist in a scene (such as in a collaborative operation involving multiple people), clustering can group their feature points to avoid trajectory confusion. Since the optical flow field may originate from one or more objects, this application uses density clustering to separate optical flow vectors with different motion characteristics and group optical flow vectors with similar motion characteristics together to form key feature point clusters. Each key feature point cluster represents a moving object.

[0096] (4) Based on all motion vectors in the key feature point cluster Calculate the average value to obtain the average vector. And calculate the average coordinates of the key feature points in the key feature point cluster. ,in, This represents the average displacement along the x-axis. This represents the average displacement along the y-axis. The average coordinates of the x-axis. This represents the average coordinate of the y-axis;

[0097] Since the calculations above yielded one or more key feature point clusters, to facilitate the analysis of the motion trajectory of independent objects, all motion vectors within each key feature point cluster are analyzed. Perform mean calculation and mean calculation for each key feature point, thereby utilizing the average coordinates. Represents the position of an individual object, and uses the average vector. It describes the motion characteristics of an independent object.

[0098] (5) Convert the average vector from a biaxial displacement representation to a representation of direction angle and uniaxial displacement, and combine it with the average coordinates. Obtain the target vector ,in:

[0099]

[0100]

[0101] In the formula, It is the direction angle. It is the displacement;

[0102] To facilitate the representation of motion of independent objects, this application converts the displacements along the x-axis and y-axis into the form of direction angles and displacements. Figure 4 This is a schematic diagram of hand movement features in one embodiment of this application. The converted hand movement features are as follows: Figure 4 As shown.

[0103] (6) Target vector based on multiple keyframes Constructing the motion trajectory of one or more objects, specifically including:

[0104] (6-1) Obtain a background image with the same viewpoint as the video data;

[0105] In this embodiment, a background image is obtained by directly using a camera and a fill light to capture the inside of the cabinet from the same angle.

[0106] (6-2) The target vectors of the multiple keyframes The motion feature image is obtained by mapping it onto the background image;

[0107] During the mapping process, As the starting point of the vector Indicates the direction of the vector. Indicates the length of the vector. Figure 5 This is a schematic diagram of a motion feature image in one embodiment of this application. The motion feature image obtained by mapping vectors is as follows: Figure 5 As shown.

[0108] (6-3) In the motion feature image, with each target vector starting position Define a radius of for the origin. The matching range, where, , This is an empirical proportional parameter;

[0109] For a single feature point, adjacent optical flow vectors can generally be directly concatenated. However, to represent the overall motion characteristics of an object, clustering and mean calculation have been performed as described above. Therefore, the resulting overall vector may not be able to be concatenated. To accurately represent the motion trajectory of a single object and avoid confusion between different objects, this application addresses this issue at each starting position. A matching range is defined, and the radius of the matching range is determined by the object's speed. Higher speeds may lead to greater deviations, thus requiring a larger matching range and therefore a larger radius. , The value can be 0.1.

[0110] (6-4) Match any two adjacent target vectors based on the matching range, wherein when the end of one of the target vectors falls into the matching range of the adjacent target vector, the two adjacent target vectors are determined to be matched;

[0111] (6-5) Fit the starting point of all matched target vectors to obtain the motion trajectory of the object.

[0112] After obtaining the object's motion trajectory, keyframes can be divided based on this trajectory, resulting in image frames for each motion stage. Recognizing each motion stage's image frames provides a clear understanding of what the user put in and took out. The specific process includes:

[0113] (1) Calculate the distance between the starting points of any two adjacent target vectors in the trajectory of the object. And based on the time interval between adjacent keyframes Calculate the velocity of an object moving along a multi-segment trajectory , ;

[0114] (2) Based on the orientation angle of each keyframe The speed of an object and collection time point Construct the feature vector of the stage to which each keyframe belongs. ;

[0115] This application is based on the orientation angle of each keyframe. The speed of an object To determine the stage to which it belongs, the feature vector of the stage is... Including direction angle The speed of an object and the time point of collection .

[0116] (3) Construct a timeline and combine the feature vectors of the stages to which multiple keyframes belong. Mapped to the time axis;

[0117] (4) Slide along the time axis based on the pre-constructed sliding window, and calculate the average direction angle and average speed within the sliding window each time the sliding window is slidable; when the average direction angle falls into a preset first angle range and the average speed is greater than or equal to a preset speed threshold, determine that the segment within the sliding window is an insertion segment; when the average direction angle falls into a preset second angle range and the average speed is greater than or equal to a preset speed threshold, determine that the segment within the sliding window is an extraction segment; when the average speed is less than a preset speed threshold, determine that the segment within the sliding window is an operation segment.

[0118] During the above process, the overall direction and speed of the object are collected over a period of time (e.g., 0.5 seconds) via a sliding window. If the average direction angle falls within a preset first angle range, and the average speed is greater than or equal to a preset speed threshold, it indicates that the object (such as a hand or consumables) is moving towards the cabinet at a relatively high speed, thus determining that the user's hand is in the insertion section. Conversely, if the average direction angle falls within a preset second angle range, and the average speed is greater than or equal to a preset speed threshold, it indicates that the object is moving towards the cabinet at a relatively high speed, thus determining that the user's hand is in the removal section.

[0119] If the movement speed is low, it is determined that the user is currently performing operations such as tidying up or putting things away in front of the cabinet. During these operations, the identified consumables are meaningless, so the image frames corresponding to this stage are filtered out.

[0120] (5) Merge consecutive insertion and extraction segments to obtain the time interval of the insertion stage and the time interval of the extraction stage, and extract the insertion stage image frame from the time interval of the insertion stage and the extraction stage image frame from the time interval of the extraction stage.

[0121] Finally, since the insertion and removal segments extracted through the sliding window are scattered, it is necessary to determine whether the user's hand is in the insertion segment at this time. Phases lasting less than 1 second should be removed, and the longer-lasting phases should be used as the final insertion or removal phase. The corresponding image frames should then be extracted from the corresponding time intervals.

[0122] The type identification module 350 is used to identify the image frames of the insertion stage and the image frames of the removal stage respectively, and obtain the type of items inserted, the initial quantity of items inserted, the type of items removed, and the initial quantity of items removed.

[0123] After extracting image frames from different stages of the user's actions, identifying these frames allows us to determine which consumables the user is using at each stage. This application utilizes AI technology for this identification process, which is as follows:

[0124] (1) Remove the non-motion background regions of the input stage image frame and the output stage image frame to obtain the input image;

[0125] In the previous text, the motion part of the keyframe was extracted using the frame difference method. Therefore, only the motion part is retained and reconstructed into the input image.

[0126] (2) Input the input image into the pre-built recognition model to obtain the type of items put in, the initial quantity of items put in, the type of items taken out, and the initial quantity of items taken out.

[0127] In this application, the recognition model performs preliminary identification of the types and quantities of consumables in the image. The method for constructing the recognition model includes:

[0128] Acquire a sample image and perform expansion processing on the sample image to obtain an expanded sample. The expansion processing includes image scaling, rotation, cropping, low-light enhancement, ghosting removal, and dynamic de-glare processing.

[0129] The sample images are then labeled to obtain training samples;

[0130] The artificial neural network is trained by backpropagation based on the training samples to obtain the recognition model.

[0131] First, collect images of objects from multiple angles, covering scenes with different angles, lighting, and backgrounds. Then, annotate the images (e.g., box out object locations, classify them). Weakly supervised learning techniques can reduce reliance on annotations, allowing for the training of a high-precision model with a small amount of labeled data.

[0132] Image processing: Image scaling, rotation, cropping, etc., to enhance the model's generalization ability. To address issues such as low light, blur, and reflection, techniques such as dark light enhancement, ghosting removal, and dynamic de-reflection are used to improve image quality.

[0133] The core models primarily employ convolutional neural networks, automatically extracting local features such as edges and textures, as well as global features such as shape and structure, through multiple convolutional layers. The training process optimizes parameters through backpropagation, enabling the model to distinguish different objects. Dynamic multi-object recognition: compressing the number of model layers allows for rapid object recognition at the terminal.

[0134] The object detection model identifies the category and locates the object's position (outputting a bounding box). In multi-object scenarios, it needs to be combined with the SLAM tracking algorithm to maintain the stability of the object in consecutive frames and avoid jitter.

[0135] Basic models (such as ResNet and MobileNet) extract common features such as color and texture.

[0136] To address product differences (such as different specifications of consumable packaging), an attention mechanism is introduced to focus on local details (such as packaging text and logo).

[0137] Coarse classification (consumable categories) → fine classification (e.g., "gauze, scissors"). Accuracy is improved by combining cross-modal learning with multi-source information such as images.

[0138] The quantity recognition module 360 ​​is used to verify the initial quantity of the put-in items and the initial quantity of the taken-out items based on the weight change value of the multiple storage units, the motion trajectory of the object in the key frame, the type of put-in items, and the type of taken-out items, so as to obtain the quantity of put-in items and the quantity of taken-out items.

[0139] After obtaining the recognition results, voltage change values, and object motion trajectories within keyframes, logical reasoning can be performed based on these results to obtain change information, specifically including:

[0140] (1) Obtain a pre-built consumable weight data table, wherein the consumable weight data table includes weight references for various consumables, and an example of the consumable weight data table is shown in the table below.

[0141] Table 1. Example of Consumable Weight Data

[0142] (2) Determine the target storage unit that produces the weight change based on the object's motion trajectory within the key frame;

[0143] When the end of the motion trajectory is within the coverage area of ​​a storage unit, that storage unit is used as the target storage unit.

[0144] (3) Extract the weight change value of the target storage unit from the weight change values ​​of the plurality of storage units. ;

[0145] weight change value The extraction process is as described above and will not be repeated here.

[0146] (4) Obtain the reference weight of the type of item put in or the type of item taken out from the consumable weight data table, and when the reference weight is less than the minimum weight threshold, take the initial quantity of the items put in and the initial quantity of the items taken out as the quantity of items put in and the quantity of items taken out, respectively; when the reference weight is greater than or equal to the minimum weight threshold, calculate the reference change value based on the initial quantity of the items put in and the reference weight of the type of items put in. Alternatively, a reference change value can be calculated based on the initial quantity of items taken out and the reference weight of the types of items taken out. ;

[0147] For lighter consumables, such as those weighing less than 20 grams, the change in total weight caused by variations is small. However, if the quantity is large, errors in the total weight variation can easily lead to calculation errors. Therefore, for consumables weighing less than 20 grams, the result of visual recognition is taken as the final accurate result.

[0148] For consumables weighing 20 grams or more, it is necessary to use the change in gravity to calibrate the image recognition results, thereby making the results more accurate.

[0149] Reference change value It is the product of the quantity of the items and their reference weight.

[0150] (5) Calculate the weight change value Compared with reference change value deviation rate , ;

[0151] (6) In the deviation rate When the deviation rate is less than or equal to a preset deviation rate threshold, the initial quantity of items put in and the initial quantity of items taken out are taken as the quantity of items put in and the quantity of items taken out, respectively; in the deviation rate When the deviation rate exceeds a preset threshold, the weight change value is calculated. The ratio of the reference weight to the target item type The target item type is either the type of item to be placed or the type of item to be taken out.

[0152] If the deviation rate is small, it indicates that the visual recognition result is correct. If the deviation rate is large, since there is a high probability of error in the quantity recognized visually, it can be considered that the quantity recognition result may be incorrect. To verify this possibility, the weight change value is calculated. The ratio of the reference weight to the target item type .

[0153] (7) In the ratio If the non-integer part is less than the target value, then the ratio will be... As an estimated quantity, the estimated quantity and the valid video data are sent to the target object. After receiving verification information from the target object, the quantity of items put in and the quantity of items taken out are obtained; in the ratio When the non-integer part is greater than or equal to the target value, an error message and the valid video data are sent to the target object. After receiving the verification information from the target object, the number of items put in and the number of items taken out are obtained.

[0154] If the ratio If the value is an integer, or the decimal part is small (less than 0.1), it can be inferred that there might be a problem with the quantity recognized visually, but no problem with the category recognition. Therefore, the ratio... As an estimated quantity, the estimated quantity and the valid video data are sent to the management personnel for verification.

[0155] If the decimal part is large, there may be a problem with the type being identified. In this case, an error will be reported and the valid video data will be sent to the administrator for verification.

[0156] The management module 370 is used to construct cabinet inventory change information based on the type of items put in, the quantity of items put in, the type of items taken out, and the quantity of items taken out, and to manage consumables based on the cabinet inventory change information.

[0157] Finally, the system obtains the user's identity information, the types and quantities of items put in, the types and quantities of items taken out, and constructs inventory change information for the cabinet. The system then adds the inventory change information to the ledger and automatically updates the inventory of the corresponding items.

[0158] This invention relates to a hospital intelligent consumables management cabinet system based on a multimodal visual large model. By collecting video data and weight change values ​​from the opening to closing of the cabinet, and filtering out invalid video data based on the duration of the video data, the system improves the recognition accuracy. Furthermore, this invention extracts optical flow information from the video data and uses this information to identify the motion trajectory of objects within the video. This decomposes the entire data into image frames representing the insertion stage (when the user's hand enters) and the removal stage (when the hand withdraws). By identifying image frames from two or more stages, the types of items inserted and removed can be determined. Combining the motion trajectory and weight change values, this invention can accurately identify the types and quantities of items inserted and removed by the user. This invention offers advantages such as accurate recognition and high information update efficiency, significantly improving the management efficiency of hospital consumables.

[0159] like Figure 6 As shown, this application also provides a hospital intelligent consumables management method based on a multimodal visual large model, including:

[0160] S610 starts timing, collecting the initial weight of multiple storage units and video when the cabinet door of the consumable management cabinet is opened, and stops timing, video collection and collecting the final weight of multiple storage units when the cabinet door of the consumable management cabinet is closed, to obtain the duration of a consumable collection cycle, video data and weight change values ​​of multiple storage units; and takes video data of consumable collection cycles with a duration greater than a preset duration threshold as valid video data;

[0161] S620, extract frames from the effective video data to obtain multiple key frames;

[0162] S630, extract optical flow information from multiple key frames, and use the optical flow information to construct the object motion trajectory within the key frames; and divide the multiple key frames into insertion stage image frames and extraction stage image frames based on the object motion trajectory.

[0163] S640, the image frames of the insertion stage and the image frames of the removal stage are identified respectively to obtain the type of items inserted, the initial quantity of items inserted, the type of items removed, and the initial quantity of items removed.

[0164] S650, based on the weight change value of the multiple storage units, the motion trajectory of the object in the key frame, the type of item put in and the type of item taken out, the initial quantity of the put in and the initial quantity of the taken out are verified to obtain the quantity of the put in and the quantity of the taken out.

[0165] S660: Construct cabinet inventory change information based on the type of items put in, the quantity of items put in, the type of items taken out, and the quantity of items taken out, and perform consumable management based on the cabinet inventory change information.

[0166] This invention discloses a hospital intelligent consumables management method based on a multimodal visual large model. It improves the recognition accuracy by collecting video data and weight change values ​​from the opening to closing of the management cabinet, and filtering out invalid video data based on the duration of the video data. Furthermore, this invention extracts optical flow information from the video data and uses this information to identify the motion trajectory of objects within the video. This decomposes the entire data into image frames representing the insertion stage (when the user's hand enters) and the removal stage (when the hand withdraws). By identifying image frames from two or more stages, the types of items inserted and removed are determined. Combining the motion trajectory and weight change values, this invention can accurately identify the types and quantities of items inserted and removed by the user. This invention offers advantages such as accurate recognition and high information update efficiency, significantly improving the management efficiency of hospital consumables.

[0167] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any one of the methods in this embodiment, wherein the method is the execution logic of this system.

[0168] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0169] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to cause the terminal to perform any of the methods in this embodiment.

[0170] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0171] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal performs the steps of the above method.

[0172] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0173] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0174] In the above embodiments, although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. The embodiments of the invention are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims.

[0175] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A hospital intelligent consumables management cabinet system based on a multimodal visual large model, characterized in that, include: The data acquisition module is used to start timing, collect the initial weight of multiple storage units and video when the cabinet door of the consumable management cabinet is opened, and stop timing, video acquisition and the final weight of multiple storage units when the cabinet door of the consumable management cabinet is closed, so as to obtain the duration of a consumable collection cycle, video data and weight change values ​​of multiple storage units. In addition, video data with a duration exceeding a preset duration threshold for consumable collection cycles will be considered as valid video data; The image extraction module is used to extract frames from the effective video data to obtain multiple keyframes; The segmentation module is used to extract optical flow information from multiple keyframes and to construct the motion trajectory of objects within the keyframes using the optical flow information. And based on the object's motion trajectory, the multiple keyframes are divided into placement stage image frames and removal stage image frames; The category recognition module is used to recognize the image frames of the insertion stage and the image frames of the removal stage respectively, and obtain the category of the inserted item, the initial quantity of the inserted item, the category of the removed item, and the initial quantity of the removed item. The quantity recognition module is used to verify the initial quantity of the put-in items and the initial quantity of the taken-out items based on the weight change value of the multiple storage units, the motion trajectory of the object in the key frame, the type of put-in items, and the type of taken-out items, so as to obtain the quantity of put-in items and the quantity of taken-out items. The management module is used to construct cabinet inventory change information based on the type of items put in, the quantity of items put in, the type of items taken out, and the quantity of items taken out, and to manage consumables based on the cabinet inventory change information.

2. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 1, characterized in that, Frames are extracted from the valid video data to obtain multiple keyframes, including: The valid video data is subjected to interval frame extraction to obtain multiple initial frames; Multiple initial frames are converted to grayscale to obtain multiple grayscale image frames; Perform a difference operation on adjacent grayscale image frames to obtain a difference image; then perform binarization on the difference image to obtain a binarized image; Extract the motion region from the binarized image, and based on the proportion of the motion region, when the proportion of the motion region is greater than a set threshold, determine that the next frame in the adjacent grayscale image frame is a dynamic frame, and use the dynamic frame as a key frame.

3. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 1, characterized in that, Optical flow information from multiple keyframes is extracted, and the object motion trajectory within the keyframes is constructed using the optical flow information, including: Extract key feature points from the keyframe, wherein the key feature points include hand joint feature points and corner points; The optical flow field in the key frame is calculated based on the key feature points to obtain the motion vector of each key feature point. ,in, This represents the displacement of the key feature point along the x-axis. This indicates the displacement of the key feature point along the y-axis; Based on motion vectors Density clustering is performed on the key feature points of each keyframe to obtain one or more key feature point clusters; Based on all motion vectors in the key feature point cluster Calculate the average value to obtain the average vector. And calculate the average coordinates of the key feature points in the key feature point cluster. ,in, This represents the average displacement along the x-axis. This represents the average displacement along the y-axis. Represents the average coordinates of the x-axis. This represents the average coordinate of the y-axis; The average vector is converted from a biaxial displacement representation to a representation of direction angle and uniaxial displacement, and then combined with the average coordinates. Obtain the target vector ,in: In the formula, It is the direction angle. It is the displacement; Target vector based on multiple keyframes Construct the motion trajectories of one or more objects.

4. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 3, characterized in that, Target vector based on multiple keyframes Constructing the motion trajectories of one or more objects, including: Obtain a background image with the same viewpoint as the video data; The target vector of the multiple key frames The motion feature image is obtained by mapping it onto the background image; In the motion feature image, with each target vector starting position Define a radius of for the origin. The matching range, where, , This is an empirical proportional parameter; Based on the matching range, any two adjacent target vectors are matched, wherein when the end of one of the target vectors falls within the matching range of the adjacent target vector, the two adjacent target vectors are determined to be matched. By fitting the starting points of all matching target vectors, the trajectory of the object is obtained.

5. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 4, characterized in that, Based on the object's motion trajectory, the multiple keyframes are divided into placement stage image frames and removal stage image frames, including: Calculate the distance between the starting points of any two adjacent target vectors in the trajectory of the object. And based on the time interval between adjacent keyframes Calculate the velocity of an object moving along a multi-segment trajectory , ; Based on the orientation angle of each keyframe The speed of an object and collection time point Construct the feature vector of the stage to which each keyframe belongs. ; Construct a timeline and combine the feature vectors of the stages to which multiple keyframes belong. Mapped to the time axis; Based on a pre-built sliding window, the sliding window slides along the time axis, and at each slide, the average direction angle and average speed within the sliding window are calculated. When the average direction angle falls within a preset first angle range and the average speed is greater than or equal to a preset speed threshold, the segment within the sliding window is determined to be an insertion segment. When the average direction angle falls within a preset second angle range and the average speed is greater than or equal to a preset speed threshold, the segment within the sliding window is determined to be an extraction segment. When the average speed is less than a preset speed threshold, the segment within the sliding window is determined to be an operation segment. The consecutive insertion and removal segments are merged to obtain the time interval of the insertion stage and the time interval of the removal stage. The insertion stage image frames are extracted from the time interval of the insertion stage and the removal stage image frames are extracted from the time interval of the removal stage.

6. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 1, characterized in that, The image frames of the insertion stage and the image frames of the removal stage are identified respectively to obtain the type of item inserted, the initial quantity of items inserted, the type of item removed, and the initial quantity of items removed, including: Remove the non-moving background regions from the input and output image frames to obtain the input image; The input image is input into a pre-built recognition model to obtain the type of items put in, the initial quantity of items put in, the type of items taken out, and the initial quantity of items taken out.

7. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 6, characterized in that, The method for constructing the recognition model includes: Acquire a sample image and perform expansion processing on the sample image to obtain an expanded sample. The expansion processing includes image scaling, rotation, cropping, low-light enhancement, ghosting removal, and dynamic de-glare processing. The sample images are then labeled to obtain training samples; The artificial neural network is trained by backpropagation based on the training samples to obtain the recognition model.

8. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 1, characterized in that, Based on the weight change values ​​of the multiple storage units, the object motion trajectory within the keyframe, and the types of items put in and taken out, the initial quantity of items put in and taken out is verified to obtain the total quantity of items put in and taken out, including: Obtain a pre-built consumable weight data table, wherein the consumable weight data table includes weight references for various consumables; The target storage unit that produces the weight change is determined based on the object's motion trajectory within the keyframe. Extract the weight change value of the target storage unit from the weight change values ​​of the plurality of storage units. ; Obtain the reference weight of the type of item put in or taken out from the consumable weight data table. When the reference weight is less than the minimum weight threshold, use the initial quantity of the put-in items and the initial quantity of the taken-out items as the put-in items quantity and the taken-out items quantity, respectively. When the reference weight is greater than or equal to the minimum weight threshold, calculate a reference change value based on the initial quantity of the put-in items and the reference weight of the put-in item type. Alternatively, a reference change value can be calculated based on the initial quantity of items taken out and the reference weight of the types of items taken out. ; Calculate the change in weight Compared with reference change value deviation rate , ; The deviation rate When the deviation rate is less than or equal to a preset deviation rate threshold, the initial quantity of items put in and the initial quantity of items taken out are taken as the quantity of items put in and the quantity of items taken out, respectively; in the deviation rate When the deviation rate exceeds a preset threshold, the weight change value is calculated. The ratio of the reference weight to the target item type The target item type is either the type of item to be placed or the type of item to be taken out. The ratio If the non-integer part is less than the target value, then the ratio will be... As an estimated quantity, the estimated quantity and the valid video data are sent to the target object. After receiving verification information from the target object, the quantity of items put in and the quantity of items taken out are obtained; in the ratio When the non-integer part is greater than or equal to the target value, an error message and the valid video data are sent to the target object. After receiving the verification information from the target object, the number of items put in and the number of items taken out are obtained.

9. The hospital intelligent consumables management cabinet system based on a multimodal visual large model according to claim 1, characterized in that, Also includes: The identity verification module is used to verify the identity information received from the user, and to unlock the consumables management cabinet when the verification is successful.

Citation Information

Patent Citations

  • Information processing method and device, equipment and storage medium

    CN114494964A

  • Goods shelf layer identification method and device for commodity change of vending machine and vending machine system

    CN118097522A

  • Large transport vehicle identification method and system based on radar and vision fusion

    CN119478860A

  • Medical consumable use compliance monitoring method and system based on Internet of Things

    CN120727231A

  • Intelligent shopping cart shopping behavior determination method

    WO2025039457A1

Cited By

  • Hospital consumable inventory management system and method based on behavior recognition

    CN121528470A