AI-based intelligent glasses medicine label identification method and system

By collecting the angular velocity and brightness distribution of the user's head in smart glasses, combining semantic matching and gaze behavior analysis, the problem of insufficient user status perception in traditional drug label recognition technology is solved, and more accurate drug recognition and information broadcasting is achieved that is more in line with user needs.

CN120544179APending Publication Date: 2025-08-26SHENZHEN SMART CLOUD TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510627508.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional drug label recognition technology lacks perceived feedback on user status changes, resulting in image out of focus and information loss, the identification content is inconsistent with user needs, and voice broadcasts cannot adapt to user's attention changes, affecting information understanding.

Method used

The user's head angular velocity information is collected through the acceleration sensor, the image acquisition equipment parameters are adjusted, and the drug label structure is identified by combining brightness distribution and texture difference analysis, the semantic matching mechanism is used to improve the correlation of drug symptoms, and the user's gaze behavior is monitored to adjust the broadcast rhythm.

Benefits of technology

It improves image acquisition stability and label recognition accuracy, optimizes the synchronization and adaptability of voice broadcasts, and improves the intelligence level of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544179A_ABST
    Figure CN120544179A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of man-machine interaction, in particular to an intelligent glasses medicine label identification method and system based on AI, and the method comprises the following steps: calling a sensor to collect the biaxial angular velocity information of a head, analyzing the posture stability, shooting a packaging image, extracting the brightness data of the image, delimiting a structure boundary, and identifying packaging characters; and constructing a drug symptom semantic matching score, identifying a drug demand, and generating a broadcast delay parameter value. According to the method, the stability and definition of image acquisition are improved by combining user posture perception with an image acquisition linkage mechanism, the accuracy of label structure identification is enhanced by using brightness distribution and texture difference analysis, the accuracy of medicine symptom association judgment is improved by adopting a semantic fitting matching mechanism, and the accuracy of medicine symptom association judgment is improved by combining user gazing behavior analysis. And the synchronism and adaptability of voice broadcasting are optimized, so that the information output rhythm better fits the receiving state of the user, and the intelligent level of interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and in particular to an AI-based smart glasses drug label recognition method and system. Background Art

[0002] The field of human-computer interaction technology encompasses the entire process of information input and output, perception response, and operational feedback between humans and computing devices. The core content of this technology includes the design of interactive interfaces for perception input methods such as image recognition, voice recognition, and touch operations, as well as multimodal interaction methods for output feedback such as graphical interface display, voice synthesis broadcast, and vibration prompts. Human-computer interaction covers multiple aspects, including the integrated design of input acquisition equipment, input response of data processing modules, user behavior modeling, and adaptation mechanisms of interactive feedback modules. It is widely used in scenarios such as wearable devices, auxiliary medical terminals, and intelligent control systems to achieve a natural and efficient information transmission process between humans and devices.

[0003] Among them, an AI-based smart glasses drug label recognition method refers to collecting drug packaging images through the camera component integrated in the smart glasses, segmenting and identifying the text information in the image based on the character recognition model constructed by the convolutional neural network, using the set drug label structure template to perform information analysis, and completing the recognition of the drug label content by matching with the preset drug database. The technical matters of this method cover image information acquisition, image preprocessing, character recognition, barcode and QR code recognition, drug label field structure analysis, label content and database information comparison, etc. The model and analysis operations required for recognition are completed by the built-in embedded processing module of the glasses, and the recognition results are output through the display and voice output module of the glasses to help users complete the perception and understanding of drug information.

[0004] Traditional drug label recognition technology adopts static collection and fixed process information analysis methods in the human-computer interaction link. Image acquisition equipment usually runs continuously based on default parameters and lacks a perception feedback mechanism for changes in user status, which can easily lead to image defocus, information missing, etc. Label recognition mostly relies on image character extraction and template comparison, which is difficult to cover the semantic changes of users' actual concerns. There is a problem of inconsistency between the recognized content and the user's real needs. A unified voice template is used for information broadcast output, ignoring the user's attention changes and visual response in the process of receiving information, which can easily cause information comprehension gaps or attention overflow, affecting the broadcast effect. When the user reads slowly, the fixed rhythm output may cause information omissions. When the eyes wander frequently, the information rhythm cannot be adjusted to match the user rhythm, resulting in difficulty for users to understand. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the existing technology and propose an AI-based smart glasses drug label recognition method and system.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for identifying drug labels using smart glasses based on AI, comprising the following steps:

[0007] S1: Call the acceleration sensor to collect the angular velocity information of the user's head in the horizontal and vertical directions, extract the angular velocity change rate and angular velocity fluctuation amplitude in each direction, analyze the stability of the user's head posture in a continuous time period, adjust the control parameters of the image acquisition device, and obtain the packaging image set;

[0008] S2: Calling the packaging image set, extracting brightness distribution data in the pharmaceutical packaging image, analyzing the brightness change trend and boundary clarity of multiple regions in the image, evaluating the texture difference between the character and barcode regions, delineating the image structure boundary, and obtaining the structure segmentation result;

[0009] S3: Calling the structural division results, extracting the indication semantic phrases in the label field, combining them with the disease keyword phrases in the user description, performing correspondence analysis on the semantic directionality, part-of-speech structure, and expression paradigm, establishing a semantic fit score between the drug field and the disease phrase, and forming a drug-disease association value;

[0010] S4: Based on the drug-disease association value, identify the user's demand for the target drug, call the medical database based on the drug name identified in the label field, extract the drug name, production batch, expiration date, usage and dosage, and precautions of the target drug, use the AI ​​big model to evaluate the drug safety of the target drug and output risk warnings, build the broadcast information output content, and obtain the output information construction record.

[0011] As a further solution of the present invention, the packaging image set includes an image frame sequence number, an image clarity identifier, and an acquisition time tag; the structural division result is specifically a character boundary position set, a barcode block coordinate, and a background interference area annotation; the drug-symptom association value specifically refers to a semantic overlap value, a part-of-speech structure adaptation coefficient, and a keyword expression distance score; and the output information construction record includes a drug name field, a medication specification field, and a voice output field.

[0012] As a further solution of the present invention, the step of obtaining the packaging image set is specifically as follows:

[0013] S111: Calling the acceleration sensor to collect the angular velocity data sequence of the horizontal axis and the vertical axis of the user's head after wearing the smart glasses, extracting the angular velocity change rate and angular velocity fluctuation amplitude in each direction, and obtaining the posture angular velocity change data;

[0014] S112: Calling the attitude angular velocity change data, obtaining the angular velocity deviation value of each time period in the continuous cycle, analyzing the data of the horizontal direction, vertical direction and vertical axis change rate, and calculating the attitude stability state evaluation value;

[0015] S113: According to the posture stability state evaluation value and the stability of the user's head posture, the control parameters of the image acquisition device are adjusted to perform image acquisition and obtain a packaging image set.

[0016] As a further solution of the present invention, the steps of obtaining the structure division result are specifically as follows:

[0017] S211: calling the packaged image set, extracting the grayscale value of each pixel in the image area, calculating the average grayscale value in each area block, and obtaining the brightness distribution value of the image area;

[0018] S212: Calling the brightness distribution value of the image region, extracting the grayscale mutation points of the region boundary pixels, calculating the grayscale mean value of the inner region of each boundary line, the grayscale mean value of the outer region, and the corresponding pixel gradient value, performing normalized gradient structure evaluation on the extracted multiple groups of boundary feature values, calculating the grayscale jump response intensity value of the image boundary, and generating a boundary clarity gradient value based on the regional grayscale mutation direction, gradient density change, and boundary response difference;

[0019] S213: Call the boundary clarity gradient value, perform morphological classification and connectivity calculation on the structural boundary lines in the image, identify pixel partition boundaries, perform structural label classification on character blocks and barcode blocks, and establish a structural division result.

[0020] As a further solution of the present invention, the step of obtaining the drug-disease correlation value is specifically as follows:

[0021] S311: Calling the structural division result, extracting the drug field in the label image, identifying the phrase components in the field that express the drug indication content, calling the language segmentation component to perform semantic block division and extract part-of-speech attributes, classifying the phrase semantic pointing structure based on the context boundary, and obtaining a set of indication semantic phrases;

[0022] S312: Based on the set of indication semantic phrases, call the disease keyword group in the user description, extract the semantic similarity score, part-of-speech matching consistency value and semantic structure synergy score between each pair of phrases, calculate the structure synergy strength index in sequence according to the phrase pair index number, and calculate the semantic fit value of the drug phrase;

[0023] S313: Call the semantic fit value of the drug phrase, establish a corresponding relationship between the drug field and the disease description, and obtain the drug-disease association value.

[0024] As a further solution of the present invention, the steps of obtaining the output information construction record are specifically as follows:

[0025] S411: calling the drug-symptom association value, identifying the user's need for the target drug, and obtaining the target drug identification value;

[0026] S412: Based on the target drug identification value, call the pharmaceutical database to extract the drug name, production batch, expiration date, usage and dosage, and precautions of the target drug to obtain a medication information data set;

[0027] S413: Based on the medication information dataset, using the AI ​​big model, evaluate the user's medication safety of the target drug and output risk warnings, construct the output text of the voice broadcast information, and generate an output information construction record.

[0028] As a further embodiment of the present invention, the method further comprises:

[0029] S5: Calling the output information to build a record, monitoring the wearer's gaze position changes during voice output in real time, analyzing the gaze dwell time and gaze point coordinate offset trajectory, evaluating the user's visual concentration state during reading, and adjusting the start interval of field announcements to generate an announcement delay adjustment parameter value;

[0030] The broadcast delay adjustment parameter values ​​specifically include a gaze dwell interval value, a coordinate stability parameter, and a broadcast interval adjustment coefficient.

[0031] As a further solution of the present invention, the step of obtaining the broadcast delay adjustment parameter value is specifically as follows:

[0032] S511: Call the output information to build a record, monitor the wearer's gaze point position in real time during the voice broadcast process, collect the coordinate value of the gaze point in each frame, record the gaze trajectory displacement path, and obtain the gaze change tracking value;

[0033] S512: Identifying fluctuation trends of the gaze point in spatial and temporal dimensions within a continuous time period based on the gaze change tracking value, calculating the stability of the gaze and the intensity of the gaze concentration during the reading process, and obtaining a visual concentration state index value;

[0034] S513: Call the visual concentration state index value, adjust the interval time of field broadcast according to the change trend of the index value, and obtain the broadcast delay adjustment parameter value.

[0035] An AI-based smart glasses drug label recognition system, the AI-based smart glasses drug label recognition system is used to execute the above-mentioned AI-based smart glasses drug label recognition method, the system comprising:

[0036] The posture detection module uses an accelerometer to collect the instantaneous angular velocity values ​​of the user's head on the horizontal and vertical axes in real time. By calculating the stability of the change trend and fluctuation amplitude of the two-axis angular velocity over a continuous period of time, the control parameters of the image acquisition device are adjusted to obtain a set of packaging images.

[0037] The image texture analysis module extracts grayscale brightness distribution data of multiple regions in the image based on the package image set, identifies the texture difference characteristics between the character and barcode regions by analyzing the regional brightness gradient change trend and boundary texture clarity, divides the character and barcode structure regions, and generates a structure division result;

[0038] Based on the structural division results, the semantic matching module extracts the indication semantic phrases in the label field and compares and analyzes them with the disease keyword phrases in the user description, including semantic directionality, part-of-speech structure, and expression paradigm. It calculates the semantic matching degree between the phrases and establishes a semantic fit score between the drug field and the disease phrase to generate a drug-disease association value.

[0039] The medication information management module identifies the user's demand for the target drug based on the drug-symptom association value, extracts the dosage, time of administration, and method of administration of the corresponding drug according to the drug name field identified in the tag structure, constructs a drug usage information set, and creates an output information construction record;

[0040] The broadcast rhythm control module builds a record based on the output information, tracks and monitors the changes in the user's gaze position during the broadcast of information in real time, analyzes the gaze dwell time and the gaze point coordinate offset trajectory, evaluates the user's visual concentration state in the reading state, and dynamically adjusts the voice broadcast interval between information fields according to the concentration state, and generates a broadcast delay adjustment parameter value.

[0041] Compared with the prior art, the advantages and positive effects of the present invention are:

[0042] In the present invention, by combining user posture perception with image acquisition linkage mechanism, the stability and clarity of image acquisition are improved, the accuracy of label structure recognition is enhanced by utilizing brightness distribution and texture difference analysis, and the accuracy of drug-symptom association judgment is improved by adopting semantic fitting matching mechanism. Combined with user gaze behavior analysis, the synchronization and adaptability of voice broadcast are optimized, so that the information output rhythm is more in line with the user's reception status, and the intelligent level of interaction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0044] Figure 2 Obtaining a flow chart for the packaging image set of the present invention;

[0045] Figure 3 Obtaining a flow chart for the structural division results of the present invention;

[0046] Figure 4 A flowchart for obtaining the drug-disease correlation value of the present invention;

[0047] Figure 5 Constructing a record acquisition flow chart for the output information of the present invention;

[0048] Figure 6 This is a flow chart for obtaining the broadcast delay adjustment parameter value of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0050] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0051] See also Figure 1 The present invention provides a technical solution: a method for identifying drug labels on smart glasses based on AI, comprising the following steps:

[0052] S1: Call the acceleration sensor to collect the angular velocity information of the user's head in the horizontal and vertical directions, extract the angular velocity change rate and angular velocity fluctuation amplitude in each direction, analyze the stability of the user's head posture in a continuous time period, adjust the control parameters of the image acquisition device, and obtain the packaging image set;

[0053] S2: Call the packaging image set, extract the brightness distribution data in the drug packaging image, analyze the brightness change trend and boundary clarity of multiple areas in the image, evaluate the texture differences between the character and barcode areas, delineate the image structure boundary, and obtain the structure segmentation results;

[0054] S3: Call the structure segmentation results, extract the indication semantic phrases in the label field, combine them with the disease keyword phrases in the user description, perform correspondence analysis on the semantic directionality, part-of-speech structure, and expression paradigm, establish a semantic fit score between the drug field and the disease phrase, and form a drug-disease association value;

[0055] S4: Based on the drug-symptom association value, identify the user's demand for the target drug, call the medical database based on the drug name identified in the tag field, extract the drug name, production batch, expiration date, usage and dosage, and precautions of the target drug, use the AI ​​big model to evaluate the drug safety of the target drug and output risk warnings, construct the broadcast information output content, and obtain the output information construction record;

[0056] S5: Call the output information to build a record, monitor the changes in the wearer's gaze position during the voice output process in real time, analyze the gaze dwell time and the gaze point coordinate offset trajectory, evaluate the user's visual concentration state in the reading state, and adjust the start interval of the field broadcast to generate the broadcast delay adjustment parameter value.

[0057] The packaging image set includes the image frame sequence number, image clarity identification, and acquisition time label. The structural division results are specifically the character boundary position set, barcode block coordinates, and background interference area annotations. The drug-disease association value specifically refers to the semantic overlap value, part-of-speech structure adaptation coefficient, and keyword expression distance score. The output information construction record includes the drug name field, medication specification field, and voice output field. The broadcast delay adjustment parameter value is specifically the gaze dwell interval value, coordinate stability parameter, and broadcast interval adjustment coefficient.

[0058] See also Figure 2 , the steps for obtaining the packaged image set are as follows:

[0059] S111: Calling the acceleration sensor to collect the angular velocity data sequence of the horizontal axis and the vertical axis of the user's head after wearing the smart glasses, extracting the angular velocity change rate and angular velocity fluctuation amplitude in each direction, and obtaining posture angular velocity change data;

[0060] The accelerometer is used to collect angular velocity data sequences along the horizontal and vertical axes of the user's head after wearing smart glasses. Raw angular velocity data points are extracted at fixed time intervals during data sampling. Two independent angular velocity sequences are constructed for dynamic change monitoring in the horizontal and vertical directions, respectively. The system's built-in data synchronization mechanism ensures that the sampling points in each direction are time-aligned. Then, a difference operation is performed on the adjacent data points in each direction. The angular velocity change rate within each time period is expressed as an absolute value. All change rates within consecutive time periods are then aggregated to form a set of horizontal and vertical change rates. To obtain more stable representative feature values, the two sets are smoothed using a sliding average. After removing sudden data points, the maximum and minimum change values ​​within each cycle are extracted, and their difference is recorded as the fluctuation amplitude of that cycle. The fluctuation amplitudes of all cycles are then aggregated into a single sequence and median extracted to establish angular velocity fluctuation amplitude indicators for the two directions. The horizontal and vertical change rates and fluctuation amplitudes are combined to form a unified 3D posture feature dataset, ultimately generating posture angular velocity change data.

[0061] S112: Call the attitude angular velocity change data to obtain the angular velocity deviation value of each period in the continuous cycle, analyze the data of the horizontal direction, vertical direction and vertical axis change rate, and use the formula:

[0062]

[0063] Calculate the attitude stability state evaluation value;

[0064] Among them, ω xi is the horizontal angular velocity value in the i-th cycle, ω yi is the vertical angular velocity value in the i-th cycle, ω zi is the angular velocity value of the vertical axis in the i-th cycle, is the average value of the horizontal angular velocity in all cycles, is the average value of the vertical angular velocity in all cycles, is the average value of the angular velocity in the vertical axis direction during all cycles, Δt i is the time interval within the i-th cycle, is the rate of change of the vertical axis angular velocity with respect to time in the i-th cycle, i is the sampling cycle index number, dt is the small time increment of the angular velocity change with time, n is the number of sampling cycles, and S is the attitude stability evaluation value;

[0065] Call the attitude angular velocity change data, collect the horizontal and vertical angular velocities of the user's head after wearing smart glasses, and record them periodically to generate a complete angular velocity data sequence. For each group of consecutive period numbers, extract the horizontal ωxi , vertical direction ω yi angular velocity value, and calculate its average angular velocity relative to the corresponding direction in each cycle The square value of the deviation and the rate of change of the vertical axis angular velocity are extracted at the same time The binding period corresponds to the time interval Δt i , performing product absolute value processing, then performing square sum operations and root extraction on the three types of offset data. Combined with the number of sampling periods n and the sum of the absolute values ​​of the average angular velocities in the three directions, performing ratio division, and performing difference processing with 1 as a constant, the final posture stability assessment value S is constructed to reflect the stability of the wearer's head movement. The specific sampling values ​​are shown in Table 1.

[0066] To further illustrate the evaluation value calculation process, we now introduce the sample data in Table 1 and calculate the mean of all angular velocity data:

[0067]

[0068] Then square the offset value for each cycle and accumulate it:

[0069] Cycle 1:

[0070]

[0071] By analogy, the total is 0.00253, which is substituted into the formula:

[0072]

[0073] The posture stability evaluation value is a normalized numerical indicator used to measure whether the three-axis head posture motion is stable within a specific time window while wearing smart glasses. This value ranges from 0 to 1. A value close to 1 indicates that the current head posture has minimal angular velocity fluctuations, minor disturbances, and stable position over multiple consecutive cycles, making it suitable for operations that rely on a static viewpoint, such as image acquisition. A value close to 0 indicates significant head directional shifts, deviations, or sudden changes, which may result in image blur or visual recognition failure. The system will delay image acquisition or issue a recapture command accordingly. This parameter serves as the core control basis for triggering modules such as image acquisition and view lock, ensuring that drug label image data is always acquired from a stable viewpoint, improving subsequent recognition accuracy and the fundamental accuracy of the interactive experience. Results indicate that the posture stability evaluation value is 0.988, which is within the set stability baseline of 0.95 to 1.0, indicating that the current head posture is stable. Posture changes have minimal impact on the control of the image acquisition device, allowing the device to enter the acquisition preparation phase.

[0074] Table 1 shows the periodic data sampling during attitude angular velocity measurement, which is used to provide the source of various parameters in the formula:

[0075] Table 1 Attitude angular velocity period sampling data table

[0076]

[0077] Table 1 lists the angular velocity and time interval monitoring values ​​for five consecutive cycles. All parameters are collected by the sensor device at a preset sampling frequency, providing the complete data source for calculating the attitude stability assessment value. The formula is beneficial in that it integrates the angular velocity change rate and the deviation in three directions to form a fused expression, forming a normalized and quantifiable stability scoring indicator. This provides multi-dimensional input support and a reliable basis for determining the stability of image acquisition by the device.

[0078] S113: adjusting control parameters of an image acquisition device according to the posture stability evaluation value and the stability of the user's head posture, performing image acquisition, and obtaining a package image set;

[0079] According to the posture stability evaluation value, the set stability judgment threshold interval is called and the size judgment operation is performed on the evaluation value. If the posture evaluation value is within the set interval, it is determined to be in a stable state. Then the image acquisition parameter configuration link is entered. According to the current state information, key parameters such as exposure time, focus sensitivity, and frame cache limit of the image acquisition device are set. The exposure time is set to a certain proportion higher than the basic value according to the stability level. The focus sensitivity is weightedly adjusted using the compensation coefficient derived from the current posture change trend. The frame cache is set to the minimum integer multiple within the limit range based on the sampling time and acquisition cycle. After all parameters are configured, they are written to the image acquisition control unit synchronously, and the optical component is called to trigger the pharmaceutical packaging image acquisition action. During the acquisition process, the timestamp and exposure confirmation status of each completed frame are recorded, and the acquisition process is terminated after the set number of frames is reached, and the packaging image set is finally obtained.

[0080] See also Figure 3 , the specific steps for obtaining the structural partitioning results are:

[0081] S211: Calling the packaged image set, extracting the grayscale value of each pixel in the image area, calculating the average grayscale value in each area block, and obtaining the brightness distribution value of the image area;

[0082] Call the packaged image set, divide each frame of the captured image into several small area blocks according to a fixed ratio, and establish a pixel matrix with 10×10 pixels per block as the division standard. Read and summarize the grayscale values ​​of all pixels in each pixel matrix point by point, call all pixel grayscale values ​​in each block for summation, and divide by the total number of pixels to obtain the average grayscale. The grayscale value range is 0 to 255. After processing a single block, a brightness value is obtained, which is recorded as the brightness representative value of the area to form the brightness distribution map of the image. For example, the image is divided into 100 area blocks, and the grayscale mean extracted from each block is a set of numerical values ​​such as

[0083] [120,123,128,…], this set is used to analyze the structural layout of the image. A two-dimensional matrix is ​​then constructed from the brightness values ​​of all regional blocks. The vertical axis represents the image row number and the horizontal axis represents the column number. Each unit corresponds to the average brightness of a block. By comparing the numerical change trends of adjacent units in the matrix, the local brightness and darkness distribution pattern is analyzed, and areas with steep grayscale changes are further extracted as potential boundaries. In addition, an additional edge compensation mechanism is set at the edge of the image. For underexposed or dark areas of the image, the brightness is stretched by the regional grayscale offset correction factor, and the regional grayscale mean is offset to the range of ±20 of the global average brightness of the image, thereby unifying the overall brightness benchmark of the image, making the boundary analysis more representative, and finally obtaining the brightness distribution value of the image area.

[0084] S212: Call the brightness distribution value of the image area, extract the grayscale mutation points of the region boundary pixels, calculate the grayscale mean value of the inner region of each boundary line, the grayscale mean value of the outer region and the corresponding pixel gradient value, and perform normalized gradient structure evaluation on the extracted multiple groups of boundary feature values ​​using the formula:

[0085]

[0086] Calculate the grayscale jump response intensity value of the image boundary, and generate the boundary clarity gradient value according to the regional grayscale mutation direction, gradient density change and boundary response difference;

[0087] Among them, B is the boundary clarity gradient value, m is the number of boundary regions involved in the calculation, and j is the index number of the boundary region. is the average grayscale value of the area inside the j-th boundary line, is the average gray value of the area outside the j-th boundary line, G j is the gradient magnitude of the jth boundary pixel;

[0088] Call the brightness distribution value of the image area, extract the grayscale mutation pixels on the boundary line of the adjacent area, summarize the grayscale values ​​of all areas within 10 pixels inside and outside each boundary line, and calculate the average grayscale of the inside Average grayscale outside Then extract the gray gradient amplitude G of the boundary pixel j , put the above parameters into the following formula to calculate the grayscale jump response intensity value of the image boundary:

[0089]

[0090] The sample data for setting three boundary points is as follows:

[0091] Table 2 Image boundary grayscale and gradient feature table

[0092] Boundary Number Inner grayscale mean Outer grayscale mean Gradient amplitude 1 128 112 30 2 135 120 45 3 142 126 28

[0093] As shown in Table 2, the grayscale and gradient sample data extracted from the three boundary areas during the image analysis process are listed for use in the formula to calculate the boundary clarity gradient value.

[0094]

[0095] Repeat this process to calculate B2≈38.78, B3≈27.62, and average the three results to get:

[0096]

[0097] Among them, the grayscale jump response intensity value of the image boundary is a comprehensive response index that reflects the degree of edge structure jump in the image. It mainly quantifies the response intensity of the visual mutation between the structural area and the background area by jointly evaluating the grayscale distribution mutation amplitude and the gradient change intensity between pixels. The higher the value of this indicator, the more drastic the grayscale jump in the local area, and the greater the gradient of the mutation position, indicating that the area is more likely to be the actual label boundary, and has structural division value in the sense of character barcode recognition. Its main role in the structural division of label images is to serve as the core reference parameter for identifying the boundary line between the character area and the barcode area, and to assist the system in accurately judging the physical demarcation position of the structural area, thereby providing a high-reliability input basis for subsequent semantic recognition and label structure matching. This value falls within the boundary clarity judgment range set by the system.

[0098] [25,40], indicating that the image has a clear boundary structure. This result will serve as the basis for subsequent structural connectivity calculations and ultimately obtain the boundary clarity gradient value.

[0099] S213: Calling the boundary definition gradient value, performing morphological classification and connectivity calculation on the structural boundary lines in the image, identifying the pixel partition boundaries, performing structural label classification on the character blocks and barcode blocks, and establishing a structural partitioning result;

[0100] The boundary clarity gradient value is called to screen the structural strength of all identified boundary lines in the image. First, all boundary lines are classified and marked as "high structural response" or "weak structural response" based on the relative position between the boundary clarity gradient value and the set clarity lower threshold of 25 and upper threshold of 40. After screening out clear boundaries, the two-dimensional arrangement of boundary pixels is analyzed to extract regions with strong connectivity and closed structure to form a connected subgraph. The grayscale symmetry, area, and contour structure of each connected region are then combined to perform preliminary type classification. If there are dense vertical grayscale stripes inside the region with a stable arrangement period, it is classified as a barcode region. If there are dense horizontal grayscale blocks in the region with small grayscale mutations between block units and bending features, it is classified as a character region. All classified regions are marked and numbered, and the region number and type are recorded in the image buffer matrix. Finally, a structure index map is output. This map is used for character recognition path optimization and barcode recognition boundary constraint generation in the subsequent recognition process. A structure annotation file is established in the form of region number and corresponding coordinate point dataset to finally obtain the structural partitioning result.

[0101] See also Figure 4 The specific steps for obtaining the drug-disease correlation value are as follows:

[0102] S311: Invoke the structural segmentation results, extract the drug field in the label image, identify the phrase components in the field that express the drug indication content, invoke the language segmentation component to perform semantic block segmentation and extract part-of-speech attributes, classify the phrase semantic pointing structure based on the context boundary, and obtain the indication semantic phrase set;

[0103] The structural segmentation results are used to extract the text content of the drug field in the label image. In practice, the text block image is first extracted using the horizontal and vertical coordinate bounding boxes of the character regions in the label image based on the structural annotation information. The image is then converted into a recognizable character sequence. After preliminary sentence segmentation, clauses containing the drug name and indication keywords are identified. For each clause, whether a typical indication verb or symptom noun combination is determined, such as "used to relieve headaches" or "indication for pharyngitis," expressions with functional descriptions are retained as phrase components. Next, the language segmentation component performs semantic segmentation on each retained phrase, segmenting the phrase, such as "relieve headaches," into two semantic blocks: "relieve" and "headache," and extracting the corresponding part-of-speech attributes as verbs and nouns. When addressing boundary issues, the context of the preceding and following sentences is considered. For example, if the preceding sentence mentions "this drug is used for," the semantic block is assigned to the "drug efficacy target" and is classified as a symptom classification group. If the context contains "non-indications," the phrase is semantically labeled as negative. In addition, some expressions are nested phrases, such as "has therapeutic effects on bronchitis and allergic cough." "Bronchitis" and "allergic cough" should be segmented and classified into two groups of phrases. After processing, each phrase is recorded and constructed into a set of indication semantic phrases.

[0104] S312: Based on the set of indication semantic phrases, call the disease keyword group in the user description, extract the semantic similarity score, part-of-speech matching consistency value and semantic structure synergy score between each pair of phrases, and calculate the structure synergy strength index in sequence according to the phrase pair index number using the formula:

[0105]

[0106] Calculate the semantic fit value of drug phrases;

[0107] Among them, R is the semantic fit value, L is the total number of phrase pairs, k is the index number of the kth phrase pair, S k Score the semantic similarity between the k-th phrase pair, P k Score the part-of-speech structure matching of the k-th drug field phrase, Q k Score the semantic expression matching of the k-th disease description phrase;

[0108] Based on the acquired indication semantic phrase set, the disease keyword phrase provided by the user is called, and all phrase pairs formed by the two are processed one by one. The semantic similarity score, part-of-speech matching consistency score and semantic structure synergy score of each phrase pair are extracted, which are respectively denoted as S k 、P k , Q kThe score is determined by comparing it with a preset similarity dictionary. For example, the score for "headache" and "migraine" is 0.87, and the score for "fever" and "reduction of fever" is 0.75. Then, an index number k is created for all phrase pairs, and the formula is used for calculation:

[0109]

[0110] Based on the sample data, there are 4 groups of phrase pairs, and the parameter values ​​are shown in Table 1:

[0111] Table 3. Table of items involved in semantic fit calculation

[0112] Phrase pair number <![CDATA[Semantic similarity score S k > <![CDATA[Part-of-speech structure matching score P k > <![CDATA[Semantic expression matching score Q k > 1 0.87 0.85 0.82 2 0.76 0.88 0.79 3 0.65 0.78 0.75 4 0.92 0.80 0.81

[0113] Refer to Table 3 and substitute the formula for the following step-by-step calculation:

[0114]

[0115] Among them, the semantic fit value is an average correlation score used to measure the semantic expression and structural form of the indication phrase in the drug label field and the user-entered symptom phrase. The larger the value, the stronger the semantic consistency and expression coordination between the drug description and the user's symptoms. This value is used to identify whether the labeled drug has semantic adaptability to the user-entered symptoms, and can then be used as a basis for subsequent matching or push logic. The semantic fit value not only supports the precise connection between the label field content and user information, but also enables semantic-guided drug adaptability judgment through the scoring results, thereby improving the accuracy and interaction rationality of drug information identification in personalized human-computer interaction scenarios. The results show that the semantic fit value between the drug label field and the user's symptom description is 1.282, and the higher the value, the stronger the semantic synergy between the two. The benefit of the formula is that by introducing two influencing factors, part-of-speech consistency and structural expression synergy, it enhances the overall recognition ability of the deep semantic structure.

[0116] S313: Calling the semantic fit value of the drug phrase, establishing a correspondence between the drug field and the symptom description, and obtaining the drug-symptom association value;

[0117] The calculated semantic fit values ​​for the drug phrases are called and the results are processed. All semantic fit values ​​are compared against the set semantic association threshold. When the fit value exceeds the threshold, the label field and the symptom phrase are constructed as a valid pair; otherwise, the label field is skipped. Subsequently, all valid field and symptom phrase combinations are grouped and counted according to the label field name, and medication context annotations are added, such as whether it is the primary medication or whether it is a concurrent adjuvant medication. Finally, the symptom reference list is integrated according to the drug field dimension to form a set of structured symptom matching tables, which are used as the output data for the drug-symptom association value.

[0118] See also Figure 5 , the specific steps for obtaining the output information construction record are:

[0119] S411: Calling the drug-disease association value, identifying the user's demand for the target drug, and obtaining the target drug identification value;

[0120] Call the semantic fit value, read the semantic matching strength parameter group between the drug phrase and the symptom phrase obtained by the previous calculation, extract the drug field identifiers whose semantic fit values ​​are greater than the matching judgment threshold in turn, and perform index arrangement operations on the field items according to the ranking value of the semantic fit. Select the drug name field corresponding to the top item in the semantic fit ranking through the structural screening action, and set it as the candidate drug content. According to the context confirmation rules set in the system, classify and judge the continuity and subject consistency of the candidate drug field in the previous and next semantic groups. If the candidate field is repeatedly associated with the same symptom phrase in the continuous segment and the subject structure does not change, then This field is identified as valid target drug content and written into the identification record list. The unique identification index number corresponding to the field is obtained and named as the target drug identification value. In actual operation, for example, the user inputs a description of the disease as "repeated headache and nasal congestion". The system establishes a semantic scoring index list for the drug field "loratadine tablets" with a semantic fit value of 0.76. Because this field is the first match in the two consecutive phrases "headache" and "nasal congestion", the "loratadine tablets" identification value is established as "DRG021-HLTD-01", which is the target drug identification value corresponding to the user's semantic input.

[0121] S412: Based on the target drug identification value, call the pharmaceutical database to extract the drug name, production batch, expiration date, usage and dosage, and precautions of the target drug to obtain a medication information data set;

[0122] First, call the medical database to perform field content retrieval operations, trigger the fuzzy matching retrieval process based on the content of the drug name, extract the first three database records by sorting by matching scores, and compare the letter and number combination extracted from the production batch field in the current image with the above candidate records one by one. The production batch number is filtered with the first 4 digits of the extracted production date field as a prefix, and then the structure is verified according to the structural requirement that the letter part should be two digits and the number part should be six digits. If the field structure does not meet the requirements, the field is invalid and removed from the candidate record. Then, extract the expiration date field, and check whether the field meets the "yyyy-mm-dd" format standard. If not, the data source is considered suspicious and marked as an invalid field. Then, based on the usage and dosage sentence content displayed on the packaging, perform keyword screening operations, segment the text content, and extract "daily", "once", "service" and other related words. "Use" and "dosage" and other keyword groups to determine whether there are missing frequency words and dosage words in the sentence. If the missing items are established, the standard usage records of the drug instructions with the same name are called in the database to fill in the information. At the same time, the precautions text area in the lower area of ​​the label is identified, and the brightness jump feature of the dense grayscale value block area in the image is used to locate the area and read all the text. Then, the sentence segments containing the keywords "prohibited", "use with caution", and "warning" are selected to form a candidate set. Subsequently, the similarity between these candidate segments and the content of the drug standard precautions segments in the database is calculated. The Jaccard similarity index is used to perform a ratio operation of the intersection and union of terms on each candidate segment and the standard segment. If the similarity value is ≥0.85, the segment is written into the precautions field as the final result. Finally, the drug name, production batch, expiration date, usage and dosage, and precautions are combined into the drug information dataset of the drug. S413: Based on the drug information dataset, the AI ​​large model is used to evaluate the user's drug safety for the target drug and output risk warnings, construct the output text of the voice broadcast information, and generate the output information construction record;

[0123] First, extract the content of the drug name field, search for its primary key index number in the medical database, and load the contraindication list, interaction information table and population use restriction items bound to the number. Compare the current disease content provided by the user, first standardize the symptom description into the corresponding ICD-10 standard code. If the user provides a natural language phrase, perform term matching and medical term normalization on the phrase to obtain the corresponding disease number, and then compare it with the code in the drug contraindication list item by item. Use the character edit distance algorithm to perform the comparison, and set the comparison condition to the edit distance d. When d < 2, the match level is considered high, and when they are equal, it is a complete match. When 2≤d<4, it is a moderate match, and d≥4 is irrelevant. Then compare the interaction table in the drug description, extract the name of the drug that needs to be avoided and find its drug code, and perform a complete match with the drug code set of other drugs taken by the user. Yes, if there is an intersection, the combination of drugs is marked as high-risk. Then the user's basic information fields are read, including age, gender, physiological status and other information, and the rules are judged against the contraindications specified in the drug instructions. For example, if the drug stipulates "forbidden for those under 18 years old", when the user's age field value is 15, the prohibited condition is triggered, and the match is recorded as a high-risk event. All the above matching records will be assigned a risk level mark. The rules are as follows: If any of the prohibited conditions matches, it is marked as high risk. If there is a combined conflict or age and other factors require caution, it is marked as medium risk. If it is only a general prompt information, it is marked as low risk, and the final risk level is calculated. The final risk level is determined by the highest level of all event item levels. Finally, the drug name, recommended usage, frequency of use and the above risk level description are integrated to form an output text for voice broadcast, and a unique numbered output information construction record is generated.

[0124] See also Figure 6 The specific steps for obtaining the broadcast delay adjustment parameter value are as follows:

[0125] S511: Call the output information to build a record, monitor the wearer's gaze point position in real time during the voice broadcast process, collect the coordinate value of the gaze point in each frame, record the gaze trajectory displacement path, and obtain the gaze change tracking value;

[0126] The output information is called to build a record, driving the eye tracking module to continuously enable the sight sampling mechanism during the broadcast. The system reads the gaze point coordinate value returned by the eye movement module in each frame sampling period, and generates a time series vector in combination with the system timestamp to build a frame-level gaze point space trajectory matrix. If the current frame image resolution is 1280×720, the system records the horizontal and vertical coordinates of the gaze point according to the time frame number, such as (620,300), (625,303), (640,310), etc., and sorts each group of coordinate points in sequence according to the frame number to form a two-dimensional space path sequence structure, and simultaneously records The Euclidean distance between each two adjacent frames of the gaze point is used to generate an array of gaze displacement lengths, and the movement amplitude and direction vector of each gaze jump are recorded. By setting the time window length, such as a 500ms window length corresponding to 25 frames, continuous gaze position offset trajectories are counted within the window, and the offset lengths and jump times are summarized. Jumps with a jump amplitude exceeding the set jump threshold are marked and written into the gaze record buffer at the same time, forming a gaze change dataset in the frame sequence dimension. Finally, the total movement length and jump frequency of the gaze displacement trajectory within the specified time window are combined to form the gaze change tracking value.

[0127] S512: Identify the fluctuation trend of the gaze point in the spatial and temporal dimensions within a continuous time period based on the gaze change tracking value, calculate the stay stability and gaze concentration intensity during the reading process, and obtain the visual concentration state index value;

[0128] According to the gaze change tracking value, the system first reads the data structure of gaze displacement length and jump number in each time window, calculates the average of the total displacement in each time window, and normalizes the jump number. Then, the mean displacement length is compared with the jump frequency ratio. When the displacement mean is less than the set stillness judgment threshold and the jump frequency is within the preset stable range, it is judged that the gaze state is stable, otherwise it is judged as visual distraction. Further, by constructing the gaze stability mean sequence of multiple groups of time windows within five seconds, the displacement mean change amplitude in each group of sequences is linearly regressed, and the slope value is extracted as the wave Dynamic trend indicator, and at the same time, the variance of the displacement data in each window is averaged to quantify the fluctuation intensity of the gaze state, set the stability index S1 = 1-mean square error / maximum allowable fluctuation, and the concentration intensity index S2 = 1-jump rate / maximum jump frequency, then finally take the weighted average of S1 and S2, and set the weights to 0.5 each, and obtain the visual concentration state index value R = 0.5*S1+0.5*S2. If S1 = 0.83 and S2 = 0.77, then R = 0.5×0.83+0.5×0.77 = 0.80, which represents the degree of visual concentration of the current user in the current time window.

[0129] S513: Call the visual focus state index value, adjust the interval time of the field broadcast according to the change trend of the index value, and obtain the broadcast delay adjustment parameter value;

[0130] The visual concentration state index value is called and compared with the field trigger interval mapping table in the broadcast control module. The mapping table is set as follows: if the visual index R is greater than 0.8, the field broadcast interval is extended by 200ms based on the standard interval. If R is between 0.6 and 0.8, the standard interval is maintained. If R is lower than 0.6, the broadcast interval is shortened by 50ms. In the current sample, the R value is 0.80, which falls into the upper critical interval. The system determines that the field trigger rhythm should be extended and the broadcast start time of the current field content "once a day" is delayed. The delay value is the basic interval of 500ms plus the delay value of 200ms, that is, the trigger rhythm is 700ms. The same calculation process is performed on the next field "Oral after dinner". The rhythm of the voice field output time point is reconstructed one by one to generate a field timestamp table. The corresponding fields, delay values, and trigger times are recorded in the following table.

[0131] Table 4 Field broadcast delay configuration table

[0132]

[0133]

[0134] Referring to Table 4, the trigger interval of each field content required by the voice output module is dynamically adjusted to obtain the broadcast delay adjustment parameter value, and the parameter is output in a table structure for calling the broadcast driver module.

[0135] An AI-based smart glasses drug label recognition system, which is used to execute the above-mentioned AI-based smart glasses drug label recognition method, includes:

[0136] The posture detection module uses an accelerometer to collect the instantaneous angular velocity values ​​of the user's head on the horizontal and vertical axes in real time. By calculating the stability of the change trend and fluctuation amplitude of the two-axis angular velocity over a continuous period of time, the control parameters of the image acquisition device are adjusted to obtain a set of packaging images.

[0137] The image texture analysis module extracts the grayscale brightness distribution data of multiple regions in the image based on the packaged image set. By analyzing the regional brightness gradient change trend and boundary texture clarity, it identifies the texture difference characteristics between the character and barcode areas, divides the character and barcode structure areas, and generates the structure division results.

[0138] Based on the structural segmentation results, the semantic matching module extracts the indication semantic phrases in the label field and compares them with the disease keyword phrases in the user description, including semantic directionality, part-of-speech structure, and expression paradigm. It calculates the semantic matching degree between the phrases and establishes a semantic fit score between the drug field and the disease phrase, generating a drug-disease association value.

[0139] The medication information management module identifies the user's demand for target drugs based on the drug-symptom association value, extracts the dosage, time of administration, and method of administration of the corresponding drug according to the drug name field identified in the tag structure, constructs a drug usage information set, and establishes an output information construction record;

[0140] The broadcast rhythm control module builds records based on the output information, tracks and monitors the changes in the user's gaze position during the broadcast of information in real time, analyzes the gaze dwell time and the gaze point coordinate offset trajectory, evaluates the user's visual concentration state in the reading state, and dynamically adjusts the voice broadcast interval between information fields according to the concentration state, generating a broadcast delay adjustment parameter value.

[0141] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for identifying drug labels on smart glasses based on AI, characterized in that: The following steps are involved: S1: Call the acceleration sensor to analyze the stability of the user's head posture in a continuous time period, adjust the control parameters of the image acquisition device, and obtain the packaging image set; S2: calling the packaged image set, delineating the image structure boundary, and obtaining the structure division result; S3: Calling the structural division result, extracting the indication semantic phrase in the label field, combining it with the symptom keyword group described by the user, establishing a semantic fit score between the drug field and the symptom phrase, and forming a drug-symptom association value; S4: Identify the user's need for the target drug based on the drug-symptom association value, construct the broadcast information output content based on the drug name, and obtain the output information construction record.

2. The AI-based smart glasses drug label recognition method according to claim 1, characterized in that: The packaging image set includes an image frame sequence number, an image clarity identifier, and an acquisition time tag; the structural division result specifically includes a character boundary position set, barcode block coordinates, and background interference area annotations; the drug-symptom association value specifically refers to a semantic overlap value, a part-of-speech structure adaptation coefficient, and a keyword expression distance score; and the output information construction record includes a drug name field, a medication specification field, and a voice output field.

3. The AI-based smart glasses drug label recognition method according to claim 1, characterized in that: The steps of obtaining the package image set are specifically as follows: S111: Calling the acceleration sensor to collect the angular velocity data sequence of the horizontal axis and the vertical axis of the user's head after wearing the smart glasses, extracting the angular velocity change rate and angular velocity fluctuation amplitude in each direction, and obtaining the posture angular velocity change data; S112: Calling the attitude angular velocity change data, obtaining the angular velocity deviation value of each time period in the continuous cycle, analyzing the data of the horizontal direction, vertical direction and vertical axis change rate, and calculating the attitude stability state evaluation value; S113: According to the posture stability state evaluation value and the stability of the user's head posture, the control parameters of the image acquisition device are adjusted to perform image acquisition and obtain a packaging image set.

4. The AI-based smart glasses drug label recognition method according to claim 3, characterized in that: The steps for obtaining the structure division result are specifically as follows: S211: calling the packaged image set, extracting the grayscale value of each pixel in the image area, calculating the average grayscale value in each area block, and obtaining the brightness distribution value of the image area; S212: Calling the brightness distribution value of the image region, extracting the grayscale mutation points of the region boundary pixels, calculating the grayscale mean value of the inner region of each boundary line, the grayscale mean value of the outer region, and the corresponding pixel gradient value, performing normalized gradient structure evaluation on the extracted multiple groups of boundary feature values, calculating the grayscale jump response intensity value of the image boundary, and generating a boundary clarity gradient value based on the regional grayscale mutation direction, gradient density change, and boundary response difference; S213: Call the boundary clarity gradient value, perform morphological classification and connectivity calculation on the structural boundary lines in the image, identify pixel partition boundaries, perform structural label classification on character blocks and barcode blocks, and establish a structural division result.

5. The AI-based smart glasses drug label recognition method according to claim 4, characterized in that: The steps for obtaining the drug-disease correlation value are specifically as follows: S311: Calling the structural division result, extracting the drug field in the label image, identifying the phrase components in the field that express the drug indication content, calling the language segmentation component to perform semantic block division and extract part-of-speech attributes, classifying the phrase semantic pointing structure based on the context boundary, and obtaining a set of indication semantic phrases; S312: Based on the set of indication semantic phrases, call the disease keyword group in the user description, extract the semantic similarity score, part-of-speech matching consistency value and semantic structure synergy score between each pair of phrases, calculate the structure synergy strength index in sequence according to the phrase pair index number, and calculate the semantic fit value of the drug phrase; S313: Call the semantic fit value of the drug phrase, establish a corresponding relationship between the drug field and the disease description, and obtain the drug-disease association value.

6. The AI-based smart glasses drug label recognition method according to claim 5, characterized in that: The steps for obtaining the output information construction record are specifically as follows: S411: calling the drug-symptom association value, identifying the user's need for the target drug, and obtaining the target drug identification value; S412: Based on the target drug identification value, call the pharmaceutical database to extract the drug name, production batch, expiration date, usage and dosage, and precautions of the target drug to obtain a medication information data set; S413: Based on the medication information dataset, using the AI ​​big model, evaluate the user's medication safety of the target drug and output risk warnings, construct the output text of the voice broadcast information, and generate an output information construction record.

7. The AI-based smart glasses drug label recognition method according to claim 1, characterized in that: The method further comprises: S5: Calling the output information to build a record, monitoring the wearer's gaze position changes during voice output in real time, analyzing the gaze dwell time and gaze point coordinate offset trajectory, evaluating the user's visual concentration state during reading, and adjusting the start interval of field announcements to generate an announcement delay adjustment parameter value; The broadcast delay adjustment parameter values ​​specifically include a gaze dwell interval value, a coordinate stability parameter, and a broadcast interval adjustment coefficient.

8. The AI-based smart glasses drug label recognition method according to claim 7, characterized in that: The steps for obtaining the broadcast delay adjustment parameter value are specifically as follows: S511: Call the output information to build a record, monitor the wearer's gaze point position in real time during the voice broadcast process, collect the coordinate value of the gaze point in each frame, record the gaze trajectory displacement path, and obtain the gaze change tracking value; S512: Identifying fluctuation trends of the gaze point in spatial and temporal dimensions within a continuous time period based on the gaze change tracking value, calculating the stability of the gaze and the intensity of the gaze concentration during the reading process, and obtaining a visual concentration state index value; S513: Call the visual concentration state index value, adjust the interval time of field broadcast according to the change trend of the index value, and obtain the broadcast delay adjustment parameter value.

9. An AI-based smart glasses drug label recognition system, characterized in that: The system is used to implement the AI-based smart glasses drug label recognition method according to any one of claims 1 to 8, and the system includes: The posture detection module uses an accelerometer to collect the instantaneous angular velocity values ​​of the user's head on the horizontal and vertical axes in real time. By calculating the stability of the change trend and fluctuation amplitude of the two-axis angular velocity over a continuous period of time, the control parameters of the image acquisition device are adjusted to obtain a set of packaging images. The image texture analysis module extracts grayscale brightness distribution data of multiple regions in the image based on the package image set, identifies the texture difference characteristics between the character and barcode regions by analyzing the regional brightness gradient change trend and boundary texture clarity, divides the character and barcode structure regions, and generates a structure division result; Based on the structural division results, the semantic matching module extracts the indication semantic phrases in the label field and compares and analyzes them with the disease keyword phrases in the user description, including semantic directionality, part-of-speech structure, and expression paradigm. It calculates the semantic matching degree between the phrases and establishes a semantic fit score between the drug field and the disease phrase to generate a drug-disease association value. The medication information management module identifies the user's demand for the target drug based on the drug-symptom association value, extracts the dosage, time of administration, and method of administration of the corresponding drug according to the drug name field identified in the tag structure, constructs a drug usage information set, and creates an output information construction record; The broadcast rhythm control module builds a record based on the output information, tracks and monitors the changes in the user's gaze position during the broadcast of information in real time, analyzes the gaze dwell time and the gaze point coordinate offset trajectory, evaluates the user's visual concentration state in the reading state, and dynamically adjusts the voice broadcast interval between information fields according to the concentration state, and generates a broadcast delay adjustment parameter value.

Citation Information

Cited By

  • Premix packaging label automatic verification method and system based on image recognition

    CN121963237A