Intelligent voice-guided self-service optometry operation system
The intelligent voice-guided self-service optometry operating system enables voice interaction, real-time physiological parameter collection, and personalized guidance, solving the problems of interaction complexity and inaccurate results of existing self-service optometry equipment, and improving the scientific nature and user experience of optometry.
Patent Information
- Application Number
- CN202511462791.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-23
AI Technical Summary
Existing self-service optometry equipment has a complex interaction method, which is difficult to adapt to elderly users and users with visual impairment. It lacks historical data utilization, cannot accurately locate visual abnormalities, and lacks personalized guidance, resulting in inaccurate optometry results and low efficiency.
The self-service optometry operating system, which adopts intelligent voice guidance, includes modules for voice input acquisition, real-time acquisition of optometry data, construction of dynamic optometry models, multi-dimensional difference analysis, and adaptive voice guidance. It generates customized voice commands through voice interaction, real-time physiological parameter acquisition, historical data analysis, and personalized guidance.
It reduces the difficulty of operation, improves the scientific nature and accuracy of optometry results, expands the applicable population, provides personalized guidance, and enhances optometry efficiency and user experience.
Smart Images

Figure CN121370046A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of self-service optometry, in particular to an intelligent voice-guided self-service optometry operation system. BACKGROUND
[0002] Optometry services have a wide and key application in the fields of ophthalmic medical treatment, optometry health care and glasses retail, and the core purpose is to obtain accurate eye vision parameters for users, and then provide a basis for subsequent vision correction scheme formulation. However, the traditional optometry mode highly depends on the operation and judgment of professional optometrists, and the entire process not only requires a very high level of professional skill of optometrists, but also needs to be equipped with complex professional optometry equipment, which makes the cost of traditional optometry services high, and also limits the expansion of its service range. Especially in the scenes of primary medical units, remote areas and small glasses retail stores, the lack of professional optometry resources is particularly prominent, and it is difficult to meet the growing demand for convenient optometry of the public.
[0003] With the rise of the self-service concept, some self-service optometry devices have gradually appeared in the market, trying to break the limitations of the traditional optometry mode. However, the existing self-service optometry devices still have many obvious defects. From the perspective of interaction, most devices rely on users to complete the optometry process through manual operation such as touch screen and key, which is difficult for elderly users, severely visually impaired users and special user groups with difficulty in hand movement, and is easy to affect the smooth progress of the optometry process due to operation errors, and even lead to inaccurate optometry results. In terms of data processing, most existing self-service optometry devices can only complete basic vision data collection and simple calculation, and lack effective use of historical optometry data. In the optometry process, the user's vision state is not constant, and may be influenced by age, eye habits, eye health status and other factors for a long time. Ignoring the reference value of historical data and relying only on single real-time collected data to draw optometry conclusions, it is difficult to fully reflect the real vision trend of the user, resulting in insufficient scientificity and reliability of the optometry results. In the analysis and feedback link of the optometry result, the existing device can usually only output a simple visual value, and cannot accurately locate the possible visual abnormal area. When the user has local visual field problems, uneven refractive errors, etc., the device cannot timely find and feed back the specific abnormal position and abnormal degree to the user, the user is difficult to clearly understand the details of his own eye problems, and it is not conducive to the subsequent formulation of targeted correction scheme. At the same time, the existing device lacks a personalized guiding mechanism, regardless of the user's optometry situation, a fixed guiding process is adopted, the guiding strategy cannot be adjusted according to the real-time optometry data and visual abnormality of the user, it is difficult to help the user better cooperate in the optometry process, which further affects the optometry efficiency and result accuracy. The existence of these problems makes the existing self-service optometry technology difficult to fully meet the user's demand for convenient, accurate and personalized optometry service, and restricts the further development of the self-service optometry field. SUMMARY
[0004] The purpose of the present application is to provide an intelligent voice guided self-service optometry operation system to solve the problems raised in the background art.
[0005] To achieve the above purpose, the present application provides an intelligent voice guided self-service optometry operation system, which comprises: A voice input acquisition module for acquiring user voice instructions and performing signal preprocessing to output digitized voice data; An optometry data real-time acquisition module for synchronously capturing user eye physiological parameters and visual test response sequences to generate a measured optometry data set; A dynamic optometry model construction module for training an optometry state prediction model based on a historical optometry data warehouse to output a visual state prediction value; A multi-dimensional difference analysis module for performing multi-dimensional deviation calculation on the visual state prediction value and the measured optometry data set to generate an optometry difference index matrix; A visual problem positioning module for performing spatial distribution mapping according to the optometry difference index matrix to output a visual abnormal area heat map; An adaptive voice guiding module for generating a customized voice guiding instruction sequence based on the visual abnormal area heat map.
[0006] Preferably, the voice input acquisition module comprises: A voice noise reduction unit for eliminating environmental noise interference and extracting a pure voice signal; A semantic feature analysis unit for performing semantic segmentation processing on the pure voice signal, extracting a key instruction feature vector, and outputting to the dynamic optometry model construction module.
[0007] Preferably, the optometry data real-time acquisition module specifically comprises: physiological parameter capturing unit: real-time measurement of pupil diameter change and corneal curvature waveform; test response recording unit: tracking of visual acuity chart selection response sequence and reaction time distribution; data set generation unit: integration of the pupil diameter change, the corneal curvature waveform, the visual acuity chart selection response sequence and the reaction time distribution, and construction of the measured test light data set.
[0008] Preferably, the dynamic refraction model construction module specifically includes: historical feature mining unit: performing multi-scale feature decomposition on the historical refraction data warehouse to extract steady-state visual features and transient reaction features; prediction model training unit: using the decomposed historical refraction data training set to train an integrated prediction network, the integrated prediction network fuses a time series prediction unit and a feature compensation unit, and dynamically updates the visual state prediction value.
[0009] Preferably, the multi-dimensional difference analysis module specifically includes: time domain cumulative bias calculation unit: aligning the visual state prediction value with the time sequence of the measured test light data set, and calculating the cumulative bias amount in the sliding window; frequency domain feature offset detection unit: decomposing physiological parameter frequency band energy distribution, and quantifying the energy spectrum density ratio of the prediction value and the measured value; response sequence similarity evaluation unit: matching the pattern structure of the visual acuity chart selection response sequence, and calculating the phase error of the sequence mutation point; index matrix generation unit: aggregating the cumulative bias amount, the energy spectrum density ratio and the phase error, and outputting the refraction difference index matrix after normalization.
[0010] Preferably, the visual problem positioning module specifically includes: visual topology modeling unit: constructing a spatial node network based on eye anatomical structure parameters, and labeling regional correlation weights; abnormal propagation simulation unit: mapping the refraction difference index matrix to the spatial node network, and calculating the abnormal propagation path through a graph structure diffusion algorithm; heat map generation unit: counting the abnormal frequency of each node, and generating the visual abnormal area heat map in combination with the regional correlation weights.
[0011] Preferably, the adaptive voice guidance module specifically includes: guidance strategy configuration unit: analyzing the probability distribution of the visual abnormal area heat map, and setting the guidance frequency and content depth; voice instruction generation unit: synthesizing real-time voice feedback sequence according to the guidance frequency and content depth.
[0012] Preferably, the system further comprises a feature map construction module: The feature map construction module performs feature linkage analysis on the measured optometry data set, determines the linkage degree of each visual feature dimension, constructs a visual feature flow map, and outputs to the multi-dimensional difference analysis module for bias calculation.
[0013] Preferably, the system further comprises a core sequence screening module: The core sequence screening module extracts key feature vectors in the visual feature flow map, calculates dimension entropy weight, screens core visual feature sequence, and inputs to the adaptive voice guidance module for instruction sequence generation.
[0014] Preferably, the system further comprises a voice classification processing module: The voice classification processing module receives the digitized voice data, performs part-of-speech sequence matching processing, and classifies into rule instructions and non-rule instructions; Rule instruction translation unit: generating standard optometry operation instructions based on preset rule mapping relationship; Non-rule instruction generation unit: using generative model to synthesize adaptive optometry response.
[0015] Compared with the prior art, the beneficial effects of the present application are: From the perspective of interaction convenience, the system is provided with a voice input acquisition module, and the user does not need to rely on manual operation, but only needs to input each operation instruction in the optometry process through voice instruction. This voice interaction method greatly reduces the threshold of optometry operation. For the elderly, users with severely impaired vision and special groups with difficulty in hand movement, it is easier to participate in the optometry process, avoiding optometry interruption or result deviation caused by manual operation errors, allowing more users to independently and smoothly complete self-service optometry, further expanding the scope of users of self-service optometry services. In terms of comprehensiveness and scientificity of optometry data, the optometry data real-time acquisition module can synchronously capture user eye physiological parameters and vision test response sequence, generate measured optometry data set containing multi-dimensional information, compared with the existing device which only collects basic vision data, it can more comprehensively record each key information in the user optometry process. At the same time, the dynamic optometry model construction module trains the optometry state prediction model based on the historical optometry data warehouse, fully utilizes the user's past vision data, combines single real-time optometry data with historical data, so that the output vision state prediction value can reflect the long-term trend of user's vision, rather than a single instantaneous state, providing more rich and more valuable data support for subsequent optometry result analysis, improving the scientificity and reliability of optometry conclusion. At the level of accurate analysis of the optometry results, the multi-dimensional difference analysis module performs multi-dimensional deviation calculation on the visual state prediction value and the measured optometry data set, generates an optometry difference index matrix, can capture the difference between the prediction value and the actual data in detail, and avoids missing key information that may be missed by traditional simple data comparison. On this basis, the visual problem positioning module performs spatial distribution mapping according to the optometry difference index matrix, and outputs a visual abnormal area heat map. This visual presentation can accurately locate the visual abnormal area of the user, let the user clearly understand the specific location and range of the eye problem, and also provides a more targeted basis for the subsequent formulation of the visual correction scheme, solving the problem that the existing device cannot accurately locate the visual abnormal area. In the aspect of guided personalization, the adaptive voice guidance module generates a customized voice guidance instruction sequence based on the visual abnormal area heat map, which can adjust the guidance content and method according to the specific visual abnormality of the user. For example, for a user with a specific visual abnormality in a certain area, the system can focus on guiding this area during optometry to help the user better cooperate with the relevant tests and ensure the accuracy of the optometry data. At the same time, after optometry, the system can also provide targeted reminders according to the user's abnormality, improving the user's optometry experience and subsequent eye care awareness. This personalized voice guidance method breaks the limitations of the fixed guidance process of existing devices and better adapts to the individual needs of different users, further improving the smoothness of the optometry process and the accuracy of the optometry results, and promoting the development of self-service optometry technology towards a more intelligent and personalized direction. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 The timing diagram of the intelligent voice-guided self-service optometry operation system described in the present application; Figure 2 The flowchart of the optometry data real-time acquisition module processing; Figure 3 The flowchart of the multi-dimensional difference analysis module processing; Figure 4 The flowchart of the visual problem positioning module processing. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0018] Please refer to Figure 1The application provides an intelligent voice-guided self-optometry operation system, which comprises a voice input collection module, an optometry data real-time collection module, a dynamic optometry model construction module, a multi-dimensional difference analysis module, a visual problem positioning module and an adaptive voice guidance module.
[0019] The voice input collection module acquires user voice instructions and converts them into digital voice data through signal preprocessing. The optometry data real-time collection module synchronously captures user eye physiological parameters and visual test response sequences to generate a measured optometry data set. The dynamic optometry model construction module trains an optometry state prediction model based on a historical optometry data warehouse and outputs visual state prediction values. The multi-dimensional difference analysis module performs multi-dimensional deviation calculation on the visual state prediction values and the measured optometry data set to generate an optometry difference index matrix. The visual problem positioning module performs spatial distribution mapping according to the optometry difference index matrix to output a visual abnormal area heat map. The adaptive voice guidance module generates a customized voice guidance instruction sequence based on the visual abnormal area heat map.
[0020] The system automates the refraction process through modular collaboration. The voice input acquisition module uses a microphone array to capture voice signals, and signal preprocessing includes analog-to-digital conversion and sampling quantization, outputting a digital audio stream. The real-time refraction data acquisition module integrates an eye tracker and a corneal topographer to record pupil diameter changes and corneal curvature waveforms in real time. Simultaneously, it records the user's response sequence and reaction time to the visual acuity chart via input devices. The dataset generation unit aligns multi-source data by timestamp and stores it as a structured dataset. The dynamic refraction model construction module connects to a historical database containing user refraction records. The feature mining unit uses wavelet transform to extract steady-state features such as baseline pupil size and transient features such as peak reaction time. The prediction model training unit uses a long short-term memory network and a convolutional neural network to build an integrated prediction network. The time-series prediction unit processes time-series data, and the feature compensation unit corrects for environmental interference. The model outputs dynamically updated prediction values. The multidimensional difference analysis module uses a time-domain cumulative deviation calculation unit to align predicted and measured time series, calculating the deviation integral within a window. The frequency-domain feature offset detection unit applies Fourier transform to analyze the energy distribution of physiological parameters and calculates spectral density differences. The response sequence similarity evaluation unit uses a dynamic time warping algorithm to match response sequence patterns and detect phase shifts. The index matrix generation unit aggregates multidimensional results and generates a standard matrix through min-max normalization. The visual problem localization module's visual topology modeling unit constructs network nodes based on ocular anatomy data, with nodes corresponding to retinal regions and association weights set based on neural density. The anomaly propagation simulation unit maps the index matrix to the network, uses graph convolutional networks to simulate anomaly propagation, calculates node influence paths, and the heatmap generation unit statistically analyzes node anomaly frequencies and generates color-coded heatmaps based on weights. The adaptive voice guidance module's guidance strategy configuration unit analyzes the heatmap probability distribution, sets the frequency parameters and content depth levels of the voice output, and the voice command generation unit uses text-to-speech synthesis technology to generate real-time voice sequences according to the strategy. The system runs on an embedded platform, with each module communicating via a bus to achieve low-latency processing.
[0021] Example 1: See Figure 2 The speech noise reduction unit of the speech input acquisition module employs an adaptive filtering-based noise cancellation method. Its core algorithm constructs an adaptive filter to estimate environmental noise and subtracts this estimate from the original input signal. This unit uses two microphones: one for acquiring the user's speech signal and the other specifically for acquiring an environmental noise reference signal. The filter coefficients are iteratively updated using a least mean square algorithm to minimize the mean square error of the output signal. This process can be expressed by the following formula: in: Indicates the first The error signal at each sampling time, i.e., the output after noise reduction; is the original mixed signal collected by the main microphone; is the first filter coefficient at the time t; is the value of the first filter coefficient at the time t; is the sampling value of the noise signal collected by the reference microphone at the time t; is the sampling value of the noise signal collected by the reference microphone at the time t; is the sampling value of the noise signal collected by the reference microphone at the time t; is the order of the filter, which determines the memory length and computational complexity of the filter. The algorithm runs in real-time on a digital signal processor as the hardware platform, with a sampling frequency set to 16 kHz. The filter order is dynamically adjusted according to the acoustic environment, typically ranging from 128 to 256 orders. After processing by this unit, the signal-to-noise ratio of the speech signal is significantly improved, providing high-quality input for semantic analysis.
[0022] The semantic feature analysis unit further processes the noise-reduced pure speech signal. This unit performs pre-emphasis processing by a first-order high-pass filter to boost high-frequency components and compensate for the attenuation of high-frequency parts during speech propagation. Subsequently, frame processing is performed to divide continuous speech signals into overlapping frames, with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. A Hamming window is applied to each frame of speech signal to reduce spectral leakage, and a fast Fourier transform is performed to convert it to the frequency domain. Mel-frequency cepstral coefficients are used to extract the feature vector of the speech, which includes calculating the power spectrum, filtering through a Mel filter bank, taking the logarithm, and performing a discrete cosine transform to obtain a 12-dimensional MFCC feature vector and its first and second order differences, forming a 39-dimensional feature vector sequence. Hidden Markov models are used to model and recognize these feature sequences. The model uses a left-to-right topology, containing 5 states and 3 Gaussian mixture models. The model training uses the Baum-Welch algorithm, and the recognition process uses the Viterbi algorithm to find the optimal state sequence. This unit can recognize a set of predefined voice commands, such as "start test", "select next line", "repeat last time", etc., and convert the recognized commands into fixed-dimensional feature vectors.
[0023] The physiological parameter capture unit of the real-time optometry data acquisition module uses a non-contact infrared eye tracker to measure the pupil diameter. This device is equipped with a high-speed infrared camera with a sampling frequency of 60 Hz and a uniform infrared light source to ensure clear imaging of the pupil edge. The collected video stream is analyzed through image processing algorithms, which are grayscale and Gaussian filtered to reduce noise, use the Canny edge detection algorithm to locate the pupil contour, and calculate the diameter of the pupil through the ellipse fitting algorithm. The output is a sequence of diameter values in pixels, which is converted to physical dimensions through a pre-calibrated conversion coefficient. The measurement of corneal curvature is completed by a corneal keratometer, which projects a set of concentric circular patterns onto the cornea and calculates the curvature radius by analyzing the shape of the reflected image. The analysis is based on the geometric optics principle of reflected light, which calculates the curvature by measuring the interval and deformation of the ring, and outputs the waveform data in the form of a sequence of curvature values changing with time.
[0024] The test response recording unit is integrated with a high-resolution touch display screen, which presents a standard logarithmic visual acuity chart to the user. The visual acuity chart symbols are dynamically presented according to the test procedure, usually starting from larger symbols and gradually becoming smaller. The user indicates the direction or content of the symbol perceived by touching the corresponding area on the screen. The unit accurately records the content of each user response, such as correctly identifying the opening direction of the "E" character, or incorrectly identifying the symbol. At the same time, a high-precision system timer is used to record the time interval between the presentation of the symbol and the user's response, with an accuracy of milliseconds. All response events and their corresponding time stamps are recorded completely.
[0025] The data set generation unit is responsible for the temporal and spatial alignment and fusion of multi-modal data from different sensors. This unit maintains a global timestamp counter, and all input data streams are labeled with hardware-generated time markers. The data fusion algorithm first resamples all data streams to a uniform time grid with a sampling interval of 16.7 milliseconds (corresponding to a sampling rate of 60 Hz). For each time point, the unit creates a data packet containing the pupil diameter measurement value, the corneal curvature measurement value, the currently displayed visual acuity chart symbol information, the user response status (no response, correct response or incorrect response) and the reaction time (if a response event occurs). These data packets are organized into a structured array or table form, which constitutes the complete measured optometry data set, and are temporarily stored in the cache for real-time reading and processing by the module. The entire data acquisition and generation process guarantees the temporal synchronization and integrity of the data, providing a reliable data basis for analysis and decision-making.
[0026] Example 2: see Figure 3The historical feature mining unit of the dynamic refraction model construction module accesses a relational database that stores a large number of historical refraction records. These records cover the complete refraction process data of users of different ages and different vision conditions, and each record contains time-series pupil diameter measurement values, corneal curvature readings, user response sequences to the vision chart, and accurate reaction time stamps. The feature mining unit preprocesses these high-dimensional time-series data, including outlier removal and data smoothing, and applies a multi-scale feature decomposition algorithm. This algorithm uses a set of pre-defined time window scales, such as 500ms, 1000ms, and 2000ms windows, to analyze the statistical properties of the data at different time granularities. For each scale, the algorithm calculates statistical quantities such as mean, variance, skewness, and kurtosis to describe the steady-state characteristics of the data, such as the baseline level of pupil size and its stability. At the same time, the algorithm also detects transient changes in the data, such as rapid contraction or expansion of the pupil diameter, and sudden prolongation of the user's reaction time. These transient features are captured by calculating the first and second order differences of the signal, and their amplitude and duration are recorded. All extracted steady-state and transient features are encoded into a structured feature vector, the dimension of which is proportional to the number of selected feature types and scales.
[0027] The prediction model training unit receives the structured feature vector output by the historical feature mining unit and uses it to train an integrated prediction network. This network consists of two main sub-modules: the time series prediction unit and the feature compensation unit. The time series prediction unit adopts a deep recurrent neural network architecture, specifically a long short-term memory network design, which contains multiple LSTM layers, each with a number of memory cells, capable of learning complex temporal dependencies in historical refraction data. The input to the network is a sequence of feature vectors within a past time window, and the output is a prediction of the vision state at future time points, such as predicting the user's ability to recognize smaller visual targets next time. The feature compensation unit is a feedforward neural network that identifies and corrects confounding factors that may affect prediction accuracy, such as accidental fluctuations in ambient light intensity or the user's immediate distraction. This unit takes current environmental sensor readings and other contextual information as additional input, and outputs a compensation vector to adjust the output of the time series prediction unit. The outputs of the two sub-networks are combined through a weighted fusion layer to produce the final, environment-corrected vision state prediction value. The training process of the entire integrated network uses the backpropagation algorithm and the stochastic gradient descent optimizer, and the loss function considers both the absolute value of the prediction error and the temporal smoothness of the error.
[0028] The time-domain cumulative bias computation unit of the multidimensional discrepancy analysis module receives the visual state prediction sequence from the dynamic refraction model construction module and the measured refraction data set sequence provided by the refraction data real-time acquisition module. The primary task of this unit is to accurately align the two sequences, which are of different origins and may have an initial time offset. It uses the dynamic time warping algorithm to compute the optimal matching path between the two sequences, thereby eliminating any systematic time delay. After the sequence alignment, the unit operates within a preset length of sliding time window, for example, set to 5 seconds. For each time point within the window, it computes the absolute difference between the predicted value and the measured value, and accumulates all these differences within the entire window, thereby obtaining a scalar value representing the total bias in the prediction during this time period - the cumulative bias amount. This calculation process is repeated as the window slides, resulting in a new sequence describing the change in bias over time.
[0029] The frequency-domain feature shift detection unit analyzes the difference between the predicted value and the measured value from the frequency perspective. This unit performs spectral analysis on the physiological parameter time series, such as the pupil diameter change curve. It applies the fast Fourier transform to convert the time-domain signal to the frequency domain, obtaining the power spectral density estimate of the signal. By comparing the power spectra of the predicted sequence and the measured sequence, the unit calculates the energy distribution difference of the two sequences in the main frequency band (e.g., the low frequency band corresponds to slow physiological regulation, and the high frequency band corresponds to instantaneous response). Specifically, it calculates the energy ratio of the two sequences in the same frequency band, and uses this ratio as a quantitative indicator of the frequency-domain feature shift. The response sequence similarity evaluation unit focuses on analyzing user behavior data. It compares the predicted response sequence (how the model thinks the user should respond) and the actual response sequence (how the user actually responds). This unit not only checks the correctness of the response content, but also focuses on analyzing the structural features of the response pattern, such as the occurrence pattern of consecutive incorrect responses, the interval rules between correct responses, etc. It uses sequence matching algorithms to identify common patterns and mutation points in the two sequences, and calculates the phase difference of these mutation points in time, i.e., whether the pattern change in one sequence is ahead of or lagging behind the other sequence.
[0030] The index matrix generating unit plays the role of an aggregator, which receives the output data from the three units of time domain, frequency domain and behavior sequence analysis: the time-varying cumulative deviation sequence, the energy spectrum density ratio of each frequency band, and the phase error metric of the response sequence. These data have different dimensions and scales. The unit respectively performs minimum-maximum normalization processing on each input data stream, linearly transforming it to a common numerical range of 0 to 1. It integrates all the normalized data by time point to construct a two-dimensional matrix structure. In this refraction difference index matrix, each row corresponds to a time point, and each column represents a specific type of difference metric (time domain deviation, frequency domain ratio of a certain frequency band, phase error, etc.). This matrix provides a comprehensive, standardized, multi-angle quantitative description of the differences between the prediction and the actual measurement, providing a data basis for problem positioning and decision guidance.
[0031] Example 3: refer to Figure 4 The visual topology modeling unit of the visual problem positioning module constructs a spatial node network representing the eye structure based on standard ophthalmic anatomy data. The network contains a series of nodes, each corresponding to a functional area on the retina, such as the fovea, parafoveal area, and more peripheral retinal partitions. The number of nodes and their spatial positions are pre-defined according to the physiological structure of the retina, usually described in polar coordinates to match the curved surface characteristics of the eyeball. Each node is assigned an associated weight, which reflects the relative importance of the area in visual function; the weight of the central visual area is significantly higher than that of the peripheral area, and its value is assigned according to the density of the neural ganglion cells and the clinical visual importance scale. The connection edges between nodes are established according to the anatomical adjacency relationship between retinal areas and the direction of nerve fiber bundles, forming a weighted undirected graph structure that fully describes the spatial topological relationship of the eye.
[0032] The anomaly propagation simulation unit receives the refraction difference index matrix generated by the multi-dimensional difference analysis module and maps the difference information contained in the matrix to the constructed spatial node network. The mapping process first aggregates the difference values at different time steps and different dimensions in the matrix to obtain a comprehensive abnormal intensity value through weighted averaging, and uses it as the initial abnormal state value of the network node at the corresponding time point or test condition. The unit starts a graph-based anomaly propagation simulation process. The propagation process follows a defined dynamics rule, and the current abnormal state of a node is affected by the state of its neighbor nodes at the previous time step. Its state update can be described by the following formula: Where: represents the abnormal state value of node at simulation time step t. It is a decay factor between 0 and 1, used to control the persistence of initial anomalous information; It is mapped to a node The initial comprehensive anomaly intensity value; Represents nodes in the network The set of all directly connected neighboring nodes; It is a connection node and nodes The weight of the edge reflects the strength of the interaction between the two regions; These are neighbor nodes. In the previous simulation time step The abnormal state value is calculated iteratively. Through iterative calculation, the abnormal state spreads throughout the network, and eventually each node obtains a stable abnormal value that reflects the global propagation effect.
[0033] The heatmap generation unit performs statistical and visualization processing on the final anomalous state values of each node after propagation simulation. This unit first normalizes the anomalous values of all nodes, transforming them into a uniform scaling range, such as 0 to 1, where 0 represents complete normality and 1 represents the highest degree of anomalousness. The normalized node anomalous values are combined with their corresponding region association weights, typically using multiplication, to amplify the anomalous values of high-weight regions (such as the fovea) in the final output, thus highlighting their importance. The combined value is defined as the final heatmap value for that visual region. These heatmap values are mapped to a predefined color coding scheme, with low values corresponding to cool tones (such as blue), high values to warm tones (such as red), and intermediate values representing gradual color transitions. Finally, based on the spatial coordinates of all nodes and the calculated color values, a heatmap of visual anomalous regions covering the entire retinal area is generated. This heatmap is output as a two-dimensional image, intuitively displaying the inferred spatial distribution of visual anomalousness and its relative severity.
[0034] The guidance strategy configuration unit of the adaptive voice guidance module analyzes the received visual anomaly region heat map to determine the specific parameters of the voice guidance. This unit calculates the pixel proportion of different color regions in the heat map and analyzes the probability distribution characteristics of the abnormal values, such as calculating the mean, variance, and higher-order moments of the abnormal values. Based on these statistical characteristics, the unit automatically sets two key operation parameters: guidance frequency and content depth. Guidance frequency determines the time interval and triggering conditions for the system to issue voice prompts, such as triggering a prompt when detecting that the abnormal value of a specific region has lasted beyond a threshold. Content depth defines the detail level and professionalism of the voice instructions. For highly abnormal or key visual area-related conditions, the system generates more detailed and instructive explanations and operation instructions; for minor or secondary abnormalities, the instructions are relatively concise. These strategy parameters are encapsulated into a configuration protocol.
[0035] The voice instruction generation unit is the execution terminal of the guidance strategy. It dynamically synthesizes real-time voice feedback sequences based on the frequency and depth parameters output by the guidance strategy configuration unit, combined with the specific stage of the current optometry process (such as testing left eye distance vision). The unit accesses a pre-recorded voice segment database containing audio files of various optometry-related instructions, prompts, and feedback sentences. The synthesis process is based on a decision logic that selects the most appropriate voice segments from the database according to the strategy parameters and real-time context, and performs splicing and combination. For non-pre-recorded content or dynamically generated information (such as specific numerical values), the unit integrates a text-to-speech synthesis engine that can convert structured text information into clear and natural voice audio in real time. The final customized voice guidance instruction sequence is played to the user through the system's audio output device, guiding the user to perform the next operation or providing necessary feedback, thereby completing the entire adaptive guidance closed loop based on visual problem positioning.
[0036] In embodiment 4, the feature map construction module receives the measured optometry data set from the optometry data real-time acquisition module. This data set records multiple-dimensional parameters such as pupil diameter, corneal curvature, visual acuity chart response type, and reaction time in time series form. The module first performs standardization preprocessing on these parameters to eliminate the influence of different dimensions, so that all feature values are within the same numerical range. Data standardization uses the minimum-maximum normalization method to linearly convert the original value of each feature dimension to the 0-1 interval. The preprocessed data is sent to the feature correlation analysis engine, which calculates the Pearson correlation coefficient between each pair of feature dimensions to evaluate the linear correlation strength between them. The larger the absolute value of the correlation coefficient, the stronger the correlation between the two feature dimensions. The analysis results are recorded in a feature correlation matrix, where the rows and columns represent the feature dimensions, and the matrix elements store the correlation coefficient values of the corresponding feature pairs.
[0037] Referring to Table 1, a simplified example of the feature correlation matrix is shown, containing correlation coefficients between five main optometric feature dimensions. In a complete system, the dimension of the matrix will be higher, containing all the monitored feature parameters. The values in the matrix reveal the intrinsic relationships between features, for example, the pupil diameter change shows a moderate negative correlation with the reaction time, while the corneal curvature fluctuation shows a weak positive correlation with the visual chart response type. These quantified correlation information provides the data basis for constructing the visual feature flow graph.
[0038] Table 1: Example of correlation coefficient matrix of optometric feature dimensions.
[0039] Feature dimensions Pupil diameter Corneal curvature Optotype response Reaction time Blink rate Pupil diameter 1.00 -0.12 -0.31 -0.45 0.28 Corneal curvature -0.12 1.00 0.18 0.09 -0.15 Optotype response -0.31 0.18 1.00 0.52 -0.22 Reaction time -0.45 0.09 0.52 1.00 -0.18 Blink rate 0.28 -0.15 -0.22 -0.18 1.00 Based on the feature correlation matrix, the module constructs the visual feature flow graph. The graph adopts the weighted directed graph structure in graph theory, where each node represents an optometric feature dimension, such as the average pupil diameter or the correct response rate. The edges between nodes represent the correlation between feature dimensions, and the direction of the edge is determined by the time sequence, for example, the edge from the reaction time node to the visual chart response node indicates that the change in reaction time will affect the response accuracy. The weight of the edge is taken from the absolute value of the correlation coefficient of the corresponding feature in the feature correlation matrix, and after threshold processing, only the significant correlation connections are retained. The graph construction algorithm automatically identifies and removes edges with weights below a preset threshold, usually set to 0.3, to maintain the simplicity and interpretability of the graph. The final generated graph not only contains the static correlation information between features, but also reflects the time sequence dynamics of feature influence in the optometric process through the directionality of the edges.
[0040] The core sequence screening module conducts in-depth analysis on the constructed visual feature flow graph and extracts the most informative feature subset. The module first calculates the centrality index of each node, including degree centrality, closeness centrality, and betweenness centrality, to evaluate the importance of each feature dimension in the graph from different angles. A node with high degree centrality means that the feature has strong correlations with many other features; a feature with high closeness centrality indicates that it is in a key position in the information transmission path; a feature with high betweenness centrality acts as a bridge between different feature groups. The module integrates these centrality indices to calculate a composite importance score for each feature dimension. The score calculation considers the relative weights of different centrality indices, and degree centrality is usually given a higher weight because it directly reflects the connectivity of the feature.
[0041] On the basis of the importance score, the module further analyzes the information entropy of each feature dimension. Information entropy quantifies the uncertainty and information content of the feature value distribution, and the higher the entropy value, the richer the information contained in the feature. The calculation uses the standard Shannon entropy formula based on the probability distribution of the feature value in the historical data. The final core feature selection criterion combines the importance score of the node and the information entropy value to produce a comprehensive weight for each feature through a linear weighting formula. The module sorts all features according to the comprehensive weight and selects the top 30% of the features as the core visual feature sequence. These selected core features usually include reaction time, correct response rate, pupil diameter change amplitude, and other dimensions that have a key impact on the results of the refraction.
[0042] The selected core visual feature sequence is sent to the adaptive voice guidance module to guide it to generate more targeted voice instructions. The features in the sequence are prioritized according to their comprehensive weights, with higher-weighted features having greater influence in voice guidance decision-making. For example, when the reaction time feature is detected to be abnormal, the system will preferentially generate voice prompts related to adjusting the test rhythm; while when the pupil diameter change is abnormal, it may trigger suggestions related to adjusting the lighting conditions. The feature sequence also contains association information between features, which enables the system to generate composite guidance strategies that take into account the interaction between multiple features. For example, when both reaction time extension and correct rate decline are detected, the system will integrate the abnormality degree and mutual relationship of the two features to generate voice instructions suggesting that the user rest or re-adjust their position. The entire process realizes a complete closed loop from raw refraction data to feature map construction, core feature selection, and finally guiding the development of voice guidance strategies.
[0043] In embodiment 5, the voice classification processing module receives the digitized voice data stream from the voice input acquisition module, which is a 16-bit pulse code modulation format audio after noise reduction and preprocessing, with a sampling rate of 16 kHz. The module first performs endpoint detection on the continuous voice stream, identifies the segments containing valid voice signals, and removes the silent segments. The detection uses a double-threshold algorithm based on short-time energy and zero-crossing rate, which can effectively distinguish between voice and background noise. The identified valid voice segments are sent to the part-of-speech sequence matching processing unit for analysis and processing. This unit integrates grammar parsing tools in natural language processing, and first performs automatic speech recognition conversion on the input voice segments, converting the audio signal into a text string. The conversion process uses an acoustic model and a language model based on deep neural networks, which can adapt to different pronunciation habits and accent variants. The converted text is processed by word segmentation, which is decomposed into a series of lexical units. Each lexical unit is assigned a part-of-speech tag, such as noun, verb, adjective, or adverb, etc. These tags are assigned based on pre-defined grammar rules and dictionaries. The entire sentence thus is represented as a part-of-speech tag sequence.
[0044] Based on the generated POS tag sequence, the module performs an instruction classification operation. The classification algorithm analyzes the pattern and structure of the POS sequence and matches it against predefined rule templates. The rule templates define the instruction syntax structure that conforms to the standard optometry operation flow, such as the "verb + noun" structure phrase (e.g. "start test", "switch lens"), or the interrogative sentence containing specific keywords (e.g. "is this clear?"). Instructions that can successfully match these rule templates are classified as regular instructions, which usually express an explicit operation intent and conform to the expected interaction logic. Instructions that cannot match any rule template are classified as irregular instructions, which may contain incomplete sentences, ambiguous expressions, or query requests beyond the preset range (e.g. "I feel the left side is a bit blurry" or "Can you make it a bit brighter?").
[0045] For the classified regular instructions, the regular instruction translation unit is responsible for converting them into standard optometry operation instructions executable by the system. This unit maintains a rule mapping table, which establishes the mapping relationship from natural language phrases to machine instructions. This mapping table contains common optometry operation commands and their corresponding parameter settings, such as "start test" mapped to the command to initialize the optometry flow, and "adjust right" mapped to the lens parameter adjustment command with a direction parameter. The translation process includes parsing the keywords and parameters in the input instruction, querying the mapping table to find the corresponding machine instruction template, and filling the extracted parameters into the corresponding positions in the template to generate a complete control instruction sequence. These standardized instructions are sent to the corresponding hardware control module through the system bus to drive the optometry equipment to perform specific operations.
[0046] The irregular instruction generation unit handles those user inputs that do not conform to the preset syntax structure. It employs a deep learning-based generative model, specifically a sequence-to-sequence neural network architecture, whose encoder encodes the input irregular instruction into a latent semantic vector, and whose decoder generates a structured system response based on the vector. The model is pre-trained on a large amount of optometry dialogue data, learning how to understand ambiguous user expressions and generate appropriate optometry-related responses. The generation process not only considers the semantic content of the instruction itself, but also incorporates the current optometry context information, such as the type of test being conducted, the data results already collected, etc. For example, when the user says "it looks a bit dark", the unit may generate the response instruction "suggest increasing the ambient light intensity" in combination with the current visual acuity chart test phase being conducted; while when the user says "I am not sure about this direction", the unit may generate the operation instruction "re-display the current optotype" or "record as an uncertain response" in combination with the size of the current test optotype. The generated response instruction is converted into a voice signal and played to the user through the system's audio output device, and may also trigger the corresponding device control operation. The voice classification processing flow realizes the understanding and response generation of user voice input, while supporting both structured standard instructions and flexible non-standard expressions. The rule instruction translation ensures quick and accurate response to explicit operation instructions, maintaining the efficiency of the optometry process; the irregular instruction generation provides understanding and feedback for user's ambiguous expressions and additional queries, enhancing the interactive flexibility and user experience of the system. The cooperative work of the two processing paths enables the system to adapt to different users' expression habits and interaction needs, providing natural and smooth voice interaction support for the self-service optometry process.
[0047] It is to be understood that the terminology used herein such as first and second, and the like, merely intend to differentiate one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0048] While the embodiments of the present application have been illustrated and described, it will be understood by those skilled in the art that various changes, modifications, alternatives, and variations can be made therein without departing from the spirit and scope of the application as defined by the appended claims and their equivalents.
Claims
1. A smart voice-guided self-service optometry operating system, characterized in that, include: Voice input acquisition module: used to acquire user voice commands, perform signal preprocessing, and output digital voice data; Real-time optometry data acquisition module: synchronously captures the user's ocular physiological parameters and visual acuity test response sequence, and generates an actual optometry dataset; Dynamic refraction model construction module: Trains a refraction status prediction model based on a historical refraction data warehouse and outputs a visual acuity status prediction value; Multidimensional difference analysis module: Performs multidimensional deviation calculation between the predicted visual acuity status value and the actual refraction dataset to generate a refraction difference index matrix; Visual problem localization module: Performs spatial distribution mapping based on the refraction difference index matrix and outputs a heat map of visual abnormality areas; Adaptive voice guidance module: Generates a customized voice guidance instruction sequence based on the heatmap of the visual anomaly region.
2. The intelligent voice-guided self-service optometry operating system according to claim 1, characterized in that, The voice input acquisition module includes: Voice noise reduction unit: Eliminates environmental noise interference and extracts pure voice signals; Semantic feature parsing unit: performs semantic segmentation processing on the pure speech signal, extracts key instruction feature vectors, and outputs them to the dynamic optometry model construction module.
3. The intelligent voice-guided self-service optometry operating system according to claim 1, characterized in that, The real-time optometry data acquisition module specifically includes: Physiological parameter capture unit: Real-time measurement of pupil diameter changes and corneal curvature waveform; Test response recording unit: tracks the visual acuity chart selection response sequence and reaction time distribution; Data set generation unit: integrates the pupil diameter change, the corneal curvature waveform, the visual acuity chart selection response sequence, and the reaction time distribution to construct the experimental light dataset.
4. The intelligent voice-guided self-service optometry operating system according to claim 1, characterized in that, The dynamic refraction model construction module specifically includes: Historical feature mining unit: Performs multi-scale feature decomposition on the historical optometry data warehouse to extract steady-state visual features and transient response features; Prediction model training unit: The integrated prediction network is trained using the decomposed historical refraction data. The integrated prediction network integrates the time series prediction unit and the feature compensation unit to dynamically update the visual state prediction value.
5. The intelligent voice-guided self-service optometry operating system according to claim 4, characterized in that, The multidimensional difference analysis module specifically includes: Temporal cumulative deviation calculation unit: Aligns the predicted visual state value with the temporal sequence of the actual test light dataset, and calculates the cumulative deviation within the sliding window; Frequency domain feature shift detection unit: decomposes the energy distribution of physiological parameters in frequency bands and quantifies the ratio of the energy spectral density of predicted values to measured values; Response sequence similarity evaluation unit: Matches the pattern structure of the selected response sequence to the visual acuity chart and calculates the phase error of the abrupt change point between sequences; Index matrix generation unit: aggregates the cumulative deviation, the energy spectral density ratio and the phase error, normalizes them and outputs the refraction difference index matrix.
6. The intelligent voice-guided self-service optometry operating system according to claim 5, characterized in that, The visual problem localization module specifically includes: Visual topology modeling unit: Construct a spatial node network based on eye anatomical parameters and label the region association weights; Anomaly propagation simulation unit: maps the refraction difference index matrix to the spatial node network and calculates the anomaly propagation path using a graph structure diffusion algorithm; Heatmap generation unit: Counts the frequency of anomalies at each node and generates a heatmap of the visually abnormal region by combining the regional association weights.
7. The intelligent voice-guided self-service optometry operating system according to claim 6, characterized in that, The adaptive voice guidance module specifically includes: Guidance strategy configuration unit: Analyzes the probability distribution of the heatmap of the visual anomaly area and sets the guidance frequency and content depth; Voice command generation unit: synthesizes a real-time voice feedback sequence based on the guidance frequency and content depth.
8. The intelligent voice-guided self-service optometry operating system according to claim 1, characterized in that, It also includes a feature map construction module: The feature map construction module performs feature link analysis on the measured light dataset to determine the connectivity of each visual feature dimension, constructs a visual feature flowchart, and outputs it to the multidimensional difference analysis module for deviation calculation.
9. The intelligent voice-guided self-service optometry operating system according to claim 8, characterized in that, It also includes a core sequence filtering module: The core sequence filtering module extracts key feature vectors from the visual feature flowchart, calculates dimensional entropy weights, filters core visual feature sequences, and inputs them into the adaptive voice guidance module for instruction sequence generation.
10. The intelligent voice-guided self-service optometry operating system according to claim 9, characterized in that, It also includes a voice classification and processing module: The speech classification processing module receives the digitized speech data, performs part-of-speech sequence matching processing, and classifies it into regular instructions and non-regular instructions. Rule-based instruction translation unit: Generates standard optometry operation instructions based on preset rule mapping relationships; Irregular instruction generation unit: Uses generative models to synthesize adaptive refraction responses.