Multi-age doctor-seeing appointment registration matching method and multi-age doctor-seeing appointment registration matching system

By collecting and analyzing user voice data through voice input devices, building a voiceprint database and performing identity confirmation, the problem of inconvenience in medical appointment registration for the elderly and minors is solved, and efficient and safe registration services for multiple age groups are realized.

CN120656671APending Publication Date: 2025-09-16CHENGDU WENJIANG MEDICAL CLOUD INTERNET HOSPITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510848976.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing medical appointment registration method is inconvenient for the elderly and minors, which reduces their medical efficiency, and telephone registration increases the manual processing burden and cost.

Method used

The user's test voice data is collected through voice input devices, voiceprint features are extracted, a voiceprint database is built, and the identity is confirmed through telephone data comparison and registration is automatically matched.

Benefits of technology

It improves registration efficiency, reduces manual operation errors, provides a convenient service experience, especially for the elderly, and enhances the security and user-friendliness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656671A_ABST
    Figure CN120656671A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital data processing, in particular to a multi-age doctor-seeing appointment registration matching method and system, and the method comprises the steps: obtaining test sound data obtained by a user through a voice input device; voiceprint features in the test sound data are extracted; associating the voiceprint features with the identity information of the user to construct a voiceprint database; obtaining telephone data of a user, comparing the telephone data in the voiceprint database to obtain target identity information, and performing voice playing on the target identity information for confirmation; and after the identity information is confirmed, collecting historical registration information of the user and registration information in the telephone data, and carrying out matching registration. Therefore, the registration efficiency is improved, errors caused by manual operation are reduced, meanwhile, the requirements of users of different age groups are also considered, and particularly, a more friendly and convenient service is provided for the elderly who are not good at using a traditional digital input mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital data processing, and in particular to a method and system for matching medical appointments for multiple age groups. Background Art

[0002] Medical appointment registration is a convenient and fast healthcare service that allows patients to select their preferred hospital, department, and doctor in advance, as well as confirm a specific appointment time. This reduces waiting times and improves healthcare efficiency. Appointments can be made in a variety of ways, including by phone, online, or through self-service devices. This approach not only helps hospitals better manage patient flow and rationally allocate medical resources, but also allows patients to choose the most appropriate appointment time based on their schedule, significantly enhancing their healthcare experience.

[0003] However, the existing appointment registration method generally uses online registration, but the online registration process is relatively complicated and not convenient for the elderly and minors, thereby reducing the medical efficiency of these people. Using telephone registration will increase the manual processing burden and increase costs. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for matching medical appointments for multiple age groups, aiming to improve registration efficiency, reduce errors caused by manual operations, and also take into account the needs of users of different age groups.

[0005] To achieve the above-mentioned object, in a first aspect, the present invention provides a method for matching medical appointments for multiple age groups, comprising obtaining test voice data obtained by a user through a voice input device, wherein the test voice data is a voice text of the user's personal information;

[0006] Extracting voiceprint features from the test sound data;

[0007] Associating the voiceprint features with the user's identity information to construct a voiceprint database;

[0008] Obtain the user's phone data and compare it with the voiceprint database to obtain the target identity information, and then play the target identity information by voice for confirmation;

[0009] After the identity information is confirmed, the user's historical registration information and the registration information in the telephone data are collected and matched.

[0010] The specific steps of obtaining the test sound data obtained by the user through the voice input device include:

[0011] Activate the voice input device and collect test sound signals in the environment in real time;

[0012] Performing noise reduction processing on the collected test sound signal;

[0013] The digital sound signal after the noise reduction process is formatted to obtain test sound data.

[0014] The specific steps of extracting the voiceprint features from the test sound data include:

[0015] Split the test sound data into short time segments of 20-40 milliseconds each and use a Hamming window to reduce spectral leakage;

[0016] Mel-frequency cepstral coefficients are used to extract acoustic feature sets that can represent the speaker from each frame;

[0017] The acoustic feature set is converted into a feature vector of fixed length.

[0018] The specific steps of extracting the acoustic feature group that can represent the speaker from each frame using the Mel-frequency cepstral coefficients include:

[0019] Perform fast Fourier transform on each segment of test sound data to convert the time domain signal into the frequency domain;

[0020] Convert the linear frequency axis to the Mel frequency axis;

[0021] Set a set of triangular filters to cover the entire spectrum range to obtain a Mel filter bank;

[0022] Pass the fast Fourier transform result through the above-mentioned Mel filter bank, calculate the sum of the energy output of each filter, and form the Mel spectrum;

[0023] Take the logarithm of each filter output in the Mel spectrum to simulate the nonlinear characteristics of the human ear's perception of sound intensity;

[0024] Perform discrete cosine transform on the logarithmized Mel spectrum to generate MFCC coefficients, and retain the first several coefficients to obtain the acoustic feature group.

[0025] The specific steps of associating the voiceprint features with the user's identity information to construct a voiceprint database include:

[0026] Collect user identity information, including name, ID number, and contact information;

[0027] Create a mapping that associates each user's voiceprint feature vector with its corresponding identity information;

[0028] The associated voiceprint feature vector and the corresponding identity information are stored.

[0029] The specific steps of obtaining the user's phone data, comparing it with the voiceprint database to obtain the target identity information, and playing the target identity information by voice for confirmation include:

[0030] Recording sound signals in real time while the user is speaking;

[0031] The received sound signal is subjected to noise reduction, frame division and windowing to obtain an audio clip;

[0032] Extracting a voiceprint feature vector from the audio clip and comparing it with records in a pre-built voiceprint database to calculate cosine similarity;

[0033] Filter target identity information based on set thresholds;

[0034] Confirm the matched target identity information.

[0035] The specific steps of extracting the voiceprint feature vector from the audio clip and comparing it with the records in the pre-built voiceprint database to calculate the cosine similarity include:

[0036] Normalize the voiceprint feature vector;

[0037] Traverse the entire voiceprint database and obtain the voiceprint feature vectors of all registered users

[0038] The cosine similarity between each candidate voiceprint feature vector and the newly input voiceprint feature vector is calculated.

[0039] The specific steps of confirming the matched identity information include:

[0040] After determining the matching target identity information, a confirmation message containing the information is generated;

[0041] Use text-to-speech technology to convert confirmation messages into voice messages;

[0042] Play the generated voice message to the user and ask whether it is correct;

[0043] Obtain user feedback to confirm the accuracy of identity information.

[0044] In a second aspect, the present invention also provides a multi-age medical appointment registration matching system, including a test data acquisition module, a voiceprint feature extraction module, a database construction module, an identity comparison module, and a registration module:

[0045] The test data acquisition module is used to acquire test voice data obtained by the user through a voice input device, wherein the test voice data is the user's personal information voice text;

[0046] The voiceprint feature extraction module is used to extract the voiceprint features in the test sound data;

[0047] The database construction module is used to associate the voiceprint features with the user's identity information to construct a voiceprint database;

[0048] The identity comparison module is used to obtain the user's phone data and compare it with the voiceprint database to obtain the target identity information, and play the target identity information by voice for confirmation;

[0049] The registration module is used to collect the user's historical registration information and the registration information in the phone data after the identity information is confirmed and match the registration.

[0050] The present invention provides a multi-age medical appointment matching method and system. The system collects user test voice data through a voice input device. This data contains the user's personal information, such as name, age, and contact information, and is obtained as text via voice input. The system then extracts voiceprint features from the collected test voice data. Voiceprint features are unique biometric features, similar to fingerprints; each person's voice is unique, making them a reliable verification method. The extracted voiceprint features are associated with the user's identity information (such as ID number, medical insurance card number, etc.) to construct a voiceprint database. This database is securely stored and used for subsequent user authentication. When a user attempts to make an appointment, the system obtains the user's phone data and compares it with the constructed voiceprint database to find the corresponding target identity information. To ensure accuracy, the system plays the target identity information through voice for the user to confirm. Once the user confirms their identity information is correct, the system begins collecting the user's historical registration information and registration-related information from the current phone data. Based on this information, the system automatically completes the registration matching process and arranges a suitable medical service time for the user. This method not only improves registration efficiency and reduces errors caused by manual operation, but also takes into account the needs of users of different age groups, especially those who are less adept at using traditional numeric input methods, providing a more user-friendly and convenient service experience. In addition, the use of voiceprint recognition technology also enhances system security and protects users' personal information from being leaked. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 This is a flow chart of a method for matching medical appointments for multiple age groups according to the present invention.

[0053] Figure 2 This is a flow chart of the present invention for obtaining test sound data obtained by a user through a voice input device.

[0054] Figure 3 It is a flow chart of extracting voiceprint features from the test sound data of the present invention.

[0055] Figure 4 This is a flow chart of the present invention for extracting an acoustic feature group that can represent a speaker from each frame using Mel-frequency cepstral coefficients.

[0056] Figure 5 This is a flow chart of the present invention for associating the voiceprint features with the user's identity information to construct a voiceprint database.

[0057] Figure 6 This is a flowchart of the present invention for obtaining user's phone data, comparing it with the voiceprint database to obtain target identity information, and playing the target identity information by voice for confirmation.

[0058] Figure 7 This is a flow chart of the present invention for extracting voiceprint feature vectors from the audio clip and comparing them with records in a pre-built voiceprint database to calculate cosine similarity.

[0059] Figure 8 This is a flow chart of the present invention for confirming the matched identity information. DETAILED DESCRIPTION

[0060] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0061] First embodiment

[0062] See also Figures 1 to 8 The present invention provides a method for matching medical appointments for multiple age groups, comprising:

[0063] S101 obtains test voice data obtained by the user through a voice input device, wherein the test voice data is the user's personal information voice text;

[0064] The specific steps include:

[0065] S201 activates the voice input device and collects the test sound signal in the environment in real time;

[0066] Make sure that the microphone or other voice input device used is properly connected and its drivers and related settings are configured. This step involves hardware-level checks and software configuration. Prompt the user to use a specific method (such as pressing a button, saying a wake-up word, or other preset actions) to activate the system to enter the voice input preparation state. Doing so ensures that the recording function is only turned on when needed, which helps protect user privacy. Once the system is activated, it will begin to collect sound signals in the environment in real time through the voice input device. This process involves converting sound waves into electrical signals and further digitizing them for computer processing.

[0067] S202 performs noise reduction processing on the collected test sound signal;

[0068] First, the collected sound signal undergoes preliminary noise reduction processing, including background noise removal and echo cancellation techniques, to improve the accuracy of subsequent processing. Common noise reduction methods include frequency domain filtering and adaptive filtering. The continuous sound signal is segmented into small segments (usually tens of milliseconds), each of which is called a frame. A window function (such as a Hamming window) is applied to each frame. This reduces spectral leakage and ensures more accurate frequency analysis.

[0069] S203 formats the digital sound signal after the noise reduction processing to obtain test sound data.

[0070] Useful acoustic features are extracted from each frame of the sound signal, such as Mel-Frequency Cepstral Coefficients (MFCCs), spectrograms, or other time-frequency representations. These features can capture the frequency and time information in the audio signal and are the basis for subsequent processing. The extracted acoustic features are organized in a certain format to form a standardized data structure. For example, the MFCC values ​​of each frame can be arranged in sequence to form a matrix or vector for easy storage and subsequent processing. Depending on actual needs, you can choose whether to save the original sound signal data. This is very useful for future analysis or as a dataset for training models. If you choose to save, you must pay attention to complying with relevant privacy protection regulations to ensure the security of the data.

[0071] Finally, the formatted test sound data is checked for completeness and accuracy to ensure that no important information is lost and that all data meets the expected standards.

[0072] S102 extracts voiceprint features from the test sound data;

[0073] The specific steps include:

[0074] S301 divides the test sound data into short time segments, each of which is 20-40 milliseconds, and uses a Hamming window to reduce spectrum leakage;

[0075] Continuous test sound data is segmented into multiple short time segments (frames), with each frame typically ranging from 20 to 40 milliseconds. This segmentation helps capture transient characteristics in the speech signal. To reduce spectral leakage caused by discontinuities at frame boundaries, a Hamming window function is applied to each frame. This is a commonly used weighted window that smooths frame edges, making frequency analysis more accurate.

[0076] S302 uses Mel-frequency cepstral coefficients to extract an acoustic feature group that can represent the speaker from each frame;

[0077] The specific steps include:

[0078] S401 performs fast Fourier transform on each segment of test sound data to convert the time domain signal into the frequency domain;

[0079] A Fast Fourier Transform (FFT) is performed on each frame of the Hamming windowed sound signal to convert the original time domain signal into a frequency domain representation. This allows us to obtain the energy distribution of different frequency components within the frame, providing a basis for further processing.

[0080] S402 converts the linear frequency axis into a Mel frequency axis;

[0081] According to the characteristics of the human auditory system, the Mel scale formula M(f) = 2595log 10 (1 + f / 700) converts the linear frequency axis to the Mel frequency axis. This step reflects the fact that the human ear has different sensitivities to different frequencies, especially in the low-frequency range.

[0082] S403 sets a set of triangular filters to cover the entire spectrum range to obtain a Mel filter bank;

[0083] A set of triangular filters is designed based on the Mel frequency axis. These filters cover the entire spectrum, but are denser in the low-frequency region and gradually become sparser in the high-frequency region. This design can better simulate the response of the human auditory system to different frequencies.

[0084] S404 passes the fast Fourier transform result through the above-mentioned Mel filter bank, calculates the sum of the energy output by each filter, and forms a Mel spectrum;

[0085] The FFT result is passed through a designed Mel filter bank and the sum of the energy output from each filter is calculated. This compresses the original spectrum information and converts it into a form more suitable for subsequent processing, namely the Mel spectrum.

[0086] S405 takes the logarithm of each filter output in the Mel spectrum to simulate the nonlinear characteristics of the human ear's perception of sound intensity;

[0087] Taking the logarithm of each filter output value in the Mel-spectrogram simulates the nonlinear characteristics of the human ear's perception of sound intensity. This operation helps to enhance the representation of weak signals and makes the subsequent feature extraction process more stable.

[0088] S406 performs discrete cosine transform on the logarithmized Mel spectrum to generate MFCC coefficients, and retains the first several coefficients to obtain an acoustic feature group.

[0089] The Mel-frequency cepstral coefficients (MFCCs) are generated by performing a discrete cosine transform (DCT) on the logarithmized Mel-frequency spectrum. Typically, only the first few coefficients (e.g., 12-13) are retained because they contain most of the information about the speaker's characteristics, while the subsequent coefficients reflect more detailed information and contain noise.

[0090] S303 converts the acoustic feature group into a feature vector of fixed length.

[0091] The extracted acoustic feature group is normalized (e.g., z-score normalization) to ensure that the mean of each dimension is 0 and the variance is 1. This step helps to improve the accuracy of similarity calculation.

[0092] If necessary, the acoustic feature set can be converted into a fixed-length feature vector using statistical methods such as averaging or maximizing, or by using a deep learning model to directly output a fixed-size vector. This step is necessary because it allows audio clips of different lengths to be compared and processed consistently. Finally, check that the converted feature vector fully and accurately reflects the main features of the original audio signal, ensuring that no key information has been lost.

[0093] S103: Associating the voiceprint features with the user's identity information to construct a voiceprint database;

[0094] The specific steps include:

[0095] S501 collects user identity information, including name, ID number, and contact information;

[0096] Before collecting any personal information, we must first obtain the user's explicit consent and provide the user with a clear privacy policy statement informing them how their data will be used, stored and protected.

[0097] Collect users' personally identifiable information through secure and reliable means (such as online forms or in-person registration). This information typically includes, but is not limited to, the user's name, ID number, and contact information (such as phone number or email address). Ensure that the information collected is accurate and complies with relevant laws and regulations, such as GDPR or other local data protection regulations.

[0098] S502 creates a mapping to associate each user's voiceprint feature vector with its corresponding identity information;

[0099] Choose an appropriate data structure to represent voiceprint features and their corresponding identity information. This can be a data structure consisting of two main parts: one for storing the voiceprint feature vector and the other for storing the user's identity information. For example, you could use a key-value pair, where the key is the user's unique identifier (such as an ID number) and the value contains the user's voiceprint feature vector and personal identity information.

[0100] For each user whose voiceprint feature vector is successfully extracted, a record is created, linking their voiceprint feature vector with the corresponding identity information. This step involves writing scripts or programs to automate large amounts of data entry work, ensuring the accuracy and consistency of each record.

[0101] After the mapping is completed, some samples are randomly selected for inspection to ensure that each user's voiceprint feature vector is correctly associated with its corresponding identity information to avoid false matching.

[0102] S503 stores the associated voiceprint feature vector and the corresponding identity information.

[0103] Select an appropriate database management system (DBMS) based on the needs of the application scenario. For sensitive data such as voiceprint features and identity information, you can choose a relational database (such as MySQL, PostgreSQL) or a non-relational database (such as MongoDB). Consider using encrypted storage technology to enhance security. Design the database table structure, which should include at least two tables - one for storing voiceprint feature vectors (including other auxiliary information such as collection time, etc.) and the other for storing user identity information. A connection can be established between these two tables through foreign keys or other mechanisms to facilitate query and management. Store the associated voiceprint feature vectors and identity information in the database. During this process, ensure that all data is properly encrypted to prevent unauthorized access. At the same time, set strict access control to allow only authorized personnel and services to access sensitive data.

[0104] S104 obtains the user's phone data and compares it with the voiceprint database to obtain the target identity information, and plays the target identity information by voice for confirmation;

[0105] The specific steps include:

[0106] S601 records the sound signal in real time while the user is speaking;

[0107] Establish a stable connection with the user through your phone system or VoIP service. Ensure good communication quality, minimizing background noise and other distractions. Prompt the user to begin speaking (for example, "Please state your name and ID number") and record the user's voice input in real time. Use high-quality recording equipment or software to ensure audio quality.

[0108] S602 performs noise reduction, frame division and windowing on the received sound signal to obtain an audio clip;

[0109] Perform preliminary noise reduction on the received sound signal to remove unnecessary background noise and improve the accuracy of subsequent feature extraction. Adaptive filters or other advanced noise reduction techniques can be used. The continuous sound signal is divided into multiple short time segments (frames), each typically 20 to 40 milliseconds long. A window function (such as a Hamming window) is applied to each frame to reduce spectral leakage, making frequency analysis more accurate. Based on actual needs, the quality of the audio segment can be further optimized, such as by adjusting the frame length, frame shift, and selecting different window functions to meet the needs of different scenarios.

[0110] S603 extracts a voiceprint feature vector from the audio clip and compares it with records in a pre-built voiceprint database to calculate cosine similarity;

[0111] The specific steps include:

[0112] S701 performs normalization processing on the voiceprint feature vector;

[0113] Normalize the voiceprint feature vectors extracted from the audio clip. This step usually uses the z-score normalization method, which is to subtract the mean from each feature value and divide it by its standard deviation. The formula is as follows:

[0114]

[0115] Where X is the original feature value, μ is the mean value of the feature, and σ is the standard deviation of the feature. This can eliminate the dimensional differences between different features and make the features have the same scale, thereby improving the accuracy of subsequent similarity calculations.

[0116] Check whether the normalized feature vector meets expectations, for example, the mean of each dimension is close to 0 and the variance is close to 1, to ensure that the data processing process does not introduce bias or lose important information.

[0117] S702 traverses the entire voiceprint database to obtain the voiceprint feature vectors of all registered users;

[0118] Ensure that appropriate permissions and security measures are in place to access the voiceprint database. This involves steps such as identity authentication and encrypted transmission to protect user privacy and data security. Based on the system design, choose an appropriate method to traverse the entire voiceprint database and obtain the voiceprint feature vectors of all registered users. This can be achieved by writing query scripts or using database management tools. For large-scale databases, consider optimizing query performance, such as using indexes to speed up the search process. If the database size is moderate and hardware resources permit, all relevant voiceprint feature vectors can be loaded into memory at once for fast access and comparison. Otherwise, data can be loaded in batches as needed.

[0119] S703 calculates the cosine similarity between each candidate voiceprint feature vector and the newly input voiceprint feature vector.

[0120] Implement a function to calculate the cosine similarity between two voiceprint feature vectors. Cosine similarity measures the similarity between two vector directions. The calculation formula is:

[0121]

[0122] Where A·B represents the dot product of two vectors, ‖A‖ and ‖B‖ are the Euclidean norms (i.e., the lengths of the vectors) of vectors A and B, respectively. The cosine similarity results range from -1 to +1, with +1 indicating the exact same direction. The closer to +1, the higher the similarity.

[0123] For each candidate voiceprint feature vector, call the similarity calculation function above and compare it with the newly input voiceprint feature vector to obtain a similarity score. Considering efficiency issues, multiple comparison tasks can be processed in parallel, especially when processing large amounts of data. Sort all calculated similarity scores from high to low to find the closest match. Set a reasonable similarity threshold to determine whether the two are considered a match. If the highest score is lower than the set threshold, it means that not enough matches were found. Record the match with the highest score and its related information (such as user ID, name, etc.) as the basis for final identity recognition.

[0124] S604 filters the target identity information based on the set threshold;

[0125] To ensure accurate identification of the user's identity, the target identity information is screened based on the calculated cosine similarity score compared with a preset threshold. The following are detailed steps: Based on the actual application scenario and test results, set a reasonable similarity threshold. This threshold determines the minimum similarity required between two voiceprint feature vectors for it to be considered a successful match. Typically, this value is determined through a large amount of experimental data and statistical analysis. For all calculated similarity scores, only those results with scores above the set threshold are retained as potential target identity information. If no score exceeds the threshold, it is considered that no suitable match has been found, and the user will be prompted to provide a new voice sample or take other measures.

[0126] S605 confirms the matched target identity information.

[0127] The specific steps include:

[0128] S801 generates a confirmation message containing the matching target identity information after determining the matching target identity information;

[0129] Once the best match and its corresponding identity information (such as name, ID number, etc.) are determined, generate a clear confirmation message. For example, "We have identified you as Zhang San, ID number 123456789012345678." This message should be clear and easy for users to understand and respond to. Depending on the specific situation, the confirmation message can be further customized, such as adding additional information or asking the user if they need any changes.

[0130] S802 converts the confirmation information into a voice message using text-to-speech technology;

[0131] Select a high-quality text-to-speech (TTS) engine to convert the generated text confirmation message into a spoken audio format. There are many mature TTS solutions available on the market, such as Google Text-to-Speech and Amazon Polly. Adjust the TTS engine settings as needed to achieve a more natural and smooth voice. This includes selecting different voice models and adjusting parameters such as speech rate and pitch to ensure that the voice message is easy to understand and sounds comfortable.

[0132] S803 plays the generated voice message to the user and asks whether it is correct;

[0133] Play the generated voice confirmation message to the user via the phone line or VoIP service. Immediately after playing the confirmation message, ask the user to confirm that the information is correct. For example, "If you are Zhang San, please press 1 to confirm. If not, please press 2 to re-enter the information."

[0134] S804 obtains user feedback to confirm the accuracy of the identity information.

[0135] Monitor the user's keystrokes or other forms of feedback (such as voice responses). Based on the user's response, determine their level of acceptance of the current identity information. If the user confirms the information is correct, proceed to the next step. If it is incorrect or the user requests re-entry, return to the initial step and recollect the voice sample. Consider various abnormal situations (such as signal loss, user non-response, etc.) and develop corresponding contingency plans. For example, you can repeatedly request user feedback within a certain period of time or provide manual customer service support to help resolve the problem.

[0136] S105 After the identity information is confirmed, the user's historical registration information and the registration information in the telephone data are collected and matched.

[0137] Specifically, after receiving confirmation feedback from the user (such as a key selection or voice response), the system will double-check whether the identity information recognized by the system is completely consistent with the user's actual information to ensure there are no errors or misunderstandings.

[0138] Based on the confirmed user's identity information, the hospital's information system or dedicated registration database is used to retrieve all of the user's historical registration records. These records include, but are not limited to, visit date, department name, doctor name, and diagnosis results. The obtained historical registration information is collated, removing duplicates to ensure data accuracy and completeness. If necessary, this information can be arranged chronologically or categorized by department, doctor, etc. for easier analysis and use.

[0139] The user's appointment request information is then extracted from the phone data. This includes the department, doctor, and desired appointment time. Automatic speech recognition (ASR) technology can be used to automatically transcribe the user's speech into text, from which the required information can be extracted.

[0140] The extracted registration request information is formatted to meet the system's input requirements. For example, the verbally described time is converted to a standard time format, and the vague department name is refined.

[0141] Based on the user's historical registration information, the system analyzes their registration patterns and preferences. For example, the user frequently registers with a certain department or doctor, the time period they typically choose (morning / afternoon), and whether they have a predisposition to specific diseases. Using this analysis and current registration request information, combined with the hospital's actual resource availability (such as doctor schedules and department hours), the system provides users with personalized registration recommendations. This step uses machine learning algorithms to improve matching accuracy and prioritize recommendations that best match the user's preferences.

[0142] Generate one or more appointment suggestions based on the matching results and present them to the user in an easy-to-understand format. This information can be presented through a voice prompt on the phone or displayed in the mobile app interface. Allow users to select their preferred appointment option and provide the opportunity to modify or add special requests (such as requesting an earlier appointment time). Ensure the entire process is user-friendly and intuitive, making it easy for users to make decisions.

[0143] Once the user selects a satisfactory appointment option, the system should reconfirm all details and officially complete the registration process. For example, "You have successfully booked an appointment with Dr. Zhang San for the cardiology clinic at 9:00 AM on February 15, 2025. Please press 1 to confirm." A confirmation message, including the specific appointment time, location, and other relevant notes, is sent to the user via SMS, email, or in-app notification.

[0144] Second embodiment

[0145] The present invention provides a multi-age medical appointment registration matching system, comprising a test data acquisition module, a voiceprint feature extraction module, a database construction module, an identity comparison module and a registration module: the test data acquisition module is used to obtain test voice data obtained by a user through a voice input device, wherein the test voice data is a voice text of the user's personal information; the voiceprint feature extraction module is used to extract voiceprint features in the test voice data; the database construction module is used to associate the voiceprint features with the user's identity information to construct a voiceprint database; the identity comparison module is used to obtain the user's telephone data and compare it in the voiceprint database to obtain target identity information, and play the target identity information by voice for confirmation; the registration module is used to collect the user's historical registration information and the registration information in the telephone data after the identity information is confirmed, and perform matching registration.

[0146] In this embodiment, the test data acquisition module is used to acquire test voice data obtained by the user through a voice input device, and the test voice data is the user's personal information voice text.

[0147] Activate the voice input device and collect test sound signals in the environment in real time. Then perform preliminary noise reduction on the received sound signals to remove unnecessary background noise. Format the digital sound signals after noise reduction to generate standardized test sound data.

[0148] The voiceprint feature extraction module is used to extract a group of acoustic features that can represent the speaker, namely, voiceprint features, from the test sound data.

[0149] The test sound data is divided into short time segments (20-40 milliseconds each) and a Hamming window is used to reduce spectral leakage. The voiceprint features are extracted from each frame using Mel-frequency cepstral coefficients (MFCCs) or other feature extraction methods. The time domain signal is converted to the frequency domain. The linear frequency axis is converted to the Mel-frequency axis. A set of triangular filters is set to cover the entire spectrum range, and the energy sum of each filter output is calculated to form a Mel-spectrum. The logarithm of each filter output in the Mel-spectrum is taken and a DCT is performed to generate the MFCC coefficients. The extracted voiceprint features are converted into a fixed-length feature vector.

[0150] The database construction module is used to associate the voiceprint features with the user's identity information to construct a voiceprint database. Specifically, it collects basic information such as the user's name, ID number, and contact information. A suitable data structure is designed to associate each user's voiceprint feature vector with its corresponding identity information. An appropriate database management system (such as MySQL or MongoDB) is selected to store the associated voiceprint feature vector and identity information to ensure data security and integrity.

[0151] The identity matching module is used to obtain the user's phone data, compare it with the voiceprint database to obtain the target identity information, and then play the target identity information for confirmation. Specifically, the module records the user's voice signal in real time as they speak. The received voice signal is subjected to noise reduction, frame division, and windowing to generate an audio clip. The voiceprint feature vector is extracted from the audio clip and compared with the records in the pre-built voiceprint database, calculating the cosine similarity. The target identity information is filtered based on a set threshold, and a confirmation message containing this information is generated. Using text-to-speech technology, this message is converted into a voice message and played back to the user, asking if it is correct. The accuracy of the identity information is confirmed based on the user's keystrokes or voice response.

[0152] The registration module is used to collect the user's historical registration information and the registration information in the telephone data after the identity information is confirmed, and match the registration. Specifically, the user's historical registration records are obtained from the hospital information system and are sorted and analyzed. The user's registration request information is automatically transcribed and formatted using voice recognition technology (ASR). Based on the user's historical registration pattern and personal preferences, as well as the current registration request information, a machine learning algorithm is used to provide personalized registration suggestions. The user is allowed to select a satisfactory registration option and a confirmation notification is sent to formally complete the registration process. The user's registration history is updated in a timely manner to continuously optimize system performance and service quality.

[0153] Through the design and implementation of the above modules, this system can effectively use voiceprint recognition technology to provide users with convenient and secure medical appointment registration services, while improving the overall efficiency of medical services and user experience.

[0154] The above disclosure is only a preferred embodiment of the present invention, and certainly cannot be used to limit the scope of the rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A method for matching medical appointments for multiple age groups. It is characterized by: The method includes: obtaining test voice data obtained by a user through a voice input device, wherein the test voice data is a voice text of the user's personal information; Extracting voiceprint features from the test sound data; Associating the voiceprint features with the user's identity information to construct a voiceprint database; Obtain the user's phone data and compare it with the voiceprint database to obtain the target identity information, and then play the target identity information by voice for confirmation; After the identity information is confirmed, the user's historical registration information and the registration information in the telephone data are collected and matched.

2. A method for matching medical appointments for multiple age groups as claimed in claim 1, characterized in that: The specific steps of obtaining the test sound data obtained by the user through the voice input device include: Activate the voice input device and collect test sound signals in the environment in real time; Performing noise reduction processing on the collected test sound signal; The digital sound signal after the noise reduction process is formatted to obtain test sound data.

3. The method for matching medical appointments for multiple age groups as claimed in claim 2, characterized in that: The specific steps of extracting the voiceprint features from the test sound data include: Split the test sound data into short time segments of 20-40 milliseconds each and use a Hamming window to reduce spectral leakage; Mel-frequency cepstral coefficients are used to extract acoustic feature sets that can represent the speaker from each frame; The acoustic feature set is converted into a feature vector of fixed length.

4. The method for matching medical appointments for multiple age groups as claimed in claim 3, characterized in that: The specific steps of extracting the acoustic feature group that can represent the speaker from each frame using the Mel-frequency cepstral coefficients include: Perform fast Fourier transform on each segment of test sound data to convert the time domain signal into the frequency domain; Convert the linear frequency axis to the Mel frequency axis; Set a set of triangular filters to cover the entire spectrum range to obtain a Mel filter bank; Pass the fast Fourier transform result through the above-mentioned Mel filter bank, calculate the sum of the energy output of each filter, and form the Mel spectrum; Take the logarithm of each filter output in the Mel spectrum to simulate the nonlinear characteristics of the human ear's perception of sound intensity; Perform discrete cosine transform on the logarithmized Mel spectrum to generate MFCC coefficients, and retain the first several coefficients to obtain the acoustic feature group.

5. The method for matching medical appointments for multiple age groups as claimed in claim 4, characterized in that: The specific steps of associating the voiceprint features with the user's identity information to construct a voiceprint database include: Collect user identity information, including name, ID number, and contact information; Create a mapping that associates each user's voiceprint feature vector with its corresponding identity information; The associated voiceprint feature vector and the corresponding identity information are stored.

6. A method for matching medical appointments for multiple age groups as claimed in claim 5, characterized in that: The specific steps of obtaining the user's phone data, comparing it with the voiceprint database to obtain the target identity information, and playing the target identity information by voice for confirmation include: Recording sound signals in real time while the user is speaking; The received sound signal is subjected to noise reduction, frame division and windowing to obtain an audio clip; Extracting a voiceprint feature vector from the audio clip and comparing it with records in a pre-built voiceprint database to calculate cosine similarity; Filter target identity information based on set thresholds; Confirm the matched target identity information.

7. A method for matching medical appointments for multiple age groups as claimed in claim 6, characterized in that: The specific steps of extracting the voiceprint feature vector from the audio clip and comparing it with the records in the pre-built voiceprint database to calculate the cosine similarity include: Normalize the voiceprint feature vector; Traverse the entire voiceprint database and obtain the voiceprint feature vectors of all registered users; The cosine similarity between each candidate voiceprint feature vector and the newly input voiceprint feature vector is calculated.

8. The method for matching medical appointments for multiple age groups according to claim 7, wherein: The specific steps of confirming the matched identity information include: After determining the matching target identity information, a confirmation message containing the information is generated; Use text-to-speech technology to convert confirmation messages into voice messages; Play the generated voice message to the user and ask whether it is correct; Obtain user feedback to confirm the accuracy of identity information.

9. A multi-age medical appointment registration matching system, applied to a multi-age medical appointment registration matching method according to any one of claims 1 to 8, characterized in that: It includes test data acquisition module, voiceprint feature extraction module, database construction module, identity comparison module and registration module: The test data acquisition module is used to acquire test voice data obtained by the user through a voice input device, wherein the test voice data is the user's personal information voice text; The voiceprint feature extraction module is used to extract the voiceprint features in the test sound data; The database construction module is used to associate the voiceprint features with the user's identity information to construct a voiceprint database; The identity comparison module is used to obtain the user's phone data and compare it with the voiceprint database to obtain the target identity information, and play the target identity information by voice for confirmation; The registration module is used to collect the user's historical registration information and the registration information in the phone data after the identity information is confirmed and match the registration.