Automatic prescription configuration system based on digital symptom analysis

By processing multi-round symptom descriptions through semantic analysis and weight decay techniques, identifying core and modifier descriptions, and generating a comprehensive description vector, the problem of inconsistency in information processing in existing technologies is solved, and the accuracy and consistency of prescription preparation are improved.

CN121725975BActive Publication Date: 2026-05-19XIAN NEW HOPE MEDICAL EQUIP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN NEW HOPE MEDICAL EQUIP CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies struggle to identify the degree of information addition and semantic changes in multiple rounds of symptom descriptions, which can easily lead to situations where historical information is either excessively retained or completely ignored, resulting in reduced discriminability and rationality of the configuration results.

Method used

The semantic analysis module identifies the core description and the modifying description, calculates the semantic strength value of the modifying description, and fuses the description vectors from different rounds based on the weight decay coefficient to generate a comprehensive description vector for prescription matching.

Benefits of technology

It improves the semantic structuring and consistency of non-standard symptom descriptions, enhances the distinguishability and practical value of symptom information, and reduces the adverse effects of information loss or redundant information on judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725975B_ABST
    Figure CN121725975B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses an automatic prescription configuration system based on digital symptom analysis, which comprises a database module, a semantic analysis module, a first coding module, a second coding module and a retrieval module; wherein: the database module is used for storing prescription configuration data and description sentences of user's each round of symptom description; the semantic analysis module is used for performing semantic segmentation on the description sentences and identifying core descriptions and modification descriptions; the first coding module is used for respectively coding each group of modification descriptions and corresponding core descriptions into description vectors; the second coding module fuses description vectors of different rounds into comprehensive description vectors; and the retrieval module generates a prescription matching list based on the comprehensive description vectors and the prescription configuration data. Through improving the semantic structuralization degree and consistent expression capability of non-standard and multi-round symptom descriptions, the application improves the distinguishability and practical value of the description information in the configuration matching process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to an automated prescription preparation system based on digital symptom analysis. Background Technology

[0002] In scenarios such as digital health assistance platforms and intelligent consultation systems, symptom-based automated prescription configuration systems are being deployed. These systems rely on users to describe their symptoms using natural language or structured input, and then generate a list of recommended prescriptions through algorithmic modules.

[0003] In real-world applications, user symptom descriptions exhibit significant non-standardization. On one hand, users often use colloquial, emotional, or vague expressions to describe their feelings, frequently incorporating time adverbs, degree words, frequency words, and subjective sentiment terms, leading to unclear semantic boundaries and uneven information density. On the other hand, different users express the same symptom in significantly different ways, and the same user's descriptions at different times may be repeated, revised, or supplemented, resulting in fragmented and dynamically changing symptom information. Most existing systems, when processing such input, typically analyze entire descriptions or single keywords, lacking an effective distinction between the core symptom expression and modifying information. This makes it difficult to accurately reflect the true state of the symptoms, reducing the discriminative power and reasonableness of the configuration results.

[0004] In multi-round interaction or phased input scenarios, existing technologies typically employ overlay or simple splicing processing strategies, relying primarily on the latest round of input or mechanically overlaying historical inputs. This makes it difficult to identify the degree of information addition and semantic change between different rounds, easily leading to situations where historical information is over-retained or completely ignored, thus affecting the consistency and stability of the overall judgment.

[0005] For example, Chinese patent CN109979558B discloses a method for analyzing symptom-drug association based on artificial intelligence technology, including the following steps: a symptom extraction module, which extracts the patient's symptom set; a symptom word vector mapping module, which maps the Chinese symptom set into a high-dimensional dense word vector set; a symptom word vector encoding module, which uses a long short-term memory network encoding model to encode the symptom word vector set, generating an information vector containing all symptom information and a set of n symptom encoded information vectors; and a suggestion prescription generation module, which uses the information vector and the symptom encoded information vector set as input, and uses a long short-term memory network generation model combined with an attention model to generate a suggested prescription containing L Chinese herbal medicines, i.e., the association between symptoms and Chinese herbal medicines.

[0006] For example, Chinese patent application CN118658591A discloses a method and system for generating traditional Chinese medicine (TCM) prescriptions based on big data processing. This method includes: using a diagnostic model to determine the user's facial complexion and tongue coating information; determining multiple matching users and their corresponding TCM prescriptions based on the user's input symptom information, facial complexion, tongue coating, and a TCM prescription database; using a prescription analysis model to determine multiple candidate principal herbs, assistant herbs, adjuvant herbs, and guiding herbs based on the user's input symptom information, facial complexion, tongue coating, and the TCM prescriptions corresponding to the multiple matching users; and using a generative adversarial network to generate the target TCM prescription. This method can quickly and accurately generate prescriptions suitable for the user.

[0007] All of the above technical solutions suffer from the problems mentioned in the background of this application: it is difficult to identify the degree of information addition and semantic changes in different rounds, and historical information is easily over-retained or completely ignored.

[0008] The information disclosed in this background section is intended only to enhance the understanding of the overall background of this application and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0009] The technical problem to be solved by this application is to overcome the shortcomings of the prior art and provide an automatic prescription configuration system based on digital symptom analysis. By improving the semantic structuring and consistency of non-standard, multi-round symptom descriptions, the system enhances the distinguishability and practical value of descriptive information in the configuration matching process.

[0010] To solve the above-mentioned technical problems, this application provides the following technical solution:

[0011] An automated prescription preparation system based on digital symptom analysis includes a database module, a semantic analysis module, a first encoding module, a second encoding module, and a retrieval module; wherein:

[0012] The database module is used to store prescription configuration data and descriptive statements of the user's symptom descriptions for each round;

[0013] The semantic analysis module is used to perform semantic segmentation on the description statement and identify the core description and the modifying description;

[0014] The first encoding module is used to calculate the semantic strength value of the modification description, and encode each group of modification descriptions and the corresponding core description into a description vector based on the semantic strength value;

[0015] The second encoding module calculates the weight decay coefficient of each round of symptom description based on the description vector, and merges the description vectors of different rounds into a comprehensive description vector based on the weight decay coefficient.

[0016] The retrieval module generates a prescription matching list based on the comprehensive description vector and prescription configuration data.

[0017] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the semantic analysis module includes a word segmentation unit and a semantic recognition unit;

[0018] The word segmentation unit is used to perform semantic segmentation on the description statement; the word segmentation unit is equipped with a word segmentation tool to segment the description statement and form a word sequence;

[0019] The semantic recognition unit is used to identify the core description and modifying description in the descriptive statement; the semantic recognition unit is configured with a sequence labeling model, which is used to identify the core description and modifying description from the word sequence, and to identify the correspondence between the core description and the modifying description;

[0020] The sequence labeling model takes a word sequence as input and outputs a labeled sequence; the labeled sequence labels the start and end points of each core description and corresponding modifier description in the corresponding word sequence.

[0021] The core description is the core expression of the symptom, and the modifier description is used to modify the strength of the core description.

[0022] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the first encoding module includes a semantic scoring unit; the semantic scoring unit is used to calculate the semantic intensity value of each modification description, specifically including:

[0023] Each modifier description is mapped to a word vector; the semantic scoring unit is also configured with a set of strength anchors; each set of strength anchors contains a set of reference modifier descriptions and a bound semantic strength value;

[0024] The similarity between each modifier description and each intensity anchor set is calculated based on word vectors; the similarity between any modifier description and any intensity anchor set is the maximum similarity between the word vector of the modifier description and the word vector of each reference description modified in the intensity anchor set;

[0025] For any modifier description, the semantic strength value of the set of strength anchors with the highest similarity is used as the semantic strength value of the modifier description.

[0026] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the first encoding module further includes a first encoding unit; the first encoding unit is configured with a first encoding strategy, used to encode each set of modification descriptions and corresponding core descriptions into description vectors respectively; the first encoding strategy specifically includes:

[0027] For any set of modifier descriptions and their corresponding core descriptions, the core descriptions are encoded into core vectors, and the semantic strength values ​​of the modifier descriptions are encoded into strength vectors; the strength vectors and core vectors are concatenated to form a description vector.

[0028] The first encoding unit is also configured with a modification weight ratio; if the number of modification descriptions corresponding to the core description is 1, the semantic strength value of the modification description is encoded into an intensity vector; if the number of modification descriptions corresponding to the core description is greater than 1, the modification descriptions are classified, and a weight value is assigned to each type of modification description based on the modification weight ratio, and the total weight value of different types of modification descriptions is 1; the weighted sum of the semantic strength values ​​of each type of modification description is calculated based on the weight value, and the weighted sum is encoded into an intensity vector.

[0029] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the second encoding module includes a fusion computing unit;

[0030] The fusion computing unit is configured with a weight allocation strategy to calculate the weight decay coefficient of each round of symptom description; the weight allocation strategy specifically includes:

[0031] The information update rate and semantic drift rate of each round of symptom description are calculated based on the description vectors.

[0032] Basic weights are set for each round of symptom description based on the corresponding information update rate and semantic drift rate, where both the information update rate and semantic drift rate are positively correlated with the basic weights;

[0033] The fusion computing unit is also equipped with a time decay function; based on the time decay function and the timestamp of each round of symptom description, the time decay factor of the most recent M rounds of symptom description is calculated respectively; M is a positive integer;

[0034] Multiply the base weight of each round in the most recent M rounds of symptom description by the corresponding time decay factor to obtain the weight decay coefficient of the most recent M rounds of symptom description.

[0035] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the calculation of the information update rate of symptom descriptions for any target round specifically includes:

[0036] The sum of the semantic intensity values ​​corresponding to the intensity vectors in all description vectors in the target round is calculated as the total information of the target round;

[0037] Extract the description vectors of the symptoms in the N consecutive rounds preceding the target round, and use them as the reference description vector for the target round; N is a positive integer;

[0038] The core vector that does not appear in the reference description vector is marked as a new description vector for the target round;

[0039] The sum of the semantic intensity values ​​corresponding to the intensity vectors of all newly added description vectors in the target round is calculated as the amount of new information in the target round.

[0040] The ratio of newly added information to the total amount of information is calculated as the information update rate for the target round.

[0041] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the calculation of the semantic drift rate of symptom descriptions in any target round specifically includes:

[0042] Extract the description vector of the symptoms from the previous round adjacent to the target round, and use it as the comparison description vector for the target round;

[0043] The description vector that appears in the comparison description vector is marked as the dominant description vector of the target round;

[0044] The information difference of each dominant description vector is calculated as the sum of the information differences, which is used as the semantic drift of the target round; the information difference of any dominant description vector is the absolute value of the difference between the corresponding semantic intensity value and the previous round.

[0045] Calculate the sum of the dominant strength values ​​of each dominant description vector as a normalization coefficient; the dominant strength value of any dominant description vector is the maximum value of the corresponding semantic strength value between the target round and the previous round.

[0046] The ratio of semantic drift to the normalization coefficient is calculated as the semantic drift rate for the target round.

[0047] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the second encoding module further includes a second encoding unit; the second encoding unit is configured with a second encoding strategy for fusing description vectors from different rounds into a comprehensive description vector; the second encoding strategy specifically includes:

[0048] The fusion strength of each core description is calculated in the most recent M rounds of symptom description. In any round of symptom description, the fusion strength of any core description is the product of the semantic strength value of the corresponding modification description and the weight decay coefficient of the corresponding round.

[0049] In the M rounds of symptom description, for any core description, only the maximum fusion intensity is retained; for any core description, the maximum fusion intensity is encoded as a fusion intensity vector, and the fusion intensity vector is concatenated with the corresponding core vector to form the description vector of the core description;

[0050] The description vectors of each core description retained in the M rounds of symptom descriptions are concatenated to form the comprehensive description vector.

[0051] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the database module includes a first storage unit and a second storage unit;

[0052] The first storage unit is used to store the descriptive statements of the user's symptom descriptions for each round; the second storage unit is used to store the prescription configuration data.

[0053] The prescription configuration data includes prescription number and adaptation description vector for each prescription; the adaptation description vector of any prescription is composed of description vectors of one or more core descriptions.

[0054] As a preferred embodiment of the automated prescription preparation system based on digital symptom analysis described in this application, the retrieval module includes a retrieval matching unit and a data output unit;

[0055] The retrieval and matching unit is configured with a matching strategy to calculate the matching score between different prescriptions and the comprehensive description vector; the data output unit is used to generate a prescription matching list; the prescription matching list includes a list of prescriptions and the matching score corresponding to each prescription in the list;

[0056] The matching strategy specifically includes: calculating the similarity between the comprehensive description vector and the adapted description vector of each prescription as the matching score between each prescription and the comprehensive description vector.

[0057] Compared with the prior art, the beneficial effects achieved by this application are as follows:

[0058] This application can effectively parse natural language symptom descriptions input by users, transforming colloquial, vague, and emotional expressions into clearly structured and semantically explicit symptom information representations, thereby improving the comprehensibility and usability of the original input in subsequent processing stages.

[0059] By distinguishing and quantifying different semantic components in symptom descriptions, this application can objectively reflect the differences in symptoms in terms of degree, frequency, or duration, avoiding the equal treatment of descriptions of different intensities in the processing results, and helping to improve the precision and consistency of symptom representation.

[0060] In response to users' phased and supplementary symptom input methods, this application can establish effective connections between multiple rounds of description, reasonably balance the impact of new input information and historical information, and reduce the adverse effects of information loss or redundant accumulation on the overall judgment. Attached Figure Description

[0061] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0062] Figure 1 A schematic diagram of the structure of the automated prescription preparation system based on digital symptom analysis provided in this application;

[0063] Figure 2 A functional diagram of the automated prescription preparation system based on digital symptom analysis provided in this application. Detailed Implementation

[0064] The technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments and specific features in the embodiments are detailed descriptions of the technical solution of this application, rather than limitations thereof. In the absence of conflict, the embodiments and technical features in the embodiments can be combined with each other.

[0065] Example 1

[0066] This embodiment introduces an automated prescription preparation system based on digital symptom analysis, referring to... Figure 1 The system includes a database module, a semantic analysis module, a first encoding module, a second encoding module, and a retrieval module; the functions of each module and the collaboration mode between modules are as follows: Figure 2 As shown.

[0067] The database module is used to store prescription configuration data and descriptive statements of the user's symptom descriptions for each round;

[0068] The semantic analysis module is used to perform semantic segmentation on the description statement and identify the core description and the modifying description;

[0069] The semantic analysis module includes a word segmentation unit and a semantic recognition unit;

[0070] The word segmentation unit is used to perform semantic segmentation on the description statement; the word segmentation unit is equipped with a word segmentation tool to segment the description statement and form a word sequence;

[0071] Optionally, the word segmentation tool can be any one of THULAC, pkuseg, HanLP, or LTP. For example, a descriptive statement is "Recently, I've been feeling a bit of leg pain." After being segmented by the word segmentation tool, the resulting word sequence is "Recently / always / feel / a bit / leg pain." Subsequent processing uses each word in the word sequence as the smallest unit, rather than the entire descriptive statement or each character as the smallest unit, thus improving processing efficiency.

[0072] The semantic recognition unit is used to identify the core description and modifying description in the descriptive statement; the semantic recognition unit is configured with a sequence labeling model, which is used to identify the core description and modifying description from the word sequence, and to identify the correspondence between the core description and the modifying description;

[0073] The sequence labeling model takes a word sequence as input and outputs a labeled sequence; the labeled sequence labels the start and end points of each core description and corresponding modifier description in the corresponding word sequence.

[0074] The core description is the core expression of the symptom, and the modifier description is used to modify the strength of the core description.

[0075] For example, a word sequence is "headache / and / recurring / feeling / somewhat / weak", where the first core description is "headache" and the corresponding modifier is "recurring"; the second core description is "weakness" and the corresponding modifier is "somewhat".

[0076] This embodiment provides an optional sequence labeling model, the specific hierarchical structure, parameter settings, and training method of which are as follows.

[0077] A sequence labeling model integrating bidirectional long short-term memory networks and conditional random fields includes:

[0078] The word embedding layer maps each segmented word to a continuous vector representation to carry its basic semantic information. The word vector parameters are trainable, and a medium-sized dimension is sufficient; specifically, 128 dimensions are used.

[0079] A bidirectional Long Short-Term Memory (LSTM) network layer is used to model the context of word sequences, enabling the model to utilize the preceding and following contextual information of the word at the current position, thereby distinguishing the functional role of the same word in different contexts. A bidirectional structure is employed, with forward LSTM and backward LSTM running in parallel. The number of hidden units in each direction is set to 128, resulting in a 256-dimensional contextual representation at each position in the word sequence. To improve generalization ability, a dropout mechanism can be introduced at this layer or its input, for example, with a dropout ratio set to 0.2.

[0080] The linear mapping layer maps the context representation of the bidirectional LSTM output to the label space, generating a score for each word on each label. The output dimension of the linear mapping is equal to the size of the label set. For example, when using BIO annotation to distinguish core descriptions, modifiers, and other terms, the number of labels can be 3.

[0081] A Conditional Random Field (CRF) layer is used to model the transition relationships between labels and perform globally optimal decoding on the word sequence. By learning label transition scores, this layer can constrain the legality of the labeled sequence, such as ensuring the correct connection between the start and end points of any core description, thereby improving the stability of continuous segment recognition. In the prediction stage, the Viterbi algorithm is used to decode the labeled sequence and output the optimal labeling result.

[0082] The training method for this model is as follows:

[0083] The training data consists of several word sequences after word segmentation, with each word position corresponding to a manually labeled tag. Tags are assigned using the BIO annotation method, corresponding to core descriptions, modifiers, and other terms. During training, the word sequences are first mapped to vector sequences via a word embedding layer, and then input into a bidirectional LSTM layer to obtain context representations. These context representations are converted into scores for each label via a linear mapping layer, and then input into a Conditional Random Field (CRF) layer. The CRF layer calculates the sequence-level log-likelihood based on the real labeled sequences, with the training objective being to maximize this log-likelihood; the corresponding loss function is the negative log-likelihood loss. Model parameters are updated via backpropagation, and the optimizer is Adam. The learning rate can be set to 0.001 as an initial value, and can be appropriately reduced, for example, to 0.0005, when training becomes unstable. The batch size is set to 16 or 32, depending on the sequence length and computational resources. To avoid gradient instability during recurrent neural network training, the gradient norm can be pruned, for example, by setting the pruning threshold to 1.0. The number of training epochs can be set to 10 to 30, and an early stopping strategy is adopted based on the validation set performance to prevent overfitting. Model evaluation prioritizes fragment-level metrics, which involves merging consecutive BIO tags into fragments and then calculating the recognition accuracy and recall for the core description and modified fragments separately. During training, the model parameters with the best fragment-level metrics on the validation set are selected as the final model for subsequent online inference and system integration.

[0084] Users often use non-standard expressions when inputting symptoms, including emotional, vague, and redundant expressions, which affects the system's accuracy in extracting keywords and subsequent configuration. This application improves the system's accuracy in understanding non-standard input by identifying core descriptions and modifying descriptions, eliminating colloquial words and redundant descriptions, and retaining core expressions. This enhances the input quality of subsequent configuration chains.

[0085] The first encoding module is used to calculate the semantic strength value of the modification description, and encode each group of modification descriptions and the corresponding core description into a description vector based on the semantic strength value;

[0086] The first encoding module includes a semantic scoring unit and a first encoding unit;

[0087] The semantic scoring unit is used to calculate the semantic strength value of each modification description, specifically including:

[0088] Each modifier description is mapped to a word vector; for example, Word2Vec, FastText, or any other Chinese word vector model can be used to map the modifier description to a word vector.

[0089] The semantic scoring unit is also configured with a set of strength anchors; each set of strength anchors contains a set of reference modifier descriptions and a bound semantic strength value;

[0090] In this embodiment, the semantic strength value is used to quantify the strength of the corresponding modifier description. For example, some examples of strength anchor sets are as follows: the set of strength anchors with a bound semantic strength value of 0.9 includes reference modifier descriptions such as "severe", "intense", and "unbearable"; the set of strength anchors with a bound semantic strength value of 0.3 includes reference modifier descriptions such as "slight", "somewhat", and "a little".

[0091] The similarity between each modifier description and each intensity anchor set is calculated based on word vectors; the similarity between any modifier description and any intensity anchor set is the maximum similarity between the word vector of the modifier description and the word vector of each reference description modified in the intensity anchor set; optionally, the similarity between word vectors is calculated based on cosine similarity.

[0092] For any modifier description, the semantic strength value of the set of strength anchors with the highest similarity is used as the semantic strength value of the modifier description.

[0093] The first encoding unit is configured with a first encoding strategy, used to encode each set of modification descriptions and corresponding core descriptions into description vectors respectively; the first encoding strategy specifically includes:

[0094] For any set of modifier descriptions and their corresponding core descriptions, the core descriptions are encoded into core vectors, and the semantic strength values ​​of the modifier descriptions are encoded into strength vectors; the strength vectors and core vectors are concatenated to form a description vector.

[0095] The first encoding unit is also configured with a modification weight ratio; if the number of modification descriptions corresponding to the core description is 1, the semantic strength value of the modification description is encoded into an intensity vector; if the number of modification descriptions corresponding to the core description is greater than 1, the modification descriptions are classified, and a weight value is assigned to each type of modification description based on the modification weight ratio, and the total weight value of different types of modification descriptions is 1; the weighted sum of the semantic strength values ​​of each type of modification description is calculated based on the weight value, and the weighted sum is encoded into an intensity vector.

[0096] Optionally, the modifiers can be categorized into intensity, duration, and frequency. Intensity describes the degree of strength, such as: somewhat, slightly, relatively, obviously, very, extremely, serious, intense, severe, etc. Frequency describes the frequency of occurrence, such as: occasionally, often, frequently, repeatedly, always, etc. Duration describes the time span or continuity, such as: continuous, consistently, long-term, several days, this period, etc. The weight ratio of intensity, duration, and frequency can be set to 12:5:3. Among them, intensity directly expresses the strength of the degree and contributes the most to the semantic intensity value, thus having the highest weight. Duration expresses how long it lasts and contributes the second most to the semantic intensity value; it is often used to emphasize that it is not accidental, so its weight is the second lowest. Frequency expresses how often it occurs and enhances the semantic intensity value, but it is easily influenced by colloquial narrative habits, so its weight is the lowest. For example, in a descriptive statement, “I’ve been experiencing chest tightness lately, and it happens frequently,” according to the above-mentioned modifier weighting ratio, “somewhat” is classified as intensity with a weight of 0.706, and “frequently” is classified as frequency with a weight of 0.294.

[0097] In this embodiment, the core vector is a fixed code for the core description. For example, the fixed code for headache is 0001, and the fixed code for chest tightness is 0002. This fixed mapping relationship remains unchanged throughout multiple rounds of symptom description and prescription configuration. When the same core description is associated with multiple modifier descriptions, this scheme does not simply accumulate the semantic intensity values. Instead, it generates a single semantic intensity value through the modifier weight ratio to avoid numerical inflation caused by redundancy. This preserves the dominant role of intensity-based modifiers while allowing modifiers such as persistence and frequency to have a controllable impact on the overall intensity.

[0098] The second encoding module calculates the weight decay coefficient of each round of symptom description based on the description vector, and merges the description vectors of different rounds into a comprehensive description vector based on the weight decay coefficient.

[0099] The second encoding module includes a fusion computing unit and a second encoding unit;

[0100] The fusion computing unit is configured with a weight allocation strategy to calculate the weight decay coefficient of each round of symptom description; the weight allocation strategy specifically includes:

[0101] The information update rate and semantic drift rate of each round of symptom description are calculated based on the description vectors.

[0102] In this embodiment, the information update rate is used to measure what proportion of valid information in the current round of input comes from core descriptions that have not appeared before. Calculating the information update rate of the symptom descriptions for any target round specifically includes:

[0103] The sum of the semantic intensity values ​​corresponding to the intensity vectors in all description vectors in the target round is calculated as the total information of the target round;

[0104] Extract the description vectors of the symptoms in the N consecutive rounds preceding the target round, and use them as the reference description vector for the target round; N is a positive integer;

[0105] The core vector that does not appear in the reference description vector is marked as a new description vector for the target round;

[0106] The sum of the semantic intensity values ​​corresponding to the intensity vectors of all newly added description vectors in the target round is calculated as the amount of new information in the target round.

[0107] The ratio of newly added information to the total amount of information is calculated as the information update rate for the target round.

[0108] In this embodiment, the semantic drift rate is used to measure the magnitude of change in the core descriptive composition and expression intensity between adjacent rounds of description. Calculating the semantic drift rate of the symptom description for any target round specifically includes:

[0109] Extract the description vector of the symptoms from the previous round adjacent to the target round, and use it as the comparison description vector for the target round;

[0110] The description vector that appears in the comparison description vector is marked as the dominant description vector of the target round;

[0111] The information difference of each dominant description vector is calculated as the sum of the information differences, which is used as the semantic drift of the target round; the information difference of any dominant description vector is the absolute value of the difference between the corresponding semantic intensity value and the previous round.

[0112] Calculate the sum of the dominant strength values ​​of each dominant description vector as a normalization coefficient; the dominant strength value of any dominant description vector is the maximum value of the corresponding semantic strength value between the target round and the previous round.

[0113] The ratio of semantic drift to the normalization coefficient is calculated as the semantic drift rate for the target round.

[0114] Basic weights are set for each round of symptom description based on the corresponding information update rate and semantic drift rate, where both the information update rate and semantic drift rate are positively correlated with the basic weights;

[0115] Optionally, a globally shared reference weight can be set, and the reference weight can be adjusted for each round of symptom description based on the information update rate and semantic drift rate to obtain the corresponding basic weight. For example, within a preset range of values, the reference weight can be increased according to the information update rate and semantic drift rate, and the larger the information update rate and semantic drift rate, the larger the corresponding basic weight.

[0116] The fusion computing unit is also equipped with a time decay function; based on the time decay function and the timestamp of each round of symptom description, the time decay factor of the most recent M rounds of symptom description is calculated respectively; M is a positive integer;

[0117] Optionally, the independent variable of the time decay function is the time difference between the timestamp of any round of symptom description and the current timestamp, and the dependent variable is the time decay factor, with a larger time difference resulting in a smaller time decay factor. The time decay function can be set as a monotonically decreasing function or a step decay function. For example, if the time difference is 0-3 days, the time decay factor is 1; if the time difference is 4-7 days, the time decay factor is 0.7; if the time difference is 7-14 days, the time decay factor is 0.3, and so on. The time decay factor is used to correct the basic weights over time, ensuring that the final weight decay coefficient simultaneously satisfies the principle of prioritizing recent inputs while retaining but weakening historical information.

[0118] Multiply the base weight of each round in the most recent M rounds of symptom description by the corresponding time decay factor to obtain the weight decay coefficient of the most recent M rounds of symptom description.

[0119] The second encoding unit is configured with a second encoding strategy for fusing description vectors from different rounds into a comprehensive description vector; the second encoding strategy specifically includes:

[0120] Calculate the fusion strength of each core description in the most recent M rounds of symptom description; in any round of symptom description, the fusion strength of any core description is the product of the semantic strength value of the corresponding modifier description (or the weighted sum of the semantic strength values ​​of multiple modifier descriptions) and the weight decay coefficient of the corresponding round;

[0121] In the M rounds of symptom description, for any core description, only the maximum fusion intensity is retained; for any core description, the maximum fusion intensity is encoded as a fusion intensity vector, and the fusion intensity vector is concatenated with the corresponding core vector to form the description vector of the core description;

[0122] The description vectors of each core description retained in the M rounds of symptom descriptions are concatenated to form the comprehensive description vector.

[0123] When users do not input all symptoms at once, but rather supplement and continuously adjust their input in multiple rounds, traditional systems often make decisions based solely on the last round of input, easily overlooking the correlation between preceding and subsequent inputs. This application improves the adaptability to use cases such as phased, multi-round descriptions and supplementary descriptions.

[0124] The retrieval module generates a prescription matching list based on the comprehensive description vector and prescription configuration data.

[0125] The database module includes a first storage unit and a second storage unit;

[0126] The first storage unit is used to store the descriptive statements of the user's symptom description for each round; optionally, one round of symptom description is performed daily on a one-day basis, and the descriptive statements and corresponding timestamps are stored in the first storage unit; the second storage unit is used to store prescription configuration data.

[0127] The prescription configuration data includes prescription number and adaptation description vector for each prescription; the adaptation description vector of any prescription is composed of description vectors of one or more core descriptions.

[0128] In this embodiment, a corresponding core description is pre-matched for each prescription, and a reference value for the fusion intensity vector is set for the core description. This value is then concatenated with the core vector of the core description to form a description vector. The adapted description vector has the same form as the integrated description vector.

[0129] The retrieval module includes a retrieval matching unit and a data output unit;

[0130] The retrieval and matching unit is configured with a matching strategy to calculate the matching score between different prescriptions and the comprehensive description vector; the data output unit is used to generate a prescription matching list; the prescription matching list includes a list of prescriptions and the matching score corresponding to each prescription in the list;

[0131] The matching strategy specifically includes: calculating the similarity between the comprehensive description vector and the adapted description vector of each prescription as the matching score between each prescription and the comprehensive description vector.

[0132] Optionally, both the composite description vector and the adapted description vector are treated as sparse vectors constructed based on the same encoding space, and cosine similarity is calculated. Core descriptions that do not appear in any vector are automatically treated as zero in the similarity calculation, without the need for padding or dimension alignment. This makes it not mandatory for the two to have the same length, nor is it required that they contain the same number of core descriptions.

[0133] Optionally, before calculating the similarity between the comprehensive description vector and the adapted description vector of each prescription, prescriptions that cover at least a few core descriptions are retrieved by using the core vector of the core description as the index key, in order to reduce the amount of subsequent calculations.

[0134] Optionally, select several prescriptions with the highest matching scores to generate a prescription matching list, and sort the prescriptions in descending order of their matching scores with the comprehensive description vector.

[0135] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0136] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of protection of this application, and these forms are all within the protection scope of this application.

Claims

1. An automated prescription preparation system based on digital symptom analysis, characterized in that: It includes a database module, a semantic analysis module, a first encoding module, a second encoding module, and a retrieval module; among which: The database module is used to store prescription configuration data and descriptive statements of the user's symptom descriptions for each round; The semantic analysis module is used to perform semantic segmentation on the description statement and identify the core description and the modifying description; The first encoding module is used to calculate the semantic strength value of the modification description, and encode each group of modification descriptions and the corresponding core description into a description vector based on the semantic strength value; The second encoding module calculates the weight decay coefficient of each round of symptom description based on the description vector, and merges the description vectors of different rounds into a comprehensive description vector based on the weight decay coefficient. The retrieval module generates a prescription matching list based on the comprehensive description vector and prescription configuration data; The second encoding module includes a fusion calculation unit; the fusion calculation unit is configured with a weight allocation strategy for calculating the weight decay coefficient of each round of symptom description; the weight allocation strategy specifically includes: The information update rate and semantic drift rate of each round of symptom description are calculated based on the description vectors. Basic weights are set for each round of symptom description based on the corresponding information update rate and semantic drift rate, where both the information update rate and semantic drift rate are positively correlated with the basic weights; The fusion computing unit is also equipped with a time decay function; based on the time decay function and the timestamp of each round of symptom description, the time decay factor of the most recent M rounds of symptom description is calculated respectively; M is a positive integer; Multiply the base weight of each round in the most recent M rounds of symptom description by the corresponding time decay factor to obtain the weight decay coefficient of the most recent M rounds of symptom description.

2. The automated prescription preparation system based on digital symptom analysis as described in claim 1, characterized in that: The semantic analysis module includes a word segmentation unit and a semantic recognition unit; The word segmentation unit is used to perform semantic segmentation on the description statement; the word segmentation unit is equipped with a word segmentation tool to segment the description statement and form a word sequence; The semantic recognition unit is used to identify the core description and modifying description in the descriptive statement; the semantic recognition unit is configured with a sequence labeling model, which is used to identify the core description and modifying description from the word sequence, and to identify the correspondence between the core description and the modifying description; The sequence labeling model takes a word sequence as input and outputs a labeled sequence; the labeled sequence labels the start and end points of each core description and corresponding modifier description in the corresponding word sequence. The core description is the core expression of the symptom, and the modifier description is used to modify the strength of the core description.

3. The automated prescription preparation system based on digital symptom analysis as described in claim 2, characterized in that: The first encoding module includes a semantic scoring unit; the semantic scoring unit is used to calculate the semantic strength value of each modification description, specifically including: Each modifier description is mapped to a word vector; the semantic scoring unit is also configured with a set of strength anchors; each set of strength anchors contains a set of reference modifier descriptions and a bound semantic strength value; The similarity between each modifier description and each intensity anchor set is calculated based on word vectors; the similarity between any modifier description and any intensity anchor set is the maximum similarity between the word vector of the modifier description and the word vector of each reference description modified in the intensity anchor set; For any modifier description, the semantic strength value of the set of strength anchors with the highest similarity is used as the semantic strength value of the modifier description.

4. The automated prescription preparation system based on digital symptom analysis as described in claim 3, characterized in that: The first encoding module further includes a first encoding unit; the first encoding unit is configured with a first encoding strategy, used to encode each set of modification descriptions and corresponding core descriptions into description vectors respectively; the first encoding strategy specifically includes: For any set of modifier descriptions and their corresponding core descriptions, the core descriptions are encoded into core vectors, and the semantic strength values ​​of the modifier descriptions are encoded into strength vectors; the strength vectors and core vectors are concatenated to form a description vector. The first encoding unit is also configured with a modification weight ratio; if the number of modification descriptions corresponding to the core description is 1, the semantic strength value of the modification description is encoded into an intensity vector; if the number of modification descriptions corresponding to the core description is greater than 1, the modification descriptions are classified, and a weight value is assigned to each type of modification description based on the modification weight ratio, and the total weight value of different types of modification descriptions is 1; the weighted sum of the semantic strength values ​​of each type of modification description is calculated based on the weight value, and the weighted sum is encoded into an intensity vector.

5. The automated prescription preparation system based on digital symptom analysis as described in claim 4, characterized in that: Calculate the information update rate of symptom descriptions for any target round, specifically including: The sum of the semantic intensity values ​​corresponding to the intensity vectors in all description vectors in the target round is calculated as the total information of the target round; Extract the description vectors of the symptoms in the N consecutive rounds preceding the target round, and use them as the reference description vector for the target round; N is a positive integer; The core vector that does not appear in the reference description vector is marked as a new description vector for the target round; The sum of the semantic intensity values ​​corresponding to the intensity vectors of all newly added description vectors in the target round is calculated as the amount of new information in the target round. The ratio of newly added information to the total amount of information is calculated as the information update rate for the target round.

6. The automated prescription preparation system based on digital symptom analysis as described in claim 5, characterized in that: Calculate the semantic drift rate of symptom descriptions for any target round, specifically including: Extract the description vector of the symptoms from the previous round adjacent to the target round, and use it as the comparison description vector for the target round; The description vector that appears in the comparison description vector is marked as the dominant description vector of the target round; The sum of the information differences of each dominant description vector is calculated as the semantic drift of the target round; the information difference of any dominant description vector is the absolute value of the difference between the corresponding semantic intensity value and the previous round. Calculate the sum of the dominant strength values ​​of each dominant description vector as a normalization coefficient; the dominant strength value of any dominant description vector is the maximum value of the corresponding semantic strength value between the target round and the previous round. The ratio of semantic drift to the normalization coefficient is calculated as the semantic drift rate for the target round.

7. The automated prescription preparation system based on digital symptom analysis as described in claim 6, characterized in that: The second encoding module further includes a second encoding unit; the second encoding unit is configured with a second encoding strategy for fusing description vectors from different rounds into a comprehensive description vector; the second encoding strategy specifically includes: The fusion strength of each core description is calculated in the most recent M rounds of symptom description. In any round of symptom description, the fusion strength of any core description is the product of the semantic strength value of the corresponding modification description and the weight decay coefficient of the corresponding round. In the M rounds of symptom description, for any core description, only the maximum fusion intensity is retained; for any core description, the maximum fusion intensity is encoded as a fusion intensity vector, and the fusion intensity vector is concatenated with the corresponding core vector to form the description vector of the core description; The description vectors of each core description retained in the M rounds of symptom descriptions are concatenated to form the comprehensive description vector.

8. The automated prescription preparation system based on digital symptom analysis as described in claim 7, characterized in that: The database module includes a first storage unit and a second storage unit; The first storage unit is used to store the descriptive statements of the user's symptom descriptions for each round; the second storage unit is used to store the prescription configuration data. The prescription configuration data includes prescription number and adaptation description vector for each prescription; the adaptation description vector of any prescription is composed of description vectors of one or more core descriptions.

9. The automated prescription preparation system based on digital symptom analysis as described in claim 8, characterized in that: The retrieval module includes a retrieval matching unit and a data output unit; The retrieval and matching unit is configured with a matching strategy to calculate the matching score between different prescriptions and the comprehensive description vector; the data output unit is used to generate a prescription matching list; the prescription matching list includes a list of prescriptions and the matching score corresponding to each prescription in the list; The matching strategy specifically includes: calculating the similarity between the comprehensive description vector and the adapted description vector of each prescription as the matching score between each prescription and the comprehensive description vector.