High-performance pinyin input method and system based on big data for medical scenarios
Patent Information
- Application Number
- CN202611037949.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明提供基于大数据用于医疗场景的高性能拼音输入方法及系统,解决相关技术中医疗场景下拼音输入候选词条排序精度不足、姓名录入缺乏针对性过滤、多模式检索效率低下以及跨诊疗状态词条预测能力缺失的技术问题
基于医疗语料库统计计算的基础频率权重,使医疗高频词汇在候选排序中相较于通用词汇具有更高的基础权重,从而在降序排列中自然排列于有序候选列表前端,减少了医护人员的选词操作次数。
Smart Images

Figure CN122837646A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information technology, specifically to a high-performance Pinyin input method and system based on big data for medical scenarios. Background Technology
[0002] In medical information systems, medical staff frequently need to enter information such as patient names, addresses, medical terms, and medical records. Pinyin input method is the most commonly used Chinese character input method in the above scenarios.
[0003] Existing general-purpose Pinyin input methods are widely used in medical scenarios. However, these methods rely on a unified retrieval strategy for all input fields, failing to differentiate between different field types such as patient name and address. Furthermore, they lack address-level association capabilities, requiring users to retype Pinyin segment by segment when entering addresses with multiple administrative divisions. On touchscreen medical terminals, rapid, continuous input frequently triggers retrieval calculations.
[0004] The aforementioned deficiencies lead to the following technical issues: medical terminology and common patient names are ranked lower in the candidate list, increasing the number of word selection operations required by medical staff; the name and address fields cannot obtain targeted search results; the input efficiency of multi-level administrative division addresses is low; and the system experiences jitter in response under high-frequency input scenarios, affecting normal use.
[0005] Furthermore, during the diagnosis and treatment process, when medical staff switch between different stages, the initial input is often a short pinyin prefix. Existing hotspot prefix caching strategies rank candidates based on the global historical selection rate across all stages, without distinguishing the current treatment state. This causes the candidate ranking returned when the cache hits after a state transition to deviate from the current treatment context. While cross-state prediction based on big data association mining can identify target terms that match the current treatment context, its inference calculation cannot be completed synchronously when the cache hits, resulting in a timing contradiction between high-performance real-time return and high-precision context prediction. Summary of the Invention
[0006] This invention provides a high-performance Pinyin input method and system based on big data for medical scenarios, solving the technical problems in related technologies such as insufficient sorting accuracy of Pinyin input candidate words, lack of targeted filtering for name input, low efficiency of multi-mode retrieval, and lack of word prediction ability across treatment statuses in medical scenarios.
[0007] This invention discloses a high-performance Pinyin input method based on big data for medical scenarios, comprising: Obtain a general single-character dictionary, a phrase dictionary, and a dictionary of characters specific to names; parse the entry information and obtain the basic frequency weights based on word frequency statistics from a medical corpus. Based on the dictionary data, generate the following Trie trees: Pinyin Trie tree, Phrase Trie tree, Continuous Pinyin Trie tree, First Letter Trie tree, Mixed Input Trie tree, and Name-specific Trie tree. Receive the input string, determine the input mode based on the current input field type, and perform a cache hit query in the corresponding cache instance according to the input mode. The input mode is either name input mode or normal input mode. Based on the input pattern, multi-pattern parallel retrieval is performed in the corresponding Trie tree using the input string to generate a candidate term set; For each term in the candidate term set, the final comprehensive weight is calculated based on the basic frequency weight, the matching degree bonus coefficient, and the scenario bonus coefficient. The terms are then sorted in descending order to generate an ordered candidate list and stored in the corresponding cache instance.
[0008] Furthermore, the calculation method of the basic frequency weight is as follows: the ratio of the frequency of each word in the medical corpus plus one to the total number of words in the corpus is taken as logarithm, and then multiplied by the normalization scaling factor to map the original frequency to the weight range calculated by integer sorting; the normalization scaling factor is determined according to the size of the corpus.
[0009] Furthermore, the mixed input Trie tree supports a mixed input mode of full pinyin syllables and initial letter abbreviations. For each Chinese character in a phrase, its complete pinyin string or the first letter of the pinyin is used as the path segment. All legal mixed representations are enumerated and inserted into the mixed input Trie tree respectively. The mixed input Trie tree also supports a partial syllable abbreviation mode. In addition to the complete pinyin string and the first letter, the syllables of the phrase are represented by the first few letters of the syllable as a syllable prefix.
[0010] Furthermore, in the name input mode, the search scope is limited to the name-specific Trie tree; type filtering is applied to the search results according to the name sub-pattern flag: when the name sub-pattern flag indicates only the surname, only the surname type entries are retained; when the name sub-pattern flag indicates only the first name, only the first name type entries are retained; when the name sub-pattern flag indicates the full name, all name type entries are returned.
[0011] Furthermore, in the name input mode, the scenario bonus coefficient is jointly determined by the name character enhancement factor, the name type differentiation weighting coefficient, and the gender tendency weighting coefficient; wherein, the name type differentiation weighting coefficient takes a positive value when the term type matches the current name sub-pattern flag, and takes zero when they do not match; the gender tendency weighting coefficient is determined by the product of the term's gender tendency weight, the patient's gender code, and the gender weight adjustment coefficient, and generates a positive bonus when the term's gender tendency matches the patient's gender, and generates a negative inhibition when they do not match; in the normal input mode, the scenario bonus coefficient is always one.
[0012] Furthermore, it also includes: performing address-level trigger word scanning on the confirmed input text buffer; when an administrative division-level identifier is scanned, setting an association trigger bonus item for the address entries associated with the next level of that level; the association trigger bonus item is added to the final comprehensive weight; when the user confirms a candidate character or phrase, a hash lookup is performed in the pre-generated character phrase association index using the confirmed content as the search key, returning the associated list of common phrase candidates and pushing it to the user interface; the character phrase association index is generated by traversing the phrase dictionary data during the system initialization phase, using Chinese characters as keys and the list of common phrases containing that character as values.
[0013] Furthermore, a debouncing process is introduced when receiving the input string. The debouncing process sets a delay timer for each input event, with the delay duration being the debouncing time window. When the user triggers multiple input events consecutively within the debouncing time window, the previous timer is canceled and a new timer is set. The subsequent retrieval process is only executed when the user does not trigger any new input events within the debouncing time window, using the latest input string. The cache instances are divided into normal input mode cache instances and name input mode cache instances, both of which are managed using a least recently used strategy.
[0014] Furthermore, it also includes: extracting keyword sequences corresponding to each treatment status from historical medical record data, generating a cross-state directed term association graph according to the time sequence of the treatment process, where nodes are medical terms, directed edges point from terms of previous states to terms of subsequent states, and edge weights are cross-state conditional probabilities; counting the cumulative number of triggers for each pinyin prefix segment and the term finally confirmed by the user, when the cumulative number of triggers exceeds the hotspot threshold, generating hotspot prefix cache entries according to the historical selection rate and writing them into the hotspot prefix cache table; when a treatment status jump is detected, extracting the previous keyword set from the confirmed input content of the previous state, performing graph traversal along the state jump direction in the cross-state directed term association graph with the previous keyword set as the starting node, calculating predictive weights according to the multi-keyword probability aggregation method, and injecting terms with predictive weights exceeding the preheating threshold into the hotspot prefix cache table.
[0015] Furthermore, it also includes: performing a matching query on the hotspot prefix cache table using the input string; when a match is found, calculating a state-aware fusion score for each candidate term in the cache entry based on historical selection rate, predictive weight, and normalized base frequency weight; generating a state-aware fast candidate list by sorting the fusion scores in descending order; determining whether the fusion score of the term ranked first in the state-aware fast candidate list exceeds a confidence threshold: if it does, directly outputting the state-aware fast candidate list as the final candidate list; if it does not exceed the threshold, retaining the state-aware fast candidate list as a priority candidate, and simultaneously performing a supplementary traversal search in the Trie tree using the input string, merging the supplementary terms with the priority candidates, and sorting and outputting them according to the final comprehensive weight; when a change in the diagnosis and treatment status is detected again, clearing the injected state-related temporary entries and state-aware bonus items, and re-performing keyword extraction and cache preheating injection with the new preceding state and current state.
[0016] This invention discloses a high-performance Pinyin input system based on big data for medical scenarios, comprising: The dictionary loading module is used to obtain a general single-character dictionary, a phrase dictionary, and a dictionary of characters specific to names, parse the entry information, and obtain the basic frequency weights based on word frequency statistics from a medical corpus. The index generation module is used to generate, based on the dictionary data, a Pinyin Trie tree, a phrase Trie tree, a continuous Pinyin Trie tree, an initial letter Trie tree, a mixed input Trie tree, and a name-specific Trie tree. The input receiving and mode determination module is used to receive the input string, determine the input mode according to the current input field type, and perform a cache hit query in the corresponding cache instance according to the input mode. The input mode is either the name input mode or the normal input mode. The multi-mode retrieval module is used to perform multi-mode parallel retrieval in the corresponding Trie tree based on the input mode and the input string to generate a candidate term set; The weight calculation and sorting module is used to calculate the final comprehensive weight for each term in the candidate term set based on the basic frequency weight, matching degree bonus coefficient and scenario bonus coefficient, and generate an ordered candidate list in descending order and store it in the corresponding cache instance.
[0017] The present invention has the following beneficial effects: Basic frequency weights calculated based on a medical corpus This gives high-frequency medical terms a higher base weight than general terms in the candidate ranking, so they naturally appear at the front of the ordered candidate list in descending order, reducing the number of word selection operations required by medical staff.
[0018] By reading the field type flags, the search scope is switched to a separate name-specific Trie tree, and a boosting factor containing name characters is applied to entries in the name pattern. Name type differentiation weighting coefficient and gender bias weighting coefficient Scene bonus coefficient This completely isolates the candidate results for the name field from those for the general field in terms of retrieval path and weight calculation, thus avoiding interference from general terms on the name candidate results.
[0019] Lenovo trigger bonus items By scanning the administrative division identifiers in the confirmed input text buffer, higher weights are automatically assigned to lower-level address entries; the character phrase association index automatically pushes related phrases after the user confirms the upper-level address, so that the user can complete the continuous address entry without re-entering the pinyin of the lower-level address, reducing the number of multi-level administrative division address entry operations.
[0020] Image stabilization time window Input events in intermediate states are filtered to ensure that the system only performs the complete retrieval process on the final state after the user input has stabilized; the dual-buffer architecture caches the retrieval results of repeated input strings using LRU, which can reduce the overhead of repeated calculations in medical scenarios where patient information is highly repetitive, thereby enabling the system to maintain stable response performance in high-frequency input scenarios.
[0021] Cross-state directed term association graph The system encodes the causal relationships between terms across different stages of diagnosis and treatment contained in big data. At each transition point in the diagnosis and treatment process, a graph traversal is performed beforehand, and the prediction results are injected into a hotspot prefix cache table, which is then integrated with the scoring. The formula uses predictive weight terms By incorporating cross-state predictive weights into candidate ranking, terms supported by the current clinical context receive a boost in the fusion score, while the ranking of globally high-frequency terms that do not conform to the current clinical state is relatively reduced. This allows for the acquisition of candidate ranking results that conform to the current clinical context when the cache is hit.
[0022] The preheating injection operation is completed before the transition between treatment states, enabling the cross-state directed term association graph. The traversal calculation and user input response are decoupled in time. When subsequent short-prefix inputs hit the hot-spot prefix cache table, a fusion score incorporating predictive weights is applied. The calculations have been pre-computed, and the response time complexity remains at [value missing]. Level. Confidence threshold The hierarchical response strategy driven by the system completely bypasses Trie tree traversal in high-confidence scenarios and performs supplementary traversal only on terms not covered by priority candidates in low-confidence scenarios, so that the overall average retrieval cost decreases as the proportion of high-confidence scenarios increases. Attached Figure Description
[0023] Figure 1 This is a flowchart of a high-performance Pinyin input method for medical scenarios based on big data, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the basic frequency weight distribution of entries in the dictionary of names provided in this embodiment of the invention; Figure 3 This is a schematic diagram of the comprehensive weight ranking of candidate terms (male patients) when the input is "wei" according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the influence of gender preference weight on scene bonus coefficient provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the causal edge weight distribution from the main complaint keyword to the diagnostic term provided in an embodiment of the present invention; Figure 6 This is a schematic diagram showing the comparison of the fusion scores of each candidate term after the prefix "guan" hits the cache, as provided in an embodiment of the present invention. Figure 7 This is a schematic diagram of the multi-keyword probability aggregation process for predictive weighting of "coronary heart disease" provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the scatter distribution of each candidate term across different scoring dimensions provided in this embodiment of the invention; Figure 9 This is a schematic diagram comparing the number of index paths in a multi-level Trie tree provided in an embodiment of the present invention. Detailed Implementation
[0024] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0025] Example 1 According to an embodiment of the present invention, a high-performance Pinyin input method based on big data for medical scenarios is provided, such as... Figure 1 As shown, it includes the following steps: Step 1: Obtain medical corpus dictionary data Three types of dictionary data are obtained from persistent storage media: a general single-character dictionary, a phrase dictionary, and a name-specific character dictionary, and the structured parsing of the entry information is completed.
[0026] Each entry in the general single-character dictionary contains the following fields: the Chinese character itself, the corresponding standard pinyin string, the first letter of the pinyin, and the basic frequency weight value. . The calculation formula is derived from word frequency statistical analysis of a large-scale medical text corpus: in, Indicates a term Frequency of occurrence in medical corpora This indicates the total number of entries in the corpus. This is the normalization scaling factor, and its value range is... The weights are determined based on the size of the lexicon to ensure that the weight values fall within a reasonable integer range.
[0027] Each entry in the phrase dictionary contains the following fields: the complete Chinese character string of the multi-character phrase, the corresponding complete Pinyin string (with each Pinyin character separated by a space), the continuous Pinyin string (with each Pinyin character directly concatenated without separators), the initial letter string, and the basic weight value. For address phrases, address hierarchy tags (provincial, municipal, district / county, and street levels) and their parent address association keys are also stored for use by subsequent address association functions.
[0028] The name-specific dictionary is stored independently of the general dictionary. Its entries include the following fields: Chinese character, corresponding pinyin string, first letter, type identifier NameType (values are SURNAME for surname or GIVEN_NAME for given name), and gender bias weight GenderBias (value range is...). ,in This indicates a strong preference for words used by women. Indicates a strong preference for masculine words. (This includes neutral-sounding characters) and basic frequency weights. The above metadata was obtained through statistical analysis of large-scale population name data.
[0029] It should be noted that the above-mentioned basic frequency weights The calculation method involves performing a logarithmic transformation on the frequency of each term in the medical corpus, followed by multiplying by a normalization scaling factor. This maps the original frequencies to a weight range suitable for integer sorting calculations. For terms with extremely low frequencies, a weight is added to the numerator. The smoothing process can avoid the problem of the logarithm tending to negative infinity.
[0030] Step 2: Generate a multi-level Trie tree index Five types of special Trie tree index structures are respectively generated based on the loaded dictionary data, so as to support subsequent efficient multi-pattern retrieval.
[0031] The Trie tree takes each character of a string as a node branch key, and the time complexity of its prefix retrieval is , where is the length of the query string, which is not affected by the scale of the word library. The data structure of all Trie tree nodes includes the following three fields: a child node hash mapping table with characters as keys and child node references as values; a Boolean flag bit IsEndOfKey, which identifies whether the node is the termination node of an entry key-value path; and an entry reference list, which stores one or more entry information corresponding to the key value when IsEndOfKey is true, so as to handle the case of homophones.
[0032] The generation methods of the five types of Trie trees are as follows: Pinyin Trie (PinyinTrie) is generated by taking the standard Pinyin string of each entry as the key-value path, and is used to support prefix matching retrieval of single-character Pinyin.
[0033] Phrase Trie (PhraseTrie) is generated by taking the complete Pinyin string (including space separators) of a phrase entry as the key-value path, and is used to support Pinyin prefix matching of multi-word phrases.
[0034] Continuous Pinyin Trie (ContinuousPinyinTrie) is generated by taking the continuous Pinyin string (without space separators) of a phrase entry as the key-value path, and is specially used to process users' continuous Pinyin input without spaces.
[0035] Initials Trie (InitialsTrie) is generated by taking the concatenated string of Pinyin initials of all Chinese characters in an entry as the key-value path, and supports the pure initial abbreviation input mode.
[0036] Mixed Input Trie (MixedInputTrie) supports the input mode mixing full Pinyin syllables and initial abbreviations. For each Chinese character in a phrase, the complete Pinyin string can be used as the path segment corresponding to the Chinese character, or only the Pinyin initial can be used as the path segment corresponding to the Chinese character. Thus, all legal mixed representation forms of the entry are enumerated and inserted into the mixed input Trie respectively. For example, the phrase "Guangzhou City" can produce multiple legal paths such as "guangzhoushi" (full Pinyin), "gzs" (all initials), "guangzs" (full Pinyin of the first two characters plus initial of the last character), all of which are inserted into the mixed input Trie.
[0037] In addition, the system independently maintains a special name Trie tree (NameTrie) for special name characters, which is generated in the same way as the pinyin Trie tree, but only includes entries from the special name character dictionary, ensuring that the name retrieval path is completely isolated from the general retrieval path.
[0038] It should be noted that the difference between the aforementioned continuous pinyin Trie tree and the phrase Trie tree is that the key path of the phrase Trie tree includes space delimiters, requiring users to separate each syllable with a space during input, while the key path of the continuous pinyin Trie tree does not contain any delimiters, which corresponds to the input habit of users directly typing all pinyin letters continuously. The two types of Trie trees are maintained in parallel to cover the input habits of different users.
[0039] In the embodiment of the present application, in order to further expand the coverage of mixed input, the mixed input Trie tree can be extended to support the partial syllable abbreviated pinyin mode. In the partial syllable abbreviated pinyin mode, in addition to the complete pinyin string and the first letter, a certain syllable of an entry can also be represented by the first several letters of the syllable (i.e., syllable prefix). For example, in addition to the complete pinyin and the first letter "z", the syllable "zhang" can also be represented by syllable prefixes such as "zh", "zha", and "zhan", so as to support the scenario where inputting "zhangs" matches the candidate word "Zhang San".
[0040] Step 3: receiving an input string and determining an input mode receiving an input event from a front-end user interface, acquiring the pinyin string InputStr currently typed by the user, and reading the current value of the global status flag IsNameMode. IsNameMode is set by the medical service system according to the field type of the current input form: when the input field is a patient name field, IsNameMode is true, and the NameSubMode sub-mode flag (with a value of SURNAME_ONLY, GIVEN_NAME_ONLY or FULL_NAME) can be further read; when the input field is another type of field, IsNameMode is false, and the normal input mode is entered. The input mode is a name input mode or a normal input mode.
[0041] Based on the input pattern, using InputStr (a combination of InputStr and the NameSubMode sub-pattern flag in the name pattern) as the key, a cache hit query is performed in the corresponding normal input pattern cache instance (_inputCache) or name input pattern cache instance (_nameInputCache). If a cache hit occurs, the ordered candidate list stored in the cache is returned directly, skipping all processing steps 4 and 5. Both cache instances use an LRU (Least Recently Used) strategy, and the default capacity limit for each instance is [value missing]. One entry.
[0042] It should be noted that the NameSubMode flag mentioned above is used to further distinguish between surname and given name in the name input mode, so as to apply differentiated processing to different types of name characters in subsequent retrieval and weight calculation, and avoid interference between surname candidates and given name candidates.
[0043] In this embodiment, to address the issue of frequent retrieval triggers when users rapidly and continuously input characters on a touchscreen device, anti-shake processing is introduced based on step 3. The anti-shake processing sets a delay timer for each input event, with the delay duration being the anti-shake time window. The default value is Milliseconds. When the user is... If multiple input events are triggered consecutively within a time window, cancel the previous timer and reset the new timer; only when the user... If no new input event is triggered within the time window, the timer expires, and a complete subsequent retrieval process is executed with the latest InputStr. The value can be adjusted according to the computing performance of the target terminal device.
[0044] Step 4: Perform multi-mode parallel retrieval Based on the input pattern determined in step 3, multi-path parallel retrieval is performed using InputStr as the query string to generate a candidate term set.
[0045] In normal input mode, the following two parallel search paths are triggered simultaneously. Path A is the exact match path: A hash lookup is performed directly in the general word hash index table and the phrase hash index table, using InputStr as the key, returning a set of terms, ExactMatchSet, that completely matches the input string. The time complexity of the hash lookup is O(log n). Path B is a prefix and mixed matching path: using InputStr as the query string, prefix searches are performed in the Pinyin Trie tree, phrase Trie tree, continuous Pinyin Trie tree, initial letter Trie tree, and mixed input Trie tree, respectively. This returns a set of all key-value paths whose prefixes are InputStr, which are then merged into a prefix matching set, PrefixMatchSet. For each term in PrefixMatchSet, its matching type label (PINYIN_PREFIX, PHRASE_PREFIX, CONTINUOUS_PREFIX, INITIALS_PREFIX, or MIXED_PREFIX) is recorded for use in the weight calculation in step 5.
[0046] In name input mode, the search scope is limited to the name-specific hash index and the name-specific Trie tree. Based on the value of the NameSubMode sub-mode flag, type filtering is applied to the search results: when the NameSubMode sub-mode flag is SURNAME_ONLY, only entries with NameType "SURNAME" are retained; when the NameSubMode sub-mode flag is GIVEN_NAME_ONLY, only entries with NameType "GIVEN_NAME" are retained; when the NameSubMode sub-mode flag is FULL_NAME, entries of all name types are returned. This filtering operation is performed on the term reference list of the Trie tree leaf nodes, without introducing additional full-scan overhead.
[0047] It should be noted that the above method of recording matching type labels refers to labeling the type of Trie tree from which each hit term comes during the Trie tree prefix search process. This allows for the application of differentiated correction coefficients to different matching types in subsequent weight calculations, reflecting the differences in the clarity of user intent under different matching modes.
[0048] Step 5: Calculate and sort the multi-dimensional dynamic weights. For all candidate terms returned by each search path, calculate the final comprehensive weight for each term. The calculation formula is: The variables are defined as follows: The base frequency weight of the term is derived from the corpus frequency statistics in step 1.
[0049] This is a matching score bonus factor, calculated based on the term's matching type and matching coverage. When a term comes from an exact matching path, ,in To perfectly match the benchmark coefficient, the value is set to... When a term comes from a prefix matching path, The calculation formula is: in The length of the current input string. For the entry The index key and value strings are the total length of the strings, both of which are dimensionless character counts. The prefix matching scaling factor has a value of For different matching types, an additional type correction coefficient is applied: the type correction coefficient for pinyin prefix matching (PINYIN_PREFIX) is [value missing]. Continuous pinyin prefix matching (CONTINUOUS_PREFIX) is The initial letter prefix matching (INITIALS_PREFIX) is Mixed input prefix matching (MIXED_PREFIX) is The type correction factor for phrase prefix matching (PHRASE_PREFIX) is: .
[0050] It should be noted that the above-mentioned type of correction coefficient applies to The method refers to calculating the prefix matching path. Multiplying the value by the correction factor corresponding to the matching type yields the final value used for that term. Value, then substitute The formula is used in subsequent calculations.
[0051] The scene-based bonus coefficient is calculated using the following formula in name input mode: in, Promotion factor for characters in the name, default value is , can Adjust within the scope; This is a weighted coefficient for name type differentiation, which is set to a value when the term NameType matches the current NameSubMode subpattern flag. When there is a mismatch, the value is taken as ; The gender bias weighting coefficient is calculated using the following formula when the patient's gender information is known in the system: in, For the entry The gender bias weight is derived from the GenderBias field stored in the name-specific dictionary in step 1, and the value range is [value range missing]. ; The patient's gender is coded, with males taking... female selection ; This is the gender weighting adjustment coefficient; the default value is [value to be filled in]. . , and All are dimensionless values, and the product of the three is... It is also a dimensionless coefficient. When and When the numbers are the same, A positive value contributes positively to the scenario-based bonus coefficient of that term; when the two values have opposite signs, A negative value negatively suppresses the scenario-based bonus coefficient for that term. This is especially true when the system does not know the patient's gender information. Values .
[0052] It should be noted that the above in the formula The range of values for . Because , , The default value is The sum of the three The range of values is .at the same time, The range of values is .therefore The minimum value is It is always a positive value. The value is always positive in name input mode and will not cause... Sign flips. In normal input mode, The constant value is .
[0053] This is an associative trigger bonus, effective only for address-related phrases. The system performs an address-level trigger word scan on the confirmed input text buffer (ConfirmedTextBuffer). The predefined trigger word set includes administrative division level identifiers such as "province," "city," "autonomous region," "district," "county," "street," "town," and "village." When a trigger word is detected... At that time, among them For the first in the set of trigger words A trigger word, which will be associated with The address terms associated with the next level down from the identified address level Set to the preset associative trigger boost value The default value is the maximum base weight value in the current lexicon. times, and With consistent dimensions, they can directly participate in addition operations. For non-address terms, The constant value is .
[0054] All candidate terms are sorted The values are sorted in descending order to generate an ordered candidate list, CandidateList. The first few values of CandidateList are then truncated after sorting. One candidate, This is the maximum number of items to be extracted from the candidate list; the default value is [value to be filled in]. The selection can be adjusted according to the screen size of the terminal device. Each candidate in the candidate list includes the Chinese characters of the term, the matching type tag, the final comprehensive weight value, and the highlighted information used for front-end display.
[0055] At the same time, InputStr and the corresponding ordered candidate list CandidateList are stored in the corresponding LRU cache instance for fast response to subsequent identical inputs.
[0056] It should be noted that the above The activation condition for the association-triggered bonus item is that after the system detects a confirmed administrative division identifier in ConfirmedTextBuffer, it automatically assigns a high weight bonus to the direct subordinate address entries of that division level. This makes the relevant address entries rank higher than non-address entries with the same basic weight in the ordered candidate list CandidateList. As a result, it automatically pushes subordinate address candidates after the user has entered the superior address, without requiring the user to re-enter the pinyin of the subordinate address.
[0057] In this embodiment, to further improve the hierarchical accuracy of address association, the following steps are included in addition to step 5: A multi-level address relationship graph of province, city, district, and street is generated and stored in a directed graph structure. Nodes represent administrative division units at various levels, directed edges represent hierarchical administrative affiliations, and the weight information of lower-level administrative division units is stored on the edges. When ConfirmedTextBuffer contains a certain administrative division unit, all its direct subordinate nodes are retrieved along the directed edges of the address relationship graph to generate a hierarchical address association candidate list, replacing the trigger word-based scanning method. The additive approach enables more precise hierarchical associations.
[0058] Step 6, perform context association extension When a user selects and confirms a candidate character or phrase (denoted as ConfirmedChar) from the ordered candidate list CandidateList, ConfirmedChar is appended to ConfirmedTextBuffer. Then, using ConfirmedChar as the search key, a hash lookup is performed in the pre-generated character / phrase association index AssociationIndex. AssociationIndex is a hash mapping structure with Chinese characters as keys and a list of common phrases containing that character as values. It is generated from phrase dictionary data during system initialization, where address phrases are pre-sorted and stored according to their frequency weights.
[0059] The search function returns a list of common phrase candidates associated with ConfirmedChar, called AssociationList. This AssociationList is then pushed to the user interface as suggested suggestions, displayed below or in the side area of the regular candidate list, allowing users to directly select subsequent phrases without having to retype the pinyin.
[0060] It should be noted that the above-mentioned character phrase association index (AssociationIndex) is generated by traversing the phrase dictionary during the system initialization phase and building a reverse index from the character to the phrase containing the character for each Chinese character in each phrase. This allows for quick retrieval of associated phrases containing any character after the user confirms it, without having to perform a full scan of the entire phrase dictionary.
[0061] In this embodiment, to further adapt to the specific needs of regional medical institutions, the following steps are added based on step 6: maintaining a regional surname weight mapping table, using the region code of the medical institution as the key and common surnames in that region and their regional weight values as the values. A regional weight correction term is introduced into the weight calculation formula. ,in This is a regional surname weighting adjustment value, which positively corrects the ranking of common local surnames in the ordered candidate list (CandidateList), enabling common local surnames to achieve higher rankings in the CandidateList.
[0062] To address the issue of medical scenario vocabulary ranking low, step 1 uses the basic frequency weights derived from medical corpora. This gives high-frequency medical terms a higher base weight than general terms in the candidate ranking, so they are naturally placed at the front of the ordered candidate list CandidateList in the descending order in step 5, reducing the number of word selection operations for medical staff.
[0063] To address the lack of differentiated processing for different field types, step 3 reads the IsNameMode flag and NameSubMode sub-pattern flag; step 4 switches the search scope to a separate name-specific Trie tree; and step 5 applies inclusion to entries under the name pattern. , and Scene bonus coefficient This completely isolates the candidate results for the name field from those for the general field in terms of retrieval path and weight calculation, thus avoiding interference from general terms on the name candidate results.
[0064] To address the inefficiency of multi-level address entry, the associative trigger enhancement item in step 5... By scanning the administrative division identifiers in the ConfirmedTextBuffer, higher weights are automatically assigned to lower-level address entries; the character phrase association index (AssociationIndex) in step 6 automatically pushes associated phrases after the user confirms the upper-level address, so that the user can complete the continuous address entry without having to re-enter the pinyin of the lower-level address.
[0065] To address the performance jitter issue caused by rapid touchscreen input, the anti-shake processing introduced in step 3 utilizes an anti-shake time window. Input events in intermediate states are filtered to ensure that the system only performs the complete retrieval process on the final state after the user input has stabilized; the dual-caching architecture (normal input mode cache instance _inputCache and name input mode cache instance _nameInputCache) caches the retrieval results of repeated input strings using LRU, which can effectively reduce the overhead of repeated calculations in medical scenarios where patient information is highly repetitive, thereby enabling the system to maintain stable response performance in high-frequency input scenarios.
[0066] Example 2 In the actual use of medical information systems, medical staff need to complete the entry of medical records in sequence according to the diagnosis and treatment process, including chief complaint, present medical history, diagnosis, and medical orders. When switching between each stage, medical staff usually start typing the content of the new stage immediately, and the initial input is often a short pinyin prefix (such as "y", "xi", "guan", etc., 1 to 3 characters).
[0067] Existing high-performance Pinyin input schemes typically employ a hotspot prefix caching strategy, maintaining a candidate list based on historical selection rates for high-frequency short prefixes, and caching them upon hit. The complexity directly returns the result, bypassing Trie tree traversal. However, the historical selection rate of the HotCache prefix cache table comes from global statistics across all diagnosis and treatment stages, without distinguishing the current diagnosis and treatment state. For example, the global high-frequency candidate for the prefix "guan" is "joint," but when medical staff jump from the stage where the patient has entered the complaint of "chest tightness + shortness of breath" to the diagnosis stage, the selection rate of the truly needed term "coronary heart disease" in the global statistics is lower than that of "joint," causing the candidate ranking returned after a cache hit to deviate from the current diagnosis and treatment context. On the other hand, although cross-state prediction based on big data association mining can identify "coronary heart disease" as a high-probability target term, its inference calculation requires traversing a large-scale association graph structure, which cannot be completed synchronously when the cache hits, causing a response timing contradiction between high-performance instant return and high-precision context prediction.
[0068] This embodiment provides a high-performance Pinyin input method based on big data for medical scenarios. It solves the problems of inaccurate hotspot prefix cache sorting and contradictory cross-state prediction response timing in the complex high-frequency scenario of short prefix input accompanied by diagnosis and treatment state transitions. This is achieved by using a cross-state directed word association graph, a hotspot prefix cache preheating injection method, and a confidence threshold-driven hierarchical response strategy.
[0069] The method described in this embodiment runs on the server or terminal device of a medical information system. It requires persistent storage medium for storing dictionary data and historical medical record data, as well as sufficient memory space for maintaining the Trie tree index structure, the HotCache hotspot prefix cache table, and the cross-state directed term association graph. The loading of dictionary data and the generation of the Trie tree index are the same as steps 1 and 2 in Example 1.
[0070] A high-performance Pinyin input method for medical scenarios based on big data, according to an embodiment of the present invention, includes the following steps: Step 2-1: Generate a cross-state directed term association graph Completed medical records are retrieved from a large-scale historical medical record database stored in persistent storage. Keyword sequences corresponding to each treatment state (chief complaint, present illness, diagnosis, and medical orders) in each record are extracted, and a cross-state directed term association graph is generated according to the chronological direction of the treatment process. .
[0071] The middle node is a medical term, and directed edges point from terms in previous states to terms in subsequent states. The edge weights are the cross-state conditional probabilities. The calculation formula is: in, These are candidate terms for subsequent states. For terms in the preceding state, This is the preliminary diagnosis and treatment status. For subsequent diagnosis and treatment status, For entries in the full medical records Appearing in state And the entry Appearing in the state that immediately follows The number of times they co-occur. For the entry Appearing in state The total number of times, both of which are dimensionless counts, and their ratio These are dimensionless probability values. Cross-state directed term association graph. The calculations are performed offline using historical medical records during the system initialization phase and then persistently stored for later retrieval at runtime.
[0072] It should be noted that the above cross-state conditional probabilities The meaning refers to the current sequence state in historical medical record big data. The term appeared in the middle. At that time, subsequent states The term appeared in the middle. The statistical probability encoded the medical causal relationships between terms at different stages of diagnosis and treatment. This probability value is related to the basic frequency weight in step 1 of Example 1. They are independent of each other and describe the statistical characteristics of the entries from different dimensions.
[0073] In this embodiment of the application, in order to improve the cross-state directed term association graph To improve the stability of low- and mid-frequency edge weights, Laplace smoothing is introduced into the edge weight calculation to smooth the numerator. add denominator add ,in For subsequent status The total number of different terms that have appeared in the sample is used to avoid large fluctuations in the edge weights of low-frequency co-occurring terms due to the scarcity of samples.
[0074] Step 2-2: Count hotspot prefixes and generate a hotspot prefix cache table. During the operation of the Pinyin input system, the cumulative number of triggers for each Pinyin prefix segment is continuously counted. And the identifier of the term selected by the user after each trigger, among which This is a prefix segment of Pinyin, and the length range of the prefix segment is... to One character, This is the maximum length of the prefix segment; the default value is [value to be filled in]. .
[0075] When the cumulative number of triggers of a certain prefix segment Exceeding the hotspot threshold At that time, calculate the historical selectability of each term in the HotCache table: in, As a candidate term, For in the prefix In all the search events triggered, the user ultimately confirms the selected term. Number of times, The cumulative number of triggers for this prefix, both being dimensionless counts, and their ratio is... This is a dimensionless probability value. (According to...) Before descending order truncation Each term and its selection rate are used to form a hotspot prefix cache entry, which is then written to the hotspot prefix cache table. . The default value is , The maximum number of terms to retain for each prefix cache entry; the default value is [value]. Both can be adjusted according to the size of the dictionary and the available memory capacity.
[0076] It should be noted that the above-mentioned hotspot prefix cache table The LRU cache instances (normal input mode cache instance_inputCache and name input mode cache instance_nameInputCache) in step 3 of Example 1 differ in function from those in the hotspot prefix cache table. The system prioritizes historical selection rates for global statistical acceleration targeting high-frequency short prefixes; LRU cache instances are managed using a least recently used strategy for immediate retrieval result reuse across any input string. Both types of caches are maintained in parallel without interference.
[0077] Steps 2-3: Detect the transition in the diagnosis / treatment status and extract the preceding keywords. Monitor changes in form field identifiers in the medical information system. When a change in the field identifier is detected... Switch to At that time, among them The current treatment status indicates a transition has occurred. The corresponding confirmed input content (i.e., the confirmed input text buffer ConfirmedTextBuffer maintained in steps 5 and 6 of Example 1) Extract keyword set from historical snapshots under the current state .
[0078] The extraction method is as follows: In the "Confirmed Input TextBuffer" state, the confirmed Chinese character string is processed for word segmentation, retaining characters with a length greater than or equal to [a certain value]. The dictionary entries are composed of Chinese characters, and stop words are filtered out to obtain... . Contains keywords ,in This represents the number of keywords extracted.
[0079] It should be noted that the above The extraction relies on the maintenance method of the ConfirmedTextBuffer in steps 5 and 6 of Embodiment 1. That is, whenever the user confirms a candidate character or phrase, the content is appended to the ConfirmedTextBuffer, from which this step reads the text. The historical fragment corresponding to the state.
[0080] Steps 2-4: Calculate predicted term weights based on the association graph and preheat the cache. by Each keyword serves as a starting node in the directed term association graph across states. Middle A one-hop graph traversal is performed on the directed edges to collect all reachable subsequent state terms, and predictive weights are calculated using a multi-keyword probability aggregation formula: in, These are candidate terms for subsequent states; This is the cross-state directed term association graph in step 2-1. The stored edge weights represent the weights in the preceding state. Keywords appearing in Subsequent state The term appeared in the middle. The conditional probability; This is the set of preceding keywords extracted in step 2-3. for The formula above is based on the assumption that each keyword's prediction of the target term is independent. It aggregates the prediction probabilities of multiple preceding keywords using a probability union, that is, it first calculates the probability of all keywords not being predicted. The product of probabilities, then... Subtracting the product yields at least one keyword support. The probability of this makes terms supported by multiple preceding keywords receive higher predictive weights.
[0081] It should be noted that when A certain keyword Related to the entry Cross-state directed term association graph When there is no directed edge in the corresponding direction, this keyword is... conditional probability Values The corresponding product factor This does not affect the aggregation results of other keywords, and the formula remains valid in this case.
[0082] For predictive weights Exceeding the preheating threshold The collection of terms, among which As a preheating threshold, extract all possible prefix segments (length) of the pinyin for each word. to ), using each prefix fragment as the key in the hotspot prefix cache table Execute the query: If the prefix already has a cached entry, predict the term and its predictive weight. The value is injected into the candidate list of the cached entry and marked as a state-aware bonus item; if there is no cached entry for the prefix, a temporary preheated cached entry is created with the predicted term as the initial candidate and marked as a state-associated temporary entry. The default value is It can be adjusted according to the business scenario.
[0083] It should be noted that the aforementioned preheating injection operation is performed before the transition between treatment states, rather than being calculated in real time when the user types in pinyin. This makes the cross-state directed term association graph... The traversal computation and subsequent user input response processing are decoupled in time—a cross-state directed term association graph. The traversal calculation is completed during the intervals between state transitions, and subsequent short prefix inputs hit the hotspot prefix cache table. There is no need to perform graph traversal again.
[0084] In this embodiment of the application, to improve the prediction coverage, the following step is also included in addition to steps 2-4: processing the cross-state directed term association graph. Predictive weight of terms reachable in one jump The preheating threshold was not exceeded. In the case of two-hop reachability, a second-hop traversal is further performed along the directed edges to collect two-hop reachable terms. The two-hop predictive weights are then calculated using the same multi-keyword probability aggregation formula. For terms exceeding the two-hop warm-up threshold... The same cache preheating injection operation is performed on the entries, where This is the second-hop warm-up threshold; the default value is [value to be filled in]. ,satisfy This is to cover indirectly related terms that involve intermediate steps in the diagnosis and treatment process.
[0085] Steps 2-5: Receive the input string and perform state-aware cache retrieval. Steps 2-5 may specifically include: Similar to step 3 of Embodiment 1, this involves receiving input events and obtaining InputStr, as well as performing a hit query on the dual-cached instance. Based on this, the prefix fragment of InputStr is used in the hotspot prefix cache table. Perform exact match queries.
[0086] If the hotspot prefix cache table is hit For each candidate term in the cached entries (including original hot cache entries and state-associated temporary entries injected in steps 2-4), ... Calculate the state-aware fusion score. (Due to historical selection rates) Predictive weights and fundamental frequency weight The three terms have different dimensions, so their dimensions need to be unified before being included in the weighted summation: Already The probability value of the interval. Already The probability value of the interval. Mean normalization based on range is applied to... An interval, after normalization, is denoted as The three unifications to After the interval, the fusion score is calculated using the following formula: in, Candidate terms; This refers to the historical selection rate statistically analyzed in step 2-2; The predictive weights calculated in steps 2-4 are assigned values for terms that have not been preheated. ; Basic frequency weight After mean normalization based on range The value after the interval; The historical selection rate is integrated with the weighting coefficient. For predictive weighting, the weighting coefficients are combined. Based on the fundamental frequency weight fusion weight coefficient, the three satisfy the following: The default values are respectively , , It can be adjusted according to the business scenario.
[0087] By fusion score Generate a state-aware fast candidate list by sorting in descending order. .
[0088] It should be noted that the above fusion score Medium predictive weights The value for the term "without preheating injection" is... This means that terms that are only supported by historical selection rates but not by predictions of the current treatment status do not have a prediction bonus in their fusion score. As a result, the ranking of high-frequency global terms that do not fit the current treatment context will naturally decrease after the status jump.
[0089] In this embodiment of the application, to support adaptive adjustment of the fusion weight coefficients, the following step is added in addition to steps 2-5: Statistically counting the types of each diagnosis and treatment status transition. Below, the term with the highest historical selection rate and its combined rating. The percentage of users who ultimately confirm the top-ranked term is used as a feedback signal to refine the algorithm using a gradient descent approach. , , Periodic updates are performed to gradually converge the fusion weight coefficients to their optimal values.
[0090] Steps 2-6: Perform hierarchical response output based on confidence thresholds. Determine the state-aware fast candidate list The top-ranked term in terms of fusion score Does it exceed the confidence threshold? ,in The combined score for the top-ranked term. This is the confidence threshold.
[0091] like Directly output a state-aware fast candidate list As the final candidate list, completely bypassing Trie tree traversal, the response time complexity is O(n log n). .
[0092] like State-aware fast candidate list As a priority candidate, a supplementary traversal search is performed on the Trie tree generated in step 2 of embodiment 1 (selected according to the input pattern determined in step 3 of embodiment 1) using InputStr as the query string. The term identifiers already included in the priority candidates are excluded, and the terms collected by the supplementary traversal are combined with the state-aware fast candidate list. The pool is expanded by merging candidates, and the final comprehensive weights are determined according to step 5 of Example 1. The formula performs precise weight calculations on all terms in the expanded candidate pool and outputs the final candidate list in descending order. The default value is It can be adjusted according to the business scenario.
[0093] It should be noted that in the above hierarchical response strategy, the supplementary traversal in the low-confidence scenario is only performed on the terms not covered by the priority candidates, rather than performing a full traversal of the entire Trie tree. Therefore, the computational scale of the supplementary traversal is affected by the state-aware fast candidate list. The constraint of the number of covered terms can effectively reduce the overall average search cost when the proportion of high-confidence scenarios is relatively high.
[0094] Steps 2-7: Update the cache and clear preheating data during state transitions. Update the hotspot prefix cache table based on the user's final confirmed selection of keywords. The corresponding prefix cache entry Count and recalculate the historical selection rate At the same time, update the LRU cache instance in step 5 of embodiment one.
[0095] Hotspot prefix cache table For cache entries that have not been triggered in the medium to long term, expiration and eviction will be performed according to the time window decay strategy: set the effective time window for cache entries. ,in The valid time window for cached entries exceeds [a certain period]. Untriggered entries are retrieved from the hotspot prefix cache table. Removed from the middle The default value is The time limit can be adjusted according to system load.
[0096] When a new transition in the diagnosis / treatment status is detected, all temporary state-associated entries and state-aware bonuses injected in steps 2-4 are cleared, and a new... and Repeat steps 2-3 and 2-4 to generate predicted term weights corresponding to the new diagnosis and treatment status and complete cache preheating injection.
[0097] It should be noted that the cache cleanup operation during the state transition only clears the temporary state-related entries and state-aware bonus items injected in steps 2-4, and does not affect the hotspot prefix cache table. The original hotspot cache entries are formed based on historical selection rate statistics to ensure that the continuous accumulation of global statistical information is not lost due to state switching.
[0098] To address the issue of inaccurate sorting of hotspot prefix cache after transitions in treatment status, the cross-state directed term association graph generated in step 2-1... The code encodes the causal relationships between terms across different diagnostic and treatment stages contained in the big data. Steps 2-4 execute the cross-state directed term association graph before the state transition. Iterate through and inject the prediction results into the hotspot prefix cache table. The fusion score in steps 2-5 Formula passed The method incorporates cross-state predictive weights into the candidate ranking, which gives a bonus to terms supported by the current diagnosis and treatment context (such as "coronary heart disease" after jumping from the chief complaint of "chest tightness + shortness of breath" to the diagnosis stage). This relatively reduces the ranking of terms that are only supported by the global historical selection rate but do not match the current diagnosis and treatment state (such as "joints"). As a result, when the cache is hit, the candidate ranking result that matches the current diagnosis and treatment context can be obtained.
[0099] To address the timing conflict between cross-state predictive inference and high-performance cached instant return, the warm-up injection operation in steps 2-4 is completed before the diagnosis / treatment state transition, rather than being triggered in real-time when the user types in pinyin. This ensures that the cross-state directed term association graph is properly configured. Traversal computation and user input response are decoupled in time—subsequent short prefix inputs hit the hotspot prefix cache table. At that time, the fusion score incorporated predictive weights The calculations have been completed in advance, and the response time complexity remains at [value missing]. level.
[0100] For overall input performance in low-confidence scenarios with a large-scale medical thesaurus, the confidence threshold in steps 2-6 is crucial. The hierarchical response strategy completely bypasses Trie tree traversal in high-confidence scenarios, while in low-confidence scenarios, it only performs exclusionary supplementary traversal on terms not covered by priority candidates, instead of a full traversal. This reduces the overall average retrieval cost as the proportion of high-confidence scenarios increases. The state transition cleanup operations in steps 2-7 ensure the hotspot prefix cache table... The predictive additives remain consistent with the current diagnostic and treatment context, avoiding interference from the prediction results of historical states on the candidate ranking of new states.
[0101] The following is an example of an application of the present invention, such as... Figure 2-9 As shown, the implementation process is as follows: On a weekday morning in 20XX, a nurse entered basic information and medical records of a newly admitted patient into a hospital's information system. The patient was male, and the system had already recorded his gender code. The nurse sequentially entered the patient's name and address fields, then proceeded to the chief complaint section, entering "chest tightness and shortness of breath," before switching to the diagnosis section and entering the short prefix "guan." The following describes the complete data flow process according to the core steps of Examples 1 and 2.
[0102] Step 1: Load medical corpus dictionary data. The system reads three types of dictionary data from persistent storage. Taking several entries in the dictionary of words specifically for names as an example, each record contains Chinese characters, pinyin string, first letter, type identifier NameType, gender bias weight GenderBias, and basic frequency weight. ,in Total number of entries in the corpus Scaling factor .
[0103] Table 1 Examples of Dictionary Entries for Name-Specific Characters Step 2 generates a multi-level Trie tree index. Based on the dictionary data from Step 1, the system constructs a Pinyin Trie tree, a phrase Trie tree, a continuous Pinyin Trie tree, an initial letter Trie tree, a mixed input Trie tree, and a name-specific Trie tree. Taking the address entry "Guangzhou City" from the phrase dictionary as an example, the valid paths inserted into the mixed input Trie tree include "guangzhoushi" (full Pinyin), "gzs" (full initial letters), "guangzs" (full Pinyin of the first two characters plus the initial letter of the last character), and "gzhoushi" (full Pinyin of the first character plus the last two characters), etc., each path pointing to the same entry reference. The name-specific Trie tree contains only entries from the name-specific dictionary and is completely isolated from the paths in the general Pinyin Trie tree.
[0104] Step 3 receives the input string and determines the input mode. The nurse's current input field is the patient's name; the system sets IsNameMode to true and NameSubMode to GIVEN_NAME_ONLY, and includes the patient's gender code. (Male). The nurse types "wei", and the system reads InputStr as "wei". Debounce timer setting delay. If the nurse does not trigger any new input within 80 milliseconds, the timer expires and the subsequent retrieval process is executed with "wei". The name input pattern cache instance (_nameInputCache) is queried using "wei+GIVEN_NAME_ONLY" as the key, but no match is found, proceeding to step 4.
[0105] Step 4 performs multi-mode parallel retrieval. The input mode is the name input mode, and the retrieval range is limited to the dedicated name hash index and the dedicated name Trie tree. Prefix search is performed in the dedicated name Trie tree with "wei" as the query string, and all entries whose Pinyin paths are prefixed with "wei" are returned; meanwhile, the entry "wei" corresponding to the query "wei" is searched by exact matching in the dedicated name hash index. NameSubMode is GIVEN_NAME_ONLY, filtering is performed on the leaf node entry reference list, and only entries with NameType being GIVEN_NAME are retained. In the retrieval results, "wei" (the Chinese character "伟") hits the exact matching path (ExactMatchSet), and other prefixed entries enter the prefix matching set (PrefixMatchSet), and the matching type tag is recorded as PINYIN_PREFIX.
[0106] Step 5 calculates multi-dimensional dynamic weights and sorts. Calculate the final comprehensive weight for each candidate entry in the retrieval results . Taking the entry "伟" as an example, it comes from the exact matching path, ; the scenario bonus coefficient , wherein , NameType being GIVEN_NAME matches NameSubMode, , , therefore ; non-address entry . The of "伟". For the entry "芳", GenderBias is -0.90, , , (Note: is a negative value. After being multiplied by a larger positive coefficient, the absolute value becomes larger, that is, the ranking is lower; here, for weight sorting, a larger value has priority, that is, a smaller absolute value ranks higher). The following table shows the weight calculation results of each candidate entry.
[0107] Table 2 Weight calculation results of candidate entries in name mode (input "wei", male patient) It should be noted that, when is a negative value, a smaller absolute value indicates a higher basic frequency; when they are both negative, a smaller absolute value indicates a higher comprehensive weight and a higher ranking. For "伟", since GenderBias is in the same direction as that of the male patient, is larger, but the absolute value increases, and the ranking is higher, which reflects the positive contribution of gender tendency weighting to the ranking. After the ordered candidate list is generated, the result is stored in _nameInputCache with "wei+GIVEN_NAME_ONLY" as the key.
[0108] After the nurse confirms the selection of the character "Wei", the system appends "Wei" to ConfirmedTextBuffer, performs a hash search using "Wei" as the key in the character-phrase association index AssociationIndex (step 6), returns a candidate list of common phrases containing "Wei", and pushes it to the interface for the nurse to select directly.
[0109] Then the nurse switches to the address field (sets IsNameMode to false and enters the normal input mode), and types "gz" after entering "Guangdong Province" in ConfirmedTextBuffer. In step 5, the system scans ConfirmedTextBuffer, detects the trigger word "Province", and assigns an association trigger bonus to municipal address entries under the province (such as "Guangzhou", "Shenzhen", etc.) , which by default takes a value of 3 times the maximum basic weight value of the current thesaurus, so that municipal address entries such as "Guangzhou" are prioritized in the candidate list, and the nurse can select the subordinate address without retyping the complete pinyin.
[0110] Thereafter, the nurse switches to the chief complaint field, enters "chest stuffiness shortness of breath" and confirms, ConfirmedTextBuffer records "chest stuffiness and shortness of breath" in the chief complaint state. When the nurse switches to the diagnosis field, the transition of diagnosis and treatment state is triggered, which is the chief complaint, and this is the diagnosis.
[0111] Steps 2-3 detect that the field identifier switches from chief complaint to diagnosis, and extract the keyword set from the historical snapshot of ConfirmedTextBuffer in the chief complaint state : word segmentation is performed on "chest stuffiness and shortness of breath", entries with a length of no less than 2 Chinese characters are retained and stop words are filtered, to obtain $K_{prev} = \{"chest stuffiness", "shortness of breath"\}$.
[0112] In step 2-4, with each keyword in as the starting node, in the cross-state directed entry association graph a one-hop graph traversal is performed along the chief complaint-diagnosis direction. Taking "coronary heart disease" as the target entry as an example, the edge weights counted from historical medical record data are shown in the following table.
[0113] Table 3 Partial edge weights in the chief complaint-diagnosis direction Calculate the predictive weight of "coronary heart disease" according to the multi-keyword probability aggregation formula: , which exceeds the preheating threshold Extract all prefix segments (length 1 to 6) of the pinyin for "coronary heart disease" ("guanxinbing"). Using each prefix segment as a key, perform a query in the HotCache table to retrieve "coronary heart disease" and its associated prefixes. Inject cache entries corresponding to the prefixes “g”, “gu”, “gua”, “guan”, “guanx”, and “guanxi”, and mark them as state-aware bonus items.
[0114] In steps 2-5, the nurse types "guan," and an exact match query is performed in the HotCache table using "guan." The query finds a cached entry (containing the original hot term "joint" and the "coronary heart disease" injected in step 2-4). A state-aware fusion score is then calculated for each candidate term in the cached entry. The three weighting coefficients are set by default. , , , After range normalization, we get .
[0115] Table 4. Calculation of fusion score after the prefix "guan" hits the cache. Steps 2-6 determine the state-aware fast candidate list The highest-ranked term in China is the fusion score for "coronary heart disease". Below the confidence threshold The system will then enter a low-confidence, graded response path. (System reservation) As a priority candidate, a supplementary traversal is performed on the Pinyin Trie tree using "guan" to exclude entries already included in the priority candidate list. These supplementary entries are then merged with the priority candidates to form an expanded candidate pool. After the formula performs precise weight calculation, the final candidate list is output in descending order, with "coronary heart disease" as the cause. The bonus is still listed at the front.
[0116] After the nurse confirms the selection of "coronary heart disease" in steps 2-7, the system updates the entries corresponding to prefixes such as "guan" in the HotCache table. Count and recalculate Simultaneously, the LRU cache instance is updated. When a nurse switches treatment fields again, the system clears all state-associated temporary entries and state-aware bonus items injected in steps 2-4, and updates the new cache. (Diagnosis) and (Medical order) Re-execute steps 2-3 and 2-4 to generate predicted term weights corresponding to the new diagnosis and treatment status and complete cache preheating injection. The original hotspot cache entries in the HotCache table, which are formed based on historical selection rate statistics, are not affected.
[0117] The entire data flow process starts with step 1, dictionary loading; step 2, index construction to form multiple types of Trie trees; step 3, debouncing and cache hit control to control the retrieval entry point; step 4, multi-mode parallel retrieval to generate a candidate set; step 5, multi-dimensional weight calculation to sort the candidate set and write it into the LRU cache; and step 6, association expansion to push related word groups. In Example 2, steps 2-1 to 2-4 complete graph traversal and cache preheating injection before the state transition; step 2-5, fusion score, is directly calculated when the cache hits; step 2-6, confidence threshold determines whether to trigger supplementary traversal; and step 2-7, maintains cache consistency. The input and output data between each step form a complete logical link, ensuring that every link from dictionary data to the final candidate list has data basis.
[0118] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A high-performance Pinyin input method based on big data for medical scenarios, characterized in that, include: Obtain a general single-character dictionary, a phrase dictionary, and a dictionary of characters specific to names; parse the entry information and obtain the basic frequency weights based on word frequency statistics from a medical corpus. Based on the dictionary data, generate the following Trie trees: Pinyin Trie tree, Phrase Trie tree, Continuous Pinyin Trie tree, First Letter Trie tree, Mixed Input Trie tree, and Name-specific Trie tree. Receive the input string, determine the input mode based on the current input field type, and perform a cache hit query in the corresponding cache instance according to the input mode. The input mode is either name input mode or normal input mode. Based on the input pattern, multi-pattern parallel retrieval is performed in the corresponding Trie tree using the input string to generate a candidate term set; For each term in the candidate term set, the final comprehensive weight is calculated based on the basic frequency weight, the matching degree bonus coefficient, and the scenario bonus coefficient. The terms are then sorted in descending order to generate an ordered candidate list and stored in the corresponding cache instance.
2. The method according to claim 1, characterized in that, The basic frequency weight is calculated as follows: the ratio of the frequency of each word in the medical corpus plus one to the total number of words in the corpus is taken as the logarithm, and then multiplied by the normalization scaling factor to map the original frequency to the weight range calculated by integer sorting. The normalization scaling factor is determined based on the size of the lexicon.
3. The method according to claim 1, characterized in that, The mixed input Trie tree supports a mixed input mode of full Pinyin syllables and initial letter abbreviations. For each Chinese character in a phrase, its complete Pinyin string or the first letter of Pinyin is used as the path segment. All legal mixed representations are enumerated and inserted into the mixed input Trie tree respectively. The hybrid input Trie tree also supports a partial syllable abbreviation mode, where the syllables of a word are represented by the first few letters of the syllable, in addition to the complete pinyin string and the first letter.
4. The method according to claim 1, characterized in that, In the name input mode, the search scope is limited to the name-specific Trie tree; the search results are filtered by type according to the name sub-pattern flag: when the name sub-pattern flag indicates only the last name, only the last name type entries are retained; when the name sub-pattern flag indicates only the first name, only the first name type entries are retained; when the name sub-pattern flag indicates the full name, all name type entries are returned.
5. The method according to claim 4, characterized in that, In the name input mode, the scenario bonus coefficient is determined by the name character enhancement factor, the name type differentiation weighting coefficient, and the gender tendency weighting coefficient. The name type differentiation weighting coefficient takes a positive value when the term type matches the current name sub-pattern flag, and takes zero when they do not match. The gender bias weighting coefficient is determined by the product of the gender bias weight of the term, the patient's gender code, and the gender weight adjustment coefficient. When the gender bias of the term matches the patient's gender, a positive bonus is generated; when they do not match, a negative inhibition is generated. In normal input mode, the scenario bonus coefficient is always one.
6. The method according to claim 1, characterized in that, Also includes: The confirmed input text buffer is scanned for address-level trigger words. When an administrative division level identifier word is scanned, an association trigger bonus item is set for the address word associated with the next level. The association trigger bonus item is added to the final comprehensive weight. Once a user confirms a candidate character or phrase, a hash lookup is performed in the pre-generated character and phrase association index using the confirmed content as the search key. This returns a list of associated common phrase candidates and pushes it to the user interface. The character phrase association index is generated by traversing the phrase dictionary data during the system initialization phase, using Chinese characters as keys and a list of common phrases containing that character as values.
7. The method according to claim 1, characterized in that, When receiving the input string, a debouncing process is introduced. The debouncing process sets a delay timer for each input event, and the delay duration is the debouncing time window. When the user triggers multiple input events consecutively within the debouncing time window, the previous timer is canceled and a new timer is set. The subsequent retrieval process is executed only when the user does not trigger any new input events within the debouncing time window, using the latest input string. The cache instances are divided into normal input mode cache instances and name input mode cache instances, both of which are managed using the least recently used strategy.
8. The method according to claim 1, characterized in that, Also includes: Extract keyword sequences corresponding to each diagnosis and treatment status from historical medical record data, and generate a cross-state directed term association graph according to the time sequence of the diagnosis and treatment process. The nodes are medical terms, the directed edges point from the terms of the previous state to the terms of the subsequent state, and the edge weights are the cross-state conditional probabilities. The cumulative number of triggers for each pinyin prefix segment and the final word selected by the user are counted. When the cumulative number of triggers exceeds the hotspot threshold, hotspot prefix cache entries are generated according to the historical selection rate and written to the hotspot prefix cache table. When a change in treatment status is detected, the set of preceding keywords is extracted from the confirmed content of the previous status. Using the set of preceding keywords as the starting node, a graph traversal is performed in the cross-state directed term association graph along the state change direction. The predictive weight is calculated according to the multi-keyword probability aggregation method. Term terms with predictive weights exceeding the preheating threshold are injected into the hotspot prefix cache table.
9. The method according to claim 8, characterized in that, Also includes: The input string is used to perform a matching query in the hotspot prefix cache table. When a match is found, a state-aware fusion score is calculated for each candidate term in the cache entry based on the historical selection rate, predictive weight, and normalized basic frequency weight. The fusion scores are then sorted in descending order to generate a state-aware fast candidate list. Determine whether the fusion score of the first-ranked term in the state-aware fast candidate list exceeds the confidence threshold: if it does, directly output the state-aware fast candidate list as the final candidate list; If the number of candidates is not exceeded, the state-aware fast candidate list is retained as the priority candidate. At the same time, a supplementary traversal search is performed in the Trie tree using the input string. The supplementary terms are merged with the priority candidates and then sorted and output according to the final comprehensive weight. When a change in treatment status is detected, the injected temporary entries associated with the status and the status-aware bonus items are cleared, and keyword extraction and cache preheating injection are re-executed with the new preceding status and the current status.
10. A high-performance Pinyin input system based on big data for medical scenarios, used to perform the method described in any one of claims 1 to 9, characterized in that, include: The dictionary loading module is used to obtain a general single-character dictionary, a phrase dictionary, and a dictionary of characters specific to names, parse the entry information, and obtain the basic frequency weights based on word frequency statistics from a medical corpus. The index generation module is used to generate, based on the dictionary data, a Pinyin Trie tree, a phrase Trie tree, a continuous Pinyin Trie tree, an initial letter Trie tree, a mixed input Trie tree, and a name-specific Trie tree. The input receiving and mode determination module is used to receive the input string, determine the input mode according to the current input field type, and perform a cache hit query in the corresponding cache instance according to the input mode. The input mode is either the name input mode or the normal input mode. The multi-mode retrieval module is used to perform multi-mode parallel retrieval in the corresponding Trie tree based on the input mode and the input string to generate a candidate term set; The weight calculation and sorting module is used to calculate the final comprehensive weight for each term in the candidate term set based on the basic frequency weight, matching degree bonus coefficient and scenario bonus coefficient, and generate an ordered candidate list in descending order and store it in the corresponding cache instance.