A human-machine interaction method and system for legal consultation
By dynamically adjusting the Hoffman tree structure in legal consultation human-computer interaction, dividing sets according to the uniformity and importance of words, and optimizing coding strategies, the efficiency of traditional Hoffman coding when introducing new topics is solved, data transmission efficiency and system adaptability are improved, and user experience is improved.
Patent Information
- Application Number
- CN202510436341.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In legal consultation human-computer interaction, traditional Hoffman coding has a very long encoding length when the new topic was introduced and the static code table cannot perceive the word frequency changes in the conversation content in real time, resulting in a significant decrease in compression efficiency.
By obtaining interactive data and counting word frequency information, calculating the distribution uniformity of words, dividing dynamic sets and static sets, adjusting the Hoffman tree nodes according to the frequency changes of words during the interaction process, optimizing coding strategies, dynamically inserting new words, and optimizing data compression and transmission.
It improves data transmission efficiency and system real-timeness, reduces computing complexity, enhances the system's robustness and resource utilization efficiency, adaptability and stability, and improves user experience.
Smart Images

Figure CN119938873B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data transmission. In particular, it relates to a human-computer interaction method and system for legal consultation. Background Art
[0002] Legal consultation is a professional legal service aimed at helping individuals or organizations solve legal problems, provide legal opinions, and guide actions. Legal consultation is carried out through face-to-face, telephone, online, or written means, and its functions are: providing legal opinions, assessing risks, guiding actions, and preventing disputes.
[0003] With the development of technology, human-computer interaction is applied in various fields, especially in chatbots and intelligent question-answering systems, significant progress has been made. In this context, combining legal consultation with human-computer interaction, with the help of artificial intelligence technology and natural language processing capabilities, can provide efficient and convenient legal services for the public and legal professionals. It not only improves the efficiency and accessibility of legal consultation but also promotes the digital transformation of the legal service industry. However, this process also faces many challenges. How to efficiently manage and process a large amount of human-computer interaction data has become a technical problem to be solved urgently. In particular, how to optimize data transmission and processing efficiency while ensuring service quality has become the focus of the industry's attention.
[0004] In the real-time interaction scenario of legal consultation, the user's conversation topic will change dynamically with the progress of the consultation. As the scale of legal consultation services continues to expand, two significant defects of traditional Huffman coding in interactive data compression have gradually emerged: First, when a new topic is introduced, some professional terms have not been pre-counted, and the coding lengths calculated based on the global word frequency of these terms are relatively long, resulting in a significant decrease in compression efficiency. Second, the static code table cannot real-time perceive the dynamic changes in the word frequency of the conversation content. As time goes by, the compression efficiency continues to decrease. Summary of the Invention
[0005] To solve the problems that when human-computer interaction occurs, the coding length becomes too long when a new topic is introduced, and the static code table cannot real-time perceive the dynamic changes in the word frequency of the conversation content, resulting in a significant decrease in compression efficiency, the present invention provides solutions in the following aspects.
[0006] In a first aspect, a human-computer interaction method for legal consultation includes: obtaining interaction data and counting word frequency information; calculating the distribution uniformity of words based on the index positions of the words in the interaction data, and dividing each word into sets based on the distribution uniformity and a preset uniformity threshold, where the sets include: a dynamic set and a static set; based on the words in the dynamic set, obtaining the importance of the words according to the change in the occurrence frequency of the words during the interaction process, adjusting the Huffman tree nodes corresponding to the words in the dynamic set, optimizing the encoding strategy, and compressing and transmitting the interaction data according to the adjusted Huffman tree.
[0007] By obtaining interaction data and counting word frequency information, and then calculating the distribution uniformity of words based on the index positions of the words in the interaction data. Based on the distribution uniformity and a preset uniformity threshold, the words are divided into a dynamic set and a static set. The words in the dynamic set obtain importance according to the change in their occurrence frequency during the interaction process, and accordingly adjust the Huffman tree nodes, optimize the encoding strategy, and finally compress and transmit the interaction data according to the adjusted Huffman tree. It significantly improves the efficiency of data transmission and the real-time performance of the system. At the same time, by reasonably dividing the dynamic set and the static set, the computational complexity is reduced, and the robustness and resource utilization efficiency of the system are improved. By monitoring the appearance of new words and reasonably inserting them into the Huffman tree, the system's ability to handle new vocabulary is ensured, further enhancing the adaptability and stability of the system. These improvements not only enhance the performance of the system but also significantly improve the user experience.
[0008] Preferably, the distribution uniformity includes:
[0009] Starting from the first occurrence of any word in the interaction data, calculate the index position interval between two adjacent occurrences in sequence, calculate the standard deviation of all index position intervals, and obtain the degree of dispersion of the index position intervals of the word;
[0010] Take the ratio between the position interval between the first occurrence and the last occurrence of the word and the total length of the interaction data as the distribution range of the word;
[0011] Normalize the ratio between the distribution range and the degree of dispersion to obtain the distribution uniformity of the word. The distribution uniformity is positively correlated with the distribution range and negatively correlated with the degree of dispersion.
[0012] By accurately calculating the degree of dispersion and the distribution range of the index position intervals of words and normalizing the ratio of the two to obtain the distribution uniformity of words, the distribution characteristics of words in the interaction data can be measured more accurately. By calculating the distribution uniformity, words can be more reasonably divided into a dynamic set and a static set, and then the adjustment strategy of the Huffman tree can be optimized. This not only improves the data compression efficiency, reduces the computational complexity, but also enhances the adaptability of the system to changes in the conversation topic, ultimately improving the user experience and system performance.
[0013] Preferably, the distribution uniformity includes:
[0014] Divide the interaction data into several paragraphs, count the number of times a word appears in each paragraph, calculate the frequency of the word in each paragraph, and use the entropy calculation formula to obtain the distribution uniformity of each word.
[0015] By dividing the interaction data into several paragraphs, counting the number of times each word appears in each paragraph and calculating its frequency, and then using the entropy calculation formula to measure the distribution uniformity of each word, the distribution of words in different paragraphs can be captured more meticulously. Through the quantification of the distribution uniformity, the system can more accurately divide words into dynamic sets and static sets, and then optimize the adjustment strategy of the Huffman tree.
[0016] Preferably, the dividing of each word into sets includes:
[0017] Divide the words with distribution uniformity less than the preset uniformity threshold into the dynamic set, and vice versa, divide the words greater than or equal to the preset uniformity threshold into the static set.
[0018] Preferably, the importance of the word includes:
[0019] Taking any word in the dynamic set as the target word, record the frequency of occurrence of the target word in each interaction and the total number of interactions, where one interaction occurs when the user in the dynamic set asks and the system answers once;
[0020] Calculate the frequency difference of the target word in adjacent interaction times to obtain the dynamic change situation;
[0021] Use the exponential decay function to obtain the time weight of each interaction and perform normalization processing, sum the dynamic change situations of all interactions multiplied by the normalized time weight to obtain the weighted frequency change, divide the weighted frequency change by the distribution uniformity and perform normalization processing to obtain the importance of the target word.
[0022] By accurately measuring the importance of words in the interaction process, especially paying attention to the dynamic changes of word frequency and time sensitivity, it provides a basis for dynamically adjusting the Huffman tree, optimizes the coding strategy, improves the data compression efficiency and the real-time performance of the system, enhances the system's adaptability to changes in the conversation topic, improves the user experience and system performance, and is especially suitable for scenarios such as legal consultation that require efficient processing of dynamic interaction data.
[0023] Preferably, the importance of the word includes:
[0024] Taking any interaction as a marked interaction and any word in the dynamic set as a target word, calculate the ratio of the occurrence frequency of the target word in the marked interaction to the total number of times the target word appears in the marked interaction data to obtain the relative frequency. Take the sum of the relative frequencies of all target words as the overall importance;
[0025] Calculate the overall importance divided by the distribution uniformity and then perform normalization to obtain the importance of the target word.
[0026] By considering the occurrence frequency of the target word in each interaction and its distribution characteristics, it is possible to more comprehensively evaluate the importance of the word, provide a more accurate basis for dynamically adjusting the Huffman tree, optimize the encoding strategy, improve the data compression efficiency and the real-time performance of the system, and enhance the system's adaptability to changes in the conversation topic.
[0027] Preferably, the importance of the word includes:
[0028] Taking any interaction as a marked interaction and any word in the dynamic set as a target word, calculate the ratio of the occurrence frequency of the target word in the marked interaction to the total number of times the target word appears in the marked interaction data to obtain the relative frequency. Take the sum of the relative frequencies of all target words as the overall importance;
[0029] Calculate the frequency proportion of the target word in all interactions. Divide the product of the frequency proportion and the overall importance by the distribution uniformity and then perform normalization to obtain the importance of the target word.
[0030] Preferably, adjusting the Huffman tree nodes corresponding to the words in the dynamic set includes:
[0031] Preset an adjustment period, traverse the Huffman tree, and mark the nodes corresponding to all the words in the dynamic set. Among them, the nodes corresponding to the words in the static set in the Huffman tree remain unchanged;
[0032] Sort all the words in the dynamic set from largest to smallest according to their importance, and fill the sorted words into the marked nodes in the Huffman tree from top to bottom in sequence, and obtain the codes corresponding to each word after adjustment.
[0033] Preferably, compressing and transmitting the interaction data according to the adjusted Huffman tree includes:
[0034] During the interaction data compression process, monitor the appearance of new words, determine whether the new words already exist in the Huffman tree. In response to the new words not existing in the Huffman tree, insert the new words into the last node of the Huffman tree. Otherwise, do not insert. After inserting the new words, complete the data compression.
[0035] Second aspect, a human-computer interaction system for legal consultation, comprising: a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned human-computer interaction method for legal consultation is implemented.
[0036] The present invention has the following effects:
[0037] 1. By dynamically adjusting the structure of the Huffman tree, the present invention enables frequently occurring words to be compressed into shorter codes, thereby reducing the bandwidth occupancy of data transmission and improving the data transmission rate. It is applicable to the dynamic changes in the word frequency distribution caused by the dynamic changes in the conversation topics in the legal consultation scenario. Through the division of the dynamic set and the static set, the system can more efficiently process newly emerging professional terms and common words, significantly improving the compression efficiency.
[0038] 2. By adjusting the words in the dynamic set, the system can dynamically generate a Huffman tree more suitable for the current conversation state according to the word frequency changes in the real-time conversation, insert the new word into the last node of the Huffman tree, and dynamically adjust its position closer to the root node when the new word appears frequently, thereby optimizing the coding strategy and improving the compression efficiency. It is beneficial for the system to quickly adapt to the changes in the conversation topics and maintain high-efficient coding and transmission performance, especially suitable for the legal consultation scenarios with high concurrency and real-time interaction.
[0039] 3. By dividing the dynamic set and the static set according to the distribution uniformity of words, and only performing frequency change analysis and Huffman tree adjustment on the dynamic set, the computational complexity is greatly reduced. Since the words in the static set have a relatively high distribution uniformity and do not need to be adjusted frequently, the computational resources are saved, which is beneficial to significantly reducing the computational burden of the system while ensuring the compression efficiency and improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the method for steps S1 - S3 in a human-computer interaction method for legal consultation according to an embodiment of the present invention.
[0041] Figure 2 is a schematic diagram of the Huffman tree after adding a new word in a human-computer interaction method for legal consultation according to an embodiment of the present invention.
[0042] Figure 3 is a structural block diagram of a human-computer interaction system for legal consultation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0044] Refer to Figure 1 , a human-computer interaction method for legal consultation includes steps S1 - S3, specifically as follows:
[0045] S1: Obtain interaction data and count word frequency information.
[0046] The interaction system records the user and system outputs to the database through the API to store the interaction data. Before word segmentation, the interaction data is preprocessed, such as removing stop words, punctuation marks, and special characters, to improve the accuracy and efficiency of word segmentation. The stored interaction data is segmented using jieba segmentation to obtain the segmentation result. Among them, jieba segmentation is widely used in the field of natural language processing, and its core function is to split a continuous Chinese interaction data into several words. Jieba segmentation can accurately perform word segmentation through the maximum probability word segmentation method based on the prefix dictionary. For the segmentation result, count the occurrence frequency of each word segment and the index position of each analysis.
[0047] It should be noted that in legal consultation services, human-computer interaction data specifically refers to all interaction information that occurs between users and legal consultation systems, including but not limited to the following three parts:
[0048] User input information: All specific interaction data content such as legal consultation questions, statements, and requirements provided by users when communicating with the system through interaction data;
[0049] System response content: Information such as legal consultation suggestions, answers, relevant regulations, and case analyses output by the system after natural language processing based on user input;
[0050] Dialogue history information: Context information that needs to be retained during the continuous dialogue process of the system, including but not limited to previous consultation questions, answer records, etc., to ensure the coherence and pertinence of subsequent exchanges.
[0051] It should be noted that in a real-time interactive legal consultation system, the conversation topics between users and the system often change dynamically. As the conversation progresses, the specific legal issues or requirements consulted by users will continue to change, resulting in the adjustment of the conversation topics. The traditional Huffman coding algorithm uses static frequency estimation of the emergency pre-global word frequency and constructs a coding table using the pre-computed word frequency distribution. This static frequency estimation method has obvious limitations. For example, as the conversation topic changes, the frequency of some words related to old topics may drop sharply, while the frequency of words related to newly emerging topics may rise rapidly. Since the Huffman coding uses a fixed coding table throughout the coding process, this frequent change in fields will lead to a significant decrease in coding efficiency and ultimately affect the compression performance of the entire real-time interactive system. The specific adjustment steps are as follows:
[0052] S2: Calculate the distribution uniformity of the word according to the index position of the word in the interaction data, and divide each word into sets based on the distribution uniformity and a preset uniformity threshold. Among them, the sets include: a dynamic set and a static set.
[0053] It should be noted that the traditional static Huffman coding table cannot perceive the dynamic changes in word frequency. If all words are counted and adjusted, it will consume a large amount of computing power. Therefore, the words are divided into a dynamic set and a static set. The words in the static set belong to common words or fixed vocabulary, and their frequencies of occurrence in the conversation change little, so they do not need to be adjusted frequently. This strategy reduces the complexity of the structure adjustment of the Huffman tree.
[0054] The distribution uniformity includes:
[0055] Starting from the first occurrence of any word in the interaction data, calculate the index position interval between two adjacent occurrences in turn, and calculate the standard deviation of all index position intervals to obtain the degree of dispersion of the index position intervals of the word;
[0056] Take the ratio between the position interval from the first occurrence to the last occurrence of the word and the total length of the interaction data as the distribution range of the word;
[0057] Normalize the ratio between the distribution range and the degree of dispersion to obtain the distribution uniformity of the word. The distribution uniformity is positively correlated with the distribution range and negatively correlated with the degree of dispersion.
[0058] Specifically, the distribution uniformity satisfies the following relational expression:
[0059] ;
[0060] In the formula, represents the distribution uniformity of the th word, represents the The interval between the first and last occurrences of a word, represents the total length of the interaction data, represents the frequency of occurrence of the word in all interaction data, represents the total number of position intervals between two adjacent occurrences of the word, represents the position interval ordinal number, represents the length of the th position interval of the word, represents the standard normalization function.
[0061] That is to say, represents the occurrence situation of the word in the entire interaction data process. If its value is larger, it indicates that the range from the first occurrence to the last occurrence of the word is wider, almost spanning the entire interaction data, indicating that the word continuously appears throughout the conversation process and has a high "persistence";
[0062] represents the difference situation of the position intervals of the word. The larger its value, the greater the difference in position intervals, that is, the more uneven the distribution of the word in the interaction process; conversely, the smaller its value, the more uniform the distribution of the word in the interaction process.
[0063] In addition, another embodiment further includes:
[0064] Divide the interaction data into several paragraphs, count the number of times the word appears in each paragraph, calculate the frequency of the word in each paragraph, and use the entropy calculation formula to obtain the distribution uniformity of each word.
[0065] Exemplarily, it can be divided by rounds, and each round of interaction between the user and the system is used as a paragraph. For example: when the user asks a question once and the system answers once, this round of interaction is used as a paragraph; it can also divide the interaction data into windows of a fixed time length, such as every 5 minutes or every 10 minutes as a paragraph.
[0066] Specifically, the distribution uniformity satisfies the following relational expression:
[0067] ;
[0068] In the formula, represents the distribution uniformity of the word, represents the number of paragraphs into which the interaction data is divided, represents the The frequency of occurrence in a paragraph or paragraphs denotes the logarithmic function with as the base.
[0069] Exemplarily, a piece of interactive data is divided into 4 paragraphs, and the number of occurrences of a certain word in each paragraph is as follows: Paragraph 1: 3 times; Paragraph 2: 1 time; Paragraph 3: 2 times; Paragraph 4: 4 times; The total number of occurrences is: times;
[0070] Calculate the frequency in each paragraph: 、 、 、 ;
[0071] Entropy: .
[0072] That is to say, if a word appears concentrated in a few paragraphs and rarely or not at all in other paragraphs, then the entropy value will be low, indicating uneven distribution, and vice versa indicating uniform distribution.
[0073] Partition the set of each word, including:
[0074] Partition the words with distribution uniformity less than the preset uniformity threshold into a dynamic set, and vice versa, partition the words greater than or equal to the preset uniformity threshold into a static set.
[0075] Exemplarily, the preset uniformity threshold is 0.5 and can be adjusted according to specific circumstances.
[0076] It should be noted that in the real-time interactive scenario of legal consultation, the partitioning of the dynamic set and the static set is an important means to optimize the Huffman coding strategy based on the distribution uniformity and frequency change characteristics of words; the dynamic set contains words with large frequency changes and uneven distribution during the interaction process. As the topic of the conversation switches, the frequency of word appearance may suddenly increase or decrease, and it is applicable to the dynamic coding strategy. The static set contains words with small changes and relatively uniform distribution during the interaction process, which are not affected by the switching of the conversation topic, and the frequency of word appearance is relatively stable, and it is applicable to the static coding strategy.
[0077] S3: Based on the words in the dynamic set, obtain the importance of the words according to the change in the occurrence frequency of the words during the interaction process, and adjust the Huffman tree nodes corresponding to the words in the dynamic set to optimize the coding strategy, and compress and transmit the interactive data according to the adjusted Huffman tree.
[0078] Example 1: The importance of words, including:
[0079] Taking any word in the dynamic set as the target word, record the frequency of occurrence of the target word in each interaction and the total number of interactions. Among them, one interaction occurs when the user asks in the dynamic set and the system answers once;
[0080] Calculate the frequency difference of the target word in adjacent interaction times to obtain the dynamic change situation;
[0081] Use the exponential decay function to obtain the time weight of each interaction and perform normalization processing. Sum the dynamic change situations of all interactions multiplied by the normalized time weight to obtain the weighted frequency change. Divide the weighted frequency change by the distribution uniformity and perform normalization processing to obtain the importance of the target word.
[0082] Specifically, the importance satisfies the following relational expression:
[0083] ;
[0084] In the formula, represents the importance of the th word, represents the total number of interactions, represents the ordinal number of interaction times, represents the th word at the th interaction, represents the th word at the th interaction, represents the distribution uniformity of the th word, represents the exponential function with the natural constant as the base, represents the standard normalization function.
[0085] That is to say, is the time weight of each interaction. The closer to the current interaction, that is to say, the closer to the higher the weight; is the normalization factor to ensure that the sum of all weights is 1, is used as the normalized time weight to measure the time sensitivity of each interaction. The closer to the current interaction, the greater the impact of the frequency change on the importance, represents the frequency change of the word. The word with an increasing frequency is considered more important, and the word with a decreasing frequency is considered less important. If , it means that the frequency of the word is increasing, indicating that the word is becoming more and more important in the current conversation. If , it indicates that the importance of the word in the current conversation has decreased.
[0086] Embodiment 2: The importance of the word includes:
[0087] Take any interaction as the marked interaction, take any word in the dynamic set as the target word, calculate the ratio of the occurrence frequency of the target word in the marked interaction to the total number of times of the target word in the marked interaction data to obtain the relative frequency, and take the sum of the relative frequencies of all target words as the overall importance;
[0088] Calculate the overall importance divided by the distribution uniformity and then perform normalization to obtain the importance of the target word.
[0089] Specifically, the importance satisfies the following relational expression:
[0090] ;
[0091] In the formula, represents the importance of the th word, represents the number of interaction data containing the th word, represents the th word in the th interaction, represents the total number of words of the th word in the th interaction data, represents the distribution uniformity of the th word, represents the function of standard normalization.
[0092] That is to say, reflects the relative frequency of the th word in the th interaction, represents the sum of the relative frequencies of the th word in all interaction data containing it, and this sum of relative frequencies reflects the overall importance of the th word in all relevant interaction data.
[0093] Embodiment 3: The importance of a word includes:
[0094] Take any interaction as the marked interaction, take any word in the dynamic set as the target word, calculate the ratio of the occurrence frequency of the target word in the marked interaction to the total number of times of the target word in the marked interaction data to obtain the relative frequency, and take the sum of the relative frequencies of all target words as the overall importance;
[0095] Calculate the frequency proportion of the target word in all interactions, divide the product of the frequency proportion and the overall importance by the distribution uniformity, and then perform normalization to obtain the importance of the target word.
[0096] Specifically, the importance satisfies the following relationship:
[0097] ;
[0098] In the formula, represents the importance of the th word, represents the number of interaction data containing the th word, represents the total number of interactions, represents the frequency of occurrence of the th word in the rd interaction, represents the total number of words in the th word in the th interaction data, represents the distribution uniformity of the th word, represents the function of standard normalization.
[0099] That is to say, represents the proportion of the th word appearing in all interaction data, that is, the frequency proportion of the th word. If is close to 1, it means that the th word appears in most interaction data and is a common word. On the contrary, if it is close to 0, it means that it is a rare word, reflecting the universality of the word.
[0100] measures the sum of the relative frequencies of the th word in all interaction data containing it, reflecting the overall importance of the word in the interaction data.
[0101] In the first embodiment, the importance of words is dynamically evaluated through time weight and frequency change. Using the real-time interaction scenario, the computational complexity is relatively high, considering the frequency change and time weight of words in different interactions; in the second embodiment, the importance of words is comprehensively evaluated through word frequency normalization and average frequency, which is applicable to the static evaluation scenario, and the computational complexity is moderate, comprehensively considering the average frequency of words in the current interaction and all relevant interactions; in the third embodiment, the importance of words is simplified through word frequency normalization and frequency proportion, which is applicable to the fast evaluation and large-scale data processing scenario, and the computational complexity is moderate, and the universality of words is quickly evaluated through the frequency proportion. Implementers can make a choice according to specific circumstances.
[0102] Adjust the Huffman tree nodes corresponding to the words in the dynamic set, including:
[0103] Preset an adjustment period, traverse the Huffman tree, and mark the nodes corresponding to the words in all dynamic sets. Among them, the nodes in the Huffman tree corresponding to the words in the static set remain unchanged;
[0104] Sort all the words in the dynamic set in descending order of importance, and fill the sorted words into the marked nodes in the Huffman tree from top to bottom in sequence, and obtain the codes corresponding to each word after adjustment.
[0105] Exemplarily, the preset adjustment period is to adjust once every two interactions, and the implementer can adjust according to the actual situation; ensure that the Huffman tree can reflect the word frequency changes in a timely manner, and at the same time avoid excessive system overhead caused by overly frequent adjustments. The word frequencies in the static set change less and do not need to be adjusted frequently. Keeping the positions of their nodes unchanged can reduce the computational complexity. Words with high importance should be placed at the upper layer of the Huffman tree to obtain a shorter coding length, thereby improving the compression efficiency. Filling the nodes in the order from top to bottom ensures that words with high importance preferentially occupy better coding positions, which can effectively improve the real-time performance and compression efficiency of the system.
[0106] Compress and transmit the interaction data according to the adjusted Huffman tree, including:
[0107] During the compression process of the interaction data, monitor the appearance of new words, judge whether the new words already exist in the Huffman tree. In response to the new words not existing in the Huffman tree, insert the new words into the last node of the Huffman tree, otherwise, do not insert. After inserting the new words, complete the data compression. Refer to Figure 2 , where A represents the original Huffman tree, B represents the Huffman tree after adding new words, a represents the last node, and b represents the added new word.
[0108] Among them, new words refer to words that do not exist in the current Huffman tree. These words may appear for the first time due to the change of the conversation topic or new concepts introduced by the user.
[0109] During the compression process of the interaction data, the system monitors the appearance of new words in real time. New words refer to words that do not exist in the current Huffman tree. These words may appear for the first time due to the change of the conversation topic or new concepts introduced by the user. After each interaction, the system dynamically updates the Huffman tree to ensure that new words can be processed in a timely manner.
[0110] When a new word is detected, the system determines whether the word already exists in the Huffman tree. If the new word does not exist in the Huffman tree, it is inserted into the last node of the Huffman tree, that is, the node with the lowest frequency or the lowest importance. The last node is chosen as the insertion position because the frequency of the new word is usually low, which is suitable for a longer coding length. After inserting the new word, the system updates the structure of the Huffman tree and recalculates the codes of all nodes to ensure that the new word can obtain an appropriate coding length.
[0111] To reduce the computational overhead, the system selects the last node as the insertion position for the new word, reducing the impact of the insertion operation on the existing tree structure. At the same time, the system can dynamically adjust the position of the new word in the Huffman tree according to the frequency change of the new word in subsequent interactions to optimize the coding strategy. If the new word appears frequently in subsequent interactions, the system can promote it to a better position to obtain a shorter coding length; if the new word appears rarely in subsequent interactions, its position at the last node can be maintained to reduce the impact on the overall coding strategy.
[0112] Through this mechanism, the system can dynamically adapt to the changes in the conversation content, ensure that new words can be effectively processed, and at the same time maintain high - efficiency coding and transmission performance.
[0113] The present invention also provides a human - machine interaction system for legal consultation. As Figure 3 shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a human - machine interaction method for legal consultation according to the first aspect of the present invention. The system also includes a communication bus and a communication interface and other components well - known to those skilled in the art. Their settings and functions are known in the art, so they will not be elaborated here.
[0114] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.
Claims
1. A human-computer interaction method for legal consultation, characterized in that, including: Obtain interaction data and count word frequency information; Calculate the distribution uniformity of words according to the index positions of words in the interaction data, including: Starting from the first occurrence of any word in the interaction data, calculate the index position intervals between adjacent occurrences in sequence, calculate the standard deviation of all index position intervals, and obtain the dispersion degree of the index position intervals of the word; Take the ratio between the position interval from the first occurrence to the last occurrence of the word and the total length of the interaction data as the distribution range of the word; Normalize the ratio between the distribution range and the dispersion degree to obtain the distribution uniformity of the word. The distribution uniformity is positively correlated with the distribution range and negatively correlated with the dispersion degree; Based on the distribution uniformity and a preset uniformity threshold, partition each word into sets. Words with large frequency changes, uneven distributions, and sudden increases or decreases in occurrence frequency with the switching of conversation topics during the interaction process are partitioned into a dynamic set, and words with small changes, relatively uniform distributions, and not affected by the switching of conversation topics during the interaction process are partitioned into a static set; Based on the words in the dynamic set, obtain the importance of words according to the change situation of the occurrence frequency of words during the interaction process, including: Take any word in the dynamic set as the target word, record the occurrence frequency of the target word in each interaction and the total number of interactions, where one interaction in the dynamic set is defined as one interaction when the user asks and the system answers once; Calculate the frequency difference of the target word in adjacent interaction times to obtain the dynamic change situation; Use an exponential decay function to obtain the time weight of each interaction and perform normalization processing, sum the dynamic change situations of all interactions multiplied by the normalized time weight, obtain the weighted frequency change, divide the weighted frequency change by the distribution uniformity and perform normalization processing to obtain the importance of the target word; Adjust the Huffman tree nodes corresponding to the words in the dynamic set, optimize the encoding strategy, and compress and transmit the interaction data according to the adjusted Huffman tree; The adjustment of the Huffman tree nodes corresponding to the words in the dynamic set includes: Preset an adjustment period, traverse the Huffman tree, and mark all the nodes corresponding to the words in the dynamic set, where the nodes corresponding to the words in the static set in the Huffman tree remain unchanged; Sort all the words in the dynamic set from largest to smallest according to importance, and fill the sorted words into the marked nodes in sequence from top to bottom in the Huffman tree, and obtain the codes corresponding to each word after adjustment; The compression and transmission of the interaction data according to the adjusted Huffman tree includes: During the compression process of the interaction data, monitor the appearance of new words, judge whether the new words already exist in the Huffman tree. In response to the new words not existing in the Huffman tree, insert the new words into the last node of the Huffman tree, otherwise, do not insert. After inserting the new words, complete the data compression.
2. The human-computer interaction method for legal consultation according to claim 1, wherein, The partitioning of each word into sets includes: Partition the words with distribution uniformity less than the preset uniformity threshold into a dynamic set, and conversely, partition the words greater than or equal to the preset uniformity threshold into a static set.
3. A human-computer interaction system for legal consultation, characterized in that, including: A processor and a memory, the memory storing computer program instructions which, when executed by the processor, implement the human-computer interaction method for legal consultation according to claim 1 or 2.
Citation Information
Patent Citations
Huffman coding method, system and device and medium
CN116505954A
Vehicle sales information optimization storage method based on data analysis
CN117278055A