Man-machine interaction method and system for legal consultation

By dynamically adjusting the Hoffman tree structure in the legal consulting human-computer interaction system, dividing word sets according to the uniformity and importance of words, and optimizing the coding strategy, the problem of traditional Hoffman coding being too long when the new topic was introduced and the static code table was unable to perceive word frequency changes in real time, significantly improving data transmission efficiency and system performance.

CN119938873AActive Publication Date: 2025-05-06GUANGDONG BOWEI CHUANGYUAN TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510436341.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the human-computer interaction scenario of legal consultation, traditional Hoffmann encoding causes the encoding length to be too long when new topics are introduced, and the static code table cannot perceive the dynamic changes in the word frequency of the conversation content in real time, resulting in a significant decrease in compression efficiency.

Method used

By obtaining interactive data and counting word frequency information, the distribution uniformity of words is calculated, and according to the distribution uniformity and preset uniformity threshold, the words are divided into dynamic sets and static sets. Words in the dynamic set acquire importance according to their frequency changes in the interaction process, and adjust the Hoffman tree nodes to optimize the coding strategy to adapt to the frequency changes of real-time dialogue word.

Benefits of technology

It significantly improves the efficiency of data transmission and the real-time nature of the system, reduces the computational complexity, improves the robustness and resource utilization efficiency of the system, enhances the system's processing ability of new vocabulary, and improves user experience and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938873A_ABST
    Figure CN119938873A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data transmission, in particular to a man-machine interaction method and system for legal consultation, and the method comprises the steps: obtaining interaction data and carrying out the statistics of word frequency information, calculating the distribution uniformity of words according to the index positions of the words in the interaction data, and carrying out the division and collection of all words based on the distribution uniformity and a preset uniform threshold value, based on the words in the dynamic set, the importance of the words is obtained according to the occurrence frequency change condition of the words in the interaction process, according to the importance of the words, Hoffman tree nodes corresponding to the words in the dynamic set are adjusted, the coding strategy is optimized, and interaction data are compressed and transmitted according to the adjusted Hoffman tree. According to the method, the change of word frequency is monitored in real time, the new word is inserted into the last node of the Huffman tree, and the position of the new word is dynamically adjusted to be closer to the root node when the new word frequently appears, so that the coding strategy is optimized, and the compression efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data transmission, and in particular to a human-computer interaction method and system for legal consultation. Background Art

[0002] Legal consultation is a professional legal service that aims to help individuals or organizations solve legal problems, provide legal opinions and guide actions. Legal consultation is carried out face-to-face, by phone, online or in writing. Its role is to provide legal opinions, assess risks, guide actions and prevent disputes.

[0003] With the development of technology, human-computer interaction has been applied in various fields, especially in chatbots and intelligent question-answering systems. In this context, legal consultation will be combined with human-computer interaction, and with the help of artificial intelligence technology and natural language processing capabilities, efficient and convenient legal services can be provided to the public and legal professionals. It not only improves the efficiency and accessibility of legal consultation, but also promotes the digital transformation of the legal services industry. However, this process also faces many challenges. How to efficiently manage and process massive amounts of human-computer interaction data has become a technical problem that needs to be solved urgently, especially how to optimize data transmission and processing efficiency while ensuring service quality, which has become the focus of industry attention.

[0004] In the real-time interactive scenario of legal consultation, the topic of user conversation will change dynamically with the progress of consultation. As the scale of legal consultation business continues to expand, traditional Huffman coding has gradually exposed two significant defects in interactive data compression: First, when new topics are introduced, some professional terms have not been pre-counted, and the encoding length of these terms calculated based on global word frequency is long, resulting in a significant decrease in compression efficiency. Second, the static code table cannot perceive the dynamic changes in the frequency of conversation content in real time, and the compression efficiency continues to decrease over time. Summary of the invention

[0005] In order to solve the problem that when a new topic is introduced during human-computer interaction, the encoding length is too long and the static code table cannot perceive the dynamic changes of the word frequency of the conversation content in real time, resulting in a significant decrease in compression efficiency, the present invention provides solutions in the following aspects.

[0006] In a first aspect, a human-computer interaction method for legal consultation includes: obtaining interaction data and counting word frequency information; calculating the distribution uniformity of the word according to the index position of the word in the interaction data, and dividing each word into a set based on the distribution uniformity and a preset uniform threshold, wherein the set includes: a dynamic set and a static set; based on the words in the dynamic set, according to the change in the frequency of occurrence of the word during the interaction process, obtaining the importance of the word, and adjusting the Huffman tree nodes corresponding to the words in the dynamic set, optimizing the encoding strategy, and compressing and transmitting the interaction data according to the adjusted Huffman tree.

[0007] By obtaining the interaction data and counting the word frequency information, the distribution uniformity of the word is calculated according to the index position of the word in the interaction data. Based on the distribution uniformity and the preset uniform threshold, the words are divided into dynamic sets and static sets. The words in the dynamic set obtain their importance according to the changes in their frequency of occurrence during the interaction process, and the Huffman tree nodes are adjusted accordingly to optimize the encoding strategy. Finally, the interaction data is compressed and transmitted according to the adjusted Huffman tree. The efficiency of data transmission and the real-time performance of the system are significantly improved. At the same time, by reasonably dividing the dynamic set and the static set, the calculation complexity is reduced, and the robustness and resource utilization efficiency of the system are improved. By monitoring the appearance of new words and reasonably inserting them into the Huffman tree, the system's ability to process new vocabulary is ensured, and the adaptability and stability of the system are further improved. These improvements not only improve the performance of the system, but also significantly improve the user experience.

[0008] Preferably, the distribution uniformity includes: For any word that appears in the interaction data, the index position intervals between two adjacent occurrences are calculated in sequence, and the standard deviation of all index position intervals is calculated to obtain the discrete degree of the index position interval of the word; The ratio between the position interval of the first and last occurrence of a word and the total length of the interaction data is taken as the distribution range of the word; The ratio between the distribution range and the discrete degree is normalized to obtain the distribution uniformity of the word. The distribution uniformity is positively correlated with the distribution range and negatively correlated with the discrete degree.

[0009] By accurately calculating the discrete degree and distribution range of the index position interval of the word, and normalizing the ratio of the two to obtain the word distribution uniformity, the distribution characteristics of the word in the interactive data can be more accurately measured. By calculating the distribution uniformity, the words can be more reasonably divided into dynamic sets and static sets, and the adjustment strategy of the Huffman tree can be optimized. This not only improves data compression efficiency and reduces computational complexity, but also enhances the system's adaptability to changes in conversation topics, ultimately improving user experience and system performance.

[0010] Preferably, the distribution uniformity includes: The interaction data is divided into several paragraphs, the number of times a word appears in each paragraph is counted, the frequency of the word in each paragraph is calculated, and the entropy calculation formula is used to obtain the distribution uniformity of each word.

[0011] By dividing the interaction data into several paragraphs, counting the number of occurrences of each word in each paragraph and calculating its frequency, and then using the entropy calculation formula to measure the distribution uniformity of each word, the distribution of words in different paragraphs can be captured more carefully. By quantifying the distribution uniformity, the system can more accurately divide words into dynamic sets and static sets, thereby optimizing the adjustment strategy of the Huffman tree.

[0012] Preferably, the dividing of the words into sets includes: Words whose distribution uniformity is less than a preset uniformity threshold are divided into a dynamic set, whereas words whose distribution uniformity is greater than or equal to the preset uniformity threshold are divided into a static set.

[0013] Preferably, the importance of the word includes: Take any word in the dynamic set as the target word, record the frequency of the target word in each interaction and the total number of interactions, where one user asks and the system answers once in the dynamic set, which is considered an interaction; Calculate the frequency difference of the target word in adjacent interactions to obtain the dynamic change situation; An exponential decay function is used to obtain the time weight of each interaction and normalize it. The dynamic changes of all interactions are multiplied by the normalized time weight and summed to obtain the weighted frequency change. The weighted frequency change is divided by the distribution uniformity and normalized to obtain the importance of the target word.

[0014] By accurately measuring the importance of words in the interaction process, with particular attention paid to the dynamic changes in word frequency and time sensitivity, this provides a basis for dynamically adjusting the Huffman tree, optimizing the encoding strategy, improving data compression efficiency and the real-time performance of the system, and enhancing the system's ability to adapt to changes in conversation topics, improving user experience and system performance. It is particularly suitable for scenarios such as legal consulting that require efficient processing of dynamic interactive data.

[0015] Preferably, the importance of the word includes: Take any interaction as the marked interaction, any word in the dynamic set as the target word, calculate the ratio of the frequency of the target word in the marked interaction to the total number of the target word in the marked interaction data, get the relative frequency, and take the sum of the relative frequencies of all words containing the target word as the overall importance; The overall importance is calculated and divided by the distribution uniformity and then normalized to obtain the importance of the target word.

[0016] By considering the frequency of occurrence of the target word in each interaction and its distribution characteristics, we can more comprehensively evaluate the importance of the word, provide a more accurate basis for dynamically adjusting the Huffman tree, optimize the encoding strategy, improve data compression efficiency and the real-time performance of the system, and enhance the system's ability to adapt to changes in conversation topics.

[0017] Preferably, the importance of the word includes: Take any interaction as the marked interaction, any word in the dynamic set as the target word, calculate the ratio of the frequency of the target word in the marked interaction to the total number of the target word in the marked interaction data, get the relative frequency, and take the sum of the relative frequencies of all words containing the target word as the overall importance; Calculate the frequency share of the target word in all interactions, divide the product of the frequency share and the overall importance by the distribution uniformity, and normalize it to get the importance of the target word.

[0018] Preferably, the step of adjusting the Huffman tree nodes corresponding to the words in the dynamic set includes: Preset an adjustment cycle, traverse the Huffman tree, and mark the nodes corresponding to the words in all dynamic sets, wherein the nodes in the Huffman tree corresponding to the words in the static set remain unchanged; Sort all the words in the dynamic set according to their importance from large to small, and fill the sorted words in the marked nodes from top to bottom in the Huffman tree, and obtain the corresponding codes of each word after adjustment.

[0019] Preferably, compressing and transmitting the interactive data according to the adjusted Huffman tree includes: During the interactive data compression process, the emergence of new words is monitored to determine whether the new words already exist in the Huffman tree. In response to the new word not existing in the Huffman tree, the new word is inserted into the last node of the Huffman tree. Otherwise, it does not need to be inserted. After the new word is inserted, the data compression is completed.

[0020] In a second aspect, a human-computer interaction system for legal consultation includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned human-computer interaction method for legal consultation is implemented.

[0021] The present invention has the following effects: 1. The present invention dynamically adjusts the structure of the Huffman tree so that frequently occurring words can be compressed into shorter codes, thereby reducing the bandwidth occupied by data transmission and improving the data transmission rate. It is applicable to the dynamic changes in the conversation topics in the legal consultation scenario, which leads to the continuous changes in the word frequency distribution. By dividing the dynamic set and the static set, the system can more efficiently process the newly emerging professional terms and common words, significantly improving the compression efficiency.

[0022] 2. By adjusting the vocabulary in the dynamic set, the system can dynamically generate a Huffman tree that is more suitable for the current dialogue state according to the word frequency changes in the real-time dialogue, insert new words into the last node of the Huffman tree, and dynamically adjust its position to be closer to the root node when the new words appear frequently, thereby optimizing the encoding strategy and improving the compression efficiency. This is conducive to the system being able to quickly adapt to changes in the dialogue topic and maintain efficient encoding and transmission performance, and is particularly suitable for high-concurrency and real-time interactive legal consulting scenarios.

[0023] 3. The present invention divides the dynamic set and the static set according to the distribution uniformity of the words, and only performs frequency change analysis and Huffman tree adjustment on the dynamic set, which greatly reduces the complexity of calculation. Since the words in the static set have a high distribution uniformity, they do not need to be adjusted frequently, thereby saving computing resources, which is beneficial to significantly reduce the computing burden of the system while ensuring compression efficiency, and improve the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a method flow chart of steps S1 to S3 in a human-computer interaction method for legal consultation in an embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram of a Huffman tree after new words are added to a human-computer interaction method for legal consultation according to an embodiment of the present invention.

[0026] Figure 3 The present invention is a structural block diagram of a human-computer interaction system for legal consultation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments.

[0028] Reference Figure 1 , a human-computer interaction method for legal consultation includes steps S1 to S3, which are as follows: S1: Obtain interaction data and count word frequency information.

[0029] The interactive system records user and system outputs to the database through the API to store interactive data. Before word segmentation, the interactive data is preprocessed, such as removing stop words, punctuation marks, and special characters to improve the accuracy and efficiency of word segmentation. Jieba word segmentation is used to segment the stored interactive data and obtain the word segmentation results. Among them, jieba word segmentation is widely used in the field of natural language processing. Its core function is to divide a continuous Chinese interactive data into several words. Jieba word segmentation can accurately segment words through the maximum probability word segmentation method based on the prefix dictionary. For the word segmentation results, the frequency of occurrence of each word segmentation and the index position of each analysis are counted.

[0030] It should be noted that in legal consulting services, human-computer interaction data refers specifically to all interactive information between users and the legal consulting system, including but not limited to the following three parts: User input information: all legal consulting questions, statements, requirements and other specific interactive data provided by the user when communicating with the system through interactive data; System response content: The system outputs legal advice, answers, relevant laws and regulations, case analysis and other information after natural language processing based on user input; Conversation history information: contextual information that the system needs to retain during an ongoing conversation, including but not limited to previous consultation questions, answer records, etc., to ensure that subsequent communications can remain coherent and targeted.

[0031] It should be noted that in the real-time interactive system for legal consultation, the topic of the conversation between the user and the system often changes dynamically. As the conversation progresses, the specific legal issues or needs of the user's consultation will continue to change, causing the topic of the conversation to be adjusted accordingly. The traditional Huffman coding algorithm urgently estimates the static frequency of global word frequencies and uses pre-calculated word frequency distributions to construct coding tables. This static frequency estimation method has obvious limitations. For example, as the topic of the conversation changes, the frequency of words related to some old topics may drop sharply, while the frequency of words related to new topics may rise rapidly. Since Huffman coding uses a fixed coding table throughout the encoding process, this frequent field change will lead to a significant decrease in encoding efficiency, which will ultimately affect the compression performance of the entire real-time interactive system. The specific adjustment steps are as follows: S2: Calculate the distribution uniformity of the word according to the index position of the word in the interaction data, and divide each word into a set based on the distribution uniformity and a preset uniformity threshold, where the set includes: a dynamic set and a static set.

[0032] It should be noted that the traditional static Huffman coding table cannot perceive the dynamic changes of word frequency. If all words are counted and adjusted, a large amount of calculation is required. Therefore, the words are divided into dynamic sets and static sets. The words in the static set are common words or fixed words. The frequency of their appearance in the conversation changes little, so they do not need to be adjusted frequently. This strategy reduces the complexity of the Huffman tree structure adjustment.

[0033] Distribution uniformity includes: For any word that appears in the interaction data, the index position intervals between two adjacent occurrences are calculated in sequence, and the standard deviation of all index position intervals is calculated to obtain the discrete degree of the index position interval of the word; The ratio between the position interval of the first and last occurrence of a word and the total length of the interaction data is taken as the distribution range of the word; The ratio between the distribution range and the discrete degree is normalized to obtain the distribution uniformity of the word. The distribution uniformity is positively correlated with the distribution range and negatively correlated with the discrete degree.

[0034] Specifically, the distribution uniformity satisfies the following relationship: ; In the formula, Indicates The distribution uniformity of words, Indicates The interval between the first and last occurrence of a word, Indicates the total length of the interaction data. Indicates The frequency of occurrence of a word in all interaction data, Indicates The total number of intervals between two consecutive occurrences of a word, Indicates the position interval ordinal number, Indicates The word The length of the position interval, Indicates The average length of the position interval of each word, Represents the standard normalization function.

[0035] That is to say, Indicates the occurrence of the word in the entire interactive data process. If the value is larger, it means that the word has a wider range from the first appearance to the last appearance, almost throughout the entire interactive data, indicating that the word continues to appear in the entire conversation process and has a higher "persistence"; It indicates the difference in the position interval of the word. The larger the value, the greater the difference in the position interval, that is, the more uneven the distribution of the word in the interaction process; conversely, the smaller the value, the more even the distribution of the word in the interaction process.

[0036] In addition, another embodiment also includes: The interaction data is divided into several paragraphs, the number of times a word appears in each paragraph is counted, the frequency of the word in each paragraph is calculated, and the entropy calculation formula is used to obtain the distribution uniformity of each word.

[0037] Exemplarily, it can be divided by rounds, with each round of interaction between the user and the system as a paragraph, for example: the user asks a question once, the system answers once, and this round of interaction is a paragraph; the interaction data can also be divided into windows of fixed time length, for example, every 5 minutes or every 10 minutes as a paragraph.

[0038] Specifically, the distribution uniformity satisfies the following relationship: ; In the formula, Indicates The distribution uniformity of words, Indicates the number of paragraphs into which the interaction data is divided. Indicates The word in The frequency of occurrence in a paragraph or paragraphs, Indicates The logarithmic function with base .

[0039] For example, a segment of interaction data is divided into 4 paragraphs, and the number of occurrences of a certain word in each paragraph is as follows: paragraph 1: 3 times; paragraph 2: 1 time; paragraph 3: 2 times; paragraph 4: 4 times; the total number of occurrences is: Second-rate; Count the frequency in each paragraph: 、 、 、 ; entropy: .

[0040] That is, if a word appears in a few paragraphs and rarely or not in other paragraphs, the entropy value will be low, indicating uneven distribution, and vice versa.

[0041] Each word is divided into sets, including: Words whose distribution uniformity is less than a preset uniformity threshold are divided into a dynamic set, whereas words whose distribution uniformity is greater than or equal to the preset uniformity threshold are divided into a static set.

[0042] Exemplarily, the preset uniformity threshold is 0.5, which can be adjusted according to specific circumstances.

[0043] It should be noted that in the real-time interactive scenario of legal consultation, the division of dynamic sets and static sets is an important means to optimize the Huffman coding strategy based on the distribution uniformity and frequency change characteristics of words; the dynamic set contains words with large frequency changes and uneven distribution during the interaction process. As the topic of the conversation switches, the frequency of word appearance may suddenly increase or decrease, which is suitable for the dynamic coding strategy. The static set contains words with small changes and more uniform distribution during the interaction process. It is not affected by the switching of conversation topics, and the frequency of word appearance is relatively stable, which is suitable for the static coding strategy.

[0044] S3: Based on the words in the dynamic set, the importance of the words is obtained according to the changes in the frequency of occurrence of the words during the interaction process, and the Huffman tree nodes corresponding to the words in the dynamic set are adjusted to optimize the encoding strategy. The interaction data is compressed and transmitted according to the adjusted Huffman tree.

[0045] Example 1: The importance of words, including: Take any word in the dynamic set as the target word, record the frequency of the target word in each interaction and the total number of interactions, where one user asks and the system answers once in the dynamic set, which is considered an interaction; Calculate the frequency difference of the target word in adjacent interactions to obtain the dynamic change situation; An exponential decay function is used to obtain the time weight of each interaction and normalize it. The dynamic changes of all interactions are multiplied by the normalized time weight and summed to obtain the weighted frequency change. The weighted frequency change is divided by the distribution uniformity and normalized to obtain the importance of the target word.

[0046] Specifically, the importance satisfies the following relationship: ; In the formula, Indicates The importance of the word Represents the total number of interactions, represents the ordinal number of interactions, Indicates The word in The frequency of occurrence in interactions, Indicates The word in The frequency of occurrence in interactions, Indicates The distribution uniformity of words, Expressed as a natural constant is the exponential function of the base, Represents the standard normalization function.

[0047] That is to say, is the time weight of each interaction, the closer to the current interaction, that is, near The higher the weight; is a normalization factor that ensures that the sum of all weights is 1, As the normalized time weight, it is used to measure the time sensitivity of each interaction. The frequency change closer to the current interaction has a greater impact on importance. Indicates the frequency change of words. Words with increasing frequency are considered more important, and words with decreasing frequency are considered less important. , indicating that the frequency of the word is increasing, indicating that the word is becoming more and more important in the current conversation. , indicating that the importance of the word in the current conversation decreases.

[0048] Example 2: The importance of words includes: Take any interaction as the marked interaction, any word in the dynamic set as the target word, calculate the ratio of the frequency of the target word in the marked interaction to the total number of the target word in the marked interaction data, get the relative frequency, and take the sum of the relative frequencies of all words containing the target word as the overall importance; The overall importance is calculated and divided by the distribution uniformity and then normalized to obtain the importance of the target word.

[0049] Specifically, the importance satisfies the following relationship: ; In the formula, Indicates The importance of the word Indicates that it contains The number of interaction data for each word, Indicates The word in The frequency of occurrence in interactions, Indicates The word in The total number of words in the interaction data, Indicates The distribution uniformity of words, Represents the standard normalization function.

[0050] That is to say, Reflects the The word in The relative frequency of interactions, Indicates The sum of the relative frequencies of a word in all interaction data containing it. This relative frequency sum reflects the The overall importance of a word in all relevant interaction data.

[0051] Example 3: The importance of words includes: Take any interaction as the marked interaction, any word in the dynamic set as the target word, calculate the ratio of the frequency of the target word in the marked interaction to the total number of the target word in the marked interaction data, get the relative frequency, and take the sum of the relative frequencies of all words containing the target word as the overall importance; Calculate the frequency share of the target word in all interactions, divide the product of the frequency share and the overall importance by the distribution uniformity, and normalize it to get the importance of the target word.

[0052] Specifically, the importance satisfies the following relationship: ; In the formula, Indicates The importance of the word Indicates that it contains The number of interaction data for each word, Represents the total number of interactions, Indicates The word in The frequency of occurrence in interactions, Indicates The word in The total number of words in the interaction data, Indicates The distribution uniformity of words, Represents the standard normalization function.

[0053] That is to say, Indicates The proportion of words appearing in all interaction data is the The frequency ratio of words, if Close to 1, indicating If a word appears in most interaction data, it is a common word. Otherwise, if it is close to 0, it is a rare word, which reflects the universality of the word.

[0054] Measure the The sum of the relative frequencies of a word in all interaction data containing it reflects the overall importance of the word in the interaction data.

[0055] In the first embodiment, the importance of words is dynamically evaluated through time weight and frequency change, using real-time interaction scenarios, with relatively high computational complexity, taking into account the frequency change and time weight of words in different interactions; in the second embodiment, the importance of words is comprehensively evaluated through word frequency normalization and average frequency, which is suitable for static evaluation scenarios, with moderate computational complexity, and comprehensively considers the average frequency of words in the current interaction and all related interactions; in the third embodiment, the importance of words is simplified through word frequency normalization and frequency proportion, which is suitable for rapid evaluation and large-scale data processing scenarios, with moderate computational complexity, and the prevalence of words is quickly evaluated through frequency proportion. Implementers can make choices based on specific circumstances.

[0056] Adjust the Huffman tree nodes corresponding to the words in the dynamic set, including: Preset an adjustment cycle, traverse the Huffman tree, mark the nodes corresponding to the words in all dynamic sets, wherein the nodes in the Huffman tree corresponding to the words in the static set remain unchanged; Sort all the words in the dynamic set according to their importance from large to small, and fill the sorted words in the marked nodes from top to bottom in the Huffman tree, and obtain the corresponding codes of each word after adjustment.

[0057] Exemplarily, the preset adjustment cycle is once every two interactions, and the implementer can make adjustments according to the actual situation; ensure that the Huffman tree can reflect the changes in word frequency in a timely manner, while avoiding too frequent adjustments that lead to excessive system overhead. The word frequency in the static set changes less and does not need to be adjusted frequently. Keeping its node position unchanged can reduce the computational complexity. Words with high importance should be placed in the upper layer of the Huffman tree to obtain a shorter coding length, thereby improving compression efficiency. Filling the nodes in order from top to bottom ensures that words with high importance take priority in better coding positions, which can effectively improve the real-time performance and compression efficiency of the system.

[0058] The interactive data is compressed and transmitted according to the adjusted Huffman tree, including: During the interactive data compression process, the appearance of new words is monitored to determine whether the new words already exist in the Huffman tree. If the new word does not exist in the Huffman tree, the new word is inserted into the last node of the Huffman tree. Otherwise, it does not need to be inserted. After the new word is inserted, the data compression is completed. Figure 2 , where A represents the original Huffman tree, B represents the Huffman tree after adding new words, a represents the last node, and b represents the added new word.

[0059] Among them, new words refer to words that do not exist in the current Huffman tree. These words may appear for the first time due to the switching of conversation topics or new concepts introduced by the user.

[0060] During the interaction data compression process, the system monitors the appearance of new words in real time. New words are words that do not exist in the current Huffman tree. These words may appear for the first time due to the switching of conversation topics or new concepts introduced by users. After each interaction, the system dynamically updates the Huffman tree to ensure that new words can be processed in a timely manner.

[0061] When a new word is detected, the system determines whether the word already exists in the Huffman tree. If the new word does not exist in the Huffman tree, it is inserted into the last node of the Huffman tree, which is the node with the lowest frequency or the lowest importance. The last node is chosen as the insertion position because the frequency of new words is usually low and suitable for longer encoding lengths. After inserting the new word, the system updates the structure of the Huffman tree and recalculates the encoding of all nodes to ensure that the new word can obtain a suitable encoding length.

[0062] In order to reduce computational overhead, the system selects the last node as the insertion position of the new word, reducing the impact of the insertion operation on the existing tree structure. At the same time, the system can dynamically adjust the position of the new word in the Huffman tree according to the frequency changes of the new word in subsequent interactions to optimize the encoding strategy. If the new word appears frequently in subsequent interactions, the system can promote it to a better position to obtain a shorter encoding length; if the new word rarely appears in subsequent interactions, it can keep its position at the last node to reduce the impact on the overall encoding strategy.

[0063] Through this mechanism, the system can dynamically adapt to changes in the content of the conversation, ensuring that new words can be effectively processed while maintaining efficient encoding and transmission performance.

[0064] The present invention also provides a human-computer interaction system for legal consultation. Figure 3 As shown, the system includes a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a human-computer interaction method for legal consultation according to the first aspect of the present invention is implemented. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, whose settings and functions are known in the art, and therefore will not be described in detail here.

[0065] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A human-computer interaction method for legal consultation, characterized in that: include: Obtain interaction data and count word frequency information; According to the index position of the word in the interactive data, the distribution uniformity of the word is calculated, and based on the distribution uniformity and a preset uniformity threshold, each word is divided into a set, wherein the set includes: a dynamic set and a static set; Based on the words in the dynamic set, the importance of the words is obtained according to the changes in the frequency of occurrence of the words in the interaction process, and the Huffman tree nodes corresponding to the words in the dynamic set are adjusted to optimize the encoding strategy. The interaction data is compressed and transmitted according to the adjusted Huffman tree.

2. A human-computer interaction method for legal consultation according to claim 1, characterized in that: The distribution uniformity includes: For any word that appears in the interaction data, the index position intervals between two adjacent occurrences are calculated in sequence, and the standard deviation of all index position intervals is calculated to obtain the discrete degree of the index position interval of the word; The ratio between the position interval of the first and last occurrence of a word and the total length of the interaction data is taken as the distribution range of the word; The ratio between the distribution range and the discrete degree is normalized to obtain the distribution uniformity of the word. The distribution uniformity is positively correlated with the distribution range and negatively correlated with the discrete degree.

3. A human-computer interaction method for legal consultation according to claim 1, characterized in that: The distribution uniformity includes: The interaction data is divided into several paragraphs, the number of times a word appears in each paragraph is counted, the frequency of the word in each paragraph is calculated, and the entropy calculation formula is used to obtain the distribution uniformity of each word.

4. A human-computer interaction method for legal consultation according to claim 1, characterized in that: The step of dividing each word into sets includes: Words whose distribution uniformity is less than a preset uniformity threshold are divided into a dynamic set, whereas words whose distribution uniformity is greater than or equal to the preset uniformity threshold are divided into a static set.

5. The human-computer interaction method for legal consultation according to claim 1, characterized in that: The importance of the words, including: Take any word in the dynamic set as the target word, record the frequency of the target word in each interaction and the total number of interactions, where one user asks and the system answers once in the dynamic set, which is considered an interaction; Calculate the frequency difference of the target word in adjacent interactions to obtain the dynamic change situation; An exponential decay function is used to obtain the time weight of each interaction and normalize it. The dynamic changes of all interactions are multiplied by the normalized time weight and summed to obtain the weighted frequency change. The weighted frequency change is divided by the distribution uniformity and normalized to obtain the importance of the target word.

6. A human-computer interaction method for legal consultation according to claim 1, characterized in that: The importance of the words mentioned include: Take any interaction as the marked interaction, any word in the dynamic set as the target word, calculate the ratio of the frequency of the target word in the marked interaction to the total number of the target word in the marked interaction data, get the relative frequency, and take the sum of the relative frequencies of all words containing the target word as the overall importance; The overall importance is calculated and divided by the distribution uniformity and then normalized to obtain the importance of the target word.

7. A human-computer interaction method for legal consultation according to claim 1, characterized in that: The importance of the words mentioned include: Take any interaction as the marked interaction, any word in the dynamic set as the target word, calculate the ratio of the frequency of the target word in the marked interaction to the total number of the target word in the marked interaction data, get the relative frequency, and take the sum of the relative frequencies of all words containing the target word as the overall importance; Calculate the frequency share of the target word in all interactions, divide the product of the frequency share and the overall importance by the distribution uniformity, and normalize it to get the importance of the target word.

8. The human-computer interaction method for legal consultation according to claim 1, characterized in that: The step of adjusting the Huffman tree nodes corresponding to the words in the dynamic set includes: Preset an adjustment cycle, traverse the Huffman tree, and mark the nodes corresponding to the words in all dynamic sets, wherein the nodes in the Huffman tree corresponding to the words in the static set remain unchanged; Sort all the words in the dynamic set according to their importance from large to small, and fill the sorted words in the marked nodes from top to bottom in the Huffman tree, and obtain the corresponding codes of each word after adjustment.

9. The human-computer interaction method for legal consultation according to claim 1, characterized in that: The step of compressing and transmitting the interactive data according to the adjusted Huffman tree includes: During the interactive data compression process, the emergence of new words is monitored to determine whether the new words already exist in the Huffman tree. In response to the new word not existing in the Huffman tree, the new word is inserted into the last node of the Huffman tree. Otherwise, it does not need to be inserted. After the new word is inserted, the data compression is completed.

10. A human-computer interaction system for legal consultation, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the human-computer interaction method for legal consultation according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Huffman coding method, system and device and medium

    CN116505954A

  • Vehicle sales information optimization storage method based on data analysis

    CN117278055A

  • Fast coding method based on Huffman coding

    CN117375631A

  • Word vector training method and device for large language model and medium

    CN118536504A

  • Data security transmission method and system for smart property

    CN118555036A