Question and answer method for power intelligent question and answer system
By dynamically adjusting the structure of the Hoffman tree, we can judge whether the Hoffman tree needs to be reconstructed based on the probability changes of each type of word segmentation in the user interactive text, solving the problem of the gradual deterioration of the compression effect of traditional Hoffman encoding, and improving the response speed and resource utilization efficiency of the power intelligent question-and-answer system.
Patent Information
- Application Number
- CN202510475037.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When the traditional power intelligent question and answer system processes massive user interactive text, the compression effect of Hoffman encoding gradually deteriorates with the increase of user interactive text, resulting in slow response speed and high resource consumption.
By obtaining the smoothing parameters and predicted probability of each category of word segmentation, the structure of the Hoffman tree is dynamically adjusted, and based on the probability changes of each category of word segmentation in the next user interaction text, it is determined whether the Hoffman tree needs to be reconstructed to improve the compression effect.
By dynamically reconstructing the Hoffman tree, the compression effect of the Hoffman coding algorithm can be effectively improved, resource consumption can be reduced, and the real-time and response speed of the system can be improved.
Smart Images

Figure CN120045530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing. More specifically, the present invention relates to a question-answering method for an electric power intelligent question-answering system. Background Art
[0002] With the rapid development of smart grid technology, the power industry is facing increasingly complex management and service requirements. The information interaction between power users and power management systems is becoming increasingly frequent. How to provide services efficiently and accurately has become a problem to be solved. Although traditional electric power intelligent question-answering systems have achieved automated services to a certain extent, they still face problems such as slow response speed, large data transmission volume, and high system resource consumption. In an electric power intelligent question-answering system, users ask questions in text, and the system needs to perform real-time data query and analysis and quickly return the corresponding answers. To meet the timely needs, the system needs to efficiently process a large amount of data and reduce the resource consumption during the transmission and calculation processes while ensuring accuracy. Therefore, for a large amount of user interaction text, by using data compression technology, redundant data can be effectively reduced, the transmission efficiency can be improved, the system burden can be alleviated. At the same time, the compressed data can accelerate the speed of question answering and improve the real-time performance of the system.
[0003] Currently, the patent application document with the publication number CN115987296A discloses a traffic energy data compression and transmission method based on Huffman coding. By obtaining a time-series discrete data set, a normal data set and an important data set are obtained; all data samples and important data samples in the time-series discrete data set are obtained; the importance degree of the important data samples is obtained according to the important data set and the normal data set; when the overall frequency of the important data samples is greater than or equal to the importance degree, the weight type of the important data samples is the overall frequency; otherwise, according to the compression volume growth rate and the total compression loss rate of the important data samples, the overall data compression loss amount of the important data samples is obtained, and then the weight type of the important data samples is obtained; encoding and compression are performed according to all important data samples and the corresponding weight types to obtain compressed data, and the compressed data is transmitted.
[0004] It is known that Huffman coding is a commonly used entropy coding method. It constructs a Huffman tree by pre-statistically calculating the occurrence frequencies of each character, making the coding length of characters with higher frequencies shorter, thereby achieving compression. For a power intelligent question-answering system, during the interaction between the user and the question-answering system, since the user's knowledge of the power system may be limited, the user will start with simple questions and gradually delve into more complex or specific questions during the interaction with the question-answering system, that is, the user's needs may change during the interaction with the question-answering system. Therefore, the keywords in the interaction information between the user and the question-answering system will gradually change. If Huffman coding is used to compress the word segmentation of the user interaction text by constructing a Huffman code, as the question-answering process is usually progressive, the probability of each type of word segmentation in the user interaction text will change, so the compression effect will gradually deteriorate as the user interaction text increases. Summary of the Invention
[0005] To solve the problem that when using Huffman coding to compress the word segmentation of the user interaction text by constructing a Huffman code, as the question-answering process is usually progressive, the probability of each type of word segmentation in the user interaction text will change, so the compression effect will gradually deteriorate as the user interaction text increases, the present invention proposes a question-answering method for a power intelligent question-answering system, and the method includes the following steps: Divide the current user interaction text into several word segmentations; obtain the difference sequence of each type of word segmentation; obtain the target word segmentation of each type of word segmentation, and the word vectors of each type of word segmentation and its target word segmentation are similar; Obtain the smoothing parameter of each type of word segmentation , represents the smoothing parameter of the i-th type of word segmentation; represents the number of target word segmentations of the i-th type of word segmentation; represents the similarity between the i-th type of word segmentation and its j-th target word segmentation; represents the sum of the similarities between the i-th type of word segmentation and all its target word segmentations; represents the standard deviation of the difference sequence of the j-th target word segmentation of the i-th type of word segmentation; represents the standard deviation of the difference sequence of the i-th type of word segmentation; represents the preset initial smoothing parameter; norm() represents the normalization function; based on the smoothing parameter, obtain the predicted probability corresponding to the next occurrence of each type of word segmentation; Through the Huffman coding algorithm, construct the current Huffman tree according to several word segmentations, and obtain the coding length corresponding to each type of word segmentation; according to the coding length corresponding to each type of word segmentation and the predicted probability corresponding to the next occurrence of each type of word segmentation, obtain the necessity of reconstructing the current Huffman tree; based on the necessity of reconstruction, determine whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text.
[0006] The innovation of the present invention lies in predicting all the corresponding probabilities when each type of word segmentation appears next according to the smoothing parameter of each type of word segmentation and the probability corresponding to each appearance of each type of word segmentation, obtaining the predicted probability corresponding to the next appearance of each type of word segmentation, and obtaining the necessity of reconstructing the Huffman tree according to the change of the probability of each type of word segmentation in the next user interaction text. If the probability of each type of word segmentation changes in the next user interaction text, it is necessary to reconstruct the Huffman tree and then compress the next user interaction text, improving the compression effect of the Huffman coding algorithm.
[0007] Preferably, the dividing the current user interaction text into several word segmentations includes: Using the Jieba word segmentation technology to segment the current user interaction text to obtain several word segmentations.
[0008] Preferably, the obtaining the difference sequence of each type of word segmentation includes: The probability corresponding to the b-th appearance of the i-th type of word segmentation ; b represents the ordinal number of the i-th type of word segmentation appearing in the current user interaction text; represents the index position corresponding to the b-th appearance of the i-th type of word segmentation in the current user interaction text; It is convenient to subsequently judge the stability degree of the probability corresponding to each appearance of each type of word segmentation according to the difference sequence of each type of word segmentation, and then adaptively adjust the smoothing parameter of each type of word segmentation.
[0009] According to the front-back order of the appearance of the i-th type of word segmentation in the current user interaction text, sort the probabilities corresponding to each appearance of the i-th type of word segmentation to obtain the cumulative probability sequence of the i-th type of word segmentation. For the u-th data in the cumulative probability sequence, record the absolute value of the difference between the u-th data and the (u - 1)-th probability data as the first difference of the u-th data; record the sequence composed of all the first differences of the data in the cumulative probability sequence of the i-th type of word segmentation as the difference sequence of the i-th type of word segmentation.
[0010] Preferably, the obtaining the target word segmentation of each type of word segmentation includes: , represents the similarity between the i-th type of word segmentation and the c-th type of word segmentation; represents the dot product; represents the norm; represents the word vector corresponding to the i-th type of word segmentation; represents the word vector corresponding to the c-th type of word segmentation; norm() represents the normalization constant; Preset a similarity threshold. If the similarity between the i-th type of word segmentation and the c-th type of word segmentation is greater than the similarity threshold, then the c-th type of word segmentation is the target word segmentation of the i-th type of word segmentation, and the target word segmentation of each type of word segmentation is obtained.
[0011] Preferably, the obtaining of the prediction probability corresponding to the next occurrence of each type of word segment includes: Using the exponential smoothing method, based on the smoothing parameter of each type of word segment and the probability corresponding to each occurrence of each type of word segment, predict all the corresponding probabilities for the next occurrence of each type of word segment, and obtain the prediction probability corresponding to the next occurrence of each type of word segment.
[0012] It is convenient to obtain the necessity of reconstructing the current Huffman tree according to the change situation of the prediction probability corresponding to the next occurrence of each type of word segment.
[0013] Preferably, the obtaining of the necessity of reconstructing the current Huffman tree includes: Obtain the coding length sequence of each type of word segment and the prediction probability sequence of each type of word segment; ; In the formula, represents the necessity of reconstructing the current Huffman tree; H represents the number of types of word segments; represents the index position of the prediction probability corresponding to the next occurrence of the i-th type of word segment in the prediction probability sequence of the i-th type of word segment; represents the index position of the coding length corresponding to the i-th type of word segment in the coding length sequence of the i-th type of word segment; || represents the absolute value symbol; exp() represents the exponential function with the natural constant as the base.
[0014] The obtained necessity of reconstructing the current Huffman tree is more accurate.
[0015] Preferably, the obtaining of the coding length sequence of each type of word segment and the prediction probability sequence of each type of word segment includes: Sort the prediction probabilities corresponding to the next occurrence of each type of word segment in descending order to obtain the prediction probability sequence of the i-th type of word segment; sort the coding lengths corresponding to each type of word segment in ascending order to obtain the coding length sequence of the i-th type of word segment.
[0016] Preferably, based on the necessity of reconstruction, determining whether to reconstruct the current Huffman tree and then encoding the next user interaction text includes: Obtain the next user interaction text. If the necessity of reconstructing the current Huffman tree is less than the reconstruction necessity T2, compress the next user interaction text according to the current Huffman tree. If the necessity of reconstructing the current Huffman tree is greater than or equal to the reconstruction necessity T2, reconstruct the current Huffman tree according to the prediction probability corresponding to the next occurrence of each type of word segment and record it as the latest Huffman tree; use the latest Huffman tree to compress the next user interaction text; and so on, compress each user interaction text.
[0017] The compression effect is improved.
[0018] The present invention has the following beneficial effects: The purpose of the present invention is to obtain the smoothing parameter of each type of word segmentation according to the probability stability corresponding to each occurrence of the i-th type of word segmentation, and predict all the corresponding probabilities when each type of word segmentation appears next according to the smoothing parameter of each type of word segmentation and the probability corresponding to each occurrence of each type of word segmentation, so as to obtain the predicted probability corresponding to each type of word segmentation when it appears next. According to the change of the probability of each type of word segmentation in the next user interaction text, the necessity of reconstructing the Huffman tree is obtained. If the probability of each type of word segmentation changes in the next user interaction text, it is necessary to reconstruct the Huffman tree and then compress the next user interaction text, thereby improving the compression effect of the Huffman coding algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] By referring to the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein: Figure 1 is a flowchart of the steps of a question-answering method for an electric power intelligent question-answering system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] The following will describe in detail the specific embodiments of the present invention with reference to the drawings.
[0022] Please refer to Figure 1 , which shows a flowchart of the steps of a question-answering method for an electric power intelligent question-answering system provided by an embodiment of the present invention. The method includes the following steps: S001. Obtain the current user interaction text, divide the current user interaction text into several word segmentations, and obtain the word vector of each word segmentation.
[0023] In the embodiment of the present invention, through the electric power intelligent question-answering system, obtain the current input text of the user's question; retrieve according to the current input text in the existing knowledge base or database to obtain the current answer text; the current input text of the user's question and the current answer text are collectively referred to as the current user interaction text; The Jieba word segmentation technology is used to segment the current user interaction text to obtain several word segments; the Word2Vec technology is used to obtain the word vectors of each word segment. Among them, the Jieba word segmentation technology and the Word2Vec technology are existing technologies, and will not be elaborated here in this embodiment.
[0024] S002. Obtain the probability corresponding to each occurrence of each type of word segment, obtain the smoothing parameter of each type of word segment, and predict all the corresponding probabilities when each type of word segment appears next time according to the probability corresponding to each occurrence of each type of word segment and the smoothing parameter of each type of word segment, so as to obtain the predicted probability corresponding to each type of word segment when it appears next time.
[0025] It should be noted that for a large amount of user Q&A text, by using data compression technology, redundant data can be effectively reduced, the transmission efficiency can be improved, the system burden can be reduced. At the same time, the compressed data can accelerate the Q&A speed and improve the real-time performance of the system. It is known that Huffman coding is a common type of entropy coding, which constructs a Huffman tree by pre-statistically counting the occurrence frequencies of each character, making the coding length of characters with higher frequencies shorter, so as to achieve compression. For the power intelligent Q&A system, during the interaction process between the user and the Q&A system, since the user's knowledge of the power system may be limited, therefore, when the user interacts with the Q&A system, they will start with simple questions and gradually delve into more complex or specific questions, that is, the user's needs may change during the interaction process with the Q&A system. Especially when they get some preliminary information, they will realize that they still need more details. Therefore, the keywords of the interaction information between the user and the Q&A system will gradually change. Therefore, if Huffman coding is used to compress the word segments of the user interaction text by constructing Huffman coding, as the Q&A process is usually step by step, the probability of each type of word segment in the user interaction text appearing will change, so the compression effect will gradually deteriorate as the user interaction text increases.
[0026] Therefore, the present invention first obtains the probability corresponding to each occurrence of each type of word segment in the current user interaction text, and then predicts all the corresponding probabilities when each type of word segment appears next time. According to the probability change situation of each type of word segment, it is judged whether it is necessary to reconstruct the Huffman tree when encoding the next user interaction text, so as to improve the compression effect.
[0027] In the embodiment of the present invention, the same word segments are recorded as one type of word segment to obtain several types of word segments. Obtain the probability corresponding to each occurrence of each type of word segment: ; In the formula, represents the probability corresponding to the b-th occurrence of the i-th type of word segmentation; b represents the ordinal number of the i-th type of word segmentation in the current user interaction text; represents the index position corresponding to the b-th occurrence of the i-th type of word segmentation in the current user interaction text; It should be noted that, for example, if the current user interaction text is abbcbbc, several word segmentations of the current user interaction text are a, bb, c, bb, c; among them, the types of word segmentations are a, bb, and c; the index position corresponding to the first occurrence of the word segmentation type bb in the current user interaction text is 2, because there are two word segmentations after the first occurrence of the word segmentation type bb in the current user interaction text; the index position corresponding to the second occurrence of the word segmentation type bb in the current user interaction text is 4.
[0028] According to the order of appearance of the i-th type of word segmentation in the current user interaction text, sort the probabilities corresponding to each occurrence of the i-th type of word segmentation to obtain the cumulative probability sequence of each type of word segmentation.
[0029] It should be noted that when predicting the probability of any type of word segmentation in the next user interaction text, by analyzing the probability change situation of this type of word segmentation in the current user interaction text, if its change degree is small, it means that a smaller smoothing parameter can be used for prediction, if its change degree is large, it means that a larger smoothing parameter can be used for prediction. In the interaction process between the user and the power intelligent question answering system, similar word segmentations usually appear together. Therefore, after a certain word segmentation appears, it means that a word segmentation similar to it may also appear. Therefore, obtain the similarity between any two types of word segmentations. If the probability change degree of the similar word segmentation of any type of word segmentation in the current user interaction text is small, it means that the probability change degree of this type of word segmentation in the current user interaction text is also small, indicating that a smaller smoothing parameter can be used for prediction.
[0030] In the embodiment of the present invention, for the u-th data in the cumulative probability sequence of the i-th type of word segmentation, the absolute value of the difference between the u-th data and the (u - 1)-th probability data is denoted as the first difference of the u-th data; the sequence composed of the first differences of all data in the cumulative probability sequence of the i-th type of word segmentation is denoted as the difference sequence of the i-th type of word segmentation, where u is greater than or equal to 2.
[0031] Obtain the similarity between any two types of word segmentations: ; In the formula, represents the similarity between the i-th type of word segmentation and the c-th type of word segmentation; represents dot product; represents the modulus length; represents the word vector corresponding to the i-th type of word segmentation; It represents the word vector corresponding to the c-th type of word segmentation; norm() represents the normalization constant; similarly, the similarity between every two types of word segmentation is obtained; A preset similarity threshold is set. If the similarity between the i-th type of word segmentation and the c-th type of word segmentation is greater than the similarity threshold, then the c-th type of word segmentation is the target word segmentation of the i-th type of word segmentation, and the target word segmentation of each type of word segmentation is obtained; Obtain the smoothing parameter of each type of word segmentation: ; In the formula, represents the smoothing parameter of the i-th type of word segmentation; represents the number of target word segmentations of the i-th type of word segmentation; represents the similarity between the i-th type of word segmentation and its j-th target word segmentation; represents the sum of the similarities between the i-th type of word segmentation and all its target word segmentations; represents the standard deviation of the difference sequence of the j-th target word segmentation of the i-th type of word segmentation; represents the standard deviation of the difference sequence of the i-th type of word segmentation; represents the preset initial smoothing parameter; in the embodiment of the present invention, the preset initial smoothing parameter = 0.5. In other embodiments, the implementer can preset the value of the initial smoothing parameter according to the specific implementation situation ; norm() represents the normalization function; represents the degree of fluctuation of the difference sequence of the i-th type of word segmentation. The smaller its value, the more stable the probability corresponding to each occurrence of the i-th type of word segmentation. Therefore, a smaller smoothing parameter can be used for prediction; represents the mean value of the degrees of fluctuation of the difference sequences of all target word segmentations of the i-th type of word segmentation. The smaller its value, the more stable the probability corresponding to each occurrence of the target word segmentation of the i-th type of word segmentation, which means that the probability corresponding to each occurrence of the i-th type of word segmentation is more stable. Therefore, a smaller smoothing parameter can be used for prediction.
[0032] Using the exponential smoothing method, according to the smoothing parameter of each type of word segmentation and the probability corresponding to each occurrence of each type of word segmentation, predict the probabilities corresponding to all occurrences of each type of word segmentation next time, and obtain the prediction probabilities corresponding to each type of word segmentation next time.
[0033] S003. Use Huffman coding to construct the current Huffman tree for the word segmentation of the current user interaction text. According to the prediction probability corresponding to each type of word segmentation next time and the current Huffman tree, obtain the necessity of reconstructing the current Huffman tree.
[0034] It should be noted that if the Huffman tree remains unchanged, as the user continues to interact with the power intelligent question-and-answer system, the frequencies of various word segmentations will change. Continuing to use the original Huffman tree for compression will result in poor compression efficiency. Therefore, the present invention determines whether the probability of each word segmentation in the next user interaction text has changed based on the predicted probability corresponding to the next occurrence of each word segmentation. If the probability has changed, the necessity of reconstructing the Huffman tree is higher.
[0035] In the embodiment of the present invention, the predicted probabilities corresponding to the next occurrence of each word segmentation are sorted in descending order to obtain the predicted probability sequence of the i-th word segmentation. Based on several word segmentations of the current user interaction text, a current Huffman tree is constructed through the Huffman coding algorithm, and each word segmentation is encoded to obtain the coding length corresponding to each word segmentation. The coding lengths corresponding to each word segmentation are sorted in ascending order to obtain the coding length sequence of the i-th word segmentation. Obtain the necessity of reconstructing the current Huffman tree: ; In the formula, represents the necessity of reconstructing the current Huffman tree; H represents the number of word segmentation categories; represents the index position of the predicted probability corresponding to the next occurrence of the i-th word segmentation in the predicted probability sequence of the i-th word segmentation; represents the index position of the coding length corresponding to the i-th word segmentation in the coding length sequence of the i-th word segmentation; || represents the absolute value symbol; exp() represents the exponential function with the natural constant as the base; it is known that the shorter the coding length of a word segmentation, the greater its probability. Therefore, the larger the value of, the more it indicates that the probability of the i-th word segmentation in the next user interaction text has changed, and the higher the necessity of reconstructing the Huffman tree.
[0036] S004. Determine whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text according to the necessity of reconstructing the current Huffman tree.
[0037] It should be noted that it is determined whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text according to the necessity of reconstructing the current Huffman tree.
[0038] In the embodiment of the present invention, a preset reconstruction necessity T2 = 0.7 is set. In other embodiments, the implementer can preset the value of the reconstruction necessity according to the specific implementation manner; First, compress the current interactive text according to the current Huffman tree, and then obtain the next user interactive text. If the necessity of reconstructing the current Huffman tree is less than the reconstruction necessity T2, compress the next user interactive text according to the current Huffman tree. If the necessity of reconstructing the current Huffman tree is greater than or equal to the reconstruction necessity T2, reconstruct the current Huffman tree according to the predicted probability corresponding to each type of word segmentation when it appears next time, and record it as the latest Huffman tree; use the latest Huffman tree to compress the next user interactive text; and so on, compress each user interactive text.
[0039] It should be noted that for the convenience of decompression, if the next user interactive text is compressed using the reconstructed Huffman tree, a delimiter needs to be added to the existing compressed data before encoding using the reconstructed Huffman tree to facilitate subsequent decompression.
[0040] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A question-answering method for an electric power intelligent question-answering system, characterized in that: include: Divide the current user interaction text into several segmented words; obtain a difference sequence of each type of segmented words; obtain a target segmented word of each type of segmented word, wherein each type of segmented word has a similar word vector to its target segmented word; Get the smoothing parameters for each type of word segmentation , Represents the smoothing parameter of the i-th category segmentation; Represents the number of target segmentations of the i-th category; Represents the similarity between the i-th category segmentation and its j-th target segmentation; Represents the sum of similarities between the i-th category segmentation and all its target segmentations; Represents the standard deviation of the difference sequence of the jth target word of the i-th category; Represents the standard deviation of the difference sequence of the i-th category of segmentation; represents the preset initial smoothing parameter; norm() represents the normalization function; based on the smoothing parameter, the predicted probability corresponding to the next occurrence of each type of word segmentation is obtained; The current Huffman tree is constructed according to several word segments through the Huffman coding algorithm to obtain the coding length corresponding to each type of word segmentation; the necessity of reconstructing the current Huffman tree is obtained according to the coding length corresponding to each type of word segmentation and the predicted probability corresponding to the next appearance of each type of word segmentation; based on the necessity of reconstruction, it is determined whether it is necessary to reconstruct the current Huffman tree and encode the next user interaction text.
2. A question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that: The current user interaction text is divided into a number of segmented words, including: Jieba word segmentation technology is used to segment the current user interaction text to obtain several word segments.
3. A question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that: The step of obtaining a difference sequence of each type of word segmentation includes: The probability corresponding to the bth occurrence of the i-th type of participle ; b represents the ordinal number of the i-th type of word in the current user interaction text; Represents the index position corresponding to the bth occurrence of the i-th type of word in the current user interaction text; According to the order in which the i-th type of word appears in the current user interaction text, the probabilities corresponding to each appearance of the i-th type of word are sorted to obtain the cumulative probability sequence of the i-th type of word. For the u-th data in the cumulative probability sequence, the absolute value of the difference between the u-th data and the u-1-th probability data is recorded as the first difference of the u-th data; the sequence composed of the first difference of all data in the cumulative probability sequence of the i-th type of word is recorded as the difference sequence of the i-th type of word.
4. A question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that: The step of obtaining a target word segmentation for each type of word segmentation includes: , represents the similarity between the i-th category participle and the c-th category participle; represents dot product; Indicates the module length; Represents the word vector corresponding to the i-th category of word segmentation; Represents the word vector corresponding to the c-th type of word segmentation; norm() represents the normalization constant; A similarity threshold is preset. If the similarity between the i-th category segmentation and the c-th category segmentation is greater than the similarity threshold, the c-th category segmentation is the target segmentation of the i-th category segmentation, and the target segmentation of each category of segmentation is obtained.
5. The question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that: The obtaining of the predicted probability corresponding to the next occurrence of each type of word segmentation includes: Using the exponential smoothing method, according to the smoothing parameters of each type of word segmentation and the probability corresponding to each type of word segmentation each time it appears, all corresponding probabilities of each type of word segmentation when it appears next are predicted, and the predicted probability corresponding to the next appearance of each type of word segmentation is obtained.
6. A question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that: The necessity of reconstructing the current Huffman tree is obtained, including: Obtain the encoding length sequence of each type of word segmentation and the predicted probability sequence of each type of word segmentation; ; In the formula, Represents the necessity of reconstructing the current Huffman tree; H represents the number of word categories; Represents the index position of the predicted probability corresponding to the next occurrence of the i-th type of word in the predicted probability sequence of the i-th type of word; Represents the index position of the code length corresponding to the i-th type of word in the code length sequence of the i-th type of word; || represents the absolute value symbol; exp() represents an exponential function with a natural constant as the base.
7. A question-answering method for an electric power intelligent question-answering system according to claim 6, characterized in that: The step of obtaining the encoding length sequence of each type of word segmentation and the prediction probability sequence of each type of word segmentation includes: According to the order from large to small, the corresponding predicted probabilities of each type of word segmentation when it appears next time are sorted to obtain the predicted probability sequence of the i-th type of word segmentation; according to the order from small to large, the corresponding coding lengths of each type of word segmentation are sorted to obtain the coding length sequence of the i-th type of word segmentation.
8. A question-answering method for an electric power intelligent question-answering system according to claim 1 or 5, characterized in that: The determining whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text based on the reconstruction necessity includes: Get the next user interaction text. If the reconstruction necessity of the current Huffman tree is less than the reconstruction necessity T2, compress the next user interaction text according to the current Huffman tree. If the reconstruction necessity of the current Huffman tree is greater than or equal to the reconstruction necessity T2, reconstruct the current Huffman tree according to the predicted probability corresponding to the next appearance of each type of word, and record it as the latest Huffman tree; use the latest Huffman tree to compress the next user interaction text; and so on, compress each user interaction text.
Citation Information
Patent Citations
Method for dynamically compressing and decompressing large file based on segmentation Huffman
CN111510156A
Transportation energy data compression and transmission method based on Huffman coding
CN115987296A
Intelligent ring information management method and system based on cloud edge collaboration
CN116153453A
Intelligent question answering method and device based on power consumer portrait and electronic equipment
CN116701584A
Data compressing apparatus, reconstructing apparatus, and its method
US6542640B1