A Q&A method for an intelligent power Q&A system

By dynamically adjusting the construction of the Hoffman tree according to the probability changes of the user's interactive text in the power intelligent question and answer system, the problem of compression effect deterioration with the interaction process is solved, and more efficient data compression and system performance improvement is achieved.

CN120045530BActive Publication Date: 2025-07-18HUBEI ANYUAN SAFETY & ENVIRONMENTAL PROTECTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510475037.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the power intelligent question and answer system, when Hoffmann encoding is used to compress the Hoffman tree for word segmentation construction of user interactive text, as the question and answer process progresses, the probability of each type of word segmentation in user interactive text changes, causing the compression effect to gradually worsen.

Method used

By obtaining the smoothing parameters of each type of word segmentation and the probability of each occurrence, predicting the probability of the next occurrence, and determining whether the Hoffman tree needs to be reconstructed based on the necessity of reconstruction of the Hoffman tree to adapt to the changes in user interaction text and improve the compression effect.

Benefits of technology

Through adaptive adjustment of the construction of the Hoffman tree, the compression efficiency is improved, the system resource consumption is reduced, and the real-time and transmission efficiency of the Q&A system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045530B_ABST
    Figure CN120045530B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing. More specifically, the present invention relates to a question-and-answer method for an intelligent power question-and-answer system. The method includes dividing the current user interaction text into a number of word segments, obtaining the word vectors of each word segment, acquiring the probability corresponding to each occurrence of each type of word segment and the smoothing parameter of each type of word segment, predicting all the probabilities corresponding to the next occurrence of each type of word segment, and obtaining the predicted probability corresponding to the next occurrence of each type of word segment; using Huffman coding to construct a current Huffman tree for the word segments of the current user interaction text, obtaining the necessity for reconstructing the current Huffman tree according to the predicted probability corresponding to the next occurrence of each type of word segment and the current Huffman tree, and based on the necessity for reconstruction, determining whether it is necessary to reconstruct the current Huffman tree and then encoding the next user interaction text. The present invention improves the compression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing. More specifically, the present invention relates to a question-answering method for an electric power intelligent question-answering system. Background Art

[0002] With the rapid development of smart grid technology, the power industry is facing increasingly complex management and service requirements. The information interaction between power users and power management systems is becoming increasingly frequent. How to provide services efficiently and accurately has become a problem to be solved. Although traditional electric power intelligent question-answering systems have realized automated services to a certain extent, they still face problems such as slow response speed, large data transmission volume, and high system resource consumption. In an electric power intelligent question-answering system, users ask questions in text, and the system needs to perform data query and analysis in real time and quickly return the corresponding answers. To meet the immediate needs, the system needs to efficiently process a large amount of data and reduce the resource consumption during the transmission and calculation processes while ensuring accuracy. Therefore, for a large amount of user interaction text, by using data compression technology, redundant data can be effectively reduced, the transmission efficiency can be improved, the system burden can be alleviated. At the same time, the compressed data can accelerate the speed of question-answering and improve the real-time performance of the system.

[0003] Currently, the patent application document with the publication number CN115987296A discloses a traffic energy data compression and transmission method based on Huffman coding. By obtaining a time-series discrete data set, a normal data set and an important data set are obtained; all data samples and important data samples in the time-series discrete data set are obtained; the importance degree of the important data samples is obtained according to the important data set and the normal data set; when the overall frequency of the important data samples is greater than or equal to the importance degree, the weight type of the important data samples is the overall frequency; otherwise, according to the compression volume growth rate and the total compression loss rate of the important data samples, the overall data compression loss amount of the important data samples is obtained, and then the weight type of the important data samples is obtained; encoding and compression are performed according to all the important data samples and the corresponding weight types to obtain compressed data, and the compressed data is transmitted.

[0004] It is known that Huffman coding is a commonly used entropy coding method. It constructs a Huffman tree by pre - counting the occurrence frequencies of each character, making the coding length of characters with higher frequencies shorter, thus achieving compression. For a power intelligent question - answering system, during the interaction between the user and the question - answering system, since the user's knowledge of the power system may be limited, the user will start with simple questions and gradually delve into more complex or specific questions during the interaction with the question - answering system. That is, the user's needs may change during the interaction with the question - answering system. Therefore, the keywords of the interaction information between the user and the question - answering system will gradually change. If Huffman coding is used to compress the word segmentation of the user interaction text by constructing Huffman coding, as the question - answering process is usually step - by - step, the occurrence probability of each type of word segmentation in the user interaction text will change, so the compression effect will gradually deteriorate as the user interaction text increases. Summary of the Invention

[0005] To solve the problem that when using Huffman coding to compress the word segmentation of the user interaction text by constructing Huffman coding, as the question - answering process is usually step - by - step, the occurrence probability of each type of word segmentation in the user interaction text will change, so the compression effect will gradually deteriorate as the user interaction text increases, the present invention proposes a question - answering method for a power intelligent question - answering system, and the method includes the following steps:

[0006] Divide the current user interaction text into several word segments; obtain the difference sequence of each type of word segment; obtain the target word segment of each type of word segment, and the word vectors of each type of word segment and its target word segment are similar;

[0007] Obtain the smoothing parameter of each type of word segment , represents the smoothing parameter of the i - th type of word segment; represents the number of target word segments of the i - th type of word segment; represents the similarity between the i - th type of word segment and its j - th target word segment; represents the sum of the similarities between the i - th type of word segment and all its target word segments; represents the standard deviation of the difference sequence of the j - th target word segment of the i - th type of word segment; represents the standard deviation of the difference sequence of the i - th type of word segment; represents the preset initial smoothing parameter; norm() represents the normalization function; based on the smoothing parameter, obtain the predicted probability corresponding to each type of word segment when it appears next;

[0008] Construct the current Huffman tree according to several word segments by using the Huffman coding algorithm, and obtain the coding length corresponding to each type of word segment; obtain the necessity of reconstructing the current Huffman tree according to the coding length corresponding to each type of word segment and the predicted probability corresponding to the next occurrence of each type of word segment; based on the necessity of reconstruction, determine whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text.

[0009] The innovation of the present invention lies in predicting all the corresponding probabilities of the next occurrence of each type of word segment according to the smoothing parameter of each type of word segment and the probability corresponding to each occurrence of each type of word segment, obtaining the predicted probability corresponding to the next occurrence of each type of word segment, obtaining the necessity of reconstructing the Huffman tree according to the change of the probability of each type of word segment in the next user interaction text. If the probability of each type of word segment changes in the next user interaction text, it is necessary to reconstruct the Huffman tree and then compress the next user interaction text, improving the compression effect of the Huffman coding algorithm.

[0010] Preferably, the dividing the current user interaction text into several word segments includes:

[0011] Segment the current user interaction text by using the Jieba word segmentation technology to obtain several word segments.

[0012] Preferably, the obtaining the difference sequence of each type of word segment includes:

[0013] The probability corresponding to the b-th occurrence of the i-th type of word segment ; b represents the ordinal number of the i-th type of word segment appearing in the current user interaction text; represents the index position corresponding to the b-th occurrence of the i-th type of word segment in the current user interaction text;

[0014] It is convenient to subsequently judge the stability degree of the probability corresponding to each occurrence of each type of word segment according to the difference sequence of each type of word segment, and then adaptively adjust the smoothing parameter of each type of word segment.

[0015] Sort the probabilities corresponding to each occurrence of the i-th type of word segment according to the front-back order of the i-th type of word segment appearing in the current user interaction text to obtain the cumulative probability sequence of the i-th type of word segment. For the u-th data in the cumulative probability sequence, record the absolute value of the difference between the u-th data and the (u - 1)-th probability data as the first difference of the u-th data; record the sequence composed of all the first differences of the cumulative probability sequence of the i-th type of word segment as the difference sequence of the i-th type of word segment.

[0016] Preferably, the obtaining the target word segment of each type of word segment includes:

[0017] , Represents the similarity between the i-th category of word segmentation and the c-th category of word segmentation; Denotes dot product; Denotes the norm; Represents the word vector corresponding to the i-th category of word segmentation; Represents the word vector corresponding to the c-th category of word segmentation; norm() represents the normalization constant;

[0018] A preset similarity threshold. If the similarity between the i-th category of word segmentation and the c-th category of word segmentation is greater than the similarity threshold, then the c-th category of word segmentation is the target word segmentation of the i-th category of word segmentation, and the target word segmentation of each category of word segmentation is obtained.

[0019] Preferably, the obtaining of the prediction probability corresponding to the next occurrence of each category of word segmentation includes:

[0020] Using the exponential smoothing method, according to the smoothing parameter of each category of word segmentation and the probability corresponding to each occurrence of each category of word segmentation, predict all the probabilities corresponding to the next occurrence of each category of word segmentation, and obtain the prediction probability corresponding to the next occurrence of each category of word segmentation.

[0021] It is convenient to obtain the necessity of reconstructing the current Huffman tree according to the change situation of the prediction probability corresponding to the next occurrence of each category of word segmentation.

[0022] Preferably, the obtaining of the necessity of reconstructing the current Huffman tree includes:

[0023] Obtain the encoding length sequence of each category of word segmentation and the prediction probability sequence of each category of word segmentation;

[0024] ;

[0025] In the formula, Represents the necessity of reconstructing the current Huffman tree; H represents the number of categories of word segmentation; Represents the index position of the prediction probability corresponding to the next occurrence of the i-th category of word segmentation in the prediction probability sequence of the i-th category of word segmentation; Represents the index position of the encoding length corresponding to the i-th category of word segmentation in the encoding length sequence of the i-th category of word segmentation; || represents the absolute value symbol; exp() represents the exponential function with the natural constant as the base.

[0026] The obtained necessity of reconstructing the current Huffman tree is more accurate.

[0027] Preferably, the obtaining of the encoding length sequence of each category of word segmentation and the prediction probability sequence of each category of word segmentation includes:

[0028] Sort the predicted probabilities corresponding to the next occurrence of each type of word segmentation in descending order to obtain the predicted probability sequence of the i-th type of word segmentation; sort the encoding lengths corresponding to each type of word segmentation in ascending order to obtain the encoding length sequence of the i-th type of word segmentation.

[0029] Preferably, based on the necessity of reconstruction, determining whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text includes:

[0030] Obtain the next user interaction text. If the reconstruction necessity of the current Huffman tree is less than the reconstruction necessity T2, compress the next user interaction text according to the current Huffman tree. If the reconstruction necessity of the current Huffman tree is greater than or equal to the reconstruction necessity T2, reconstruct the current Huffman tree according to the predicted probabilities corresponding to the next occurrence of each type of word segmentation, and denote it as the latest Huffman tree; use the latest Huffman tree to compress the next user interaction text; and so on, compress each user interaction text.

[0031] The compression effect is improved.

[0032] The present invention has the following beneficial effects: The purpose of the present invention is to obtain the smoothing parameter of each type of word segmentation according to the probability stability corresponding to each occurrence of the i-th type of word segmentation, predict all the corresponding probabilities of the next occurrence of each type of word segmentation according to the smoothing parameter of each type of word segmentation and the probability corresponding to each occurrence of each type of word segmentation, obtain the reconstruction necessity of the Huffman tree according to the change situation of the probability of each type of word segmentation in the next user interaction text. If the probability of each type of word segmentation changes in the next user interaction text, it is necessary to reconstruct the Huffman tree and then compress the next user interaction text, improving the compression effect of the Huffman coding algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0034] Figure 1 is a flowchart of the steps of a question-answering method for an electric power intelligent question-answering system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present invention.

[0036] The following will specifically describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0037] Please refer to Figure 1 , which shows a step flow chart of a question-and-answer method for an electric power intelligent question-and-answer system provided by an embodiment of the present invention. The method includes the following steps:

[0038] S001. Obtain the current user interaction text, divide the current user interaction text into several word segments, and obtain the word vectors of each word segment.

[0039] In the embodiment of the present invention, through the electric power intelligent question-and-answer system, obtain the current input text of the user's question; retrieve according to the current input text in the existing knowledge base or database to obtain the current answer text; the current input text of the user's question and the current answer text are collectively referred to as the current user interaction text;

[0040] Use the Jieba word segmentation technology to segment the current user interaction text to obtain several word segments; use the Word2Vec technology to obtain the word vectors of each word segment; among them, the Jieba word segmentation technology and the Word2Vec technology are existing technologies, and will not be elaborated here in this embodiment.

[0041] S002. Obtain the probability corresponding to each occurrence of each type of word segment, obtain the smoothing parameter of each type of word segment, and predict all the corresponding probabilities when each type of word segment appears next time according to the probability corresponding to each occurrence of each type of word segment and the smoothing parameter of each type of word segment, so as to obtain the predicted probability corresponding to each type of word segment when it appears next time.

[0042] It should be noted that for a large amount of user question-and-answer texts, by using data compression technology, redundant data can be effectively reduced, the transmission efficiency can be improved, the system burden can be reduced, and at the same time, the compressed data can accelerate the speed of question-and-answer and improve the real-time performance of the system. It is known that Huffman coding is a common type of entropy coding, which constructs a Huffman tree by pre-statistically counting the occurrence frequencies of each character, making the coding length of characters with higher frequencies shorter, so as to achieve compression;

[0043] For a power intelligent question - answering system, during the interaction between the user and the question - answering system, since the user's knowledge of the power system may be limited, when the user interacts with the question - answering system, the user will start with simple questions and gradually delve into more complex or specific questions. That is, the user's needs may change during the interaction with the question - answering system. Especially when they get some preliminary information, they will realize that they need more details. Therefore, the keywords of the interaction information between the user and the question - answering system will gradually change. Thus, if Huffman coding is used to compress the word segmentation of the user interaction text to construct a Huffman code, as the question - answering process is usually progressive, the probability of each type of word segmentation in the user interaction text will change, so the compression effect will gradually deteriorate as the user interaction text increases.

[0044] Therefore, the present invention first obtains the probability corresponding to each occurrence of each type of word segmentation in the current user interaction text, and then predicts all the corresponding probabilities for the next occurrence of each type of word segmentation. According to the probability change situation of each type of word segmentation, it is judged whether it is necessary to reconstruct the Huffman tree when encoding the next user interaction text, thereby improving the compression effect.

[0045] In the embodiment of the present invention, the same word segmentations are recorded as one type of word segmentation, and several types of word segmentations are obtained;

[0046] Obtain the probability corresponding to each occurrence of each type of word segmentation:

[0047] ;

[0048] In the formula, represents the probability corresponding to the b - th occurrence of the i - th type of word segmentation; b represents the ordinal number of the i - th type of word segmentation in the current user interaction text; represents the index position corresponding to the b - th occurrence of the i - th type of word segmentation in the current user interaction text;

[0049] It should be noted that, for example, if the current user interaction text is abbcbbc, the several word segmentations of the current user interaction text are a, bb, c, bb, c; among them, the types of word segmentations are a, bb, and c; when the word segmentation type bb appears for the first time in the current user interaction text, the corresponding index position is 2, because there are two word segmentations after the word segmentation type bb appears for the first time in the current user interaction text; when the word segmentation type bb appears for the second time in the current user interaction text, the corresponding index position is 4.

[0050] According to the order of appearance of the i - th type of word segmentation in the current user interaction text, sort the probabilities corresponding to each occurrence of the i - th type of word segmentation to obtain the cumulative probability sequence of each type of word segmentation.

[0051] It should be noted that when predicting the probability of any type of word segmentation in the next user interaction text, by analyzing the change in the probability of this type of word segmentation in the current user interaction text, if the degree of change is small, it indicates that a smaller smoothing parameter can be used for prediction; if the degree of change is large, it indicates that a larger smoothing parameter can be used for prediction. During the interaction process between the user and the power intelligent question answering system, similar word segmentations usually co-occur. Therefore, after a certain word segmentation appears, it indicates that word segmentations that are relatively similar to it may also appear. Therefore, obtaining the similarity between any two types of word segmentations, if the degree of change in the probability of the similar word segmentation of any type of word segmentation in the current user interaction text is small, it indicates that the degree of change in the probability of this type of word segmentation in the current user interaction text is also small, indicating that a smaller smoothing parameter can be used for prediction.

[0052] In the embodiment of the present invention, for the u-th data in the cumulative probability sequence of the i-th type of word segmentation, the absolute value of the difference between the u-th data and the (u - 1)-th probability data is denoted as the first difference of the u-th data; the sequence formed by all the first differences of the data in the cumulative probability sequence of the i-th type of word segmentation is denoted as the difference sequence of the i-th type of word segmentation, where u is greater than or equal to 2.

[0053] Obtain the similarity between any two types of word segmentations:

[0054] ;

[0055] In the formula, represents the similarity between the i-th type of word segmentation and the c-th type of word segmentation; represents dot product; represents the norm; represents the word vector corresponding to the i-th type of word segmentation; represents the word vector corresponding to the c-th type of word segmentation; norm() represents the normalization constant; similarly, obtain the similarity between each two types of word segmentations;

[0056] Preset a similarity threshold. If the similarity between the i-th type of word segmentation and the c-th type of word segmentation is greater than the similarity threshold, then the c-th type of word segmentation is the target word segmentation of the i-th type of word segmentation, and obtain the target word segmentation of each type of word segmentation;

[0057] Obtain the smoothing parameter of each type of word segmentation:

[0058] ;

[0059] In the formula, represents the smoothing parameter of the i-th type of word segmentation; represents the number of target word segmentations of the i-th type of word segmentation; represents the similarity between the i-th type of word segmentation and its j-th target word segmentation; represents the sum of the similarities between the i-th type of word segmentation and all its target word segmentations; The standard deviation of the difference sequence of the j-th target word segment of the i-th type of word segment; The standard deviation of the difference sequence of the i-th type of word segment; Represents a preset initial smoothing parameter; in the embodiments of the present invention, the preset initial smoothing parameter = 0.5. In other embodiments, the implementer can preset the value of the initial smoothing parameter according to the specific implementation situation; The value of norm() represents a normalization function; Represents the degree of fluctuation of the difference sequence of the i-th type of word segment. The smaller its value, the more stable the probability corresponding to each occurrence of the i-th type of word segment. Therefore, a smaller smoothing parameter can be used for prediction; Represents the mean value of the degree of fluctuation of the difference sequences of all target word segments of the i-th type of word segment. The smaller its value, the more stable the probability corresponding to each occurrence of the target word segment of the i-th type of word segment, which means that the probability corresponding to each occurrence of the i-th type of word segment is more stable. Therefore, a smaller smoothing parameter can be used for prediction.

[0060] Using the exponential smoothing method, according to the smoothing parameter of each type of word segment and the probability corresponding to each occurrence of each type of word segment, predict the probabilities corresponding to all occurrences of each type of word segment the next time they appear, and obtain the predicted probabilities corresponding to each type of word segment the next time they appear.

[0061] S003. Use Huffman coding to construct the current Huffman tree for the word segments of the current user interaction text, and obtain the necessity of reconstructing the current Huffman tree according to the predicted probabilities corresponding to each type of word segment the next time they appear and the current Huffman tree.

[0062] It should be noted that if the Huffman tree remains unchanged, as the user continues to interact with the power intelligent question-answering system, the frequencies of various types of word segments will change. Continuing to use the original Huffman tree for compression will result in poor compression efficiency. Therefore, in the present invention, according to the predicted probabilities corresponding to each type of word segment the next time they appear, it is judged whether the probabilities of each type of word segment in the next user interaction text have changed. If they have changed, the necessity of reconstructing the Huffman tree is higher.

[0063] In the embodiments of the present invention, sort the predicted probabilities corresponding to each type of word segment the next time they appear in descending order to obtain the predicted probability sequence of the i-th type of word segment;

[0064] Through the Huffman coding algorithm, construct the current Huffman tree according to several word segments of the current user interaction text, and encode each type of word segment to obtain the encoding length corresponding to each type of word segment. Sort the encoding lengths corresponding to each type of word segment in ascending order to obtain the encoding length sequence of the i-th type of word segment;

[0065] Obtain the necessity of reconstructing the current Huffman tree:

[0066] ;

[0067] In the formula, represents the necessity of reconstructing the current Huffman tree; H represents the number of categories of word segmentation; represents the index position of the predicted probability corresponding to the next occurrence of the i-th category of word segmentation in the predicted probability sequence of the i-th category of word segmentation; represents the index position of the coding length corresponding to the i-th category of word segmentation in the coding length sequence of the i-th category of word segmentation; || represents the absolute value symbol; exp() represents the exponential function with the natural constant as the base; it is known that the shorter the coding length of a category of word segmentation, the greater its probability. Therefore the larger the value of, the more it indicates that the probability of the i-th category of word segmentation in the next user interaction text has changed, and the higher the necessity of reconstructing the Huffman tree.

[0068] S004. According to the necessity of reconstructing the current Huffman tree, determine whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text.

[0069] It should be noted that according to the necessity of reconstructing the current Huffman tree, it is determined whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text.

[0070] In the embodiment of the present invention, the preset reconstruction necessity T2 = 0.7. In other embodiments, the implementer can preset the value of the reconstruction necessity according to the specific implementation method;

[0071] First, compress the current interaction text according to the current Huffman tree, then obtain the next user interaction text. If the necessity of reconstructing the current Huffman tree is less than the reconstruction necessity T2, compress the next user interaction text according to the current Huffman tree. If the necessity of reconstructing the current Huffman tree is greater than or equal to the reconstruction necessity T2, reconstruct the current Huffman tree according to the predicted probability corresponding to the next occurrence of each category of word segmentation, and denote it as the latest Huffman tree; use the latest Huffman tree to compress the next user interaction text; and so on, compress each user interaction text.

[0072] It should be noted that for the convenience of decompression, if the next user interaction text is compressed using the reconstructed Huffman tree, it is necessary to add a delimiter to the existing compressed data before encoding with the reconstructed Huffman tree to facilitate subsequent decompression.

[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A question-and-answer method for an intelligent power question-and-answer system, characterized in that, Including: Dividing the current user interaction text into several word segments; obtaining the difference sequence of each type of word segment; Obtaining the target word segment of each type of word segment, where the word vectors of each type of word segment and its target word segment are similar; The obtaining of the difference sequence of each type of word segment includes: The probability corresponding to the b-th occurrence of the i-th class of word segmentation ; b represents the ordinal number of the i-th class of word segmentation in the current user interaction text; represents the index position corresponding to the b-th occurrence of the i-th class of word segmentation in the current user interaction text; According to the front-back order of the occurrence of the i-th type of word segment in the current user interaction text, sorting the probabilities corresponding to each occurrence of the i-th type of word segment to obtain the cumulative probability sequence of the i-th type of word segment. For the u-th data in the cumulative probability sequence, taking the absolute value of the difference between the u-th data and the (u - 1)-th probability data as the first difference of the u-th data; the sequence composed of the first differences of all data in the cumulative probability sequence of the i-th type of word segment is denoted as the difference sequence of the i-th type of word segment; Obtain the smoothing parameter for each type of word segmentation , represents the smoothing parameter for the i-th type of word segmentation; represents the number of target word segmentations for the i-th type of word segmentation; represents the similarity between the i-th type of word segmentation and its j-th target word segmentation; represents the sum of the similarities between the i-th type of word segmentation and all its target word segmentations; represents the standard deviation of the difference sequence of the j-th target word segmentation of the i-th type of word segmentation; represents the standard deviation of the difference sequence of the i-th type of word segmentation; represents the preset initial smoothing parameter; norm() represents the normalization function; based on the smoothing parameter, obtain the predicted probability corresponding to the next occurrence of each type of word segmentation Constructing the current Huffman tree according to several word segments through the Huffman coding algorithm to obtain the coding length corresponding to each type of word segment; obtaining the reconstruction necessity of the current Huffman tree according to the coding length corresponding to each type of word segment and the predicted probability corresponding to the next occurrence of each type of word segment; based on the reconstruction necessity, judging whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text, and the encoded text is applied to the question-answering method of the power intelligent question-answering system to speed up the question-answering speed.

2. The question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that, The dividing of the current user interaction text into several word segments includes: Using the Jieba word segmentation technology to segment the current user interaction text to obtain several word segments.

3. The question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that, The obtaining of the target word segment of each type of word segment includes: , represents the similarity between the i-th category of word segmentation and the c-th category of word segmentation; denotes dot product; denotes the norm; represents the word vector corresponding to the i-th category of word segmentation; represents the word vector corresponding to the c-th category of word segmentation; norm() represents the normalization constant; Presetting a similarity threshold. If the similarity between the i-th type of word segment and the c-th type of word segment is greater than the similarity threshold, then the c-th type of word segment is the target word segment of the i-th type of word segment, and the target word segments of each type of word segment are obtained.

4. The question and answer method for an electric power intelligent question and answer system according to claim 1, wherein The obtaining of the predicted probability corresponding to the next occurrence of each type of word segment includes: Using the exponential smoothing method to predict the probability corresponding to the next occurrence of each type of word segment according to the smoothing parameter of each type of word segment and the probability corresponding to each occurrence of each type of word segment, and obtaining the predicted probability corresponding to the next occurrence of each type of word segment.

5. The question-answering method for an electric power intelligent question-answering system according to claim 1, characterized in that, The obtaining of the reconstruction necessity of the current Huffman tree includes: Obtaining the coding length sequence of each type of word segment and the predicted probability sequence of each type of word segment; ; In the formula, represents the necessity for reconstructing the current Huffman tree; H represents the number of categories of word segmentation; represents the index position of the predicted probability corresponding to the next occurrence of the i-th category of word segmentation in the predicted probability sequence of the i-th category of word segmentation; represents the index position of the coding length corresponding to the i-th category of word segmentation in the coding length sequence of the i-th category of word segmentation; || represents the absolute value symbol; exp() represents the exponential function with the natural constant as the base.

6. The question-answering method for an electric power intelligent question-answering system according to claim 5, characterized in that, The obtaining of the coding length sequence of each type of word segment and the predicted probability sequence of each type of word segment includes: Sorting the predicted probabilities corresponding to the next occurrence of each type of word segment in descending order to obtain the predicted probability sequence of the i-th type of word segment; sorting the coding lengths corresponding to each type of word segment in ascending order to obtain the coding length sequence of the i-th type of word segment.

7. The question-answering method for an intelligent power question-answering system according to claim 1 or 4, characterized in that The judging, based on the reconstruction necessity, whether it is necessary to reconstruct the current Huffman tree and then encode the next user interaction text includes: Obtain the next user interaction text. When the necessity for reconstructing the current Huffman tree is less than the reconstruction necessity T2, compress the next user interaction text according to the current Huffman tree. When the necessity for reconstructing the current Huffman tree is greater than or equal to the reconstruction necessity T2, reconstruct the current Huffman tree according to the predicted probability corresponding to the next occurrence of each type of word segmentation, and record it as the latest Huffman tree; use the latest Huffman tree to compress the next user interaction text; and so on, compress each user interaction text.

Citation Information

Patent Citations

  • Transportation energy data compression and transmission method based on Huffman coding

    CN115987296A

  • Intelligent ring information management method and system based on cloud edge collaboration

    CN116153453A