Knowledge base construction method for assisting digital employee question and answer

Through multimodal data acquisition and efficient processing technology, a knowledge base for assisting digital employees' Q&A is built, which solves the problems of inefficiency and inaccurate data processing in the existing technology, and realizes efficient and accurate knowledge base construction and dynamic updates, and supports intelligent Q&A based on semantic search.

CN120011575APending Publication Date: 2025-05-16BEIJING INTEGRAL TIMES TECH CO LTD

Patent Information

Application Number
CN202510078293.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art relies on manual sorting and manual input in the construction of knowledge bases, which are inefficient and error-prone, especially when processing large amounts of unstructured data, lack flexibility and adaptability.

Method used

A knowledge base construction method that assists digital employee Q&A is proposed. Through multimodal data collection, information entropy analysis, interactive verification, dynamic correction, semantic extraction and cluster optimization, high-quality knowledge entries are generated, and the knowledge base is dynamically updated under a distributed architecture to support digital employee Q&A based on semantic search.

Benefits of technology

It significantly improves the efficiency and accuracy of knowledge base construction, reduces the burden of redundant data processing, improves the quality of data input, solves the problems of data noise interference and inaccurate analysis results, and realizes the dynamic update and efficient question-and-answer functions of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011575A_ABST
    Figure CN120011575A_ABST
Patent Text Reader

Abstract

The invention relates to the field of knowledge base construction, and discloses a knowledge base construction method for assisting digital employee question answering, and the method comprises the following steps: collecting internal multi-modal data of an enterprise, including document data, voice data and metadata; carrying out information amount analysis on the collected data, and screening high-value data through an information entropy theory; performing interactive verification on the screened multi-modal data to analyze data consistency; dynamically correcting the verified data, and reducing data noise in combination with a prediction model and an observation model; and performing semantic extraction and clustering optimization on the corrected data to generate knowledge entries. Through the closed-loop process of multi-modal data collection, dynamic correction, semantic extraction, distributed updating and intelligent question and answer, the data processing efficiency, the knowledge base quality and the question and answer accuracy are remarkably improved, and the problems that in the prior art, unstructured data processing efficiency is low, updating is lagged, and dynamic optimization capacity is lacked are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge base construction, and in particular to a knowledge base construction method for assisting digital employees in question and answering. Background Art

[0002] As information technology continues to develop, more and more companies are introducing digital employees to perform daily office tasks. These virtual assistants rely on a powerful knowledge base to provide decision support and automate processing.

[0003] However, most of the knowledge base construction in the current market relies on manual sorting and manual input, which is not only inefficient but also prone to errors. With the rapid development of machine learning technology, especially the breakthrough of deep learning technology, the automation of knowledge acquisition and processing has become possible. This progress has significantly improved the efficiency and accuracy of knowledge base construction, providing a more reliable foundation for enterprise intelligent services; it is inefficient when processing large amounts of unstructured data, and lacks sufficient flexibility and adaptability for processing professional knowledge within a specific industry or organization. Summary of the invention

[0004] In view of the shortcomings of the existing technology, the present invention proposes a knowledge base construction method to assist digital employee question and answer, which can efficiently and accurately extract various types of unstructured and semi-structured data from the internal system of the enterprise, and generate a high-quality knowledge base through preprocessing, so as to better support the digital employee intelligent question and answer function based on a large model.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for constructing a knowledge base to assist digital employee question and answer, comprising the following steps: Collect multimodal data within the enterprise, including document data, voice data, and metadata; Analyze the amount of information on the collected data and filter out high-value data using information entropy theory; Conduct cross-validation on the screened multimodal data to analyze data consistency; Dynamically correct the verified data and combine the prediction model with the observation model to reduce data noise; Perform semantic extraction and clustering optimization on the corrected data to generate knowledge entries; Dynamically update the generated knowledge entries to the knowledge base under a distributed architecture; Leverage an updated knowledge base to support digital employee Q&A based on semantic search.

[0006] Preferably, the information entropy analysis measures the amount of information in the data by calculating the probability distribution of the data, wherein the screening condition is that data above a preset information entropy threshold is processed first.

[0007] Preferably, the interactive verification of the multimodal data comprises the following steps: Obtain time series data X from multimodal data sources = {x1, x2, ..., x n} and Y={y1,y2,...,y n}, respectively representing the data sampling values ​​from different sensors; the covariance matrix between time series data is calculated to measure the linear correlation between data, and its calculation formula is: Among them, μ X and μ Y represent the means of the time series X and Y respectively, represents the expectation operation; Compare the value of the covariance matrix with the preset consistency threshold, select the multimodal data pairs with covariance values ​​higher than the threshold, and mark them as data with high consistency; Data enhancement or elimination is performed on data pairs that do not meet the consistency threshold to ensure that the data entering the subsequent steps meets the consistency requirements.

[0008] Preferably, the dynamic correction includes using a dynamic prediction and observation correction model to calculate the optimal state estimate of the data, and adjusting the error range through a gain parameter to correct the uncertainty in the multimodal data.

[0009] Preferably, the dynamically corrected data must satisfy preset physical constraints, including continuity or conservation constraints; and the data that does not satisfy the constraints is marked or resampled.

[0010] Preferably, the semantic extraction adopts a variational inference model to extract the latent semantics of unstructured data by constructing an approximate posterior distribution and a latent semantic space; the semantic clustering utilizes a Gaussian mixture model to perform clustering optimization on the extracted semantic data.

[0011] Preferably, the dynamic update of the knowledge base is based on a consistency protocol, data synchronization between multiple nodes is achieved through distributed log replication, and content that needs to be updated first is screened according to the change in information entropy of knowledge items.

[0012] Preferably, the knowledge base supports digital employee Q&A through a semantic search mechanism, wherein the semantic search is sorted according to the semantic relevance between the input query content and the knowledge items, and an attention mechanism is used to improve the contextual adaptability of the Q&A results.

[0013] Preferably, the knowledge base dynamically adjusts the content of knowledge items based on user feedback, wherein the feedback information includes the correctness score of the question and answer results, and optimizes the semantic adaptability of the knowledge items through a training model.

[0014] The present invention also provides a knowledge base construction system for assisting digital employee question and answer, comprising: Data collection module, used to collect multimodal data within the enterprise; Information volume analysis module, used to analyze the information entropy of the collected data and filter out high-value data; Data interaction verification module, used to perform consistency analysis on multimodal data; Dynamic correction module, used for dynamic prediction and correction of multimodal data; Semantic extraction and clustering module, used to extract semantics and cluster data to optimize and generate knowledge items; Distributed knowledge base module, used to store the generated knowledge entries and implement dynamic updates under the consistency protocol; The Q&A interaction module is used to support digital employee Q&A by leveraging the knowledge base through semantic search and receiving user feedback to optimize the knowledge base content.

[0015] The present invention provides a method for constructing a knowledge base to assist digital employees in answering questions. It has the following beneficial effects: 1. The present invention adopts a technical solution based on multimodal data collection and information entropy analysis. Through multimodal collection and high-value data screening of internal enterprise document data, voice data and metadata, it achieves the technical effect of effectively reducing the burden of redundant data processing and improving the quality of data input. Compared with the technical solution in the prior art that relies on manual sorting and direct use of original data, it solves the problems of more data redundancy, low processing efficiency and large data noise interference.

[0016] 2. The present invention realizes the interactive verification of multimodal data by introducing the method of covariance matrix and normalized correlation coefficient, and performs consistency verification on the collected data in combination with the time alignment algorithm, thereby achieving the technical effect of significantly improving the reliability of multimodal data and the accuracy of time synchronization. Compared with the technical solution of the prior art that lacks consistency verification and time alignment mechanism for multimodal data, the problem of inaccurate analysis results caused by data deviation is solved.

[0017] 3. The present invention adopts Kalman filter dynamic correction technology to perform real-time state prediction and correction on multimodal data, and combines physical consistency verification to reduce noise interference and ensure data continuity, achieving the technical effect of improving data accuracy and time series consistency. Compared with the technical solution in the prior art that does not dynamically adjust unstructured data, it solves the problem of random noise and uncertainty in the data.

[0018] 4. The present invention extracts the potential semantic features of unstructured data through variational inference technology, and uses Gaussian mixture model to cluster and optimize the extracted features, thus achieving the technical effect of generating high-quality knowledge items and reducing data redundancy. Compared with the technical solution of the prior art that only relies on keyword matching for semantic processing, it solves the problems of inaccurate semantic understanding and poor organization of knowledge items.

[0019] 5. The present invention realizes the dynamic update of the knowledge base based on the Raft distributed consistency protocol, and achieves the technical effect of improving the real-time update capability of the knowledge base and giving priority to high-value items through entropy value change analysis and multi-dimensional priority evaluation mechanism. Compared with the technical solution in the prior art where the knowledge base update cycle is long and difficult to dynamically adjust, the problem of slow response of the knowledge base and untimely content update is solved.

[0020] 6. The present invention constructs an efficient interaction mechanism between the knowledge base and the question-answering system through semantic search technology, and introduces a dynamic adjustment method for optimizing the question-answering results through user feedback, thereby achieving the technical effect of improving the retrieval accuracy and user experience of the question-answering system. Compared with the technical solutions of one-way retrieval and fixed answers in the prior art, the problem of the lack of dynamic optimization and insufficient semantic understanding ability of the question-answering system is solved.

[0021] 7. The present invention realizes a closed-loop process from data collection, screening, verification to dynamic correction, semantic extraction, knowledge base update and intelligent question and answer through a full-process technical solution, achieving the technical effect of improving the overall intelligence level of the system and adapting to the needs of complex enterprise environments. Compared with the technical solutions in the prior art that separate data processing and knowledge base construction and have a low degree of intelligence, it solves the problem that the enterprise intelligent question and answer system has insufficient comprehensive requirements for data processing efficiency and knowledge base accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION

[0023] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] Please refer to the attached Figure 1 , an embodiment of the present invention provides a method for constructing a knowledge base to assist digital employee question and answer, comprising the following steps: S1. Collect multimodal data within the enterprise, including document data, voice data, and metadata; First, it is necessary to obtain multimodal data from within the enterprise as the basic data source for building the knowledge base. The multimodal data acquisition of the method of the present invention is a key step in realizing the efficient construction of the knowledge base. The data includes document data, voice data, and metadata, covering the diverse information forms widely existing in the enterprise. It should be noted that there may be large format differences, time synchronization problems, and noise interference between multimodal data. Therefore, during the acquisition process, it is necessary to combine appropriate sensor technology and data preprocessing mechanisms to ensure the integrity and consistency of the data.

[0025] In the present invention, the collected multimodal data provides an important basis for the subsequent steps, wherein the data source and processing flow are as follows: In this embodiment, the image sensor deployed within the enterprise is first used to collect document data, including scanned contracts, PDF files of rules and regulations, etc. These documents usually exist in an unstructured form. The collected image data is stored in the form of a pixel matrix, and each pixel can be represented as a three-channel RGB value, which can be described in mathematical form as: I(x,y)=[R(x,y),G(x,y),B(x,y)] Among them, I(x,y) represents the pixel point at the position (x,y) of the image, R(x,y), G(x,y), and B(x,y) represent the intensity values ​​of the red, green, and blue channels of the pixel point, respectively, and the value range is usually between 0 and 255. As an option, the embedded text content can be directly extracted through PDF parsing technology; if the embedded text cannot be directly extracted, the optical character recognition (OCR) technology is used to recognize the text information in the document. Optical character recognition technology usually extracts the character outline through edge detection algorithms, and combines deep learning models to achieve character classification and restoration.

[0026] In this embodiment, voice data such as meeting records and voice commands of the enterprise are also collected through voice sensors. The collected voice data is stored in the form of discrete time series, which is expressed as: A(t)={a1,a2,...,a n} Among them, A(t) represents the amplitude sequence of the audio signal at time t, a irepresents the amplitude of the i-th sampling point, and n is the total number of sampling points of the audio signal. The speech data is first subjected to noise reduction and standardization processing by the preprocessing module. The noise reduction process can use a bandpass filtering method to filter out frequency components below 300Hz and above 3400Hz to enhance the clarity of the speech signal. Then, the audio signal is converted into text through automatic speech recognition (ASR) technology. The implementation steps of ASR technology include feature extraction of audio signals, acoustic model training, and speech-to-text mapping.

[0027] The feature extraction of audio signals usually uses Mel-frequency cepstral coefficients (MFCC), and the calculation formula is: Among them, MFCC k represents the kth cepstral coefficient, X i Represents the spectrum energy value, N is the total number of channels for spectrum decomposition, k is the index of the cepstral coefficient, and the value range is 1 to N.

[0028] In one possible implementation, the auxiliary sensor is used to record metadata related to data collection, including timestamp, collection device status, and environmental variables. Specifically, the timestamp data is recorded in the ISO8601 standard format, expressed as: T capture =YYYY-MM-DD HH:MM:SS Among them, T capture The data collection timestamp is YYYY-MM-DD, which represents the date part, and HH:MM:SS, which represents the time part. The metadata record provides necessary context information for subsequent data processing. For example, in the multimodal data synchronization link, the timestamp can be used to align the time series of data in different modalities, and the device status information can be used to monitor the reliability of the collection device.

[0029] As an option, a synchronization calibration module can be added to the multimodal data acquisition to achieve time alignment of the multimodal data by cross-referencing the timestamp data. For example, the voice data and image data can be aligned by comparing the acquisition timestamp T audio and T image Synchronization is performed to ensure that the two types of data have the same time reference point. In another possible implementation, a multi-channel sensor system can also be introduced to collect multi-modal data at the same time to ensure time synchronization from the hardware level.

[0030] It should be noted that the collected multimodal data needs to be classified and stored by type and source. For example, image data is stored in a dedicated document database, voice data is stored in an audio database, and metadata is stored in a relational database. The overall data storage structure can be expressed as: D={D image,D audio ,D metadata} Among them, D image ,D audio ,D metadata Represents the storage collection of image data, voice data and metadata respectively.

[0031] In some embodiments, a compression algorithm can be introduced for image data, for example, a lossless compression format (such as PNG) can be used to reduce storage requirements while ensuring image quality; a silence detection algorithm can be introduced for voice data to eliminate invalid parts in the audio to reduce storage and processing burdens.

[0032] It is understandable that data collection is a basic step in realizing knowledge base construction, and the quality of the collection process directly determines the effect of subsequent information entropy analysis and consistency verification. The present invention provides reliable data support for the subsequent dynamic knowledge base construction through multimodal data collection technology.

[0033] It should be further explained that after the collection is completed, the integrity and correctness of the data can be automatically checked through the verification algorithm. The detection includes checking the validity of the timestamp, the continuity of the sampling points, and the abnormal status of the collection device. All abnormal data will be marked and stored in the pending list to avoid affecting the accuracy of subsequent processing steps.

[0034] S2. Analyze the amount of information on the collected data and select high-value data through information entropy theory; First, it is necessary to analyze the information content of the multimodal data collected in step S1 to screen out high-value data and ensure the accuracy and efficiency of subsequent processing. This step quantifies the information entropy of the data, and prioritizes the data with high information content as valid data and enters the subsequent steps, while low information content or redundant data is marked as secondary data or temporarily stored to reduce resource waste and optimize the performance of subsequent processing links. It should be noted that the calculation of information entropy is a quantitative indicator based on the distribution characteristics of data, which aims to evaluate the uncertainty and information content contained in the data.

[0035] In this embodiment, the collected multimodal data are analyzed one by one by using the information entropy theory. Information entropy is used to measure the uncertainty of data, and its calculation formula is: Where H(X) represents the information entropy of the random variable X, in bits; x i represents the i-th possible value of data X; P(x i ) means X takes the value x i The probability of ; n represents the total number of possible values ​​of the data. i) is calculated based on the normalized frequency distribution as follows: Among them, count(x i ) represents data x i As an option, for document data, the data feature can be word frequency distribution, and the vocabulary set is defined as W = {w1,w2,...,w m}, where m represents the total number of words in the document. The probability of each word P(w i ) is calculated as: Among them, count(w i ) represents the word w i The frequency of occurrence in the document. Then, the total information entropy H(document) of the document is calculated according to the above information entropy formula.

[0036] For speech data, its characteristics can be expressed as the spectral distribution of the audio signal, and the spectral features are extracted using short-time Fourier transform. The frequency component set is defined as F = {f1, f2, ..., f N}, where N represents the total number of frequency components. The frequency component f i The probability P(f i ) is expressed as: Among them, |F(f i )| 2 Represents the frequency component f i The total information entropy H(audio) of the speech signal is calculated according to the above formula.

[0037] In one possible implementation, the joint distribution of multimodal data is included in the analysis scope, and the joint information entropy H(X,Y) is calculated, and the formula is: Among them, H(X,Y) represents the joint information entropy of multimodal data X and Y; P(x i ,y j ) means that variables X and Y take values ​​of x at the same time i and j The joint probability is calculated as: Among them, count(x i ,y j ) represents x i and j The co-occurrence frequency of .

[0038] As an option, a predefined threshold H can be set for the entropy value of each piece of data. threshold , the filtering rules are: H(X)>H threshold Among them, H threshold It represents the minimum value requirement of information entropy, and only the data that meets this condition is retained to proceed to the next step. It can be understood that H threshold The choice can be adjusted according to the actual scenario. For example, when the data quality requirement is high, a higher entropy threshold can be set to filter out better quality data.

[0039] In some embodiments, statistical indicators such as standard deviation and kurtosis may be combined to further analyze data characteristics. The specific formula is as follows: Among them, σ X represents the standard deviation of data X; κ X Indicates the kurtosis of data X; μ X Represents the mean of data X, and the calculation formula is: By combining these indicators, the accuracy of data screening can be further improved.

[0040] It should be noted that the data marked as secondary during the screening process will be temporarily stored in a specific storage module and provide a reference for subsequent anomaly detection or supplementary analysis. It can be understood that the information volume analysis provides high-quality data input for the subsequent steps of the present invention, while significantly reducing the interference of redundant data. Through the calculation of information entropy and joint information entropy, the importance of data can be evaluated in a quantitative manner, which has good adaptability and scalability.

[0041] S3. Interactively verify the screened multimodal data to analyze data consistency; First, in order to ensure the consistency and reliability of multimodal data, it is necessary to analyze the characteristics of different data sources and verify their relevance in time and content. Multimodal data (such as image data, voice data, and metadata) may cause data distortion or deviation due to differences in acquisition equipment or changes in the acquisition environment. This step quantitatively analyzes the consistency of the data through the covariance matrix, screens out valid data with high relevance, removes or corrects abnormal data, and provides high-quality input for subsequent dynamic correction.

[0042] In this embodiment, the covariance matrix is ​​used to calculate the linear correlation between two random variables to measure the relationship between different modal data. The calculation formula of the covariance matrix is: Where: C(X,Y) represents the covariance of random variables X and Y; represents mathematical expectation; μ X and μ Y are the means of random variables X and Y, respectively, and the calculation formula is: Where: x i and i are the values ​​of X and Y at the i-th sampling point respectively; n represents the total number of data samples.

[0043] When the covariance value is positive, it means that the variables are positively correlated; when the covariance value is negative, it means that the variables are negatively correlated; the closer the covariance value is to zero, the weaker the correlation between the variables.

[0044] As an option, the normalized correlation coefficient can be combined to calculate the standardized correlation of multimodal data to eliminate the dimensional effects of different modal eigenvalues. The calculation formula of the normalized correlation coefficient is: Where: ρ(X,Y) is the correlation coefficient between variables X and Y, and its value range is [-1, 1]; σ X and σ Y are the standard deviations of variables X and Y, respectively, and the calculation formula is: It should be noted that when the normalized correlation coefficient value is 1, it means that the variables are completely positively correlated, when it is -1, it means that the variables are completely negatively correlated, and when it is 0, it means that the variables have no linear correlation.

[0045] Specifically, in this embodiment, for the verification of image data and speech data, their key features are extracted to construct the covariance matrix. The feature X of the image data is represented as the local mean of the pixel intensity matrix, and the feature Y of the speech data is represented as the short-time spectrum energy value of the audio signal. Assume that the mean of the image feature is x i , the audio feature value is y i , then the covariance calculation formula between the two is: It should be noted that the covariance value is higher than the set consistency threshold C threshold Data pairs with covariance values ​​lower than the threshold are considered to have high consistency and can enter the subsequent processing steps; data pairs with covariance values ​​lower than the threshold are marked as abnormal data.

[0046] In one possible implementation, the multimodal data can be synchronized by a time alignment algorithm to ensure the consistency of the time dimension. For example, comparing the timestamp T of the image data image and the timestamp T of the voice data audio, the time offset ΔT between the two is defined as: ΔT=|T image -T audio | When ΔT is less than the preset synchronization error threshold ΔT threshold When , the data pair is considered to meet the time consistency condition and can enter the covariance calculation stage.

[0047] For example, this step can be applied in the scenario of enterprise contract approval. The image data can come from the document image of the contract scan, and the voice data can come from the audio features of the relevant meeting minutes. By calculating the covariance of the key content extracted from the contract document and the keyword energy distribution in the meeting minutes, the consistency of the contract content and the meeting minutes is ensured.

[0048] As an option, the covariance results can be dynamically weighted to adapt to the consistency verification requirements of different scenarios. For example, in high-precision application scenarios, a higher consistency threshold C threshold The data can be strictly screened, while in the scenario of rapid response, the threshold can be appropriately relaxed.

[0049] It should be further explained that the data marked as abnormal will not be discarded directly, but stored in the abnormal data storage module for subsequent manual review or automatic repair. In the repair stage, the abnormal data can be repaired by using an interpolation algorithm, which can fit the missing values ​​based on the least squares method.

[0050] It can be understood that the multimodal data interactive verification in this step provides high-quality input data for subsequent dynamic correction, significantly improving the reliability and effectiveness of subsequent processing. The use of the covariance matrix provides a rigorous mathematical basis for data consistency verification, giving the verification process good theoretical support and operability.

[0051] S4, dynamically correct the verified data and combine the prediction model and observation model to reduce data noise; After the interactive verification of multimodal data, in order to further eliminate data noise, compensate for possible deviations in the acquisition process and improve the overall quality of the data, it is necessary to dynamically correct the verified data. The core of dynamic correction is to predict and correct the data through Kalman filtering technology, and adjust the state estimate value in combination with real-time observation data. It should be noted that dynamic correction of data can not only reduce noise interference, but also improve the continuity of data in the time dimension, laying a high-quality foundation for subsequent semantic extraction and knowledge item generation.

[0052] In this embodiment, a Kalman filter model is used to implement dynamic correction, and the input data is dynamically optimized through two stages: state prediction and observation correction. The basic state model of Kalman filtering can be expressed as: x k =Ax k-1 +Bu k +w k z k =Hx k +v k Where: x k represents the state variables of the system at time k; A represents the state transfer matrix, which describes the state change of the system from time k-1 to time k; B represents the control matrix, which describes the control input u k Impact on system status; k Represents process noise, which follows a zero-mean normal distribution Q is the process noise covariance matrix; z k represents the observed value of the system at time k; H represents the observation matrix, which describes the mapping relationship between state variables and observed variables; v k represents observation noise, which follows a zero-mean normal distribution R is the observation noise covariance matrix.

[0053] In the state prediction stage, the system model is used to predict the current state value and error covariance. The specific formula is: P k|k-1 =AP k-1|k-1 A T +Q in: represents the predicted state at time k; P k|k-1 represents the forecast error covariance matrix.

[0054] In the observation correction stage, according to the observed value z k Correct the predicted state, and the correction formula is: K k =P k|k-1 H T (HP k|k-1 H T +R) -1 P k|k =(IK k H)P k|k-1 Where: K k is the Kalman gain, which represents the contribution of observation data to state correction; is the corrected state estimate; P k|k is the corrected error covariance matrix; I is the identity matrix.

[0055] Specifically, in the dynamic correction process, the pixel mean of the image data can be used as the state variable x k , the spectral energy value of the speech data is taken as the observation value z k Through Kalman filtering technology, combined with the time series characteristics of image and voice data, noise interference is reduced and the continuity and accuracy of data are improved.

[0056] As an option, physical consistency verification can be introduced in the dynamic correction stage to perform physical constraint checks on the corrected data. For example, for data with continuity characteristics, the following constraints can be used: in: It represents the divergence of the data field and reflects the continuity of the data in space; x, y, t are the spatial and temporal coordinates respectively.

[0057] It should be noted that data that does not meet the physical consistency constraint will be marked as abnormal and enter the abnormality handling module for resampling or correction. In a possible implementation, a trend analysis function can be added to the dynamically corrected data to predict future states using historical data. For example, the corrected data x can be analyzed using a time series model. k Long-term trend modeling is performed to assist the subsequent semantic extraction step. The implementation of trend analysis can be based on the autoregressive model (AR), and its mathematical model is: x k =φ1x k-1 +φ2x k-2 +...+φ p x k-p +∈ k Where: φ1, φ2, ..., φ p represents the autoregressive coefficient; p is the model order; ∈ k represents the white noise term.

[0058] It should be further explained that the dynamic correction in the present invention is not only applicable to image and voice data, but can also be extended to other types of multimodal data, such as physical measurement data from sensors or log data from enterprise management systems. It can be understood that the introduction of dynamic correction technology effectively reduces the random noise and uncertainty of the data, providing more reliable basic data for the subsequent generation of knowledge items. The iterative optimization mechanism of Kalman filtering provides a solid theoretical support for dynamic correction of data, and has good adaptability and scalability.

[0059] S5. Perform semantic extraction and clustering optimization on the corrected data to generate knowledge items; After dynamic correction, in order to further transform the multimodal data into high-quality knowledge items that are suitable for knowledge base construction, the data needs to be semantically extracted and clustered. This step extracts the latent semantic features of the data by combining variational inference technology, and uses the Gaussian mixture model (GMM) to semantically cluster the extracted features. It should be noted that the purpose of semantic extraction is to represent the semantic information of unstructured data as a low-dimensional vector space, and the purpose of clustering optimization is to classify similar semantic items into a group to improve the organization and retrieval efficiency of knowledge base items.

[0060] In this embodiment, semantic extraction is performed on the dynamically corrected multimodal data using variational inference technology. Variational inference is used to estimate the potential semantic distribution of data, and its goal is to maximize the following variational lower bound: in: represents the variational lower bound; q φ (z|x) represents the approximate posterior distribution of the latent semantic variable z; p θ (x|z) represents the probability of generating data x given the latent semantic variable z; p(z) represents the prior distribution of the latent semantic variable, which is usually a standard normal distribution. D KL Represents the KL divergence, which is used to measure the gap between the true posterior distribution and the approximate posterior distribution.

[0061] It should be noted that by optimizing the above objective function, the semantic representation of the data can be obtained in the low-dimensional latent space.

[0062] As an option, deep learning models can be used to φ (z|x) and p θ (x|z) is parameterized. For example, a variational autoencoder (VAE) structure is used to generate an approximate posterior distribution of latent semantic variables using the encoder: Where: μ φ (x) and Σ φ (x) are the mean and covariance of the latent semantic distribution, usually output by a neural network.

[0063] Then, the decoder is used to reconstruct the distribution of the original data x for the latent semantic variable z: Where: f θ (z) represents the output activity of the decoder; σ 2 represents the variance of the observation noise.

[0064] After the semantic extraction is completed, the Gaussian mixture model is further used in this embodiment to perform cluster optimization on the extracted semantic representation. The goal of the Gaussian mixture model is to divide the semantic representation into K categories, and its probability density function is expressed as: Where: K represents the number of clustering categories; π k represents the weight of the kth Gaussian component, satisfying and∑ k denote the mean and covariance matrix of the kth Gaussian component respectively.

[0065] Specifically, the expectation maximization (EM) algorithm is used to optimise the parameters π of the Gaussian mixture model. k , μ k and∑ k Make an estimate. The iterative process of the EM algorithm includes: Step E: Calculate the posterior probability that the data point belongs to the kth category; Step M: Update the parameters of the Gaussian mixture model based on the posterior probability.

[0066] As an option, a semantic similarity optimization mechanism can be introduced to the clustering results, measuring the distance between the data point and the cluster center by cosine similarity: Among them: x and y represent the semantic vectors of two data points; x·y represents the inner product of two vectors; ∥x∥ and ∥y∥ represent the modulus lengths of the two vectors respectively.

[0067] It should be noted that the compactness and interpretability of clustering results can be improved through semantic similarity optimization.

[0068] In a possible implementation, structured knowledge items can be generated based on the clustering results. Specifically, data items belonging to the same category are organized into a knowledge module, and the meta-information of the knowledge module includes the semantic representation of the cluster center, the category label, and the data item list. For example, in the enterprise contract management scenario, items related to financial reimbursement can be clustered into one module, and items related to project approval can be clustered into another module.

[0069] It is understandable that, through semantic extraction and clustering optimization, the present invention can effectively reduce data redundancy and enhance the organizational structure of knowledge base entries. This process uses variational inference and Gaussian mixture model, combined with semantic similarity optimization mechanism, to provide a high-quality semantic foundation for the construction of the knowledge base, while significantly improving the performance of subsequent retrieval and question-answering systems. The latent space representation of variational inference has good versatility and scalability, and the parameter optimization method of Gaussian mixture model ensures the accuracy and stability of clustering results.

[0070] S6, dynamically updating the generated knowledge entries to the knowledge base under the distributed architecture; In order to adapt to the dynamic changes in the enterprise knowledge base, the present invention designs a knowledge base dynamic update method based on a distributed architecture. Through the distributed consistency protocol, the synchronous update between multi-node knowledge bases is realized, and at the same time, combined with dynamic entropy analysis, high-value knowledge items are updated first. It should be noted that this step can not only ensure the consistency and real-time performance of the knowledge base content, but also significantly improve the adaptability of the knowledge base in high-concurrency scenarios.

[0071] In this embodiment, the Raft protocol is used to implement dynamic updates of the distributed knowledge base. The Raft protocol is a consensus algorithm used to achieve consistency in a distributed system. Its basic idea is to ensure that there is a unique master node in the system through an election mechanism, and the master node is responsible for processing log replication and data synchronization.

[0072] Specifically, the Raft protocol includes the following core stages: Election phase: When the system starts or the master node fails, other nodes in the cluster vote to elect a new master node. It should be noted that each node will initiate a vote after a random timeout, and the node that first obtains the majority of votes will be elected as the new master node.

[0073] Log replication phase: The master node broadcasts the updated knowledge entries to the slave nodes in the form of logs. After receiving the logs, each slave node stores them in the local log.

[0074] Consistency commit phase: When a majority of nodes confirm a log entry, the master node marks the entry as committed and applies it to the local knowledge base.

[0075] The replication process of log entries can be expressed as: Log nodei =Log leader Among them: Log nodei Represents the log copy of the i-th slave node; Log leader Represents the log of the master node.

[0076] It should be noted that the correctness of log replication is ensured by the heartbeat mechanism between nodes. The master node sends heartbeat signals to the slave nodes regularly. If the slave node does not receive the heartbeat signal within the specified time, it considers the master node invalid and re-triggers the election.

[0077] After the log copy is completed, the present embodiment also combines the dynamic entropy value analysis mechanism to give priority to updating high-value knowledge items. The information entropy change of the knowledge item is defined as: ΔH=|H new -H old| Where: ΔH represents the change in information entropy of knowledge items; H new represents the information entropy after the entry is updated; H old Represents the information entropy of the entry before it is updated.

[0078] Specifically, knowledge items with larger entropy changes are updated first, while items with smaller entropy changes are delayed or remain unchanged. This strategy effectively improves the dynamic adaptability of the knowledge base and ensures timely updating of key knowledge items.

[0079] As an option, the priority of updating knowledge items can be evaluated in multiple dimensions. For example, based on the dynamic entropy analysis, combined with the access frequency F of the knowledge items, access and usage time T usage Calculate the comprehensive priority P priority : P priority =w1·ΔH+w2·F access +w3·T usage Among them: w1, w2, w3 are weight parameters, which can be adjusted according to actual needs; P priority Indicates the comprehensive priority. The larger the value, the higher the priority.

[0080] It should be noted that, through the above-mentioned multi-dimensional evaluation mechanism, the update order of knowledge items can be determined more accurately, thereby improving the overall performance of the knowledge base.

[0081] In a possible implementation, the dynamic update of knowledge items is also combined with distributed storage technology. Specifically, the storage nodes of the knowledge base are distributed on multiple physical servers in the form of shards, and each shard stores knowledge items of a specific category or a specific field. For example, in an enterprise financial management scenario, knowledge items related to financial reimbursement can be stored on shard S1, while knowledge items related to the approval process can be stored on shard S2.

[0082] The access mechanism of distributed storage can be expressed as: D={S1,S2,...,S n} Where: D represents the complete knowledge base; S i Represents the i-th shard.

[0083] The advantage of distributed storage is that it significantly improves the query speed and scalability of the knowledge base, especially when dealing with large-scale concurrent access.

[0084] It is understandable that the log replication mechanism and dynamic entropy value analysis method implemented by the Raft protocol effectively improve the consistency and update efficiency of the knowledge base content. At the same time, the multi-dimensional update priority evaluation and distributed storage architecture provide a good foundation for the dynamic expansion of the knowledge base. The implementation of this step ensures the performance stability and efficiency of the knowledge base in high-concurrency access and dynamic update scenarios.

[0085] S7. Use an updated knowledge base to support digital employee Q&A based on semantic search.

[0086] After the dynamic update of the knowledge base is completed, in order to realize the intelligent question-answering function based on semantic search, it is necessary to establish an interaction mechanism between the knowledge base and the question-answering system. This step uses semantic search and context optimization technology to enable the question-answering system to quickly and accurately retrieve entries related to user questions from the knowledge base. At the same time, the user feedback mechanism is introduced to dynamically optimize the question-answering results, further improving the accuracy and adaptability of the question-answering system. It should be noted that the interaction between the knowledge base and the question-answering system is not only a one-way data query process, but also involves feedback-based knowledge base optimization and dynamic supplementation.

[0087] In this embodiment, the interaction between the knowledge base and the question-answering system implements semantic search based on the attention mechanism. The attention mechanism generates a priority ranking of search results by calculating the correlation between user questions and knowledge items. Specifically, the core calculation formula of the attention mechanism is: Where: Q represents the query vector (the semantic representation of the user's question); K represents the key vector (the semantic representation of the knowledge item); V represents the value vector (the content of the knowledge item); d k represents the dimension of the key vector; softmax is a normalization function used to map the relevance score to a probability distribution.

[0088] It should be noted that by calculating the dot product similarity of Q and K and normalizing the results, the attention mechanism can effectively screen out the knowledge items that are most relevant to the user's questions.

[0089] As an option, the user question can be context-enhanced in the semantic search stage. Specifically, the user question is combined with the previous question and answer to construct the question context C, which is expressed by the following formula: C=Concat(Q current ,Q previous ,A previous ) Where: Q current Indicates the current question; Q previous Indicates the previous question; A previousIndicates the answer to the previous question; Concat represents a vector concatenation operation.

[0090] By constructing context-enhanced question representations, question-answering systems can better understand user intent and improve the accuracy of semantic search.

[0091] In this embodiment, the semantic representation of the user question is generated using a deep learning model, such as BERT (Bidirectional Encoder Representation) or Transformer model. These models are trained to generate a high-dimensional vector representation Q of the user question. The specific formula is: Q=f model (user question ) Where: f model Represents a deep learning model; user question The text content representing the user's question.

[0092] The semantic representation K and content V of knowledge items are generated and stored in a similar way during the knowledge base construction phase.

[0093] In a possible implementation, the priority of the search results can be combined with the semantic similarity and the access frequency of the entry F access Optimize. The comprehensive scoring formula is: S=α·Sim(Q,K)+β·log(1+F access ) Where: S represents the comprehensive score of the search results; Sim(Q,K) represents the semantic similarity between the query vector and the key vector; F access Indicates the access frequency of the entry; α and β are weight parameters, which can be adjusted according to the scenario.

[0094] For example, in an internal Q&A scenario, the questions asked by users may involve specific rules for the contract approval process. The Q&A system retrieves relevant entries from the knowledge base through semantic search and optimizes the search results based on the context of the question (such as the specific contract type the user asked about previously). The final answer generated not only includes the content of the relevant knowledge entry, but may also combine the entry meta-information (such as the most recent update date) to provide users with a more comprehensive answer.

[0095] In order to further optimize the question-answering results, a user feedback mechanism is introduced in this embodiment. The user's satisfaction score for the answer is R feedback Defined as: Where: R feedback represents the average satisfaction score; r i represents the rating of the i-th user; m represents the number of users participating in the rating.

[0096] Based on user ratings, the knowledge base dynamically adjusts the weight and content of entries. For example, for answers with low ratings, the quality of answers can be improved by retraining the semantic model or updating the content of knowledge entries.

[0097] It should be further explained that this step not only realizes the efficient interaction between the knowledge base and the question-answering system, but also realizes the dynamic optimization of the knowledge base through the feedback mechanism. It is understandable that this closed-loop optimization mechanism can significantly improve the accuracy and user experience of the question-answering system, while ensuring the timeliness and relevance of the knowledge base content. The semantic search technology combined with the application of deep learning models provides solid theoretical and technical support for the implementation of the question-answering system, with good adaptability and scalability.

[0098] Please see attached Figure 2 The present invention also provides a knowledge base construction system for assisting digital employee question and answer, comprising: Data collection module, used to collect multimodal data within the enterprise, including document data, voice data and metadata; Document data includes contract documents, rules and regulations, financial reimbursement forms, etc., which are acquired using image sensors or text parsing technology; voice data includes meeting minutes, voice commands, etc., which are collected by voice sensors and converted into text using automatic speech recognition (ASR) technology; Metadata records the data’s timestamp, acquisition device status, and environmental variables, providing contextual support for data processing; This module organizes multimodal data in a unified time serialization manner to provide consistent input for subsequent processing.

[0099] The information volume analysis module is used to perform information entropy analysis on the collected data, filter high-value data and eliminate redundant or low-information data; Based on the information entropy theory, the probability distribution of data is quantitatively analyzed to screen high-value data whose information entropy is greater than the set threshold; the word frequency statistics method is used for document data, and the spectrum feature extraction method is used for speech data to generate probability distribution; It can be extended to joint entropy analysis, combining the joint probability distribution of multimodal data to further optimize the screening effect; Through the screening of this module, the possibility of low-value data entering the subsequent processing flow is significantly reduced, thereby improving the overall processing efficiency of the system.

[0100] Data interaction verification module, used to perform consistency analysis on the filtered multimodal data to ensure time synchronization and content relevance between data sources; The covariance matrix was used to calculate the linear correlation of multimodal data, and the normalized correlation coefficient was used to quantify the consistency; A time alignment algorithm is introduced to ensure synchronization in the time dimension by comparing the differences in timestamps of multimodal data; Data that fails verification is marked or sent to the exception processing module to eliminate data points that do not meet consistency requirements; This module ensures the logical and temporal consistency of multimodal data, providing high-quality input for subsequent dynamic correction.

[0101] Dynamic correction module, used to dynamically predict and correct multimodal data, reduce noise interference and improve data accuracy and continuity; State prediction and observation correction, including state transfer and noise compensation, are achieved through Kalman filtering technology; Introduce physical consistency verification to check the continuity or conservation constraints of the corrected data; Dynamically adjust abnormal data points by interpolation correction or resampling; This module ensures the high reliability of multimodal data through a dynamic optimization mechanism and provides accurate input data for the semantic extraction stage.

[0102] The semantic extraction and clustering module is used to perform semantic extraction and clustering optimization on the corrected data to generate high-quality knowledge items that are suitable for the knowledge base; Extract the latent semantic features of data based on variational inference technology and represent unstructured data as low-dimensional vectors; The Gaussian mixture model (GMM) is used to cluster semantic vectors, generate knowledge items and organize them into categorized modules; the semantic similarity optimization mechanism can be combined to improve the compactness of clustering and the interpretability of items; This module significantly improves the quality of knowledge entries through deep semantic analysis and provides structured input for knowledge base construction.

[0103] The distributed knowledge base module is used to store the generated knowledge entries and realize dynamic updates through a distributed consistency protocol.

[0104] Log replication and multi-node consistency are implemented based on the Raft protocol to ensure the reliability and real-time performance of the knowledge base in a distributed environment; Introduce dynamic entropy value analysis to give priority to updating knowledge items with large changes in information entropy; Supports distributed storage and sharding mechanisms, dividing knowledge items into different storage nodes by category or field to optimize query efficiency and scalability; This module ensures the dynamic adaptability and query performance of the knowledge base through an efficient update mechanism and distributed storage architecture.

[0105] A question-and-answer interaction module, which is used to support digital employee question-and-answer sessions using a knowledge base through semantic search and receive user feedback to optimize knowledge base content; The attention mechanism is used to implement semantic search, and priority ranking is generated based on the similarity between the query vector and the knowledge item vector; Supports context enhancement, combining user questions with previous questions and answers to construct enhanced semantic representation; Introduce a user feedback mechanism to dynamically adjust the weight or content of knowledge items based on answer satisfaction; The retrieval strategy of the question-answering system can be optimized by combining access frequency and dynamic priority evaluation mechanism; Through semantic search and feedback optimization, this module realizes the accurate answer and dynamic adjustment functions of the question-and-answer system, significantly improving the user experience.

[0106] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a knowledge base to assist digital employee question and answer, characterized in that: The following steps are involved: Collect multimodal data within the enterprise, including document data, voice data, and metadata; Analyze the amount of information on the collected data and filter out high-value data using information entropy theory; Conduct cross-validation on the screened multimodal data to analyze data consistency; Dynamically correct the verified data and combine the prediction model with the observation model to reduce data noise; Perform semantic extraction and clustering optimization on the corrected data to generate knowledge entries; Dynamically update the generated knowledge entries to the knowledge base under a distributed architecture; Leverage an updated knowledge base to support digital employee Q&A based on semantic search.

2. A method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The information entropy analysis measures the amount of information in the data by calculating the probability distribution of the data, wherein the screening condition is that data above a preset information entropy threshold is processed first.

3. A method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The interactive verification of the multimodal data comprises the following steps: Obtain time series data X from multimodal data sources = {x1, x2, ..., x n } and Y={y1,y2,...,y n }, respectively representing the data sampling values ​​from different sensors; the covariance matrix between time series data is calculated to measure the linear correlation between data, and its calculation formula is: Among them, μ X and μ Y They represent the means of time series X and Y respectively, and E represents the expected operation; Compare the value of the covariance matrix with the preset consistency threshold, select the multimodal data pairs with covariance values ​​higher than the threshold, and mark them as data with high consistency; Data enhancement or elimination is performed on data pairs that do not meet the consistency threshold to ensure that the data entering the subsequent steps meets the consistency requirements.

4. A method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The dynamic correction includes using a dynamic prediction and observation correction model to calculate the optimal state estimate of the data and adjusting the error range through a gain parameter to correct the uncertainty in the multimodal data.

5. The method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The dynamically corrected data must meet preset physical constraints, including continuity or conservation constraints; data that does not meet the constraints is marked or resampled.

6. A method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The semantic extraction adopts a variational inference model to extract the latent semantics of unstructured data by constructing an approximate posterior distribution and a latent semantic space; the semantic clustering utilizes a Gaussian mixture model to perform clustering optimization on the extracted semantic data.

7. A method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The dynamic update of the knowledge base is based on a consistency protocol, data synchronization between multiple nodes is achieved through distributed log replication, and the content that needs to be updated first is screened according to the change in information entropy of the knowledge items.

8. The method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The knowledge base supports digital employee question and answering through a semantic search mechanism, where the semantic search is sorted according to the semantic relevance between the input query content and the knowledge items, and an attention mechanism is used to improve the contextual adaptability of the question and answer results.

9. The method for constructing a knowledge base to assist digital employee question and answer according to claim 1, characterized in that: The knowledge base dynamically adjusts the content of knowledge items based on user feedback, where the feedback information includes the correctness score of the question and answer results, and optimizes the semantic adaptability of the knowledge items through training models.

10. A knowledge base construction system for assisting digital employee question and answer, characterized in that: include: Data collection module, used to collect multimodal data within the enterprise; Information volume analysis module, used to analyze the information entropy of the collected data and filter out high-value data; Data interaction verification module, used to perform consistency analysis on multimodal data; Dynamic correction module, used for dynamic prediction and correction of multimodal data; Semantic extraction and clustering module, used to extract semantics and cluster data to optimize and generate knowledge items; Distributed knowledge base module, used to store the generated knowledge entries and implement dynamic updates under the consistency protocol; The Q&A interaction module is used to support digital employee Q&A by leveraging the knowledge base through semantic search and receiving user feedback to optimize the knowledge base content.

Citation Information

Patent Citations

  • Multi-modal machine translation method based on variational reasoning and multi-task learning

    CN112016332A

  • Distributed storage method and system for massive industrial small files

    CN115827560A

  • Method and system for quickly constructing industry question and answer knowledge base

    CN117290489A

  • Intelligent document question and answer method based on large model

    CN117932018A

  • Intelligent wearable device based on health monitoring and early warning of electric power operating personnel

    CN118512159A

Cited By

  • Supplier portrait visual display method and system based on knowledge graph

    CN120744128A

  • Supplier portrait visual display method and system based on knowledge graph

    CN120744128B

  • Knowledge question-answering method and system based on multi-agent collaboration and distributed consensus

    CN120851226A

  • Intelligent student question and answer method and system based on large language model and rapid retrieval

    CN121009178A