A multimodal knowledge graph construction and retrieval system and method
By designing cascading data processing and management units in the multimodal knowledge graph construction and retrieval system, and distinguishing data processing methods based on the data change rate, the existing system's low efficiency and low real-time nature are solved, and efficient and real-time knowledge graph construction is achieved.
Patent Information
- Application Number
- CN202210122075.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-02-09
AI Technical Summary
The existing multimodal knowledge graph construction and retrieval system has problems of low efficiency and low real-time performance.
A multimodal knowledge graph construction and retrieval system is designed, and the cascading knowledge data acquisition and processing unit, a knowledge graph construction management unit and a knowledge graph application service unit are used to collect data using the multimodal data acquisition unit, and distinguish high-speed update and slow update data according to the real-time change rate of the data. The fast data estimation fusion program and the slow data processing fusion program are called for data fusion respectively.
It realizes high efficiency and real-time knowledge graph construction, can effectively process rapidly changing data, and improve the overall comprehensive efficiency of the system.
Smart Images

Figure CN114741466B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graphs, and specifically to a multimodal knowledge graph construction and retrieval system and method. Background Art
[0002] Knowledge Graph, also known as knowledge domain visualization or knowledge domain mapping map, is a series of various graphics that show the development process and structural relationship of knowledge. It uses visualization technology to describe knowledge resources and their carriers, and mine, analyze, construct, draw and display knowledge and their mutual connections. Knowledge Graph is a modern theory that combines the theories and methods of applied mathematics, graphics, information visualization technology, information science and other disciplines with metrological citation analysis, co-occurrence analysis and other methods, and uses visualized graphs to vividly display the core structure, development history, frontier fields and overall knowledge architecture of the discipline to achieve the purpose of multidisciplinary integration. Knowledge Graph can provide practical and valuable reference for disciplinary research.
[0003] The existing multimodal knowledge graph construction and retrieval system and method have the technical problems of low efficiency and low real-time performance. The present invention provides a multimodal knowledge graph construction and retrieval system and method to solve the above technical problems. Summary of the invention
[0004] The technical problem to be solved by the present invention is the low efficiency and low real-time technical problem existing in the prior art. A new multimodal knowledge graph construction and retrieval system is provided, which has the characteristics of high efficiency and high real-time performance.
[0005] In order to solve the above technical problems, the technical solutions adopted are as follows:
[0006] A multimodal knowledge graph construction and retrieval system, comprising a cascaded knowledge data acquisition and processing unit, a knowledge graph construction management unit and a knowledge graph application service unit;
[0007] The knowledge data acquisition and processing unit is used for acquiring and transmitting data, including a multimodal data acquisition unit;
[0008] The knowledge graph construction management unit is used for the construction and update management of the knowledge graph; the construction of the knowledge graph includes building an ontology according to business needs, completing knowledge fusion according to data content and ontology structure, associating labeled data with the ontology, and completing the construction of the knowledge graph model;
[0009] The knowledge graph application service unit includes knowledge retrieval unit, knowledge association and recommendation unit, and knowledge question and answer unit;
[0010] The knowledge fusion execution includes the following steps:
[0011] Step S1, calculate the real-time change rate of each multimodal data, and distinguish the high-speed update data P according to the predefined change rate threshold n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r Calculate the change value of
[0012] Step S2: The change value estimated by the fast data estimation fusion program exceeds a predefined threshold or the slow update data P r If the change value exceeds the predefined threshold, at least two knowledge graph construction models are called to complete the knowledge graph model construction;
[0013] Step S3, the knowledge graph model constructed in step S2 is formed into a final knowledge graph model according to a predefined voting strategy.
[0014] Working principle of the present invention: When constructing a knowledge graph for multimodal data as original data, the decision and method for processing the data can be effectively distinguished according to the rate of change, thereby improving real-time performance and efficiency. The present invention calculates the real-time rate of change of each multimodal data and distinguishes the high-speed update data P according to the predefined rate of change threshold. n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r The change value is calculated to achieve efficient knowledge graph construction; the multimodal data acquisition unit includes a text data acquisition unit, an image data acquisition unit, an audio data acquisition unit and a video data acquisition unit.
[0015] Further, calling the fast data estimation and fusion program to perform data estimation and fusion includes:
[0016] Step R1, define Among them, {x1,x2,...x k ,x K} is the K independent data sample observation values in the historical high-speed update data sample, k = 1, 2, 3...K, j and w are predefined parameters, w1, w2,...w k is the set of real numbers;
[0017] Step R2, through y k =μ+αt k +ε k , μ = log(2γ), calculate the characteristic index α and the dispersion coefficient γ; where, ε k is the error term coefficient with the same distribution but independent with a predefined mean of 0, t k =log|w k |, K is the number of historical samples;
[0018] Step R3, by z k =δw k +ε k , calculate the location parameter δ, where
[0019] z k =arctan(Im(w k ) / Re(w k )), ε k It is the error term coefficient with the same distribution but independent with a predefined mean of 0;
[0020] Step R4, substitute the characteristic index α, dispersion coefficient γ, and location parameter δ obtained in steps R2 and R3 into φ(w)=exp{jδw-γ|w| ∝}, and perform Fourier transform to obtain the probability density function f(x), completing the high-speed update of data P n The data estimation fusion.
[0021] Further, calling a fast data estimation and fusion program to perform data estimation and fusion also includes:
[0022] Step R5, confirm As a fast data estimation fusion procedure, whether the estimated change value exceeds a predefined threshold T max Indicators; A is the real-time estimated and fused data value, is a parameter estimated by historical high-speed update data samples, T max is the detection threshold corresponding to the pre-defined fusion rate.
[0023] Furthermore, the multimodal knowledge graph construction and retrieval system includes a plurality of the knowledge data acquisition and processing units, a plurality of knowledge graph construction management units, and a plurality of knowledge graph application service units;
[0024] Step A1, select multiple knowledge data acquisition and processing units, multiple knowledge graph construction and management units, and multiple knowledge graph application service units to form a real-time system;
[0025] Step A2, optionally selecting adjacent front and rear stages, defining the unit of the front stage as a primary unit, and the unit of the rear stage as a secondary unit;
[0026] Step A3, defining the real-time system performance model as H=H1·H2·H3·H4·H5, where H1 is effectiveness, H2 is processing efficiency, H3 is system load rate, H4 is data processing accuracy, and H5 is system failure rate;
[0027] Step A4, H4 is predefined, H5 is the real-time system failure rate calculated based on historical conditions, and is calculated according to the following formula: H2=PH 21 +(1-P)H 21 H 22 , H3=(NH 31 +NH 32 ) / (N+M);
[0028] Where W = PW1 + (1-P) (W1 + W2), T = PT1 + (1-P) (T1 + T2), t is the total time that the data is in the primary unit and the secondary unit, P is the predefined probability that the data enters the secondary unit from the primary unit, and the primary unit processing efficiency Secondary Unit Processing Efficiency Primary unit load factor Secondary Unit Load Factor N is the number of primary units, M is the number of secondary units, R is an integer, P R According to the predefined average data volume of the primary unit Find, Q R Based on the average data volume of the predefined secondary unit It is found that W1 = L1 / λ is the average response time of the primary unit data, and W2 = L2 / λH 21 P is the secondary unit response time; T1 = 1 / μ1 is the average service time of the primary unit data, T2 = 1 / μ2 is the average service time of the secondary unit data; μ1 and μ2 are the parameters of the exponential distribution, and λ is the predefined Poisson parameter;
[0029] Step A5, calculate the overall performance value of the real-time system, determine the size of the overall performance value, if it is greater than a predefined threshold, return to step A1 to reselect a new real-time system.
[0030] The present invention also provides a multimodal knowledge graph construction and retrieval method, the method comprising:
[0031] Step 1: The multimodal data collection unit collects knowledge data and preprocesses the knowledge data to distinguish data categories, establish data identifiers, generate standard data items, and determine whether the standard data items exist in the knowledge graph database. If so, the identifier is obtained for indexing. If not, the standard data items are stored.
[0032] Step 2: Build an ontology according to business needs, and construct a mapping relationship between standard data items and ontology to complete the preliminary construction of the knowledge graph model;
[0033] Step 3: Complete knowledge fusion processing based on data content and ontology structure and update the knowledge graph, including:
[0034] Step S1, calculate the real-time change rate of each multimodal data, and distinguish the high-speed update data P according to the predefined change rate threshold n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r Calculate the change value of
[0035] Step S2: The change value estimated by the fast data estimation fusion program exceeds a predefined threshold or the slow update data P r If the change value exceeds the predefined threshold, at least two knowledge graph construction models are called to complete the knowledge graph model construction;
[0036] Step S3: The knowledge graph model constructed in step S2 is converted into the final knowledge graph model according to the predefined voting strategy.
[0037] Step 4: The knowledge graph application service unit calls the knowledge graph to complete the business according to business needs.
[0038] Further, calling the fast data estimation and fusion program to perform data estimation and fusion includes:
[0039] Step R1, define Among them, {x1,x2,...x k ,x K} is the K independent data sample observation values in the historical high-speed update data sample, k = 1, 2, 3...K, j and w are predefined parameters, w1, w2,...w k is the set of real numbers;
[0040] Step R2, through y k =μ+αt k +εk , μ = log(2γ), calculate the characteristic index ∝ and the dispersion coefficient γ; where, ε k is the error term coefficient with the same distribution but independent with a predefined mean of 0, t k =log|w k |, K is the number of historical samples;
[0041] Step R3, by z k =δw k +ε k , calculate the position parameter δ, where z k =arctan(Im(w k ) / Re(w k ), ε k It is the error term coefficient with the same distribution but independent with a predefined mean of 0;
[0042] Step R4, substitute the characteristic index ∝, dispersion coefficient γ, and location parameter δ obtained in steps R2 and R3 into φ(w)=exp{jδw-γ|w| ∝}, and perform Fourier transform to obtain the probability density function f(x), completing the high-speed update of data P n The fitted estimates of fusion.
[0043] Further, calling a fast data estimation and fusion program to perform data estimation and fusion also includes:
[0044] Step R5, confirm As a fast data estimation fusion procedure, whether the estimated change value exceeds a predefined threshold T max Indicators; A is the real-time estimated and fused data value, is a parameter estimated by historical high-speed update data samples, T max is the detection threshold corresponding to the pre-defined fusion rate.
[0045] Furthermore, the multimodal knowledge graph construction and retrieval method further includes:
[0046] Step A1, select multiple knowledge data acquisition and processing units, multiple knowledge graph construction and management units, and multiple knowledge graph application service units to form a real-time system;
[0047] Step A2, optionally selecting adjacent front and rear stages, defining the unit of the front stage as a primary unit, and the unit of the rear stage as a secondary unit;
[0048] Step A3, defining the real-time system performance model as H=H1·H2·H3·H4·H5, where H1 is effectiveness, H2 is processing efficiency, H3 is system load rate, H4 is data processing accuracy, and H5 is system failure rate;
[0049] Step A4, H4 is predefined, H5 is the real-time system failure rate calculated based on historical conditions, and is calculated according to the following formula: H2=PH 21 +(1-P)H 21 H 22 , H3=(NH 31 +NH 32 ) / (N+M);
[0050] Where W = PW1 + (1-P) (W1 + W2), T = PT1 + (1-P) (T1 + T2), t is the total time that the data is in the primary unit and the secondary unit, P is the predefined probability that the data enters the secondary unit from the primary unit, and the primary unit processing efficiency Secondary Unit Processing Efficiency Primary unit load factor Secondary Unit Load Factor N is the number of primary units, M is the number of secondary units, R is an integer, P R According to the predefined average data volume of the primary unit Find, Q R Based on the average data volume of the predefined secondary unit It is found that W1 = L1 / λ is the average response time of the primary unit data, and W2 = L2 / λH 21 P is the secondary unit response time; T1 = 1 / μ1 is the average service time of the primary unit data, T2 = 1 / μ2 is the average service time of the secondary unit data; μ1 and μ2 are the parameters of the exponential distribution, and λ is the predefined Poisson parameter;
[0051] Step A5, calculate the overall performance value of the real-time system, determine the size of the overall performance value, if it is greater than a predefined threshold, return to step A1 to reselect a new real-time system.
[0052] Beneficial effects of the present invention: When constructing a knowledge graph for multimodal data as the original data, the present invention can effectively distinguish the decision and method for processing the data according to the change rate, thereby improving real-time performance and efficiency. The present invention calculates the real-time change rate of each multimodal data and distinguishes the high-speed update data P according to the predefined change rate threshold. n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r, directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r The change value is calculated to achieve efficient knowledge graph construction. For fast-changing data, high-speed and high-precision fitting is achieved through the unique algorithm of the present invention. The overall comprehensive efficiency of the system is evaluated in real time, and the composition of the system is adjusted to achieve efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0054] Figure 1 , Schematic diagram of multimodal knowledge graph construction and retrieval system.
[0055] Figure 2 ,Schematic diagram of knowledge fusion steps. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0057] Example 1
[0058] This embodiment provides a multimodal knowledge graph construction and retrieval system. Figure 1 , the multimodal knowledge graph construction and retrieval system includes a cascaded knowledge data acquisition and processing unit, a knowledge graph construction management unit and a knowledge graph application service unit;
[0059] The knowledge data acquisition and processing unit is used for acquiring and transmitting data, including a multimodal data acquisition unit;
[0060] The knowledge graph construction management unit is used for the construction and update management of the knowledge graph; the construction of the knowledge graph includes building an ontology according to business needs, completing knowledge fusion according to data content and ontology structure, associating labeled data with the ontology, and completing the construction of the knowledge graph model;
[0061] The knowledge graph application service unit includes knowledge retrieval unit, knowledge association and recommendation unit, and knowledge question and answer unit;
[0062] like Figure 2 ,The knowledge fusion execution includes the following steps:
[0063] Step S1, calculate the real-time change rate of each multimodal data, and distinguish the high-speed update data P according to the predefined change rate threshold n and slow update data P r ; For high-speed update data P n, call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r Calculate the change value of
[0064] Step S2: The change value estimated by the fast data estimation fusion program exceeds a predefined threshold or the slow update data P r If the change value exceeds the predefined threshold, at least two knowledge graph construction models are called to complete the knowledge graph model construction;
[0065] Step S3, the knowledge graph model constructed in step S2 is formed into a final knowledge graph model according to a predefined voting strategy.
[0066] In this embodiment, when constructing a knowledge graph for multimodal data as the original data, the decision and method for processing the data can be effectively distinguished according to the change rate, thereby improving real-time performance and efficiency. The present invention calculates the real-time change rate of each multimodal data and distinguishes the high-speed update data P according to the predefined change rate threshold. n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r The change value is calculated to achieve efficient knowledge graph construction.
[0067] Specifically, the multimodal data acquisition unit includes a text data acquisition unit, an image data acquisition unit, an audio data acquisition unit and a video data acquisition unit.
[0068] Preferably, in order to efficiently fit, estimate, and integrate fast data, this embodiment uses a special fast data estimation and fusion program to estimate and fuse data. Of course, existing data fusion methods can also be used. The method of this embodiment includes:
[0069] Step R1, define Among them, {x1,x2,...x k ,x K} is the K independent data sample observation values in the historical high-speed update data sample, k = 1, 2, 3...K, j and w are predefined parameters, w1, w2,...w k is the set of real numbers;
[0070] Step R2, through y k=μ+αt k +ε k , μ = log(2γ), calculate the characteristic index ∝ and the dispersion coefficient γ; where, ε k is the error term coefficient with the same distribution but independent with a predefined mean of 0, t k =log|w k |, K is the number of historical samples;
[0071] Step R3, by z k =δw k +ε k , calculate the position parameter δ, where z k =arctan(Im(w k ) / Re(w k ), ε k It is the error term coefficient with the same distribution but independent with a predefined mean of 0;
[0072] Step R4, substitute the characteristic index ∝, dispersion coefficient γ, and location parameter δ obtained in steps R2 and R3 into φ(w)=exp{jδw-γ|w| ∝}, and perform Fourier transform to obtain the probability density function f(x), completing the high-speed update of data P n The fitted estimates of fusion.
[0073] Specifically, calling a fast data estimation and fusion program to perform data estimation and fusion also includes:
[0074] Step R5, confirm As a fast data estimation fusion procedure, whether the estimated change value exceeds a predefined threshold T max Indicators; A is the real-time estimated and fused data value, is a parameter estimated by historical high-speed update data samples, T max is the detection threshold corresponding to the pre-defined fusion rate.
[0075] Preferably, in order to prevent the inefficiency caused by real-time system degradation and failure, preferably, the knowledge data acquisition processing unit includes multiple, the knowledge graph construction management unit includes multiple, and the knowledge graph application service unit includes multiple;
[0076] Step A1, select multiple knowledge data acquisition and processing units, multiple knowledge graph construction and management units, and multiple knowledge graph application service units to form a real-time system;
[0077] Step A2, optionally selecting adjacent front and rear stages, defining the unit of the front stage as a primary unit, and the unit of the rear stage as a secondary unit;
[0078] Step A3, defining the real-time system performance model as H=H1·H2·H3·H4·H5, where H1 is effectiveness, H2 is processing efficiency, H3 is system load rate, H4 is data processing accuracy, and H5 is system failure rate;
[0079] Step A4, H4 is predefined, H5 is the real-time system failure rate calculated based on historical conditions, and is calculated according to the following formula: H2=PH 21 +(1-P)H 21 H 22 , H3=(NH 31 +NH 32 ) / (N+M);
[0080] Where W = PW1 + (1-P) (W1 + W2), T = PT1 + (1-P) (T1 + T2), t is the total time that the data is in the primary unit and the secondary unit, P is the predefined probability that the data enters the secondary unit from the primary unit, and the primary unit processing efficiency Secondary Unit Processing Efficiency Primary unit load factor Secondary Unit Load Factor N is the number of primary units, M is the number of secondary units, R is an integer, P R According to the predefined average data volume of the primary unit Find, Q R Based on the average data volume of the predefined secondary unit It is found that W1 = L1 / λ is the average response time of the primary unit data, and W2 = L2 / λH 21 P is the secondary unit response time; T1 = 1 / μ1 is the average service time of the primary unit data, T2 = 1 / μ2 is the average service time of the secondary unit data; μ1 and μ2 are the parameters of the exponential distribution, and λ is the predefined Poisson parameter;
[0081] Step A5, calculate the overall performance value of the real-time system, determine the size of the overall performance value, if it is greater than a predefined threshold, return to step A1 to reselect a new real-time system.
[0082] This embodiment also provides a multimodal knowledge graph construction and retrieval method, the method comprising:
[0083] Step 1: The multimodal data collection unit collects knowledge data and preprocesses the knowledge data to distinguish data categories, establish data identifiers, generate standard data items, and determine whether the data exists in the knowledge graph database. If it exists, the identifier is obtained for indexing. If it does not exist, it is stored.
[0084] Step 2: Build an ontology according to business needs, and construct a mapping relationship between standard data items and ontology to complete the preliminary construction of the knowledge graph model;
[0085] Step 3: Complete knowledge fusion processing based on data content and ontology structure and update the knowledge graph, including:
[0086] Step S1, calculate the real-time change rate of each multimodal data, and distinguish the high-speed update data P according to the predefined change rate threshold n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r Calculate the change value of
[0087] Step S2: The change value estimated by the fast data estimation fusion program exceeds a predefined threshold or the slow update data P r If the change value exceeds the predefined threshold, at least two knowledge graph construction models are called to complete the knowledge graph model construction;
[0088] Step S3: The knowledge graph model constructed in step S2 is converted into the final knowledge graph model according to the predefined voting strategy.
[0089] Step 4: The knowledge graph application service unit calls the knowledge graph to complete the business according to business needs.
[0090] Preferably, based on the conventional data fusion method, the fast data estimation fusion procedure of this embodiment includes:
[0091] Step R1, define Among them, {x1,x2,...x k ,x K} is the K independent data sample observation values in the historical high-speed update data sample, k = 1, 2, 3...K, j and w are predefined parameters, w1, w2,...w k is the set of real numbers;
[0092] Step R2, through y k =μ+αt k +ε k , μ = log(2γ), calculate the characteristic index ∝ and the dispersion coefficient γ; where, ε k is the error term coefficient with the same distribution but independent with a predefined mean of 0, tk =log|w k |, K is the number of historical samples;
[0093] Step R3, by z k =δw k +ε k , calculate the position parameter δ, where z k =arctan(Im(w k ) / Re(w k ), ε k It is the error term coefficient with the same distribution but independent with a predefined mean of 0;
[0094] Step R4, substitute the characteristic index ∝, dispersion coefficient γ, and location parameter δ obtained in steps R2 and R3 into φ(w)=exp{jδw-γ|w| ∝}, and perform Fourier transform to obtain the probability density function f(x), completing the high-speed update of data P n The fitted estimates of fusion.
[0095] Preferably, calling a fast data estimation and fusion program to perform data estimation and fusion also includes:
[0096] Step R5, confirm As a fast data estimation fusion procedure, whether the estimated change value exceeds a predefined threshold T max Indicators; A is the real-time estimated and fused data value, is a parameter estimated by historical high-speed update data samples, T max is the detection threshold corresponding to the pre-defined fusion rate.
[0097] Preferably, in order to improve the real-time efficiency of the system and prevent failures or system function degradation, the multimodal knowledge graph construction and retrieval method further includes:
[0098] Step A1, select multiple knowledge data acquisition and processing units, multiple knowledge graph construction and management units, and multiple knowledge graph application service units to form a real-time system;
[0099] Step A2, optionally selecting adjacent front and rear stages, defining the unit of the front stage as a primary unit, and the unit of the rear stage as a secondary unit;
[0100] Step A3, defining the real-time system performance model as H=H1·H2·H3·H4·H5, where H1 is effectiveness, H2 is processing efficiency, H3 is system load rate, H4 is data processing accuracy, and H5 is system failure rate;
[0101] Step A4, H4 is predefined, H5 is the real-time system failure rate calculated based on historical conditions, and is calculated according to the following formula: H2=PH 21 +(1-P)H 21 H 22 , H3=(NH 31 +NH 32 ) / (N+M);
[0102] Where W = PW1 + (1-P) (W1 + W2), T = PT1 + (1-P) (T1 + T2), t is the total time that the data is in the primary unit and the secondary unit, P is the predefined probability that the data enters the secondary unit from the primary unit, and the primary unit processing efficiency Secondary Unit Processing Efficiency Primary unit load factor Secondary Unit Load Factor N is the number of primary units, M is the number of secondary units, R is an integer, P R According to the predefined average data volume of the primary unit Find, Q R Based on the average data volume of the predefined secondary unit It is found that W1 = L1 / λ is the average response time of the primary unit data, and W2 = L2 / λH 21 P is the secondary unit response time; T1 = 1 / μ1 is the average service time of the primary unit data, T2 = 1 / μ2 is the average service time of the secondary unit data; μ1 and μ2 are the parameters of the exponential distribution, and λ is the predefined Poisson parameter;
[0103] Step A5, calculate the overall performance value of the real-time system, determine the size of the overall performance value, if it is greater than a predefined threshold, return to step A1 to reselect a new real-time system.
[0104] In this embodiment, when constructing a knowledge graph for the original data of multimodal data, the decision and method for processing the data can be effectively distinguished according to the change rate, thereby improving real-time performance and efficiency. The present invention calculates the real-time change rate of each multimodal data and distinguishes the high-speed update data P according to the predefined change rate threshold. n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P rThe change value is calculated to achieve efficient knowledge graph construction. For fast-changing data, high-speed and high-precision fitting is achieved through the unique algorithm of the present invention. The overall comprehensive efficiency of the system is evaluated in real time, and the composition of the system is adjusted to achieve efficiency.
[0105] Although the above describes the illustrative specific embodiments of the present invention so that those skilled in the art can understand the present invention, the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, all inventions and creations using the concepts of the present invention are protected.
Claims
1. A multimodal knowledge graph construction and retrieval system, characterized by: The multimodal knowledge graph construction and retrieval system includes a cascaded knowledge data acquisition and processing unit, a knowledge graph construction management unit, and a knowledge graph application service unit; The knowledge data acquisition and processing unit is used for acquiring and transmitting data, including a multimodal data acquisition unit; The knowledge graph construction management unit is used for the construction and update management of the knowledge graph; the construction of the knowledge graph includes building an ontology according to business needs, completing knowledge fusion according to data content and ontology structure, associating labeled data with the ontology, and completing the construction of the knowledge graph model; The knowledge graph application service unit includes knowledge retrieval unit, knowledge association and recommendation unit, and knowledge question and answer unit; The knowledge fusion execution includes the following steps: Step S1, calculate the real-time change rate of each multimodal data, and distinguish the high-speed update data P according to the predefined change rate threshold n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r Calculate the change value of Step S2: The change value estimated by the fast data estimation fusion program exceeds a predefined threshold or the slow update data P r If the change value exceeds the predefined threshold, at least two knowledge graph construction models are called to complete the knowledge graph model construction; Step S3, the knowledge graph model constructed in step S2 is used to form a final knowledge graph model according to a predefined voting strategy; The multimodal data acquisition unit includes a text data acquisition unit, an image data acquisition unit, an audio data acquisition unit and a video data acquisition unit.
2. The multimodal knowledge graph construction and retrieval system according to claim 1, characterized in that: Calling the fast data estimation and fusion program to perform data estimation and fusion includes: Step R1, define Among them, {x1,x2,...x k ,x K } is the K independent data sample observation values in the historical high-speed update data sample, k = 1, 2, 3...K, j and w are predefined parameters, w1, w2,...w k is the set of real numbers; Step R2, through y k =μ+αt k +ε k , μ = log(2γ), calculate the characteristic index α and the dispersion coefficient γ; where, ε k is the error term coefficient with the same distribution but independent with a predefined mean of 0, t k =log|w k |, K is the number of historical samples; Step R3, by z k =δw k +ε k , calculate the position parameter δ, where z k =arctan(Im(w k ) / Re(w k )), ε k It is the error term coefficient with the same distribution but independent with a predefined mean of 0; Step R4, the characteristic index α, dispersion coefficient γ, and location parameter δ obtained in steps R2 and R3 are substituted into φ(w)=exp{jδw-γ|w| ∝ }, and perform Fourier transform to obtain the probability density function f(x), completing the high-speed update of data P n The data estimation fusion.
3. The multimodal knowledge graph construction and retrieval system according to claim 2, characterized in that: Call the fast data estimation and fusion program to perform data estimation and fusion, including: Step R5, confirm As a fast data estimation fusion procedure, whether the estimated change value exceeds a predefined threshold T max Indicators; A is the real-time estimated and fused data value, is a parameter estimated by historical high-speed update data samples, T max is the detection threshold corresponding to the pre-defined fusion rate.
4. The multimodal knowledge graph construction and retrieval system according to claim 1, characterized in that: The multimodal knowledge graph construction and retrieval system includes a plurality of knowledge data acquisition and processing units, a plurality of knowledge graph construction and management units, and a plurality of knowledge graph application service units; Step A1, select multiple knowledge data acquisition and processing units, multiple knowledge graph construction and management units, and multiple knowledge graph application service units to form a real-time system; Step A2, optionally selecting adjacent front and rear stages, defining the unit of the front stage as a primary unit, and the unit of the rear stage as a secondary unit; Step A3, defining the real-time system performance model as H=H1·H2·H3·H4·H5, where H1 is effectiveness, H2 is processing efficiency, H3 is system load rate, H4 is data processing accuracy, and H5 is system failure rate; Step A4, H4 is predefined, H5 is the real-time system failure rate calculated based on historical conditions, and is calculated according to the following public statement H2=PH 21 +(1-P)H 21 H 22 , H3=(NH 31 +NH 32 ) / (N+M); Where W = PW1 + (1-P) (W1 + W2), T = PT1 + (1-P) (T1 + T2), t is the total time that the data is in the primary unit and the secondary unit, P is the predefined probability that the data enters the secondary unit from the primary unit, and the primary unit processing efficiency Secondary Unit Processing Efficiency M)Q R ; Primary unit load factor Secondary Unit Load Factor N is the number of primary units, M is the number of secondary units, R is an integer, P R According to the predefined average data volume of the primary unit Find, Q R Based on the average data volume of the predefined secondary unit It is found that W1 = L1 / λ is the average response time of the primary unit data, and W2 = L2 / λH 21 P is the secondary unit response time; T1 = 1 / μ1 is the average service time of the primary unit data, T2 = 1 / μ2 is the average service time of the secondary unit data; μ1 and μ2 are the parameters of the exponential distribution, and λ is the predefined Poisson parameter; Step A5, calculate the overall performance value of the real-time system, determine the size of the overall performance value, if it is greater than a predefined threshold, return to step A1 to reselect a new real-time system.
5. A multimodal knowledge graph construction and retrieval method, characterized by: The multimodal knowledge graph construction and retrieval method is based on the multimodal knowledge graph construction and retrieval system according to any one of claims 1 to 4, and the method comprises: Step 1: The multimodal data collection unit collects knowledge data and preprocesses the knowledge data to distinguish data categories, establish data identifiers, generate standard data items, and determine whether the standard data items exist in the knowledge graph database. If the standard data items exist, the identifier is obtained for indexing. If not, the standard data items are stored. Step 2: Build an ontology according to business needs, and construct a mapping relationship between standard data items and ontology to complete the preliminary construction of the knowledge graph model; Step 3: Complete knowledge fusion processing based on data content and ontology structure and update the knowledge graph, including: Step S1, calculate the real-time change rate of each multimodal data, and distinguish the high-speed update data P according to the predefined change rate threshold n and slow update data P r ; For high-speed update data P n , call the fast data estimation and fusion program to perform data estimation and fusion; for the slow update data P r , directly call the slow data processing fusion program for calculation and fusion, and directly update the slow data P r Calculate the change value of Step S2: The change value estimated by the fast data estimation fusion program exceeds a predefined threshold or the slow update data P r If the change value exceeds the predefined threshold, at least two knowledge graph construction models are called to complete the knowledge graph model construction; Step S3: The knowledge graph model constructed in step S2 is converted into the final knowledge graph model according to the predefined voting strategy. Step 4: The knowledge graph application service unit calls the knowledge graph to complete the business according to business needs.
6. The multimodal knowledge graph construction and retrieval method according to claim 5, characterized in that: Calling the fast data estimation and fusion program to perform data estimation and fusion includes: Step R1, define Among them, {x1,x2,...x k ,x K } is the K independent data sample observation values in the historical high-speed update data sample, k = 1, 2, 3...K, j and w are predefined parameters, w1, w2,...w k is the set of real numbers; Step R2, through y k =μ+αt k +ε k , μ = log(2γ), calculate the characteristic index α and the dispersion coefficient γ; where, ε k is the error term coefficient with the same distribution but independent with a predefined mean of 0, t k =log|w k |, K is the number of historical samples; Step R3, by z k =δw k +ε k , calculate the position parameter δ, where z k =arctan(Im(w k ) / Re(w k )), ε k It is the error term coefficient with the same distribution but independent with a predefined mean of 0; Step R4, substitute the characteristic index α, dispersion coefficient γ, and location parameter δ obtained in steps R2 and R3 into φ(w)=exp{jδw-γ|w| ∝ }, and perform Fourier transform to obtain the probability density function f(x), completing the high-speed update of data P n The data estimation fusion.
7. The multimodal knowledge graph construction and retrieval method according to claim 5, characterized in that: Call the fast data estimation and fusion program to perform data estimation and fusion, including: Step R5, confirm As a fast data estimation fusion procedure, whether the estimated change value exceeds a predefined threshold T max Indicators; A is the real-time estimated and fused data value, is a parameter estimated by historical high-speed update data samples, T max is the detection threshold corresponding to the pre-defined fusion rate.
8. The multimodal knowledge graph construction and retrieval method according to claim 5, characterized in that: The multimodal knowledge graph construction and retrieval method further includes: Step A1, select multiple knowledge data acquisition and processing units, multiple knowledge graph construction and management units, and multiple knowledge graph application service units to form a real-time system; Step A2, optionally selecting adjacent front and rear stages, defining the unit of the front stage as a primary unit, and the unit of the rear stage as a secondary unit; Step A3, defining the real-time system performance model as H=H1·H2·H3·H4·H5, where H1 is effectiveness, H2 is processing efficiency, H3 is system load rate, H4 is data processing accuracy, and H5 is system failure rate; Step A4, H4 is predefined, H5 is the real-time system failure rate calculated based on historical conditions, and is calculated according to the following public statement H2=PH 21 +(1-P)H 21 H 22 , H3=(NH 31 +NH 32 ) / (N+M); Where W = PW1 + (1-P) (W1 + W2), T = PT1 + (1-P) (T1 + T2), t is the total time that the data is in the primary unit and the secondary unit, P is the predefined probability that the data enters the secondary unit from the primary unit, and the primary unit processing efficiency H 21 =1- Secondary Unit Processing Efficiency M)Q R ; Primary unit load factor Secondary Unit Load Factor N is the number of primary units, M is the number of secondary units, R is an integer, P R According to the predefined average data volume of the primary unit Find, Q R Based on the average data volume of the predefined secondary unit So, W1 = L1 / λ is the average response time of primary unit data, W2 = L2 / λH 21 P is the secondary unit response time; T1 = 1 / μ1 is the average service time of the primary unit data, T2 = 1 / μ2 is the average service time of the secondary unit data; μ1 and μ2 are the parameters of the exponential distribution, and λ is the predefined Poisson parameter; Step A5, calculate the overall performance value of the real-time system, determine the size of the overall performance value, if it is greater than a predefined threshold, return to step A1 to reselect a new real-time system.
Citation Information
Patent Citations
A knowledge graph system construction method
CN109697233A
Limited domain-oriented knowledge graph updating method and system
CN111914550A