Method for constructing front-end engineering design database applicable to FPSO project

By constructing a knowledge graph for the FEED phase of the FPSO project, the degree of correlation and content similarity of document data were determined, which solved the problem of missing data correlation in the database and improved the database query quality and design efficiency.

CN120744138BActive Publication Date: 2025-11-04DALIAN JINCHENG YANGFAN OCEAN ENG DESIGN CO LTD

Patent Information

Application Number
CN202511179391.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-04
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing technologies suffer from a lack of data correlation when establishing the front-end engineering design (FEED) database for FPSO projects. This leads to logical breaks, disconnect between physical and logical objects, and difficulty in forming reliable and clear index relationships, thus affecting design efficiency and quality.

Method used

By constructing a knowledge graph for the FEED phase of an FPSO project, the degree of data association and content similarity between document data is determined. Based on the index association degree, related documents are sorted to establish clear index relationships and improve the quality of database queries.

Benefits of technology

It establishes reliable indexing relationships between document data, improving the efficiency and quality of FPSO project design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744138B_ABST
    Figure CN120744138B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data retrieval, in particular to a construction method of a front-end engineering design database suitable for an FPSO project, which comprises the following steps: determining associated documents of document data according to the data correlation degree between any two document data in the FEED stage of the FPSO project; obtaining the index correlation degree of the document data and each associated document of the document data according to the data correlation degree of the document data and each associated document of the document data and the initial index quantity difference; sorting the associated documents based on the index correlation degree to obtain the index result display priority of the document data and the reliable and clear index relationship between the data, so that the query quality of each document data in the construction process of the database is improved, and finally the design efficiency and quality of the FPSO project are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data retrieval, and particularly relates to a construction method of a front-end engineering design database suitable for an FPSO project. BACKGROUND

[0002] FPSO (Floating Production Storage and Offloading) is an important facility for offshore oil and gas development, and plays a crucial role in offshore oil and gas exploitation. Front-end engineering design (FEED) of the FPSO project is a key link of the offshore oil and gas development project, aiming to provide a feasible and reasonable design scheme for the project, and to lay a foundation for subsequent detailed design and construction. Database establishment of FEED is a prerequisite for standardization of the design process.

[0003] The existing method has the deficiency of lack of association in the database establishment of FEED, such as logical break between data generated in different stages of the FEED design process, disconnection between physical objects and logical objects, and difficulty in forming reliable and clear index relationship between different data, resulting in many conflicts in collaborative design, difficulty in knowledge reuse, and thus seriously restricting the design efficiency and quality of the FPSO project. SUMMARY

[0004] In order to solve the technical problem that different data are difficult to form a reliable and clear index relationship, the purpose of the present application is to provide a construction method of a front-end engineering design database suitable for an FPSO project, and the technical scheme adopted is as follows:

[0005] The present application provides a construction method of a front-end engineering design database suitable for an FPSO project, comprising:

[0006] According to the data association degree between any two document data in the FEED stage of the FPSO project, the associated documents of the document data are determined;

[0007] According to the data association degree of the document data and each associated document thereof, and the difference in the number of initial indexes, the index association degree of the document data and each associated document thereof is obtained;

[0008] The associated documents are sorted based on the index association degree, and the index result display priority of the document data is obtained, which is used to indicate the database establishment.

[0009] In an exemplary embodiment, the data association degree acquisition process comprises:

[0010] According to the constructed knowledge graph of the FEED stage of the FPSO project, distribution similarity of the first document data and the second document data in a plurality of sub-knowledge graphs is determined; the first document data and the second document data are any two document data; the knowledge graph comprises the plurality of sub-knowledge graphs;

[0011] According to the number of same segmented words of the first document data and the second document data in the co-occurring sub-knowledge graph and the attribute consistency of the same segmented words, content similarity of the first document data and the second document data is obtained.

[0012] The data correlation degree of the first document data and the second document data is obtained by fusing the distribution similarity and the content similarity.

[0013] In an exemplary embodiment, the process of obtaining the distribution similarity comprises:

[0014] A first sub-knowledge graph sequence and a second sub-knowledge graph sequence, and a first number and a second number are determined; the first sub-knowledge graph sequence is composed of sub-knowledge graphs in which segmented words in the first document data appear; the second sub-knowledge graph sequence is composed of sub-knowledge graphs in which segmented words in the second document data appear, the first number is the number of sub-knowledge graphs in the first sub-knowledge graph sequence, and the second number is the number of sub-knowledge graphs in the second sub-knowledge graph sequence;

[0015] According to the number difference of the first number and the second number, and the sequence similarity of the first sub-knowledge graph sequence and the second sub-knowledge graph sequence, the distribution similarity is obtained; the distribution similarity is inversely related to the number difference, and is positively related to the sequence similarity.

[0016] In an exemplary embodiment, the sequence similarity is a Jaccard correlation coefficient of the first sub-knowledge graph sequence and the second sub-knowledge graph sequence.

[0017] In an exemplary embodiment, the process of obtaining the content similarity comprises:

[0018] The number of same segmented words of the first document data and the second document data in a target co-occurring sub-knowledge graph, and the attribute consistency of each same segmented word in the target co-occurring sub-knowledge graph are determined; the target co-occurring sub-knowledge graph is any one of the sub-knowledge graphs in which the segmented words of the first document data and the second document data co-occur;

[0019] According to the same number of word segmentation features and attribute consistency, content similarity of the first document data and the second document data with respect to the target co-located sub-knowledge graph is obtained; the content similarity and the same number of word segmentation features and attribute consistency are positively correlated; the same number of word segmentation features is obtained from the same number of word segmentation;

[0020] The content similarity of the first document data and the second document data with respect to all the co-located sub-knowledge graphs is fused to obtain the content similarity.

[0021] In an exemplary embodiment, the sub-knowledge graph is composed of a plurality of triples;

[0022] The attribute consistency acquisition process includes: if the same word segmentation of the first document data and the second document data has the same position in the triples in the target co-located sub-knowledge graph, the attribute consistency of the same word segmentation is a first value; if the same word segmentation of the first document data and the second document data has different positions in the triples in the target co-located sub-knowledge graph, the attribute consistency of the same word segmentation is a second value; the first value is greater than the second value.

[0023] In an exemplary embodiment, the same number of word segmentation feature acquisition process includes:

[0024] A minimum number of word segmentation is determined, which is the minimum value of the number of word segmentation of the first document data in the target co-located sub-knowledge graph and the number of word segmentation of the second document data in the target co-located sub-knowledge graph;

[0025] The ratio of the same number of word segmentation to the minimum number of word segmentation is taken as the same number of word segmentation feature.

[0026] In an exemplary embodiment, the determination of the associated document of the document data includes: comparing the data association degree of the document data with each of the other document data with a preset threshold, and taking the document data greater than the preset threshold as the associated document of the document data.

[0027] In an exemplary embodiment, the sorting of the associated document based on the index association degree includes: sorting each associated document of the document data in descending order of the index association degree.

[0028] In an exemplary embodiment, the database establishment method further includes: using the jieba word segmentation tool to perform word segmentation processing on the document data to obtain each word segmentation of the document data.

[0029] The present application has the following beneficial effects: first, the associated documents of the document data are determined according to the data correlation degree between any two document data of the FEED stage of the FPSO project, then the index correlation degree of the document data and each associated document thereof is obtained according to the data correlation degree of the document data and each associated document thereof and the difference in the initial index quantity between the two, the associated documents can be sorted according to the index correlation degree, the index result display priority of the document data is obtained, and the reliable and clear index relationship between the document data is obtained, thereby improving the query quality of each document data in the construction process of the database, and finally improving the design efficiency and quality of the FPSO project. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is a flowchart of a construction method of a front-end engineering design database suitable for an FPSO project provided by an embodiment of the present application;

[0031] Figure 2 is a flowchart of the acquisition of the data correlation degree provided by an embodiment of the present application;

[0032] Figure 3 is a flowchart of the acquisition of the distribution similarity provided by an embodiment of the present application;

[0033] Figure 4 is a flowchart of the acquisition of the content similarity provided by an embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined purpose, the specific embodiments, structures, features and effects of the present application are described in detail below in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The data information collected in the present application is obtained with the full authorization.

[0036] The present embodiment provides a construction method of a front-end engineering design database suitable for an FPSO project, which is suitable for the database establishment in the FEED stage of the FPSO project.

[0037] FEED data acquisition: FPSO integrates production processing, storage, transportation and life and power supply, and can preliminarily process offshore crude oil or natural gas, then store it, and then transport it to the shore terminal through a shuttle tanker. In a marine oil development project, the construction of FPSO usually involves the FEED stage, which is used to determine the overall scheme, technical parameters, equipment configuration, etc. of the FPSO, and to provide guidance for subsequent detailed design and construction. Among them, the FEED stage of the FPSO will go through project initiation, process design, structural design, equipment selection, performance evaluation and other processes, which contain various types of digital information and document information, such as long-term statistical target sea area weather data (wind speed, wind direction, air temperature, etc.), wave data (significant wave height, maximum wave height, period, wave direction, etc.), current data (current speed, flow direction, etc.) and other data-based data, as well as oilfield development plan, core process calculation book and other types of document data.

[0038] First, the various text data of the FEED stage (such as process calculation book, equipment specification book, etc. scattered in different subsystems) are modeled into a unified knowledge base through "entity-relation", which can be analyzed for logical relationship between data when establishing the FEED stage database. Therefore, the existing method is used to first construct the knowledge graph of the FEED stage of the FPSO project. The knowledge graph includes multiple sub-knowledge graphs, that is, multiple sub-knowledge graphs are constructed, and in an exemplary embodiment, the multiple sub-knowledge graphs include a process design sub-knowledge graph, a structure design sub-knowledge graph, a safety and environmental protection sub-knowledge graph, etc. Moreover, each sub-knowledge graph is composed of multiple triples, and the structure of each triple is (head entity-relation-tail entity).

[0039] As shown in Figure 1 The construction method of the front-end engineering design database for the FPSO project provided by the embodiment includes the following steps:

[0040] Step S1: determining the associated documents of the document data according to the data correlation degree between any two document data of the FEED stage of the FPSO project;

[0041] Step S2: obtaining the index correlation degree of the document data and each associated document thereof according to the data correlation degree of the document data and each associated document thereof and the difference in the number of initial indexes;

[0042] Step S3: sorting the associated documents based on the index correlation degree to obtain the index result display priority of the document data, which is used to indicate the database establishment.

[0043] The following will be described in detail in combination with the drawings.

[0044] Step S1: determining the associated documents of the document data according to the data correlation degree between any two document data of the FEED stage of the FPSO project.

[0045] Since the FEED stage of the FPSO project contains multiple document data, such as oilfield development plan, core process calculation book, equipment specification book, etc., the associated documents of each document data are determined by obtaining the data correlation degree between any two document data, and the associated documents represent other document data that are more closely associated with the document data. In an exemplary embodiment, as shown in FIG. 2, a specific process for obtaining the data correlation degree is given as follows: Figure 2

[0046] Step S11: determining the distribution similarity of the first document data and the second document data in the multiple sub-knowledge graphs according to the constructed knowledge graph of the FEED stage of the FPSO project.

[0047] First, the jieba word segmentation tool is used to perform word segmentation processing on each document data to obtain each word segmentation of each document data.

[0048] For ease of illustration, it is assumed that the first document data and the second document data are any two document data in the FEED stage of the FPSO project.

[0049] The more similar the distribution of the word segmentation in the first document data and the second document data in the multiple sub-knowledge graphs is, the higher the distribution similarity of the first document data and the second document data in the multiple sub-knowledge graphs is. In an exemplary embodiment, as shown in FIG. 3, a specific process for obtaining the distribution similarity is given as follows: Figure 3

[0050] Step S111: determining the first sub-knowledge graph sequence and the second sub-knowledge graph sequence, and the first quantity and the second quantity.

[0051] It should be understood that for any one document data, the document data includes multiple word segmentations, and each word segmentation of the document data may appear in more than one sub-knowledge graph. Then, the sub-knowledge graphs in which the word segmentations of the first document data appear are determined, thereby obtaining the first sub-knowledge graph sequence, and the first sub-knowledge graph sequence is composed of the sub-knowledge graphs in which the word segmentations in the first document data appear. At the same time, the number of sub-knowledge graphs contained in the first sub-knowledge graph sequence is obtained, which is defined as the first quantity. Similarly, the sub-knowledge graphs in which the word segmentations of the second document data appear are determined, thereby obtaining the second sub-knowledge graph sequence, and the second sub-knowledge graph sequence is composed of the sub-knowledge graphs in which the word segmentations in the second document data appear. At the same time, the number of sub-knowledge graphs contained in the second sub-knowledge graph sequence is obtained, which is defined as the second quantity.

[0052] ​​Step S112: Based on the difference between the first quantity and the second quantity, and the sequence similarity between the first sub-knowledge graph sequence and the second sub-knowledge graph sequence, obtain the distribution similarity.

[0053] The difference between the first and second quantities is obtained by calculating the absolute value of the difference between them. The smaller the difference between the first and second quantities, the closer the number of word segments appearing in the first and second document data across multiple sub-knowledge graphs, the more similar the distributions of the first and second document data, and the higher their distribution similarity. Therefore, the distribution similarity between the two is inversely correlated with the difference between the first and second quantities.

[0054] The sequence similarity between the first and second sub-knowledge graph sequences is obtained to represent the similarity of the sub-knowledge graphs appearing in the word segmentation of the first and second document data. In an exemplary embodiment, the sequence similarity between the first and second sub-knowledge graph sequences is the Jaccard correlation coefficient between them. The Jaccard correlation coefficient is the intersection-union ratio (IUR) of the two sequences. Specifically, the intersection and union of the sub-knowledge graphs in the first and second sub-knowledge graph sequences are obtained, and the ratio of the intersection to the union is calculated. The larger the IUR, the more identical sub-knowledge graphs are in the two sequences, and the more similar they are. Therefore, the distributional similarity between the first and second document data is positively correlated with the sequence similarity between the first and second sub-knowledge graph sequences.

[0055] Based on the difference between the first and second quantities, and the sequence similarity between the first and second sub-knowledge graph sequences, the distributional similarity between the first and second document data is obtained. A specific quantification method is given below:

[0056] ;

[0057] in, Indicates the first The document data and the first Distribution similarity of document data Indicates the first The number of sub-knowledge graphs appearing in the word segmentation of a document's data. Indicates the first The number of sub-knowledge graphs appearing in the word segmentation of a document's data. This represents an exponential function with base e. Indicates the first The sequence of sub-knowledge graphs appearing in the word segmentation of a document's data. Indicates the first the sub-knowledge graph sequence appeared by the segmentation of the first document data, the Jaccard correlation coefficient of and . Therefore, the greater the value of , the greater the distribution similarity of the segmentation of the first document data and the second document data in multiple sub-knowledge graphs.

[0058] Step S12: obtaining the content similarity of the first document data and the second document data according to the number of same segmentation in the common appearing sub-knowledge graph of the first document data and the second document data, and the attribute consistency of the same segmentation.

[0059] Since the content distribution of each segmentation in the sub-knowledge graph is not considered when analyzing the distribution similarity of the first document data and the second document data in multiple sub-knowledge graphs, for example, the same segmentation of the first document data and the second document data in the same sub-knowledge graph is respectively in different positions in the triple, for example, one is in the head entity and the other is in the tail entity, which indicates that the sub-knowledge graph where the segmentation of the first document data and the second document data is located is the same, but the content represented by the segmentation is different, that is, the attribute of the same segmentation is inconsistent. Therefore, according to the number of same segmentation in the common appearing sub-knowledge graph of the first document data and the second document data, and the attribute consistency of the same segmentation, the content similarity of the first document data and the second document data is obtained. In an exemplary embodiment, as shown in Figure 4 , a specific process for obtaining the content similarity is given as follows:

[0060] Step S121: determining the number of same segmentation of the first document data and the second document data in the target common sub-knowledge graph, and the attribute consistency of each same segmentation in the target common sub-knowledge graph.

[0061] Set each sub-knowledge graph in which the segmentation of the first document data and the second document data commonly appears as each common sub-knowledge graph, then if at least one same segmentation of the first document data and the second document data exists in a certain sub-knowledge graph, the sub-knowledge graph is the common sub-knowledge graph of the first document data and the second document data, thereby obtaining the multiple common sub-knowledge graphs corresponding to the first document data and the second document data.

[0062] For ease of illustration, set the target common sub-knowledge graph as any one of the sub-knowledge graphs in which the first document data and the second document data commonly appear, that is, the target common sub-knowledge graph is any one of the multiple common sub-knowledge graphs corresponding to the first document data and the second document data. ​

[0063] For the target co-located sub-knowledge graph of the first document data and the second document data, the number of identical segmented words of the first document data and the second document data in the target co-located sub-knowledge graph is determined, and then the attribute consistency of each identical segmented word in the target co-located sub-knowledge graph is obtained. Since the sub-knowledge graph is composed of a plurality of triples, the structure of each triple is (head entity-relation-tail entity), therefore, in an exemplary embodiment, the attribute consistency obtaining process includes: for any one identical segmented word of the first document data and the second document data in the target co-located sub-knowledge graph, if the position of the identical segmented word in the triples in the target co-located sub-knowledge graph is the same, the attribute consistency is a first value; if the position of the identical segmented word in the triples in the target co-located sub-knowledge graph is different, the attribute consistency is a second value. For example: if the position of the identical segmented word in the triples in the target co-located sub-knowledge graph is the same as the head entity, the attribute consistency is the first value; if the position of the identical segmented word in the triples in the target co-located sub-knowledge graph is one head entity and the other tail entity, the attribute consistency is the second value. The first value is greater than the second value, which is used to represent that the attribute consistency corresponding to the same position of the identical segmented word in the triples in the target co-located sub-knowledge graph is greater than the attribute consistency corresponding to the different position of the identical segmented word in the triples in the target co-located sub-knowledge graph, and the attribute consistency corresponding to the same position of the identical segmented word in the triples in the target co-located sub-knowledge graph is higher. In this embodiment, the first value is set to 1, and the second value is set to 0.

[0064] Step S122: According to the identical segmented word quantity feature and the attribute consistency, the content similarity performance of the first document data and the second document data with respect to the target co-located sub-knowledge graph is obtained.

[0065] According to the number of identical segmented words of the first document data and the second document data in the target co-located sub-knowledge graph, the identical segmented word quantity feature is obtained. The identical segmented word quantity feature is used to represent the level of the number of identical segmented words, in an exemplary embodiment, the number of segmented words of the first document data appearing in the target co-located sub-knowledge graph and the number of segmented words of the second document data appearing in the target co-located sub-knowledge graph are obtained, and then the minimum value is determined from the two segmented word quantities as the minimum segmented word quantity. Then, the ratio of the number of identical segmented words of the first document data and the second document data in the target co-located sub-knowledge graph to the minimum segmented word quantity is calculated, and the ratio is taken as the identical segmented word quantity feature. It should be understood that the number of identical segmented words of the first document data and the second document data in the target co-located sub-knowledge graph is less than or equal to the minimum segmented word quantity, and the numerical value range of the identical segmented word quantity feature is 0-1.

[0066] The greater the number of identical word segments in the first and second document data, the more identical word segments there are, and the more similar the content of the first and second document data is to the target coexisting sub-knowledge graph. This indicates a higher content similarity, which is positively correlated with the number of identical word segments. Similarly, the higher the consistency of the attributes of each identical word segment, the more similar the content of the first and second document data is to the target coexisting sub-knowledge graph. This also indicates a higher content similarity, which is positively correlated with the consistency of the attributes of each identical word segment. Therefore, based on the number of identical word segments and attribute consistency, the content similarity between the first and second document data to the target coexisting sub-knowledge graph is obtained. Since the content similarity is related to the position of identical word segments in the triples of the target coexisting sub-knowledge graph, it is also called triple similarity. In an exemplary embodiment, a specific quantification method for content similarity is given below:

[0067] ;

[0068] in, Indicates the first The document data and the first Content similarity representation of document data in the b-th co-occurring sub-knowledge graph Indicates the first The document data and the first The number of identical word segments in the b-th co-occurring sub-knowledge graph for each document. This represents the minimum number of word segments, i.e., the number of segments. The number of word segments appearing in the b-th co-occurrence sub-knowledge graph of document data is related to the number of words in the b-th co-occurrence sub-knowledge graph. The minimum number of word segments appearing in the b-th co-occurring sub-knowledge graph for a given document. This indicates the number of identical word segments. Indicates the first The document data and the first The attribute consistency of the i-th identical word in the b-th co-occurring sub-knowledge graph of document data is determined by the fact that if the i-th identical word is in the same position in the triples of the b-th co-occurring sub-knowledge graph, then... =1, otherwise If it is 0, then, Indicates the first The document data and the first The average value of attribute consistency of the same word segment in the b-th co-occurrence sub-knowledge graph of document data, with a value range of 0-1. The larger the value, the more similar the positions of all identical word segments in the triplet.

[0069] Step S123: fusing the content similarity of the first document data and the second document data with respect to the content similarity of all sub-knowledge graphs that coexist, to obtain the content similarity.

[0070] Through step S122, the content similarity of the first document data and the second document data with respect to each coexisting sub-knowledge graph is obtained, and the content similarity with respect to each coexisting sub-knowledge graph is fused, specifically, the average value of the content similarity with respect to each coexisting sub-knowledge graph is calculated, and the result is the content similarity of the first document data and the second document data with respect to all sub-knowledge graphs that coexist, that is, the content similarity of the first document data and the second document data is obtained, and the calculation formula is as follows:

[0071] ;

[0072] Among them, represents the content similarity of the first document data and the first document data, and B represents the number of coexisting sub-knowledge graphs of the first document data and the first document data.

[0073] Step S13: fusing the distribution similarity and the content similarity to obtain the data correlation degree of the first document data and the second document data.

[0074] For any two document data of the FEED stage of the FPSO project, if the distribution similarity of the two document data in multiple sub-knowledge graphs is greater, it means that the information of the two document data in the knowledge graph is more similar; at the same time, if the content similarity of the two document data is greater, it means that the information expressed by the content of the two document data in the knowledge graph is also more similar. Therefore, by combining the distribution similarity and the content similarity of the two document data in multiple sub-knowledge graphs, the data correlation degree of the two document data of the FEED stage of the FPSO project can be obtained. And the higher the distribution similarity and the content similarity of the two document data, the higher the data correlation degree of the two document data.

[0075] In an exemplary embodiment, the distribution similarity and the content similarity of the first document data and the second document data are weighted and summed, and the weight is 0.5, and the result is the data correlation degree of the first document data and the second document data, and the essence is: calculating the average value of the distribution similarity and the content similarity of the first document data and the second document data, and the result is the data correlation degree of the first document data and the second document data.

[0076] Since each document data corresponds to one or more index records in the data storage process, it is convenient to quickly locate data according to the query requirements. However, when the user queries, there are too many index records for each document data, and the indexes between different document data do not have linkage. If the correlation between two document data is high, the query result cannot accurately find the required data.

[0077] After constructing the database in the FEED stage of the FPSO project by the existing method, multiple indexes of each document data are obtained, and the number of indexes is determined, which is defined as the initial index number of each document data. The more the initial index number is, the more resources the database invests to improve the query efficiency of the document data, and the higher the importance of the document data in the database is.

[0078] For any document data, the data correlation degree between the document data and other document data is obtained through the above process. A threshold is preset, which is used to compare the data correlation degree between the document data and other document data, so as to filter the data correlation degree with a larger value, that is, to obtain other document data that is more associated with the document data. The value range of the preset threshold is 0-1, and the specific value is set according to the actual filtering needs. The larger the value of the preset threshold is, the fewer the other document data that meets the condition is. In this embodiment, 0.6 is taken as an example.

[0079] The data correlation degree between the document data and other document data is compared with the preset threshold, and the other document data corresponding to the data correlation degree greater than the preset threshold is obtained. The other document data corresponding to the data correlation degree greater than the preset threshold is determined as the associated document of the document data. In the subsequent search process of the database, the index correlation degree of the associated document of the document data and the document data is set to be higher.

[0080] Step S2: According to the data correlation degree between the document data and each associated document thereof, and the difference in the initial index number, the index correlation degree between the document data and each associated document thereof is obtained.

[0081] The initial index count of the document data and the initial index count of each associated document of the document data are obtained, thereby obtaining the difference in the initial index count between the document data and each of its associated documents. The difference in the initial index count is the absolute value of the difference in the initial index counts. That is, for any associated document of the document data, the absolute value of the difference between the initial index count of the document data and the initial index count of the associated document is calculated as the difference in the initial index count between the document data and the associated document. The smaller the difference in the initial index count, the closer the initial index counts of the two documents are, and the stronger the index correlation between the two documents. At the same time, the stronger the data correlation between the two documents, the stronger the index correlation between the two documents. Therefore, based on the data correlation degree between the document data and each of its associated documents, and the difference in the initial index count, the index correlation degree between the document data and each of its associated documents is obtained. The index correlation degree is positively correlated with the data correlation degree and inversely correlated with the difference in the initial index count. In an exemplary embodiment, the first... Taking the data of the first document as an example, let the first document be... There are V related documents for each document, and the v-th related document is the... For any associated document of the document data, then the first... The formula for calculating the index correlation between a document and the v-th associated document is as follows:

[0082] ;

[0083] in, Indicates the first The index correlation between each document and the v-th associated document. Indicates the first The degree of correlation between the data of document v and the data of associated document v. Indicates the first The initial number of indexes for each document's data. This represents the initial number of indexes for the v-th associated document. Indicates the first The difference between the initial index count of the vth document and the vth associated document.

[0084] Step S3: Sort the related documents based on the index relevance to obtain the index results display priority of the document data, which is used to indicate the database establishment.

[0085] According to the The degree of correlation between the data of a document and the index of each of its associated documents, for the first document... The associated documents of the document data are sorted. In an exemplary embodiment, the documents are sorted in descending order of index relevance. The document data is sorted according to the index relevance of each associated document, and then the document is sorted according to the index relevance. Sort the associated documents of the first document data to obtain the second document. A sequence of associated documents for each document. In a sequence of associated documents for a given document, the earlier the associated document appears, the closer it is to the document in question. The stronger the index correlation of each document's data. The associated document sequence of the document data is used as the first document data. The index results for each document are displayed with priority; the earlier a related document appears, the higher its priority.

[0086] The index results are displayed with priority to indicate the database creation process. During subsequent database queries, when the user selects the index... When querying data for document number 1, enter the document number 2. When indexing a document, first index the document data. The system displays data from each document. If the user is not satisfied with the search results, they can enter the document again. When indexing a document's data, then follow the first index. The order of the related documents in the sequence of document data will be used to display each related document as a search result to the user (for example, when the user searches a second time, the first two documents in the related document sequence will be displayed as search results). Furthermore, through the above analysis, this invention can improve the query quality during the database construction process by analyzing the degree of data correlation between different document data during the FEED phase of an FPSO project, and then setting corresponding index correlation degrees for different document data, thereby improving the efficiency and quality of FPSO project design.

[0087] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0088] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A method for constructing a front-end engineering design database applicable to an FPSO project, characterized in that, The application relates to a method for establishing a database of FPSO (Floating Production Storage and Offloading) project data, and belongs to the technical field of data processing. According to the data correlation degree between any two document data in the FEED stage of an FPSO project, the associated documents of the document data are determined; the data correlation degree is obtained by determining the distribution similarity of first document data and second document data in a plurality of sub-knowledge graphs according to a constructed knowledge graph of the FEED stage of the FPSO project; the first document data and the second document data are any two document data; the knowledge graph comprises the plurality of sub-knowledge graphs; the content similarity of the first document data and the second document data is obtained according to the number of same words in the common sub-knowledge graphs and the attribute consistency of the same words. The distribution similarity and the content similarity are fused to obtain the data correlation degree of the first document data and the second document data. According to the data correlation degree of the document data and each associated document thereof and the difference between the initial index number of the document data and the initial index number of each associated document thereof, the index correlation degree of the document data and each associated document thereof is obtained. The associated documents are sorted based on the index correlation degree to obtain the index result display priority of the document data, which is used for indicating the database establishment.

2. The method for constructing a front-end engineering design database applicable to an FPSO project according to claim 1, characterized in that, The obtaining process of the distribution similarity comprises the following steps. The first sub-knowledge graph sequence and the second sub-knowledge graph sequence and the first number and the second number are determined; the first sub-knowledge graph sequence is composed of the sub-knowledge graphs in which the words in the first document data appear; the second sub-knowledge graph sequence is composed of the sub-knowledge graphs in which the words in the second document data appear; the first number is the number of the sub-knowledge graphs in the first sub-knowledge graph sequence; and the second number is the number of the sub-knowledge graphs in the second sub-knowledge graph sequence. The distribution similarity is obtained according to the number difference of the first number and the second number and the sequence similarity of the first sub-knowledge graph sequence and the second sub-knowledge graph sequence; the distribution similarity is inversely related to the number difference and is positively related to the sequence similarity.

3. The method for constructing a front-end engineering design database applicable to an FPSO project according to claim 2, characterized in that, The sequence similarity is the Jaccard correlation coefficient of the first sub-knowledge graph sequence and the second sub-knowledge graph sequence.

4. The method for constructing a front-end engineering design database applicable to an FPSO project according to claim 1, characterized in that, The obtaining process of the content similarity comprises the following steps. The number of same words of the first document data and the second document data in a target common sub-knowledge graph and the attribute consistency of the same words in the target common sub-knowledge graph are determined; the target common sub-knowledge graph is any one of the sub-knowledge graphs in which the words of the first document data and the second document data commonly appear. The content similarity performance of the first document data and the second document data for the target common sub-knowledge graph is obtained according to the same word number feature and the attribute consistency; the content similarity performance is positively related to the same word number feature and the attribute consistency; the same word number feature is obtained from the same word number; The content similarity is obtained by fusing the content similarity performances of the first document data and the second document data for all the common sub-knowledge graphs.

5. The method for constructing a FEED database for FPSO projects according to claim 4, wherein, The sub-knowledge graph is composed of multiple triplets; The attribute consistency acquisition process comprises: if the same same part of speech of the first document data and the second document data is in the same position in the triplet in the target co-located sub-knowledge graph, the attribute consistency of the same same part of speech is a first value; if the same same part of speech of the first document data and the second document data is in different positions in the triplet in the target co-located sub-knowledge graph, the attribute consistency of the same same part of speech is a second value; the first value is greater than the second value.

6. The method for constructing a FEED database for FPSO projects according to claim 4, wherein, The acquisition process of the same part of speech quantity feature comprises: determining a minimum part of speech quantity, the minimum part of speech quantity being the minimum value of the number of parts of speech of the first document data appearing in the target co-located sub-knowledge graph and the number of parts of speech of the second document data appearing in the target co-located sub-knowledge graph; the ratio of the same part of speech quantity to the minimum part of speech quantity is taken as the same part of speech quantity feature.

7. The method for constructing a FEED database for FPSO projects according to claim 1, wherein, The method for determining the associated document of the document data comprises: comparing the data association degree of the document data with each of the other document data with a preset threshold, and taking the document data corresponding to the data association degree greater than the preset threshold as the associated document of the document data.

8. The method for constructing a FEED database for FPSO projects according to claim 1, wherein, The sorting of the associated document based on the index association degree comprises: sorting each associated document of the document data in the order of the index association degree from large to small.

9. The method for constructing a front-end engineering design database applicable to an FPSO project according to claim 1, wherein, The database establishment method further comprises: performing part of speech processing on the document data using a jieba part of speech tool to obtain each part of speech of the document data.

Citation Information

Patent Citations

  • Document-based retrieval method and device

    CN113094519A

  • Document retrieval method based on knowledge graph and related equipment thereof

    CN114780746A

Cited By

  • FPSO front-end engineering design database construction method based on multi-dimensional data fusion

    CN122240587A