Cross-modal electronic file retrieval method and device based on common knowledge graph

By constructing a cross-modal electronic file search method based on common knowledge graphs, the complexity of cross-domain power BIM model search is solved, efficient and accurate BIM model search and image matching are achieved, and the design, construction and operation and maintenance efficiency of power engineering is improved.

CN120277231APending Publication Date: 2025-07-08DATANG HUIZHOU THERMAL POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510477995.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to retrieve multiple cross-domain power BIM models at the same time, and it is impossible to effectively integrate multimodal information, resulting in inefficient retrieval efficiency.

Method used

A cross-modal electronic file retrieval method based on common knowledge graphs is constructed, key geometric features are extracted by generating multi-view rendering images, cluster recognition and define geometric words, multi-modal common knowledge graphs are constructed, and embedded learning is used to calculate similarity scores and sort them.

Benefits of technology

It realizes rapid and accurate retrieval of BIM models and related images of power engineering, improves work efficiency, reduces costs, and enhances the generalization ability and adaptability of the model, and improves the accuracy of the search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277231A_ABST
    Figure CN120277231A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal electronic file retrieval method and device based on a common knowledge graph, and relates to the field of three-dimensional shape retrieval, and the method comprises the steps: generating a multi-view rendering image of an electric power BIM model according to a three-dimensional modeling tool, extracting key geometric features of each component, carrying out the clustering recognition, defining a recognition result as geometric words, and carrying out the clustering recognition of the key geometric features; taking geometric words as nodes of the multi-modal common knowledge graph; constructing a multi-modal common knowledge graph containing a plurality of entities based on the geometric words, and defining various types of edges to connect the plurality of entities; performing embedding learning on nodes of the multi-modal common knowledge graph according to the graph convolutional network to generate entity embedding; according to the generated entity embedding, calculating similarity scores between query objects and candidate objects of different modalities, and sorting the similarity scores; and returning the most matched electric power BIM model according to a sorting result. Therefore, the BIM model and related images can be quickly and accurately retrieved, the working efficiency is improved, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of three-dimensional shape retrieval, and in particular to a cross-modal electronic document retrieval method and device based on a common knowledge graph. Background Art

[0002] With the rapid development of the Internet, knowledge graphs are widely used in many fields. The power engineering field is also in urgent need of relevant technologies to achieve efficient retrieval and matching of power BIM models to promote secondary development. To solve the problem of BIM model retrieval in power engineering, many methods have emerged, such as extracting rendering views from power images through multi-view convolutional neural networks (MVCNN) and finally integrating them into shape descriptors; PointNet++ uses point cloud data to create BIM three-dimensional representations; multi-surround CNN architecture considers hierarchical links between views for 3D shape retrieval, etc. However, these traditional methods focus on cross-modal feature learning and global structural information descriptor creation, both of which rely on parameter learning, extensive training data sets and model design, and cannot achieve simultaneous retrieval of multiple cross-domain power BIM models.

[0003] Therefore, there is an urgent need for a method that can simultaneously retrieve multiple cross-domain power BIM models. Summary of the invention

[0004] The present invention provides a cross-modal electronic file retrieval method and device based on a common knowledge graph, aiming to solve the complexity of retrieval of three-dimensional models of BIM of electric power engineering projects.

[0005] According to a first aspect of the present disclosure, a cross-modal electronic document retrieval method based on a common knowledge graph is provided, comprising: Generate multi-view rendering images of the power BIM model using 3D modeling tools, and extract key geometric features of each component using image segmentation technology; Clustering and identifying the key geometric features, defining the identification results as geometric words, and using the geometric words as nodes of a multimodal common knowledge graph; Building a multimodal common knowledge graph containing multiple entities based on the geometric words, and defining multiple types of edges to connect the multiple entities; Performing embedding learning on the nodes of the multimodal common knowledge graph according to a graph convolutional network to generate entity embedding; Based on the generated entity embeddings, the similarity scores between the query objects and candidate objects of different modalities are calculated and ranked; The results are sorted according to the similarity score and the most matching power BIM model is returned.

[0006] In the aspects and any possible implementation manners described above, a further implementation manner is provided. The multiple entities include a model entity, an image entity, a part entity, and a geometric word entity. The multiple types of edges include a binding edge for connecting a three-dimensional shape and a rendered image, a geometric edge for connecting a shape part and a geometric word, and a category edge for representing a category relationship.

[0007] In the aspects and any possible implementation manners described above, a further implementation manner is provided. Generating a multi-view rendered image of a power BIM model according to a three-dimensional modeling tool includes: Making the power BIM model stand upright along a fixed axis by performing PCA on the normal vectors, rotating the camera around the fixed axis at a fixed angle to obtain a multi-view rendered image, where the normal vector is a vector perpendicular to the surface of a three-dimensional object in the power engineering BIM model; or, Constructing a regular dodecahedron with a three-dimensional object in the power engineering BIM model as a reference, and making the center of the regular dodecahedron coincide with the object center, deploying virtual cameras at the vertices of the dodecahedron to obtain a multi-view rendered image.

[0008] In the aspects and any possible implementation manners described above, a further implementation manner is provided. Embedding learning is performed on the nodes of the multi-modal common knowledge graph according to a graph convolutional network to generate entity embeddings, including: Defining the multi-modal common knowledge graph as an undirected weighted graph and defining a node feature matrix, where each node is represented by an N-dimensional feature vector; Learning the embedding vectors of the nodes through GCN; Setting an optimization objective function and minimizing the direct distance between nodes through the optimization objective function so that similar nodes are closer in the embedding space; Enhancing the connections between entities by using category edges and geometric word entities to complete entity embedding.

[0009] In the aspects and any possible implementation manners described above, a further implementation manner is provided. Calculating and sorting the similarity scores between query objects and candidate objects of different modalities according to the generated entity embeddings, including: Extracting the embedding vectors of the shape entity, model entity, part entity, and geometric word entity of the power engineering related to the query object; According to a pre-defined similarity measurement method between different entities, calculating the similarity measurement between each entity of the query object and the candidate object respectively. The similarity measurement methods between different entities include the similarity measurement between the model entity, image entity, part entity, and geometric word entity between the query object and the candidate object; The weighted combination of the similarity measures between the different entities is obtained to get the similarity between the query object and the candidate object.

[0010] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. If the query object is a model, the similarity calculation between the query object and the candidate object is performed using the following formula: ; s.t. ; where, represents the similarity score between the query object and the candidate object, and respectively represent the model entities of the query object and the candidate object, and respectively represent the image entities of the query object and the candidate object, and respectively represent the partial entities of the query object and the candidate object, and respectively represent the geometric word entities of the query object and the candidate object, , , , respectively represent the similarity scores of the model entity, the image entity, the partial entity, and the geometric word entity of the query object and the candidate object, , , and all represent weight values.

[0011] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. If the query object is an image, the similarity calculation between the candidate object and the query object is performed using the following formula: ; s.t. ; where, represents the similarity score between the query object and the candidate object, and respectively represent the model entities of the query object and the candidate object, and respectively represent the partial entities of the query object and the candidate object, and respectively represent the geometric word entities of the query object and the candidate object, , , respectively represent the similarity scores of the model entity, the partial entity, and the geometric word entity of the query object and the candidate object, , and both represent weight values.

[0012] For the aspects and any possible implementation manners described above, a further implementation manner is provided. When calculating the similarity between a query object and partial entities of a candidate object, a bipartite graph matching method is adopted.

[0013] For the aspects and any possible implementation manners described above, a further implementation manner is provided. By using a GCN to learn the embedding vectors of nodes, it further includes: Determine whether it is necessary to introduce category edges; If it is necessary, set category edges in the graph structure to put the model in a supervised mode, and use the category label information corresponding to the category edges as a supervision signal to perform embedding learning on the nodes of the multimodal common knowledge graph. If it is not necessary to introduce category edge information, remove the category edges in the graph structure to put the model in an unsupervised mode. The model performs embedding learning on the nodes through self-organized learning according to the topological structure of the graph and the node feature matrix.

[0014] According to a second aspect of the present disclosure, a cross-modal electronic document retrieval device based on a common knowledge graph is provided, including: An image extraction and segmentation module, configured to generate multi-view rendering images of a power BIM model according to a 3D modeling tool, and extract key geometric features of each component through image segmentation technology; A geometric word definition module, configured to perform clustering recognition on the key geometric features, define the recognition result as a geometric word, and use the geometric word as a node of the multimodal common knowledge graph; A multimodal common knowledge graph construction module, configured to construct a multimodal common knowledge graph including multiple entities based on the geometric words, and define multiple types of edges to connect the multiple entities; A node embedding learning module, configured to perform embedding learning on the nodes of the multimodal common knowledge graph according to a graph convolutional network to generate entity embeddings; A similarity calculation module, configured to calculate similarity scores between query objects and candidate objects of different modalities based on the generated entity embeddings and sort them; A result output module, configured to return the most matching power BIM model according to the sorted result of the similarity scores.

[0015] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.

[0016] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the method as described above.

[0017] Compared with the prior art, the present disclosure achieves the following beneficial effects: (1) The present disclosure innovatively applies the knowledge graph to the retrieval of power engineering BIM models, constructs a multimodal common knowledge graph for power engineering, can effectively integrate multimodal information in power engineering, and successfully solves the complex challenges of cross-modal retrieval. In this way, in the stages of power engineering design, construction, and operation and maintenance, the required BIM models and related images can be quickly and accurately retrieved, improving work efficiency and reducing costs.

[0018] (2) The present disclosure performs representation learning on the entities in the common knowledge graph based on the graph convolutional network (GCN). This strategy can flexibly adapt to supervised and unsupervised conditions, effectively enhancing the consistency of the feature vectors of similar BIM model components, strengthening the association between power engineering images and BIM models, enabling the model to more accurately capture and express the features of different modal data, improving the generalization ability and adaptability of the model, and better coping with the complex and changeable power engineering data.

[0019] (3) The present disclosure introduces a similarity measurement method. During the retrieval process, by comprehensively considering the similarities of various entities and performing weighted fusion, the similarity between the query and the candidate object can be more accurately measured, thereby improving the accuracy of the retrieval results and providing more valuable reference information for power engineering personnel.

[0020] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 shows a flowchart of a cross-modal electronic file retrieval method according to an embodiment of the present disclosure; Figure 2 shows a block diagram of a cross-modal electronic file retrieval device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0023] In addition, the term "and / or" in this document merely describes an association relationship between associated objects and indicates that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.

[0024] Embodiment 1 As Figure 1 shown is a flowchart of a cross-modal electronic document retrieval method 100 based on a common knowledge graph according to the present disclosure. The method includes: Step S110: Generate multi-view rendering images of a power BIM model according to a 3D modeling tool, and extract key geometric features of each component through image segmentation technology.

[0025] In some embodiments, generating multi-view rendering images of a power BIM model according to a 3D modeling tool may include the following two methods. First, make the substation BIM model stand upright along a fixed axis (such as the z-axis), and then place the camera around this axis at a fixed angle and at an angle of 30° to the ground plane and facing the center of the shape to obtain multi-view rendering images within a 360° range. For example, for the BIM 3D model of a power tower, this method can be used. These images show the appearance and spatial layout of components such as transformers, switchgear, and transmission lines in the substation from different angles.

[0026] Second, sample multiple views from the real space. Specifically, construct a dodecahedron based on the 3D objects in the power engineering BIM model, and make the center of the dodecahedron coincide with the center of the object. Twenty virtual cameras are deployed at the vertices of the dodecahedron shape to obtain multi-view rendering images. For example, for photos of substations at the actual power engineering site, this method can be used.

[0027] In some embodiments, after obtaining the multi-view rendered images, advanced image segmentation techniques can be used to process the obtained multi-view rendered images. For example, through deep learning algorithms, each component in the image is precisely segmented to extract key geometric features such as the circular contour of the transformer, the rectangular structure of the switchgear cabinet, and the linear features of the transmission line, and these features are recorded in digital form, such as coordinates, shape parameters, etc.

[0028] Step S120: Cluster and identify the key geometric features, define the identification results as geometric words, and use the geometric words as nodes of the multi-modal common knowledge graph.

[0029] In some embodiments, the K-means clustering algorithm is used to cluster and identify the extracted key geometric features to determine the labels and descriptors of the geometric words. For example, assume that after clustering analysis, the circular contour feature of the transformer is clustered into the geometric word "circular - large - electrical equipment", and the rectangular structure of the switchgear cabinet is clustered into the geometric word "rectangular - metal - electrical cabinet body", etc. These geometric words, as nodes of the multi-modal common knowledge graph, effectively summarize the key shape features of the power components.

[0030] Step S130: Construct a multi-modal common knowledge graph containing multiple entities based on the geometric words, and define multiple types of edges to connect the multiple entities.

[0031] In some embodiments, the entity types include model entities, image entities, part entities, and geometric word entities, where, One is the model entity. The model entity is the entire power engineering BIM model. Taking the power transformer as an example, PointNet++ is used to extract its feature vectors, and the extracted feature vectors are integrated into the knowledge graph. At the same time, in order to ensure the unified dimension of the power engineering model descriptors, principal component analysis (PCA) is needed to reduce the dimension of the feature vectors.

[0032] The second is the image entity. This entity is derived from the multi-view rendered images of the power BIM model collected in the above step S110 or the photos taken of the actual substation, and it can further represent the real images in the knowledge graph for cross-modal BIM model retrieval.

[0033] The third is the part entity. According to the rendered or real images, a model trained on the dataset is used to segment each power engineering image (such as an image of a power line tower) into multiple parts. The segmented parts are used as part entities to represent the attributes of each shape in the knowledge graph.

[0034] The fourth is the geometric word entity. The geometric word entity acts as a key element in the power common knowledge graph to bridge the gaps between different fields or modalities.

[0035] Furthermore, construct the edges of the knowledge graph, define the edges according to different types, so as to ensure that the representation of the edges for the 3D shape meets the requirements of entity learning. The definition of the edges is divided into three categories: One is the gap edge. The gap edge establishes connections between the 3D model of the power engineering BIM, its rendered image, the rendered image and the segmented part, and the real image and the segmented part. This kind of connection can reflect the geometric and visual features of the power engineering BIM model.

[0036] The second is the geometric edge: The geometric edge connects part of the entity to its respective geometric word entity, enabling the two types of entities to establish corresponding connections and reflecting the geometric structure of the power engineering components.

[0037] The third is the category edge: The category edge is used to capture the relationships between categories, and such edges represent prior knowledge. The category edge enables our method to function in a supervised manner.

[0038] The above three types of edges together construct the edges of the power common knowledge graph.

[0039] According to the above definition, construct a multi-modal common knowledge graph. Present each entity in the form of a node, and connect them according to the definition of the edges to form a complex network structure. For example, connect the model entity of the transformer to the corresponding rendered image through the binding edge, then connect the part entity of the transformer to the geometric word entity of "circular - large - electrical equipment" through the geometric edge, and establish a relationship with other electrical equipment categories through the category edge, so as to comprehensively display the information of the substation components and their mutual relationships.

[0040] In summary, the common knowledge graph established through the above process can comprehensively capture the geometric structure and classification information of the power BIM model.

[0041] Step S140: Perform embedding learning on the nodes of the multi-modal common knowledge graph according to the graph convolutional network to generate entity embeddings.

[0042] In some embodiments, although the 3D shape knowledge graph can capture the structure of the 3D model of the power engineering BIM and directly reveal the correlation between visual information and the 3D shape. However, it does not completely solve the problem of weak correlation in space between real power images and power BIM model components. Node (graph) embedding, that is, by defining the problem and the objective function, learning node embeddings through GCN, enhancing the consistency of the feature vectors of similar BIM models, and establishing connections between real power images and power BIM model components, can solve the above problems.

[0043] In some embodiments, first define the multi-modal common knowledge graph: Define the multi-modal common knowledge graph as an undirected weighted graph, , where represents nodes, represents edges. Furthermore, divide the node set into five different parts: . Among them, represents the model entity set, represents the set of rendered image entities, represents the set of real image entities, represents the set of partial shape entities, represents the set of geometric word entities. The feature matrix of nodes is defined as , where each node is represented by an N-dimensional feature vector, which reflects the characteristics of the node in the power engineering. represents the number of entities in the knowledge graph. Therefore, define the last node in the common knowledge graph as to ensure that each node has an E-dimensional embedding.

[0044] Secondly, determine the optimization objective. The optimization objective aims to measure the connection relationship between nodes and is achieved by minimizing an objective function. Specifically, the objective function is defined as (1) where represents that there is a direct edge between nodes and in the knowledge graph, indicates that there is no edge between the two nodes.

[0045] (2) (3) In the above equation (2), defines the probability that there is an edge between nodes and node . In equation (3), defines the probability that there is no edge between nodes and node . Among them, represents 's embedding vector, represents 's embedding vector, is the sigmoid function.

[0046] Furthermore, based on GCN, conduct entity embedding learning. GCN plays an irreplaceable role in embedding the common knowledge graph, especially in enhancing the representation learning process of shape entities and image entities. The traditional GCN structure for learning node embeddings is defined as: (4) Among them, is the final entity feature matrix in graph G, is the index of the domain sharing layer, A is the original adjacency matrix, and the graph adjacency matrix is calculated as follows: (5) The learning weight of the l-th layer of GCN is , and the embedding is generated by combining all entity embeddings.

[0047] The final objective optimization function is defined as follows: (6) is the entity set in graph G, and it has a clear method to reach node , is the node set that has no access to node .

[0048] After determining the objective function, use the classical backpropagation algorithm for optimization to minimize the direct distance between nodes, so as to make similar nodes closer in the embedding space. For example, for transformer model entities and image entities of the same type, the distance in the embedding space will gradually shrink through optimization. At the same time, use category edges and geometric word entities to enhance the connection between entities and complete entity embedding.

[0049] Step S150: Calculate and sort the similarity scores between query objects and candidate objects of different modalities according to the generated entity embeddings.

[0050] When retrieving a specific type of transmission line tower BIM model, extract the entity embedding vectors related to the query object (model or image). According to different entity similarity measurement methods, calculate the similarity between the query object and different entities of the candidate object respectively, such as calculating the similarity of model entities, image entities, part entities, and geometric word entities, and perform weighted combination on the similarity measurements between the different entities to obtain the similarity score between the query object and the candidate object. Specifically, If the query object is the power equipment BIM three-dimensional model Q, extract the entities directly related to it (including image entities, part entities, geometric word entities, etc.), and these entities can be represented by a series of embeddings. Specifically, it can be represented by to represent the image entity set, to represent the part entity set, Represents a geometric word entity set, with f representing the embedding of each entity. The goal is to calculate how similar Q and candidate object M are.

[0051] Specifically: For any two entities and , we use the following equation (7) to define the similarity between them: (7) where is the classical cosine distance, whose range is , and it is normalized to the interval [0,1] through the above equation.

[0052] First, it is necessary to calculate the similarity between the embedding vector of the query model Q and the embedding vector of the candidate shape , where , and substituting into the above formula (7) can calculate it.

[0053] Second, it is necessary to calculate the similarity of the image entity set . Based on the extracted embedding vectors, use the cosine distance formula to calculate the similarity between the corresponding image entities in and . Then organize these similarity values into an n×m similarity matrix , = , and the value range is between . The closer the value is to 1, the more similar the features of the two image entities are; the closer the value is to -1, the greater the difference. Then take the average of all elements of the matrix Sim as the value of , specifically the calculation formula is, (8) Third, it is necessary to calculate the similarity of the partial entity set . The bipartite graph matching method can be used to measure the similarity of different partial entity sets. Regard the partial entities in and as the two vertex sets of the bipartite graph. By finding the optimal matching, the matching score is obtained. The matching score is the sum of two pairs. Finally, the result is normalized by combining the average value.

[0054] Fourth, it is necessary to calculate the similarity of the geometric word entity set : (9) where and They are the geometric word sets of the candidate shape m and the query model Q respectively. In this way, the similarity degree between the two geometric word entity sets is measured.

[0055] Finally, the similarity calculation between the query object Q and the candidate object M adopts the following formula (10): (10) s.t. ; Among them, represents the similarity score between the query object and the candidate object, and represent the model entities of the query object and the candidate object respectively, and represent the image entities of the query object and the candidate object respectively, and represent the partial entities of the query object and the candidate object respectively, and represent the geometric word entities of the query object and the candidate object respectively, , , , represent the similarity scores of the model entity, image entity, partial entity and geometric word entity of the query object and the candidate object respectively, , , and all represent weight values.

[0056] If the query object is the image I, when calculating the comprehensive similarity between the query image I and the candidate object M, the process is the same as calculating the comprehensive similarity between the query model Q and the candidate shape M. The specific process will not be elaborated. Finally, the similarity calculation between the query object I and the candidate object M adopts the following formula (11): (11) s.t. ; Among them, represents the similarity score between the query object and the candidate object, and represent the model entities of the query object and the candidate object respectively, and represent the partial entities of the query object and the candidate object respectively, and represent the geometric word entities of the query object and the candidate object respectively, , , respectively represent the similarity scores of the model entity, partial entity, and geometric word entity between the query object and the candidate object, 、 and both represent weight values.

[0057] Step S160: Return the most matching power BIM model according to the sorting result of the similarity scores.

[0058] In some embodiments, according to the sorting result of the similarity scores, return the most matching transmission line tower BIM model. In practical applications, power engineering construction personnel can use these retrieval results to quickly find the transmission line tower models that meet the construction requirements, improve construction efficiency, and reduce design and construction costs.

[0059] According to the above embodiments of the present disclosure, applying the knowledge graph paradigm to the retrieval of power engineering BIM models and constructing a common knowledge graph for power engineering can effectively integrate multi-modal information in power engineering, successfully solve the complex challenges of cross-domain and cross-modal retrieval. In this way, in the stages of power engineering design, construction, and operation and maintenance, the required BIM models and related images can be quickly and accurately retrieved, improving work efficiency and reducing costs. In the process of constructing the common knowledge graph for power engineering, the present disclosure uses image segmentation technology and dataset training models to obtain shape parts and geometric words, and bridges the gap between different modalities and domains through geometric words; uses the edge definition method to restrict the common knowledge graph for power engineering, thereby optimizing the structure and geometric features of the knowledge graph; uses graph convolutional neural networks for node embedding operations to complete the construction of the common knowledge graph for power engineering, effectively solving the problem of weak correlation in space between real power engineering images and BIM three-dimensional model components. Perform representation learning on the entities in the common knowledge graph based on the graph convolutional network (GCN). This strategy can flexibly adapt to supervised and unsupervised conditions, effectively enhance the consistency of the feature vectors of similar BIM model components, strengthen the association between power engineering images and BIM models, enable the model to more accurately capture and express the features of different modality data, improve the generalization ability and adaptability of the model, and better cope with the complex and changeable power engineering data. By introducing a similarity measurement method, in the retrieval process, by comprehensively considering the similarity of multiple entities and performing weighted fusion, the similarity between the query and the candidate object can be measured more accurately, thereby improving the accuracy of the retrieval results and providing more valuable reference information for power engineering personnel.

[0060] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.

[0061] The above is the introduction to the method embodiments. The following further illustrates the solution of the present disclosure through device embodiments.

[0062] Figure 2 The block diagram of a cross-modal electronic document retrieval device 200 based on a common knowledge graph according to an embodiment of the present disclosure is shown. As Figure 2 shown, the device 200 includes: An image extraction and segmentation module 210, configured to generate multi-view rendering images of a power BIM model according to a 3D modeling tool, and extract key geometric features of each component through image segmentation technology; A geometric word definition module 220, configured to perform clustering recognition on the key geometric features, define the recognition result as a geometric word, and use the geometric word as a node of the multi-modal common knowledge graph; A multi-modal common knowledge graph construction module 230, configured to construct a multi-modal common knowledge graph including multiple entities based on the geometric words, and define multiple types of edges to connect multiple entities; A node embedding learning module 240, configured to perform embedding learning on the nodes of the multi-modal common knowledge graph according to a graph convolutional network to generate entity embeddings; A similarity calculation module 250, configured to calculate and sort the similarity scores between query objects and candidate objects of different modalities according to the generated entity embeddings; A result output module 260, configured to return the most matching power BIM model according to the sorting result of the similarity scores.

[0063] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0064] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0065] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitations are imposed herein.

[0066] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A cross-modal electronic document retrieval method based on a common knowledge graph, characterized in that Including: Generating multi - perspective rendering images of the power BIM model according to a 3D modeling tool, and extracting key geometric features of each component through image segmentation technology; Performing clustering recognition on the key geometric features, defining the recognition result as a geometric word, and using the geometric word as a node of the multi - modal common knowledge graph; Constructing a multi - modal common knowledge graph containing multiple entities based on the geometric words, and defining multiple types of edges to connect the multiple entities; Performing embedding learning on the nodes of the multi - modal common knowledge graph according to the graph convolutional network to generate entity embeddings; Calculating and sorting the similarity scores between query objects and candidate objects of different modalities according to the generated entity embeddings; Returning the most matching power BIM model according to the sorted similarity score results.

2. The method according to claim 1, characterized in that, The multiple entities include model entities, image entities, part entities, and geometric word entities, and the multiple types of edges include binding edges for connecting 3D shapes and rendering images, geometric edges for connecting shape parts and geometric words, and category edges for representing category relationships.

3. The method according to claim 1, wherein The generating multi - perspective rendering images of the power BIM model according to a 3D modeling tool includes: Making the power BIM model stand upright along a fixed axis by performing PCA on the normal vectors, rotating the camera around the fixed axis at a fixed angle to obtain multi - perspective rendering images; where the normal vector is a vector perpendicular to the surface of the 3D object in the power engineering BIM model, or, Constructing a regular dodecahedron with the 3D object in the power engineering BIM model as a reference, and making the center of the regular dodecahedron coincide with the object center, deploying virtual cameras at the vertices of the dodecahedron to obtain multi - perspective rendering images.

4. The method according to claim 1, characterized in that, Performing embedding learning on the nodes of the multi - modal common knowledge graph according to the graph convolutional network to generate entity embeddings, including: Defining the multi - modal common knowledge graph as an undirected weighted graph, and defining a node feature matrix, where each node is represented by an N - dimensional feature vector; Learning the embedding vectors of the nodes through GCN; Setting an optimization objective function, and minimizing the direct distance between nodes through the optimization objective function so that similar nodes are closer in the embedding space; Enhancing the connection between entities by using category edges and geometric word entities to complete entity embedding.

5. The method according to claim 1, wherein The calculating and sorting the similarity scores between query objects and candidate objects of different modalities according to the generated entity embeddings includes: Extracting the embedding vectors of the shape entities, model entities, part entities, and geometric word entities of the power engineering related to the query object; Calculating the similarity measures between each entity of the query object and the candidate object respectively according to the pre - defined similarity measurement methods between different entities, where the similarity measurement methods between different entities include the similarity measurements between the model entity, image entity, part entity, and geometric word entity between the query object and the candidate object; Weightedly combining the similarity measures between the different entities to obtain the similarity between the query object and the candidate object.

6. The method according to claim 5, wherein If the query object is a model, the similarity calculation between the query object and the candidate object adopts the following formula: ; s.t. ; Among them, represents the similarity score between the query object and the candidate object, and respectively represent the model entities of the query object and the candidate object, and respectively represent the image entities of the query object and the candidate object, and respectively represent the part entities of the query object and the candidate object, and respectively represent the geometric word entities of the query object and the candidate object, , , , respectively represent the similarity scores of the model entity, image entity, part entity, and geometric word entity between the query object and the candidate object, , , and all represent weight values.

7. The method according to claim 5, wherein If the query object is an image, the similarity between the candidate object and the query object is calculated using the following formula: ; s.t. ; Among them, represents the similarity score between the query object and the candidate object, and respectively represent the model entities of the query object and the candidate object, and respectively represent the partial entities of the query object and the candidate object, and respectively represent the geometric word entities of the query object and the candidate object, 、 、 respectively represent the similarity scores of the model entity, partial entity and geometric word entity of the query object and the candidate object, 、 and all represent weight values.

8. The method according to any one of claims 6 or 7, characterized in that When calculating the similarity between the partial entities of the query object and the candidate object, the bipartite graph matching method is adopted.

9. The method according to claim 4, wherein Learning the embedding vectors of nodes through GCN further includes: Judging whether it is necessary to introduce category edges; If necessary, set category edges in the graph structure to make the model in the supervised mode, and use the category label information corresponding to the category edges as the supervision signal to perform embedding learning on the nodes of the multimodal common knowledge graph. If it is not necessary to introduce category edge information, remove the category edges in the graph structure to make the model in the unsupervised mode. The model embeds the nodes through self-organization learning based on the topological structure of the graph and the node feature matrix.

10. A cross-modal electronic document retrieval device based on a common knowledge graph, characterized in that, Including: An image extraction and segmentation module, configured to generate multi-view rendering images of the power BIM model according to a three-dimensional modeling tool, and extract key geometric features of each component through image segmentation technology; A geometric word definition module, configured to perform clustering recognition on the key geometric features, define the recognition results as geometric words, and use the geometric words as nodes of the multimodal common knowledge graph; A multimodal common knowledge graph construction module, configured to construct a multimodal common knowledge graph containing multiple entities based on the geometric words, and define multiple types of edges to connect the multiple entities; A node embedding learning module, configured to perform embedding learning on the nodes of the multimodal common knowledge graph according to a graph convolutional network to generate entity embeddings; A similarity calculation module, configured to calculate and sort the similarity scores between query objects and candidate objects of different modalities according to the generated entity embeddings; A result output module, configured to return the most matching power BIM model according to the sorting result of the similarity scores.