An object recognition method, apparatus, electronic device, and storage medium

By constructing a semantic graph and extracting image features from the text data of the objects to be identified, and then integrating it with a pre-set knowledge structure graph, the accuracy problem of early-stage hidden diseases is solved, and the effectiveness and accuracy of object identification are improved.

CN115410062BActive Publication Date: 2026-04-03CETHIK GRP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to identify latent diseases in their early stages, leading to reduced accuracy and effectiveness in object identification.

Method used

By constructing a semantic graph from the text data of the object to be identified, extracting image feature information, and fusing it with a pre-set knowledge structure graph, the object recognition result is determined by combining the semantic structure graph and image feature information.

Benefits of technology

It improves the accuracy and effectiveness of object identification, especially in identifying hidden diseases, enhancing the effectiveness of information fusion and the accuracy of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410062B_ABST
    Figure CN115410062B_ABST
Patent Text Reader

Abstract

This application discloses an object recognition method, apparatus, electronic device, and storage medium. The method includes: constructing a semantic graph from text data corresponding to the object to be recognized to obtain a semantic structure graph; extracting features from image data corresponding to the object to be recognized to obtain image feature information; fusing information from the semantic structure graph with a preset knowledge structure graph to obtain a first fusion result; and fusing information from the image feature information with the preset knowledge structure graph to obtain a second fusion result. Based on the first and second fusion results, determining the object recognition result corresponding to the object to be recognized from the knowledge structure graph. This method can combine the fusion results of semantic information and knowledge information, as well as the fusion results of image information and knowledge information, to recognize the object, thereby improving the accuracy and effectiveness of object recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an object recognition method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the advent of the era of medical big data, knowledge interconnection has received widespread attention. How to extract useful medical knowledge from massive amounts of medical data is key to medical diagnostic big data analysis.

[0003] In existing technologies, some diseases are somewhat concealed in their early stages, making it difficult to obtain corresponding medical knowledge through existing big data analysis methods. This hinders the early identification of concealed diseases and reduces the accuracy and effectiveness of object identification. Summary of the Invention

[0004] This application provides an object recognition method, apparatus, electronic device, and storage medium that can improve the accuracy and effectiveness of object recognition.

[0005] On the one hand, this application provides an object recognition method, the method comprising:

[0006] A semantic graph is constructed from the text data corresponding to the object to be identified to obtain a semantic structure graph. The semantic structure graph represents the first feature information of the object to be identified and the association information with the first feature description information.

[0007] Feature extraction is performed on the image data corresponding to the object to be identified to obtain image feature information;

[0008] Based on the semantic structure graph, information is fused with a preset knowledge structure graph to obtain a first fusion result; the knowledge structure graph represents the second feature information corresponding to multiple identified objects, and the association information with the second feature description information.

[0009] Based on the image feature information, information fusion is performed with the knowledge structure graph to obtain a second fusion result;

[0010] Based on the first fusion result and the second fusion result, the object recognition result corresponding to the object to be identified is determined from the knowledge structure graph.

[0011] In an optional embodiment, the step of fusing information based on the semantic structure graph and a preset knowledge structure graph to obtain a first fusion result includes:

[0012] Determine the association structure graph that is associated with the semantic structure graph from the knowledge structure graph;

[0013] Information fusion is performed on the semantic structure graph and the association structure graph to obtain the first fusion result;

[0014] The step of fusing information based on the image feature information and the knowledge structure graph to obtain the second fusion result includes:

[0015] The image feature information and the associated structure diagram are fused to obtain the second fusion result.

[0016] In an optional embodiment, the information fusion of the semantic structure graph and the association structure graph to obtain the first fusion result includes:

[0017] Determine the degree of correlation between semantic nodes in the semantic structure graph and knowledge nodes in the association structure graph to obtain the first fusion weight;

[0018] Based on the first fusion weight, the semantic nodes in the semantic structure graph are subjected to weighted fusion processing to obtain the first fusion result.

[0019] In an optional embodiment, the step of fusing the image feature information and the associated structure map to obtain the second fusion result includes:

[0020] Determine the correlation between the image feature information and the knowledge nodes in the association structure graph to obtain the second fusion weight;

[0021] Based on the second fusion weight, the image feature information is subjected to weighted fusion processing to obtain the second fusion result.

[0022] In an optional embodiment, before fusing information between the semantic structure graph and the association structure graph to obtain the first fusion result, the method further includes:

[0023] Information enhancement is performed on the semantic structure graph, the association structure graph, and the image feature information to obtain semantic graph enhancement information corresponding to the semantic structure graph, association graph enhancement information corresponding to the association structure graph, and image enhancement information corresponding to the image feature information.

[0024] The step of fusing information from the semantic structure graph and the association structure graph to obtain the first fusion result includes:

[0025] The semantic graph enhancement information and the association graph enhancement information are fused to obtain the first fusion result;

[0026] The step of fusing the image feature information and the associated structure map to obtain the second fusion result includes:

[0027] The image enhancement information and the correlation graph enhancement information are fused to obtain the second fusion result.

[0028] In an optional embodiment, the step of performing information enhancement on the semantic structure graph, the association structure graph, and the image feature information respectively to obtain semantic graph enhancement information corresponding to the semantic structure graph, association graph enhancement information corresponding to the association structure graph, and image enhancement information corresponding to the image feature information includes:

[0029] Determine the first correlation between the semantic nodes in the semantic structure graph and the object to be identified;

[0030] Based on the first relevance, the semantic nodes are updated to obtain the semantic graph enhancement information;

[0031] Determine the second relevance between the knowledge nodes in the association structure graph and the object to be identified;

[0032] Based on the second relevance, the knowledge nodes are updated to obtain enhanced association graph structure information;

[0033] Determine the third correlation between the image feature information and the object to be identified;

[0034] Based on the third correlation, the image feature information is updated to obtain image enhancement information.

[0035] In an optional embodiment, the first relevance includes semantic node weights and semantic association weights, and determining the first relevance between the semantic nodes in the semantic structure graph and the object to be identified includes:

[0036] Attention processing is performed on each semantic node to obtain the semantic node weight corresponding to each semantic node;

[0037] Attention processing is performed on the associated semantic nodes of each semantic node to obtain the semantic association weight corresponding to each semantic node;

[0038] The step of updating the semantic nodes based on the first relevance to obtain the semantic graph enhancement information includes:

[0039] Based on the semantic association weights, the associated semantic nodes are subjected to weighted fusion processing to obtain the first node update information corresponding to each semantic node;

[0040] Based on the first node update information and the semantic node weights, node update processing is performed on each semantic node to obtain the semantic graph enhancement information.

[0041] In an optional embodiment, the second relevance includes knowledge node weight and knowledge association weight, and determining the second relevance between the knowledge node in the association structure graph and the object to be identified includes:

[0042] Attention processing is performed on each knowledge node to obtain the knowledge node weight corresponding to each knowledge node;

[0043] Attention processing is performed on the associated knowledge nodes of each knowledge node to obtain the knowledge association weight corresponding to each knowledge node.

[0044] The step of updating the knowledge nodes based on the second relevance to obtain enhanced association graph structure information includes:

[0045] Based on the knowledge association weights, the associated knowledge nodes are weighted and fused to obtain the second node update information corresponding to each knowledge node;

[0046] Based on the update information of the second node and the weight of the knowledge node, each knowledge node is updated to obtain the knowledge graph enhancement information.

[0047] In an optional embodiment, determining the third correlation between the image feature information and the object to be identified includes:

[0048] Attention processing is performed on the image feature information to obtain the third relevance.

[0049] The process of updating the image feature information based on the third correlation to obtain image enhancement information includes:

[0050] Based on the third relevance, the image feature information is updated to obtain the image enhancement information.

[0051] On the other hand, an object recognition device is provided, the device comprising:

[0052] The semantic graph construction module is used to construct a semantic graph from the text data corresponding to the object to be identified, thereby obtaining a semantic structure graph. The semantic structure graph represents the first feature information of the object to be identified and the association information between the first feature description information and the object to be identified.

[0053] The image feature extraction module is used to extract features from the image data corresponding to the object to be identified, and obtain image feature information.

[0054] The first information fusion module is used to perform information fusion with the preset knowledge structure graph based on the semantic structure graph to obtain a first fusion result; the knowledge structure graph represents the second feature information corresponding to multiple identified objects, and the association information with the second feature description information.

[0055] The second information fusion module is used to perform information fusion with the knowledge structure graph based on the image feature information to obtain a second fusion result;

[0056] An object recognition module is used to determine the object recognition result corresponding to the object to be recognized from the knowledge structure graph based on the first fusion result and the second fusion result.

[0057] On the other hand, an electronic device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the object recognition method described above.

[0058] On the other hand, a computer-readable storage medium is provided, the storage medium including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement an object recognition method as described above.

[0059] On the other hand, a computer program product is provided, including a computer program that, when executed by a processor, implements the object recognition method described above.

[0060] This application provides an object recognition method, apparatus, electronic device, and storage medium. The method includes: constructing a semantic graph from text data corresponding to the object to be recognized to obtain a semantic structure graph; extracting features from image data corresponding to the object to be recognized to obtain image feature information; fusing information from the semantic structure graph with a preset knowledge structure graph to obtain a first fusion result; and fusing information from the image feature information with the preset knowledge structure graph to obtain a second fusion result. Based on the first and second fusion results, determining the object recognition result corresponding to the object to be recognized from the knowledge structure graph. This method can combine the fusion results of semantic information and knowledge information, as well as the fusion results of image information and knowledge information, to recognize the object, thereby improving the accuracy and effectiveness of object recognition. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a schematic diagram illustrating an application scenario of an object recognition method provided in an embodiment of this application;

[0063] Figure 2 A flowchart illustrating an object recognition method provided in an embodiment of this application;

[0064] Figure 3 This is a flowchart illustrating information fusion after obtaining an association structure graph in an object recognition method provided in this application embodiment;

[0065] Figure 4 This is a flowchart illustrating the process of obtaining a first fusion result in an object recognition method provided in an embodiment of this application.

[0066] Figure 5 This is a schematic diagram illustrating the intermodal fusion of semantic structure graph, association structure graph, and image feature information in an object recognition method provided in this application embodiment;

[0067] Figure 6 A flowchart illustrating the second fusion result obtained in an object recognition method provided in this application embodiment;

[0068] Figure 7 A flowchart illustrating information enhancement followed by information fusion in an object recognition method provided in this application embodiment;

[0069] Figure 8 A flowchart illustrating the information enhancement of semantic structure graph, association structure graph, and image feature information in an object recognition method provided in this application embodiment;

[0070] Figure 9 This is a flowchart illustrating information enhancement of a semantic structure graph in an object recognition method provided in an embodiment of this application;

[0071] Figure 10 This is a schematic diagram illustrating node updates in a semantic structure graph in an object recognition method provided in an embodiment of this application.

[0072] Figure 11 This is a flowchart illustrating information enhancement of an association structure graph in an object recognition method provided in an embodiment of this application;

[0073] Figure 12This is a schematic diagram illustrating node updates in an association structure graph within an object recognition method provided in an embodiment of this application.

[0074] Figure 13 This is a schematic diagram illustrating feature updating of image feature information in an object recognition method provided in an embodiment of this application;

[0075] Figure 14 This is a schematic diagram illustrating the application of an object recognition method to identify ophthalmic diseases or treatment plans for ophthalmic diseases, as provided in an embodiment of this application.

[0076] Figure 15 This is a schematic diagram of the structure of an object recognition device provided in an embodiment of this application;

[0077] Figure 16 This is a schematic diagram of the hardware structure of a device for implementing the method provided in the embodiments of this application. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0079] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Furthermore, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.

[0080] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0081] First, the relevant terms involved in the embodiments of this application are explained as follows:

[0082] BERT algorithm: short for Bidirectional Encoder Representation from Transformers, is a natural language processing framework proposed by Google that can pre-train deep bidirectional representations from unlabeled text by conditional computation common to both left and right contexts.

[0083] Please see Figure 1 This illustration shows an application scenario diagram of an object recognition method provided in this application embodiment. The application scenario includes a client 110 and a server 120. The client 110 transmits user-input text data and image data to the server 120. The server 120 constructs a semantic graph from the text data corresponding to the object to be recognized, obtaining a semantic structure graph, and extracts features from the image data corresponding to the object to be recognized, obtaining image feature information. Based on the semantic structure graph, the server 120 performs information fusion with a preset knowledge structure graph to obtain a first fusion result, and then performs information fusion with the knowledge structure graph based on the image feature information to obtain a second fusion result. Based on the first and second fusion results, the server 120 determines the object recognition result corresponding to the object to be recognized from the knowledge structure graph. The server 120 transmits the object recognition result to the client 110.

[0084] In this embodiment, the client 110 includes physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, and smart wearable devices, and may also include software running on the physical device, such as applications. The operating system running on the physical device in this embodiment may include, but is not limited to, Android, iOS, Linux, Unix, and Windows.

[0085] In this embodiment, server 120 may include a standalone server, a distributed server, or a server cluster consisting of multiple servers. Server 120 may include a network communication unit, a processor, and memory, etc.

[0086] Please see Figure 2 It demonstrates an object recognition method applicable to the server side, the method comprising:

[0087] S210. Construct a semantic graph from the text data corresponding to the object to be identified to obtain a semantic structure graph. The semantic structure graph represents the first feature information of the object to be identified and the association information with the first feature description information.

[0088] In one optional embodiment, the object to be identified is an object with concealment or an object processing method related to an object with concealment. Examples include chronic eye diseases with insidious onset such as glaucoma, high myopia, diabetic retinopathy, age-related macular degeneration, and thyroid-associated ophthalmopathy, as well as corresponding treatment plans.

[0089] By extracting text features from the text data corresponding to the object to be identified, and combining this with the positional information of each word in the text data, we can obtain first feature information, first feature description information, and the association information between the first feature information and the first feature description information. The positional information of each word indicates its position in a paragraph or sentence. The association information can include the relationship between the first feature information and the relationship between the first feature information and the first feature description information. The first feature description information can be the information in the text data that describes the first feature information.

[0090] By using the first feature information and the first feature description information as semantic nodes, and the relationships between the first feature information and the relationships between the first feature information and the first feature description information as edges, a semantic structure graph can be constructed.

[0091] In scenarios involving disease-assisted identification, the object to be identified can be a disease type, treatment plan, etc., while the text data can be medical record information, including name, gender, age, past medical history, drug allergy history, characteristics, chief complaint, auxiliary examination results, biochemical indicators, and present illness history. The first feature information can be the symptoms of the disease to be identified, and the first feature description information can be the duration, cause, and other symptom description information of the symptom. For example, if the text description is "headache lasting two hours," then the text feature extraction result corresponding to "headache" is the first feature information, and the text feature extraction result corresponding to "two hours" is the first feature description information. The association information between "headache" and "two hours" is the symptom duration. Or, if the text description is "headache accompanied by dizziness," then the text feature extraction results corresponding to "headache" and "dizziness" are both first feature information, and the association information between "headache" and "dizziness" is the accompanying symptom.

[0092] In the text feature extraction process, the BERT algorithm can be used to embed word vectors into the text data. Combined with the position of each word in the paragraph or sentence, a word vector group is obtained. This word vector group includes the first feature information, the first feature description information, and the association information between the first feature information and the first feature description information.

[0093] The text data corresponding to the object to be identified can be input into a preset semantic graph construction model for semantic graph construction. The semantic graph construction model can include a text feature extraction layer and a semantic graph generation layer. The text data corresponding to the object to be identified is input into the text feature extraction layer for text feature extraction, obtaining first feature information, first feature description information, and the association information between the first feature information and the first feature description information. The first feature information, the first feature description information, and the corresponding association information are input into the semantic graph generation layer for semantic graph generation, which can obtain a semantic structure graph.

[0094] S220. Extract features from the image data corresponding to the object to be identified to obtain image feature information;

[0095] In an optional embodiment, the image data can be multi-source image data, which refers to various image data acquired from different image acquisition devices. The feature dimensions of the image feature information corresponding to different types of image data are the same. The image data corresponding to the object to be identified can be input into a preset image feature extraction model for feature extraction. This image feature extraction model can be a model corresponding to a deep learning-based image recognition method.

[0096] In scenarios where disease identification is assisted, if the object to be identified is an ophthalmic disease, the image data can include various types of data such as fundus photographs, fluorescein angiography, orbital CT scans, and corneal topography.

[0097] S230. Based on the semantic structure graph, information is fused with a preset knowledge structure graph to obtain a first fusion result; the knowledge structure graph represents the second feature information corresponding to multiple identified objects, and the association information with the second feature description information.

[0098] In one optional embodiment, the identified object can be a historically identified object or an object obtained through relevant knowledge information, such as textbooks, papers, and encyclopedia entries. In the scenario of assisting in disease identification, the identified object can include objects in a disease knowledge base and objects in a historical case database.

[0099] The semantic structure graph and the preset knowledge structure graph can be input into the first information fusion model for information fusion to obtain the first fusion result.

[0100] S240. Based on image feature information, information fusion is performed with the knowledge structure graph to obtain a second fusion result;

[0101] In an optional embodiment, image feature information and knowledge structure graph can be input into a second information fusion model for information fusion to obtain a second fusion result.

[0102] In an optional embodiment, please refer to Figure 3 Based on the semantic structure graph, information is fused with a pre-defined knowledge structure graph to obtain the first fusion result, which includes:

[0103] S310. Determine the association structure graph that is related to the semantic structure graph from the knowledge structure graph;

[0104] S320. Perform information fusion on the semantic structure graph and the association structure graph to obtain the first fusion result;

[0105] Based on image feature information, information fusion is performed with the knowledge structure graph to obtain the second fusion result, which includes:

[0106] S330. Information fusion is performed on the image feature information and the associated structure map to obtain the second fusion result.

[0107] In an optional embodiment, an association structure graph associated with the semantic structure graph can be determined from the knowledge structure graph. This association structure graph can be fused with the semantic graph to obtain a first fusion result, and it can also be fused with image feature information to obtain a second fusion result.

[0108] The semantic nodes and edges in the semantic structure graph are transformed into knowledge query information. Based on this knowledge query information, an association structure graph is determined within the knowledge structure graph. The knowledge query information can be connected to the knowledge structure graph. Knowledge nodes associated with the knowledge query information are identified from the knowledge structure graph, and the association structure graph is constructed by connecting these knowledge nodes with their corresponding edges. The vector representations of nodes in the association structure graph can be the same as those of nodes in the semantic structure graph, and the vector representations of edges in the association structure graph can also be the same as those of edges in the semantic structure graph. For example, if the vectors constituting the semantic structure graph are BERT word vectors, then the vectors constituting the association structure graph are also BERT word vectors.

[0109] For example, in scenarios where disease identification is assisted, when determining the association structure graph in the knowledge structure graph based on knowledge query information, one can query historical medical record information associated with medical record information and other knowledge information associated with medical record information, and construct the association structure graph through the nodes corresponding to historical medical record information and the nodes corresponding to other knowledge information.

[0110] By querying the knowledge nodes associated with the knowledge structure graph through the semantic structure graph, each object recognition can be performed using only a portion of the knowledge structure graph, thereby reducing the computational load of information fusion and improving the efficiency of object recognition.

[0111] In an optional embodiment, please refer to Figure 4 Information fusion is performed on the semantic structure graph and the association structure graph to obtain the first fusion result, which includes:

[0112] S410. Determine the degree of correlation between semantic nodes in the semantic structure graph and knowledge nodes in the association structure graph to obtain the first fusion weight;

[0113] S420. Based on the first fusion weight, perform weighted fusion processing on the semantic nodes in the semantic structure graph to obtain the first fusion result.

[0114] In an optional embodiment, determining the relevance between semantic nodes in the semantic structure graph and knowledge nodes in the association structure graph yields a first fusion weight. The first fusion weight represents the importance of each semantic node in the semantic structure graph to each knowledge node in the association structure graph. Each knowledge node in the association structure graph is traversed, and the currently traversed knowledge node is taken as the current knowledge node. The following operations are performed on the current knowledge node: determining the relevance between each semantic node in the semantic structure graph and the current knowledge node, thus obtaining the first fusion weight corresponding to the current knowledge node. Based on the first fusion weight corresponding to the current knowledge node, a weighted fusion process is performed on the semantic nodes in the semantic structure graph to obtain the semantic knowledge node corresponding to the current knowledge node. The semantic knowledge node corresponding to each knowledge node is taken as the first fusion result.

[0115] The degree of relevance between semantic nodes in the semantic structure graph and knowledge nodes in the association structure graph can be determined using an attention mechanism, and a weighted fusion can be performed to obtain the first fusion result. The specific formula is shown below:

[0116]

[0117]

[0118] in, As the first fusion weight, These are semantic nodes in the semantic structure graph. This represents a knowledge node in the association structure graph. W d W8 and W9 are parameter information. The values ​​of these parameter information can be determined through model training. This is the first fusion result.

[0119] q represents the query text representation of the object to be identified. During each object recognition process, the query text representation corresponding to the current object to be identified is obtained and added to each attention process. The query text representation can be obtained by vector encoding the preset object query text. For example, in a scenario for assisted disease identification, the corresponding query text representation can be obtained by vector encoding the object query text such as "need to determine the disease type" or "need to determine the treatment plan".

[0120] Square brackets denote vector concatenation operations. The above formula corresponds to the additive model in the attention mechanism. When determining the relevance between semantic nodes and knowledge nodes using the attention mechanism, different models such as the dot product model, the scaled dot product model, and the bilinear model can also be applied.

[0121] By using an attention mechanism, the knowledge nodes in the association structure graph and the query text representation are concatenated to obtain the first concatenated information. The degree of correlation between the semantic nodes in the semantic structure graph and the first concatenated information is determined to obtain the first fusion weight. The first fusion result is then calculated based on the first fusion weight.

[0122] In an optional embodiment, please refer to Figure 5 ,like Figure 5 The diagram illustrates the intermodal fusion of a semantic structure graph, an association structure graph, and image feature information. The semantic structure graph includes semantic nodes a1, a2, and a3. Using an attention mechanism, the knowledge node v in the knowledge structure graph is concatenated with the query text representation q to obtain the first concatenated information. The correlation between semantic node a1 and the first concatenated information is determined, resulting in the first fusion weight for semantic node a1. The correlation between semantic node a2 and the first concatenated information is also determined, resulting in the first fusion weight for semantic node a2. Finally, the correlation between semantic node a3 and the first concatenated information is determined, resulting in the first fusion weight for semantic node a3. Based on these first fusion weights, a weighted sum of semantic nodes a1, a2, and a3 is performed to obtain the first fusion result V1.

[0123] The first fusion weight is determined by the degree of correlation between semantic nodes and knowledge nodes, and the semantic nodes are fused by the first fusion weight to obtain the first fusion result. This first fusion result contains semantically related knowledge, which improves the effectiveness of information fusion. Furthermore, it can be combined with the second fusion result in subsequent steps to determine the object recognition result, thereby improving the accuracy of the object recognition result.

[0124] In an optional embodiment, please refer to Figure 6 The image feature information and the associated structure map are fused to obtain the second fusion result, which includes:

[0125] S610. Determine the degree of correlation between image feature information and knowledge nodes in the association structure graph to obtain the second fusion weight;

[0126] S620. Based on the second fusion weight, the image feature information is weighted and fused to obtain the second fusion result.

[0127] In an optional embodiment, determining the correlation between image feature information and knowledge nodes in the association structure graph yields a second fusion weight. The second fusion weight represents the importance of image feature information to each knowledge node in the association structure graph. Each knowledge node in the association structure graph is traversed, and the currently traversed knowledge node is taken as the current knowledge node. The following operations are performed on the current knowledge node: determining the correlation between each image feature information and the current knowledge node, thus obtaining the second fusion weight corresponding to the current knowledge node. Based on the second fusion weight corresponding to the current knowledge node, the image feature information is weighted and fused to obtain the image knowledge node corresponding to the current knowledge node. The image knowledge node corresponding to each knowledge node is then used as the second fusion result.

[0128] The correlation between image feature information and knowledge nodes in the association structure graph can be determined through an attention mechanism, and a weighted fusion is then performed to obtain the second fusion result. The specific formula is as follows:

[0129]

[0130]

[0131] in, As the second fusion weight, Image feature information, This represents a knowledge node in the association structure graph. W d W', W8', and W9' are parameter information. The values ​​of these parameter information can be determined through model training. This is the second fusion result. q represents the query text representation of the object to be identified.

[0132] Square brackets denote vector concatenation operations. The above formula corresponds to the additive model in the attention mechanism. When determining the correlation between image feature information and knowledge nodes through the attention mechanism, different models such as the dot product model, the scaled dot product model, and the bilinear model can also be applied.

[0133] By using an attention mechanism, the knowledge nodes and query text representations in the association structure graph are concatenated to obtain the second concatenation information. The correlation between the image feature information and the second concatenation information is determined to obtain the second fusion weight, and the second fusion result is calculated based on the second fusion weight.

[0134] In an optional embodiment, please refer to Figure 5 ,like Figure 5The diagram illustrates the intermodal fusion of semantic structure graph, association structure graph, and image feature information. The image feature information includes image feature information b1, image feature information b2, and image feature information b3. Through an attention mechanism, after concatenating the knowledge node v in the knowledge structure graph with the query text representation q, a second concatenated information is obtained. The correlation between image feature information b1 and the first concatenated information is determined, resulting in a second fusion weight for image feature information b1. The correlation between image feature information b2 and the second concatenated information is determined, resulting in a second fusion weight for image feature information b2. The correlation between image feature information b3 and the second concatenated information is determined, resulting in a second fusion weight for image feature information b3. Based on the second fusion weights, a weighted sum of image feature information b1, image feature information b2, and image feature information b3 is performed to obtain the second fusion result V2.

[0135] The second fusion weight is determined by the correlation between image feature information and knowledge nodes, and the semantic nodes are fused by the second fusion weight to obtain the second fusion result. This allows the second fusion result to contain knowledge related to the image, improving the effectiveness of information fusion. Furthermore, it can be combined with the first fusion result in subsequent steps to determine the object recognition result, thereby improving the accuracy of the object recognition result.

[0136] S250. Based on the first fusion result and the second fusion result, determine the object recognition result corresponding to the object to be identified.

[0137] In an optional embodiment, nodes in the knowledge structure graph are updated based on the first fusion result and the second fusion result to obtain the target structure graph. The object recognition result corresponding to the object to be identified is determined from the target nodes in the target structure graph. If an association structure graph is determined from the knowledge structure graph based on the semantic structure graph, nodes in the association structure graph are updated based on the first fusion result and the second fusion result to obtain the target structure graph. If information enhancement information is obtained after enhancing the association structure graph, nodes in the association graph enhancement information are updated to obtain the target structure graph. The specific formula is as follows:

[0138] Where t∈{SF,IF,o},

[0139]

[0140] in, The first fusion result corresponding to each knowledge node. The second fusion result corresponding to each knowledge node. This refers to the target node in the target structure graph. q represents the knowledge node in the association structure graph. q is the query text representation of the object to be identified. W e W 10 and W 11 The parameter information can be obtained through model training.

[0141] The target structure graph can be input into a pre-trained classifier to obtain classification information for each target node. The target node corresponding to the target classification information is then used as the object recognition result. This classifier can be a binary classifier, classifying each target node in the target knowledge structure graph to obtain classification information for each node, which can be a classification probability value. The classification information for each target node is then sorted in ascending order to determine the target classification information, which can be the maximum probability value among the classification probabilities. The target node corresponding to this target classification information is then used as the object recognition result.

[0142] In an optional embodiment, please refer to Figure 7 Before fusing information from the semantic structure graph and the association structure graph to obtain the first fusion result, the method further includes:

[0143] S710. Perform information enhancement on the semantic structure graph, the association structure graph, and the image feature information respectively to obtain semantic graph enhancement information corresponding to the semantic structure graph, association graph enhancement information corresponding to the association structure graph, and image enhancement information corresponding to the image feature information;

[0144] Information fusion is performed on the semantic structure graph and the association structure graph to obtain the first fusion result, which includes:

[0145] S720. Information fusion is performed on the semantic graph enhancement information and the association graph enhancement information to obtain the first fusion result;

[0146] The image feature information and the associated structure map are fused to obtain the second fusion result, which includes:

[0147] S730. Information fusion is performed on the image enhancement information and the correlation graph enhancement information to obtain the second fusion result.

[0148] In an optional embodiment, information related to the object to be identified in the semantic structure graph, the association structure graph, and the image feature information can be enhanced to obtain semantic graph enhancement information corresponding to the semantic structure graph, association graph enhancement information corresponding to the association structure graph, and image enhancement information corresponding to the image feature information. Node updates can be performed on the semantic nodes in the semantic structure graph and the knowledge nodes in the association structure graph to enhance information.

[0149] After enhancing the semantic structure graph, association structure graph, and image feature information separately, the enhanced information of the semantic graph and the association structure graph are fused based on the aforementioned method for information fusion of the semantic structure graph and the association structure graph, resulting in a first fusion result. Then, based on the aforementioned method for information fusion of image feature information and association structure graph, the enhanced information of the image and the enhanced information of the association graph are fused, resulting in a second fusion result.

[0150] By enhancing the semantic structure graph, association structure graph, and image feature information respectively, and then using the enhanced information for information fusion and object recognition, the importance of information related to the object to be identified in the semantic structure graph, association structure graph, and image feature information can be increased, thereby improving the accuracy of object recognition.

[0151] In an optional embodiment, please refer to Figure 8 Information enhancement is performed on the semantic structure graph, the association structure graph, and the image feature information respectively, resulting in semantic graph enhancement information corresponding to the semantic structure graph, association graph enhancement information corresponding to the association structure graph, and image enhancement information corresponding to the image feature information, including:

[0152] S810. Determine the first correlation between semantic nodes in the semantic structure graph and the object to be identified;

[0153] S820. Based on the first relevance, perform node update processing on the semantic nodes to obtain semantic graph enhancement information;

[0154] S830. Determine the second relevance between knowledge nodes in the association structure graph and the object to be identified;

[0155] S840. Based on the second relevance, perform node update processing on the knowledge nodes to obtain enhanced information on the association graph structure;

[0156] S850. Determine the third correlation between image feature information and the object to be identified;

[0157] S860. Based on the third correlation, update the image feature information to obtain image enhancement information.

[0158] In an optional embodiment, an attention mechanism is used to determine the first relevance between semantic nodes in the semantic structure graph and the object to be identified. A higher first relevance indicates a stronger correlation between the semantic node and the object. Through graph convolution operations, based on the first relevance, the semantic nodes are updated, and the semantic nodes and their associated semantic nodes are fused to enhance the node representation of the semantic nodes, thus obtaining semantic graph enhancement information. Associated semantic nodes can be adjacent nodes of the semantic nodes.

[0159] By employing an attention mechanism, a second relevance is determined between knowledge nodes in the association graph and the object to be identified. A higher second relevance indicates a stronger correlation between the knowledge node and the object. Then, through graph convolution operations, node updates are performed on the knowledge nodes based on this second relevance. This process merges knowledge nodes with their associated knowledge nodes, enhancing the node representation of each knowledge node and yielding augmented knowledge graph information. Associated knowledge nodes can be adjacent nodes of a knowledge node.

[0160] An attention mechanism is used to determine the third correlation between image feature information and the object to be identified. The higher the third correlation, the more relevant the image feature information is to the object to be identified. Based on the third correlation, the image feature information is updated to obtain image enhancement information.

[0161] The correlation between the semantic structure graph, the association structure graph, and the image feature information and the object to be identified is calculated separately. Information enhancement is performed based on the correlation, thereby increasing the importance of information related to the object to be identified in the semantic structure graph, the association structure graph, and the image feature information, and improving the effectiveness of information enhancement.

[0162] In an optional embodiment, please refer to Figure 9 The first relevance includes semantic node weights and semantic association weights. Determining the first relevance between semantic nodes in the semantic structure graph and the object to be identified includes:

[0163] S910. Perform attention processing on each semantic node to obtain the semantic node weight corresponding to each semantic node;

[0164] S920. Perform attention processing on the associated semantic nodes of each semantic node to obtain the semantic association weights corresponding to each semantic node;

[0165] Based on the first relevance, the semantic nodes are updated to obtain semantic graph enhancement information, including:

[0166] S930. Based on semantic association weights, perform weighted fusion processing on the associated semantic nodes to obtain the first node update information corresponding to each semantic node;

[0167] S940. Based on the first node update information and the semantic node weights, perform node update processing on each semantic node to obtain semantic graph enhancement information.

[0168] In an optional embodiment, attention processing is performed on each semantic node in the semantic structure graph to calculate the attention score for each semantic node. This attention score is then used as the semantic node weight for each semantic node. The semantic node weight represents the importance of each semantic node to the object to be identified, and the specific formula is as follows:

[0169]

[0170] Where, α i For semantic node weights, is a semantic node in the semantic structure graph, and q is the query text representation of the object to be identified. a W1 and W2 are parameter information that can be obtained through model training.

[0171] Attention processing is applied to the semantic association nodes of each semantic node, and the attention association score corresponding to the semantic association nodes of each semantic node is calculated. This attention association score is used as the semantic association weight for each semantic node. The semantic association weight represents the importance of each semantic association node to the object to be identified and its corresponding semantic node. The specific formula is as follows:

[0172]

[0173] Where, β ni For semantic association weights, Includes semantic association nodes and edges connecting semantic association nodes to their corresponding semantic nodes; q' includes the query text representation of the object to be identified and the semantic nodes corresponding to the semantic association nodes. W b W3 and W4 are parameter information that can be obtained through model training.

[0174] The above formula corresponds to the additive model in the attention mechanism. When determining the first relevance using the attention mechanism, different models such as the dot product model, the scaled dot product model, and the bilinear model can also be applied.

[0175] Based on semantic association weights, weighted fusion processing is performed on associated semantic nodes to obtain the first node update information corresponding to each semantic node. Semantic nodes are then weighted based on their respective weights. The weighted semantic nodes and their corresponding first node update information are concatenated, and a non-linear processing is applied to the concatenated information to obtain semantic update nodes. These semantic update nodes constitute the semantic graph enhancement information. The specific formula is as follows:

[0176]

[0177]

[0178] Where, β ni For semantic association weights, α i For semantic node weights, m i Update information for the first node. This includes semantically related nodes and the edges connecting semantically related nodes to their corresponding semantic nodes. Semantic update nodes are used to enhance information in semantic graphs. This represents a semantic node in the semantic structure graph. W5 represents parameter information, which can be obtained through model training. Square brackets indicate vector concatenation operations.

[0179] In an optional embodiment, please refer to Figure 10 ,like Figure 10 The diagram illustrates node updates in a semantic structure graph. The graph includes three nodes: semantic node c1, semantic node c2, and semantic node c3. Semantic nodes c2 and c3 are adjacent to semantic node c1 and are semantically related nodes of c1. Semantic nodes c2 and c1 are connected by edge r1, and semantic nodes c3 and c1 are connected by edge r2. The semantic node weight of semantic node c1 is calculated. Based on the query text representation q, the features of semantic node c1, the features of semantic node c2, and the features of edge r1, the semantic association weight between semantic nodes c1 and c2 is calculated. Similarly, based on the query text representation q, the features of semantic node c1, the features of semantic node c3, and the features of edge r2, the semantic association weight between semantic nodes c1 and c3 is calculated. Using these semantic association weights, semantic nodes c2 and c3 are fused to obtain the first node update information. Based on this first node update information and the semantic node weights, semantic node c1 is updated.

[0180] By using attention computation and graph convolution, the semantic association nodes of each semantic node are fused together and then fused with each weighted semantic node. This allows the semantic nodes to be updated based on the weights of the nodes and edges, improving the effectiveness of information augmentation for semantic nodes.

[0181] In an optional embodiment, please refer to Figure 11 The second relevance includes knowledge node weight and knowledge association weight. Determining the second relevance between knowledge nodes and the object to be identified in the association structure graph includes:

[0182] S1110. Perform attention processing on each knowledge node to obtain the knowledge node weight corresponding to each knowledge node;

[0183] S1120. Perform attention processing on the associated knowledge nodes of each knowledge node to obtain the knowledge association weight corresponding to each knowledge node;

[0184] Based on the second relevance, node update processing is performed on the knowledge nodes to obtain enhanced information for the association graph structure, including:

[0185] S1130. Based on the knowledge association weight, perform weighted fusion processing on the associated knowledge nodes to obtain the second node update information corresponding to each knowledge node;

[0186] S1140. Based on the update information of the second node and the weight of the knowledge node, update each knowledge node to obtain the knowledge graph enhancement information.

[0187] In an optional embodiment, attention processing is performed on each knowledge node in the association structure graph to calculate the attention score for each knowledge node. This attention score is then used as the weight of each knowledge node. The knowledge node weight represents the importance of each knowledge node to the object being identified, and the specific formula is as follows:

[0188]

[0189] Where, α' i For knowledge node weights, W represents a knowledge node in the association structure graph, where q is the query text representation of the object to be identified. a W1' and W2' are parameter information, which can be obtained through model training.

[0190] Attention processing is applied to the knowledge-related nodes of each knowledge node, and the attention association score corresponding to each knowledge node's knowledge-related nodes is calculated. This attention association score is used as the knowledge association weight for each knowledge node. The knowledge association weight represents the importance of each knowledge-related node to the object to be identified and its corresponding knowledge node. The specific formula is as follows:

[0191]

[0192] Where, β' ni For knowledge association weight, This includes knowledge-related nodes and the edges connecting them to their corresponding knowledge nodes. q' includes the query text representation of the object to be identified and the knowledge node corresponding to the knowledge-related node. W b W', W3', and W4' are parameter information that can be obtained through model training.

[0193] The above formula corresponds to the additive model in the attention mechanism. When determining the first relevance using the attention mechanism, different models such as the dot product model, the scaled dot product model, and the bilinear model can also be applied.

[0194] Based on knowledge association weights, a weighted fusion process is performed on associated knowledge nodes to obtain the update information of the second node corresponding to each knowledge node. Based on the knowledge node weights, the knowledge nodes are weighted. The weighted knowledge node and its corresponding second node update information are concatenated, and a non-linear processing is applied to the concatenated information to obtain the knowledge update nodes. These knowledge update nodes constitute the knowledge graph enhancement information. The specific formula is as follows:

[0195]

[0196]

[0197] Where, β' ni For knowledge association weights, α' i m' represents the weight of the knowledge node. i Update information for the second node. This includes knowledge-related nodes and the edges connecting each knowledge-related node to its corresponding knowledge node. Enhance the knowledge update nodes in the knowledge graph. These represent knowledge nodes in the association structure graph. Square brackets indicate vector concatenation operations.

[0198] In an optional embodiment, please refer to Figure 12 ,like Figure 12 The diagram illustrates node updates in a knowledge structure graph. The graph includes three nodes: knowledge node d1, knowledge node d2, and knowledge node d3. Knowledge nodes d2 and d3 are adjacent to knowledge node d1 and are knowledge-related nodes of d2. Knowledge nodes d2 and d1 are connected by edge r3, and knowledge nodes d3 and d1 are connected by edge r4. The knowledge node weight of knowledge node d1 is calculated. Based on the query text representation q, the features of knowledge node d1, the features of knowledge node d2, and the features of edge r3, the knowledge-related weight between knowledge nodes d1 and d2 is calculated. Similarly, based on the query text representation q, the features of knowledge node d1, the features of knowledge node d3, and the features of edge r4, the knowledge-related weight between knowledge nodes d1 and d3 is calculated. By fusing the knowledge-related weights, the update information for the second node is obtained. Based on the update information and the knowledge node weights, knowledge node d1 is updated.

[0199] By using attention computation and graph convolution, the knowledge-related nodes of each knowledge node are fused together, and then fused with each weighted knowledge node. This allows the knowledge nodes to be updated based on the weights of the nodes and edges, improving the effectiveness of information augmentation for knowledge nodes.

[0200] In an optional embodiment, please refer to Figure 13 Determining the third correlation between image feature information and the object to be identified includes:

[0201] S1310. Perform attention processing on the image feature information to obtain the third relevance;

[0202] Based on the third correlation, the image feature information is updated to obtain image enhancement information, including:

[0203] S1320. Based on the third correlation, perform feature update processing on the image feature information to obtain image enhancement information.

[0204] In an optional embodiment, attention processing is performed on the image feature information to calculate the attention score corresponding to each image feature information. This attention score is then used as the third relevance for each image feature information. The third relevance can be the image feature weight, representing the importance of each image feature information to the object to be identified. The specific formula is as follows:

[0205]

[0206] Where, α k As the third relevance, q represents the image feature information, and q represents the query text representation of the object to be identified. c W6 and W7 are parameter information that can be obtained through model training.

[0207] When the image feature information is from multiple sources, it can be treated as discrete nodes with no correlation between them. This allows for node updates to obtain image enhancement information. Based on third-party relevance, the image feature information is weighted. The weighted image feature information is then concatenated with the query text representation of the object to be identified, and the concatenated information undergoes non-linear processing to obtain the image enhancement information. The specific formula is as follows:

[0208]

[0209] Where, α k As the third relevance, Enhance information in the image. This represents image feature information. Square brackets indicate vector concatenation operations. q represents the query text representation of the object to be identified.

[0210] By using attention calculation to fuse each weighted image feature information, the image feature information can be updated, thereby improving the effectiveness of information enhancement for image feature information.

[0211] In an optional embodiment, the method can be applied to scenarios where ophthalmic diseases or treatment plans for ophthalmic diseases are determined in an auxiliary manner. For example... Figure 14 As shown, Figure 14 This diagram illustrates the application of the aforementioned object recognition method to identify ophthalmic diseases or treatment plans. In scenarios where ophthalmic diseases or treatment plans are being identified, the object to be identified is the ophthalmic disease or its treatment plan.

[0212] Medical record data is acquired and input as text data into the text feature extraction layer of the semantic graph construction model for text feature extraction. The text data is divided into multiple words and converted into vector representations to obtain first feature information, first feature description information, and the association information between the first feature information and the first feature description information. The first feature information, the first feature description information, and the corresponding association information are then input into the semantic graph generation layer of the semantic graph construction model for semantic graph generation, resulting in a semantic structure graph.

[0213] After extracting textual features from medical record data, identification can be performed. If the medical record data includes an ophthalmological disease type, the object to be identified can be an ophthalmological disease treatment plan; if the medical record data does not include an ophthalmological disease type, the object to be identified can be an ophthalmological disease type. The object to be identified can also be obtained through user configuration.

[0214] Using the ophthalmology knowledge graph as a knowledge structure graph, we query candidate knowledge nodes that are associated with the semantic structure graph in the knowledge structure graph, and construct an association structure graph by connecting the candidate knowledge nodes and the edges corresponding to the candidate knowledge nodes.

[0215] By using multi-source inspection image data as image data and inputting the image data into an image feature extraction model for feature extraction, image feature information can be obtained.

[0216] Attention is calculated for each semantic node in the semantic structure graph to determine its corresponding semantic node weight. Attention is also calculated for each semantic node's associated semantic nodes to determine their associated semantic node weights. Based on these associated weights, a weighted fusion process is performed on the associated semantic nodes to obtain the first node update information for each semantic node. Based on this first node update information and the semantic node weights, the semantic nodes are updated by adding the importance of adjacent semantic nodes to the object to be identified, as well as the importance of each semantic node to the object to be identified. This increases the weight of semantic nodes strongly correlated with the object to be identified in the semantic structure graph, resulting in semantic graph enhancement information.

[0217] Attention is calculated for each knowledge node in the association graph to determine its weight. Attention is also calculated for each knowledge node's associated knowledge nodes to determine their associated weights. Based on these weights, a weighted fusion process is performed on the associated knowledge nodes to obtain the updated information for the second node corresponding to each knowledge node. Using this updated information and the knowledge node weights, the knowledge nodes are updated by adding the importance of adjacent knowledge nodes to the object to be identified, as well as the importance of each knowledge node itself to the object. This increases the weight of knowledge nodes strongly correlated with the object to be identified in the association graph, resulting in enhanced knowledge graph information.

[0218] Attention is calculated for each image feature to determine the third relevance. Based on the third relevance, the image feature information is updated by adding the importance of the image feature to the object to be identified. This increases the weight of image features strongly correlated with the object to be identified, resulting in image enhancement information.

[0219] Information fusion is performed on the semantic graph augmentation information and the association graph augmentation information to obtain the first fusion result. The current knowledge node can be determined from the knowledge nodes in the association graph augmentation information. Through an attention mechanism, the relevance between each semantic node in the semantic graph augmentation information and the current knowledge node in the association graph augmentation information is determined, obtaining the first fusion weight corresponding to the current knowledge node. Based on the first fusion weight, a weighted fusion is performed on each semantic node in the semantic graph augmentation information, which is equivalent to updating the current knowledge node, to obtain the semantic knowledge node corresponding to the current knowledge node. The semantic knowledge node corresponding to each knowledge node in the association graph augmentation information is the first fusion result. The specific formula is as follows:

[0220]

[0221]

[0222] in, As the first fusion weight, To enhance the information of nodes in a semantic graph, To enhance the information of nodes in the association graph. W d W8 and W9 are parameter information. The values ​​of these parameter information can be determined through model training. This is the first fusion result.

[0223] Information fusion is performed on image enhancement information and correlation graph enhancement information to obtain a second fusion result. The current knowledge node can be determined from the knowledge nodes of the correlation graph enhancement information. Through an attention mechanism, the correlation between each image enhancement information and the current knowledge node in the correlation graph enhancement information is determined, obtaining the second fusion weight corresponding to the current knowledge node. Based on the second fusion weight, weighted fusion is performed on each image enhancement information, which is equivalent to updating the current knowledge node, thus obtaining the image knowledge node corresponding to the current knowledge node. The image knowledge node corresponding to each image enhancement information is the second fusion result. The specific formula is as follows:

[0224]

[0225]

[0226] in, As the second fusion weight, To enhance image information, To enhance the information of nodes in the association graph. W d W', W8', and W9' are parameter information. The values ​​of these parameter information can be determined through model training. This is the second fusion result. q represents the query text representation of the object to be identified.

[0227] Based on the first and second fusion results, the nodes in the association structure graph are updated to obtain the target structure graph. The specific formula is as follows:

[0228] Where t∈{SF,IF,o},

[0229]

[0230] in, The first fusion result corresponding to each knowledge node. The second fusion result corresponding to each knowledge node. This refers to the target node in the target structure graph. Nodes in the association graph are used to enhance information. q represents the query text representation of the object to be identified. W e W 10 and W 11 The parameter information can be obtained through model training.

[0231] The target nodes in the target structure graph are classified, and the target node corresponding to the maximum probability value among the classification probability values ​​is determined. The target node corresponding to the maximum probability value is used as the object recognition result for the object to be identified. This object recognition result is the target ophthalmic disease type or the target ophthalmic disease treatment plan, thus enabling the early identification of insidious ophthalmic diseases.

[0232] This embodiment proposes an object recognition method, which includes: constructing a semantic graph from the text data corresponding to the object to be recognized, obtaining a semantic structure graph; extracting features from the image data corresponding to the object to be recognized, obtaining image feature information; fusing information from the semantic structure graph with a preset knowledge structure graph to obtain a first fusion result; fusing information from the image feature information with the preset knowledge structure graph to obtain a second fusion result; and determining the object recognition result corresponding to the object to be recognized from the knowledge structure graph based on the first and second fusion results. This method can combine the fusion results of semantic information and knowledge information, as well as the fusion results of image information and knowledge information, to recognize the object to be recognized, thereby improving the accuracy and effectiveness of object recognition.

[0233] This application also provides an object recognition device; please refer to [link to relevant documentation]. Figure 15 The device includes:

[0234] The semantic graph construction module 1510 is used to construct a semantic graph from the text data corresponding to the object to be identified, and obtain a semantic structure graph. The semantic structure graph represents the first feature information of the object to be identified and the association information with the first feature description information.

[0235] The image feature extraction module 1520 is used to extract features from the image data corresponding to the object to be identified, and obtain image feature information.

[0236] The first information fusion module 1530 is used to perform information fusion with a preset knowledge structure graph based on the semantic structure graph to obtain a first fusion result; the knowledge structure graph represents the second feature information corresponding to multiple identified objects and the association information with the second feature description information.

[0237] The second information fusion module 1540 is used to perform information fusion with the knowledge structure graph based on image feature information to obtain a second fusion result.

[0238] The object recognition module 1550 is used to determine the object recognition result corresponding to the object to be recognized from the knowledge structure graph based on the first fusion result and the second fusion result.

[0239] In an optional embodiment, the first information fusion module includes:

[0240] The association structure graph determination unit is used to determine the association structure graph associated with the semantic structure graph from the knowledge structure graph;

[0241] The first fusion result determination unit is used to perform information fusion on the semantic structure graph and the association structure graph to obtain the first fusion result;

[0242] The second information fusion unit includes:

[0243] The second fusion result determination unit is used to perform information fusion on image feature information and correlation structure map to obtain the second fusion result.

[0244] In an optional embodiment, the first fusion result determination unit includes:

[0245] The first fusion weight determination unit is used to determine the degree of correlation between semantic nodes in the semantic structure graph and knowledge nodes in the association structure graph, and to obtain the first fusion weight.

[0246] The first weighted fusion unit is used to perform weighted fusion processing on the semantic nodes in the semantic structure graph based on the first fusion weight to obtain the first fusion result.

[0247] In an optional embodiment, the second fusion result determination unit includes:

[0248] The second fusion weight determination unit is used to determine the degree of correlation between image feature information and knowledge nodes in the association structure graph, and to obtain the second fusion weight.

[0249] The second weighted fusion unit is used to perform weighted fusion processing on image feature information based on the second fusion weight to obtain the second fusion result.

[0250] In an optional embodiment, the apparatus further includes:

[0251] The information enhancement module is used to enhance the semantic structure graph, the association structure graph and the image feature information respectively, to obtain the semantic graph enhancement information corresponding to the semantic structure graph, the association graph enhancement information corresponding to the association structure graph and the image enhancement information corresponding to the image feature information.

[0252] The first information fusion module includes:

[0253] The first enhanced information fusion unit is used to fuse the semantic graph enhanced information and the association graph enhanced information to obtain the first fusion result;

[0254] The second information fusion module includes:

[0255] The second enhanced information fusion unit is used to fuse image enhancement information and correlation graph enhancement information to obtain a second fusion result.

[0256] In an optional embodiment, the information enhancement module includes:

[0257] The first relevance determination unit is used to determine the first relevance between the semantic nodes in the semantic structure graph and the object to be identified.

[0258] The first node update unit is used to perform node update processing on semantic nodes based on the first relevance to obtain semantic graph enhancement information;

[0259] The second relevance determination unit is used to determine the second relevance between the knowledge nodes in the association structure graph and the object to be identified.

[0260] The second node update unit is used to update the knowledge nodes based on the second relevance to obtain enhanced information about the association graph structure.

[0261] The third correlation determination unit is used to determine the third correlation between image feature information and the object to be identified.

[0262] The image feature update unit is used to update the image feature information based on the third correlation to obtain image enhancement information.

[0263] In an optional embodiment, the first relevance includes semantic node weights and semantic association weights, and the first relevance determining unit includes:

[0264] The semantic node weight determination unit is used to perform attention processing on each semantic node to obtain the semantic node weight corresponding to each semantic node.

[0265] The semantic association weight determination unit is used to perform attention processing on the associated semantic nodes of each semantic node to obtain the semantic association weight corresponding to each semantic node.

[0266] The first node update unit includes:

[0267] The first node update information determination unit is used to perform weighted fusion processing on the associated semantic nodes based on the semantic association weight to obtain the first node update information corresponding to each semantic node.

[0268] The semantic graph enhancement information determination unit is used to perform node update processing on each semantic node based on the first node update information and the semantic node weights to obtain semantic graph enhancement information.

[0269] In an optional embodiment, the second relevance includes knowledge node weights and knowledge association weights, and the second relevance determining unit includes:

[0270] The knowledge node weight determination unit is used to perform attention processing on each knowledge node to obtain the knowledge node weight corresponding to each knowledge node.

[0271] The knowledge association weight determination unit is used to perform attention processing on the associated knowledge nodes of each knowledge node to obtain the knowledge association weight corresponding to each knowledge node.

[0272] The second node update unit includes:

[0273] The second node update information determination unit is used to perform weighted fusion processing on the associated knowledge nodes based on the knowledge association weight to obtain the second node update information corresponding to each knowledge node.

[0274] The knowledge graph augmentation information determination unit is used to update each knowledge node based on the second node update information and the knowledge node weights to obtain knowledge graph augmentation information.

[0275] In an optional embodiment, the third relevance determination unit includes:

[0276] The attention processing unit is used to perform attention processing on image feature information to obtain a third relevance.

[0277] The image feature update unit includes:

[0278] The image enhancement information determination unit is used to perform feature update processing on image feature information based on the third correlation to obtain image enhancement information.

[0279] The apparatus provided in the above embodiments can execute the method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiments can be found in an object recognition method provided in any embodiment of this application.

[0280] This embodiment also provides a computer-readable storage medium storing computer-executable instructions, which are loaded by a processor and executed by the object recognition method described above in this embodiment.

[0281] This embodiment also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the object identification described above.

[0282] This embodiment also provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program adapted to be loaded by the processor and executed as described above in this embodiment of an object recognition method.

[0283] The device may be a computer terminal, a mobile terminal, or a server, and may also participate in constituting the apparatus or system provided in the embodiments of this application. For example... Figure 16 As shown, server 16 may include one or more processors 1602 (shown as 1602a, 1602b, ..., 1602n in the figure) 1602 (processor 1602 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPLD), a memory 1604 for storing data, and a transmission device 1606 for communication functions. In addition, it may also include: input / output interfaces (I / O interfaces) and network interfaces. Those skilled in the art will understand that... Figure 16 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 16 may also include components that are more... Figure 16 The more or fewer components shown, or having the same Figure 16 The different configurations shown.

[0284] It should be noted that the aforementioned one or more processors 1602 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the server 16.

[0285] The memory 1604 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method described in the embodiments of this application. The processor 1602 executes various functional applications and data processing by running the software programs and modules stored in the memory 1604, thereby realizing the above-described method for generating temporal behavior capture boxes based on self-attention networks. The memory 1604 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1604 may further include memory remotely located relative to the processor 1602, and these remote memories can be connected to the server 16 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0286] The transmission device 1606 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 16. In one example, the transmission device 1606 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet.

[0287] This specification provides the operational steps of the methods described in the embodiments or flowcharts, but more or fewer operational steps may be included based on conventional or non-inventive labor. The steps and order listed in the embodiments are merely one possible execution order among many steps and do not represent the only execution order. In actual system or interrupt product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0288] The structure shown in this embodiment is only a partial structure related to the solution of this application and does not constitute a limitation on the device to which the solution of this application is applied. Specific devices may include more or fewer components than shown, or combinations of certain components, or arrangements of different components. It should be understood that the methods, apparatuses, etc., disclosed in this embodiment can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or unit modules through some interfaces.

[0289] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0290] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0291] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An object recognition method, characterized in that, The method includes: A semantic graph is constructed from the text data corresponding to the object to be identified to obtain a semantic structure graph. The semantic structure graph represents the first feature information of the object to be identified and the association information with the first feature description information. Feature extraction is performed on the image data corresponding to the object to be identified to obtain image feature information; Based on the semantic structure graph, information is fused with a preset knowledge structure graph to obtain a first fusion result; the knowledge structure graph represents the second feature information corresponding to multiple identified objects, and the association information with the second feature description information. Based on the image feature information, information fusion is performed with the knowledge structure graph to obtain a second fusion result; Based on the first fusion result and the second fusion result, the object recognition result corresponding to the object to be identified is determined from the knowledge structure graph; The first fusion result obtained by fusing information based on the semantic structure graph and a preset knowledge structure graph includes: Determine the association structure graph that is associated with the semantic structure graph from the knowledge structure graph; Determine the first correlation between the semantic nodes in the semantic structure graph and the object to be identified; Based on the first relevance, the semantic nodes are updated to obtain the semantic graph enhancement information; Determine the second relevance between the knowledge nodes in the association structure graph and the object to be identified; Based on the second relevance, the knowledge nodes are updated to obtain enhanced association graph structure information; Determine the third correlation between the image feature information and the object to be identified; Based on the third correlation, the image feature information is updated to obtain image enhancement information; The semantic graph enhancement information and the association graph enhancement information are fused to obtain the first fusion result; The step of fusing information based on the image feature information and the knowledge structure graph to obtain the second fusion result includes: The image enhancement information and the correlation graph enhancement information are fused to obtain the second fusion result.

2. The object recognition method according to claim 1, characterized in that, The first fusion result obtained by fusing information based on the semantic structure graph and a preset knowledge structure graph includes: Determine the association structure graph that is associated with the semantic structure graph from the knowledge structure graph; Information fusion is performed on the semantic structure graph and the association structure graph to obtain the first fusion result; The step of fusing information based on the image feature information and the knowledge structure graph to obtain the second fusion result includes: The image feature information and the associated structure diagram are fused to obtain the second fusion result.

3. The object recognition method according to claim 2, characterized in that, The step of fusing information from the semantic structure graph and the association structure graph to obtain the first fusion result includes: Determine the degree of correlation between semantic nodes in the semantic structure graph and knowledge nodes in the association structure graph to obtain the first fusion weight; Based on the first fusion weight, the semantic nodes in the semantic structure graph are subjected to weighted fusion processing to obtain the first fusion result.

4. The object recognition method according to claim 2, characterized in that, The step of fusing the image feature information and the associated structure map to obtain the second fusion result includes: Determine the correlation between the image feature information and the knowledge nodes in the association structure graph to obtain the second fusion weight; Based on the second fusion weight, the image feature information is subjected to weighted fusion processing to obtain the second fusion result.

5. The object recognition method according to claim 1, characterized in that, The first relevance includes semantic node weights and semantic association weights. Determining the first relevance between the semantic nodes in the semantic structure graph and the object to be identified includes: Attention processing is performed on each semantic node to obtain the semantic node weight corresponding to each semantic node; Attention processing is performed on the associated semantic nodes of each semantic node to obtain the semantic association weight corresponding to each semantic node; The step of updating the semantic nodes based on the first relevance to obtain the semantic graph enhancement information includes: Based on the semantic association weights, the associated semantic nodes are subjected to weighted fusion processing to obtain the first node update information corresponding to each semantic node; Based on the first node update information and the semantic node weights, node update processing is performed on each semantic node to obtain the semantic graph enhancement information.

6. The object recognition method according to claim 1, characterized in that, The second relevance includes knowledge node weight and knowledge association weight. Determining the second relevance between the knowledge nodes in the association structure graph and the object to be identified includes: Attention processing is performed on each knowledge node to obtain the knowledge node weight corresponding to each knowledge node; Attention processing is performed on the associated knowledge nodes of each knowledge node to obtain the knowledge association weight corresponding to each knowledge node. The step of updating the knowledge nodes based on the second relevance to obtain enhanced association graph structure information includes: Based on the knowledge association weights, the associated knowledge nodes are weighted and fused to obtain the second node update information corresponding to each knowledge node; Based on the update information of the second node and the weight of the knowledge node, each knowledge node is updated to obtain knowledge graph enhancement information.

7. The object recognition method according to claim 1, characterized in that, Determining the third correlation between the image feature information and the object to be identified includes: Attention processing is applied to the image feature information to obtain the third relevance; updating the image feature information based on the third relevance to obtain image enhancement information includes: Based on the third relevance, the image feature information is updated to obtain the image enhancement information.

8. An object recognition device, characterized in that, The device includes: The semantic graph construction module is used to construct a semantic graph from the text data corresponding to the object to be identified, thereby obtaining a semantic structure graph. The semantic structure graph represents the first feature information of the object to be identified and the association information between the first feature description information and the object to be identified. The image feature extraction module is used to extract features from the image data corresponding to the object to be identified, and obtain image feature information. The first information fusion module is used to perform information fusion with the preset knowledge structure graph based on the semantic structure graph to obtain a first fusion result; the knowledge structure graph represents the second feature information corresponding to multiple identified objects, and the association information with the second feature description information. The second information fusion module is used to perform information fusion with the knowledge structure graph based on the image feature information to obtain a second fusion result; An object recognition module is used to determine the object recognition result corresponding to the object to be recognized from the knowledge structure graph based on the first fusion result and the second fusion result. The first fusion result obtained by fusing information based on the semantic structure graph and a preset knowledge structure graph includes: Determine the association structure graph that is associated with the semantic structure graph from the knowledge structure graph; Determine the first correlation between the semantic nodes in the semantic structure graph and the object to be identified; Based on the first relevance, the semantic nodes are updated to obtain the semantic graph enhancement information; Determine the second relevance between the knowledge nodes in the association structure graph and the object to be identified; Based on the second relevance, the knowledge nodes are updated to obtain enhanced association graph structure information; Determine the third correlation between the image feature information and the object to be identified; Based on the third correlation, the image feature information is updated to obtain image enhancement information; The semantic graph enhancement information and the association graph enhancement information are fused to obtain the first fusion result; The step of fusing information based on the image feature information and the knowledge structure graph to obtain the second fusion result includes: The image enhancement information and the correlation graph enhancement information are fused to obtain the second fusion result.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement claim 1.

7. An object recognition method as described in any one of the following.

10. A computer-readable storage medium, characterized in that, The storage medium includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement claim 1.

7. An object recognition method as described in any one of the following.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements claim 1.

7. The object recognition method described in any one of the following.

Citation Information

Patent Citations

  • Knowledge-fused multi-modal false news identification method and device

    CN113946683A

  • Text recognition method and device, electronic equipment and readable storage medium

    CN114298054A