Knowledge graph construction method, knowledge graph construction device, medium and electronic equipment
By constructing a visual tag knowledge graph, obtaining the relationships between visual tags and adding attributes, the reliability problem in the field of visual tags is solved, and accurate visual tag processing and rich knowledge information are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI JINSHENG COMM TECH CO LTD
- Filing Date
- 2023-03-23
- Publication Date
- 2026-05-26
AI Technical Summary
The lack of reliable visual labeling knowledge graphs in existing technologies prevents their widespread application in the field of vision.
A visual tag knowledge graph is constructed by acquiring multiple visual tags, determining the relationships between them, generating nodes and edges, and adding node attributes and edge attributes based on the characteristics of the visual tags.
An accurate and reliable visual tag knowledge graph was constructed, enriching the knowledge information. It is applicable to various visual tag processing scenarios, improving the utilization efficiency and coverage of visual tags.
Smart Images

Figure CN116304104B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a knowledge graph construction method, a knowledge graph construction apparatus, a computer-readable storage medium, and an electronic device. Background Technology
[0002] Knowledge graphs aim to model and record the relationships and knowledge between things using graph structures, enabling more accurate object-level search. In recent years, with the rapid development of natural language processing, deep learning, graph data processing, and many other fields, knowledge graph technologies have been widely applied in search engines, intelligent question answering, language understanding, recommendation computing, big data decision analysis, and many other areas. Knowledge graphs have become an indispensable technology for realizing cognitive-level artificial intelligence. However, existing technologies typically focus only on how to apply existing knowledge graphs to specific scenarios, lacking accurate and reliable visual labeling knowledge graphs that can be widely applied in the visual domain. Summary of the Invention
[0003] This disclosure provides a knowledge graph construction method, a knowledge graph construction apparatus, a computer-readable storage medium, and an electronic device, thereby at least partially solving the problem of the lack of reliable and effective visual tag knowledge graphs in the field of visual tagging in the prior art.
[0004] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0005] According to a first aspect of this disclosure, a knowledge graph construction method is provided, comprising: acquiring a plurality of visual tags and determining the relationship between the visual tags; generating nodes based on the visual tags, generating edges between the nodes based on the relationship between the visual tags, to construct a visual tag knowledge graph; adding node attributes to the nodes according to the features of the visual tags, and adding edge attributes to the edges according to the relationship type between the visual tags.
[0006] According to a second aspect of this disclosure, a knowledge graph construction apparatus is provided, comprising: a tag acquisition module for acquiring multiple visual tags and determining the relationship between the visual tags; a graph generation module for generating nodes based on the visual tags and generating edges between the nodes based on the relationship between the visual tags, so as to construct a visual tag knowledge graph; and an attribute addition module for adding node attributes to the nodes according to the features of the visual tags and adding edge attributes to the edges according to the relationship type between the visual tags.
[0007] According to a third aspect of this disclosure, an image description method is provided, comprising: acquiring a visual tag knowledge graph, the visual tag knowledge graph being constructed by the knowledge graph construction method described in the first aspect; obtaining a main tag of the image to be described based on the result of recognizing the image to be described; mapping the main tag using the visual tag knowledge graph to obtain a mapped tag; and generating description information of the image to be described based on the main tag or a tag phrase formed by the main tag and the mapped tag.
[0008] According to a fourth aspect of this disclosure, an image description apparatus is provided, comprising: a knowledge graph acquisition module for acquiring a visual tag knowledge graph, the visual tag knowledge graph being constructed by the knowledge graph construction method described in the first aspect; a main tag acquisition module for obtaining a main tag of the image to be described based on the result of recognizing the image to be described; a mapping tag acquisition module for mapping the main tag using the visual tag knowledge graph to obtain a mapping tag; and a description information acquisition module for generating description information of the image to be described based on the main tag or a tag phrase formed by the main tag and the mapping tag.
[0009] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the knowledge graph construction method of the first aspect, the image description method of the third aspect, and possible implementations thereof.
[0010] According to a sixth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor. The processor is configured to execute the knowledge graph construction method of the first aspect, the image description method of the third aspect, and possible implementations thereof, by executing the executable instructions.
[0011] The technical solution disclosed herein has the following beneficial effects:
[0012] This exemplary embodiment proposes a method for constructing a knowledge graph. It obtains multiple visual tags and determines the relationships between them. Nodes are generated based on the visual tags, and edges are generated between these nodes based on the relationships between them, thus constructing a visual tag knowledge graph. Node attributes are added to the nodes based on the characteristics of the visual tags, and edge attributes are added to the edges based on the types of relationships between the visual tags. On one hand, this exemplary embodiment proposes a method for constructing a knowledge graph that uses visual tags as nodes and the relationships between them as edges to build an accurate and reliable visual tag knowledge graph, thereby providing a trustworthy processing method for the visual tag domain. On the other hand, this exemplary embodiment adds corresponding attributes to the visual tag nodes and the types of relationships between them to further improve the visual tag knowledge graph and enrich its knowledge information. Furthermore, this exemplary embodiment can construct a visual tag knowledge graph for visual tag processing through a simple and convenient process, and has a wide range of applications.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0015] Figure 1 A schematic diagram of a system architecture in this exemplary embodiment is shown;
[0016] Figure 2 This diagram illustrates a flowchart of a knowledge graph construction method in this exemplary embodiment;
[0017] Figure 3 A schematic diagram illustrating a category definition of a visual label in this exemplary embodiment;
[0018] Figure 4 A schematic diagram illustrating another category definition of a visual label in this exemplary embodiment;
[0019] Figure 5 A flowchart illustrating an image description method in this exemplary embodiment is shown;
[0020] Figure 6 A flowchart illustrating another image description method in this exemplary embodiment is shown;
[0021] Figure 7This diagram illustrates a structural block diagram of a knowledge graph construction apparatus according to this exemplary embodiment;
[0022] Figure 8 This diagram illustrates a structural block diagram of an image description device according to an exemplary embodiment of the present invention.
[0023] Figure 9 A structural diagram of an electronic device according to this exemplary embodiment is shown. Detailed Implementation
[0024] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0025] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0026] The exemplary embodiments of this disclosure first provide a method for constructing a knowledge graph. The following, in conjunction with... Figure 1 The system architecture of the operating environment for this exemplary embodiment will be described.
[0027] refer to Figure 1 As shown, the system architecture 100 may include a terminal 110 and a server 120. The terminal 110 may be an electronic device such as a mobile phone, tablet computer, or smart wearable device. The server 120 generally refers to the backend system providing the knowledge graph generation in this exemplary embodiment, and may be a single server or a cluster of multiple servers. The terminal 110 and the server 120 can be connected via a wired or wireless communication link for data interaction.
[0028] In one implementation, the knowledge graph construction method of this exemplary embodiment can be executed by terminal 110. For example, terminal 110 can obtain visual tags from a local photo album, the network, or other cloud platforms, determine the relationships between the visual tags, and then execute the method for constructing a visual tag knowledge graph.
[0029] In one implementation, the knowledge graph construction method of this exemplary embodiment can be executed by server 120. For example, server 120 can obtain multiple visual tags and the relationships between the visual tags from terminal 110, and then execute the visual tag knowledge graph construction method; or terminal 110 can upload multiple visual tags to server 120, and server 120 can determine the relationships between the visual tags and execute the visual tag knowledge graph construction method, etc.
[0030] As can be seen from the above, in this exemplary embodiment, the execution subject of the knowledge graph construction method can be the aforementioned terminal 110 or server 120, and this disclosure does not limit it in this regard.
[0031] The following is combined with Figure 2 This section explains the process of constructing knowledge graphs. (References) Figure 2 As shown, the knowledge graph construction method may include the following steps S210 to S230:
[0032] Step S210: Obtain multiple visual labels and determine the relationship between the visual labels.
[0033] In this context, a visual label refers to a classification label for an object with visual characteristics. The object can be a physical object, scene, building, event, etc., within an image. An object with visual characteristics is one that can be distinguished from other things based on visual perception. For example, "cat" and "dog" have obvious visual features and can have corresponding visual labels for cat or dog. While "wedding" doesn't have a directly corresponding physical object, it has a certain degree of visual distinguishability; for example, a photo showing a wedding in a church can have the visual label "wedding." "You" or "I," on the other hand, do not possess visual characteristics and do not have corresponding visual labels. In this exemplary embodiment, a visual label can indicate a class of visually characteristic objects. For example, multiple images containing dogs could be images of dogs being walked, images of pet stores, or close-up images of a particular dog, all of which could correspond to the visual label "dog."
[0034] The relationships between visual tags can be used to reflect the type of relationship between visual tags, or the degree of relevance, etc. For example, the visual tags "cat" and "dog" are both hyponyms of the visual tag "pet", and the visual tags "balloon" and "wedding" have a high degree of relevance, and so on.
[0035] In this exemplary embodiment, visual tags can be obtained from multiple sources, such as images extracted from a local photo album, images downloaded from the cloud, or collected from an existing visual tag library, etc.
[0036] In an exemplary embodiment, step S210, obtaining multiple visual labels, may include:
[0037] Obtain candidate tags, which include at least one of the following: requirement tags in visual business, tags supported by image classification function, search terms that meet preset search conditions in image search, and words obtained by segmenting image description statements.
[0038] Visual labels are selected from candidate labels.
[0039] Considering that the content in the visual tag knowledge graph needs to be highly relevant to the visual content of the image, this exemplary embodiment can determine candidate tags from multiple sources and then filter them to obtain visual tags.
[0040] Specifically, candidate tags can be requirement tags in visual business, such as tags under the album classification requirement in the terminal or device, such as pet tags, portrait tags, or building tags; or tags under the search requirement in the terminal or device, such as portrait tags proposed when a user searches for portraits in the album, or landscape tags proposed when searching for scenery.
[0041] Candidate tags can also be tags supported by the image classification function. For example, the terminal album can perform image recognition based on all or some of the existing images on the local device, classify the images according to the image recognition results, and assign a category name to each category. Then, the category name can be used as a candidate tag.
[0042] Candidate tags can also be search terms that meet preset search conditions in image search. These preset search conditions can be those that satisfy high-frequency search terms, specifically including a search frequency greater than a preset frequency threshold, or a preset number of search terms ranking highly in terms of frequency. For example, if a user frequently searches for "dog," "cat," "rabbit," and "rose" when performing image searches, then the top three most frequent search terms, such as "dog," "cat," and "rabbit," can be used as candidate tags. The settings for the preset frequency threshold or the preset number of top-ranked terms, as well as other preset search conditions, can be defined according to actual needs, and this disclosure does not impose specific limitations on them.
[0043] Candidate tags can also be words obtained from segmenting the image description. For example, if an image is analyzed and the description is "a girl walking her dog by a lake," then segmenting it can yield words such as "girl," "lakeside," "walking dog," or "dog." Candidate tags can also be obtained based on these segments. For instance, all segments can be used as candidate tags, the top-ranking segments can be used, or a few segments can be randomly selected. The analyzed image can come from various databases, such as caption databases or other image databases; this disclosure does not impose specific limitations on this.
[0044] In one exemplary embodiment, the above-described filtering of visual labels from candidate labels includes at least one of the following filtering methods:
[0045] Filter out candidate labels that do not have visual features from the candidate labels;
[0046] Determine the visual discriminability between different candidate labels. If the visual discriminability between multiple candidate labels is lower than the visual discriminability threshold, then remove duplicates from these multiple candidate labels.
[0047] Filter out candidate tags whose commonness is lower than a preset commonness threshold;
[0048] Filter out candidate tags that do not have a specific purpose from the candidate tags.
[0049] To improve the accuracy and effectiveness of the obtained visual labels, after obtaining candidate labels, they can be screened according to specific criteria to select the visual labels required by this exemplary embodiment from the candidate labels. The specific screening method can include at least one of the above methods, that is, screening can be performed by any one of the above methods, or screening can be performed by multiple methods together. When multiple methods are selected for screening, each screening process can be performed in a certain order, or the screening process can be performed simultaneously, etc. This disclosure does not make specific limitations in this regard.
[0050] Specifically, candidate labels that do not have visual features are filtered out from the candidate labels. For example, the candidate label "cat" can be retained as a visual label, while labels such as time, location without specific markers, person, or other labels that cannot be determined by purely visual features need to be filtered out, such as "today", "Shanghai", and "I".
[0051] Determine the visual distinguishability between different candidate labels. If the visual distinguishability between multiple candidate labels is lower than the visual distinguishability threshold, then deduplicate these multiple candidate labels. The visual distinguishability threshold can be a criterion used to determine whether two objects are easily distinguishable. When the visual distinguishability of two objects is higher than the threshold, it means they are easily distinguishable and can be used as two labels. When the visual distinguishability of two objects is lower than the threshold, it means they are not easily distinguishable and may be the same or similar objects. For example, "Poodle" and "Teddy Dog" have low visual distinguishability and point to the same category. In this case, multiple candidate labels can be deduplicated, retaining only "Poodle" or only "Teddy Dog," etc.
[0052] Considering that visual labels should have a certain degree of commonality, candidate labels with a commonality level below a preset commonality threshold can be filtered out. This process can be considered as filtering very rare objects, scenes, or names that do not conform to daily habits. For example, "royo racing" is relatively rare and will be filtered out, while "cowboy" is relatively more common and can be retained; "Chihuahua" is very rare and will be filtered out, while "Chihuahua" is a common name and can be added, etc. In this exemplary embodiment, whether a candidate label is common can be characterized by the frequency index of the label's occurrence. The preset commonality threshold is the frequency threshold. The judgment of common words or uncommon words can be determined statistically based on the frequency of words appearing in the existing word library or the number of times objects in the image library appear. For example, in the existing word library, when the frequency of a certain word is lower than the preset commonality threshold, the word is considered to have a commonality level below the preset commonality threshold and is therefore an uncommon word. In addition, the determination of common or uncommon words can be achieved by statistically analyzing the interval between occurrences of the same word over a period of time. If the interval exceeds a preset duration, it indicates that the word does not appear frequently and is considered an uncommon word. The preset commonness threshold can then be characterized by the interval between occurrences. When the commonness of a candidate tag is lower than the preset commonness threshold, it must be filtered out in this step.
[0053] Furthermore, visual tags typically possess search or recall value for users; that is, users may need to locate images related to the tag in certain situations for recall, sharing, etc. Therefore, candidate tags that lack a specific purpose can be filtered out. Specific purposes can include sharing, recording, etc. For example, various objects or products fit the sharing purpose, and screenshots fit the recording purpose, and thus can be retained.
[0054] Step S220: Generate nodes based on visual tags, and generate edges between nodes based on the relationships between visual tags, in order to construct a visual tag knowledge graph.
[0055] This exemplary embodiment uses visual tags as nodes and the relationships between visual tags as edges connecting the nodes to construct a visual tag knowledge graph. In the visual tag knowledge graph, each node, or entity, can represent a distinguishable category tag. This can be a concrete visual feature, such as a pet, or an abstract visual feature concept, such as a wedding. Each edge can represent a relationship between categories, thereby accurately establishing a knowledge database about visual tags through the visual tag knowledge graph. In this exemplary embodiment, each node can have a corresponding ID (unique identifier).
[0056] Step S230: Add node attributes to nodes based on the characteristics of the visual labels, and add edge attributes to edges based on the relationship type between the visual labels.
[0057] In this exemplary embodiment, each node and edge in the constructed visual label knowledge graph can have its corresponding attributes, namely node attributes and edge attributes.
[0058] In an exemplary embodiment, the features of the above-mentioned visual tag may include: the category definition of the visual tag, the example image of the visual tag, the visual tag score, the visual tag commonness, the user usage frequency of the visual tag, the foreign language name of the visual tag, the scenario of the visual tag, the sentence component to which the visual tag belongs, the part of speech of the visual tag, the error correction word of the visual tag, the synonym of the visual tag, the near-synonym of the visual tag, and the associated search terms of the visual tag.
[0059] Add node attributes to nodes based on the characteristics of the visual labels, including:
[0060] Add one or more features of the visual label to the associated data of the node as node attributes.
[0061] This exemplary embodiment can first determine multiple features of visual labels, and then add one or more of the features of the visual labels to the associated data of the nodes as node attributes, that is, configure corresponding feature attributes for each node. The specific features of the visual labels and their descriptions are shown in Table 1 below:
[0062] Table 1
[0063]
[0064]
[0065] In an exemplary embodiment, the knowledge graph construction method described above may further include the following steps:
[0066] Identify objects in an image and determine the proportion of each object in the image;
[0067] If the proportion of an object in the image exceeds a preset threshold, the category definition of the visual label is determined based on the object.
[0068] Images typically include multiple objects or elements. For example, an image containing both a cat and a dog might be assigned a visual tag for the cat and a visual tag for the dog. To improve the effectiveness of image visual tags and ensure the reliability of knowledge graph construction, this exemplary embodiment can first identify objects in the image, such as people, plants, buildings, and animals, and then determine their proportion in the image. When the proportion of an identified object in the image exceeds a preset threshold, a visual tag can be determined based on that object. Figure 3 In the image shown, a cat is identified. When the cat's width and height occupy a certain proportion of the image, it is considered to have a significant position in the image, and its visual label can be defined as "cat," etc. Figure 4 In the image shown, the cat is obscured by an obstacle, and its proportion in the image does not meet the preset rules. Therefore, it is not required to assign a visual label to the cat. The proportion of an object in the image can include the ratio of the area of the object to the area of the image, or the ratio of the object's width and height to the width and height of the image, or the ratio of the number of pixels in the object's area to the number of pixels in the image, etc.
[0069] The features of some of the visual labels mentioned above can be obtained in the manner shown in Table 2 below:
[0070] Table 2
[0071]
[0072]
[0073] When combined with manual review, this process can be used to verify the correctness of existing relationships and determine whether new relationships need to be added. The process may include multiple raters voting on the same attribute or relationship; if more than a preset threshold percentage of raters agree with the attribute or relationship, then the attribute / relationship is considered correct. To ensure the quality of the results, this exemplary embodiment can also calculate the consistency of the raters, i.e., calculate the ICC (Interclass Correlation Coefficient) between each rater and the others. If it exceeds a preset threshold, the rater's result is considered acceptable; otherwise, the rater's result is discarded. The preset threshold can be set according to actual needs; for example, the preset threshold for ICC can be set to 0.4.
[0074] In an exemplary embodiment, the aforementioned relationship types include: hypernyms, hyponyms, related words, inference mapping words, and similar tags;
[0075] Add edge attributes to the edges based on the relationship type between the visual labels, including:
[0076] Based on the relationship type between the two visual labels, determine the category information of the edge between the two nodes corresponding to the two visual labels, and add the edge category information to the edge attribute.
[0077] The category information of the edge between the two nodes corresponding to the two visual labels refers to the information determined from the defined relationship types that matches the relationship between the two nodes. This information reflects the relationship type, the nature of the relationship, and the degree of association between the two nodes. This exemplary embodiment can determine the category information of the edge between the two nodes corresponding to the two visual labels based on the relationship type, and then add the edge category information to the edge's attributes. In this exemplary embodiment, the edge between the two nodes can be a directed edge, and the corresponding relationship type can be a hypernym or hyponym, etc.; the edge between the two nodes can also be a non-directed edge, and the corresponding relationship type can be a related word, etc.
[0078] In one exemplary embodiment, the edge attributes may further include label conditional probabilities, and the knowledge graph construction method may further include:
[0079] Calculate the conditional probability that when one of the two visual labels appears, the other also appears, to obtain the label conditional probability of the two visual labels, and add the label conditional probability to the edge attribute of the edge between the two nodes corresponding to the two visual labels.
[0080] In this exemplary embodiment, in addition to qualitatively reflecting the relationship type of the edge between two nodes through the above-mentioned correlation, the relationship between the edge between two nodes can also be actually measured by quantitative calculation and statistical analysis of the conditional probability that when one of the two visual labels appears, the other also appears.
[0081] The specific relationship types and descriptions are shown in Table 3 below:
[0082] Table 3
[0083]
[0084] In one exemplary embodiment, the relationship type of the aforementioned partial edges can be obtained as shown in Table 4 below:
[0085] Table 4
[0086]
[0087] In summary, this exemplary embodiment acquires multiple visual tags and determines the relationships between them; it generates nodes based on the visual tags and edges between nodes based on the relationships between them to construct a visual tag knowledge graph; it adds node attributes to nodes based on the characteristics of the visual tags and adds edge attributes to edges based on the relationship types between the visual tags. On one hand, this exemplary embodiment proposes a method for constructing a knowledge graph, which can use visual tags as nodes and the relationships between visual tags as edges to construct an accurate and reliable visual tag knowledge graph, thereby providing a trustworthy processing method for the visual tag domain. On the other hand, this exemplary embodiment adds corresponding attributes to visual tag nodes and the relationship types between visual tags to further improve the visual tag knowledge graph and enrich its knowledge information. Furthermore, this exemplary embodiment can construct a visual tag knowledge graph for visual tag processing through a simple and convenient process, and has a wide range of applications.
[0088] After generating the visual tag knowledge graph, on the one hand, the visual tag knowledge graph can structure the relevant knowledge of visual tags and apply it to various application scenarios to improve the utilization efficiency and coverage of visual tags; on the other hand, in this exemplary embodiment, when adding node attributes to the nodes of each visual tag according to the characteristics of the visual tag, different effects can be achieved according to the different characteristics. For example, the feature of synonyms of visual tags can expand the coverage of tag recognition capabilities without spending additional model computing resources; the error correction words and the associated search terms of visual tags can provide error correction and intent guessing capabilities for tags; the sentence component features to which the visual tag belongs can provide content-controllable image description capabilities. Compared with the existing technology of directly generating the entire descriptive text information through the caption model (image description model), the tagging model (tag model) outputs tags and sorts them according to components, which is more controllable. The foreign language name features of visual labels can be used for comparison with other databases or models, and for outputting foreign language labels. Furthermore, when adding edge attributes to edges based on the relationship types between visual labels, the corresponding effects can be achieved in different application scenarios for different relationship types. For example, the relationship type of inference mapping words can expand the coverage of label recognition capabilities without spending additional model computing resources; the relationship type of related words can provide search append recommendation word candidates, providing prior knowledge for label recognition capabilities, thereby improving recognition accuracy.
[0089] The visual tag knowledge graph generated above can be applied to scenarios involving search intent association. Specifically, when users conduct keyword searches, due to spelling errors, differences in expression, or other reasons, the entered keywords may not be visually relevant, or the keywords may differ from the actual search intent. Through the node attributes in the visual tag knowledge graph of this exemplary embodiment, such as correction words for visual tags and associated search terms for visual tags, the user's search intent can be associated, significantly improving the coverage of search results and user experience. For example, based on correction words, "ID card" can be determined based on "sfz" or "ID card certificate"; based on associated search terms for visual tags, a medical testing query code can be determined based on "medical testing," a grade certificate can be determined based on "level four," and an ID card image can be determined based on the ID card number, etc.
[0090] The visual tag knowledge graph generated above can also be applied to search term recommendation scenarios. After a user enters a search term and receives search results, they may add keywords for further searching. These keywords are mostly highly related to the already output keywords, so the visual tag knowledge graph can be used to recommend additional keywords. For example, based on "food," one can associate it with "barbecue" and "hot pot"; based on "bag," one can associate it with "yellow," "backpack," and "shoulder bag"; based on "running," one can associate it with "gym" and "park"; based on "dog," one can associate it with "grass" and "walking the dog," and so on.
[0091] Furthermore, this exemplary embodiment also provides another image description method, such as... Figure 5 As shown, the specific process may include the following:
[0092] Step S510: Obtain the visual tag knowledge graph, which is constructed by the knowledge graph construction method described above;
[0093] Step S520: Based on the recognition results of the image to be described, obtain the main label of the image to be described;
[0094] Step S530: Map the main label using the visual label knowledge graph to obtain the mapped label;
[0095] Step S540: Generate descriptive information for the image to be described based on the main label or a tag phrase formed by the main label and the mapping label.
[0096] The image to be described refers to the image whose content is to be described through word segmentation, keywords, or tags. It can be any image that requires description. For example, in an accessible labeling scenario, tags reflecting the image content are typically extracted from the image to be described and broadcast via voice or other means to meet the reading needs of people with special needs. The image to be described is the image to be described. The main tag refers to some or all of the tags identified in the image to be described. For example, from an image to be described, we can identify "person," "wedding dress," "wedding," "skirt," "clothes," "man," "woman," etc., all of which can be broadcast as main tags in an accessible labeling scenario. The main tag can be various nouns as shown above. Furthermore, considering the diversity of content in an image, the tag can also be a verb, such as "walking the dog," "strolling," etc. This exemplary embodiment can pre-train an image recognition model to process the image to be described to obtain the initial main tags. Typically, when the image to be described is complex, the initial main labels will be numerous and disorganized. Directly outputting them in this case may affect the user's judgment of the true meaning of the image being described, resulting in a poor user experience. For example, in the scenario of accessible label broadcasting, the accessible broadcasting function for visually impaired people needs to broadcast the visual labels contained in the image, such as labels for people, objects, scenes, or events. When there are many and disorganized labels, users may not be able to obtain image information directly or quickly.
[0097] At this point, the main label can be processed using the visual label knowledge graph obtained in step S510, such as filtering, sorting, deduplication, or scoring.
[0098] Specifically, after obtaining the initial image labels, the aforementioned visual label knowledge graph can be used to map the main labels to obtain mapped labels. Mapped labels can include inference labels with a causal relationship to the main labels, such as mapping "dog walking" to "dog"; they can also include inference labels with a hierarchical relationship to the main labels, such as "dress" being the parent label of "clothes". The obtained mapped labels may be existing main labels; for example, during image recognition, the main label "dog" has already been identified, but during mapping, the label "dog" is also inferred based on "dog walking". Mapped labels may also be new main labels, which can be retained for later use. During inference mapping, the labels and related information of the hierarchical relationships of each main label can be configured in the attributes of the main labels for subsequent use or to improve the attributes of the labels. For example, each main label can include attributes related to its parent or child labels, or the part-of-speech tag.
[0099] In this exemplary embodiment, to ensure the rationality of the output labels and make them easy for users to understand, the main label can be replaced with a user-friendly label after it is obtained. For example, when determining "Poodle," since users commonly refer to it as "Teddy Dog," it can be replaced with "Teddy Dog." Then, the mapping label of the main label is obtained based on the edge attributes in the aforementioned visual label knowledge graph.
[0100] Finally, descriptive information for the image to be described can be generated based on the main tag or a tag phrase formed by the main tag and the mapping tag. For example, when the mapping tag is a parent tag of an existing main tag, the existing main tag can be updated based on the mapping tag, and the updated main tag can be used as the descriptive information for the image to be described. When the mapping tag is different from the existing main tag and is a new tag, the main tag and the mapping tag can be used together as the descriptive information for the image to be described. Alternatively, descriptive information for the image to be described can be generated based on a tag phrase formed by the main tag and the mapping tag.
[0101] In scenarios where tags can be broadcast in an accessible manner, the broadcast can be based on the descriptive information of the image to be described. For example, in a tag phrase formed by the main tag and the mapped tag, only the main tag can be broadcast to avoid broadcasting multiple similar tags, which would result in a cluttered broadcast effect.
[0102] In summary, in this exemplary embodiment, a visual tag knowledge graph is obtained, constructed by the aforementioned knowledge graph construction method; based on the recognition result of the image to be described, the main tag of the image to be described is obtained; the main tag is mapped using the visual tag knowledge graph to obtain a mapped tag; and descriptive information of the image to be described is generated based on the main tag or a tag phrase formed by the main tag and the mapped tag. This exemplary embodiment proposes an image description method that can infer the mapped tag of the main tag through the logical relationship between tags in the constructed visual tag knowledge graph, and determine the descriptive information based on the main tag or the main tag and the mapped tag. Compared with obtaining descriptive information directly through image recognition, where the tags in the descriptive information lack hierarchy and the relationships between tags are chaotic, the descriptive information obtained by this exemplary embodiment is more concise, accurate, and reliable.
[0103] In an exemplary embodiment, step S540 described above may include:
[0104] Sentence components of main tags or mapping tags are obtained from the visual tag knowledge graph. If there is a main tag or mapping tag whose sentence component is a predicate, the main tags or tag phrases are sorted according to the sentence components of the main tags or mapping tags, and descriptive information is generated based on the sorted main tags or tag phrases.
[0105] When describing an image, descriptive tags can typically include various components, such as verbs and nouns. Different components can form different sentence elements, such as subject, predicate, and object. For example, a description of an image could be "A girl is reading a book on the lawn," where "girl" is the subject, "lawn" is an adverbial, "reading" is the predicate, and "book" is the object. This exemplary embodiment can obtain the sentence components of the main tag or mapping tag and sort them based on these components to generate descriptive information. For instance, if the main tag or mapping tag contains a predicate, it can be sorted in the order of subject, predicate, and other components. If the main tag or mapping tag does not contain a predicate, it can be sorted in a random order or in order of subjective scores from high to low. In scenarios where tags are broadcast in an accessible manner, tags can be broadcast based on the sorted descriptive information to ensure the logical consistency of the tag broadcast and facilitate user understanding of the image information.
[0106] In an exemplary embodiment, the edge attributes of the aforementioned visual label knowledge graph include hypernyms; the mapping labels include first-level mapping labels; the aforementioned step S520 may include:
[0107] In the visual tag knowledge graph, find the outgoing edges of the main tag that include the hypernym attribute, and use the nodes corresponding to the outgoing edges as the first-level mapping tags of the main tag.
[0108] In a visual tag knowledge graph, some edges between nodes have direction. For example, edges with a hierarchical relationship can have direction. When pointing from node A to node B, the edge connecting the two nodes is an outgoing edge of node A and an incoming edge of node B. In this exemplary embodiment, outgoing edges of the main tag with edge attributes including hypernyms can be searched in the visual tag knowledge graph. The node corresponding to the outgoing edge is then used as the first-level mapping tag of the main tag. That is, based on the relationship between the edges between nodes, the node with a superior relationship to that node is searched. The first-level mapping tag can be considered as the hypernym of the current main tag.
[0109] Considering different hierarchical relationships, the main tag may include multiple superordinate terms. In an exemplary embodiment, the mapping tag may further include second-level mapping tags to N-level mapping tags, where N is a positive integer not less than 2; step S520 may further include:
[0110] Starting with the first-level mapping tag of the main tag, the visual tag knowledge graph is used to map each level of the main tag to obtain the next level of mapping tag of the main tag.
[0111] In other words, this exemplary embodiment can find the multi-level mapping tags of the current main tag in the visual knowledge graph. The specific mapping tags to be obtained can be set according to actual needs. The multi-level mapping tags can be continuous. For example, the first-level mapping tag of "wedding dress" is "dress", the second-level mapping tag is "skirt", and the third-level mapping tag is "clothes". All the parent mapping tags of "wedding dress" can be found, or only the second-level and third-level mapping tags can be selected. This disclosure does not make specific limitations in this regard.
[0112] In an exemplary embodiment, before generating descriptive information for the image to be described based on the main tag or a tag phrase formed by the main tag and the mapping tag, the image description method may further include:
[0113] If the main tag does not conform to the word rules of image description, then if the first-level mapping tag of the main tag conforms to the word rules of image description, the first-level mapping tag of the main tag will replace the main tag.
[0114] The word rules for image description refer to the preset rules for words that can be output as main tags. Considering that recognition deviations, over-recognition, or overly coarse recognition may occur during image recognition, this exemplary embodiment can pre-define the word rules for image description to ensure the rationality of the output words. That is, not all main tags can be output to the user for image description. For example, if the main tag for the image to be described is "white wedding dress", but this main tag does not conform to the word rules for image description, and the first-level to third-level mapping tags for "white wedding dress" can include ("wedding dress", "skirt", "clothes"), then if the first-level mapping tag "wedding dress" of the main tag conforms to the word rules for image description, the first-level mapping tag of the main tag can be used to replace the main tag.
[0115] In an exemplary embodiment, before generating descriptive information for the image to be described based on the main tag or a tag phrase formed by the main tag and the mapping tag, the image description method may further include:
[0116] Remove duplicates from the main tag, but not from tag phrases.
[0117] Finally, duplicate tags with the same main tag can be deduplicated to ensure the validity of the output tags. Tag phrases can be deduplicated, for example, if the main tag "evening dress" has a mapping tag ("clothes") and the main tag "wedding dress" also has the same mapping tag ("clothes"), then "evening dress" ("clothes") and "wedding dress" ("clothes") do not need to be deduplicated.
[0118] Figure 6 A flowchart of an image description method in this exemplary embodiment is shown, which may specifically include:
[0119] Step S610: Obtain the visual tag knowledge graph;
[0120] Step S620: Based on the recognition results of the image to be described, obtain the main label of the image to be described;
[0121] Step S630: Replace the main label of the image to be described with a user-friendly label word;
[0122] Step S640: Map the main label using the visual label knowledge graph to obtain the upper-level mapping label or the causal mapping label; wherein the upper-level mapping label may include multi-level mapping labels, and the causal mapping label may include labels that have a causal reasoning relationship with the main label, such as "walking the dog" can infer "dog".
[0123] Step S650: Obtain sentence components of the main tag or mapping tag from the visual tag knowledge graph, and judge the sentence components;
[0124] Step S660: If there is a main label or mapping label that is a predicate in the sentence, sort the main label or label phrase according to the sentence components of the main label or mapping label.
[0125] Step S670: If there is no main label or mapping label that is a predicate in the sentence, the order of the main label or label phrase is determined according to the subjective score.
[0126] Step S680: Remove duplicates from the main tag, but do not remove duplicates from tag phrases;
[0127] Step S690: Generate descriptive information for the image to be described based on the main label or a tag phrase formed by the main label and the mapping label.
[0128] The following example illustrates the application scenario of deduplication of accessibility broadcast labels. For an image to be described, the main labels shown in Table 5 below can be obtained through an image recognition model:
[0129] Table 5
[0130] Main tag Confidence aldult 0.777 people 0.933 wedding 0.709 wedding dress 0.846 skirt 0.846 clothing 0.846 man 0.727 woman 0.705
[0131] The main labels are output in a random order. These main labels may include their own mapping labels, such as "wedding dress", "skirt", "clothes", etc. The main labels and mapping labels have the same confidence level.
[0132] In this exemplary embodiment, when the main label is processed using the image description method described above, it can also be sorted according to the priority order of sentence components, subjective scores, and label confidence levels to obtain the processed main label and label phrases. Specifically, when the sentence components include a predicate, the sorting can be based on the order of subject, predicate, and others. When the orange component does not include a predicate, the sorting can be based on subjective scores. The criteria for subjective scoring can be set as needed, for example, based on Table 6 below:
[0133] Table 6
[0134] Sentence components Rate Subject 1 (Mapping 1, Mapping 2) 1 Subject 2 2 Predicate 1 3 Verb-object phrase 1 4 Object 1 (Mapping 3) 5 Object 2 6 Adverbial 1 7
[0135] After processing the main labels in Table 5 above, we can obtain the results in Table 7 below:
[0136] Table 7
[0137] main tag or tag phrase Confidence Man (person) 0.727 Woman (person) 0.705 Adult (person) 0.777 Wedding dress (skirt, dress) 0.846 wedding 0.709
[0138] In this context, "person" serves as a superordinate mapping to "man," "woman," and "adult," and can form tag phrases with these terms. Similarly, "skirt" and "clothes" serve as superordinate mappings to "wedding dress," and can form tag phrases with "wedding dress." These mapping tags can be added after the main tag and are not listed separately. Furthermore, the main tags can be spoken in the ordered sequence. It should be noted that tag phrases can only be spoken as non-mapping main tags; for example, only "man," "woman," "adult," and "wedding dress" can be spoken, without speaking "skirt" and "clothes." This demonstrates that in this scenario, by integrating the mapping tags with the main tag and outputting the fine-grained main tag as the speaking tag, without outputting the superordinate mapping tags, the accuracy and effectiveness of the information description in the image broadcast can be further guaranteed.
[0139] Exemplary embodiments of this disclosure also provide a knowledge graph construction apparatus. For example... Figure 7 As shown, the knowledge graph construction device 700 may include: a tag acquisition module 710, used to acquire multiple visual tags and determine the relationship between the visual tags; a graph generation module 720, used to generate nodes based on the visual tags and generate edges between nodes based on the relationship between the visual tags, so as to construct a visual tag knowledge graph; and an attribute addition module 730, used to add node attributes to nodes according to the features of the visual tags and add edge attributes to edges according to the relationship type between the visual tags.
[0140] In one exemplary embodiment, the tag acquisition module includes: a tag acquisition unit, configured to acquire candidate tags, the candidate tags including at least one of the following: requirement tags in visual services, tags supported by image classification functions, search terms that meet preset search conditions in image search, and words obtained by segmenting image description statements; and a tag filtering unit, configured to filter out visual tags from the candidate tags.
[0141] In an exemplary embodiment, the label filtering unit is configured to perform at least one of the following filtering methods: filtering out candidate labels that do not have visual features from candidate labels; determining the visual distinguishability between different candidate labels, and if the visual distinguishability between multiple candidate labels is lower than a visual distinguishability threshold, deduplicating the multiple candidate labels; filtering out candidate labels whose commonness is lower than a preset commonness threshold from candidate labels; and filtering out candidate labels that do not have a specific purpose from candidate labels.
[0142] In an exemplary embodiment, the features of the visual tag include: a category definition of the visual tag, an example image of the visual tag, a visual tag score, a visual tag commonness, a user usage frequency of the visual tag, a foreign language name of the visual tag, a scene of the visual tag, a sentence component to which the visual tag belongs, a part of speech of the visual tag, a correction word of the visual tag, a synonym of the visual tag, a near-synonym of the visual tag, and an associated search term of the visual tag; the attribute adding module includes: a node attribute adding unit, used to add one or more of the features of the visual tag to the associated data of the node as node attributes.
[0143] In one exemplary embodiment, the knowledge graph construction apparatus further includes: a category definition determination unit, configured to identify objects in an image and determine the proportion of objects in the image; if the proportion of objects in the image exceeds a preset threshold, then determine the category definition of visual labels based on the objects.
[0144] In an exemplary embodiment, the relationship types include: hypernyms, hyponyms, related words, inference mapping words, and similar tags; the attribute adding module includes: an edge attribute adding unit, used to determine the category information of the edge between the two nodes corresponding to the two visual tags according to the relationship type to which the relationship between the two visual tags belongs, and add the category information of the edge to the edge attribute.
[0145] In an exemplary embodiment, the edge attribute further includes label conditional probability; the knowledge graph construction apparatus further includes: a conditional probability determination unit, used to calculate the conditional probability that when one of the two visual labels appears, the other also appears, to obtain the label conditional probability of the two visual labels, and to add the label conditional probability to the edge attribute of the edge between the two nodes corresponding to the two visual labels.
[0146] Exemplary embodiments of this disclosure also provide an image description device, such as... Figure 8 As shown, the image description device 800 may include: a knowledge graph acquisition module 810, used to acquire a visual tag knowledge graph, which is constructed by the aforementioned knowledge graph construction method; a main tag acquisition module 820, used to obtain the main tag of the image to be described based on the recognition result of the image to be described; a mapping tag acquisition module 830, used to map the main tag using the visual tag knowledge graph to obtain a mapping tag; and a description information acquisition module 840, used to generate description information of the image to be described based on the main tag or a tag phrase formed by the main tag and the mapping tag.
[0147] In an exemplary embodiment, the node attributes of the nodes in the visual tag knowledge graph include sentence components; the description information acquisition module includes: a sorting unit, used to acquire sentence components of main tags or mapping tags from the visual tag knowledge graph; if there is a main tag or mapping tag whose sentence component is a predicate, the main tag or tag phrase is sorted according to the sentence components of the main tag or mapping tag, and description information is generated based on the sorted main tag or tag phrase.
[0148] In an exemplary embodiment, the edge attributes of the visual tag knowledge graph include hypernyms; the mapping tags include first-level mapping tags; the mapping tag acquisition module includes: a first mapping unit, used to find outgoing edges of the main tag in the visual tag knowledge graph whose edge attributes include hypernyms, and to use the nodes corresponding to the outgoing edges as the first-level mapping tags of the main tag.
[0149] In an exemplary embodiment, the mapping tag further includes secondary mapping tags to N-level mapping tags, where N is a positive integer not less than 2; the mapping tag acquisition module further includes: a second mapping unit, used to map each level mapping tag of the main tag sequentially using the visual tag knowledge graph, starting from the first-level mapping tag of the main tag, to obtain the next-level mapping tag of the main tag.
[0150] In one exemplary embodiment, the image description apparatus further includes: a replacement unit, configured to, before generating description information of the image to be described based on the main label or a tag phrase formed by the main label and the mapping label, replace the main label with the first-level mapping label of the main label if the main label does not conform to the word rules of image description, provided that the first-level mapping label of the main label conforms to the word rules of image description.
[0151] In one exemplary embodiment, the image description apparatus further includes a deduplication unit, configured to deduplicate the main label and not deduplicate the tag phrase before generating description information of the image to be described based on the main label or a tag phrase formed by the main label and the mapping label.
[0152] The specific details of each part of the above-mentioned device have been described in detail in the method section of the implementation, and therefore will not be repeated here.
[0153] Exemplary embodiments of this disclosure also provide a computer-readable storage medium, which can be implemented as a program product, including program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure, such as executing... Figure 2 , Figure 5 or Figure 6 The program product may be a portable compact disc read-only memory (CD-ROM) containing program code and may run on a terminal device, such as a personal computer. However, the program product disclosed herein is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in connection with an instruction execution system, apparatus, or device.
[0154] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory, read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0155] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0156] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0157] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0158] Exemplary embodiments of this disclosure also provide an electronic device. The electronic device may include a processor and a memory, the memory storing executable instructions for the processor, and the processor configured to execute the aforementioned knowledge graph construction method and image description method by executing the executable instructions.
[0159] The following is based on Figure 9 Taking a mobile terminal 900 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 9 The structure can also be applied to fixed types of equipment.
[0160] like Figure 9 As shown, the mobile terminal 900 may specifically include: a processor 901, a memory 902, a bus 903, a mobile communication module 904, an antenna 1, a wireless communication module 905, an antenna 2, a display screen 906, a camera module 907, an audio module 908, a power module 909, and a sensor module 910.
[0161] Processor 901 may include one or more processing units, such as: AP (Application Processor), modem processor, GPU (Graphics Processing Unit), ISP (Image Signal Processor), controller, encoder, decoder, DSP (Digital Signal Processor), baseband processor and / or NPU (Neural-Network Processing Unit), etc.
[0162] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 900 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0163] The processor 901 can be connected to the memory 902 or other components via the bus 903.
[0164] The memory 902 can be used to store executable program code, which includes instructions. The processor 901 executes various functional applications and data processing of the mobile terminal 900 by running the instructions stored in the memory 902. The memory 902 can also store application data, such as images, videos, and other files.
[0165] The communication functions of the mobile terminal 900 can be implemented through a mobile communication module 904, antenna 1, a wireless communication module 905, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 904 can provide 3G, 4G, and 5G mobile communication solutions for use on the mobile terminal 900. The wireless communication module 905 can provide wireless communication solutions such as Wi-Fi, Bluetooth, and near-field communication for use on the mobile terminal 900.
[0166] The display screen 906 is used to implement display functions, such as displaying the user interface, images, and videos. The camera module 907 is used to implement shooting functions, such as capturing images and videos. The audio module 908 is used to implement audio functions, such as playing audio and capturing voice. The power module 909 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 910 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 910 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 900 and output inertial sensing data.
[0167] Those skilled in the art will understand that various aspects of this disclosure can be implemented as systems, methods, or program products. Therefore, various aspects of this disclosure can be embodied in entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuit,” “module,” or “system.” Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0168] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is defined only by the appended claims.
Claims
1. A method for constructing a knowledge graph, characterized in that, include: Obtain multiple visual labels and determine the relationships between the visual labels; Nodes are generated based on the visual tags, and edges are generated between the nodes based on the relationships between the visual tags, in order to construct a visual tag knowledge graph; Node attributes are added to the nodes based on the characteristics of the visual labels, and edge attributes are added to the edges based on the relationship type between the visual labels; The acquisition of multiple visual labels includes: Obtain candidate tags, which include at least one of the following: requirement tags in visual business, tags supported by image classification function, search terms that meet preset search conditions in image search, and words obtained by segmenting image description statements. The visual labels are selected from the candidate labels; The step of filtering the visual labels from the candidate labels includes at least one of the following filtering methods: Filter out candidate labels that do not have visual features from the candidate labels; Determine the visual discriminability between different candidate labels. If the visual discriminability between multiple candidate labels is lower than the visual discriminability threshold, then remove duplicates from these multiple candidate labels. Filter out candidate tags whose commonness is lower than a preset commonness threshold from the candidate tags; Filter out candidate tags that do not have a specific purpose from the candidate tags.
2. The method according to claim 1, characterized in that, The features of the visual tags include: the category definition of the visual tag, the example image of the visual tag, the visual tag score, the visual tag commonness, the user usage frequency of the visual tag, the foreign language name of the visual tag, the scene of the visual tag, the sentence component to which the visual tag belongs, the part of speech of the visual tag, the error correction word of the visual tag, the synonym of the visual tag, the near-synonym of the visual tag, and the associated search terms of the visual tag. Adding node attributes to the node based on the features of the visual label includes: One or more features of the visual label are added to the associated data of the node as node attributes.
3. The method according to claim 2, characterized in that, The method further includes: Identify objects in an image and determine the proportion of said objects in the image; If the proportion of the object in the image exceeds a preset threshold, the category definition of the visual label is determined based on the object.
4. The method according to claim 1, characterized in that, The relationship types include: hypernyms, hyponyms, related words, inference mapping words, and similar tags; Adding edge attributes to the edges based on the relationship type between the visual labels includes: Based on the relationship type between the two visual labels, determine the category information of the edge between the two nodes corresponding to the two visual labels, and add the category information of the edge to the edge attribute.
5. The method according to claim 1, characterized in that, The edge attribute also includes label conditional probability; the method further includes: Calculate the conditional probability that when one of the two visual labels appears, the other also appears, to obtain the label conditional probability of the two visual labels, and add the label conditional probability to the edge attribute of the edge between the two nodes corresponding to the two visual labels.
6. An image description method, characterized in that, include: A visual tag knowledge graph is obtained, wherein the visual tag knowledge graph is constructed by the knowledge graph construction method according to any one of claims 1 to 5; Based on the recognition results of the image to be described, the main label of the image to be described is obtained; The main tag is mapped using the visual tag knowledge graph to obtain the mapped tag; Based on the main tag or the tag phrase formed by the main tag and the mapping tag, the description information of the image to be described is generated.
7. The method according to claim 6, characterized in that, The node attributes of the visual tag knowledge graph include sentence components; the generation of descriptive information for the image to be described based on the main tag or the tag phrase formed by the main tag and the mapping tag includes: Sentence components of the main tag or the mapping tag are obtained from the visual tag knowledge graph. If there is a main tag or mapping tag whose sentence component is a predicate, the main tag or the tag phrase is sorted according to the sentence components of the main tag or the mapping tag, and the description information is generated based on the sorted main tag or the tag phrase.
8. The method according to claim 6, characterized in that, The edge attributes of the visual tag knowledge graph include hypernyms; the mapping tags include first-level mapping tags; the process of mapping the main tag using the visual tag knowledge graph to obtain the mapping tags includes: In the visual tag knowledge graph, find the outgoing edge of the main tag that includes the hypernym attribute, and use the node corresponding to the outgoing edge as the first-level mapping tag of the main tag.
9. The method according to claim 8, characterized in that, The mapping labels also include second-level mapping labels to N-level mapping labels, where N is a positive integer not less than 2; The step of mapping the main label using the visual label knowledge graph to obtain the mapped label further includes: Starting from the first-level mapping tag of the main tag, the visual tag knowledge graph is used to map each level mapping tag of the main tag in sequence to obtain the next level mapping tag of the main tag.
10. The method according to claim 8, characterized in that, Before generating descriptive information for the image to be described based on the main tag or a tag phrase formed by the main tag and the mapping tag, the method further includes: If the main tag does not conform to the word rules of image description, then if the first-level mapping tag of the main tag conforms to the word rules of image description, the main tag is replaced by the first-level mapping tag of the main tag.
11. The method according to claim 8, characterized in that, Before generating descriptive information for the image to be described based on the main tag or a tag phrase formed by the main tag and the mapping tag, the method further includes: The main tag is deduplicated, but the tag phrases are not deduplicated.
12. A knowledge graph construction device, characterized in that, include: The tag acquisition module is used to acquire multiple visual tags and determine the relationship between the visual tags; The graph generation module is used to generate nodes based on the visual tags and generate edges between the nodes based on the relationships between the visual tags, so as to construct a visual tag knowledge graph. An attribute adding module is used to add node attributes to the nodes according to the features of the visual labels, and to add edge attributes to the edges according to the relationship type between the visual labels; The acquisition of multiple visual labels is configured as follows: Obtain candidate tags, which include at least one of the following: requirement tags in visual business, tags supported by image classification function, search terms that meet preset search conditions in image search, and words obtained by segmenting image description statements. The visual labels are selected from the candidate labels; The step of filtering the visual labels from the candidate labels is configured to be at least one of the following filtering methods: Filter out candidate labels that do not have visual features from the candidate labels; Determine the visual discriminability between different candidate labels. If the visual discriminability between multiple candidate labels is lower than the visual discriminability threshold, then remove duplicates from these multiple candidate labels. Filter out candidate tags whose commonness is lower than a preset commonness threshold from the candidate tags; Filter out candidate tags that do not have a specific purpose from the candidate tags.
13. An image description device, characterized in that, include: A knowledge graph acquisition module is used to acquire a visual tag knowledge graph, wherein the visual tag knowledge graph is constructed by the knowledge graph construction method according to any one of claims 1 to 5; The main label acquisition module is used to obtain the main label of the image to be described based on the recognition result of the image to be described; The mapping label acquisition module is used to map the main label using the visual label knowledge graph to obtain the mapping label; The description information acquisition module is used to generate description information for the image to be described based on the main label or a tag phrase formed by the main label and the mapping label.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the knowledge graph construction method according to any one of claims 1 to 5 and the image description method according to any one of claims 6 to 11.
15. An electronic device, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the knowledge graph construction method of any one of claims 1 to 5 and the image description method of any one of claims 6 to 11 by executing the executable instructions.