Multimodal data-driven commodity graph construction method and system, and electronic device
The product atlas construction method driven by multimodal data addresses the shortcomings of single-modal data in traditional methods. By extracting image boundary features and verifying field matching, coreference conflicts are identified and nodes are reinforced, thereby improving the semantic accuracy and structural integrity of the product atlas.
Patent Information
- Application Number
- CN202511164700.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Traditional product atlas construction methods lack fine-grained structural feature analysis of single-modal data, resulting in incomplete semantics, recognition errors, and inaccurate product classification, which affects the accuracy of attribute recognition and knowledge reasoning. Furthermore, insufficient field completion affects data quality and structural integrity.
By using a multimodal data-driven approach, nodes with incomplete modal structures in the product atlas are obtained, image foreground segmentation and boundary feature extraction are performed, and matching verification is carried out in combination with structured fields to identify coreference conflicts, refine class target labels, and identify missing nodes by field presence rate and standard deviation coefficient. A list of field-enhanced nodes is constructed to ensure the structural consistency of the atlas construction.
It improves the semantic accuracy and entity modeling quality of the product graph, enhances the recognition and distinguishability of class tags, improves field coverage and structural integrity, and ensures the coordination and consistency of the construction process.
Smart Images

Figure CN120744190B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of commodity graph, in particular to a multi-modal data driven commodity graph construction method, system and electronic device. BACKGROUND
[0002] The technical field of commodity graph involves the structured representation and semantic modeling of commodity information. The core lies in constructing a graph structure with commodities as nodes and attribute relationships as edges, realizing the unified modeling, management and semantic understanding of commodities. This technical field covers key links such as commodity entity recognition, attribute extraction, category mapping, commodity deduplication, attribute standardization, and relationship reasoning. The goal is to realize the fusion of commodity data across platforms and categories and the connection of knowledge, supporting downstream applications such as commodity search, recommendation, price comparison, and precision marketing. Current research in this field is gradually expanding in the direction of multi-source heterogeneous data fusion, graph neural network modeling, knowledge enhancement, and incremental graph construction.
[0003] Among them, the multi-modal data driven commodity graph construction method is a commodity graph construction scheme based on the joint participation of text, image, structured data and other modal data. This scheme aims to solve the problems of incomplete semantics and recognition errors caused by relying on a single modality in the traditional commodity graph construction process. By jointly modeling modal information such as commodity description text, commodity pictures, and category labels, accurate extraction of commodity entities, attribute recognition, and category normalization are achieved, and a consistent and complete semantic commodity graph is constructed, which can be widely used in commodity management, semantic search, and intelligent recommendation scenarios on e-commerce platforms.
[0004] Traditional construction methods lack analysis of fine-grained structural features in single modal content during the commodity graph construction process. In particular, in image modal processing, explicit expression of image regions through foreground boundary structure parameters is not achieved, leading to misjudgment in field matching. In addition, current category labels are based on text semantics and historical graph labels, making it difficult to achieve reasonable refinement when category label expressions are ambiguous or boundaries overlap, resulting in inaccurate commodity classification and affecting the accuracy of attribute recognition and knowledge reasoning. In the field completion aspect, it relies on template rules and does not introduce field existence rate analysis to locate key missing fields, resulting in insufficient coverage of structured content and restricting the performance of graph construction in data quality and structural integrity. SUMMARY
[0005] The purpose of the present application is to solve the shortcomings in the prior art, and a multi-modal data driven commodity graph construction method, system and electronic device are proposed.
[0006] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows: a multi-modal data driven commodity graph construction method, comprising the following steps:
[0007] S1: acquire a node of which a modal structure construction in a commodity atlas is not completed, collect a commodity image bound to the node, perform foreground segmentation on the image, aggregate image region features to form a multi-dimensional feature vector, and generate a commodity image boundary feature set;
[0008] S2: based on the commodity image boundary feature set, call an attribute item with geometric meaning in a structured field set associated with the node in the commodity atlas, convert a field value into a descriptive boundary target and map it to an image foreground region, judge whether a matching region exists in the image region corresponding to the target field value, and generate a graph node coreference conflict annotation item;
[0009] S3: based on the graph node coreference conflict annotation item, perform edge density estimation and color cluster number statistics on the target region, judge whether the class target label bound to the current commodity node exists between the vector and the density deviation, and generate a refined class target label candidate item set;
[0010] S4: call the refined class target label candidate item set, calculate the field existence rate and obtain the field standard deviation coefficient for each modal field, judge whether the current commodity node is lower than the average value of the existence rate and falls into the standard deviation positive deviation interval under the field, and generate a graph construction field reinforcement node list.
[0011] The application improves that the commodity image boundary feature set includes an edge detection point set, a boundary continuous region numerical value, a closed region quantity index, a color block quantity value and a boundary line segment length average value, the graph node coreference conflict annotation item is specifically a field item identification list, an image region coordinate index, a field mapping failure mark set and a multi-field overlap conflict identification, the refined class target label candidate item set includes a sub-label set, a dense deviation threshold, a main label identification item and a candidate label sorting table, and the graph construction field reinforcement node list is specifically a reinforcement node number, a field missing modal identification, an existence rate offset table and a standard deviation overrun record.
[0012] The application improves that the acquisition step of the commodity image boundary feature set is specifically:
[0013] S111: acquire a node of which a modal structure construction in a commodity atlas is not completed, collect a commodity image bound to the node, perform foreground segmentation on the image, aggregate image region features to form a multi-dimensional feature vector, and generate a commodity image boundary feature set;
[0014] S112: based on the foreground contour region coordinate value set, calling a Canny edge detection operator to perform contour point extraction on the contour range according to the calibrated region boundary, calculating a contour point density value and a number of boundary continuous regions for each contour region, and identifying a number of complete closed regions based on a boundary closed point arrangement manner, to establish a boundary structure statistical index set corresponding to each image sample;
[0015] S113: calling the boundary structure statistical index set to perform block recognition based on a region color aggregation feature in the contour coordinate range, performing main color proportion analysis and color category distinguishability operation on the recognized region, counting a number of color aggregation blocks, calculating an average value of contour line segment length, aggregating image structure data to form a multi-dimensional feature vector, and generating a commodity image boundary feature set.
[0016] The application improves that the acquisition step of the atlas node co-reference conflict labeling item is specifically:
[0017] S211: based on the commodity image boundary feature set, according to the image structure parameters of each image, calling the attribute field marked as a geometric expression type in the node association field set in the commodity atlas, performing structure semantic analysis on the field values of the shape field, external structure field and material layout field, converting the field values into boundary target label structure, establishing a one-to-one mapping relationship between the field and the boundary target, and generating an attribute boundary matching item set;
[0018] S212: based on the attribute boundary matching item set, according to the boundary target label structure, extracting the corresponding image region coordinate range, judging whether there is an intersection overlap between the coordinate coverage relationship between the regions pointed by the field label structure, sequentially performing region pointing consistency detection on the field label structure, performing region boundary repetition judgment and intersection degree measurement on the mapped field, calculating a field region coincidence coefficient, comparing a field repetition region threshold value, performing conflict screening according to whether the field region pointing structure is exclusive, and generating a field intersection conflict measurement result;
[0019] S213: based on the field intersection conflict measurement result, according to the field region repetition item, the no matching item and the corresponding image coordinate point information, screening the items in the field item that do not correspond to the image boundary or overlap with other field regions, extracting the field item and the corresponding image region coordinate index, establishing an abnormal mapping table of the field pointing within the node, and generating an atlas node co-reference conflict labeling item.
[0020] The application improves that the acquisition step of the refined class target label candidate item set is specifically:
[0021] S311: based on the graph node co-reference conflict annotation item, according to the annotated field corresponding image region coordinates, taking the region boundary coordinates in the image as a limited region, performing contour edge recognition on the commodity image in the limited region, obtaining an edge point density value by dividing the total number of contour edge pixel points by the area of the region, recording the number of edge line segments and the number of edge length change sections of the region, and establishing image region edge density information;
[0022] S312: calling the image region edge density information, according to the corresponding image region, performing main color extraction and difference distinguishing processing on the pixel color channel data in the region, identifying color aggregation regions and counting color principal component clustering values, constructing a double-parameter density feature vector combining the edge density value and the clustering value, calculating to obtain a region density deviation measurement value, comparing the deviation value with a standard density vector corresponding to the class target label bound to the current commodity node, judging whether it exceeds the density recognition threshold, and obtaining density deviation information;
[0023] S313: based on the density deviation information, according to the label item marked as exceeding the recognition threshold, extracting the category hierarchy structure corresponding to the original class target label of the commodity node, obtaining a same layer sub-label set in which the current label is located from the structure, and comparing the deviation value with the preset density standard vector of the sub-label item by item, selecting the label item with the smallest difference value as the main label identification, establishing a mapping relationship between the sub-label set and the main label, and generating a refined class target label candidate item set.
[0024] The application improves that the obtaining step of the graph construction field reinforced node list is specifically:
[0025] S411: calling the refined class target label candidate item set, according to the commodity node set bound to each sub-label, counting the filling state of three modal fields of image field, text field and structure field in each node, calculating the ratio of the number of nodes in which each modal field is effectively filled to the total number of set nodes, and generating a modal field existence rate value set;
[0026] S412: based on the modal field existence rate value set, according to the existence rate of each modal field, calculating the discreteness of the existence distribution of the field among multiple nodes in the set, obtaining the sample standard deviation value of each modal field, and constructing a standard deviation coefficient set together with the field existence rate and the standard deviation, calculating to obtain a field existence deviation factor, judging whether the field existence deviation factor exceeds the field consistency deviation threshold, screening the nodes whose field existence rate is lower than the average value and the deviation factor falls into the positive deviation interval, and obtaining field consistency offset indicator information;
[0027] S413: Based on the field consistency offset indicator information, according to the node item marked as low padding and high offset, the commodity node number and the modal field type information are extracted, the field missing node index table is constructed, and the graph construction field reinforcement node list is generated.
[0028] The method further comprises the following steps:
[0029] S5: Based on the graph construction field reinforcement node list, the field organization order and the node link structure constraint template are extracted from the graph configuration template according to the belonging category target label item, the boundary topology connection table of the commodity entity node is constructed by combining the mapping index of the existing field value in the standard structure, and the commodity graph construction scheduling instruction set is generated.
[0030] The commodity graph construction scheduling instruction set comprises a field organization sequence, a graph boundary connection table, a link structure type mapping and a target graph writing instruction.
[0031] The method further comprises the following steps:
[0032] S511: Based on the graph construction field reinforcement node list, the node filled image field, text field and structure field content are extracted according to the number of each commodity node, and the field name is positionally mapped and numbered according to the standard field sequence order defined in the graph template, so as to correspond each field content to the preset position number in the category structure, and generate the modal field mapping sequence information.
[0033] S512: The modal field mapping sequence information is called, the corresponding field organization order structure and node link structure constraint rule are extracted from the graph configuration template according to the standard position number of each field and the category label bound by the current node, the field connectivity is judged, the connectable field pairs are screened based on the field node allowed connection type matrix, the sequential link structure between fields is established, and the field organization connection graph structure measurement information is generated.
[0034] S513: Based on the field organization connection graph structure measurement information, the field boundary connection direction is calculated and the graph topology position index is calibrated according to the field pair link rule and the field node sequence number, the node field connection edge is sorted according to the link structure, the node boundary topology table is constructed, the structure generation command group is established for each node, and the commodity graph construction scheduling instruction set is generated.
[0035] The multi-modal data driven commodity graph construction system is used to realize the multi-modal data driven commodity graph construction method, and the system comprises:
[0036] The product feature recognition module acquires nodes in the product map that have not completed modal structure construction, collects product images bound to the nodes, performs foreground segmentation on the images, aggregates image region features to form multidimensional feature vectors, and generates product image boundary feature sets.
[0037] The conflict annotation module, based on the product image boundary feature set, calls the geometrically meaningful attribute items in the structured field set associated with nodes in the product atlas, transforms the field values into descriptive boundary targets and maps them to the foreground region of the image, determines whether there is a matching region in the image region corresponding to the target field value, and generates atlas node co-referencing conflict annotation items;
[0038] Based on the core-pointing conflict annotations of the graph nodes, the target refinement analysis module estimates the edge density and counts the number of color clusters in the target area, determines whether there is a density deviation between the target class tag bound to the current product node and the vector, and generates a refined target class tag candidate set.
[0039] The reinforcement node identification module calls the refined class target candidate option set, calculates the field existence rate for each modal field and obtains the field standard deviation coefficient, determines whether the current product node is lower than the mean existence rate under the field and falls into the positive skewed interval of the standard deviation, and generates a graph to construct a list of field reinforcement nodes.
[0040] The scheduling instruction construction module constructs a list of field-enhanced nodes based on the graph, extracts the field organization order and node link structure constraint template from the graph configuration template according to the class target tag, and constructs the boundary topology connection table of commodity entity nodes by combining the mapping index of existing field values in the standard structure, thereby generating a commodity graph construction scheduling instruction set.
[0041] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the multimodal data-driven commodity map construction system as described above.
[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0043] In the present application, by extracting the structural parameters of the foreground region of the commodity image and aggregating the boundary features, multi-dimensional coding of the key region in the image is realized, the matching degree of the field value and the image structure is identified in combination with the field mapping verification mechanism, the co-reference conflict label is generated in the case of matching exception or multiple field overlap, the fine splitting of the class target label is guided based on the regional statistical index, the main label is determined after constructing the candidate label set taking the density deviation as the evaluation basis, the recognition and differentiation of the class target label are effectively enhanced, the field vacancy node is identified by calculating the field existence rate and the standard deviation coefficient and added to the reinforcement queue, the field coverage and the structural integrity in the graph are improved, the scheduling instruction set is formed in combination with the field organization template and the connection structure constraint, the structural consistency and the field coordination in the construction process are ensured, and the semantic accuracy and the entity modeling quality of the commodity graph are improved. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 The method flowchart of the present application is shown in the following figure:
[0045] Figure 2 The flowchart of obtaining the boundary feature set of the commodity image in the present application is shown in the following figure:
[0046] Figure 3 The flowchart of obtaining the co-reference conflict label item of the graph node in the present application is shown in the following figure:
[0047] Figure 4 The flowchart of obtaining the refined class target label candidate set in the present application is shown in the following figure:
[0048] Figure 5 The flowchart of obtaining the graph construction field reinforcement node list in the present application is shown in the following figure:
[0049] Figure 6 The flowchart of obtaining the commodity graph construction scheduling instruction set in the present application is shown in the following figure. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and not to limit the present application.
[0051] In the description of the present application, it should be understood that the orientations or positional relationships indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like are based on the orientations or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, in the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0052] Referring to Figure 1 The application provides a technical solution: a multi-modal data-driven commodity atlas construction method, comprising the following steps:
[0053] S1: Obtain a node in the commodity atlas for which the modal structure construction is not completed, collect a commodity image bound to the node, perform foreground segmentation on the image, extract edge detection points of a boundary contour of a foreground region, a number of boundary continuous regions, a number of closed regions, and label a number of color blocks and an average length of boundary line segments in the extracted region, aggregate image region features to form a multi-dimensional feature vector, and generate a commodity image boundary feature set;
[0054] S2: Based on the commodity image boundary feature set, according to image structure parameters, call attribute items with geometric meanings in a structured field set associated with the node in the commodity atlas, including a shape field, an external structure field, and a material layout field, convert field values into descriptive boundary targets and map them to the image foreground region, determine whether there is a matching region in the image region corresponding to the target field value, if multiple fields correspond to the same image boundary region or the field value cannot match any boundary structure, record the field item and the image coordinate point index, and generate a graph node coreference conflict annotation item;
[0055] S3: Based on the graph node coreference conflict annotation item, according to the field corresponding image region coordinates, limit the region as a category recognition input range, perform edge density estimation and color cluster number statistics on the target region, combine the two statistical values into a density feature vector, determine whether there is a density deviation between the current commodity node bound category label and the vector, if there is a deviation, split the current label corresponding category set into a sub-label set, obtain a main label according to the minimum dense deviation principle, and generate a refined category label candidate item set;
[0056] Edge density estimation refers to the number of image contour edge pixel points in a unit region, which is usually counted in a unit pixel region after detection by Canny, Sobel, etc. Color cluster number statistics can extract the number of color distribution clusters in the foreground region of the image by K-means algorithm, indicating the color complexity of the image detail region.
[0057] S4: Call the refined category label candidate item set, according to the commodity node set corresponding to each sub-label, count the field filling situation of the three types of modal fields of image fields, text fields, and structure fields in the set, calculate the field existence rate of each modal field and obtain the field standard deviation coefficient, determine whether the current commodity node is lower than the average value of the existence rate and falls into the standard deviation positive deviation interval, if the condition is met, add it to the reinforcement queue, and generate a graph construction field reinforcement node list;
[0058] The field existence rate refers to the proportion of the filled nodes to the total nodes of the modal field in the commodity node set, and the standard deviation coefficient is the ratio of the standard deviation to the mean value, which can measure the deviation degree of the field distribution, and the higher the standard deviation coefficient is, the more uneven the distribution of the field in the graph is;
[0059] S5: constructing a field reinforcement node list based on the graph, extracting the modal field content currently possessed by each node, and extracting the field organization order and node link structure constraint template from the graph configuration template according to the belonging class target label item, combining the mapping index of the existing field value in the standard structure, constructing the boundary topology connection table of the commodity entity node, and generating the commodity graph construction scheduling instruction set;
[0060] The field organization order refers to the logical presentation and storage order of the fields, and the node link structure constraint template defines the connection rules and edge types between the nodes of a certain type of commodity;
[0061] The commodity image boundary feature set includes an edge detection point set, a boundary continuous region numerical value, a closed region quantity index, a color block quantity value, and a boundary line segment length average value. The graph node co-reference conflict label item specifically includes a field item identification list, an image region coordinate index, a field mapping failure marker set, and a multi-field overlap conflict identification. The refined class target label candidate item set includes a sub-label set, a dense deviation threshold, a main label identification item, and a candidate label sorting table. The graph construction field reinforcement node list specifically includes a reinforcement node number, a field missing modal identification, an existence rate deviation amount table, and a standard deviation overrun record. The commodity graph construction scheduling instruction set includes a field organization sequence, a graph boundary connection table, a link structure type mapping, and a target graph writing instruction.
[0062] Please refer to Figure 2 The acquisition steps of the commodity image boundary feature set are as follows:
[0063] S111: acquiring the nodes of which the modal structure construction is not completed in the commodity graph, collecting the commodity image corresponding to each node, and performing foreground segmentation based on salient region extraction on each image. After segmentation, the main commodity contour region in the image is located and the background interference information is removed. The main region boundary is judged according to the gray gradient difference range and the region connectivity inside and outside the contour closed region, and the foreground contour region coordinate value set is generated;
[0064] To obtain the node set of which the modal structure in the commodity atlas is not completed, firstly, the atlas database needs to be traversed to screen out nodes with a field integrity score lower than 0.5 and eliminate the nodes that have been labeled, so as to determine the current incomplete atlas entity. The commodity images corresponding to each node are collected, and the image paths are extracted from the image resource library by matching the commodity ID, and then loaded into the processing cache area after decoding. The foreground segmentation operation based on the salient region extraction is performed on the image, the foreground mask is identified by color contrast, the salient weight center point is calibrated on each image, and the center periphery of 200*200 pixels is set as the segmentation window. The pixel points in the window are extracted and mean filtering is performed on them to remove local texture noise. Then, the gray transfer curve is established according to the difference of each row of pixel values, the edge segments with an absolute value of continuous difference greater than 30 are marked, the gray gradient mutation region is recorded, and the gray gradient range is calculated. For example, for IMG001 image, the gray range of this region is 48, and the "gray gradient mutation threshold" set here is 30. The setting basis is the lower limit of the average brightness difference between the image background region and the main commodity. The statistical results of the gray contrast experiment of 50 standard sample images show that the average gray difference between the background and the main commodity edge is 31.6, and the standard deviation is 6.2. Therefore, the value obtained by subtracting the standard deviation from the mean value is taken as the separable critical value, that is, 30. If the change range exceeds 40, it is judged that the region is the main commodity region. Then, the closed region is counted and the connected region judgment is performed. If a single region has a connection with more than 4 pixel blocks, it is defined as an independent connected region. The setting of the connectedness reference value "4" is based on the average statistical value of the boundary block of the main component contour in the furniture commodity. The average number of adjacent contour blocks possessed by the main component boundary is 4.3, so 4 is taken as the judgment threshold for region connectivity recognition. Thus, 5 connected regions are identified in IMG001 image, and the contour structure point sequence is determined as the main contour boundary. Further, the coordinate pairs are obtained according to the boundary start and end points, the main commodity boundary start and end coordinates of each node are extracted, and finally the foreground contour region coordinate value set is generated.
[0065] S112: Based on the foreground contour region coordinate value set, according to the calibrated region boundary, a Canny edge detection operator is called to perform contour point extraction on the contour range. The contour point density value and the number of boundary continuous regions are calculated for each contour region, and the number of complete closed regions is identified based on the arrangement mode of the boundary closed points. The boundary structure statistical index set corresponding to each image sample is established.
[0066] Based on the calibrated boundary range of the foreground contour region coordinate value set, the upper and lower limits of the Canny operator threshold are selected as 50 and 150 respectively, and edge detection operation is performed on all pixel points in the contour region. It is judged whether the 8-neighborhood gray scale gradient of each pixel point exceeds the set lower limit and connects the upper limit critical point to generate an edge line segment. The total number of pixel points of all recognized line segments is counted as the basic data of contour point density. The detected edge point number is divided by the corresponding boundary region pixel number using image width normalization ratio to obtain the normalized contour point density value. For example, the boundary pixel area of the contour region after processing the IMG002 image is 160x140, and the number of edge pixel points is 1720, so its density value is about 0.077. The strike vector angle difference of the contour line segment is calculated. If the angle difference of the line segment is less than 5 degrees continuously, the line segment is merged into a continuous boundary segment. The number of boundary continuous segments is counted as the number of boundary continuous regions. The start and end points of the image edge coordinate point sequence are matched. If the Euclidean distance is less than 4 pixels, it is determined as a closed structure, and the number of closed regions is accumulated. The "closed judgment threshold" is set to 4 pixels. This threshold is estimated by analyzing the average start and end offset of the boundary of 500 closed regions in the sample set. The standard offset range is 2.7 to 3.8 pixels, and the upper limit is taken and increased by 10% tolerance, that is, 4 pixels, to avoid the risk of false merging. Finally, the boundary structure statistical index set is established.
[0067] S113: Call the boundary structure statistical index set, execute block recognition based on regional color aggregation features in the contour coordinate range, perform main color proportion analysis and color category distinguishability operation on the recognized region, count the number of color aggregation blocks, calculate the average length of the contour line segment, aggregate the image structure data to form a multi-dimensional feature vector, and generate the commodity image boundary feature set.
[0068] The structural characteristic values such as the contour point density, the number of boundary continuous regions and the number of closed regions of each image in the call boundary structure statistical index set are called, and color clustering analysis is performed on the corresponding coordinate range, K-means clustering is performed on the image region based on the pixel RGB value, the number of clustering centers is initially set to 4 and iterated to stable clustering centers, and the setting basis of the initial value "4" of the clustering is that, according to the color main region division experiment of common commodity images, it is found that most commodities have background color, main material color, local structure color and shadow area, so 4 is taken as the initial clustering number, and after the clustering of the IMG003 image converges, three color main block regions are finally formed, the main color channel proportion of each aggregation block is counted and the main color proportion is calculated, and the RGB value difference matrix is constructed for the clustering pair with a difference amplitude of more than 20, and the RGB color difference judgment threshold value "20" is set based on the average color span between the human eye resolution limit value and the commodity main color matching, after the color difference comparison test, the recognition accuracy reaches more than 91% when the RGB difference is greater than 20, so the value is set to 20, and the average color difference value between the clusters is counted as the color differentiation value, the contour line segments are sorted according to their lengths and the average value is calculated to obtain the average contour line segment length, the total length of the contour line segments in the IMG003 image is 780 pixels, the number of line segments is 42, and the average value is about 18.6 pixels, finally the color aggregation block number, the main color proportion, the color differentiation, the contour point density, the number of boundary continuous regions, the number of closed regions and the average contour length are constructed into a multi-dimensional vector to generate a commodity image boundary feature set.
[0069] Table 1 shows the processed image in the execution of the boundary extraction and color recognition process. The table lists the key parameter values, including color clustering number, boundary structure parameters and color distribution structure, which provides a basis for filling the modal field in the subsequent graph node construction. The source of each parameter setting is combined with the image data feature distribution and experimental verification process to ensure the structural rigor and implementability of the scheme.
[0070]
[0071] As shown in Table 1, the table lists the key parameter values of the processed image in the execution of the boundary extraction and color recognition process, including the color clustering number, the boundary structure parameters and the color distribution structure, which provides a basis for filling the modal field in the subsequent graph node construction. The source of each parameter setting is combined with the image data feature distribution and experimental verification process to ensure the structural rigor and implementability of the scheme.
[0072] Referring to Figure 3 , the steps for obtaining the graph node coreference conflict labeling item are as follows:
[0073] S211: Based on the commodity image boundary feature set, according to the image structure parameters of each image, the attribute fields marked as geometric expression type in the node association field set in the commodity graph are called, the field values of the shape field, the external structure field and the material layout field are subjected to structural semantic analysis, the field values are converted into boundary target label structure, a one-to-one mapping relationship between the field and the boundary target is established, and an attribute boundary matching item set is generated;
[0074] To obtain the image structure parameters of each image in the product image boundary feature set, the first step is to read the joint index of the image field and attribute field within the graph node. After matching the image number with the node ID, the image to be processed is located. The set of structural fields bound to the image is extracted, and the shape field, external structure field, and material layout field in the field set are sequentially subjected to structural semantic parsing. During the parsing process, the text content of the field value needs to be extracted using standard vocabulary and converted into structural identifiers in the defined boundary target label set. For example, the field "border: right angle" is converted into "RectFrame", and the field "material coverage: local symmetry" is converted into "SymLayout". After extracting the structural labels, the label index number is found through the boundary mapping template. Then, the structural feature data extracted from the product image boundary feature set is called. Based on the field semantic target, image region blocks with similar boundary closure, symmetry, and local contour direction in the image structure are found. Each field label is bound to the image region coordinate segment, and an index list from the field target to the image region is established. The corresponding number index table of field-image region correspondence is recorded, and an attribute boundary matching itemset is generated.
[0075] S212: Based on attribute boundary matching itemsets, extract the corresponding image region coordinate range according to the boundary target label structure, determine whether there is an overlap in the coordinate coverage relationship between the regions pointed to by the field label structure, perform region orientation consistency detection on the field label structure in sequence, and perform region boundary repetition judgment and crossover measurement on the mapped fields, using the formula:
[0076] ;
[0077] The calculation obtains the field region overlap coefficient, compares it with the field repetition region threshold, performs conflict screening based on whether the field region points to the structure exclusively, and generates the field cross conflict measurement result.
[0078] in, This is the field region overlap coefficient, representing the degree of spatial overlap between the regions pointed to by all fields. The total number of target field boundaries indicates the number of field boundaries involved in the region comparison. For the first The normalized values of the boundary areas of each target field are normalized to the standard dimensions of the image region. For the first The normalized value of the boundary area of each target field, defined as follows: Used with the first The regions form an intersection. For the first The field region and the first The normalized value of the area of intersection of the regions in each field. For the first A normalized value of the total length of the field image region boundary line segment, normalized by the image width, is the first A normalized value of the region width estimation value of the image region, normalized by the image width;
[0079] According to the boundary target label structure in the attribute boundary matching item set, the image region coordinate range is extracted, and it is judged whether there is an intersection overlap between the coordinate coverage relationship between the regions pointed to by each field label structure. The region pointing consistency detection is performed on each field label structure in turn, the field region coordinate set is called, and its boundary attribute information is recognized. The intersection area calculation and boundary shape reconstruction operation are performed on the intersection structure between each field, and the formula is used:
[0080] ;
[0081] The operation logic of the formula is as follows: the formula numerator part represents the sum of the intersection areas between the regions of each field, reflecting the overlap degree between the fields; the first and second items of the denominator are the sum of the total areas of the regions participating in the field, and the third item is the spatial scale penalty item of the boundary size item. The field structure with large area is punished by the square root of the sum of the boundary line segment length and the width, and the ratio of the area proportion and the structure redundancy is formed as a whole. The higher the ratio is, the more intensive the regions pointed to by the fields are overlapped, and there is a risk of region repetition conflict. The field region coincidence degree coefficient is obtained by operation, which is compared with the preset field repeated region threshold value. The threshold value is set to 0.35, which is set according to the average overlap rate of the image region and the average interval between nodes. If the field coincidence degree is higher than the threshold value, it means that the region pointed to by the field cannot be exclusive, and there is a risk of conflict mapping, which needs to be included in the conflict judgment range.
[0082] In order to more intuitively explain the calculation process, an example is introduced: there are three field region pairs, whose normalized areas are , , , the intersection area is , , , the length of the boundary line segment is , , , the width is , , , the formula numerator is , the first item of the denominator is , the second item is the same value , and the third item is:
[0083] ;
[0084] In summary:
[0085] ;
[0086] The result is less than the threshold value of 0.35, indicating that there is no conflict mapping between the node fields, and the calculation shows that the distribution of the pointing area between the fields is reasonable, no serious overlap occurs, the area intersection does not reach the judgment threshold, and the current field combination can retain the original structure without modification. The result is used to establish the field intersection conflict measurement result.
[0087] Table 2 Field area overlap calculation example table
[0088]
[0089] As shown in Table 2, the overlap coefficient between fields A1 and A2 is within the judgment range, and no conflict needs to be marked.
[0090] wherein, represents the field area overlap coefficient, dimensionless, reflecting the spatial overlap relationship of the fields; 、 is the normalized area of the boundary area of the first 、 field, the original area unit is the number of pixels, and the image standard size is normalized; represents the normalized value of the intersection area of the first 、 field area; and respectively represent the normalized values of the boundary line segment length and width of the first field area, the unit is pixel length and pixel width, and the image width is normalized; represents the total number of field boundary targets.
[0091] The advantage of the formula is that by introducing the size penalty term of the boundary structure, combining the area intersection area and the global structure area to form a ratio, the risk of misidentification of image boundary label repetition is effectively controlled, and the uniqueness of the structure field pointing is consistent with the atlas node.
[0092] S213: Based on the field intersection conflict measurement result, according to the field area repetition, no matching item and the corresponding image coordinate point information, the item in the field item that is not corresponding to the image boundary or overlapping with other field area is screened, the field item and the corresponding image area coordinate index are extracted, the node internal structure field pointing abnormal mapping table is established, and the atlas node co-reference conflict marked item is generated;
[0093] Based on the number of field area overlaps in the field intersection conflict measure result, all field items are subjected to conflict determination. If the mapping area of a field overlaps with two or more other field areas and the intersection total area accounts for more than a set threshold of 0.2, it is determined to be a multi-field overlap anomaly. At the same time, it is detected whether there is a no-match area, i.e. it fails to locate any closed contour area in the image structure, and its corresponding intersection area is 0 and there is no coordinate pointing mark, and it is marked as a no-match item. For all marked fields, the field number and the associated image area coordinate point index are extracted, and the abnormal field and its image matching failure identification are formed into a field anomaly mapping matrix. The output node structure field exists in the image structure expression. The semantic overlap of the field set is generated, and the atlas node coreference conflict annotation item is generated.
[0094] Please refer to Figure 4 The acquisition step of the refined class object tag candidate set is specifically:
[0095] S311: Based on the atlas node coreference conflict annotation item, the image area coordinates corresponding to the annotated fields are used as the limited area. The contour edge of the goods image in the limited area is recognized. The total number of contour edge pixels is counted and divided by the area to obtain the edge point density value. At the same time, the number of edge line segments and the number of edge length change segments are recorded. The image area edge density information is established.
[0096] Based on the image area coordinates corresponding to the fields annotated in the atlas node coreference conflict annotation item, the corresponding image area is intercepted as a limited area, which is used as the basis for all subsequent edge processing. The image pixel matrix is called and grayscale processing is applied to convert the original image into a grayscale image. Then the Sobel operator is executed to calculate the gradient of each pixel point in the region. The pixels with a gradient amplitude greater than 50 are selected as edge points. The total number of edge pixels is ratioed with the area of the limited area to obtain the edge point density value. The density value is calculated as follows: 1786 edge points are extracted from a selected 256x256 pixel area, and the density value is 0.0273. Then the image edge is scanned, and the number of continuous edge line segments in the region is calculated to obtain 21 segments. At the same time, the number of segments with a pixel length change rate greater than 15% in the continuous line segment is counted to be 5 segments. Based on the above indicators, the number of edge line segments, the number of edge length change segments, and three values are recorded to establish the image area edge density information. The sample area data is shown in Table 3.
[0097] Table 3 Image area edge density information table
[0098]
[0099] As shown in Table 3, the edge point density information of image number IMG_01 has been constructed, forming a basic structure indicator set for subsequent dense feature vector calculation.
[0100] S312: Call the image region edge density information, according to the corresponding image region, the color channel data of the pixel in the region is extracted and the difference is distinguished, the color aggregation area is identified and the color principal component clustering value is counted, the double parameter density feature vector is constructed combined with the edge density value and the clustering value, using the formula:
[0101] ;
[0102] The operation obtains the region density deviation measurement value, and the deviation value is compared with the standard density vector corresponding to the class target label bound to the current commodity node. Whether it exceeds the density recognition threshold is judged, and the density deviation information is obtained;
[0103] Wherein, The density deviation measurement value between the image region and the category standard, The edge density value of the image region, The standard edge density reference value corresponding to the category label, The standard deviation of the edge density sample of the image region, The standard deviation of the edge density sample of the category standard, The color clustering number of the image region, The color clustering number of the category standard;
[0104] Call the image region corresponding to the image region edge density information, respectively execute the main color proportion statistics for RGB channel, set the maximum three colors of pixel frequency as the main color, and use K-means method to initially set Pixel clustering, after clustering convergence, the number of clustering clusters is counted, and the color aggregation number of the current image region is obtained 6, the standard category color clustering number is compared 4; the edge density value , the standard edge density reference value , the sample standard deviation , the standard sample standard deviation Together into the formula:
[0105] ;
[0106] In the formula, the first term The absolute difference between the edge density value of the current image region and the category standard value is the numerator, and the square root of the sum of the square of the standard deviation of the two is the denominator, which measures the deviation degree of the structural edge complexity; the second term The difference degree of color cluster number, the denominator plus 1 processing avoids the problem of division by 0, and reduces the influence on the overall measurement when the number of clusters is small; the overall formula is the joint measurement of the structure density and the color distribution difference, and the obtained The larger the value, the stronger the deviation of the image region from the standard category image in structure and color distribution.
[0107] Substitute the sample region parameters into the calculation as follows:
[0108] Difference value molecule part:
[0109] Standard deviation square root:
[0110] The first result:
[0111] The second result:
[0112] Synthesis
[0113] The meaning of the formula calculation result is: the current image region has a combined deviation degree of 1.115 in edge structure density and color cluster distribution compared with the standard image region under the category label to which it belongs. This index is an evaluation parameter of the comprehensive image structure complexity and color dispersion.
[0114] The "density identification threshold" used to determine whether the region has abnormal density in this step is set to 1.0, and the setting basis is: the edge density value and the cluster number of all product node images in the standard image set are calculated under the category to which they belong Distribution is used as a reference, and the upper limit of the distribution is extracted within a 95% confidence interval, which is determined by the mean + 1.5 times the standard deviation, forming a boundary reference for judging whether it deviates from the category standard distribution. The specific setting process is: the average of the sample in the standard category image set is 0.68, and the standard deviation is 0.21, so the threshold is set to , which is rounded up to 1.0 to simplify the application process.
[0115] Since the result is greater than the threshold 1.0, it is marked as a dense abnormal item, and forms a dense deviation information.
[0116] S313: Based on the density deviation information, the category hierarchy corresponding to the original category label of the commodity node is extracted according to the label item marked as exceeding the recognition threshold, the same layer sub-label set in which the current label is located is obtained from the structure, and each item is compared according to the deviation value and the preset density standard vector of the sub-label, the label item with the smallest difference value is selected as the main label identification, the mapping relationship between the sub-label set and the main label is established, and the refined category label candidate set is generated;
[0117] According to the label item marked as exceeding the recognition threshold in the density deviation information, the current binding category label item of the commodity node is "outdoor shoes", which is located in the "shoes-movement shoes / outdoor shoes / leisure shoes" level in the category tree structure, the same layer label set is obtained as "movement shoes", "leisure shoes", "outdoor shoes", the standard density vector of each sub-label is called respectively as: movement shoes: [0.0329, 5]; Leisure shoes: [0.0295, 4]; Outdoor shoes: [0.0341, 4]; The Euclidean distance is calculated with the current image area vector [0.0273, 6] respectively, and the results are:
[0118] The distance from the movement shoes: ;
[0119] The distance from the leisure shoes: ;
[0120] The distance from the outdoor shoes: ;
[0121] The above results show that "movement shoes" is closest to the current image feature, so "movement shoes" is selected as the main label, and it is added to the main label identification item, and a mapping relationship is established with the sub-labels "outdoor shoes" and "leisure shoes", and a refined category label candidate set is generated.
[0122] Please refer to Figure 5 , the acquisition steps of the atlas construction field reinforcement node list are as follows:
[0123] S411: Call the refined category label candidate set, according to the commodity node set bound by each sub-label, count the filling state of the image field, the text field and the structure field of each node, calculate the ratio of the number of nodes filled effectively in each modal field to the total number of set nodes, and generate the modal field existence rate value set;
[0124] Call the set of product nodes bound by each sub-label in the target label candidate set of the refinement class, take each node as a traversal object, access its image field, text field and structure field one by one, and judge the filling state of each field, wherein the judgment basis of the image field filling is whether the field stores an image file index number, the text field is based on whether the field content length exceeds a set threshold, for example, more than 5 characters is considered to be valid filling, and the structure field is based on whether it contains structured position information data, for example, a field containing coordinate points or size units is defined as filled. For a set containing a total of 150 product nodes, whether each modal field is filled is recorded in Boolean form as 1 or 0, the number of nodes filled in the image field is 132, the number of nodes filled in the text field is 145, and the number of nodes filled in the structure field is 123. After dividing by the total number of nodes 150, the existence rate values of each field are 0.88, 0.97 and 0.82, respectively, to form the modal field existence rate value set.
[0125] S412: Based on the modal field existence rate value set, the existence distribution of the field among multiple nodes in the set is calculated according to the existence rate of each modal field, the sample standard deviation value of each modal field is obtained, and the standard deviation coefficient set is constructed by the field existence rate and the standard deviation. The formula is:
[0126] ;
[0127] The operation obtains the field existence deviation factor, judges whether the existence deviation factor exceeds the field consistency deviation threshold, selects the nodes with an existence rate lower than the average value and a deviation factor falling into the positive deviation interval, and obtains the field consistency offset indicator information.
[0128] wherein, represents the field existence deviation factor, and represents the outlying degree of the current node in the modal field filling performance, is the normalized value of the modal field existence rate of the current product node, and the calculation method is the normalized representation of the Boolean value of whether the node is filled in the set, is the average existence rate normalized value of the modal field in the corresponding node set, indicating the average level of the frequency of filling the field, is the normalized value of the standard deviation of the existence rate of the modal field in the node set, used to measure the discreteness of the filling state, is a non-zero constant, used to balance the denominator value to prevent the standard deviation from being zero, is an adjustment coefficient, controlling the influence strength of the offset term on the whole indicator, is the field filling state normalized value of the node, taking the value of 0 or 1, is the normalized value of the average value of the filling state of the corresponding modal field in the node set, To refine the total number of nodes under the current class target sign;
[0129] Based on the foregoing modal field presence rate value set, the presence rate of each modal field is extracted as a Boolean value vector formed between different nodes, and the sample standard deviation is calculated for each field. For example, the Boolean filling state vector of the image field is a 0 / 1 sequence with a length of 150, of which the number of 1 is 132. According to the standard deviation formula, the standard deviation is 0.33. The standard deviations of the text field and the structure field are 0.17 and 0.39 respectively. Then, the presence rate value is paired with the corresponding standard deviation value to build a standard deviation coefficient set. On this basis, the following formula is used to calculate the field existence deviation factor:
[0130] ;
[0131] In the formula, the first term represents the degree of deviation of the current node field filling rate from the average filling rate of the set, and the standard deviation is a quantitative control of the fluctuation range to avoid extreme value out of control. The second term measures the overall stability fluctuation of the current field in the node set, which plays a role in adjusting the weight of overall fluctuation. The two are added to form a composite evaluation function, and the absolute value is used to ensure the directionality and uniformity. The greater the field existence deviation factor, the more the filling of the node in the field deviates from the general state of the set.
[0132] The field consistency deviation threshold is set with reference to the standard distribution range of the existence rate deviation factor of the three types of fields in the whole node set. According to statistics, 95% of the node deviation factors are concentrated in the interval , so the threshold is set to 0.45 as the upper limit criterion for identifying outlier nodes, and the value has a monotonic convergence characteristic with the number of nodes , the fluctuation range is controlled within ±0.02. Taking node C113 as an example, the image field existence rate is 0, the set average is 0.88, and the standard deviation is 0.33. Set , , the field Boolean deviation average difference is 0.42, which is calculated as follows:
[0133] ;
[0134] The value is much higher than the threshold value 0.45, so it is marked as a high deviation node. The result generated in this section is the field consistency deviation index information, which is used for subsequent screening of low filling nodes.
[0135] Table 4 Field existence rate and deviation factor sample table
[0136]
[0137] As shown in Table 4, the node C113 bias factor is much higher than the threshold standard, which is an outlier node.
[0138] S413: Based on the field consistency offset indicator information, according to the node item marked as low filling and high offset, the commodity node number and modal field type information are extracted, the field missing node index table is constructed, and the graph construction field reinforcement node list is generated.
[0139] According to the node item marked as the existence rate below the average value and the bias factor higher than the set threshold in the field consistency offset indicator information, the commodity node number and modal field type information are extracted from each marked node, and the field missing record corresponding to each type of field is constructed, and the missing index table is formed by the node number and the field name, such as {C113, image field}, {C120, text field}, {C143, structure field}, and finally the above missing index is aggregated to form the field missing node index table, and the graph construction field reinforcement node list is generated accordingly, which is used as the input data source of the subsequent reinforcement scheduling instruction.
[0140] Please refer to Figure 6 The acquisition step of the commodity graph construction scheduling instruction set is specifically:
[0141] S511: Based on the graph construction field reinforcement node list, according to the number of each commodity node, the filled image field, text field and structure field content of the node are extracted, and the field name is mapped and numbered according to the standard field sequence order defined in the graph template, and each field content is corresponded to the preset position number in the category structure, and the modal field mapping sequence information is generated;
[0142] Based on the number of each commodity node in the graph construction field reinforcement node list, the original content of the filled fields such as image field, text field and structure field corresponding to the commodity node is read one by one, the field name is standardized and the empty field item is removed, and according to the category label information currently bound to the commodity node, the standard field sequence structure defined in the graph template for the category is called, for example, for the "sports shoes" category, the standard field order is image field, material composition field, origin field, function description field and size specification field in turn, for each filled field, the matching is performed, the corresponding position number in the standard structure is located, the image field is 1, the function description field is 4, and the size specification field is 5, so the modal field mapping sequence number of the current node is In actual operation, a key-value mapping is formed for each field content and standard number position, and a triple structure is constructed: field name, field content, and position number, all triples are arranged in ascending order of number, and finally output as modal field mapping sequence information, which is used as the input of subsequent graph structure calibration.
[0143] S512: Call the modal field mapping sequence information, according to the standard position number of each field, extract the corresponding field organization order structure and node link structure constraint rules from the atlas configuration template according to the class target label bound by the current node, judge the connectivity between fields, and filter the connectable field pairs based on the node allowed connection type matrix, establish the sequential link structure between fields, and generate field organization connection graph structure metric information;
[0144] Call the standard position number of each field in the modal field mapping sequence information, read the corresponding field relationship combination in sequence, query the field organization order structure and node link structure constraint rules defined in the atlas configuration template, such as "image field" and "function description field" allowed to establish graph-text link, the number order needs to meet the condition that the former is less than the latter, if it is met, it is added to the effective link candidate set; Further, according to the field node allowed connection type matrix, the connectability is judged, for example, the image field number is 1, the function description field number is 4, and the connection matrix value is 1, indicating that it can be directly linked, and the connection pair is constructed ; In this way, all connectable field pairs are filtered out, and each connectable field pair is indexed by the standard node sequence number, and recorded as a connection item. The example field organization connection graph structure is: , and record the link direction, node order difference, and connection level difference between each pair, generate field organization connection graph structure metric information, which includes the connectable pairs between fields, link level difference value and structure direction.
[0145] Table 5 Field Connection Graph Structure Example Table
[0146]
[0147] As shown in Table 5, field 1→4 and 4→5 are legal link paths, forming a continuous connection link. Field 1→5 does not have a direct link relationship and is illegal in the current configuration template.
[0148] S513: Based on the field organization connection graph structure metric information, according to the field pair link rule and field node sequence number, calculate the field boundary connection direction and mark the atlas topology position index, sort the node field connection edge according to the link structure, construct the node boundary topology table, and for each node Establish a structure generation command group, generate a product atlas construction scheduling instruction set;
[0149] According to the field organization connection graph structure measurement information, the link rule and the field node sequence number are linked, the boundary link direction parameter in each legal connection pair is calculated, the direction value is defined as the front node number is less than the rear node number, which is recorded as forward, and vice versa, and the topological arrangement position of each connection edge is determined according to this, and the topological position is calibrated under the standard atlas structure, such as the boundary direction of the image field connection text field is recorded as (1, 4), and the topological position is located in the front half area, after constructing the field connection table, all connection edges are sorted in ascending order of field sequence, forming a topological connection edge table, and each edge is provided with corresponding field serial number, start and end field number, boundary direction and other index items, each node constructs an independent topological boundary table, and all field edge pairs are arranged in ascending order of field number, and combined with the structure generation rule preset in the atlas configuration template, a structure construction command item is generated for each node, for example, the instruction format is: link field 1 and field 4, the node position is in the topological index 2-3 area, and the record is command , all node commands are generated to form a complete structure, and the final output is a commodity atlas construction scheduling instruction set, which drives the subsequent atlas entity construction engine to perform node merging, edge connection and topological mapping operation to form a structural atlas pattern.
[0150] The multi-modal data driven commodity atlas construction system is used to realize the multi-modal data driven commodity atlas construction method, and the system comprises:
[0151] The commodity feature recognition module acquires the nodes in the commodity atlas which have not completed the modal structure construction, collects the commodity images bound to the nodes, performs foreground segmentation on the images, aggregates the image region features to form a multi-dimensional feature vector, and generates a commodity image boundary feature set;
[0152] The conflict labeling module calls the attribute items with geometric meaning in the structured field set associated with the nodes in the commodity atlas based on the commodity image boundary feature set, converts the field values into descriptive boundary targets and maps them to the image foreground region, judges whether there is a matching region corresponding to the image region of the target field value, and generates atlas node coreference conflict labeling items;
[0153] The target refinement analysis module performs edge density estimation and color cluster number statistics on the target region based on the atlas node coreference conflict labeling items, judges whether the class target label bound to the current commodity node exists deviation between the vectors, and generates a refined class target label candidate item set;
[0154] The reinforcement node recognition module calls the refined class target label candidate item set, calculates the field existence rate for each modal field and acquires the field standard deviation coefficient, judges whether the current commodity node is lower than the mean value of the existence rate and falls into the standard deviation positive deviation interval under the field, and generates a graph construction field reinforcement node list;
[0155] The scheduling instruction construction module constructs a node list based on the graph construction field, extracts a field organization order and a node link structure constraint template from the graph configuration template according to the class target label item, combines the mapping index of the existing field value in the standard structure, constructs a boundary topology connection table of the commodity entity node, and generates a commodity graph construction scheduling instruction set.
[0156] An electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the multi-modal data driven commodity graph construction system as described above when executing the computer program.
[0157] The above is only a preferred embodiment of the present application, and does not limit the present application in other forms. Any person skilled in the art can use the disclosed technical content to make changes or modifications as equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application without departing from the technical solution content of the present application still belongs to the protection scope of the technical solution of the present application.
Claims
1. A multi-modal data-driven commodity graph construction method, characterized in that, The method comprises the following steps: S1: obtaining a node with an incomplete modal structure in a commodity atlas, collecting a commodity image bound to the node, performing foreground segmentation on the image, aggregating image region features to form a multi-dimensional feature vector, and generating a commodity image boundary feature set; S2: based on the commodity image boundary feature set, calling attribute items with geometric meaning in a structured field set associated with the node in the commodity atlas, converting field values into descriptive boundary targets and mapping them to image foreground regions, judging whether there is a matching region in the image region corresponding to the target field value, and generating atlas node coreference conflict annotation items; S3: based on the atlas node coreference conflict annotation items, performing edge density estimation and color cluster number statistics on the target region, judging whether there is a density deviation between the current commodity node bound class target label and the vector, and generating a refined class target label candidate set; S4: calling the refined class target label candidate set, calculating the field existence rate for each modal field and obtaining the field standard deviation coefficient, judging whether the current commodity node is lower than the average of the existence rate and falls into the standard deviation positive deviation interval, and generating a graph construction field reinforcement node list.
2. The multi-modal data-driven commodity graph construction method of claim 1, wherein, The commodity image boundary feature set includes edge detection point set, boundary continuous region numerical value, closed region quantity index, color block quantity value and boundary line segment length average value, the atlas node coreference conflict annotation item is specifically a field item identification list, image region coordinate index, field mapping failure marker set and multi-field overlap conflict identification, the refined class target label candidate set includes sub-label set, dense deviation threshold, main label identification item and candidate label sorting table, and the graph construction field reinforcement node list is specifically a reinforcement node number, field missing modal identification, existence rate offset table and standard deviation overrun record.
3. The multi-modal data-driven commodity graph construction method of claim 2, wherein, The acquisition step of the commodity image boundary feature set is specifically: S111: obtaining a node with an incomplete modal structure in a commodity atlas, collecting a commodity image corresponding to each node, and performing foreground segmentation on each image, locating the main commodity contour region in the image after segmentation and removing background interference information, judging the main region boundary according to the gray scale gradient difference range and region connectivity inside and outside the contour closed region, and generating a foreground contour region coordinate value set; S112: based on the foreground contour region coordinate value set, calling Canny edge detection operator to perform contour point extraction on the contour range according to the calibrated region boundary, calculating the contour point density value and the number of boundary continuous regions for each contour region, and identifying the number of complete closed regions based on the arrangement mode of the boundary closed points, and establishing a boundary structure statistical index set corresponding to each image sample; S113: calling the boundary structure statistical index set, performing block recognition based on region color aggregation features within the contour coordinate range, performing main color proportion analysis and color category distinguishability operation on the recognized region, counting the number of color aggregation blocks, calculating the average value of contour line segment length, and aggregating image structure data to form a multi-dimensional feature vector, and generating a commodity image boundary feature set.
4. The multi-modal data-driven commodity graph construction method of claim 3, wherein, The acquisition step of the atlas node coreference conflict annotation item is specifically: S211: Based on the set of commodity image boundary features, according to the image structure parameters of each image, calling the attribute fields marked as geometric expression type in the node association field set in the commodity graph atlas, performing structural semantic analysis on the field values of the shape field, external structure field and material layout field, converting the field values into boundary target label structure, establishing a one-to-one mapping relationship between the field and the boundary target, and generating a set of attribute boundary matching items; S212: Based on the set of attribute boundary matching items, according to the boundary target label structure, extracting the corresponding image region coordinate range, judging whether there is an intersection overlap between the coordinate coverage relationship between the region pointed to by the field label structure, sequentially performing region pointing consistency detection on the field label structure, performing region boundary repetition judgment and intersection degree measurement on the mapped field, calculating the field region coincidence coefficient, comparing the field repeated region threshold value, performing conflict screening according to whether the field region pointing structure is exclusive, and generating a field intersection conflict measurement result; S213: Based on the field intersection conflict measurement result, according to the field region repeated items, no matching items and corresponding image coordinate point information, screening the items in the field item that do not correspond to the image boundary or overlap with other field regions, extracting the field item and the corresponding image region coordinate index, establishing an intra-node structure field pointing abnormal mapping table, and generating a graph node co-reference conflict annotation item.
5. The multi-modal data-driven commodity graph construction method of claim 4, wherein, The obtaining step of the set of refined class label candidate items is specifically: S311: Based on the graph node co-reference conflict annotation item, according to the annotated field corresponding image region coordinates, taking the region boundary coordinates in the image as a limited region, performing contour edge recognition on the commodity image in the limited region, calculating the total number of contour edge pixels and dividing the area to obtain the edge point density value, recording the number of edge line segments and the number of edge length change sections, and establishing image region edge density information; S312: Calling the image region edge density information, according to the corresponding image region, performing main color extraction and difference processing on the pixel color channel data in the region, identifying color aggregation regions and counting color principal component clustering values, constructing a two-parameter density feature vector combining the edge density value and the clustering value, calculating the region density deviation measurement value, comparing the deviation value with the standard density vector corresponding to the class label bound to the current commodity node, judging whether it exceeds the density recognition threshold, and obtaining the density deviation information; S313: Based on the density deviation information, according to the label item marked as exceeding the recognition threshold, extracting the category hierarchy structure corresponding to the original class label of the commodity node, obtaining the same layer sub-label set where the current label is located from the structure, and comparing the deviation value with the preset density standard vector of the sub-label item by item, selecting the label item with the smallest difference value as the main label identifier, establishing the mapping relationship between the sub-label set and the main label, and generating a set of refined class label candidate items.
6. The multi-modal data-driven commodity graph construction method of claim 5, wherein, The obtaining step of the graph construction field reinforced node list is specifically: S411: Call the set of refined class target tag candidates, according to the set of product nodes bound by each sub-tag, count the filling state of the image field, the text field and the structure field of each node, calculate the ratio of the number of nodes filled effectively to the total number of nodes in the set, and generate a set of modal field existence rate values; S412: Based on the set of modal field existence rate values, the existence distribution of the field in multiple nodes in the set is calculated according to the existence rate of each modal field, the sample standard deviation value of each modal field is obtained, the standard deviation coefficient set is constructed by the field existence rate and the standard deviation, the field existence deviation factor is obtained, and it is judged whether the existence deviation factor exceeds the field consistency deviation threshold, the nodes with low filling and high deviation factor are screened, and the field consistency offset index information is obtained; S413: Based on the field consistency offset index information, according to the node item marked as low filling and high offset, the product node number and modal field type information are extracted, the field missing node index table is constructed, and the field supplement node list for graph construction is generated.
7. The multi-modal data-driven commodity graph construction method of claim 6, wherein, The method further comprises the following steps: S5: Based on the field supplement node list for graph construction, the field organization order and node link structure constraint template are extracted from the graph configuration template according to the class target tag item, the boundary topology connection table of the product entity node is constructed by combining the mapping index of the existing field value in the standard structure, and the product graph construction scheduling instruction set is generated; The product graph construction scheduling instruction set includes field organization sequence, graph boundary connection table, link structure type mapping and target graph writing instruction.
8. The multi-modal data-driven commodity graph construction method of claim 7, wherein, The acquisition step of the product graph construction scheduling instruction set is specifically: S511: Based on the field supplement node list for graph construction, the image field, the text field and the structure field content filled by each product node are extracted according to the number of each product node, and the field name is mapped and numbered according to the standard field sequence order defined in the graph template, so that each field content corresponds to the preset position number in the category structure, and the modal field mapping sequence information is generated; S512: Call the modal field mapping sequence information, according to the standard position number of each field, according to the class target tag bound by the current node, extract the corresponding field organization order structure and node link structure constraint rule from the graph configuration template, judge the connectivity between fields, and filter the connectable field pairs based on the node allowed connection type matrix between fields, establish the sequence link structure between fields, and generate the field organization connection graph structure metric information; S513: Based on the field organization connection graph structure metric information, according to the field pair link rule and the field node sequence number, the field boundary connection direction is calculated and the graph topology position index is calibrated, the node field connection edge is sorted according to the link structure, the node boundary topology table is constructed, and the structure generation command group is established for each node, and the product graph construction scheduling instruction set is generated.
9. A multi-modal data driven commodity graph construction system, characterized in that, The system is used to realize the multi-modal data-driven commodity graph construction method of any one of claims 1-8, and the system comprises: The commodity feature recognition module acquires a node in the commodity graph for which a modal structure construction is not completed, collects a commodity image bound to the node, performs foreground segmentation on the image, aggregates image region features to form a multi-dimensional feature vector, and generates a commodity image boundary feature set; The conflict labeling module, based on the commodity image boundary feature set, calls attribute items with geometric meanings in a structured field set associated with a node in the commodity graph, converts field values into descriptive boundary targets and maps them to image foreground regions, judges whether there is a matching region in the image region corresponding to the target field value, and generates a graph node coreference conflict labeling item; The target refinement analysis module, based on the graph node coreference conflict labeling item, performs edge density estimation and color cluster number statistics on the target region, judges whether there is a density deviation between the current commodity node bound class target label and the vector, and generates a set of refined class target label candidates; The reinforcement node identification module calls the set of refined class target label candidates, calculates a field existence rate and acquires a field standard deviation coefficient for each modal field, judges whether the current commodity node is below the mean value of the existence rate and falls into the standard deviation positive deviation interval under the field, and generates a graph construction field reinforcement node list; The scheduling instruction construction module, based on the graph construction field reinforcement node list, extracts a field organization order and a node link structure constraint template from a graph configuration template according to the class target label item, combines the mapping index of the existing field value in the standard structure, constructs a boundary topology connection table of the commodity entity node, and generates a commodity graph construction scheduling instruction set.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor executes the computer program to realize the multi-modal data-driven commodity graph construction system of claim 9.
Citation Information
Patent Citations
Multi-modal commodity knowledge graph construction method
CN112528042A
Traditional Chinese medicine massage multi-modal knowledge graph construction method based on large model and comparative learning
CN118136264A