An e-commerce main picture-oriented text-driven multi-commodity intelligent synthesis method and system

By constructing a sensitive adjacency matrix and an isolated potential field, the problem of visual block instability caused by SKU inventory changes in the generation of multi-product main images is solved, achieving high efficiency in local editing and controllability in product display.

CN122453979APending Publication Date: 2026-07-24XIAMEN FINGERPRINT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN FINGERPRINT TECH CO LTD
Filing Date
2026-06-26
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing multi-product main image generation solutions lack effective control over the identity preservation, local editing constraints, and structural stability during subsequent addition, deletion, and modification processes of multiple product entities within the same main image. In particular, it is difficult to accurately distinguish between visual blocks and non-target blocks when SKUs change due to inventory.

Method used

By constructing a sensitive adjacency matrix and an editable mask set, combined with visual anchor points and risk propagation analysis, an isolation potential field is built to enable local editing operations, ensuring the stability and independence of the product layout.

Benefits of technology

It enables only partial adjustments to visual blocks when SKU inventory changes, avoiding full image redrawing, thus improving the responsiveness of the e-commerce main image synthesis process and the consistency of product display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122453979A_ABST
    Figure CN122453979A_ABST
Patent Text Reader

Abstract

The application discloses an e-commerce main picture-oriented text-driven multi-commodity intelligent synthesis method and system, relates to the technical field of image processing, and constructs a sensitive adjacent matrix and an editable mask by analyzing SKU inventory state and combined co-occurrence relationship, obtains actual independence by combining visual anchor point space-time evolution, and constructs a collision probability field and a risk propagation diagram, forms an isolated potential field with direction selection characteristics based on a risk role result, and accordingly executes a local editing operation. The application utilizes binding iterative updating to maintain the continuous stability of SKU identity and visual anchor point binding, so that local editing only needs to adjust the area covered by the to-be-edited anchor point and its buffer isolation belt, avoids whole picture redrawing, ensures that an explicit pixel-level segmentation interface is reserved between visual anchor points, suppresses search weight attenuation caused by edge fusion, and meets the review requirements of a platform on commodity main body independence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a text-driven intelligent synthesis method and system for e-commerce main images. Background Technology

[0002] With the development of e-commerce, intelligent marketing design, and generative artificial intelligence technology, digital content automatic generation technology is gradually being applied to the field of product display. Especially in the scenario of intelligent generation of e-commerce main images, for the display needs of product combinations with multiple SKUs under the same SPU, it is necessary to coordinate the processing of product text information, inventory information, product image materials, and operational strategy information, and complete the layout of multiple product subjects within a unified screen.

[0003] Existing multi-product main image generation solutions typically employ fixed template layouts, rule-driven layouts, or diffusion models to directly generate product combinations. While these systems can generate corresponding visual content based on product text descriptions, they lack effective control over maintaining the identity of multiple product entities within the same main image, local editing constraints, and structural stability during subsequent additions, deletions, and modifications. For example, when a SKU needs to be replaced, deleted, or added due to inventory changes, existing solutions often re-execute the entire generation process or only make local adjustments based on spatial distance, making it difficult to accurately distinguish the visual blocks corresponding to the target SKU from non-target visual blocks that should remain unchanged. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a text-driven intelligent multi-product synthesis method and system for e-commerce main images, solving the problems mentioned in the background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] Firstly, a text-driven intelligent multi-product composition method for e-commerce main images includes:

[0007] Receive raw data of multiple SKUs under SPU, generate initial feature vectors corresponding to SKU nodes, and construct a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationship of each SKU.

[0008] Load product image materials, obtain visual anchor points, and combine them with the initial feature vectors corresponding to SKU nodes to construct a binding matrix and an editable mask in each layout iteration;

[0009] By analyzing the spatiotemporal evolution of each visual anchor point, the actual independence is obtained. Combined with the sensitive adjacency matrix, a collision probability field and risk role results are constructed for risk propagation analysis. Based on independence and risk identification, the lock flag is updated and supplemented with editable state codes. The risk role results include potential adhesion centers and potential adhesion propagation chains.

[0010] Edge tensors are extracted based on visual anchor point boundaries, diffusion tensors are constructed by combining risk role results, and an isolation potential field with direction selection characteristics is constructed by analyzing the correspondence between edge structure direction and risk propagation path.

[0011] Based on the supplemented editable state code, the anchor point to be edited is determined, and local editing operations are performed in combination with the isolation potential field and risk role results to complete the multi-product layout adjustment of the e-commerce main image.

[0012] Secondly, a text-driven multi-product intelligent synthesis system for e-commerce main images includes:

[0013] The SKU feature modeling module is used to receive multiple sets of raw SKU data under SPU, generate initial feature vectors corresponding to SKU nodes, and construct a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationship of each SKU.

[0014] The anchor binding module is used to load product image materials, obtain visual anchor points, and combine them with the initial feature vectors corresponding to the SKU nodes to construct a binding matrix and an editable mask in each layout iteration.

[0015] The independent risk module is used to obtain the actual independence by analyzing the spatiotemporal evolution of each visual anchor point. Combined with the sensitive adjacency matrix, it constructs a collision probability field and risk role results for risk propagation analysis. Based on independence and risk identification, it locks the flags to update and supplement the editable state code. The risk role results include potential adhesion centers and potential adhesion propagation chains.

[0016] The isolation potential energy construction module is used to extract edge tensors based on visual anchor point boundaries, construct diffusion tensors by combining risk role results, and construct an isolation potential energy field with direction selection characteristics by analyzing the correspondence between edge structure direction and risk propagation path.

[0017] The local adjustment module is used to determine the anchor point to be edited based on the supplemented editable state code, and to perform local editing operations in combination with the isolation potential field and risk role results to complete the multi-product layout adjustment of the e-commerce main image.

[0018] The above-described solution of the present invention has at least the following beneficial effects:

[0019] By constructing a bidirectional binding matrix between SKU text entities and visual anchors, and combining lifecycle trajectories and binding iteration updates to achieve cross-iteration identity tracking, this addresses the architectural flaw of existing methods that require a complete image redraw due to SKU configuration changes. Building upon this, a collision probability field is constructed based on an inventory-sensitive adjacency matrix and the first-order difference of independence. Kernel density estimation is used to analyze the distribution density of boundary points, and the kernel density bandwidth is dynamically reduced for anchors triggering warnings, achieving a centralized representation of collision risk. Furthermore, by identifying adhesion centers and propagation chains through a risk propagation graph, the propagation chain positions are mapped to normal diffusion factors. The preset normal diffusion coefficient is scaled, ensuring that the isolation potential energy maintains high-speed diffusion along the edge tangential direction, while normal diffusion is adaptively suppressed by the risk level, forming a direction-selective isolation potential energy field. This mechanism allows a single local add / delete / modify operation to only adjust the limited area covered by the anchor to be edited and its buffer isolation zone, fundamentally avoiding redundant calculations for a full image redraw and improving the response efficiency of SKU inventory changes during e-commerce main image synthesis.

[0020] In the anisotropic diffusion equation guided by multi-scale edge tensors, the normal diffusion coefficient of the isolation potential field is linked to the position of the propagation chain in the risk propagation map. This results in smaller normal diffusion coefficients for potential adhesion centers and nodes at the beginning and middle of the chain, forming a narrow and steep isolation potential energy distribution. Conversely, larger normal diffusion coefficients are obtained for the end of the chain and risk-free nodes, forming a wide and gentle isolation potential energy distribution. This differentiated isolation strategy ensures that the product edges remain continuous in the tangential direction and are dynamically adjusted by the risk level in the normal direction, preventing edge fusion from making it difficult for the platform's image retrieval model to separate independent product features. Simultaneously, the buffer isolation zone acts as an elastic constraint boundary to absorb displacement impacts, maintaining the minimum interval between the anchor point to be edited and its neighboring anchor points during local editing operations. This ensures that the generated main image meets the platform's implicit review requirements for the independence of the product subject at each editing stage, thereby suppressing the search weight decay problem caused by pixel adhesion. Attached Figure Description

[0021] Figure 1 This is a flowchart of the method of the present invention;

[0022] Figure 2 This is a system structure diagram of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] like Figure 1As shown, embodiments of the present invention provide a text-driven intelligent multi-product synthesis method for e-commerce main images, including:

[0025] S100: Receives multiple sets of raw SKU data under SPU, generates initial feature vectors corresponding to SKU nodes, and constructs a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationship of each SKU.

[0026] S200: Load product image materials, obtain visual anchor points, and combine them with the initial feature vectors corresponding to the SKU nodes to construct a binding matrix and an editable mask in each layout iteration;

[0027] S300: By analyzing the spatiotemporal evolution of each visual anchor point, the actual independence is obtained. Combined with the sensitive adjacency matrix, a collision probability field and risk role results are constructed for risk propagation analysis. Based on independence and risk identification, the lock flag is updated and supplemented with editable state codes. The risk role results include potential adhesion centers and potential adhesion propagation chains.

[0028] S400: Extract edge tensors based on visual anchor point boundaries, construct diffusion tensors by combining risk role results, and construct an isolation potential field with direction selection characteristics by analyzing the correspondence between edge structure direction and risk propagation path.

[0029] S500: Based on the supplemented editable state code, determine the anchor point to be edited, and perform local editing operations in combination with the isolation potential field and risk role results to complete the multi-product layout adjustment of the e-commerce main image.

[0030] In this embodiment of the invention, in S100, the SKU is processed by feature vectorization and a sensitive adjacency matrix is ​​established, so that the system can understand the combination constraints and inventory associations between SKUs, providing a data foundation for subsequent binding and layout.

[0031] S200 uses visual anchor point binding to map text features to image features, achieving precise matching between SKU nodes and image blocks, and providing anchor point references for layout iteration;

[0032] S300 constructs a collision probability field and risk role results based on the spatiotemporal evolution of visual anchor points and independence analysis, enabling potential adhesion centers and propagation chains to be identified in advance, thereby updating the editable state code to lock key areas and prevent misoperation.

[0033] S400 extracts the edge tensor and combines it with the risk role results to construct a diffusion tensor, forming a direction-selective isolation potential field to ensure that high-risk areas do not interfere with surrounding blocks during local editing.

[0034] The S500 utilizes updated editable state codes and isolated potential fields to perform local addition, deletion, and modification operations, enabling precise layout adjustments of multiple products in the main image.

[0035] For example, when a certain SKU needs to be replaced with another style due to insufficient inventory, the system can guide adjustments only to the visual block containing the SKU and adjacent affected areas through risk roles and isolation constraints, without regenerating the entire main image. This avoids misalignment or overlap of non-target SKU blocks, thus maintaining consistency and controllability in the display of multiple products in a dynamic inventory environment. The entire process, through closed-loop processing of feature mapping, binding, risk analysis, and potential constraints, enables highly automated and intelligent visual layout and partial editing control of e-commerce main images in scenarios with multiple SKUs and changing inventory.

[0036] In a preferred embodiment of the present invention, S100: Receive multiple sets of original SKU data under the SPU pushed by the ERP system, perform multimodal feature representation preprocessing on the original SKU data, and generate an initial feature vector corresponding to each SKU through semantic feature extraction and specification feature mapping; this vector is a high-dimensional vector obtained by semantic embedding of text attributes and numerical mapping of specification parameters, reflecting the comprehensive semantic features, numerical specification features and potential combination constraint information of the SKU, and is the attribute representation of the SKU node in the subordinate hypergraph;

[0037] Using SKU entities carrying initial feature vectors as SKU nodes and the hierarchical relationships connecting multiple SKU nodes within the same set as hyperedges, a hierarchical hypergraph is constructed; this graph provides a structured foundation for subsequent inventory-sensitive correlation analysis, editable mask generation, and visual anchor binding;

[0038] SKU raw data includes text attribute information, specification parameter information, and package combination relationships;

[0039] Multimodal feature representation preprocessing is performed on the raw SKU data. First, the text attribute information is standardized, including product title, attribute key-value pairs, and specification description. Text standardization involves unifying units and symbols, standardizing capitalization, and cleaning up redundant spaces and special characters to make the text meet the input requirements of the semantic model. Then, the standardized text is input into the BERT model for semantic feature extraction, resulting in a 768-dimensional semantic embedding vector for each SKU. This vector represents the contextual semantic information of the SKU. The 768-dimensional semantic embedding is then mapped to an 896-dimensional feature space through a learnable linear projection layer to meet the requirements of subsequent feature fusion and network input dimension.

[0040] Meanwhile, the numerical specifications of the SKU, such as size, weight, and number of accessories, are normalized to avoid feature imbalance caused by differences in units. Then, the normalized specifications are mapped to a 128-dimensional specification embedding vector through a multilayer perceptron (MLP). This embedding vector can encode the numerical features into a high-dimensional vector representation, making it compatible with semantic features.

[0041] Next, the 896-dimensional semantic projection vector is concatenated with the 128-dimensional specification embedding vector and input into a learnable fusion weight matrix. The final 896-dimensional SKU initial feature vector is obtained through the ReLU nonlinear activation function. This vector simultaneously represents the semantic information, specification features, and potential combinatorial constraint information of the SKU, and is used as an SKU node in the subsequent construction of the subordinate hypergraph. This completes the entire process of multimodal feature representation preprocessing, semantic feature extraction, and specification feature mapping of the original SKU data, providing a unified high-dimensional representation and computational foundation for the hypergraph nodes.

[0042] Numerical specifications refer to quantifiable numerical information of SKUs, such as size, weight, and quantity, which can be directly used for calculation or normalization. Specification descriptions are textual attribute descriptions or descriptive information, such as material being aluminum alloy and color being dark blue.

[0043] Attribute key-value pairs refer to the structured information pairs of SKUs recorded in the ERP system, such as the color being red and the model number being X200. These can be used to parse SKU features and serialize the input model.

[0044] An ERP system, or Enterprise Resource Planning system, is an information management system used by enterprises to uniformly manage business processes such as procurement, inventory, production, sales, and finance. In this solution, the ERP system is mainly responsible for storing and pushing basic product data, inventory data, and product combination relationship data, providing the original data source for subsequent SKU modeling.

[0045] The BERT model is a pre-trained language model based on the Transformer architecture. It acquires text semantic features through bidirectional context learning and is applied to tasks such as text classification, semantic matching, information extraction, and product understanding. This solution uses the BERT model to semantically encode text information such as product titles and attribute descriptions to extract product semantic feature vectors.

[0046] SPU stands for Standardized Product Unit, a standardized abstract entity used in e-commerce to represent the same type of goods, corresponding to a major product category. For example, the iPhone 16 Pro can be considered an SPU, defining the product itself rather than its specific sales specifications.

[0047] SKU stands for Inventory Unit, a specific sales unit under SPU, used to differentiate products by different specifications, colors, capacities, or configurations. For example, under the same SPU, iPhone 16 Pro, different configurations such as black 256GB, white 512GB, and blue 1TB correspond to different SKUs. Each SKU has independent inventory, price, and sales data. SKU nodes are composed of SKU entities, with node attributes being initial feature vectors. Hyperedges are formed by set relationships, and a hyperedge can connect multiple SKU nodes simultaneously. The dependent hypergraph is composed of the set of SKU nodes and the set of dependent hyperedges.

[0048] The learnable linear projection layer is a trainable matrix mapping that maps the 768-dimensional semantic vectors extracted by the BERT model to an 896-dimensional space to match the dimensionality requirements of subsequent feature fusion.

[0049] The learnable fusion weight matrix is ​​the weight parameters of the final unified 896-dimensional initial feature vector obtained by linearly mapping the 1024-dimensional vector obtained by concatenating the semantic projection vector and the dimensional embedding vector through a trainable matrix and adding ReLU activation.

[0050] Text attribute information refers to the descriptive or categorical characteristics of SKUs, such as textual information like product title, color, and material; specification parameter information includes specification descriptions and numerical specification parameters.

[0051] The package combination relationship refers to the combination logic between SKUs, such as the relationship between main components and accessories, and optional accessories, indicating that the SKUs belong to the same sales combination unit;

[0052] Read the real-time inventory data corresponding to each SKU node, and regard the SKU nodes connected by the same hyperedge in the subordinate hypergraph as associated nodes with a combined co-occurrence relationship; construct a sensitive adjacency matrix for association constraint analysis based on the real-time inventory data between associated nodes; and analyze the current SKU inventory status, operation strategy status and manual intervention marking status for each SKU node to perform editability filtering and generate an editable mask set for limiting the editing scope.

[0053] Identify whether there is a co-occurrence relationship between SKU nodes based on the hyperedge connections in the subordinate hypergraph; for SKU node pairs with co-occurrence relationships, further analyze their inventory status differences and inventory balance changes, and based on the inventory status values, use the formula... Calculate the strength of sensitive associations between nodes, where The sensitive adjacency matrix element represents the strength of the sensitive association. A larger value indicates a greater difference in inventory status among different SKU nodes. For example, if at least one SKU has low inventory and both belong to the same set, then visual or logical binding needs to be strengthened to reflect the enhanced association. It is used for association constraint analysis. In subsequent SKU combination binding or local editing processes, sensitive adjacency matrix elements can be used as weights to adjust the strength of visual binding or collision probability between SKUs, ensuring that SKUs with low inventory are given priority when combining or operating.

[0054] This is the Sigmoid function, used to map inventory discrepancies to continuous correlation strength values; The sensitivity coefficient is used to control the sigmoid function's sensitivity to inventory discrepancies. A larger value indicates a stronger response to inventory status differences. It is determined by performing a grid search on the validation set.

[0055] and For inventory status values ​​at different SKU nodes, The value represents the inventory coordination missing quantity. The larger the value, the more likely there is a shortage of inventory at least one SKU node. In this case, the inventory sensitivity correlation strength between the corresponding SKU nodes will be increased to reflect the importance of the SKU to the overall display structure of the package.

[0056] and Indexing different SKU nodes; This is an indicator function that indicates whether different SKU nodes co-occur in the same set combination, i.e., whether they are related nodes. If they are, output 1; otherwise, output 0. For associated nodes;

[0057] When the inventory status value is lower than the preset replenishment threshold, higher than the preset clearance threshold, or when there is a manual editing request, the corresponding SKU is marked as an editable node and an editing identifier is assigned to the node. Finally, an editable mask set is generated to limit the range of SKUs that subsequent local editing operations are allowed to operate on, thereby avoiding the editing behavior from affecting the association structure of non-target SKU nodes.

[0058] Real-time inventory data refers to the current available inventory quantity of each SKU in the system, which is used to reflect the real-time status of the inventory; among which, the inventory status value is a value after normalizing the real-time inventory data, which is used to quantify the inventory level of the SKU, so as to measure the inventory synergy or difference between SKUs when calculating the inventory sensitive adjacency matrix.

[0059] To avoid identity confusion in multiple iterations, identity codes are generated based on the initial feature vector, sensitive adjacency matrix, and editable mask set through standardization, quantization, and subordinate path splicing analysis.

[0060] For the initial feature vector of each SKU node, global mean and standard deviation standardization is first performed. Each element in the feature vector is subtracted from the corresponding mean and divided by the standard deviation. Then, the discrete integer vector is obtained by quantization through the element-wise rounding function.

[0061] The quantized integer vector is then concatenated with the SKU subordinate path string and the current timestamp to form a unified sequence, and then input into the SHA-256 hash function to generate an identity code, so that each SKU node has a unique and traceable identity identifier throughout the entire process;

[0062] During the generation process, all identity codes and their initial states are written into the global binding word, and new state information is appended in subsequent iterations to realize the dynamic evolution tracking of SKU node identities. This hierarchical hash identity code is used to generate an identity index, providing a unified identity credential for SKU nodes in subsequent visual anchor binding, local editing operations, drift monitoring, and state backtracking, so that each SKU node remains identifiable and unique in multiple iterations.

[0063] The SKU dependency path string describes the hierarchical relationship and structural position of the current SKU in the dependency hypergraph. For example, a product set includes a main component, a mobile phone, and an accessory, a charger. The corresponding dependency path string can be represented as: mobile phone: mobile phone; charger: mobile phone / charger. Its function is to distinguish two SKUs with very similar semantic features but in different combination structures. For example, charger in set A: mobile phone A / charger; charger in set B: tablet B / charger.

[0064] In this embodiment of the invention, by performing multimodal feature representation preprocessing on the original data of multiple SKUs under SPU, fusing semantic and specification features to generate a high-dimensional initial feature vector, and constructing a subordinate hypergraph and sensitive adjacency matrix based on the package combination relationship, it is possible to perform structured analysis on the combination constraints and inventory status differences between SKUs. At the same time, an editable mask set is generated to limit the scope of subsequent operations, ensuring that each SKU node maintains its unique identity and traceability in multiple iterations.

[0065] By utilizing global standardization, quantization, and the identity coding generated by subordinate path concatenation, even SKUs with similar semantic features can be accurately distinguished in different package structures. For example, in a mobile phone package, chargers of the same model generate different identity codes because they belong to different main components, thereby preventing confusion during the binding or editing process.

[0066] By combining real-time inventory data and a sensitive adjacency matrix, the system can dynamically assess the inventory synergy or differences between SKUs, prioritizing SKUs with insufficient inventory for inclusion in the editable scope. This ensures that the combination structure of other SKU nodes is not disrupted during partial editing or addition / deletion operations. This identity indexing mechanism enables SKU nodes to be accurately identified and managed in subsequent visual anchor binding, layout iteration, and status updates, thereby improving the controllability and accuracy of multi-product main image generation and updates. At the same time, it provides direct data support for inventory-driven partial addition, deletion, and modification operations, realizing intelligent layout management of complex product combinations.

[0067] In a preferred embodiment of the present invention, S200: Load the corresponding product image material according to the identity code, and use the mask region convolutional neural network to perform instance segmentation on the product subject in the product image material to obtain the boundary information of each visual anchor point and the pixel-level instance mask;

[0068] Visual feature vectors of boundary information within each visual anchor point are extracted and combined with the initial feature vectors. Cosine similarity is then used to calculate an alignment score for binding and matching, thus constructing a binding matrix. This score quantifies the degree of matching between the SKU semantic description and the visual anchor point features, guiding the system to determine the most likely visual anchor point corresponding to each SKU node. This achieves accurate binding of text features and visual information and cross-modal mapping. Elements in the binding matrix represent the alignment score between each SKU node and each visual anchor point; if an SKU node and a visual anchor point have no corresponding material, the alignment score is 0.

[0069] Convolutional Neural Networks with Masked Regions (CNNs) are deep learning models used for image instance segmentation. They can detect the bounding box of each object in an image and generate a pixel-level binary mask for each object, accurately marking the region occupied by that object. They are commonly used in scenarios such as product matting and medical image segmentation. Specifically, the product image to be processed is input into the backbone feature extraction network of this network, which typically employs a combination of residual networks and feature pyramid networks. The image sequentially passes through convolutional layers, batch normalization layers, and activation function layers, extracting feature maps at multiple scales. The feature pyramid network generates a multi-level feature pyramid from high-resolution low-semantic to low-resolution high-semantic through top-down paths and lateral connections, ensuring that even small-sized products can be effectively detected.

[0070] Then, the feature maps at each level of the feature pyramid are fed into the region proposal network. Next, for each candidate region output by the feature pyramid, a feature map of appropriate scale needs to be selected from the feature pyramid for feature extraction. Subsequently, the fixed-size feature map output by the region of interest alignment is input into two parallel fully connected branches: one branch is responsible for the fine classification and refinement regression of the bounding boxes, outputting the probability of each candidate region belonging to each product category and the further adjusted bounding box coordinates; the other branch is responsible for mask prediction. This branch consists of a fully convolutional network, which takes the feature map output by the region of interest alignment as input and outputs a 28×28×C binary mask, where C is the number of product categories, each channel corresponds to one category, and each pixel value in the channel represents the probability that the position belongs to that product category. After passing through the Sigmoid activation function, the final instance mask is obtained by binarization through a threshold (usually 0.5).

[0071] Finally, post-processing is performed on all detected product instances. First, Non-maximum Suppression (NMS) is applied to remove redundant overlapping detection boxes, retaining the unique optimal detection result for each product. For each retained instance, its refined bounding box and corresponding pixel-level binary mask are output. This mask has the same size as the original image; pixels with a value of 1 in the mask represent the product, and pixels with a value of 0 represent the background. By traversing the connected components in the mask, the independent pixel regions of each product can be extracted for subsequent steps to calculate block independence and edge isolation fields.

[0072] Boundary information uses four normalized values ​​to represent the top-left corner position and width and height of the rectangle, which is used to roughly determine the spatial position of the target in the image; the pixel-level instance mask is a binary matrix with the same resolution as the image, where each pixel marks whether the pixel belongs to the SKU instance, which is used to accurately determine the specific region of the SKU in the image. These two outputs together define the visual anchor point, that is, the spatial region of each SKU in the image, and at the same time provide a localization reference for subsequent cross-modal feature extraction.

[0073] Visual anchors are pixel regions and their geometric information that belong to a single product entity and are extracted from an image through instance segmentation. They typically include bounding boxes and instance masks. As visual blocks of SKU text entities, they establish a two-way binding relationship with SKU nodes and serve as the basic operation unit for local addition, deletion, and modification.

[0074] Subsequently, to map each visual anchor point to a fixed-dimensional feature space, the masked region convolutional neural network extracts features from each boundary region through a region of interest alignment operation. The specific steps are as follows: First, the region corresponding to the bounding box in the feature map extracted by the network backbone is segmented; then, this region is divided into a fixed number of grid cells, and bilinear interpolation is performed on the convolutional features within each grid cell to ensure a fixed spatial resolution and minimize alignment error.

[0075] Finally, the features of all grid cells are summarized or pooled to obtain a single fixed-length vector representation, namely the visual feature vector. This vector is a set of floating-point value vectors that can describe the color, texture, shape and spatial structure of the visual block in the numerical space. Its function is to uniformly map the irregularly sized and shaped blocks in the image to a computable and comparable feature space, providing an operable data representation for cross-modal analysis and binding.

[0076] Region of Interest (ROI) alignment is a key operation in masked region convolutional neural networks. It is used to solve the spatial position offset problem caused by the rounding operation in traditional ROI pooling. ROI alignment uses bilinear interpolation to accurately align the feature maps of candidate regions of different sizes to a fixed size, avoiding pixel-level positional errors and thus improving the accuracy of mask prediction.

[0077] The initial feature vectors need to be projected onto the same dimensional space as the visual features;

[0078] Product image materials refer to the original visual resource data stored in the product material library that corresponds to the current SKU identity code, and are used for subsequent instance segmentation and visual block extraction.

[0079] The Hungarian algorithm is applied to the binding matrix to solve for the optimal match, that is, to determine the optimal visual block matching relationship corresponding to each SKU node, thereby obtaining the globally optimal one-to-one binding scheme; the visual feature vectors of the visual anchor points are re-extracted in conjunction with the layout iteration process, and the binding matrix is ​​updated together with the optimal match of the previous iteration and the current visual feature vector.

[0080] The optimal matching result is a binary binding matrix obtained by the Hungarian algorithm. Each element is 0 or 1, indicating whether a certain SKU node is bound to a certain visual block. 1 indicates binding and 0 indicates unbinding. It is used as a unique and definite mapping between SKU nodes and visual blocks to ensure cross-modal identity consistency and avoid identity confusion or duplicate allocation of visual blocks when local modifications or layout adjustments are made.

[0081] Iteratively update the formula The binding matrix is ​​continuously updated. This formula is derived from the principle of recursive least squares filtering. It balances historical binding with current visual observation, so that the binding matrix maintains identity continuity and visual adaptability during continuous layout adjustment. It avoids SKU identity drift caused by local editing, block scaling, position change or visual style adjustment, thereby ensuring that subsequent lifecycle tracking, independence monitoring and local editing operations always act on the correct visual object.

[0082] in For the first In the round layout iteration, SKU node With visual anchors The alignment score is used to determine the correspondence between SKU identity and visual blocks. The larger the value, the better the feature matching between the node and the anchor point, and the higher the binding credibility.

[0083] The historical weighting coefficient is used to control the contribution ratio of historical binding results to the current binding matrix, balancing historical inertia and current visual observations. It is determined by performing a grid search on the validation set, for example, traversing within the search range [0.5, 0.9] with a step size of 0.1 for each group. The system performs a complete layout iteration process, including binding updates, independence monitoring, and local editing, evaluating the stability of the binding matrix in consecutive iterations and its response speed to changes in visual features to select the optimal approach. The value achieves the best balance between maintaining identity continuity and adapting to real visual changes; For the SKU node in the previous layout iteration With visual anchors The alignment score, used to provide a reference for historical identity continuity, is the inertial part of the binding matrix; For visual anchor indexing, For layout iteration index; This is the initial feature vector, used for alignment with visual features; This represents the transpose operation, used to transpose a column vector. Convert to row vectors for matrix or vector dot product operations; The learnable projection matrix (896×256) is used to map SKU text features to the visual feature space, making cross-modal features comparable. It is obtained using an end-to-end supervised training method. For the first Visual anchor points in round iteration Visual feature vector (256 dimensions); The product of Euclidean norms of vectors is used to normalize the dot product and calculate the cosine similarity. This means that after projecting the initial feature vector of the SKU onto the visual feature space, it is matched and calculated with the current visual anchor point feature to obtain the original alignment score of the two.

[0084] The layout iteration process is a continuous process of adjusting the layout that occurs during subsequent intelligent generation of product main images, local additions, deletions and modifications, material replacement, and visual optimization.

[0085] Based on the editable mask set, it is determined whether the SKU node is eligible for editing. Then, combined with the binding matrix in each layout iteration, the current binding status of the corresponding visual anchor is obtained. A lifecycle trajectory is created for the binding pair formed between each SKU node and the visual anchor, and a corresponding editable state code is assigned, including a stable state and an editable state. The trigger condition for the stable state is that the inventory status value of the corresponding SKU node is within the normal range, that is, neither too low nor too high, and the user has not actively marked the SKU as pending editing. The trigger condition for the editable state is that the corresponding SKU node belongs to the editable mask set.

[0086] stable, editable, and locked are specific status values ​​in the editable status code assigned to each binding pair, representing the current permission level that the binding pair is allowed to be modified.

[0087] Since the binding matrix provides a stable identity anchoring benchmark, the optimal match recalculated after each iteration may temporarily match the same visual anchor point with different SKUs due to layout changes. Updating the state based on temporary matches would cause frequent state jumps and confusion in attribution. However, the binding matrix records the unique correspondence established initially and is smoothly updated through recursive least squares filtering during iterative tracking, ensuring that the editable state is always attached to a fixed SKU and visual anchor point binding pair, so that subsequent local editing operations have a clear and stable operation object. Therefore, the state is obtained by combining the established binding matrix rather than the optimal match in each layout iteration.

[0088] The current binding state of a visual anchor point is the actual state of the anchor point bound to that SKU during the layout iteration for each SKU node and visual anchor point, i.e. whether the block is allowed to be modified, locked, or kept stable.

[0089] The lifecycle trajectory refers to the sequence of records of the state evolution of a visual anchor point in each layout iteration from the first binding to the current iteration round. Specifically, it includes the bounding box position, instance mask, editable state, and timestamp information for each iteration round.

[0090] In practical applications, by establishing a stable and traceable cross-modal binding relationship between SKU nodes and visual anchors, unified mapping management of product text entities and image visual objects is achieved.

[0091] First, instance segmentation technology is used to accurately extract the main product region from the product image, obtaining boundary information and pixel-level instance masks, so that each SKU has a corresponding visual anchor point. Then, the alignment score is calculated by combining the initial feature vector and the visual feature vector, and the globally optimal binding result is obtained through the Hungarian algorithm, thereby avoiding the problem of multiple SKUs competing for the same visual block or the same block being repeatedly assigned.

[0092] Based on this, the binding matrix and lifecycle trajectory continuously record the position changes, state changes and identity relationships of visual anchors during the layout iteration process, so that visual objects can maintain identity continuity after undergoing local editing, scaling, replacement or rearrangement.

[0093] For example, when operations staff need to replace an SKU with a new product of the same type that is out of stock, the system can accurately locate the visual anchor point corresponding to the SKU based on the existing binding matrix, update only the target block, and not mistakenly modify other product main bodies;

[0094] Meanwhile, because the lifecycle trajectory continuously records the evolution of the visual anchor point, its identity can still be accurately tracked even after multiple rounds of layout adjustments, thus avoiding problems such as the drift of the correspondence between visual blocks and SKUs, confusion of status ownership, or misalignment of editing objects.

[0095] In a preferred embodiment of the present invention, S300: Based on the life cycle trajectory, the spatiotemporal evolution of each visual anchor point is analyzed, and the actual independence of each visual anchor point is calculated using the image connected component theory.

[0096] Perform first-order difference calculation on the actual independence to quantify trend changes. When the first-order difference is continuously negative, that is, in two or more adjacent layout iterations, the sign of the first-order difference is continuously negative, marking the possible risk of decreased independence and triggering an early warning.

[0097] The spatiotemporal evolution of each visual anchor point is analyzed. First, the center coordinate sequence and scale sequence of each visual anchor point are input into a bidirectional LSTM temporal encoder, with a hidden layer dimension of 64, to generate a trajectory embedding vector to capture the motion pattern and change trend of the block over time. Then, based on the connected component theory, the evolution is performed using the formula... Calculate the actual degree of independence for each anchor point, where visual anchor point The actual degree of independence reflects the degree of adhesion that has occurred at the current moment; the value range is (0,1], where a value of 1 indicates complete independence, and the smaller the value, the more severe the adhesion, which is used to determine whether isolation or correction needs to be triggered. This represents the total number of visual anchor points. visual anchor point In the The area of ​​independent pixels in the layout iteration is determined by marking the connected components of the instance mask of the visual anchor point and counting the number of pixels with a value of 1 within the mask, which is the total area of ​​independent pixels belonging to that anchor point.

[0098] visual anchor point With visual anchors The area of ​​the overlapping pixels is determined by superimposing the mask of the visual anchor point with the mask of the adjacent anchor point at the pixel level, and counting the overlapping area of ​​the two masks, that is, the area of ​​the pixel positions that are simultaneously marked as 1. This area is the area of ​​the overlapping pixels. and Indexing different visual anchor points;

[0099] The center coordinate sequence refers to the sequence formed by arranging the center point positions recorded in each round of multiple layout iterations of a visual anchor point in chronological order; the scale sequence refers to the sequence formed by arranging the width and height recorded in each round of iterations of the same visual anchor point in chronological order.

[0100] A bidirectional LSTM time encoder is a variant of a recurrent neural network that contains two LSTM layers, one forward and one backward. The forward layer reads the sequence from front to back, and the backward layer reads the sequence from back to front. The final hidden states of the two directions are concatenated as the output to capture the contextual dependencies before and after each time step in the sequence.

[0101] The trajectory embedding vector is a fixed-length vector output by inputting the center coordinate sequence and scale sequence into a bidirectional LSTM. It is 128-dimensional, which is convenient for compressing and representing the motion trajectory and scale change pattern of the visual anchor point over the entire time span.

[0102] Subtract the actual independence value of the previous round from the actual independence value of the current round to obtain the result of performing a first-order difference calculation on the actual independence;

[0103] Based on the boundary information of each visual anchor point, the kernel density estimation method is used to analyze the distribution density of the boundary points of each visual anchor point on the image plane. The contribution of the boundary points is weighted and attenuated by the sensitive adjacency matrix to generate the probability distribution of block overlap at each coordinate to construct a collision probability field. The visual anchor point that triggers the warning will use a reduced bandwidth to increase the local collision probability density, thereby reflecting its potential risk concentration area and realizing the centralization of the probability distribution.

[0104] Specifically, based on the boundary information of each visual anchor point, the boundary points of each visual anchor point are determined. These boundary points are the coordinates of the corresponding visual anchor point on the bounding box. Then, the kernel density of the boundary point set for each visual anchor point is estimated using the formula... Calculate the probability of two blocks overlapping at each location on the image plane; This is the collision probability field. The larger the value, the more boundary points from different visual anchor points are gathered around this location, indicating that a collision has occurred.

[0105] It should be noted that triggering the warning is to improve the sensitivity of the collision probability field constructed in this round and accelerate the generation of the subsequent isolation zone, but the construction of the collision probability field will be executed regardless of whether the warning is triggered.

[0106] Next, a sensitive adjacency matrix is ​​introduced as a decay weight for the contribution of boundary points. To correct the collision probability field;

[0107] in The corrected collision probability field represents the collision probability density at the normalized coordinates (x,y) where visual anchor point boundaries overlap or cluster. It is used to provide the mean collision probability of the internal region of each visual anchor point when calculating the adaptive threshold in the subsequent calculation, so as to dynamically adjust the independence requirement of the anchor point according to the density of the surrounding boundaries. This represents the total number of sampling points at all visual anchor point boundaries. This represents the total number of visual anchor points. visual anchor point The number of sampling points on the boundary As a denominator, it serves a normalization function, ensuring The integral over the entire plane is approximately 1; To determine the sensitive correlation strength between visual anchors, subtract this value from 1 and then multiply it together. This reduces the attenuation factor of anchors adjacent to multiple high-inventory-related anchors, thus suppressing the collision contribution of their boundary points. visual anchor point A single sampling point on the boundary represents a two-dimensional coordinate vector; It is an exponential function with base e. The spatial coordinates of the location where the collision probability is to be calculated; For this location and boundary sampling point The Euclidean distance between them; This is the multiplication operator;

[0108] Pi is a constant. visual anchor point The kernel density estimation bandwidth is dynamically determined based on whether an early warning is triggered: when a visual anchor point triggers an early warning, the bandwidth of the boundary points of that anchor point is adjusted to... By reducing the bandwidth, the collision probability distribution of the anchor point becomes more concentrated, and the local peak value is higher, thereby increasing the average collision probability in the region inside the visual anchor point, and thus increasing the adaptive threshold, making the system more sensitive to the anchor point; if no warning is triggered, the base bandwidth is used; the base bandwidth is determined according to Silverman's rule of thumb. This is the adjusted bandwidth; Basic bandwidth;

[0109] in The reduction factor is used to proactively reduce the kernel density estimation bandwidth of a visual anchor point when it triggers an alert, thus achieving a forward-looking correction. Its value is determined by performing a grid search on the validation set, typically ranging from 0.5 to 0.95, selecting the value that best minimizes the F1 score of the independent detection, with a value range less than 1. The validation set consists of a batch of pre-labeled e-commerce main image data, where the independence status of each product in each image has been manually labeled, and the inventory data of each product is also recorded. Indicates Centered on, with a bandwidth of The two-dimensional Gaussian kernel function value reflects the position (x,y) affected by the sampling points. The degree of influence increases with distance; the closer the distance, the higher the value. The farther the distance, the lower the value. This represents the total influence of all sampling points at the corresponding anchor point on the position (x, y);

[0110] Multiply the total impact mentioned above by the sensitivity attenuation factor. Indicates if anchor point When an anchor point is adjacent to a number of high-inventory-related anchor points, the collision contribution of its boundary points is suppressed. For two products with high specific inventory correlation (such as main components and supporting accessories), commercially they tend to be displayed together, allowing a certain degree of visual proximity or even slight overlap, and should not be regarded as high collision risk. Therefore, by reducing the contribution of such anchor points to the collision probability field through attenuation factors, the collision probability field can better reflect the collision risk between dissimilar products that need to be strictly isolated, rather than similar or supporting products that are commercially allowed to be close.

[0111] The collision probability field not only considers the spatial boundary proximity, but also suppresses the collision contribution between anchors with high commercial relevance through the inventory-sensitive adjacency matrix. At the same time, it reduces the bandwidth when triggering an early warning to increase the local peak value of risk anchors, so that the final adaptive threshold can comprehensively reflect the three dimensions of spatial geometry, commercial relevance and trend risk.

[0112] The image plane is the two-dimensional pixel space where the product image material is located, that is, the product image plane after being uniformly scaled to 1024×1024 pixels. On this plane, each coordinate point corresponds to a pixel position on the image. The collision probability field is the probability density of multiple visual anchor point boundaries overlapping at this pixel position.

[0113] Based on the collision probability field, the mean collision probability density is obtained, and after performing a nonlinear mapping on the mean collision probability density, a risk coefficient for threshold correction is obtained. ,in The risk coefficient has a value range of [0,1), which represents the current spatial compression risk level of the visual anchor point. The larger the value, the higher the risk, and the higher the corresponding adaptive independence threshold. It is used to characterize the local collision risk state of the visual anchor point in the current layout. It is an exponential function with base e, used to implement saturation mapping, when When the exponent term is close to 1, Approaching 0; when At that time, the exponent term is close to 0. Approximately 1; where e is the natural constant, approximately 2.718; It is the average value of the collision probability field in the region inside the visual anchor point, that is, the average collision probability density, which reflects the density of the boundary around the anchor point. The reference density value is used to... Scaling to a range that matches the exponential function allows you to take the median of all anchor points;

[0114] An independent threshold range is preset, and the risk coefficient is used as a mapping ratio factor to calculate the corresponding position of the risk coefficient in the independent threshold range, thereby obtaining the adaptive threshold corresponding to the visual anchor point, which is used to guide layout adjustment and correction operations. Among them, the larger the risk coefficient, the closer its corresponding interval mapping position is to the upper limit of the preset independent threshold; the smaller the risk coefficient, the closer its corresponding interval mapping position is to the lower limit of the preset independent threshold, so that the adaptive threshold can be dynamically adjusted according to the current collision risk state.

[0115] Through formula Calculate and obtain the adaptive threshold corresponding to each visual anchor point, where For adaptive threshold, and The lower and upper limits are preset independent threshold intervals; This is the risk correction amount, representing the proportional scaling of the threshold span using the risk coefficient.

[0116] The preset independent threshold intervals are obtained using the mean-standard deviation method;

[0117] The boundary point set and collision probability field of each visual anchor point are extracted. For any two visual anchor points, the boundary points are connected pairwise and integrated. The average value is used to calculate the propagation weight. The propagation weight is compared with a preset propagation threshold to determine the existence of a directed edge. If the propagation weight exceeds the preset threshold, a possible adhesion propagation path exists between the two anchor points. A directed edge exists in the graph, and the visual anchor points are used as graph nodes, with the weight of the directed edge representing the corresponding propagation weight, forming a risk propagation graph. This graph, through the structure of nodes and directed edges, reflects the spatial distribution and propagation network of potential adhesion risks in the entire layout. It can identify key anchor points where adhesion may occur in clusters, predict the trend of adhesion spreading along the path in the graph, and guide layout optimization and local editing operations. This supports the prediction and control of adhesion risks during multiple iterations, thereby maintaining the spatial independence of visual blocks and layout stability.

[0118] For any two different visual anchors in the risk propagation map, firstly, extract the set of boundary points for each anchor. These boundary points represent the contour coordinates occupied by the anchor in the image. Then, connect each pair of boundary points in the two boundary point sets and accumulate the corresponding collision probability density at each position on the connection. Integrate the collision probability density along the line and then average all point pairs to obtain the propagation weight from one visual anchor to another. This weight reflects the spatial continuity and strength of the potential adhesion between the two anchors. That is, if there is a continuous high-probability collision area between the boundaries of the two anchors, the larger the weight, the higher the probability that the adhesion will propagate along this path in the future.

[0119] Based on the risk propagation graph, potential adhesion centers are identified through out-degree and in-degree analysis of graph nodes, and potential adhesion propagation chains are identified through depth-first search. The risk role results are then summarized, providing clear intervention targets and priorities for subsequent layout optimization and local editing operations.

[0120] The risk role results include potential adhesion centers and potential adhesion propagation chains, which are used to determine priority intervention areas, propagation blocking nodes, and local location correction targets in the subsequent layout adjustment process;

[0121] Calculate the sum of the in-degree and out-degree of each graph node. The graph node with the highest degree is marked as a potential adhesion center, indicating that there are multiple high-collision-probability propagation paths around this anchor point, and it is most likely to become the core area where multiple anchor points merge in the future.

[0122] For each directed edge, its direction represents the collision risk from the anchor point. To the anchor point The main trend of dissemination, if The mainstream direction of dissemination is Towards , indicating anchor point The boundary area is being anchored. Collision probability field compression; and From the anchor point To the anchor point Propagation and from the anchor point To the anchor point The propagation weight of the message;

[0123] In the risk propagation graph, starting from the potential adhesion center, a depth-first search is performed along the outgoing edge direction to identify directed paths with a length greater than or equal to 2, thus forming a potential adhesion propagation chain.

[0124] If the actual independence is lower than the adaptive threshold, and the risk role result of the graph node belongs to the potential adhesion center or the head or middle position of the potential adhesion propagation chain, then the locking flag is locked, and the editable state code will be updated and supplemented; otherwise, the locking flag is unlocked, and no update processing is performed.

[0125] After updating and supplementing the editable status codes, the status codes include stable, editable, and locked. The triggering conditions for the stable status after the update and supplement are that the inventory of the SKU is within the normal range, the user has not actively marked it, and it is not in a locked state. The triggering conditions for the editable status after the update and supplement are the same as before the update and supplement.

[0126] The locking condition is that the actual independence is lower than the adaptive threshold; at this time, even if the inventory conditions are met or the user marks it, the visual anchor point cannot be modified.

[0127] Potential sticking centers represent the core areas most likely to merge with multiple anchor points in the future. Temporarily locking these centers means that during local editing operations, these anchor points will not be selected for deletion, scaling, or replacement, and their isolation potential field strength will be enhanced. The purpose of this is to prevent accidental operations triggered by user or inventory events, such as mistakenly deleting key products forming sticking centers or incorrect scaling that accelerates the sticking and deteriorates, thus disrupting ongoing interventions, before the system completes the isolation enhancement and rebinding correction of these anchor points.

[0128] The risk role result of the graph node belongs to the potential adhesion center, indicating that the anchor point needs to be temporarily locked to prevent accidental operation before the isolation enhancement and rebinding correction are completed;

[0129] Since the head and middle positions are key transmission nodes for the sticky propagation, they need to be locked first to interrupt the propagation path. Therefore, when the risk role result of a graph node belongs to a potential sticky propagation chain, the tail node is not locked temporarily because it is squeezed by other nodes rather than actively squeezing others; the middle position is a graph node excluding the head and tail nodes.

[0130] If the actual independence is lower than the adaptive threshold, it means that the pixel-level adhesion of the anchor point has exceeded the system's tolerance limit based on its spatial compression risk. Locking the anchor point at this point means suspending any editing operations on it, as modifications might worsen the adhesion. Simultaneously, it triggers the editing priority flow of subsequent steps, restoring its independence by readjusting block positions, strengthening isolation zones, etc. The lock is lifted once the independence rises above the threshold. Locking is a protective measure to prevent potentially worsening operations on the problematic anchor point before the problem is resolved.

[0131] In this embodiment of the invention, by analyzing the life cycle trajectory and spatial evolution of visual anchor points, combining the image connected component theory to calculate the actual independence of each anchor point, and further constructing a collision probability field based on boundary point kernel density estimation and inventory-sensitive adjacency matrix, the potential adhesion risk of visual blocks is quantified and proactively identified.

[0132] By performing a nonlinear mapping on the mean collision probability density, a risk coefficient is generated and mapped to an adaptive independent threshold, enabling the system to dynamically adjust the independence requirements for each anchor point. Then, a risk propagation graph is constructed based on the propagation weight and preset propagation threshold. Potential adhesion centers and propagation chains are identified through graph node out-degree, in-degree, and depth-first search, thereby providing locking flags and priority intervention basis for editing operations.

[0133] For example, when the visual anchor points between the main component (phone) and multiple accessories of a set show a tendency to stick together during continuous iteration, the system can automatically calculate the collision probability field, update the adaptive threshold, and mark the main component anchor point as locked to prevent accidental operation from damaging the potential sticking control. At the same time, it can prioritize adjusting the position of intermediate chain nodes or enhance the isolation potential energy to ensure the stability and independence of the overall layout.

[0134] Under this mechanism, partial addition, deletion and modification operations of multiple product main images can be carried out based on precise quantified risk data and traceable status, avoiding visual block drift and layout abnormalities caused by single-round operations or inventory changes, and realizing safe management and forward-looking intervention of SKUs and visual blocks in continuous iteration.

[0135] In a preferred embodiment of the present invention, S400: a multi-scale Gabriel filter is used to extract multi-directional edge features from the visual anchor point to obtain the edge tensor used for the construction of isolation potential energy;

[0136] The local image region covered by the instance mask of each visual anchor point is cropped and scaled to a size of 256×256 pixels, then converted into a grayscale image. Three Gaussian derivative filters at different scales are used to perform convolution operations on the grayscale image: at each scale, the first derivative response in four directions (horizontal, vertical, and two diagonal directions) is calculated to obtain the edge intensity map for each direction. The edge intensity maps of the four directions at the same scale are stacked along the channel dimension to generate a multi-directional edge response tensor for that scale. Finally, the edge response tensors generated at the three scales are concatenated along the channel dimension to form a four-dimensional edge tensor, whose dimensions are the height of the feature map × width × number of scales × the value at each scale. The number of directions; where the height and width are both 256, the number of scales is 3, and the standard deviation parameters of the corresponding Gaussian filters are 0.01, 0.02 and 0.04, respectively, corresponding to the detection of fine, medium and macro edge features in the normalized coordinate system, so that the edge tensor can simultaneously encode the local detail edges and the overall shape contour of the product; the standard deviation parameter of the Gaussian filter is used to control the width of the Gaussian smoothing kernel. The smaller the value, the narrower the coverage of the Gaussian kernel and the lower the smoothness, which can preserve the fine edge details in the image; the larger the value, the wider the coverage of the Gaussian kernel and the higher the smoothness, which will suppress noise and small-scale textures and only retain macro edge contours;

[0137] The number of directions at each scale is 4, corresponding to the four directions of horizontal, vertical, main diagonal, and secondary diagonal; therefore, the total number of channels is the number of scales × the number of directions at each scale = 12. Each element in this tensor is a floating-point value, representing the edge response intensity at the corresponding pixel position, corresponding scale, and corresponding direction.

[0138] The edge tensor is used to determine the principal direction of the diffusion tensor in the subsequent anisotropic isolation potential field. Specifically, by analyzing the principal direction of the edge tensor at each pixel position, the eigenvector direction of the diffusion tensor is determined, so that the diffusion proceeds along the edge tangent and is suppressed perpendicular to the edge normal, thereby achieving edge-preserving isolation diffusion. That is, the isolation potential field decays outward along the product edge contour, rather than spreading uniformly in all directions, ensuring that the isolation band acts precisely on the product boundary rather than intruding into the product interior.

[0139] The multi-scale Gabor filter is a multi-scale Gabor filter that can extract edge features of an image at multiple scales and in multiple directions.

[0140] Based on the edge tensor, determine the edge tangent direction and normal direction at each pixel location;

[0141] The tangential diffusion coefficient and normal diffusion coefficient are preset and used as diffusion reference benchmarks in the tangential and normal directions of the edge, respectively. Combined with the risk role results, the normal diffusion factor is dynamically generated based on the propagation risk level corresponding to each visual anchor point. The normal diffusion factor is used to scale the preset normal diffusion coefficient to obtain the normal diffusion coefficient corresponding to the visual anchor point. The normal diffusion coefficient is used to characterize the propagation ability of isolation potential energy perpendicular to the edge contour direction.

[0142] The propagation risk level corresponding to each visual anchor point is determined based on the risk role result. Specifically, when the anchor point belongs to the potential adhesion center, its propagation risk level is the highest risk. If it belongs to the potential adhesion propagation chain, it is further subdivided according to its position in the chain: if it is the first node of the chain, that is, the source of propagation, it is defined as the highest risk; if it is a node in the chain, it is defined as medium to high risk; if it is a node at the end of the chain, it is defined as low risk. If the anchor point does not meet any of the above conditions, it is defined as no risk, and the normal diffusion factor is set to 1, that is, no additional suppression is performed.

[0143] If the visual anchor point is a potential adhesion center and located in the potential adhesion propagation chain, the propagation risk level is determined according to the highest level among them;

[0144] Through formula The normal diffusion factor is calculated and obtained, where The normal diffusion factor corresponding to the visual anchor point is used to scale and modulate the preset normal diffusion coefficient. The larger the value, the larger the scaled normal diffusion coefficient, and the weaker the normal diffusion suppression, so that the isolation potential energy forms a wide and gentle distribution, generating a dispersed and gentle repulsive force at the propagation end. and The lower and upper limits of the normal diffusion factor threshold are determined using a grid search method. This represents the position number of the visual anchor point on the chain. The number of visual anchors contained in the chain; This represents the normalized position of the node in the chain; the preset normal diffusion coefficient is 1.

[0145] Construct a diffusion tensor at each pixel location based on preset tangential and normal diffusion coefficients;

[0146] The diffusion tensor is a 2×2 symmetric positive definite matrix, which is constructed through eigenvalue decomposition. The matrix has four floating-point numbers as its four elements, and both of its eigenvalues ​​are positive. This is the diffusion tensor, used to control the diffusion rate in different directions; The preset tangential diffusion coefficient is set to 1, which represents the diffusion rate along the tangential direction of the edge. The normal diffusion coefficient, ranging from 0 to 1, represents the diffusion rate along the normal direction of the edge. It is a tangential unit vector. The direction is parallel to the main direction of the edge; It is the normal unit vector. The direction is perpendicular to the edge, that is, pointing outwards or inwards; The outer product of the tangent vectors is a 2×2 matrix with elements of . , ( , =1,2) are used to extract the components along the tangential direction; The outer product of the normal vectors is a 2×2 matrix with elements of... It is used to extract the component along the normal direction. This is the transpose symbol for a matrix or vector, converting a column vector into a row vector, which is used to calculate the outer product.

[0147] The normal diffusion coefficient is used to control the diffusion rate of the isolation potential energy field along the direction perpendicular to the edge. The larger the value, the faster the isolation potential energy diffuses away from the boundary, forming a wider isolation zone; conversely, the isolation potential energy mainly diffuses tangentially along the edge, and diffusion in the direction perpendicular to the edge is suppressed. By reducing the normal diffusion coefficient, it is possible to prevent the isolation potential energy from eroding into the product, while controlling the intensity of the outward thrust.

[0148] The edge tangential direction angle is the angle between the edge tangential direction and the horizontal axis of the image. and for The cosine and sine values;

[0149] The cross product of the tangential vector and the cross product of the normal vector are used to project any vector onto the tangential and normal directions, respectively, thereby enabling the diffusion tensor to independently control the propagation rate in different directions.

[0150] The preset tangential diffusion coefficient is 1 and remains constant because the isolation potential energy needs to flow freely along the edge contour of the goods to maintain the continuity of the isolation zone, and does not need to be adjusted according to the risk level. That is, the diffusion in the direction of the edge contour will not cause the goods to erode each other, so it always spreads at the maximum rate.

[0151] The eigenvalues ​​include the preset tangential diffusion coefficient and the normal diffusion coefficient;

[0152] Using the diffusion tensor as a directional constraint in the propagation of isolation information, with the visual anchor point boundary as the high potential energy region and the image boundary as the low potential energy region, the anisotropic diffusion solution algorithm is used to simulate the diffusion process of isolation information from the high potential energy region to the low potential energy region, thereby obtaining the isolation potential energy distribution of the corresponding visual anchor point in the image space, forming an isolation potential energy field, which is used to apply dynamic isolation constraints to the risk role results.

[0153] Isolation information refers to the scalar signal of isolation push intensity, which uses the potential energy value of each pixel position in the isolation potential energy field as its carrier. That is, each element in the field is an instantiation of isolation information at the corresponding pixel position. Its function is to describe the distribution of isolation intensity radiating outward from the visual anchor point boundary. The higher the potential energy, the stronger the isolation push needs to be applied at that position.

[0154] Image boundary refers to the outer contour of the entire image. When solving the diffusion equation, the image boundary is set as a low potential energy region with a potential energy value of 0 because the isolation push only needs to occur between the items inside the image. When the image boundary is reached, the isolation effect should decay to zero, and there is no need to push the items out of the frame.

[0155] Anisotropic diffusion solution algorithms refer to algorithms used to solve steady-state anisotropic diffusion equations. Numerical methods include the finite difference method, which discretizes partial differential equations into a system of linear equations, and the conjugate gradient iterative algorithm for solving this system of equations, with the convergence condition being that the relative residual is less than 1 / 3. Its input is the diffusion tensor and boundary conditions, and its output is the numerical solution of the isolation potential field; the boundary conditions are obtained by setting the potential energy value of the high potential energy region to 1 and the potential energy value of the low potential energy region to 0.

[0156] in The isolation potential field corresponding to the visual anchor point is a scalar function defined on the image plane. , representing the potential energy value at position (x,y); for The gradient is a two-dimensional vector field pointing to... The direction of fastest growth, with its magnitude representing the rate of change; The product of the diffusion tensor and the gradient yields a two-dimensional vector representing the flux after directional modulation. This vector reflects the magnitude and direction of the flow of the isolation potential energy along different directions on the image plane. Specifically, a large tangential component indicates smooth transmission along the edge contour, while a small normal component indicates obstructed transmission in the direction perpendicular to the edge. The sum of the partial derivatives with respect to the flux is the divergence operator, representing the degree of convergence or divergence of the flux. The diffusion term is positive when the potential energy at a point is higher than that of its neighborhood, indicating a net outflow of potential energy; and negative when the potential energy at a point is lower than that of its neighborhood, indicating a net inflow of potential energy. This term describes the redistribution trend of potential energy caused by spatial inhomogeneity. This is the attenuation coefficient, used to control the natural attenuation rate of potential energy as distance increases; The term is a reaction term or a zero-order term, representing the exponential decay of potential energy. It reflects the degree to which potential energy naturally decays with increasing distance. That is, the larger the potential energy value, the faster the decay rate, exhibiting an exponential decrease.

[0157] The steady-state anisotropic diffusion equation is derived from partial differential equations;

[0158] The isolation potential field is a two-dimensional matrix with the same size as the original image. Each element is a floating-point number with a value range of [0,1], representing the isolation potential value at the corresponding pixel position. It provides a quantitative basis for isolation push in subsequent steps. That is, by calculating the gradient at a certain position in the field, the direction and intensity of the isolation push that position should receive are obtained, which is used to guide the fine-tuning of the visual anchor point boundary and the control of the spacing between adjacent anchor points.

[0159] In this embodiment of the invention, multi-scale Gabriel filters are used to extract multi-directional edge features from visual anchors, generating an edge tensor that can encode local details and overall contour information. Based on this tensor, the edge tangential direction and normal direction of each pixel are determined. Combined with preset tangential and normal diffusion coefficients and the risk level of the visual anchor, a normal diffusion factor is dynamically generated, thereby constructing a two-dimensional anisotropic diffusion tensor to simulate the propagation of isolation information in the image plane and form an isolation potential field.

[0160] The potential energy field uses the visual anchor point boundary as the high potential energy region and the image boundary as the low potential energy region. By solving the steady-state anisotropic diffusion equation, the isolation potential energy is allowed to diffuse freely along the edge tangential direction, while it is modulated in the normal direction according to the risk level, thus achieving precise isolation of potential adhesion areas.

[0161] For example, when the main mobile phone component of a set is located in the central area of ​​multiple accessories, the potential energy field can generate an isolation zone that smoothly decays along the edge of the main component, preventing other accessories from approaching the core boundary. At the same time, it is restricted in the normal direction to prevent the isolation effect from intruding into the interior of the product. This ensures that each visual anchor point maintains independence and spatial stability during local editing or layout adjustments, thereby providing a quantitative and operable basis for isolation control for subsequent local addition, deletion and modification operations, and realizing dynamic constraint and safety management of high-risk areas.

[0162] In a preferred embodiment of the present invention, S500: retrieve the visual anchor points corresponding to the editable states in the updated and supplemented set of editable masks, record them as anchor points to be edited, and retrieve the visual anchor points adjacent to the anchor points to be edited, record them as neighboring anchor points;

[0163] Apply a linear enhancement transformation to the isolation potential energy field of the anchor point to be edited and its neighboring anchor points, and reset the isolation potential energy field on its boundary to 1;

[0164] After the enhancement operation is completed, the middle area between the anchor point to be edited and its adjacent anchor points is used as a buffer zone to absorb the displacement impact generated by the local editing operation.

[0165] The potential energy value of each pixel in the isolation potential energy field of the anchor point to be edited is multiplied by (1 + enhancement factor) to achieve a linear enhancement transformation. This facilitates a rapid increase in the repulsion strength of the anchor point boundary to its external neighbors without needing to resolve the diffusion equation. This provides greater leeway for subsequent local deletion, addition, or scaling operations and prevents accidental adhesion with adjacent anchor points due to insufficient isolation during the operation. Simultaneously, resetting the boundary potential energy value to 1 ensures that the physical significance of the anchor point as the starting point for isolation is not compromised.

[0166] The enhancement factor is used to temporarily enhance the isolation strength of the visual anchor point to be edited and its adjacent anchor points under the drive of inventory events, so as to reserve space for subsequent local editing operations and prevent new adhesion from being generated during the operation; its value ranges from 0 to 1, and is specifically determined by the grid search method; for adjacent anchor points, the potential energy value of each pixel in the isolation potential energy field of the adjacent anchor point is multiplied by (1 + enhancement factor * 0.5).

[0167] For adjacent visual anchors whose distance from the anchor to be edited is less than a preset distance threshold, this is specifically represented as the set of all pixels on the line segment formed by the two nearest points on the boundary of the two anchors. When the isolation potential energy field of the anchor to be edited is enhanced, the potential energy value of this region forms a bulge higher than its independent attenuation value due to the combined influence of the two enhanced potential energy fields, forming a transitional high potential energy barrier. This barrier is used to absorb the impact force generated by the displacement or deformation of the anchor to be edited during local editing operations, preventing the operation effect from being directly transmitted to the more peripheral stable anchors, thereby avoiding a chain reaction of layout disorder. In short, the buffer isolation zone acts as a shock-absorbing layer between the area to be edited and the stable area.

[0168] In this invention, all parameters are dimensionless by using dimensionless processing technology to remove their dimensions, and all thresholds can be obtained by the mean-standard deviation method.

[0169] Based on the risk role results, the operation priority is determined, and local editing operations are performed on the anchor points to be edited in sequence according to the real-time inventory data and operation priority. Among them, for any point on the buffer isolation zone, any point on the boundary of the anchor point to be edited must not be located on the same side of the buffer isolation zone close to the anchor point to be edited after the local editing operation is performed.

[0170] Based on the risk role results, the transmission risk level corresponding to each visual anchor point is determined, and the operation priority is assigned according to the risk level. For example, the highest risk corresponds to the first priority, medium-high risk corresponds to the second priority, and low risk corresponds to the third priority. For the first priority, this type of anchor point is the source of adhesion transmission and has the highest risk. Prioritizing its handling can effectively block the extension of the transmission chain.

[0171] Local editing operations include deletion, addition, or scaling.

[0172] After a deletion operation, all anchor points adjacent to the deleted anchor point should be moved closer together along the direction of the blank area to fill the gap. However, during the moving process, the constraint of the buffer zone must be maintained and cannot be crossed. For blocks that are far from the deleted block, no movement is required. The goal of this step is not to eliminate all blank areas, but to maintain the compactness and visual balance of the layout through local fine-tuning, while preventing a chain reaction caused by large-scale movement.

[0173] In this embodiment of the invention, the anchor point to be edited and its neighboring anchor points are determined by retrieving the set of editable masks, and a linear enhancement is applied to their isolation potential field. At the same time, the potential value is reset in the boundary region, thereby forming a buffer isolation zone between the region to be edited and the adjacent stable region, effectively absorbing the displacement impact generated by the local editing operation and preventing accidental adhesion.

[0174] In practical applications, this method prioritizes the operation of visual anchors based on the risk role results, prioritizing high-risk propagation sources and sequentially performing deletion, addition, or scaling operations, while adhering to buffer zone constraints to ensure that local adjustments do not disrupt the stability of the overall layout.

[0175] like Figure 2 As shown, embodiments of the present invention also provide a text-driven multi-product intelligent synthesis system for e-commerce main images, comprising:

[0176] The SKU feature modeling module is used to receive multiple sets of raw SKU data under SPU, generate initial feature vectors corresponding to SKU nodes, and construct a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationship of each SKU.

[0177] The anchor binding module is used to load product image materials, obtain visual anchor points, and combine them with the initial feature vectors corresponding to the SKU nodes to construct a binding matrix and an editable mask in each layout iteration.

[0178] The independent risk module is used to obtain the actual independence by analyzing the spatiotemporal evolution of each visual anchor point. Combined with the sensitive adjacency matrix, it constructs a collision probability field and risk role results for risk propagation analysis. Based on independence and risk identification, it locks the flags to update and supplement the editable state code. The risk role results include potential adhesion centers and potential adhesion propagation chains.

[0179] The isolation potential energy construction module is used to extract edge tensors based on visual anchor point boundaries, construct diffusion tensors by combining risk role results, and construct an isolation potential energy field with direction selection characteristics by analyzing the correspondence between edge structure direction and risk propagation path.

[0180] The local adjustment module is used to determine the anchor point to be edited based on the supplemented editable state code, and to perform local editing operations in combination with the isolation potential field and risk role results to complete the multi-product layout adjustment of the e-commerce main image.

[0181] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0182] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A text-driven intelligent multi-product synthesis method for e-commerce main images, characterized in that, The method includes: Receive raw data of multiple SKUs under SPU, generate initial feature vectors corresponding to SKU nodes, and construct a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationship of each SKU. Load product image materials, obtain visual anchor points, and combine them with the initial feature vectors corresponding to SKU nodes to construct a binding matrix and an editable mask in each layout iteration; By analyzing the spatiotemporal evolution of each visual anchor point, the actual independence is obtained. Combined with the sensitive adjacency matrix, a collision probability field and risk role results are constructed for risk propagation analysis. Based on independence and risk identification, the lock flag is updated and supplemented with editable state codes. The risk role results include potential adhesion centers and potential adhesion propagation chains. Edge tensors are extracted based on visual anchor point boundaries, diffusion tensors are constructed by combining risk role results, and an isolation potential field with direction selection characteristics is constructed by analyzing the correspondence between edge structure direction and risk propagation path. Based on the supplemented editable state code, the anchor point to be edited is determined, and local editing operations are performed in combination with the isolation potential field and risk role results to complete the multi-product layout adjustment of the e-commerce main image.

2. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 1, characterized in that, It receives raw data from multiple SKUs under an SPU, generates initial feature vectors corresponding to SKU nodes, and constructs a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationships of each SKU, including: Receive raw data of multiple SKUs under SPU pushed by ERP system, perform multimodal feature representation preprocessing on the raw SKU data, and generate the initial feature vector corresponding to each SKU through semantic feature extraction and specification feature mapping; Using SKU entities carrying initial feature vectors as SKU nodes, and using the hierarchical relationships connecting multiple SKU nodes within the same set as hyperedges, construct a hierarchical hypergraph; Read the real-time inventory data corresponding to each SKU node, and regard the SKU nodes connected by the same hyperedge in the subordinate hypergraph as associated nodes; construct a sensitive adjacency matrix for association constraint analysis based on the real-time inventory data between associated nodes; and perform editability filtering on each SKU node to generate an editable mask set for limiting the editing range. Based on the initial feature vector, sensitive adjacency matrix, and editable mask set, identity codes are generated through standardization, quantization, and subordinate path splicing analysis.

3. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 2, characterized in that, Load product image assets, obtain visual anchor points, and combine them with the initial feature vectors corresponding to the SKU nodes. In each layout iteration, construct a binding matrix and an editable mask, including: Load the corresponding product image material according to the identity code, and use the mask region convolutional neural network to perform instance segmentation on the product subject in the product image material to obtain the boundary information of each visual anchor point and the pixel-level instance mask; Visual feature vectors of boundary information within each visual anchor point are extracted and combined with the initial feature vectors. Cosine similarity is used to calculate the alignment score for binding matching in order to construct the binding matrix. The Hungarian algorithm is applied to the binding matrix to solve for the optimal match. The visual feature vectors of the visual anchors are re-extracted in conjunction with the layout iteration process, and the binding matrix is ​​updated together with the optimal match of the previous iteration and the current visual feature vector. Based on the set of editable masks, the current binding state of the corresponding visual anchor is obtained by combining the binding matrix in each round of layout iteration. A life cycle trajectory is created for the binding pair formed between each SKU node and the visual anchor and the corresponding editable state code is assigned, including stable state and editable state.

4. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 3, characterized in that, By analyzing the spatiotemporal evolution of each visual anchor point, the actual independence is obtained. Combined with the sensitive adjacency matrix, a collision probability field for risk propagation analysis is constructed, including: Based on the life cycle trajectory, the spatiotemporal evolution of each visual anchor point is analyzed, and the actual independence of each visual anchor point is calculated using the image connected component theory. Perform first-order difference calculation on the actual degree of independence, and trigger an early warning when the first-order difference is continuously negative; Based on the boundary information of each visual anchor point, the kernel density estimation method is used to analyze the distribution density of the boundary points of each visual anchor point on the image plane. The contribution of the boundary points is weighted and attenuated by the sensitive adjacency matrix to generate the probability distribution of block overlap at each coordinate to construct a collision probability field. The visual anchor point that triggers the warning will use a reduced bandwidth to increase the local collision probability density and realize the centralization of the probability distribution.

5. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 4, characterized in that, The process of constructing risk role outcomes and supplementing them with editable state codes includes: Based on the collision probability field, the mean collision probability density is obtained, and after performing a nonlinear mapping on the mean collision probability density, a risk coefficient for threshold correction is obtained. Preset independent threshold ranges and use the risk coefficient as a mapping ratio factor to calculate the corresponding position of the risk coefficient in the independent threshold range, and obtain the adaptive threshold corresponding to the visual anchor point, which is used to guide layout adjustment and correction operations. Extract the boundary point set and collision probability field of each visual anchor point, perform pairwise integration on the boundary points of any two visual anchor points, and calculate the propagation weight by taking the average value. Compare the propagation weight with the preset propagation threshold to determine the existence of directed edges. If the propagation weight exceeds the preset propagation threshold, then there are directed edges. Use the visual anchor points as graph nodes to form a risk propagation graph. Based on the risk propagation graph, potential adhesion centers are identified through out-degree and in-degree analysis of graph nodes, and potential adhesion propagation chains are identified through depth-first search to summarize the risk role results. If the actual independence is lower than the adaptive threshold, and the risk role result of the graph node belongs to the potential adhesion center or the head or middle position of the potential adhesion propagation chain, then the locking flag is locked, and the editable state code will be updated and supplemented; otherwise, the locking flag is unlocked, and no update processing is performed.

6. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 5, characterized in that, Edge tensors are extracted based on visual anchor point boundaries, and diffusion tensors are constructed by combining risk role results. Furthermore, by analyzing the correspondence between edge structure directions and risk propagation paths, an isolation potential field with direction-selective properties is constructed, including: A multi-scale Gabriel filter is used to extract multi-directional edge features from visual anchor points in order to obtain the edge tensor used for the construction of isolation potential energy. Based on the edge tensor, determine the edge tangent direction and normal direction at each pixel location; The tangential diffusion coefficient and normal diffusion coefficient are preset and used as diffusion reference benchmarks in the tangential and normal directions of the edge, respectively. Combined with the risk role results, the normal diffusion factor is dynamically generated based on the propagation risk level corresponding to each visual anchor point. The normal diffusion factor is used to scale the preset normal diffusion coefficient to obtain the normal diffusion coefficient corresponding to the visual anchor point. The normal diffusion coefficient is used to characterize the propagation ability of isolation potential energy perpendicular to the edge contour direction. Construct a diffusion tensor at each pixel location based on preset tangential and normal diffusion coefficients; Using the diffusion tensor as a directional constraint in the propagation of isolation information, with the visual anchor point boundary as the high potential energy region and the image boundary as the low potential energy region, the anisotropic diffusion solution algorithm is used to simulate the diffusion process of isolation information from the high potential energy region to the low potential energy region, thereby obtaining the isolation potential energy distribution of the corresponding visual anchor point in the image space, forming an isolation potential energy field, which is used to apply dynamic isolation constraints to the risk role results.

7. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 6, characterized in that, Based on the supplemented editable state code, the anchor point to be edited is determined, and local editing operations are performed in conjunction with the isolated potential field and risk role results, including: Retrieve the visual anchor point corresponding to the editable state within the editable mask set, and record it as the anchor point to be edited. Also retrieve the visual anchor point adjacent to the anchor point to be edited, and record it as the neighboring anchor point. Apply a linear enhancement transformation to the isolation potential energy field of the anchor point to be edited and its neighboring anchor points, and reset the isolation potential energy field on its boundary to 1; After the enhancement operation is completed, the middle area between the anchor point to be edited and its neighboring anchor points is used as a buffer zone to absorb the displacement impact caused by the local editing operation.

8. The text-driven intelligent multi-product synthesis method for e-commerce main images according to claim 7, characterized in that, Based on the supplemented editable state code, the anchor point to be edited is determined, and local editing operations are performed in combination with the isolated potential field and risk role results. This also includes: Based on the risk role results, the operation priority is determined, and local editing operations are performed on the anchor points to be edited in sequence according to the real-time inventory data and operation priority. Among them, for any point on the buffer isolation zone, any point on the boundary of the anchor point to be edited must not be located on the same side of the buffer isolation zone close to the anchor point to be edited after the local editing operation is performed.

9. A text-driven multi-product intelligent synthesis system for e-commerce main images, used to implement the text-driven multi-product intelligent synthesis method for e-commerce main images as described in any one of claims 1 to 8, characterized in that, include: The SKU feature modeling module is used to receive multiple sets of raw SKU data under SPU, generate initial feature vectors corresponding to SKU nodes, and construct a sensitive adjacency matrix and an editable mask set for association constraint analysis by analyzing the inventory status and co-occurrence relationship of each SKU. The anchor binding module is used to load product image materials, obtain visual anchor points, and combine them with the initial feature vectors corresponding to the SKU nodes to construct a binding matrix and an editable mask in each layout iteration. The independent risk module is used to obtain the actual independence by analyzing the spatiotemporal evolution of each visual anchor point, and to construct a collision probability field and risk role results for risk propagation analysis by combining the sensitive adjacency matrix. It also locks the flag based on independence and risk identification to update and supplement the editable state code. The risk role results include potential adhesion centers and potential adhesion propagation chains; The isolation potential energy construction module is used to extract edge tensors based on visual anchor point boundaries, construct diffusion tensors by combining risk role results, and construct an isolation potential energy field with direction selection characteristics by analyzing the correspondence between edge structure direction and risk propagation path. The local adjustment module is used to determine the anchor point to be edited based on the supplemented editable state code, and to perform local editing operations in combination with the isolation potential field and risk role results to complete the multi-product layout adjustment of the e-commerce main image.