Training method, prediction method and related equipment for generating e-commerce scene graph model with harmonious commodity size
By training and adjusting the e-commerce scene graph model, the relative size of the main product and other products is controlled, which solves the problem of inconsistency when generating product images on e-commerce platforms and achieves a more efficient product display effect.
Patent Information
- Application Number
- CN202511330543.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-01-02
AI Technical Summary
When generating product images on e-commerce platforms, the size of unknown products may not match that of other products, resulting in a mismatch between the product and the background in the generated image, which increases costs and reduces the display effect.
By obtaining real e-commerce scene image samples, extracting scene image description text and product segmentation images, training a model to control the relative size relationship between the main product and other products, randomly removing the size labels of some other products, and adjusting the model to generate harmonious e-commerce scene images.
It improves the coordination of generated e-commerce scene images, reduces costs, and enhances the product display effect on e-commerce platforms.
Smart Images

Figure CN121259518A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a training method, prediction method and related equipment for generating a graph model of e-commerce scene with harmonious product sizes. Background Technology
[0002] With the rapid development of the internet and e-commerce, more and more merchants are selling goods online. However, product images on e-commerce platforms often require professional photographers for shooting and editing, which is an expensive cost for small businesses or individual shops. At the same time, improving product display to boost sales has always been a challenge for e-commerce platforms. Using AI image generation technology to automatically generate product display images suitable for e-commerce platforms can effectively improve efficiency and significantly reduce costs.
[0003] Generating product display images that fit e-commerce scenarios has the following problems: For unknown products, their size does not match the size of other products, resulting in a mismatch between the physical size of the product and other products in the final generated image. Summary of the Invention
[0004] The main objective of this application is to propose a training method, prediction method, and related equipment for generating e-commerce scene graph models with harmonious product sizes, aiming to achieve the generation of e-commerce scene graphs with harmonious product sizes.
[0005] To achieve the above objectives, one aspect of this application proposes a method for training a graph model of an e-commerce scene to generate harmonious product sizes. The method includes the following steps: Obtain real e-commerce scene image samples, and extract real scene image description text based on the real e-commerce scene image samples; the real scene image description text includes product physical size label, product name label and description label; Extract real product segmentation images based on the real e-commerce scene image samples; The real product segmentation image and the real scene image description text are used as the first input, and the real e-commerce scene image sample is used as the output. The preset model is trained based on the first input and the output. The physical size labels of the main products in the real scene description text in the first input are retained, and some physical size labels of other products are randomly removed to determine the second input. The trained preset model is adjusted according to the second input and the output to obtain the final model for generating e-commerce scene diagrams with harmonious product sizes.
[0006] In some embodiments, the step of extracting a real product segmentation image based on the real e-commerce scene image sample includes: Based on the product name tag, and using a target detection method, the position of the product is extracted from the real e-commerce scene image sample; Based on the location of the product, a real product segmentation image is segmented from the real e-commerce scene image sample.
[0007] In some embodiments, extracting real scene description text from the real e-commerce scene image sample includes: The real e-commerce scene image samples are input into the multimodal large model to extract the product name tags; The description tags are obtained by describing the real e-commerce scene image sample based on the product name tags as the main description objects; Determine the physical size label of the product based on the actual product segmentation diagram; The actual scene image description text is formed based on the product name label, the description label, and the product physical size label.
[0008] In some embodiments, determining the physical size label of the product based on the actual product segmentation diagram includes: Based on the three-dimensional point cloud data and global scale factor of the predicted scene using the real product segmentation map, the horizontal and vertical endpoint coordinates of the real product in the three-dimensional point cloud data are determined based on the real product segmentation map. The physical size label of the product is determined based on the endpoint coordinates in the horizontal direction, the endpoint coordinates in the vertical direction, and the global scale factor.
[0009] To achieve the above objectives, another aspect of this application proposes a method for generating a graph model prediction method for e-commerce scenarios with harmonious product sizes, comprising: Obtain the segmented white background image of the predicted product, the estimated physical size of the predicted product, and the e-commerce scene image description of the predicted product. Based on the estimated physical size of the predicted product and the e-commerce scene image description of the predicted product, form the scene image description text of the predicted product. The segmented white background image of the predicted product and the scene description text of the predicted product are input into the final model for generating a harmonious e-commerce scene image of the product size, and a physically harmonious e-commerce scene image is obtained.
[0010] To achieve the above objectives, another aspect of this application proposes a training device for generating an e-commerce scene graph model with harmonious product sizes. The training device includes: The description text determination module is used to obtain real e-commerce scene image samples and extract real scene image description text based on the real e-commerce scene image samples; the real scene image description text includes product physical size label, product name label and description label; The segmentation map determination module is used to extract real product segmentation maps based on the real e-commerce scene image samples. The initial training module is used to take the real product segmentation image and the real scene image description text as the first input, and the real e-commerce scene image sample as the output, and train the preset model based on the first input and the output. The adjustment module is used to retain the physical size labels of the main products in the real scene image description text in the first input, randomly remove some physical size labels of other products to determine the second input, and adjust the trained preset model according to the second input and the output to obtain the final model for generating e-commerce scene images with harmonious product sizes.
[0011] To achieve the above objectives, another aspect of this application proposes a predictive apparatus for generating an e-commerce scene graph model with harmonious product sizes. The predictive apparatus includes: The data acquisition module is used to acquire the segmented white background image of the predicted product, the estimated physical size of the predicted product, and the e-commerce scene image description of the predicted product, and to form the scene image description text of the predicted product based on the estimated physical size of the predicted product and the e-commerce scene image description of the predicted product. The prediction module is used to input the segmented white background image of the predicted product and the scene image description text of the predicted product into the final model for generating a product-size-harmonious e-commerce scene image, thereby obtaining a physically-size-harmonious e-commerce scene image.
[0012] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the training method or prediction method described above.
[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method or prediction method described above.
[0014] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the training method or prediction method described above.
[0015] The embodiments of this application include at least the following beneficial effects: This application provides a training method, prediction method, and related equipment for generating e-commerce scene graph models with harmonious product sizes. The training method adds physical size labels to the scene graph description text of real e-commerce scene graph samples, uses real product segmentation images and scene graph description text to perform initial training on a preset model, controls the relative size relationship between the main product and other products through physical size labels, then retains the physical size labels of the main product in the real scene graph description text mentioned in the first input, randomly removes some physical size labels of other products, and adjusts the trained preset model again to obtain a final model sensitive to the physical size of the main product, thereby enabling the final model to generate e-commerce scene graphs with harmonious physical sizes for unknown products. The prediction method forms scene graph description text for predicted products based on the estimated physical size of the predicted products and the e-commerce scene graph description of the predicted products; and inputs the segmentation white background image of the predicted products and the scene graph description text of the predicted products into the aforementioned final model for generating e-commerce scene graphs with harmonious product sizes to obtain e-commerce scene graphs with harmonious physical sizes. Attached Figure Description
[0016] Figure 1 This is a flowchart of a method for training an e-commerce scene graph model to generate harmonious product sizes, provided in an embodiment of this application. Figure 2 This is a flowchart for determining the actual product segmentation diagram provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the determination of a real scene description text, as provided in an embodiment of this application. Figure 4 This is a flowchart illustrating the process of determining the physical size label of a product, as provided in an embodiment of this application. Figure 5 This is a flowchart of a method for generating a graph model prediction method for e-commerce scenarios with harmonious product sizes, provided in an embodiment of this application. Figure 6 This application provides a training device for generating an e-commerce scene graph model with harmonious product sizes. Figure 7 This is a schematic diagram of the structure of an e-commerce scene graph model prediction device for generating harmonious product sizes, provided in an embodiment of this application. Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] The training and prediction methods for generating harmonious e-commerce scene graph models of product size provided in this application relate to the field of information technology. These methods can be applied to terminals, servers, or software running on either a terminal or server. In some embodiments, the terminal may be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software may be an application implementing the training and prediction methods for generating harmonious e-commerce scene graph models of product size, but is not limited to these forms.
[0022] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0023] Figure 1 This is an optional flowchart of a method for training an e-commerce scene graph model to generate harmonious product sizes, provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S104.
[0024] Step S101: Obtain real e-commerce scene image samples, and extract real scene image description text based on the real e-commerce scene image samples; the real scene image description text includes product physical size label, product name label and description label; Step S102: Extract real product segmentation images based on real e-commerce scene image samples; Step S103: Take the real product segmentation image and the real scene image description text as the first input, and take the real e-commerce scene image sample as the output, and train the preset model based on the first input and the output. Step S104: Retain the physical size labels of the main products in the real scene image description text in the first input, and randomly remove some physical size labels of other products to determine the second input. Adjust the trained preset model according to the second input and output to obtain the final model for generating e-commerce scene images with harmonious product sizes.
[0025] It should be noted that the preset model and the corresponding loss function are determined according to the actual application. This embodiment does not impose specific restrictions. For example, the preset model adopts the Kontext flux model, and the loss function is the corrected flow matching loss.
[0026] Real e-commerce scene image samples refer to actual photographed e-commerce scene images, while real product segmentation images refer to images containing only a specific target product. First, real scene image description text is extracted from the real e-commerce scene image samples. This description text includes, but is not limited to, product physical size labels, product name labels, and description labels. The description labels are determined based on the actual application, and this embodiment does not impose specific limitations. Then, real product segmentation images are extracted from the real e-commerce scene image samples. Next, the real product segmentation images and real scene image description text are used as the first input, and the real e-commerce scene image samples are used as the output. A preset model is trained based on the first input and output for initial training. Finally, the physical size labels of the main product in the real scene image description text from the first input are retained, while some other product physical size labels are randomly removed to determine the second input. The trained preset model is then adjusted based on the second input and output to obtain the final model for generating e-commerce scene images with harmonious product sizes. This improves the sensitivity of the final model to the physical size of the main product and enhances the consistency of subsequently predicted e-commerce scene images.
[0027] In some embodiments, see Figure 2 Based on real e-commerce scenario image samples, extract real product segmentation images, including: Step S201: Based on the product name tag, extract the product location from real e-commerce scene image samples using object detection methods; Step S202: Segment the real product image from the real e-commerce scene image sample based on the product's location.
[0028] In a specific embodiment, each product name n in the main item list of the real e-commerce scene image sample is used as a prompt, and the position of the product in the real e-commerce scene image sample is found using a prompt-based object detection method (such as Grounding DINO), thus obtaining the position p of each product in the image of the real e-commerce scene image sample. The position p of the product in the image is then used as a prompt to input into a segmentation method that supports position as prompt (such as segmentanything), thus obtaining a segmentation map of the product in the real e-commerce scene image sample. The corresponding real product segmentation map s is then extracted from the real e-commerce scene image sample using the segmentation map. The real product segmentation map is a white background image, and all real product segmentation maps are represented by S.
[0029] In some embodiments, see Figure 3 Based on real e-commerce scenario image samples, extract real scenario image description text, including: Step S301: Input real e-commerce scene image samples into the multimodal large model and extract product name tags; Step S302: Describe the real e-commerce scenario image sample based on the product name tag as the main description object to obtain the description tag; Step S303: Determine the physical size label of the product based on the actual product segmentation diagram; Step S304: Generate a realistic scene image description text based on the product name label, description label, and product physical size label.
[0030] We collect real e-commerce scene image sample data, denoted as Y, and use a multimodal large model (such as Doubao-Seed-1.6) to extract the main items contained in the real e-commerce scene image samples, obtaining a list n of main items contained in each real e-commerce scene image sample. We then describe each real e-commerce scene image sample using each item as the main descriptive object, obtaining image description text t for each product. For example, if an image contains two main items, there will be two corresponding descriptions. All description texts for real e-commerce scene image samples are denoted as T. The description text follows a specified format for optimal effect, such as product name + placement + background description + style, etc.
[0031] Determine the physical size label of the product based on the actual product segmentation image, and add the obtained physical size label to the subject name mentioned in the above-obtained image description text T. For example, a skin care product (h:xx,w:xx) is placed on a table (h:xx,w:xx), resulting in a new image description text T1. T1 contains the image description text corresponding to each item.
[0032] In some embodiments, see Figure 4 Determine the physical size label of the product based on the actual product segmentation diagram, including: Step S401: Based on the 3D point cloud data and global scale factor of the predicted scene using the real product segmentation map, determine the horizontal and vertical endpoint coordinates of the real product in the 3D point cloud data using the real product segmentation map. Step S402: Determine the physical size label of the product based on the horizontal endpoint coordinates, the vertical endpoint coordinates, and the global scale factor.
[0033] A monocular geometric estimation model is used to predict 3D point clouds of real-world product scene images, and a global scale factor is combined to measure physical size. First, the predicted 3D point cloud of the real-world product scene image retains affine invariance (i.e., the shape and relative proportions of the point cloud are accurate, but the global scale and offset are unknown). This step ensures that the relative sizes and distance relationships between objects are accurate. Based on the affine-invariant point cloud, an additional global scale factor is predicted (learned from global features of the image via an MLP module). This factor represents the scaling factor that converts the affine-invariant point cloud to a physical scale. Finally, the physical scale is calculated by multiplying the coordinates of each point in the affine-invariant point cloud by the global scale factor, resulting in a 3D point cloud expressed in physical units (e.g., meters), thus directly measuring the physical size of the objects.
[0034] The calculation process for the final physical scale can be illustrated with a scenario example. Scenario setting: Suppose a picture of a table is taken, and the actual length of the table (physical unit: meter) is estimated.
[0035] Step 1: Predict affine invariant point clouds The image of the table is input into MoGe-2, which outputs the affine invariant point cloud of the table and its full scale factor. The point cloud coordinates are relative and have no physical units, but their internal relative proportions are accurate. For example, the coordinates of the points at the left and right ends of the table in the point cloud are p1=(0,0,0) and p2=(10,0,0), respectively, and the distance between the two points is 10 relative units. This "relative unit" has no actual physical meaning, but it reflects that "the distance between the two ends of the table is 5 times the height of the table legs" (assuming the height of the table legs is 2 relative units in the point cloud).
[0036] Step 2: Predict the global scale factor The model predicts the global scale factor from global image features using an MLP module, assumed to be s = 0.1 meters per relative unit. This factor means that 1 relative unit = 0.1 meters.
[0037] Step 3: Calculate the 3D point cloud at the physical scale Multiplying the coordinates of each point in the affine invariant point cloud by the global scale factor s yields the coordinates in physical units: the physical coordinates of the two ends of the table are: p1′=(0×0.1,0×0.1,0×0.1)=(0,0,0) meters, p2′=(10×0.1,0×0.1,0×0.1)=(1,0,0) meters. The distance between the two points is 1-0=1 meter, meaning the actual length of the table is 1 meter.
[0038] By preserving the relative scale (10 relative units) of the affine invariant point cloud and combining it with a global scale factor (0.1 meters / relative unit), a directly measurable result in the physical world (1 meter) is finally obtained. This process ensures the relative accuracy of the object's internal structure and achieves the conversion from "relative scale" to "physical scale" through the scale factor. The estimated value is in cm, rounded to the nearest whole number and divided by 10, and represented as (h:xx, w:xx) (e.g., 23cm is ultimately represented as 2, 49cm is ultimately represented as 5).
[0039] To achieve the above objectives, please refer to Figure 5 Another aspect of this application proposes a method for generating a graph model prediction method for e-commerce scenarios with harmonious product sizes, including: Step S501: Obtain the segmented white background image of the predicted product, the estimated physical size of the predicted product, and the e-commerce scene image description of the predicted product. Based on the estimated physical size of the predicted product and the e-commerce scene image description of the predicted product, form the scene image description text of the predicted product. Step S502: Input the segmentation white background image of the predicted product and the scene image description text of the predicted product into the final model of the e-commerce scene image used to generate the product size harmony, and obtain the e-commerce scene image with physical size harmony.
[0040] First, the segmented white-background image of the predicted product is obtained as a white-background image of the product, along with the corresponding estimated physical size of the predicted product and the desired e-commerce scene image description. Then, a large language model is used to generate a prompt that matches the final model input based on the obtained physical size of the predicted product and the desired e-commerce scene image description. This includes adjusting the description format and adding physical labels to the product. Here, the large language model refers to a text-to-text model (such as Deepseek), which primarily uses a large language model to generate descriptions in a specified format. The user provides the desired background description plus the physical size of the product, generating a description like "skincare product (h:xx, w:xx) placed on a table," ensuring the description format matches the one used during training. This mainly involves preprocessing the physical size description according to the dataset construction, then transforming it into (h:xx, w:xx) and placing it after the corresponding product name in the description. Finally, the white-background product image and the corresponding prompt are input into the trained final model to obtain an e-commerce scene image with harmonious physical size.
[0041] Please see Figure 6 This application embodiment also provides a training device for generating an e-commerce scene graph model with harmonious product sizes, which can implement the above-mentioned training method. The training device includes: The description text determination module is used to obtain real e-commerce scene image samples and extract real scene image description text based on the real e-commerce scene image samples; the real scene image description text includes product physical size label, product name label and description label; The segmentation map determination module is used to extract real product segmentation maps based on real e-commerce scene image samples; The initial training module is used to take real product segmentation images and real scene image description text as the first input, and real e-commerce scene image samples as the output, and train the preset model based on the first input and output. The adjustment module is used to retain the physical size labels of the main products in the real scene image description text in the first input, and randomly remove some physical size labels of other products to determine the second input. Based on the second input and output, the trained preset model is adjusted to obtain the final model for generating e-commerce scene images with harmonious product sizes.
[0042] It is understood that the content of the above training method embodiments is applicable to this training device embodiment. The specific functions implemented by this training device embodiment are the same as those of the above training method embodiments, and the beneficial effects achieved are also the same as those achieved by the above training method embodiments.
[0043] Please see Figure 7 Another aspect of this application proposes an e-commerce scene graph model prediction device for generating harmonious product sizes, which can implement the above-mentioned prediction method. The prediction device includes: The data acquisition module is used to acquire the segmented white background image of the predicted product, the estimated physical size of the predicted product, and the e-commerce scene image description of the predicted product. Based on the estimated physical size of the predicted product and the e-commerce scene image description of the predicted product, the module generates the scene image description text of the predicted product. The prediction module is used to input the segmented white background image of the predicted product and the scene image description text of the predicted product into the final model of the e-commerce scene image that generates harmonious product size, so as to obtain the e-commerce scene image with harmonious physical size.
[0044] It is understood that the content of the above prediction method embodiments is applicable to the prediction device embodiments. The specific functions implemented by the prediction device embodiments are the same as those of the above prediction method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0045] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the training method or prediction method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0046] It is understood that the content of the above training method or prediction method embodiments is applicable to the embodiments of this device. The specific functions implemented by the embodiments of this device are the same as those of the above training method or prediction method embodiments, and the beneficial effects achieved are also the same as those achieved by the above training method or prediction method embodiments.
[0047] Please see Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the methods described in the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0048] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described training method or prediction method.
[0049] It is understood that the content of the above training method or prediction method embodiments is applicable to this storage medium embodiment. The specific functions implemented by this storage medium embodiment are the same as those of the above training method or prediction method embodiments, and the beneficial effects achieved are also the same as those achieved by the above training method or prediction method embodiments.
[0050] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described training method or prediction method.
[0051] It is understood that the content of the above training method or prediction method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above training method or prediction method embodiments, and the beneficial effects achieved are also the same as those achieved by the above training method or prediction method embodiments.
[0052] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0053] This application provides a training method, prediction method, and related equipment for generating e-commerce scene graph models with harmonious product sizes. The training method involves adding physical size labels to the scene graph description text of real e-commerce scene graph samples. A preset model is initially trained using real product segmentation images and scene graph description text. The physical size labels control the relative size relationship between the main product and other products. Then, the physical size labels of the main product in the real scene graph description text are retained, while some other product physical size labels are randomly removed. The trained preset model is then adjusted again to obtain a final model sensitive to the physical size of the main product, enabling the final model to generate e-commerce scene graphs with harmonious product sizes for unknown products. The prediction method generates scene graph description text for predicted products based on the estimated physical size of the predicted products and the predicted e-commerce scene graph description. The segmentation white background image of the predicted product and the predicted scene graph description text are then input into the aforementioned final model for generating e-commerce scene graphs with harmonious product sizes, resulting in e-commerce scene graphs with harmonious product sizes.
[0054] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0055] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0056] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0057] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0058] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0059] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0060] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0061] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0062] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0063] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0064] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for training a graph model of an e-commerce scene to generate harmonious product sizes, characterized in that, The method includes the following steps: Obtain real e-commerce scene image samples, and extract real scene image description text based on the real e-commerce scene image samples; the real scene image description text includes product physical size label, product name label and description label; Extract real product segmentation images based on the real e-commerce scene image samples; The real product segmentation image and the real scene image description text are used as the first input, and the real e-commerce scene image sample is used as the output. The preset model is trained based on the first input and the output. The physical size labels of the main products in the real scene description text in the first input are retained, and some physical size labels of other products are randomly removed to determine the second input. The trained preset model is adjusted according to the second input and the output to obtain the final model for generating e-commerce scene diagrams with harmonious product sizes.
2. The method according to claim 1, characterized in that, The step of extracting real product segmentation images based on the real e-commerce scene image samples includes: Based on the product name tag, and using a target detection method, the position of the product is extracted from the real e-commerce scene image sample; Based on the location of the product, a real product segmentation image is segmented from the real e-commerce scene image sample.
3. The method according to claim 1, characterized in that, The step of extracting real scene description text based on the real e-commerce scene image sample includes: The real e-commerce scene image samples are input into the multimodal large model to extract the product name tags; The description tags are obtained by describing the real e-commerce scene image sample based on the product name tags as the main description objects; Determine the physical size label of the product based on the actual product segmentation diagram; The actual scene image description text is formed based on the product name label, the description label, and the product physical size label.
4. The method according to claim 3, characterized in that, The step of determining the physical size label of the product based on the actual product segmentation diagram includes: Based on the three-dimensional point cloud data and global scale factor of the predicted scene using the real product segmentation map, the horizontal and vertical endpoint coordinates of the real product in the three-dimensional point cloud data are determined based on the real product segmentation map. The physical size label of the product is determined based on the endpoint coordinates in the horizontal direction, the endpoint coordinates in the vertical direction, and the global scale factor.
5. A method for predicting e-commerce scene graph models to generate harmonious product sizes, characterized in that, include: Obtain the segmented white background image of the predicted product, the estimated physical size of the predicted product, and the e-commerce scene image description of the predicted product. Based on the estimated physical size of the predicted product and the e-commerce scene image description of the predicted product, form the scene image description text of the predicted product. The segmented white background image of the predicted product and the scene image description text of the predicted product are input into the final model for generating an e-commerce scene image with harmonious product size as described in any one of claims 1-4, to obtain an e-commerce scene image with harmonious physical size.
6. A training device for generating e-commerce scene graph models with harmonious product sizes, characterized in that, The device training includes: The description text determination module is used to obtain real e-commerce scene image samples and extract real scene image description text based on the real e-commerce scene image samples; the real scene image description text includes product physical size label, product name label and description label; The segmentation map determination module is used to extract real product segmentation maps based on the real e-commerce scene image samples. The initial training module is used to take the real product segmentation image and the real scene image description text as the first input, and the real e-commerce scene image sample as the output, and train the preset model based on the first input and the output. The adjustment module is used to retain the physical size labels of the main products in the real scene image description text in the first input, randomly remove some physical size labels of other products to determine the second input, and adjust the trained preset model according to the second input and the output to obtain the final model for generating e-commerce scene images with harmonious product sizes.
7. A device for predicting e-commerce scene graph models with harmonious product sizes, characterized in that, The prediction device includes: The data acquisition module is used to acquire the segmented white background image of the predicted product, the estimated physical size of the predicted product, and the e-commerce scene image description of the predicted product, and to form the scene image description text of the predicted product based on the estimated physical size of the predicted product and the e-commerce scene image description of the predicted product. The prediction module is used to input the segmented white background image of the predicted product and the scene image description text of the predicted product into the final model for generating a product-size-harmonious e-commerce scene image as described in any one of claims 1-4, so as to obtain a physically-size-harmonious e-commerce scene image.
8. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1-5.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.