Image generation method and system, electronic equipment and storage medium

Through the image processing model, the style feature information of the target image is analyzed, and the matching of the display objects in the scene is automatically determined, which solves the problem of low image generation efficiency and realizes efficient and personalized image generation.

CN120298527APending Publication Date: 2025-07-11ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510442928.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, image generation efficiency is low, the design process is cumbersome and time-consuming, and lacks automation and intelligent support, making it difficult to meet the needs of large-scale and high-frequency online design.

Method used

By acquiring the target image, analyzing style feature information using the image processing model, determining the spatial position relationship between the display object and the matching object, and rendering the matching image into the target scene image, realizing automated style analysis, rapid retrieval of matching images and optimized image rendering.

Benefits of technology

Generate a large number of stylized, personalized and high-quality matching images in a short period of time, improving user experience and commercial operation efficiency, and promoting the digitalization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298527A_ABST
    Figure CN120298527A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and system, electronic equipment and a storage medium. The method comprises the steps that a target image is acquired, and the image content of the target image comprises at least one display object; the image processing model is utilized to analyze the style feature information of the target image to obtain at least one matching image, the style feature information of the target image is used for representing at least one attribute feature of the display object, and the image content of the matching image comprises a matching object matched with the at least one attribute feature; determining a spatial position relationship between the display object and the matching object; and rendering the matching image and the target image into the target scene image according to the spatial position relationship. The technical problem of low image generation efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology. Specifically, it relates to a method, system, electronic device, and storage medium for generating images. Background Art

[0002] Currently, with the rise of e-commerce and digital lifestyles, more and more users (consumers) choose to purchase and design products on online platforms. For example, the matching of home furnishings (furniture), the matching of clothing and accessories, etc. To improve the user experience and promote product sales, e-commerce platforms have generally introduced matching design tools and services, aiming to assist users in intuitively previewing and designing images of matching results that meet user needs.

[0003] In related technologies, corresponding design tools are provided for professional designers, including various item pictures (such as furniture, clothing) and scene templates (such as home scenes, outdoor scenes, etc.). Designers are allowed to freely select item pictures and adjust the position and layout of the item pictures by dragging in the preset scene templates to create personalized design solutions (such as images of home decoration designs and clothing matching designs that meet user needs). However, the above method relies on the subjective judgment and manual operation of designers. The design process is cumbersome and time-consuming, and it is difficult to meet the large-scale and high-frequency online design requirements. In addition, due to the lack of automation and intelligent support, designers often need to spend a lot of energy to ensure the consistency of style and the rationality of layout, which undoubtedly increases the design cost and cycle and reduces work efficiency. Therefore, there is still a technical problem of low efficiency in generating images.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a method, system, electronic device, and storage medium for generating images to at least solve the technical problem of low efficiency in image generation.

[0006] According to one aspect of the embodiments of this application, a method for generating an image is provided. The method may include: obtaining a target image, where the image content of the target image includes at least one display object; analyzing the style feature information of the target image using an image processing model to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes matching objects that match at least one attribute feature; determining the spatial position relationship between the display object and the matching object; and rendering the matching image and the target image into the target scene image according to the spatial position relationship.

[0007] According to another aspect of the embodiments of the present application, a method for determining an image processing model is provided. The method may include: obtaining an image sample set, where the image sample set includes a target image sample set and a candidate image sample set, the image content of the target image sample includes at least one display object sample, and the image content of the candidate image sample set includes at least one candidate object sample; using the image sample set to perform contrastive learning on a deep learning model to obtain an image processing model, where the image processing model is used to analyze the style feature information of the target image to obtain at least one matching image, the image content of the target image includes at least one display object, the style feature information of the target image is used to represent at least one attribute feature of the display object, the image content of the matching image includes a matching object that matches at least one attribute feature, and the matching image and the target image are used to be rendered to a target scene image based on the spatial position relationship between the display object and the matching object.

[0008] According to another aspect of the embodiments of the present application, another method for generating an image is provided. The method may include: identifying an image of a product object to be matched on a product matching platform, where the image content of the product object image includes at least one product object; using a product matching model to analyze the style feature information of the product object image to obtain at least one matching image, where the style feature information of the product object image is used to represent at least one attribute feature of the product object, and the image content of the matching image includes a matching object that matches at least one attribute feature; determining the spatial position relationship between the product object and the matching object; rendering the matching image and the product object image to a target scene image according to the spatial position relationship; and returning the rendered target scene image to the product matching platform.

[0009] According to another aspect of the embodiments of the present application, another method for generating an image is provided. The method may include: in response to an input operation on an operation interface, displaying a target image on the operation interface, where the image content of the target image includes at least one display object; in response to an image generation instruction on the operation interface, displaying a rendered target scene image on the operation interface, where the rendered target scene image is obtained by rendering the matching image and the target image to the target scene image according to the spatial position relationship between the display object and the matching object, the matching image is obtained by using an image processing model to analyze the style feature information of the target image, the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one attribute feature.

[0010] According to another aspect of the embodiments of the present application, another method for generating an image is provided. The method may include: obtaining a target image by invoking a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the target image, and the image content of the target image includes at least one display object; analyzing the style feature information of the target image by using an image processing model to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes matching objects that match the at least one attribute feature; determining the spatial position relationship between the display object and the matching object; rendering the matching image and the target image into a target scene image according to the spatial position relationship; and outputting the rendered target scene image by invoking a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the rendered target scene image.

[0011] According to another aspect of the embodiments of the present application, an image generation system is provided. The system may include: a client for uploading a target image, where the image content of the target image includes at least one display object; and a server for analyzing the style feature information of the target image by using an image processing model to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes matching objects that match the at least one attribute feature; determining the spatial position relationship between the display object and the matching object; rendering the matching image and the target image into a target scene image according to the spatial position relationship; where the client is used to display the rendered target scene image.

[0012] According to another aspect of the embodiments of the present application, an electronic device is further provided. The electronic device may include a memory and a processor: the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the above computer-executable instructions are executed by the processor, the method for generating an image according to the embodiments of the present application is implemented.

[0013] According to another aspect of the embodiments of the present application, a processor is further provided. The processor is used to run a program, and when the program runs, the method for generating an image according to the embodiments of the present application is executed.

[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, and when the program runs, the device where the storage medium is located is controlled to execute the method for generating an image according to the above embodiments of the present application.

[0015] According to another aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes a computer program which, when executed by a processor, implements the method for generating an image in the above embodiments of the present application.

[0016] In the embodiments of the present application, if it is necessary to match the display object in a certain target image, the target image can be obtained, and the image processing model is used to analyze the style to which the display object in the target image belongs, so as to obtain the corresponding number of matching images that match the style. The spatial position relationship between the display object and the matching object in the matching image can be determined, and thus, according to the spatial position relationship, the corresponding matching image and the target image are rendered into the target scene image including the display scene.

[0017] In the embodiments of the present application, the target image can be analyzed by a pre-trained image processing model to obtain the style of the display object, and it can be automatically determined how the display object is matched in the display scene in this style. For example, the number of matching images in this style and the matching object in them match the display object and conform to the belonging style. That is, by introducing the image processing model for automatic style analysis, rapid retrieval of matching images, automatic determination of spatial position relationship, and optimized image rendering and synthesis technology, it is possible to generate a large number of stylized, personalized, and high-quality matching images in a short time, greatly promoting the digital process of the matching design industry, improving the user experience and business operation efficiency, thus achieving the technical effect of improving the generation efficiency of images and solving the technical problem of low efficiency of image generation.

[0018] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. Description of the Drawings

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0020] Figure 1 is a schematic diagram of an application scenario of a method for generating an image according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of a method for generating an image according to an embodiment of the present application;

[0022] Figure 3 is a flowchart of a method for determining an image processing model according to an embodiment of the present application;

[0023] Figure 4It is a flowchart of another method for generating an image according to an embodiment of the present application;

[0024] Figure 5 It is a flowchart of another method for generating an image according to an embodiment of the present application;

[0025] Figure 6 It is a flowchart of another method for generating an image according to an embodiment of the present application;

[0026] Figure 7 It is a schematic diagram of an image generation system according to an embodiment of the present application;

[0027] Figure 8 It is a schematic diagram of a home improvement intelligent design tool in a related art;

[0028] Figure 9 It is a schematic diagram of a whole-house design platform in another related art;

[0029] Figure 10(a) is a schematic diagram of a target furniture according to an embodiment of the present application;

[0030] Figure 10(b) is a schematic diagram of a matching furniture that is consistent with the style of the target furniture according to an embodiment of the present application;

[0031] Figure 10(c) is a schematic diagram of a non-matching furniture that is inconsistent with the style of the target furniture according to an embodiment of the present application;

[0032] Figure 11 It is a schematic diagram of a system for determining whether the styles of furniture match according to an embodiment of the present application;

[0033] Figure 12(a) is a schematic diagram of a scene picture according to an embodiment of the present application;

[0034] Figure 12(b) is a schematic diagram of a white-background picture corresponding to the scene picture according to an embodiment of the present application;

[0035] Figure 13 It is a schematic diagram of a method for constructing a training data set according to an embodiment of the present application;

[0036] Figure 14(a) is a schematic diagram of an example of a new Chinese style seed furniture according to an embodiment of the present application;

[0037] Figure 14(b) is a schematic diagram of another example of a new Chinese style seed furniture according to an embodiment of the present application;

[0038] Figure 14(c) is a schematic diagram of another example of a new Chinese style seed furniture according to an embodiment of the present application;

[0039] FIG. 15(a) is a schematic diagram of an example of modern-style seed furniture according to an embodiment of the present application;

[0040] FIG. 15(b) is a schematic diagram of another example of modern-style seed furniture according to an embodiment of the present application;

[0041] FIG. 15(c) is a schematic diagram of another example of modern-style seed furniture according to an embodiment of the present application;

[0042] FIG. 16(a) is a schematic diagram of a main furniture retrieval method according to an embodiment of the present application;

[0043] FIG. 16(b) is a schematic diagram of a random furniture selection retrieval method according to an embodiment of the present application;

[0044] Figure 17 is a schematic diagram of rendering a set of rendering images into a sample room according to an embodiment of the present application;

[0045] Figure 18 is a schematic diagram of obtaining a matching result by superimposing a set of rendering images and a sample room according to an embodiment of the present application;

[0046] Figure 19 is a schematic diagram of a device for generating an image according to an embodiment of the present application;

[0047] Figure 20 is a schematic diagram of a device for determining an image processing model according to an embodiment of the present application;

[0048] Figure 21 is a schematic diagram of another device for generating an image according to an embodiment of the present application;

[0049] Figure 22 is a schematic diagram of another device for generating an image according to an embodiment of the present application;

[0050] Figure 23 is a schematic diagram of another device for generating an image according to an embodiment of the present application;

[0051] Figure 24 is a block diagram of the structure of a computer terminal according to an embodiment of the present application;

[0052] Figure 25 is a block diagram of an electronic device for an image generation method according to an embodiment of the present application;

[0053] Figure 26 is a hardware block diagram of a computer terminal (or mobile device) for implementing an image generation method according to an embodiment of the present application;

[0054] Figure 27It is a structural block diagram of a computing environment for an image generation method according to an embodiment of the present application. Detailed implementation manners

[0055] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0056] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to the above process, method, product or device.

[0057] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0058] Contrastive Learning, a self-supervised learning method that learns data feature representations by constructing positive and negative sample pairs, emphasizing the aggregation of similar samples and the separation of dissimilar samples;

[0059] White background image, a picture of a furniture product taken against a pure white background, used to eliminate background interference and highlight the main features of the product;

[0060] Triplet data, a data structure containing the target furniture, matching furniture (positive sample), and non-matching furniture (negative sample), used for contrastive learning training;

[0061] Embedding Space / vector, mapping high-dimensional data to a low-dimensional continuous vector space through a deep learning model, used to measure semantic similarity;

[0062] Three-Dimensional (abbreviated as 3D) show flat, a three-dimensional model of a virtual indoor scene, used for immersive visualization of furniture matching.

[0063] According to an embodiment of the present application, a method for generating an image is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0064] The method for generating an image provided by the embodiments of the present application can be applied to Figure 1 the application scenarios shown, but is not limited thereto. In the application scenarios shown Figure 1 as such, the server 10 can be a cloud. The server 10 can be connected to the database 40. An image processing model can be deployed in the database 40, and candidate images can also be stored. Thus, when performing the method for generating an image of the present application, the image processing model can be called from the database 40 to process the target image, and a matching image can be determined from the candidate images. The server 10 can be connected to one or more client devices (clients) 20 through a network 30, such as, for example, a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client devices 20 here can include, but are not limited to: smart phones, tablet computers, laptop computers, handheld computers, personal computers, smart home devices, vehicle-mounted devices, etc. The client devices can also be referred to as terminal devices and client devices. The client devices 20 together constitute the client opposite to the server 10. An operation interface for the user to operate can be deployed on the graphical user interface of the client device 20. The client device 20 can interact with the user through the graphical user interface to implement the method for generating an image provided by the embodiments of the present application.

[0065] In the embodiments of the present application, the system composed of the playback client 20 and the server 10 can execute the following steps: If it is necessary to process the target image to be processed, the user can perform corresponding operations on the operation interface of the client 20 and input the corresponding target image. The above target request can be transmitted to the server 10 through the network 30.

[0066] Specifically, the following steps can be executed in the server 10:

[0067] Step S102, obtain the target image; Step S104, analyze the style feature information of the target image by using the image processing model to obtain at least one matching image; Step S106, determine the spatial position relationship between the display object and the matching object; Step S108, render the matching image and the target image into the target scene image according to the spatial position relationship.

[0068] In the above process, the rendered target scene image obtained in server 10 can be sent to client 20 through network 30. The target scene image can be loaded and displayed on client 20.

[0069] In this embodiment, the target image can be analyzed by a pre-trained image processing model to obtain the style of the display object, and how to match the display object in the display scene in this style can be automatically determined. For example, the number of matching images in this style and the matching objects therein match the display object and conform to the belonging style. That is, by introducing an image processing model for automated style analysis and matching strategy generation, fast retrieval of matching images, automatic determination of spatial position relationships, and optimized image rendering and synthesis technologies, this application makes it possible to generate a large number of stylized, personalized, and high-quality matching images in a short time, greatly promoting the digital process of the matching design industry, improving the user experience and business operation efficiency, thereby achieving the technical effect of improving the generation efficiency of images and solving the technical problem of low efficiency in image generation.

[0070] The embodiments of this application propose the following method. In the above application scenario, this application provides, from the server side, an Figure 2 image generation method as shown. Figure 2 It is a flowchart of an image generation method according to an embodiment of this application. As Figure 2 shown, it may include the following steps:

[0071] Step S202, obtain a target image.

[0072] In the technical solution provided in step S202 of this application above, the image content of the target image includes at least one display object. The display object can be an item. That is, the target image can be an image including the item to be matched. For the application scenario of an e-commerce platform, the item to be matched can be a commodity, such as a home commodity (furniture), clothing, etc. The display object can also be referred to as an anchor. The display object can be the target furniture. It should be noted that the display object is not limited to a single type and can be a mixture of multiple home elements. For example, a sofa, a coffee table, a lamp, a decorative painting, etc. The target image can be an original image uploaded by the user according to their own matching design needs; it can also be an intermediate product obtained by processing the above original image (such as extracting a white-background image containing the display object and deleting the remaining objects other than the display object) to ensure the accuracy of the matching design; it can also be an image to be matched and designed selected from a pre-set image set according to the user's needs.

[0073] It should be noted that the above-mentioned target images and their acquisition methods are only for illustrative purposes and are not specifically limited here. As long as it is a method capable of acquiring a target image containing the display object to be collocated, and the target image containing the above-mentioned display object, they are all within the protection scope of the embodiments of the present application, and no further examples will be given here.

[0074] In this embodiment, if it is necessary to collocate a certain display object, a target image containing the display object can be acquired.

[0075] For example, when a user browses an e-commerce platform and is interested in a certain home furnishing product (such as a sofa) and hopes to know other furniture that matches it. At this time, the user can upload a picture containing the furniture, which can be either the product display picture on the product detail page, or the real-life photo of the user's own home, or even the inspiration picture found from channels such as social media and design websites. After the server receives the above picture, it will identify the display object therein and start the subsequent image generation process with this as the anchor point.

[0076] For another example, on the merchant side, in order to better display products or improve the conversion rate of products, merchants can provide some scene pictures containing multiple display objects as target images. For example, a furniture brand may include a picture of a fully furnished living room in its product catalog, and the multiple display objects therein can include sofas, coffee tables, TV cabinets, etc. After the merchant uploads such pictures, multiple display objects can be identified, and collocation suggestions can be generated for each piece of furniture, thereby enriching the scene display on the product detail page and enhancing the shopping experience.

[0077] As an optional example, the e-commerce platform can pre-prepare a series of home furnishing scene template pictures with different styles for generating collocation suggestions. The user selects a certain style template and identifies the display objects in the above style template. For example, a modern-style living room design contains a minimalist-designed sofa. Based on the display objects in this style template, collocation recommendations that match this style can be generated, such as the recommended coffee tables, lamps, and decorations. The above method is particularly applicable to the situation where the user has not yet clearly defined a specific product but has a preference for a certain style or scene.

[0078] It should be noted that the above methods and processes for obtaining the target image are only for illustrative purposes and are not specifically limited here. Any process that can be used to obtain the target image including the display object to be matched is within the protection scope of the embodiments of the present application. The methods for obtaining the target image in the above various examples all have their application scenarios and advantages, jointly constituting a flexible and variable matching recommendation method, which can meet the needs of different users in different situations. The above diverse input methods not only improve the user experience but also expand the scope of system applications, promoting the intelligent and scenario-based expression of indoor furniture matching. During the implementation of the present application, multiple methods can be combined to intelligently select the appropriate way to obtain the target image according to user behavior and platform strategies.

[0079] Step S204: Analyze the style feature information of the target image by using an image processing model to obtain at least one matching image.

[0080] In the technical solution provided in step S204 of the present application, the style feature information of the target image is used to represent at least one attribute feature (style) of the corresponding display object. For example, features such as color, material, and size. The image content of the matching image includes matching objects that match at least one attribute feature. The image processing model can also be referred to as a matching model. If it is for an indoor furniture matching scenario, the image processing model can be an indoor furniture matching model. This is only for illustrative purposes and does not specifically limit the image processing models for different scenarios. Different styles correspond to different matching strategies. The matching strategy can also be referred to as a matching method, which can include rules for color, material, size, layout, and the number of matching objects to be arranged (target quantity) under the style to which the display object belongs. The target quantity can be determined and analyzed through attribute features, that is, based on the attribute features of the display object in the target image (such as the color, material, size, etc. of the display object in the target image), the corresponding matching objects and the number of matching objects can be matched for the display object. For a home matching scenario, the matching objects can be matching furniture, also known as complementary furniture. For example, if the display object is a sofa, the matching object can be a complementary coffee table in style. The image processing model can be obtained by performing contrastive learning training on a deep learning model.

[0081] In this embodiment, after obtaining the target image to be processed, the image processing model can be used to analyze the target image to obtain the corresponding number of matching images for the style to which the display object in the target image belongs.

[0082] Optionally, based on the style feature information of the target image, an image processing model is used to generate a series of matching images in terms of style. For the indoor furniture matching scenario, the above process is particularly complex because it involves the understanding of design intentions and the comprehensive application of color, material, size, layout, and quantity rules.

[0083] Optionally, the image processing model can perform in-depth analysis on the target image to extract the style feature information of the displayed object. The style feature information can include, but is not limited to, color distribution, material texture, design lines, decorative elements, etc. The above style feature information jointly depicts the style of the displayed object in the target image. For example, Nordic style, industrial style, or Oriental classical style, etc.

[0084] Optionally, based on the extracted style feature information, corresponding matching strategies can be applied. The matching strategy is a set of rules preset for a specific style, aiming to ensure that the matched image is consistent with the target image style in terms of color, material, size, layout, etc. For example, for a displayed object with an obvious industrial style, the matching strategy may tend to use matching objects with metal and dark-toned materials, as well as the morphological design of simple straight lines. The matching strategy also stipulates the target number of matching objects to be arranged. For example, in the design of a small living room, only one coffee table and two single chairs may be needed as matching objects, while in a larger space, more furniture may be needed to complete the scene.

[0085] Optionally, the image processing model will retrieve the matching object images with the same style in the database according to the style feature information of the target image. For example, a contrast learning model can be used to calculate the vector similarity between the target image and the candidate matching object image in the embedding space.

[0086] Optionally, based on the detailed rules in the matching strategy, the retrieved matching object images will be further screened to retain those images that meet the requirements in terms of color, material, size, and layout.

[0087] Optionally, this embodiment uses the image processing model to perform style analysis on the target image, and formulates and applies a matching strategy according to the analysis result of the style to generate a matching image that matches the style.

[0088] Optionally, the images in the embodiments of the present application used to analyze the style in the target image can include identifying the color, material, design elements, layout method, etc. of the displayed object in the target image, so as to determine the style (decoration style, clothing style) to which the displayed object belongs. If the displayed object is furniture, the corresponding style can be modern minimalist, French retro, new Chinese, etc.

[0089] Optionally, once the style of the display object is recognized, a matching strategy can be formulated according to this style. The matching strategy can be a set of rules for guiding how to select other items that match the display object. The above rules can cover multiple dimensions such as color coordination, material complementarity, size ratio, spatial layout, and quantity suggestions. For example, for a modern minimalist style sofa, the matching strategy may include choosing a coffee table with equally simple lines, maintaining color unity, and avoiding overly complex or obtrusive decorations, etc.

[0090] Optionally, according to the matching strategy formulated above, suitable matching objects can be screened out from a pre-established database, that is, other items (such as furniture, decorations, clothing, etc.) that match the style of the display object in the target image and can form an aesthetic match.

[0091] For example, the image processing model can compare the key styles of the target image, such as color distribution, material texture, design lines, etc., with the styles in the images in the database, identify the relatively close style categories, and use the images corresponding to these styles as the matching images for the target image. The matching strategy can include color matching rules, such as using the color wheel theory or the principles of design color psychology, to ensure that the colors of the matching objects are coordinated with the display object, avoid color conflicts, and enhance the harmony of the overall visual effect. It can also include material matching rules, which can suggest using materials of different textures to increase the sense of hierarchy, or choosing similar materials to maintain style consistency. It can also include size ratio matching rules, where the size of the matching object needs to match the display object, following certain proportion principles, and avoiding visual incoordination in space due to being too large or too small. It can also include spatial layout rules, that is, considering the layout of the items to ensure that the display object and the matching object are reasonably distributed in space, forming a comfortable and practical layout plan. It can also include quantity suggestion rules, that is, the quantity of the matching object is also part of the matching strategy, and the quantity of the matching object can be suggested according to the size of the space, the type of the display object, and personal preferences, avoiding being too many or too few.

[0092] It should be noted that the rules included in the above matching strategy are only for illustrative purposes and are not specifically limited herein. As long as the matching strategy can determine the matching objects that match the display object, it is within the protection scope of the embodiments of this application.

[0093] In summary, in the embodiments of the present application, determining the matching image through the image processing model is a highly automated and deep learning-based matching recommendation process. It not only considers the characteristics of the display object itself but also comprehensively takes into account various factors such as style features, colors, materials, sizes, and layouts, providing users with matching suggestions that not only meet personal preferences but also have a professional design sense. Through the above steps, the shopping experience of users on the e-commerce platform can be significantly improved, promoting more efficient product selection and matching decisions.

[0094] Step S206: Determine the spatial position relationship between the display object and the matching object.

[0095] In the technical solution provided in step S206 of the present application, the spatial position relationship refers to the relative position, direction, and distance between the display object and the matching object in a three-dimensional space. In the field of interior design, it includes but is not limited to: position, that is, the specific coordinates of the display object and the matching object in the room, such as the position of the sofa and the coffee table, which needs to consider the layout and size of the room; direction, that is, the placement direction of the display object and the matching object to ensure their coordination with each other and with the room structure, such as the orientation of the bed and the hanging angle of the painting; distance, that is, the physical distance between the display object and the matching object, which can ensure the reasonable use of space and visual comfort, avoiding furniture being too crowded or too scattered; hierarchical relationship, that is, the occlusion relationship between the display object and other matching objects visually, ensuring the prominence of important objects and the sense of hierarchy in space.

[0096] In this embodiment, after determining the matching image corresponding to the target image using the image processing model, the spatial position relationship between the display object and the matching object can be determined.

[0097] Optionally, if the target scene image corresponding to the display object and the display scene to which the matching display object needs to be displayed is determined, a deep analysis can be performed on the target scene image to understand its spatial layout and potential rules of furniture placement. Using image processing technology, accurately locate the positions of the display object and the matching object in the target image in the 3D space of the above display scene.

[0098] For example, for the home matching scene, the general rules and common sense in the fields of interior design and furniture matching can be applied, such as "the sofa is usually placed against the wall and the coffee table is placed directly in front of the sofa", "the bedside table should be placed on both sides of the bed", and the layout guiding principles considering the passage space and usability, to determine the spatial position relationship between the matching object and the display object.

[0099] It should be noted that the determination of the spatial position relationship can take into account conditions such as the style to which the display object belongs, the collocation aesthetics, and the spatial limitations of the actual display scene. The specific determination process and limiting conditions of the spatial position relationship are not restricted here. As long as the collocation principle and deployment principle of the style to which the display object belongs can be satisfied, and the spatial position relationship with beautiful collocation can be achieved, it is within the protection scope of the embodiments of the present application.

[0100] In the embodiments of the present application, the display object and the collocation object are integrated into a 3D display scene to ensure that the position, direction, and distance relationships between the display object and the collocation object meet the aesthetic and functional standards of interior design. The above process not only requires a profound understanding of the indoor space but also relies on advanced image processing and 3D rendering technologies, which is the core step to realize the scene-based furniture collocation and display.

[0101] Step S208: Render the collocation image and the target image into the target scene image according to the spatial position relationship.

[0102] In the technical solution provided in step S208 of the present application above, the image content of the target scene image includes the display scene. The display scene can be a 3D model apartment. The rendered target scene image can be the finally completed scene collocation picture.

[0103] In this embodiment, after determining the spatial position relationship between the display object and the collocation object, the collocation image and the target image can be rendered into the target scene image according to the spatial position relationship.

[0104] Optionally, the target image, the collocation image, and their spatial position relationship in the target scene image are integrated together, and the final target scene image is generated through 3D rendering technology. The key to the above steps lies in how to transform the two-dimensional collocation image and target image, as well as according to the determined spatial position relationship, into a visual representation in a three-dimensional scene, so as to create a real and beautiful scene collocation picture.

[0105] For example, mapping the display object and the collocation object in the two-dimensional space to the three-dimensional space is the key step to generate the target scene image. The above process involves 3D modeling of the objects in the two-dimensional image. That is to say, each piece of furniture or decoration in the target image and the collocation image can be transformed into a three-dimensional model for positioning and rendering in the 3D scene. For planar decorations such as hanging pictures and carpets, although complete 3D modeling is not required, perspective transformation can be performed to ensure that the above planar decorations look natural and fitting in the display scene.

[0106] Optionally, a two-dimensional image (such as a matching image, a target image) can be converted into a three-dimensional view, and perspective transformation and projection processing can be performed on the three-dimensional image. Perspective transformation ensures that the position, size, and angle of objects in the three-dimensional image are correctly reflected in the three-dimensional space, while projection projects the above 3D model into the scene to ensure that the 3D model is coordinated with the hard decoration background, other objects, and lighting conditions in the scene. For flat objects such as carpets and hanging pictures, accurate perspective projection of the flat objects is particularly required to make the flat objects look as if they are naturally placed in the scene.

[0107] For example, to make the scene matching picture more realistic, the lighting conditions in the scene and the material parameters of the objects can also be adjusted. Lighting not only affects the visibility and color performance of objects, but is also the key to creating atmosphere and style. Material adjustment ensures that the surface texture and texture of the 3D model match the original visual effects in the target image and the matching image, thereby enhancing the realism of the scene.

[0108] Optionally, after positioning, perspective transformation, and material adjustment of each object (display object, matching object) are completed, the above elements can be synthesized into an initial scene image using 3D rendering technology. The rendering process can include the overlay of multiple layers. First, the hard decoration background is rendered, then the 3D models of furniture and decorations are overlaid, and finally the overall light and shadow effect is adjusted. The goal of synthesis is to create a coherent, beautiful, and scene matching image that conforms to the design style.

[0109] In the embodiment of the present application, this embodiment realizes the transformation from abstract item matching suggestions to specific and visual scene matching pictures, provides an intuitive decision-making basis for users, and also provides an effective tool for e-commerce platforms to enhance the user shopping experience and promote commodity sales. The above process not only requires high-precision image processing and 3D rendering technology, but also reflects a deep understanding of the matching design principle and user needs.

[0110] Through the above steps S202 to S208 of this application, if it is necessary to match the display object in a certain target image, the target image can be obtained, and the image processing model can be used to analyze the style of the display object in the target image, and the target number of matching images that match the style can be obtained. The spatial position relationship between the display object and the matching object in the matching image can be determined, and then the corresponding matching image and the target image can be rendered into the target scene image including the display scene according to the spatial position relationship. In the embodiment of this application, by introducing an image processing model for automated style analysis, rapid retrieval of matching images, automatic determination of spatial position relationships, and optimized image rendering and synthesis technologies, it is possible to generate a large number of stylized, personalized, and high-quality matching images in a short time, greatly promoting the digital process of the matching design industry, enhancing the user experience and business operation efficiency, thus achieving the technical effect of improving the generation efficiency of images and solving the technical problem of low efficiency in image generation.

[0111] The above method of this embodiment will be further introduced below.

[0112] As an optional implementation manner, in step S204, using the image processing model to analyze the style feature information of the target image to obtain at least one matching image includes: using the image processing model to analyze the style feature information of the target image, and determining, from multiple candidate images, the target candidate images that meet the style feature information of the target image, where the image content of the candidate images includes candidate objects; using the image processing model, determining the target candidate images whose attribute features match the attribute features of the display object as the matching images, where the matching object is the candidate object whose attribute features match the attribute features of the display object.

[0113] In this embodiment, in the process of using the image processing model to analyze the target image to obtain the target number of matching images that match the style feature information of the target image, the image processing model can be used to analyze the target image, so as to determine, from the candidate images, the target candidate images that meet the style feature information. The target candidate images whose attribute features match the attribute features of the display object can be determined as the matching images by using the image processing model, so as to obtain the target number of matching images. Wherein, the image content of the candidate images includes candidate objects. The candidate objects can be alternative furniture. The matching object can be the candidate object whose attribute features match the attribute features of the display object.

[0114] Optionally, use the image processing model to perform in-depth analysis on the target image, and then screen the target images that meet the specific style feature information from the candidate images to generate matching images that are consistent with the original design in style and attributes.

[0115] Optionally, the image processing model comprehensively analyzes the target image to identify the style feature information of the displayed object. For example, style feature information that can represent this style, such as color combination, material texture, design lines, decorative elements, etc. When analyzing the target image, the image processing model can also extract the attribute features of the displayed object, such as size, shape, function category, etc. The above information is crucial for subsequent attribute feature matching.

[0116] Optionally, the image processing model can screen out target candidate images that are consistent with the style of the target image from the candidate image pool in the database based on the matching degree between style feature information. For example, the similarity between image embedding vectors can be calculated to determine whether images belong to the same style.

[0117] Optionally, for the target candidate images initially screened out, the attribute features (such as color, size, shape, etc.) of the candidate objects in the image content can be further checked and compared with the attribute features of the displayed object in the target image. Only when the attribute features of the candidate object match the attribute features of the displayed object, the candidate image may become a matching image.

[0118] Optionally, in the process of determining the target number of matching images, according to the matching strategy or user requirements, determining the number of matching images to be generated directly affects the richness of the scene and the comprehensiveness of the final matching recommendation. After completing the matching of style feature information and the screening of attribute features, sufficient images will be selected from the target candidate images to meet the requirements of the predetermined target number. The objects in the selected matching images, that is, the matching objects, will be furniture or decorations that match the displayed object in multiple attributes such as style, color, size, and function.

[0119] For example, if the displayed object in the target image is a blue-toned sofa, the image processing model will preferentially select a coffee table or a decorative painting with a blue tone as the target candidate image. In a target image scene with a small space, the image processing model can select matching objects with a moderate size that do not take up too much space, such as small lamps or wall shelves, instead of large bookcases or dining tables. If the displayed object in the target image is a desk, the matching strategy may tend to select a bookshelf or a table lamp as the target candidate image, rather than a sofa or a dining table that is not related to the function of the desk. Ensure that the matching objects belong to the same scene type. For example, the bed in the bedroom should not be matched with the tableware in the kitchen. The above situations require the image processing model to accurately identify scene features to ensure the logic and rationality of the matching.

[0120] It should be noted that the above process and method for determining a matching image from candidate images using an image processing model, as well as the attribute features, are only for illustrative purposes and are not specifically limited here. Any process and method that can ensure high-quality matching objects for the display object are within the protection scope of the embodiments of the present application.

[0121] In the embodiments of the present application, based on deep learning and image processing technologies, matching objects that match the display object in style and attributes are accurately selected from a large number of candidate images, and then matching images that meet the user's expectations and personal preferences are generated. The above process not only improves the accuracy and personalization of matching recommendations, but also greatly simplifies the process of furniture selection and matching by users on e-commerce platforms, enhancing the overall shopping experience. Through the above method, a set of matching images that highly match the target image in style and attributes can be intelligently generated, which not only strengthens the intelligence of indoor furniture matching, but also significantly improves the user experience and shopping satisfaction. The above process makes full use of the advantages of the image processing model and realizes the automatic conversion from design to physical matching suggestions, which is of great significance for promoting the upgrade of the furniture shopping experience in the e-commerce scenario.

[0122] As an optional implementation manner, the style feature information of the target image is analyzed using an image processing model, and a target candidate image that meets the style feature information of the target image is determined from multiple candidate images, including: using the image processing model to determine the similarity between the style feature information of the target image and the style feature information of the candidate image, where the style feature information of the candidate image is used to represent at least one attribute feature of the candidate object; and determining the candidate image with a similarity greater than the similarity threshold as the target candidate image.

[0123] In this embodiment, in the process of analyzing the target image using the image processing model and determining a target candidate image that meets the style feature information from the candidate images, the image processing model can be used to determine the similarity between the style feature information of the target image and the style feature information of the candidate image. The candidate image with a similarity greater than the similarity threshold can be determined as the target candidate image. The style feature information of the target image is used to represent the style to which the corresponding display object belongs. The style feature information of the candidate image is used to represent at least one attribute feature of the candidate object, that is, the style to which the candidate object belongs.

[0124] Optionally, this embodiment elaborates on how to perform style matching on the target image and candidate images using an image processing model.

[0125] Optionally, during the extraction of style feature information, an image processing model (which can be a pre-trained deep learning model, such as a contrastive learning-based model) will perform in-depth analysis on the target image to extract its style feature information. The above style feature information can include color tone, design elements, material texture, line style, etc. The above style feature information together constitutes a unique style identifier of the target image. The same process can also be applied to candidate images to obtain the style feature information of candidate objects.

[0126] Optionally, the image processing model can calculate the similarity between the style feature information of the target image and the style feature information of each candidate image. The above process can be carried out in the embedding space, and the proximity of the styles of the two images can be quantified by comparing the distance or angle between two vectors (for example, using cosine similarity). The calculation result of the similarity reflects the matching degree of the target image and the candidate image in terms of style.

[0127] Optionally, in order to ensure the accuracy of the matching, a similarity threshold can be set. The above similarity threshold can be dynamically adjusted according to the specific requirements and context of the application (such as the strictness of furniture matching, the requirement for style consistency, etc.). The level of the threshold directly affects the screening criteria for target candidate images. Too low a threshold may lead to a mixed style, while too high a threshold may reduce the number of matching results and affect the diversity and richness of the matching.

[0128] Optionally, after calculating the style similarity between the candidate image and the target image, the candidate images with similarity lower than the similarity threshold can be filtered out. Only the candidate images with similarity higher than or equal to the similarity threshold can be determined as target candidate images. This means that the target candidate images have a high degree of fit with the displayed object in the target image in terms of style and are potential matching objects.

[0129] Optionally, once the target candidate images are preliminarily screened, the attribute features of the target candidate images in terms of color, material, size, etc. can be further considered to determine the final matching images. The above process may involve additional image recognition technologies for more detailed analysis and matching of attribute features to ensure that the matching objects are not only consistent in style but also complementary and coordinated in details.

[0130] In the embodiments of the present application, through the above steps, it is possible to effectively screen out target candidate images that match the target image in style from a large number of candidate images, and then generate matching images with consistent style and coordinated details, providing personalized and scenario-based furniture matching suggestions for users. The above process not only demonstrates the powerful ability of the deep learning model in image style recognition but also reflects the efficiency and accuracy of the image processing model in intelligent matching and automated decision-making.

[0131] As an alternative implementation, the image processing model includes a feature comparison model. The style feature information of the target image includes a first style feature vector, and the first style feature vector is used to represent the style semantic features of the corresponding display object in the embedding space. Using the image processing model to determine the similarity between the style feature information of the target image and the style feature information of the candidate image includes: obtaining a second style feature vector of the candidate image, where the second style feature vector is used to represent the style semantic features of the corresponding candidate object in the embedding space; using the feature comparison model to determine the similarity between the first style feature vector and the second style feature vector.

[0132] In this embodiment, in the process of using the image processing model to determine the similarity between the style feature information of the target image and the style feature information of the candidate image, the second style feature vector of the candidate image can be obtained. Using the feature comparison model in the image processing model, determine the similarity between the first style feature vector and the second style feature vector. Among them, the image processing model may include a feature comparison model. The feature comparison model may be a cosine distance determination module. The style feature information of the target image may include a first style feature vector. The first style feature vector can be used to represent the style semantic features of the corresponding display object in the embedding space, and can also be called a matching vector. The second style feature vector can be used to represent the style semantic features of the corresponding candidate object in the embedding space, and can be a vector of alternative furniture.

[0133] Optionally, this embodiment illustrates how the image processing model accurately determines the style similarity between the target image and the candidate image through a feature comparison model (for example, a cosine distance determination module).

[0134] Optionally, the image processing model performs deep feature extraction on the target image and the candidate image, and converts the style feature information of the image into vector representations, that is, the first style feature vector and the second style feature vector. The above vectors are multi-dimensional numerical representations in the Embedding space and can capture the style semantic features of the image. The style feature information of the target image is encoded as the first style feature vector, and the style feature information of each candidate image is encoded as the corresponding second style feature vector. In this way, each image is mapped to a point in the embedding space, representing the mathematical abstraction of its style features.

[0135] Optionally, once the style feature vectors of the target image and the candidate image are extracted, a feature comparison model (such as a cosine distance determination module) is used to calculate the similarity between the first style feature vector and the second style feature vector. The cosine distance determines the similarity in direction between two vectors by calculating the cosine similarity between them. Specifically, the dot product of the first style feature vector and the second style feature vector can be calculated and divided by the product of their magnitudes. The value range of the cosine similarity is [-1, 1]. The closer the value is to 1, the more similar the directions of the two vectors are, that is, the closer the style features of the two images are. If the value is close to 0, it means that the directions of the two vectors are almost orthogonal, that is, the image styles are quite different.

[0136] Optionally, the similarity determined by the feature comparison model can be used as a basis for evaluating whether a candidate image is suitable as a match for the target object. A candidate image whose similarity reaches or exceeds a certain similarity threshold is considered to be stylistically matched with the target image and is thus determined as the target candidate image. The above process not only relies on the powerful capabilities of the image processing model but also on the accurate calculation of vector similarity by the cosine distance determination module. Through the above feature comparison, the consistency and coordination in style during the matching process can be ensured, and highly personalized and scene-based furniture matching suggestions can be generated for users.

[0137] In the embodiments of the present application, through the above steps, the style similarity between the target image and the candidate image can be accurately calculated in the embedding space by using a feature comparison model (such as a cosine distance determination module), thereby achieving a consistent match in style. In the above process, the application of deep learning technology plays a core role and provides technical support for automated and intelligent furniture matching.

[0138] As an alternative implementation, the attribute features of the display object at least include one of the following: the color attribute, size attribute, and function attribute of the display object. Using the image processing model, the target candidate image whose attribute features match those of the display object is determined as the matching image, including: using the image processing model, determining as the matching image the target candidate image corresponding to the candidate object whose similarity between the color attribute and the color attribute of the display object is greater than the similarity threshold, and / or the size attribute matches the size attribute of the display object, and / or the function attribute matches the function attribute of the display object.

[0139] In this embodiment, in the process of using an image processing model to determine a target candidate image with the attribute features of a target quantity matching the attribute features of a display object as a matching image of the target quantity, the image processing model can be used to determine a target candidate image corresponding to a candidate object whose similarity between the color attribute and the color attribute of the display object is greater than a similarity threshold, and / or whose size attribute matches the size attribute of the display object, and / or whose function attribute matches the function attribute of the display object as a matching image. Among them, the attribute features of the display object at least include one of the following: the color attribute, the size attribute, and the function attribute of the display object.

[0140] Optionally, this embodiment elaborates on how to use an image processing model to accurately locate and screen target candidate images based on the color attribute, size attribute, and function attribute of the display object, and then determine the matching image.

[0141] Optionally, the image processing model can perform in-depth analysis on the display object in the target image and extract its attribute features, including color attribute, size attribute, and function attribute. The color attribute may involve the main tone, hue, saturation, etc. of the color; the size attribute includes the length, width, and height of the display object, as well as the relationship with the spatial proportion; the function attribute refers to the purpose of use of the display object, such as a sofa, a coffee table, etc.

[0142] Optionally, for a candidate image, the image processing model also needs to extract the attribute features of the candidate object therein and compare them with the attribute features of the display object in the target image.

[0143] For example, using mathematical calculations in the color space to determine the similarity between the color attribute of the candidate object and the color attribute of the display object. If the similarity exceeds a preset similarity threshold, the candidate object is regarded as a potential matching object in terms of color. Check whether the size attribute of the candidate object matches the size attribute of the display object. This may include direct size comparison (such as numerical comparison of length, width, and height), and indirect matching based on proportion and spatial layout. Confirm whether the function attribute of the candidate object conforms to the function attribute of the display object. For example, if the display object is a dining table, then the candidate objects that match it may include dining chairs, sideboards, etc.

[0144] Optionally, according to the matching results of the above attribute features, select images that meet at least one attribute matching condition from the candidate images. Candidate objects that meet the attribute matching conditions (color, size, and function) will be regarded as suitable matching objects. If only some conditions are met, weight adjustment will be performed according to the application scenario and user requirements to determine the final matching image.

[0145] It should be noted that the above process can achieve flexibility and diversity by setting different attribute matching weights and similarity thresholds. For example, in scenarios where high style consistency is pursued, the weight of color similarity may be increased, while the matching requirements for size and functional attributes may be relatively loose.

[0146] For example, to accurately calculate color similarity, the extracted color attributes can be converted to a standard color space because the color space is not intuitively suitable for color similarity calculation. When matching size attributes, the extracted size information may need to be standardized in a certain proportion to eliminate the influence brought by different image resolutions and perspective differences. For the matching of functional attributes, natural language processing techniques may be needed to understand the text information describing the candidate object to ensure that its function is coordinated with the displayed object. For the case of multiple attribute features, a comprehensive evaluation mechanism can be designed to integrate the matching results of color, size, and functional attributes, and determine the matching paired images through a comprehensive score or decision tree. To adapt to different application scenarios and user preferences, the similarity threshold and attribute matching weight can be dynamically adjusted to generate paired images that meet the current requirements.

[0147] In the embodiment of the present application, through the above steps, the information of the displayed object and the candidate object in terms of color, size, and functional attributes can be accurately identified and matched by using an image processing model, so as to intelligently screen and determine suitable paired images. The above process not only relies on deep learning and image processing technologies, but also reflects the intelligent and automated capabilities of the image processing model in attribute matching and decision-making.

[0148] As an optional implementation manner, in step S204, an image processing model is used to analyze the style feature information of the target image to obtain at least one paired image, including: obtaining foreground information from the target image, where the foreground information includes the displayed object; using the image processing model to analyze the foreground information to obtain the style feature information of the target image; using the image processing model to determine the paired image based on the style feature information of the target image.

[0149] In this embodiment, in the process of using the image processing model to analyze the target image to obtain the target number of paired images that match the style feature information of the target image, the foreground information containing the displayed object can be obtained from the target image. The image processing model can be used to analyze the foreground information to obtain the style feature information. The image processing model is used to determine the target number of paired images based on the style feature information of the target image. Among them, the foreground information can also be called a white background image.

[0150] Optionally, this embodiment outlines the process of using an image processing model to perform multi-level analysis on a target image to determine a matching collocation image for the style features of the target image.

[0151] Optionally, a clear view of the display object, i.e., the so-called foreground information, is extracted from the target image. This can be achieved through object detection and image segmentation techniques. In the embodiments of this application, the foreground information can be presented in the form of a white-background image, which means that the background interference is removed, highlighting the style feature information of the display object. The white-background image can capture the details of the display object more precisely, including shape, material, color, etc., which is crucial for subsequent style feature analysis.

[0152] Optionally, the image processing model can perform in-depth analysis on the obtained white-background image to extract the style feature information of the display object. The above process can involve various image recognition and analysis techniques, including but not limited to deep learning models, feature extraction networks, etc. The extraction of style feature information usually focuses on aspects such as the design elements, color matching, and line style of the display object. The above style feature information together constitutes a mathematical description of the style to which the display object belongs.

[0153] Optionally, once the style feature information of the display object is obtained, the image processing model can be used to search for a matching collocation image in the candidate image pool. The above search process can rely on the similarity calculation of feature vectors, especially metrics such as cosine similarity. By comparing the distances between the style feature vectors of different images, the degree of style matching can be measured.

[0154] In the embodiments of this application, in order to accurately obtain the white-background image of the display object, object detection and image segmentation techniques can be adopted. The above techniques can identify specific objects in the target image and separate them from other elements to generate a clean white-background image. Deep-level style feature information is captured from the white-background image, even if the above style feature information is not easily perceptible intuitively. The similarity is calculated using the distance or angle information of the feature vectors, and then the candidate images are sorted, and the images ranked at the top are selected as the collocation images. The above process involves the construction and query optimization of the vector database to support efficient image retrieval. The image processing model can have the ability to dynamically adjust the number of targets and determine the number of targets of the final collocation image according to different application scenarios (such as generating collocation suggestions, creating diverse scenarios, etc.) and user requirements (such as the richness and diversity of the collocation).

[0155] Generally speaking, through the above methods, it is possible to effectively extract the style feature information of the display object from the target image, and use the image processing model to intelligently find the matching image with the same style in the candidate images, so as to realize furniture matching recommendations or scene designs with consistent styles and coordinated details. The above process integrates image processing, deep learning, and intelligent decision-making technologies, effectively improving the user shopping experience and technical effects in the e-commerce scenario.

[0156] As an alternative implementation, the image processing model includes a feature extraction model. The style feature information of the target image includes a first style feature vector, which is used to represent the style semantic features of the corresponding display object in the embedding space. Using the image processing model, the foreground information is analyzed to obtain the style feature information of the target image, including: using the feature extraction model to identify the style semantic features of the display object from the foreground information; using the feature extraction model to map the style semantic features of the display object into the embedding space to obtain the first style feature vector.

[0157] In this embodiment, in the process of using the image processing model to analyze the foreground information to obtain the style feature information, the feature extraction model in the image processing model can be used to identify the style semantic features of the display object from the foreground information, and the feature extraction model is used to map the style semantic features of the display object into the embedding space to obtain the first style feature vector. Among them, the feature extraction model can be a feature extractor, and the style semantic features can be matching-related semantics.

[0158] Optionally, this embodiment elaborates on the process of using the feature extractor in the image processing model to analyze and obtain the style semantic features of the display object from the foreground information of the target image, and then map them into the embedding space to form the first style feature vector.

[0159] Optionally, the feature extraction model can be a deep neural network, which is responsible for identifying the style semantic features from the foreground information (i.e., the display object image with background interference removed). The style semantic features can include design elements (such as lines, shapes), color combinations, material textures, pattern styles, etc., which together constitute the semantic description of the style to which the display object belongs. The feature extraction model is trained to automatically identify and understand the above style semantic features, even if the above style semantic features may be very subtle or complex visually.

[0160] Optionally, once the feature extraction model identifies the style semantic features of the display object, the next task is to map the above style semantic features into the embedding space to generate the first style feature vector. The embedding space is a high-dimensional continuous vector space, where each vector represents the feature representation of an object. Vectors of similar objects are closer in the space, while vectors of dissimilar objects are farther apart. The process of converting style semantic features into the first style feature vector involves feature encoding and dimensionality reduction techniques. The first style feature vector thus becomes a compact mathematical representation of the style features of the display object, facilitating subsequent similarity calculation and matching.

[0161] Optionally, the feature extraction model learns how to extract features crucial for style identification during the training process. The above process requires training on a large number of images labeled with specific styles to optimize the model's feature recognition ability. When constructing the embedding space, it can be ensured that vectors of similar styles cluster in the space, while vectors of dissimilar styles are separated. This can be achieved through contrastive learning, self-supervised, or semi-supervised learning methods to learn the intrinsic representation of style features without relying on a large number of explicit labels. To improve computational efficiency, the feature vectors may be normalized and dimensionally reduced before being mapped into the embedding space. In the embedding space, cosine similarity, Euclidean distance, or other metrics can be used to calculate the similarity between vectors, thereby finding the matching objects for the first style feature vector among the candidate images.

[0162] In the embodiment of the present application, by using the feature extraction model to identify style semantic features and mapping the style semantic features into the embedding space to form the first style feature vector, a deep understanding and mathematical representation of the style features of the target image can be achieved. The above representation not only contains intuitive style information but also integrates deeper semantic features, providing a solid foundation for subsequent image style matching and furniture matching. The above process demonstrates the powerful function of deep learning technology in the field of image feature analysis and style recognition.

[0163] As an optional implementation manner, in step S204, analyzing the style feature information of the target image by using an image processing model to obtain at least one matching image includes: determining a generation strategy corresponding to the target image based on the style feature information of the target image, where the generation strategy is used to represent the rule for comparing the target image and candidate images to generate a target number of matching images, and the image content of the candidate images includes candidate objects; controlling the image processing model to compare the style feature information of the target image with the style feature information of the candidate images according to the generation strategy to obtain a target number of matching images, where the matching objects are candidate objects whose style feature information matches the style features of the target image.

[0164] In this embodiment, in the process of analyzing a target image using an image processing model to obtain a target number of matching images that match the style feature information of the target image, a generation strategy corresponding to the target image can be determined based on the style feature information of the target image. According to the generation strategy, the image processing model is controlled to compare the style feature information of the target image with the style feature information of the candidate images, and a target number of matching images are obtained. Among them, the generation strategy can be used to represent the rules for comparing the target image and the candidate images to generate a target number of matching images, and the generation strategy can also be called a retrieval scheme. The matching object is a candidate object whose style feature information matches the style feature information of the target image.

[0165] Optionally, this embodiment illustrates how to use an image processing model to determine a generation strategy based on the style feature information of the target image and generate matching images that match the style of the target image according to the generation strategy.

[0166] Optionally, the generation strategy is essentially a set of rules or algorithms that guide the image processing model on how to screen out matching images that match the style features of the target image from the candidate images. The above-mentioned generation strategy can formulate specific comparison and retrieval rules based on the first style feature vector of the target image and other style feature information, such as color tone, design elements, spatial layout, etc.

[0167] For example, if the target image belongs to the minimalist style, the generation strategy may preferentially compare candidate objects in the candidate images with simple lines and single colors; for a target image in the Baroque style, the generation strategy may be more inclined to find matching objects with complicated lines, rich colors, and strong decorativeness.

[0168] Optionally, after determining the generation strategy, the image processing model can be controlled to compare the style feature information of the target image and the candidate images according to this strategy. The image processing model can extract features from the candidate objects in the candidate images to generate a second style feature vector, and the second style feature vector reflects the style semantic features of the candidate objects. Using the rules defined in the generation strategy, calculate the similarity between the first style feature vector of the target image and the second style feature vector of the candidate image, such as cosine similarity, Euclidean distance, etc. According to the set similarity threshold, target number, and other possible parameters (such as size, functional attribute matching), screen out eligible candidate objects as potential matching objects.

[0169] Optionally, after comparison and screening, the image processing model selects a target number of matching images from the candidate images according to the generation strategy, and the candidate objects in the above-mentioned matching images match the displayed objects in the target image in style. The selection of the target number may be based on the actual needs of the user, such as creating a set of diverse matching schemes or focusing on generating a few high-quality matching images. Finally, the generated matching images are displayed to the user to help them more intuitively understand the consistency and coordination in style between different furniture products and the displayed objects in the target image.

[0170] For example, the generation strategy can be customized according to different application scenarios and objectives, including factors such as style preferences, user historical behavior, and seasonal changes, to improve the relevance of the matching results and user satisfaction.

[0171] In the embodiment of the present application, by determining the generation strategy and controlling the image processing model to compare the style feature information of the target image and the candidate images, the above embodiment realizes the accurate and personalized generation of matching images. The above process not only relies on the feature extraction and vector representation capabilities of deep learning, but also incorporates strategy and flexibility, and can adapt to diverse application scenarios and user needs, thereby enhancing the satisfaction and practicality of the overall matching effect.

[0172] As an optional implementation manner, the generation strategy includes a first generation strategy. The first generation strategy is used to represent the rules for comparing the target image and the candidate images, and comparing the matching images and the candidate images to generate a target number of matching images. According to the generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate images to obtain a target number of matching images, including: controlling the image processing model according to the first generation strategy to compare the style feature information of the target image with the style feature information of the candidate images to obtain a first number of matching images, where the first number is less than the target number; controlling the image processing model to compare the style feature information of the first number of matching images with the style feature information of the candidate images to obtain a second number of matching images, where the second number is less than the target number; determining the target number of matching images based on the first number of matching images and the second number of matching images.

[0173] In this embodiment, in the process of controlling the image processing model to compare the style feature information of the target image with the style feature information of the candidate image according to the generation strategy to obtain the target number of paired images, the image processing model can be controlled according to the first generation strategy to compare the style feature information of the target image with the style feature information of the subsequent images to obtain the first number of paired images. The image processing model can be controlled to compare the style feature information of the first number of paired images with the style feature information of the candidate image to obtain the second number of paired images. The target number of paired images can be determined based on the first number of paired images and the second number of paired images. Among them, the generation strategy includes the first generation strategy. The first generation strategy can be used to represent the rule of comparing the target image and the candidate image, as well as comparing the paired image and the candidate image to generate the target number of paired images. That is, some of the paired images determined under the first generation strategy can participate in the process of determining the remaining paired images until the target number of paired images is determined. The first generation strategy can be random selection of furniture retrieval.

[0174] Optionally, this embodiment elaborates on a phased retrieval and comparison process, which details how to gradually generate the target number of paired images to ensure the diversity and quality of the pairing results.

[0175] Optionally, for the first stage under the first generation strategy, the target image and the candidate image can be compared to generate the first number of paired images. In the first stage, according to the first generation strategy, the image processing model is controlled to compare the style feature information of the target image with the style feature information of each candidate image in the candidate image library. The first generation strategy can take various forms, such as random selection of furniture retrieval, that is, randomly select a certain number of images from the candidate images for preliminary comparison. By calculating the distance or similarity between the style feature vectors of the target image and the candidate image, the system can screen out the initial batch of paired images, and the number of this batch of paired images is the first number, which is less than the target number. The above steps are aimed at quickly locating the preliminary candidate pairing objects that match the style of the target image.

[0176] Optionally, for the second stage under the second generation strategy, a comparison can be made between the first quantity of paired images and candidate images to generate the second quantity of paired images. After determining the first quantity of paired images, the second stage is executed, and the image processing model is continuously used to further compare the style feature information of the first quantity of paired images with the style feature information of the remaining candidate images. The purpose of the above steps is to continue to expand the paired set on the basis of the already selected pairings, making it more rich and coordinated. By comparing the styles of the preliminary paired images and candidate images, the second batch of paired images can be screened out, and the quantity of this batch of paired images is the second quantity, which is still less than the target quantity. Compared with the first stage, the retrieval in the second stage may be more refined, taking into account the influence of the existing paired images and further optimizing the pairing scheme.

[0177] Optionally, combining the paired images of the first quantity and the second quantity, comprehensively evaluate the diversity and aesthetics of the pairings, and finally generate a set of paired images with the target quantity. The above process may involve further screening, sorting, and optimizing the paired images to ensure that the final paired set not only conforms to the style characteristics of the target images but also meets the requirements of the target quantity.

[0178] Optionally, the random selection furniture retrieval method under the first generation strategy can quickly narrow the search scope, reduce the computational cost, and at the same time provide a diverse set of basic paired images for subsequent comparison. By comparing the first quantity of paired images with the remaining candidate images, the paired set can be gradually optimized, increasing its diversity while maintaining style consistency. When finally determining the paired images with the target quantity, multiple strategies such as clustering analysis, reinforcement learning, and rule matching can be combined to achieve a paired combination that meets the requirements under the established style characteristics. To improve the pairing effect, the system may dynamically adjust the retrieval parameters such as the similarity threshold and the filtering conditions for candidate images according to user feedback, design trends, and resource limitations, in order to achieve better pairing results.

[0179] In the embodiments of this application, through the above-mentioned phased retrieval and comparison mechanism under the first generation type, it is possible to accurately locate the paired images that match the style of the target images among a large number of candidate images, ensuring both the quality of the pairing results and increasing their diversity, providing richer and more personalized furniture pairing suggestions for the users of the e-commerce platform. This process makes full use of the feature comparison ability of the image processing model and also demonstrates the important applications of strategy planning and decision optimization in intelligent pairing.

[0180] As an alternative implementation, based on the first quantity of paired images and the second quantity of paired images, determining the target quantity of paired images includes: in response to the sum of the first quantity and the second quantity being less than the target quantity, determining the second quantity as the first quantity and returning to perform the following steps until the sum of the first quantity and the second quantity is equal to the target quantity, and determining the paired images of the first quantity and the paired images of the second quantity as the paired images of the target quantity: controlling an image processing model to compare the style feature information of the paired images of the first quantity with the style feature information of candidate images to obtain the paired images of the second quantity.

[0181] In this embodiment, in the process of determining the paired images of the target quantity based on the paired images of the first quantity and the paired images of the second quantity, if the sum of the first quantity and the second quantity is less than the target quantity, the second quantity can be determined as the first quantity and returned to perform the following steps until the sum of the first quantity and the second quantity is equal to the target quantity, and determining the paired images of the first quantity and the paired images of the second quantity as the paired images of the target quantity: controlling an image processing model to compare the style feature information of the paired images of the first quantity with the style feature information of candidate images to obtain the paired images of the second quantity.

[0182] Optionally, this embodiment elaborates on how to gradually expand the paired set (the paired images of the target quantity) through an iterative approach until the target quantity is reached during the process of generating the paired images of the target quantity.

[0183] Optionally, before starting the iterative process, it can be checked whether the total number of the current paired images of the first quantity and the second quantity has reached the target quantity. If the sum of the total number of the paired images of the first quantity and the second quantity is less than the target quantity, then the system needs to further expand the paired set. If the current total number of paired images is less than the target quantity, update the paired images of the second quantity to the paired images of the first quantity. This means using the above initially confirmed paired images as a new retrieval starting point and performing a comparison of style feature information again using the image processing model.

[0184] Optionally, in the above stage, an image processing model can be controlled to compare the style feature information of the paired images of the first quantity with the remaining images in the candidate image library. The above comparison process aims to find new images that are stylistically coordinated with the existing paired images to further enrich the paired set. The paired images of the second quantity will be screened out according to the comparison results, and the number of these paired images may vary according to the specific configuration of the system and the size of the target quantity.

[0185] Optionally, continue the above iterative process of updating and retrieval, update the newly generated second quantity of paired images into the first quantity of paired images again, and then perform the comparison and retrieval again until the sum of the total number of paired images of the first quantity and the second quantity is equal to the target quantity. The above iterative process ensures the gradual expansion and optimization of the paired set, so that the finally generated paired set not only includes images that closely match the style feature information of the target image, but also takes into account the coordination between paired images. While reaching the target quantity, it ensures the diversity and coordination of the pairing.

[0186] In the embodiment of the present application, through the above iterative retrieval and update mechanism, a paired image set that not only conforms to the style feature of the target image but also meets the target quantity requirement can be gradually generated. The above process not only reflects the application of deep learning and image processing technologies in the field of feature comparison and retrieval, but also demonstrates the important role of strategy planning and iterative optimization in improving the quality of the pairing result (paired image).

[0187] As an alternative implementation, the generation strategy includes a second generation strategy. The second generation strategy is used to represent the rule for comparing the target image and the candidate image to generate the paired images of the target quantity. According to the generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate image to obtain the paired images of the target quantity, including: according to the second generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate image to obtain the target candidate image, where the similarity between the style feature information of the target candidate image and the style feature information of the target image is greater than the similarity threshold; use the image processing model to determine the target candidate image whose attribute features of the target quantity match the attribute features of the display object as the paired images of the target quantity.

[0188] In this embodiment, when controlling the image processing model to compare the style feature information of the target image with the style feature information of the candidate image according to the generation strategy to obtain the paired images of the target quantity, the image processing model can be controlled according to the second generation strategy to compare the style feature information of the target image with the style feature information of the candidate image to obtain the target candidate image. The target candidate image whose attribute features of the target quantity match the attribute features of the display object can be determined as the paired images of the target quantity by using the image processing model. Among them, the generation strategy can include the second generation strategy. The second generation strategy can be used to represent the rule for comparing the target image and the candidate image to generate the paired images of the target quantity, that is, under the second generation strategy, the paired images are found only through the target image, and there is no need to find the remaining paired images through other paired images. The second generation strategy can be the main furniture retrieval.

[0189] Optionally, this embodiment elaborates a second generation strategy based on main furniture retrieval for efficiently generating a target number of matching images that match the style features of the target image.

[0190] Optionally, according to the second generation strategy, control the image processing model to compare the style feature information between the target image and the candidate image. The above process involves the application of a deep learning model, converting the target image and the candidate image into high-dimensional feature vectors, and then calculating the similarity between the above feature vectors in the embedding space.

[0191] Optionally, based on the similarity evaluation between the style feature information of the target image and the candidate image respectively, further screen out the target candidate image set, and the style feature information of the above target candidate image has a similarity greater than the preset similarity threshold with the style feature information of the target image. The target candidate image set not only matches the target image in style, but may also be screened based on other attribute features, such as size, material, functional characteristics, etc., to ensure the feasibility and coordination of the matching.

[0192] Optionally, after screening out the target candidate images, use the image processing model to further analyze the attribute features of the matching images and match them with the attribute features of the objects shown in the target image to determine the final target number of matching image sets. The above matching process may involve the quantification and comparison of attribute features, such as color tone, size ratio, design elements, etc. Those target candidate images that are coordinated with the objects shown in terms of attributes will be preferentially selected to generate the final matching image set until its quantity is equal to the target number.

[0193] In the embodiment of the present application, the second generation strategy based on main furniture retrieval can efficiently generate the target number of matching images while ensuring style consistency, providing users with fast and accurate furniture matching recommendations. The above process highlights the role of the deep learning model in image feature representation and comparison, and at the same time demonstrates the importance of attribute feature matching in ensuring the practicality of the matching, jointly improving the quality of furniture matching services and user experience in the e-commerce scenario.

[0194] As an optional implementation manner, in step S208, the image content of the target scene image includes the display scene of the object to be displayed and the matching object. According to the spatial position relationship, rendering the matching image and the target image into the target scene image includes: obtaining the three-dimensional object models corresponding to the target image and the matching image respectively, and the three-dimensional scene model corresponding to the display scene; rendering the three-dimensional object models into the three-dimensional scene model according to the spatial position relationship to obtain the rendered three-dimensional scene model; determining the rendered target scene image based on the rendered three-dimensional scene model.

[0195] In this embodiment, in the process of rendering the matching image and the target image into the target scene image according to the spatial position relationship, the three-dimensional object models of the target image and the matching image respectively can be obtained, as well as the three-dimensional scene model corresponding to the displayed scene in the target scene image. The three-dimensional object models can be rendered into the three-dimensional scene model according to the spatial position relationship to obtain the rendered three-dimensional scene model. Based on the rendered three-dimensional scene model, the rendered target scene image is determined. The image content of the target scene image may include the displayed scene of the object to be displayed and the matching object.

[0196] Optionally, this embodiment illustrates how to integrate the displayed object in the target image and the matching object in the matching image into the displayed scene of the target scene image through three-dimensional modeling and rendering techniques to generate an intuitive and realistic matching image.

[0197] Optionally, before rendering the matching image and the target image into the target scene image, the three-dimensional object model of the displayed object in the target image and the three-dimensional object model of the matching object in the matching image can be obtained. The above models can be created by 3D scanning, modeling software or selected from existing 3D model libraries. The above three-dimensional object models provide detailed information such as the shape, size, texture and material of the displayed object and the matching object in the three-dimensional space. At the same time, the three-dimensional scene model corresponding to the target scene image can also be obtained. The three-dimensional scene model represents the spatial layout of the displayed scene, including static environmental elements such as walls, floors, ceilings, etc. and their spatial position and size information. The three-dimensional scene model can be created based on virtual reality or augmented reality technology or selected from a pre-built 3D scene database.

[0198] Optionally, according to the relative position relationship between the displayed object in the target image and the matching object in the matching image, these three-dimensional object models will be placed and rendered into the three-dimensional scene model. The above process involves multiple links such as precise position positioning, perspective adjustment, and lighting simulation to ensure that the layout of the three-dimensional object models in the three-dimensional scene model is reasonable, the perspective is appropriate, and the lighting effect is realistic, highly similar to the actual placement effect.

[0199] Optionally, after the above three-dimensional object model addition and rendering process, a rendered three-dimensional scene model will be obtained, which includes the display object of the target image and the matching object of the matching image. Subsequently, the above three-dimensional scene model can be converted into a two-dimensional image, that is, the finally rendered target scene image. The above conversion process can be completed by a 3D rendering engine, which can render the elements in the three-dimensional scene model, including physical phenomena such as dynamic lighting, shadow effects, reflection and refraction, into a two-dimensional image in a high-fidelity manner. The finally generated target scene image will show the complete scene effect after matching, enabling users to clearly preview the placement layout and matching effect of furniture in the actual space.

[0200] In the embodiment of the present application, by obtaining the three-dimensional object model and the three-dimensional scene model and performing three-dimensional rendering based on the spatial position relationship, the above implementation method can generate a target scene image with high quality and strong immersion, providing users with an intuitive and realistic furniture matching preview, and greatly enhancing the attractiveness and interaction experience of furniture matching services in the e-commerce scenario.

[0201] As an optional implementation method, rendering the three-dimensional object model into the three-dimensional scene model according to the spatial position relationship to obtain a rendered three-dimensional scene model includes: in response to the image content of either the target image or the matching image including the target object, performing a deformation process on the three-dimensional object model corresponding to the target object, where in any dimension in the three-dimensional space, the size of the target object is smaller than the size threshold; adjusting the spatial position relationship corresponding to the three-dimensional object model after the deformation process to obtain an adjusted spatial position relationship, where the spatial position relationship corresponding to the three-dimensional object model after the deformation process is the spatial position relationship between the three-dimensional object models corresponding to the display object and the matching object other than the three-dimensional object model corresponding to the target object in the three-dimensional object model; rendering the three-dimensional object model after the deformation process into the three-dimensional scene model according to the adjusted spatial position relationship.

[0202] In this embodiment, during the process of rendering the three-dimensional object model into the three-dimensional scene model according to the spatial position relationship, it can be determined whether the image content of either the target image or the matching image includes the target object. If there is a target object, a deformation process can be performed on the three-dimensional object model corresponding to the target object. Adjust the spatial position relationship between the three-dimensional object model after the deformation process and other three-dimensional object models to obtain an adjusted spatial position relationship. Thus, the corresponding three-dimensional object model can be rendered into the three-dimensional scene model according to the adjusted spatial position relationship. Among them, in any dimension in the three-dimensional space, the size of the target object is smaller than the size threshold, that is, the target object can be approximately regarded as a two-dimensional object, such as a mural, a carpet, etc.

[0203] Optionally, this embodiment details how to integrate these two-dimensional or nearly two-dimensional objects (such as murals, carpets, etc.) into a three-dimensional scene model through deformation processing and adjustment of spatial position relationships when processing a target image or a matching image containing a target object, so as to generate a realistic and coordinated scene rendering effect.

[0204] Optionally, through image recognition technology, it is determined whether a specific target object is included in the target image or the matching image. Since the size of the above-mentioned target object is less than a specific size threshold in any dimension in the three-dimensional space, it can be regarded as a two-dimensional or nearly two-dimensional object, such as hanging pictures, carpets, ornaments, etc. The above judgment step is crucial for subsequent model deformation and position adjustment, ensuring the pertinence and effectiveness of the processing flow.

[0205] Optionally, after the target object is recognized, a deformation process will be performed on the three-dimensional model corresponding to the target object. The above processing mainly targets the features of the target object such as size, shape, and texture to ensure that its performance in the three-dimensional scene is consistent with that in the target image or the matching image.

[0206] For example, for a carpet, the size ratio and texture mapping of the three-dimensional model can be adjusted according to the placement position and shape of the carpet in the target image to adapt to the floor layout of the three-dimensional scene while maintaining the original design style and details of the carpet.

[0207] Optionally, after the deformation process is completed, the spatial position relationship between the three-dimensional model of the target object after deformation and other three-dimensional object models (including display objects and matching objects) in the scene can be further adjusted. The above adjustment process may involve the following aspects: position alignment, that is, ensuring that the position of the target object in the three-dimensional scene is consistent with that in the target image or the matching image, such as the hanging height and angle of a hanging picture, the placement position and direction of a carpet; collision detection, that is, through physical engine technology, to avoid collisions between the target object and other objects, ensuring the rationality and feasibility of the scene layout, such as the appropriate distance between the carpet and furniture, and between the hanging picture and the wall; spatial coordination, that is, based on aesthetic principles and spatial design knowledge, adjusting the relative position and posture between the target object and other objects in the scene to achieve visual harmony and beauty, such as the visual relationship between the hanging picture and the sofa, and the matching of the size of the carpet with the room space.

[0208] Optionally, after completing the adjustment of the spatial position relationship, the three-dimensional model of the target object after deformation processing can be integrated into the three-dimensional scene model according to the adjusted relative position for final rendering. The above process involves the processing of details such as lighting, shadows, and materials to enhance the realism and immersion of the scene. After the rendering is completed, the user can preview the actual effect of the target object in the matching scene through the generated three-dimensional scene model. Whether it is the decorative effect of the hanging picture or the laying texture of the carpet, it can be intuitively displayed.

[0209] In the embodiment of the present application, through deformation processing and intelligent adjustment of the spatial position relationship, the problem of natural integration of two-dimensional or nearly two-dimensional objects in a three-dimensional scene is solved, providing a more realistic and coordinated scene preview experience for users, and further improving the accuracy and satisfaction of the furniture matching recommendation service.

[0210] As an alternative implementation manner, in step S206, the image content of the target scene image includes the display scene of the object to be displayed and the matching object. Determining the spatial position relationship between the display object and the matching object includes: obtaining the attribute characteristics of both the display object and the matching object, and the layout strategy of the display scene, where the attribute characteristics at least include one of the following: the functional attributes of the display object and the matching object, and the dimensional attributes of the display object and the matching object in three-dimensional space. The layout strategy is used to represent the rules for laying out the display object and the matching object in a display scene that matches the style characteristic information of the display object; based on the layout strategy, the attribute characteristics, and the layout information of the target scene image, determining the spatial position relationship, where the layout information is used to represent the structure and dimensions of the display scene in three-dimensional space.

[0211] In this embodiment, during the process of determining the spatial position relationship between the display object and the matching display object, the attribute characteristics between the display object and the matching display object, and the layout strategy of the display scene can be obtained. Thus, according to the layout strategy, the attribute characteristics, and the layout information of the target scene image, the spatial position relationship is determined. Among them, the image content of the target scene image includes the display scene of the matching object and the display object to be displayed. The attribute characteristics can at least include one of the following: the functional attributes of the display object and the matching object, and the dimensional attributes of the display object and the matching object in three-dimensional space. The layout strategy can be used to represent the rules for laying out the display object and the matching object in a display scene that matches the style characteristic information of the display object. The layout information can be used to represent the structure and dimensions of the display scene in three-dimensional space. The layout strategy ensures that the layout of the furniture meets the requirements of a specific style, maintaining the consistency and beauty of the overall design. For example, in the new Chinese style, a symmetrical layout is preferred, while in the modern style, an asymmetrical but smooth layout can be adopted.

[0212] Optionally, this embodiment illustrates how to determine the spatial position relationship between the display object and the matching object during the process of generating the target scene image, so as to ensure that the final matching scene not only has a consistent style, but also has a reasonable layout and high space utilization rate.

[0213] Optionally, obtain the attribute characteristics of the display object and the matching object. The above attribute characteristics at least include functional attributes and dimensional attributes. The functional attribute indicates the use of the furniture, such as sofas, dining tables, beds, etc., while the dimensional attribute describes the actual size and shape of the furniture, which is crucial for reasonable layout in space. In addition, it is also possible to master the layout strategy of the display scene. The layout strategy refers to a set of rules based on a specific style, guiding the layout method of the display object and the matching object in the three-dimensional space. For example, in the new Chinese style, the layout strategy may emphasize symmetry, balance, and leaving blank spaces, while in the modern minimalist style, it may focus on an open and fluid asymmetric layout.

[0214] Optionally, the layout information of the target scene image provides details about the structure and size of the display scene in the three-dimensional space. By parsing the layout information, the system can understand the size of the room, the positions of the walls, doors, and windows, etc., as well as the layout of the existing furniture, laying a foundation for the subsequent determination of the spatial position relationship. The analysis of the layout information helps to identify the feasible placement areas in the display scene, avoid conflicts between the matching object and other items or fixed structures, and ensure the rationality of the layout.

[0215] Optionally, based on the obtained attribute characteristics, layout strategy, and layout information of the target scene image, the spatial position relationship between the display object and the matching object in the three-dimensional space will be determined. The above process can ensure that the functional attributes of the matching object are coordinated with the display object. For example, avoid placing a dining table in the bedroom. According to the dimensional attributes of the display object and the matching object, adjust the size, shape, and placement position of the furniture to adapt to the space limitations of the display scene, avoiding congestion or emptiness. Follow the layout strategy to ensure that the placement method of the furniture meets the requirements of a specific style, such as the symmetric layout in the new Chinese style, or the creation of a natural and relaxed atmosphere in the Nordic style. Consider the movement paths of people in the display scene and the dynamic lines in the use of furniture, avoid layout obstacles, and improve the circulation and utilization efficiency of the space.

[0216] In the embodiment of the present application, through the above method, a reasonable and beautiful spatial position relationship between the display object and the matching object can be effectively determined, providing strong technical support for the furniture matching recommendation service in the e-commerce scenario, not only improving the decision-making efficiency of users to purchase furniture, but also enhancing the pleasant experience of scenario-based shopping.

[0217] From the perspective of model training, the embodiment of the present application also provides a method for determining an image processing model. Figure 3It is a flowchart of a method for determining an image processing model according to an embodiment of the present application. As Figure 3 shown, the method may include:

[0218] Step S302, obtain an image sample set.

[0219] In the technical solution provided in step S302 of the present application above, the image sample set includes a target image sample set and a candidate image sample set. The image content of the target image sample includes at least one display object sample, and the image content of the candidate image sample set includes at least one candidate object sample.

[0220] In this embodiment, the above method is the data preparation stage in the entire matching generation method, and its core goal is to collect and prepare an image sample set for training an indoor matching model. The above steps are crucial for subsequent model training and matching ability generation.

[0221] Optionally, the target image sample set contains target image samples, and the image content of the above target image samples contains at least one display object sample. The display object sample can be any furniture item (such as a sofa, coffee table, chair, etc.). In the above stage, images containing furniture displays can be collected through various channels. The above images can be sourced from product pictures on e-commerce platforms, interior design magazines, social media sharing, designer portfolios, etc. The acquisition of the target image sample set is not only an accumulation of data volume, but more importantly, to ensure the diversity and representativeness of the samples, covering furniture of various styles, sizes, functions, and materials, as well as the display effects of the above furniture in different environments and layouts, which helps to train an indoor matching model that can adapt to multiple scenarios.

[0222] Optionally, in contrast to the target image sample set, the candidate image sample set contains candidate image samples, and the image content of the above image samples contains at least one candidate object sample. The candidate object sample can be other furniture or ornaments used to match the display object. The purpose of obtaining the candidate image sample set is to provide potential matching options for the deep learning model to be trained. Through contrastive learning, the model can learn to identify which candidate objects match the display objects in the target image in terms of style, function, size, and layout, and which do not. Therefore, the candidate image sample set should contain a rich variety of furniture types and styles to ensure the comprehensiveness and richness of the training data.

[0223] Step S304, use the image sample set to perform contrastive learning on the deep learning model to obtain an image processing model.

[0224] In the technical solution provided in step S304 of the present application, the image processing model is used to analyze the style feature information of the target image to obtain at least one matching image. The image content of the target image includes at least one display object, and the style feature information of the target image is used to represent at least one attribute feature of the display object. The image content of the matching image includes matching objects that match at least one attribute feature. The matching image and the target image are rendered into the target scene image based on the spatial position relationship between the display object and the matching object.

[0225] In this embodiment, the above method is a key stage for training a deep learning model based on contrast learning to generate an image processing model. This model will be used to analyze the style feature information in the target image and retrieve matching images that match the target style based on this information.

[0226] Optionally, the primary task of contrast learning for the deep learning model is to utilize an image sample set. The image sample set consists of a target image sample set and a candidate image sample set. Each target image in the target image sample set contains at least one display object, and this display object carries specific style feature information. The candidate image sample set includes various potential matching objects, and the styles of the above matching objects vary, some of which match the style of the display object in the target image, while some do not.

[0227] Optionally, the purpose of model training is to learn to distinguish which matching objects match the style of the display object and which do not. Through contrast learning, the model compares positive sample pairs (i.e., display objects and matching objects with matching styles) and negative sample pairs (i.e., display objects and matching objects with non-matching styles), minimizes the distance between positive sample pairs, and maximizes the distance between negative sample pairs, thereby forming a clear style boundary in the feature space of the model.

[0228] Optionally, the deep learning model trained through contrast learning will ultimately be transformed into an image processing model. The output target of this model is to generate matching images that match a specific style, and the quantity is determined by the requirement (i.e., the target quantity), and the target quantity is associated with the style. This means that users can request the model to generate the corresponding number of matching images according to the style and quantity they need.

[0229] Optionally, the application scenarios of the above image processing model mainly involve two aspects: style analysis and collocation image generation, and scene rendering based on spatial position relationships. The user can upload a target image containing a display object, and the model will analyze the style feature information of the target image and retrieve collocation objects that match the style based on this information to generate a corresponding number of collocation images. The generated collocation images and the target image can be used for rendering into a target scene image, where the target scene image can be a 3D show flat or a background image of any virtual space. The model will consider the position relationships of the display object and the collocation objects in space to ensure that the collocation images are both style-consistent and reasonably arranged during scene rendering, forming a coordinated visual effect.

[0230] Through the above steps S302 to S304 of this application, an image sample set is obtained, and the deep learning model is subjected to contrast learning using the image sample set to obtain an image processing model. The target image can be analyzed by the pre-trained image processing model to obtain the style of the display object and automatically determine how the display object is collocated in the display scene in this style. For example, the number of collocation images in this style and the collocation objects in them match the display object and conform to the belonging style. That is, by introducing the image processing model for automated style analysis, rapid retrieval of collocation images, automatic determination of spatial position relationships, and optimized image rendering and synthesis technologies, this application makes it possible to generate a large number of stylized, personalized, and high-quality collocation images in a short time, greatly promoting the digital process of the collocation design industry, improving the user experience and business operation efficiency, thereby achieving the technical effect of improving the generation efficiency of images and solving the technical problem of low efficiency of image generation.

[0231] The above method of this embodiment will be further introduced below.

[0232] As an optional implementation manner, step S304, using the image sample set to perform contrast learning on the deep learning model to obtain an image processing model, includes: using the deep learning model to determine a collocation image sample set from the candidate image sample set, where the image content of the collocation image sample set includes at least one collocation object sample, and the collocation object sample is a candidate object sample whose attribute features match the attribute features of the display object sample; using at least one display object sample, the corresponding collocation object sample, and the candidate object samples other than the collocation object sample in at least one candidate object sample to perform contrast learning on the deep learning model to obtain an image processing model.

[0233] In this embodiment, in the process of using an image sample set to perform contrastive learning on a deep learning model to obtain an image processing model, the deep learning model can be used to determine a matching image sample set from a candidate image sample set. The deep learning model is subjected to contrastive learning using a display object sample, a corresponding matching object sample, and candidate object samples other than the above, to obtain an image processing model. Among them, the display object sample can be an anchor. The matching object sample can be a positive sample. The candidate object samples other than the matching object sample can be negative samples. The above anchor, positive sample, and negative sample can form training data (compatibility training data) Triplet: Among them, x can be used to represent the anchor; can be used to represent the positive sample; can be used to represent the negative sample.

[0234] Optionally, this embodiment can train a deep learning model through contrastive learning to generate an image processing model. This process particularly focuses on the model's ability to understand and analyze style feature information, so as to accurately generate a matching image that matches the style of the display object in the target image.

[0235] Optionally, the deep learning model is used to screen and determine a matching image sample set from the candidate image sample set. The above screening process depends on the model's preliminary style recognition ability, and attempts to find matching object samples with styles that match the display object sample in the target image from a large number of candidate images. The matching object sample becomes the positive sample in the subsequent training stage, and is used to characterize the characteristic of matching the style of the display object.

[0236] For example, assume that the display object in the target image is a new Chinese style sofa. The model will retrieve matching objects with the same new Chinese style from the candidate image sample set, such as square tables, tea sets, calligraphy and paintings, etc. These objects constitute the matching image sample set.

[0237] Optionally, a display object sample is selected from the target image sample set, a matching object sample that matches its style is selected from the matching image sample set, and a candidate object sample that does not match the style of the display object is randomly selected from the candidate image sample set to jointly construct a training data Triplet: The construction of Triplet data is the key to contrastive learning. By comparing between the anchor, positive sample, and negative sample, it helps the model learn the ability to distinguish and recognize style compatibility. The distance between the positive sample and the anchor should be smaller than the distance between the positive sample and the negative sample, prompting the model to form a clear style boundary in the feature space.

[0238] Optionally, using the constructed Triplet data, start the contrastive learning of the deep learning model. The goal of the above stage is to optimize the model parameters so that they can effectively capture the feature representations related to style compatibility in the feature space. The training process is completed through the following steps: Feature extraction, extracting high-dimensional feature vectors from the display object samples, positive samples, and negative samples in each Triplet, and these feature vectors characterize the style and content features of the images; Contrastive loss calculation, calculating the contrastive loss based on the distance between the feature vectors, such as using the triplet loss, which encourages the model to reduce the distance between the positive sample and the anchor, while increasing the distance between the negative sample and the anchor. Backpropagation and weight update, feeding back the contrastive loss into the model, adjusting the weight parameters of the model through the backpropagation algorithm, and optimizing the feature representation until the model converges or reaches a predetermined number of training epochs. Model evaluation and tuning, during the training process, continuously evaluate the model performance, such as the accuracy and recall rate of style matching, and adjust the model architecture and training strategy according to the evaluation results to achieve accurate style recognition.

[0239] Optionally, after being trained by contrastive learning, the deep processing model will become a mature image processing model. The capabilities of the above image processing model are mainly reflected in: being able to analyze the style feature information in the input target image and identify the style attributes of the display object. Being able to retrieve the matching object samples with compatible styles from the candidate image sample set based on the style feature information of the target image and generate a set of matching images. While retrieving the matching images, it can also analyze the potential positional relationship between the display object and the matching object in space, providing a reference for subsequent scene rendering.

[0240] In the embodiment of the present application, an image processing model is successfully trained using contrastive learning technology. Based on the analysis of style features, this model can accurately generate matching images that match the style of the display object in the target image. It not only improves the automation level of furniture matching but also provides more accurate and personalized scene-based recommendation services for e-commerce platforms, thus greatly enhancing the user experience and shopping satisfaction.

[0241] As an alternative implementation, a matching image sample set is determined from a candidate image sample set using a deep learning model, including: analyzing the image sample set using the deep learning model to obtain a first style feature vector sample of a display object sample and a second style feature vector sample of a candidate object sample, where the first style feature vector sample is used to represent the style semantic features of the display object sample in the embedding space, and the second style feature vector sample is used to represent the style semantic features of the candidate object sample in the embedding space; using the deep learning model to compare the first style feature vector sample and the second style feature vector sample to obtain a comparison result; and determining the candidate image sample set corresponding to the candidate object sample whose comparison result is that the similarity between the first style feature vector sample and the second style feature vector sample is greater than the similarity threshold as the matching image sample set.

[0242] In this embodiment, in the process of determining a matching image sample set from a candidate image sample set using a deep learning model, the image sample set can be analyzed using the deep learning model to obtain a first style feature vector sample of a display object sample and compare it with a second style feature vector sample of a candidate object sample to obtain a comparison result. The candidate image sample set corresponding to the candidate object sample whose comparison result is that the similarity between the first style feature vector sample and the second style feature vector sample is greater than the similarity threshold is determined as the matching image sample set. The comparison result can be the distance between the first style feature vector and the second style feature vector.

[0243] Optionally, this embodiment describes the process of screening and determining a matching image sample set from a candidate image sample set using a deep learning model. By analyzing the style features of the images and calculating the similarity of the style feature vectors of different furniture images in the embedding space, it is possible to automatically identify which furniture images match the style of the display object in the target image.

[0244] Optionally, the image sample set is analyzed using a deep learning model. The deep learning model can be a pre-trained image feature extractor that can convert an image into a fixed-length feature vector to represent the style and content information of the image. After obtaining the style feature vectors of the display object and the candidate object, the feature vectors will be compared to calculate the similarity. The similarity can be quantified by the distance between the feature vectors, such as using distance metrics such as Euclidean distance, cosine similarity, or Manhattan distance.

[0245] For example, can be used to represent the distance between the furniture in two images, where f θ can be the feature extractor. In the embodiments of the present application, Given a triple data, the deep learning model can determine which of the two alternative furniture pieces matches the target furniture.

[0246]

[0247] Among them, d0 can be used to represent f θ The calculated distance (similarity) between the anchor point and the positive sample; d1 can be used for f θ The calculated distance (similarity) between the anchor point and the negative sample; It can be used to represent the supervision signal.

[0248] During the training process, the deep learning model can continuously optimize the feature extractor to obtain an embedding space containing collocation-related semantics, so that the distance between two matching furniture in the embedding space is relatively close, and the distance between two non-matching furniture in the embedding space is relatively far.

[0249] Optionally, according to the similarity calculation result, a collocation image sample set will be determined. When the similarity between the second style feature vector sample of the candidate object and the first style feature vector sample of the display object exceeds the preset similarity threshold, the candidate image sample corresponding to the candidate object will be determined as a part of the collocation image sample set. Traverse each sample in the candidate image sample set, calculate its similarity with the display object, and retain those samples with similarity higher than the threshold. Through the above screening, a collocation image sample set is obtained, and the images in it represent furniture that matches the style of the display object.

[0250] As an optional implementation manner, step S302, obtaining the image sample set, includes: using the search prompt word to obtain the first initial image sample set, where the search prompt word is at least used to describe the characteristics of the image content of the initial image samples in the first initial image sample set; using the prompt information to perform screening processing on the first initial image sample set to obtain the screened first initial image sample set, where the prompt information is used to determine whether the initial image samples in the first initial image sample set can be used to train the deep learning model; determining at least one initial image sample with image quality higher than the image quality threshold in the screened first initial image sample set as the second initial image sample set; extracting the foreground information sample set from the second initial image sample set and determining the foreground information sample set as the image sample set, where the image content of the foreground information sample set includes the display object sample and the candidate object sample.

[0251] In this embodiment, during the process of obtaining the image sample set, the first initial image sample set can be obtained by using search prompt words. The first initial image sample set can be screened by using prompt information to obtain the screened first initial image sample set. At least one initial image sample with an image quality higher than the image quality threshold in the screened first initial image sample set can be determined as the second initial image sample set. The foreground information sample set can be extracted from the second initial image sample set. And the foreground information sample set is determined as the image sample set. Among them, the search prompt words can be at least used to describe the characteristics of the image content of the initial image samples in the first initial image sample set, and the search prompt words can be query words. The prompt information (Prompt) can be used to determine whether the initial image samples in the first initial image sample set are used to train the deep learning model. The image content of the foreground information sample set can include the display object samples and candidate object samples. The foreground information sample set can be a collection of white background images of furniture. For example, the pictures of the bounding boxes (abbreviated as bbox) of each piece of furniture.

[0252] Optionally, this embodiment describes the process of obtaining and preprocessing the image sample set from the Internet for training the deep learning model.

[0253] Optionally, the selected keywords or phrases are used to describe the content features of the expected images, such as furniture types, styles, etc. For example, "modern style sofa", "Nordic style coffee table". By querying the above keywords through a search engine, a preliminary screening data set composed of a large number of pictures crawled from the Internet is obtained. These pictures may contain the target furniture and its matching environment, and initially cover a variety of styles and scenarios.

[0254] Optionally, the prompt information (Prompt) can be a set of predefined conditions or rules for evaluating whether a picture meets the requirements for training the deep learning model. For example, "whether it contains non-furniture elements (such as humans, pets)", "whether it is a product advertisement rather than a real scene". Based on the multimodal model and manual review, check whether the pictures meet the conditions specified in the prompt information, eliminate irrelevant or substandard pictures, reduce noise, and improve the quality of subsequent training data.

[0255] Optionally, at least one initial image sample with an image quality higher than the image quality threshold in the screened first initial image sample set is determined as the second initial image sample set. The image quality threshold can be a preset standard for evaluating factors such as picture resolution and clarity. Pictures below the threshold may not provide sufficient training information due to lack of details.

[0256] Optionally, the second set of initial image samples is a more pure and high-quality set of pictures retained after double screening of quality and content. The above steps lay a good data foundation for subsequent feature extraction and in-depth analysis. From the second set of initial image samples, a set of foreground information samples is extracted, and the set of foreground information samples is determined as the set of image samples.

[0257] Optionally, through image processing techniques (such as object detection algorithms), the specific parts of the furniture in the picture are located and cropped, the background is removed, and independent and clear furniture images, that is, bounding box (bbox) pictures, are obtained. The set formed by collecting foreground information (bbox pictures) contains white-background pictures of display object samples and candidate object samples, and the above white-background pictures will be used to construct the training data of the matching model.

[0258] In the embodiment of the present application, a high-quality set of image samples for training a deep learning model is gradually established through precise search term selection, strict picture screening, and efficient foreground information extraction. Each step aims to remove data that does not match the target and retain the most relevant and highest-quality images, so as to ensure that the model can learn features that truly reflect the furniture matching style during the training process. Finally, through the training of the deep learning model, knowledge of indoor furniture matching can be learned from the carefully prepared set of image samples, and functions such as intelligent furniture matching generation and recommendation can be realized.

[0259] The effectiveness of the above method lies in the ingenious use of large-scale data resources on the Internet. Through automated and semi-automated image screening and preprocessing means, the efficiency and accuracy of data preparation are greatly improved, the dependence on manual annotation is reduced, the cost is lowered, and the speed of model iteration is accelerated. At the same time, since the model can be trained on diverse image samples, the learned matching rules will be more generalizable and can adapt to different furniture styles and matching scenarios.

[0260] The embodiment of the present application also provides a method for generating an image, Figure 4 which is a flowchart of another method for generating an image according to the embodiment of the present application, as Figure 4 shown. The method may include:

[0261] Step S402, identifying an image of a product object to be matched from a product matching platform, where the image content of the product object image includes at least one product object. The product matching platform may be an e-commerce platform or other pre-set platforms or application programs for providing users with matching designs.

[0262] Step S404: Analyze the style feature information of the product object image by using the product matching model to obtain at least one matching image, where the style feature information of the product object image is used to represent at least one attribute feature of the product object, and the image content of the matching image includes matching objects that match at least one attribute feature.

[0263] Step S406: Determine the spatial position relationship between the product object and the matching object.

[0264] Step S408: Render the matching image and the product object image to the target scene image according to the spatial position relationship.

[0265] Step S410: Return the rendered target scene image to the product matching platform.

[0266] Through the above steps S402 to S410 of this application, the product object image to be matched is identified from the product matching platform; the style feature information of the product object image is analyzed by using the product matching model to obtain at least one matching image; the spatial position relationship between the product object and the matching object is determined; the matching image and the product object image are rendered to the target scene image according to the spatial position relationship; the rendered target scene image is returned to the product matching platform, thereby achieving the technical effect of improving the generation efficiency of the image and solving the technical problem of low efficiency of image generation.

[0267] The embodiment of this application also provides a method for generating an image from the perspective of the client. Figure 5 It is a flowchart of another method for generating an image according to the embodiment of this application. As Figure 5 shown, the method may include:

[0268] Step S502: Respond to the input operation on the operation interface and display the target image on the operation interface, where the image content of the target image includes at least one display object.

[0269] Step S504: Respond to the image generation instruction on the operation interface and display the rendered target scene image on the operation interface, where the rendered target scene image is obtained by rendering the matching image and the target image to the target scene image according to the spatial position relationship between the display object and the matching object. The matching image is obtained by analyzing the style feature information of the target image by using the image processing model. The style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes matching objects that match at least one attribute feature.

[0270] Through the above steps S502 to S504 of the present application, in response to an input operation on the operation interface, a target image is displayed on the operation interface; in response to an image generation instruction on the operation interface, a rendered target scene image is displayed on the operation interface, thereby achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0271] The embodiment of the present application also provides a method for generating an image from the Software as a Service (SAAS) side. Figure 6 It is a flowchart of another method for generating an image according to the embodiment of the present application, as Figure 6 shown, the method may include:

[0272] Step S602, obtain a target image by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the target image, and the image content of the target image includes at least one display object.

[0273] Step S604, analyze the style feature information of the target image using an image processing model to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one attribute feature.

[0274] Step S606, determine the spatial position relationship between the display object and the matching object.

[0275] Step S608, render the matching image and the target image into the target scene image according to the spatial position relationship.

[0276] Step S610, output the target scene image by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the rendered target scene image.

[0277] Through the above steps S602 to S610 of the present application, obtain a target image by calling a first interface; analyze the style feature information of the target image using an image processing model to obtain at least one matching image; determine the spatial position relationship between the display object and the matching object; render the matching image and the target image into the target scene image according to the spatial position relationship; output the target scene image by calling a second interface, thereby achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0278] According to the embodiment of the present application, an embodiment of an image generation system is also provided. Figure 7 It is a schematic diagram of an image generation system according to the embodiment of the present application, asFigure 7 As shown in Figure 7 , the image generation system 700 may include: a client 701 and a server 702.

[0279] The client 701 is used to upload a target image, where the image content of the target image includes at least one display object.

[0280] The server 702 is used to analyze the style feature information of the target image by using an image processing model to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one attribute feature; determine the spatial position relationship between the display object and the matching object; and render the matching image and the target image into the target scene image according to the spatial position relationship.

[0281] In this embodiment, an image generation system is provided. The target image is uploaded through the client 701. The server 702 uses an image processing model to analyze the style feature information of the target image to obtain at least one matching image; determines the spatial position relationship between the display object and the matching object; and renders the matching image and the target image into the target scene image according to the spatial position relationship. The client 701 can display the finally generated target scene image, thereby achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0282] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application, such as the data for inspection, are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0283] In the embodiment of this application, taking the indoor furniture matching as an example, the above content will be explained in detail.

[0284] Indoor furniture matching refers to the rational selection and combination of furniture in an indoor space according to functional requirements, aesthetic principles, space limitations, and personal preferences to create a harmonious, comfortable, and practical living environment. It is not just about the simple placement of furniture but an overall design process that comprehensively considers various factors such as spatial layout, style coordination, color matching, and material selection. The quality of furniture matching directly affects the living experience of the occupants and the utilization efficiency of the space. Therefore, this link occupies a crucial position in interior design. With the improvement of people's living standards and the increasing requirements for living environments, furniture matching has gradually evolved from traditional single-functional needs to personalized, scenario-based, and intelligent directions. Especially in the context of the rapid development of e-commerce platforms, consumers are increasingly inclined to purchase furniture products through online channels. However, although online shopping brings convenience, it also makes it difficult for consumers to intuitively feel the matching effect of furniture with the actual space, which poses new challenges for the matching of furniture products.

[0285] Currently, the matching of indoor furniture products mainly relies on the experience and aesthetic ability of professional designers. Designers usually provide design solutions based on customer needs through personalized customization, taking into account spatial dimensions, functional requirements, and current trends. However, there are many problems with the above-mentioned related technologies: Firstly, the work efficiency of designers is relatively low, often requiring a large amount of time for plan adjustment and optimization; Secondly, the cost of manual design is relatively high. Especially for ordinary consumers, hiring a professional designer may be a significant expense; Thirdly, there are significant differences in design standards and styles among different designers, resulting in uneven final presentation effects; Finally, the entire operation process is complex, involving multiple communications and revisions, which is prone to causing resource waste and time delays.

[0286] Indoor furniture product matching is an important technology for the scenario-based expression of furniture products on e-commerce platforms. In the above fields, the research on related technologies and methods is insufficient. Therefore, there is an urgent need for an improved solution in this field to solve the above problems.

[0287] Figure 8 It is a schematic diagram of a home improvement intelligent design tool in related technologies, such as Figure 8As shown, this design tool customized for home decoration designers integrates a series of rich functions and resources, aiming to simplify the design process and improve creative efficiency. Its core feature is to provide a large number of furniture picture libraries and home scene templates, covering spaces with different styles and uses, such as modern minimalist, classical European, rural pastoral, etc. Designers can, based on the above resources, freely arrange the positions of furniture in the scene through intuitive drag-and-drop operations and preview the design effects in real time, so as to create personalized home design solutions that meet the needs of customers. The tool also incorporates advanced image processing technology, allowing users to upload custom furniture pictures or scene pictures, which are automatically recognized and segmented and integrated into the design project. In addition, various design elements such as lighting, color, and texture are provided to enhance the realism and aesthetics of the scene and help designers better express their creativity.

[0288] Although this design tool provides home decoration designers with an efficient and intuitive design experience, it still faces the following challenges in practical applications: Although the database contains a large number of furniture and scene pictures, it may still not cover design requirements, especially for non-mainstream or customized furniture, lacking corresponding picture materials. The furniture styles and home trends in the market change rapidly, and the picture library in the tool may not be updated in time, making it difficult for designers to apply the latest furniture styles and popular trends. When designers select different combinations of furniture and scenes from the tool, it is difficult to ensure the consistency and coordination of the overall style, especially in the absence of professional matching guidance. Provide home decoration design, scene layout and rendering services, and at the same time provide users with intelligent matching solutions for soft furnishings. Its technology can provide some soft furnishings in a 3D show flat, but it is currently mainly targeted at designers and decoration companies and does not have a scene-based expression for furniture products on e-commerce platforms. Therefore, there are still technical problems with low image generation efficiency.

[0289] Figure 9 is a schematic diagram of a whole-house design platform in another related technology, such as Figure 9 As shown, the DIY whole-house design function of the e-commerce platform is based on 3D show flats and furniture products deployed in the platform, providing users with a personalized home design experience and accurate product recommendations. The core of the above technology lies in using 3D modeling and rendering technology, combined with furniture matching algorithms, to achieve real-time placement and preview of furniture in a virtual space, enabling users to intuitively see the matching effects of furniture in an actual home environment, so as to make more informed purchase decisions. The e-commerce platform pre-creates a series of 3D show flats, which cover different home styles and space layouts, such as modern style, Nordic style, etc., as well as typical home environments such as living rooms, bedrooms, and kitchens. In each show flat, various furniture products are pre-placed and adapted according to their styles and functional requirements. These furniture products are available for sale on the platform and cover various household items from sofas, coffee tables to beds, wardrobes, etc.

[0290] Although the above technologies have played a positive role in enhancing user engagement and purchase experience, in practical applications, the provided whole-house designs are pre-made, and users can only purchase corresponding products according to their own needs, unable to achieve customized matching results according to user needs. Therefore, there is still the technical problem of low image generation efficiency.

[0291] Furthermore, the embodiments of the present application provide a furniture product matching generation method based on contrast learning, which is applied to generate a furniture product matching set and a scene matching picture, and can also be used for online furniture matching recommendation. By making full use of 3D showflat data and 2D matching data, an indoor matching model is obtained based on contrast learning, which to a certain extent solves the problems of low efficiency and high cost caused by the long-term dependence on designers for indoor furniture matching. By collecting pictures on the Internet, after preprocessing by a multi-modal model and manual cleaning, the training data of this style can be obtained. Based on contrast learning, the matching method of the style can be learned unsupervised to generate a furniture product matching set. A method for generating a scene matching picture based on the 3D showflat background rendering picture and the 3D model rendering picture is proposed, thus achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0292] The above method of this embodiment will be further introduced below.

[0293] In the embodiments of the present application, to generate a large number of indoor furniture matching results, there must be a product set containing sufficient samples, and there must be potential matchable objects in this product set. The task of the indoor matching model is to discover potential matchable objects and combine them into a set of matching results. Therefore, the quantity and quality of furniture determine the upper limit of the matching results. After having an available furniture product set, what needs to be done is to train a good matching model so that the matching model can identify appropriate object combination relationships according to aesthetics, function, space constraints, and common sense.

[0294] The atomic operation of indoor furniture matching is to judge whether each pair of furniture matches, or to match one furniture with another. Therefore, the indoor matching problem can be completed by the matching of each pair of furniture. Given a product one, there is a product two that matches product one as a positive sample; there is also a product three that does not match product one as a negative sample. Therefore, contrast learning is used to train the matching model. Contrast learning is a self-supervised learning method that learns the feature representation of data by comparing the similarities and differences between samples. Different from related supervised learning, contrast learning does not need to rely on manually labeled tag data, which gives it significant advantages in data expansion and model expansion.

[0295] Similar to image similarity, for the field of furniture matching, a measure of image compatibility can be defined. The measure of image compatibility can be obtained through contrastive learning, and the key lies in the collection and construction of compatible training data. By collecting public data on the Internet and performing cleaning and post-processing, compatible training data can be obtained. Each piece of training data in the compatible training data is a Triplet: Among them, x can be used to represent the anchor point; can be used to represent the positive sample; can be used to represent the negative sample, and each photo is a white-background picture of furniture.

[0296] Figure 10(a) is a schematic diagram of a target furniture according to an embodiment of the present application. As shown in Figure 10(a), this coffee table can be the target furniture, that is, the anchor point. Figure 10(b) is a schematic diagram of a matching furniture with the same style as the target furniture according to an embodiment of the present application. As shown in Figure 10(b), it can be a matching sofa corresponding to the above-mentioned coffee table. Figure 10(c) is a schematic diagram of a non-matching furniture with a style inconsistent with the target furniture according to an embodiment of the present application. As shown in Figure 10(c), it can be a non-matching sofa corresponding to the above-mentioned coffee table.

[0297] Figure 11 is a schematic diagram of a system for determining whether the styles of furniture match according to an embodiment of the present application. As Figure 11 shown, this system may include a cosine distance determination module 1101, a cosine distance determination module 1102, and a hinge loss determination module 1103 (Hinge Loss). Among them, the cosine distance determination module 1101 is used to determine the cosine distance (similarity) d0 between the positive sample and the anchor point. The cosine distance determination module 1102 is used to determine the cosine distance (similarity) d1 between the anchor point and the negative sample. can be used to represent the supervision signal.

[0298] can use to represent the distance between the furniture in two images. Among them, f θ can be a feature extractor. In the embodiment of the present application, given a triplet data, the deep learning model can determine which of the two alternative furniture matches the target furniture.

[0299]

[0300] During the training process, the deep learning model can continuously optimize the feature extractor to obtain an embedding space containing matching-related semantics, so that the two matching furniture are close in the embedding space, and the two non-matching furniture are far apart in the embedding space.

[0301] Optionally, there are various styles in interior collocations, such as the cream style, the wabi-sabi style, the mid-century modern style, etc. To generate furniture product collocations in a specific style, the Triplets of that style (the target furniture of that style, the matching products of that style, and non-matching products) need to be included in the training set. Therefore, it is necessary to construct the training data of that style. However, what exists in the network are all scene pictures of certain styles, and the white-background pictures required for the training of the matching model are very rare. Therefore, it is necessary to construct Triplet training data from the scene pictures.

[0302] Figure 12(a) is a schematic diagram of a scene picture according to an embodiment of the present application. As shown in Figure 12(a), it can be a scene picture inside a cream-style sofa. Figure 12(b) is a schematic diagram of the corresponding white-background picture of a scene picture according to an embodiment of the present application. As shown in Figure 12(b), the white-background picture of the cream-style sofa can be extracted from the field-frequency picture as the Triplet training data.

[0303] Optionally, the data used in the embodiments of the present application comes from publicly available data on the Internet. To obtain publicly available data from the Internet, it is necessary to construct query words. Taking the mid-century modern style and the French retro style as examples, some query words can be constructed to obtain the corresponding pictures from the search engine.

[0304] Table 1 Examples of query words for the mid-century modern style and the French retro style

[0305]

[0306]

[0307] Table 1 is an example of query words for the mid-century modern style and the French retro style in the embodiments of the present application. As shown in Table 1, it shows how to search for query words for relevant pictures of the mid-century modern style and the French retro style from the search engine. By entering the above query words in the search engine, pictures meeting the corresponding style can be queried.

[0308] The data obtained from the search engine according to the query words includes a large number of irrelevant pictures, including non-indoor scenes, advertising texts, mannequins, etc., which affect the quality of the final data set. Therefore, a multi-modal model can be used for preprocessing to roughly screen out the pictures that do not meet the requirements, reducing the workload for subsequent manual selection. What needs to be done in manual review is to select the scene pictures with the correct style and good collocation from the preprocessed pictures. The following are the Prompts when using the multi-modal model for preprocessing:

[0309] 1. Is it an indoor scene?

[0310] 2. Is it a graphic advertisement / promotional picture / social media advertisement?

[0311] 3. Is it a close-up shot of an object on the desktop?

[0312] 4. Does it contain a human body?

[0313] 5. Is it a Collage?

[0314] Check the given picture, answer the above questions, and return yes or no

[0315] Return in JSON format

[0316] {"Is it an indoor scene?": answer, "Is it a graphic advertisement / promotional picture / social media advertisement?": answer, "Is it a close-up shot of an object on the desktop?": answer, "Does it contain a human body?": answer, "Is it a Collage?": answer, "Is there a rotation angle?": answer}

[0317] In this embodiment, a batch of scene pictures with a specific style can be obtained during the collection of the matching style data set. According to the scene pictures, the bbox pictures of each piece of furniture can be obtained through object detection. At this time, a set of white background pictures of furniture is required, and this set should have enough white background pictures so that according to the bbox pictures, consistent or similar white background pictures can be retrieved through picture similarity search.

[0318] Figure 13 It is a schematic diagram of a method for constructing a training data set according to an embodiment of the present application, as Figure 13 shown, the method may include the following steps:

[0319] Step S1301, perform object detection on the scene images of a specific style.

[0320] In this embodiment, object detection aims to identify and locate multiple object instances in the image, that is, to detect what objects are contained in the image and the positions of these objects in the image. In the embodiment of the present application, object detection is used to process scene images of a specific style, which is the first step in the process of constructing a training data set. Specifically, the object detection algorithm will identify various furniture and decorations in these scene images, such as sofas, coffee tables, lamps, etc., and give the bounding box (bbox) of each object.

[0321] The key to the above steps lies in selecting or training a sufficiently accurate object detection model to ensure that it can effectively identify the furniture and other objects in the scene image. If the matching model is not precise enough, it may lead to furniture being misidentified or missed, thereby affecting the quality of the subsequent training data. Therefore, the accuracy and generalization ability of the matching model are the key points in this stage, and additional training and optimization can be carried out for diverse scene images and furniture of different styles.

[0322] Step S1302, perform picture similarity search.

[0323] In this embodiment, the picture similarity retrieval is carried out after object detection, aiming to find white-background pictures or other types of standardized pictures similar to the target objects in the scene image from a large picture library. The above steps are crucial for constructing Triplet data because in contrastive learning, a set of data containing the target furniture, matching furniture (positive samples), and non-matching furniture (negative samples) is required for model training. By retrieving the similarity, it can be ensured that the positive samples and the target furniture are visually coordinated, while the negative samples are different from the target furniture in terms of style, color, shape, etc.

[0324] Picture similarity retrieval is usually based on deep learning feature extraction. A pre-trained deep model is used to extract feature vectors from each picture, and then these vectors are compared in the feature space to determine the similarity between pictures.

[0325] It is worth noting that in order to construct high-quality training data, picture similarity retrieval not only needs to consider the similarity of visual appearance, but also can consider factors such as the functionality of objects and style consistency. This means that the retrieval algorithm may need to combine multiple information sources, including but not limited to image features, text descriptions, user feedback, etc., to comprehensively evaluate the similarity degree of pictures.

[0326] In summary, the above method identifies furniture from scene images of a specific style through object detection, and constructs Triplet data for training through picture similarity retrieval. The two steps work together to provide the basic data support for the furniture product matching generation method based on contrastive learning, and are key components of the entire training process.

[0327] Optionally, a white-background picture can be retrieved for each object included in each scene picture. The white-background pictures within the same scene are considered matching white-background pictures, and the white-background pictures in different scenes are considered non-matching white-background pictures. Therefore, by sampling two white-background pictures from one scene picture respectively as, and sampling one white-background picture from another scene picture as, the Triplet required for training can be constructed.

[0328] In this embodiment, after the indoor matching model is trained, the feature extractor in the model can extract vectors related to matching from the target furniture pictures. According to this vector, vector retrieval is performed on the furniture product set in the embedding space, and the furniture corresponding to the nearest vector obtained is the one that matches the target furniture. By continuously matching other furniture with the furniture in the matching set, the indoor furniture matching task can be completed step by step.

[0329] Figure 14(a) is a schematic diagram of an example of new Chinese style seed furniture according to an embodiment of the present application. As shown in Figure 14(a), it can be a new Chinese style sofa, and the cushion of this sofa is a one-piece type. Figure 14(b) is a schematic diagram of another example of new Chinese style seed furniture according to an embodiment of the present application. As shown in Figure 14(b), it can be another new Chinese style sofa, and the cushion of this sofa is a two-piece type and includes three cushions. Figure 14(c) is a schematic diagram of another example of new Chinese style seed furniture according to an embodiment of the present application. As shown in Figure 14(c), it can be another new Chinese style sofa, and the two cushions of this sofa are separated by a wooden cabinet.

[0330] Figure 15(a) is a schematic diagram of an example of modern style seed furniture according to an embodiment of the present application. As shown in Figure 15(a), it can be a modern style sofa, and this sofa is vertical and consists of three-piece individual small cushions and a chaise longue. Figure 15(b) is a schematic diagram of another example of modern style seed furniture according to an embodiment of the present application. As shown in Figure 15(b), it can be another modern style sofa, and the cushion of this sofa is a three-piece type and corresponds to three cushions. Figure 15(c) is a schematic diagram of another example of modern style seed furniture according to an embodiment of the present application. As shown in Figure 15(c), it can be another modern style sofa, and the cushion of this sofa is a two-piece type and corresponds to three cushions.

[0331] In the embodiment of the present application, the combination containing 0 furniture is named an empty combination. When starting from an empty combination, there is no target furniture required for the combination. Therefore, the white background picture of the furniture in the training data can be used as the seed furniture or the seed furniture of this style can be manually configured. Starting from the seed furniture, first retrieve a piece of furniture and put it into the combination set, and then match other furniture with the furniture in the combination set. Therefore, the retrieval schemes generated by the combination set can be roughly divided into the following two types: retrieval according to the main furniture and random selection of furniture for retrieval. In this way, different combination sets can be generated by selecting different seed furniture.

[0332] Figure 16(a) is a schematic diagram of a main furniture retrieval method according to an embodiment of the present application. As shown in Figure 16(a), the main furniture retrieval method can be a method of retrieving various other furniture that matches its style from one main furniture. For example, the picture of the main furniture can be initialized, that is, from the picture containing the above-mentioned main furniture, the white background picture of the main furniture is detected by target detection, and the white background picture is retrieved for compatibility to obtain other furniture that matches it, such as chairs, coffee tables, murals, stools, etc.

[0333] FIG. 16(b) is a schematic diagram of a random furniture retrieval method according to an embodiment of the present application. As shown in FIG. 16(b), the random furniture retrieval method can be a method of starting from a main piece of furniture and then retrieving the remaining furniture by the other furniture retrieved from the main piece of furniture. For example, the picture of the main piece of furniture can be initialized, that is, the white-background picture of the main piece of furniture can be detected from the picture containing the above-mentioned main piece of furniture. And perform a matching retrieval on the white-background picture to retrieve a matching stool. The stool can be further subjected to a matching retrieval to obtain matching sofa seats, murals, and so on.

[0334] Figure 17 is a schematic diagram of rendering a set of rendering images into a sample room according to an embodiment of the present application. As Figure 17 shown, what is generated by the matching set is a set of furniture, but the real matching effect cannot be seen in the scene, and it is also very difficult to imagine the real matching effect based on a set. However, the scene-based expression of furniture can give users of e-commerce platforms a more immersive experience. Therefore, assembling the retrieved matching set into the scene is an important ability. For example, for the living room, the set of living room rendering images to be rendered may include furniture such as table lamps, sofas, coffee tables, and armchairs. For the bedroom, the set of bedroom living room collections to be rendered may include furniture such as wardrobes, windows, table lamps, and stools.

[0335] Figure 18 is a schematic diagram of obtaining a matching result by superimposing a set of rendering images and a sample room according to an embodiment of the present application. As Figure 18 shown, to generate a scene matching picture, some preconditions are required: a 3D sample room and 3D models. Based on 3D sample rooms of different styles, hard decoration backgrounds of different styles and a set of rendering layers of 3D models in the sample room can be rendered. The set of furniture products mentioned above is the set of rendering layers of 3D models in the sample room here. After superimposing the hard decoration background and the layers in the rendering image matching set, a scene matching picture can be generated. Since hanging pictures and carpets are 2D objects in the scene, the white-background pictures are directly used for superimposition. It should be noted that when superimposing the carpet, the superimposition position needs to be calculated and perspective projection is performed.

[0336] In the embodiments of the present application, if it is necessary to match the display object in a certain target image, the target image can be obtained, and the image processing model is used to analyze the style of the display object in the target image to obtain a matching strategy that matches the style. And according to the matching strategy, at least one matching image corresponding to the target image is determined. The spatial position relationship between the display object and the matching object in the matching image can be determined, so that according to the spatial position relationship, the corresponding matching image and the target image are rendered into the initial scene image including the display scene to obtain the final target scene image. In the embodiments of the present application, by introducing an image processing model for automated style analysis, rapid retrieval of matching images, automatic determination of spatial position relationships, and optimized image rendering and synthesis technologies, it is possible to generate a large number of stylized, personalized, and high-quality matching images in a short time, greatly promoting the digital process of the matching design industry, improving the user experience and business operation efficiency, thereby achieving the technical effect of improving the generation efficiency of images and solving the technical problem of low efficiency in image generation.

[0337] According to an embodiment of the present application, there is also provided an image generation device for implementing the above Figure 2 image generation method shown.

[0338] Figure 19 is a schematic diagram of an image generation device according to an embodiment of the present application, as Figure 19 shown, the image generation device 1900 may include: a first acquisition unit 1902, a first determination unit 1904, a second determination unit 1906, and a first rendering unit 1908.

[0339] The first acquisition unit 1902 is configured to acquire a target image.

[0340] The first determination unit 1904 is configured to analyze the style feature information of the target image by using an image processing model to obtain at least one matching image.

[0341] The second determination unit 1906 is configured to determine the spatial position relationship between the display object and the matching object.

[0342] The first rendering unit 1908 is configured to render the matching image and the target image into the target scene image according to the spatial position relationship.

[0343] Here, the first acquisition unit 1902, the first determination unit 1904, the second determination unit 1906, and the first rendering unit 1908 correspond to steps S202 to S208 in the above embodiments. The instances and application scenarios implemented by the four units are the same as those of the corresponding steps, but are not limited to the content disclosed in the above embodiments. It should be noted that the above units may be hardware components or software components stored in a memory (for example, memory 2404) and processed by one or more processors (for example, processors 2402a, 2402b,..., 2402n). The above units may also be part of a device and can run in the computer terminal A provided in the following embodiments.

[0344] According to an embodiment of the present application, there is also provided an apparatus for implementing the Figure 3 determination method of the image processing model shown above for determining an image processing model.

[0345] Figure 20 is a schematic diagram of an apparatus for determining an image processing model according to an embodiment of the present application, as Figure 20 shown. The apparatus 2000 for determining the image processing model may include: a second acquisition unit 2002 and a contrast learning unit 2004.

[0346] The second acquisition unit 2002 is configured to acquire an image sample set.

[0347] The contrast learning unit 2004 is configured to perform contrast learning on a deep learning model by using the image sample set to obtain an image processing model.

[0348] Here, it should be noted that the above second acquisition unit 2002 and contrast learning unit 2004 correspond to steps S302 to S304 in the above embodiments. The instances and application scenarios implemented by the two units are the same as those of the corresponding steps, but are not limited to the content disclosed in the above embodiments. It should be noted that the above units may be hardware components or software components stored in a memory (for example, memory 2404) and processed by one or more processors (for example, processors 2402a, 2402b,..., 2402n). The above units may also be part of a device and can run in the computer terminal A provided in the following embodiments.

[0349] According to an embodiment of the present application, there is also provided an apparatus for implementing the Figure 4 image generation method shown above for generating an image.

[0350] Figure 21 is a schematic diagram of another apparatus for generating an image according to an embodiment of the present application, as Figure 21As shown in the figure, the image generation device 2100 may include: an identification unit 2102, a third determination unit 2104, a fourth determination unit 2106, a second rendering unit 2108, and a return unit 2110.

[0351] The identification unit 2102 is configured to identify an image of a product object to be paired from a product matching platform.

[0352] The third determination unit 2104 is configured to analyze the style feature information of the product object image by using a product matching model to obtain at least one paired image.

[0353] The fourth determination unit 2106 is configured to determine the spatial position relationship between the product object and the paired object.

[0354] The second rendering unit 2108 is configured to render the paired image and the product object image to a target scene image according to the spatial position relationship.

[0355] The return unit 2110 is configured to return the rendered target scene image to the product matching platform.

[0356] Here, the above-mentioned identification unit 2102, third determination unit 2104, fourth determination unit 2106, second rendering unit 2108, and return unit 2110 correspond to steps S402 to S410 in the above embodiment. The instances and application scenarios implemented by the five units and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above units may be hardware components or software components stored in a memory (for example, memory 2404) and processed by one or more processors (for example, processors 2402a, 2402b..., 2402n). The above units may also be part of a device and can run in the computer terminal A provided in the following embodiment.

[0357] According to an embodiment of the present application, there is also provided an image generation device for implementing the above Figure 5 shown image generation method.

[0358] Figure 22 is a schematic diagram of another image generation device according to an embodiment of the present application. As Figure 22 shown, the image generation device 2200 may include: a first display unit 2202 and a second display unit 2204.

[0359] The first display unit 2202 is configured to display a target image on the operation interface in response to an input operation on the operation interface.

[0360] A second display unit 2204, configured to respond to an image generation instruction acting on the operation interface and display a rendered target scene image on the operation interface.

[0361] Here, the first display unit 2202 and the second display unit 2204 correspond to steps S502 to S504 in the above embodiment. The two units have the same implemented examples and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above embodiment. It should be noted that the above units may be hardware components or software components stored in a memory (for example, memory 2404) and processed by one or more processors (for example, processors 2402a, 2402b,..., 2402n). The above units may also be part of a device and can run in the computer terminal A provided in the following embodiment.

[0362] According to an embodiment of the present application, there is also provided an image generation device for implementing the above Figure 6 shown image generation method.

[0363] Figure 23 is a schematic diagram of another image generation device according to an embodiment of the present application. As Figure 23 shown, the image generation device 2300 may include: a third acquisition unit 2302, a fifth determination unit 2304, a sixth determination unit 2306, a third rendering unit 2308, and an output unit 2310.

[0364] The third acquisition unit 2302 is configured to acquire a target image by invoking a first interface.

[0365] The fifth determination unit 2304 is configured to analyze the style feature information of the target image by using an image processing model to obtain at least one matching image.

[0366] The sixth determination unit 2306 is configured to determine the spatial position relationship between the display object and the matching object.

[0367] The third rendering unit 2308 is configured to render the matching image and the target image into the target scene image according to the spatial position relationship.

[0368] The output unit 2310 is configured to output the rendered target scene image by invoking a second interface.

[0369] Here, the above-mentioned third acquisition unit 2302, fifth determination unit 2304, sixth determination unit 2306, third rendering unit 2308, and output unit 2310 correspond to steps S602 to S610 in the above-mentioned embodiment. The functions of the five units are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the content disclosed in the above-mentioned embodiment. It should be noted that the above units may be hardware components or software components stored in a memory (for example, memory 2404) and processed by one or more processors (for example, processors 2402a, 2402b,..., 2402n). The above units may also be part of a device and can run in the computer terminal A provided in the following embodiment.

[0370] In the image generation device, if it is necessary to match the display object in a certain target image, the target image can be obtained, and the image processing model can be used to analyze the style of the display object in the target image to obtain the target number of matching images that match the style. The spatial position relationship between the display object and the matching object in the matching image can be determined, so that the corresponding matching image and the target image can be rendered into the target scene image including the display scene according to the spatial position relationship. In the embodiment of the present application, by introducing an image processing model for automated style analysis, rapid retrieval of matching images, automatic determination of spatial position relationships, and optimized image rendering and synthesis technologies, it is possible to generate a large number of stylized, personalized, and high-quality matching images in a short time, greatly promoting the digital process of the matching design industry, improving the user experience and business operation efficiency, thereby achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0371] Embodiments of the present application can provide a computer terminal, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0372] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.

[0373] In this embodiment, the above computer terminal can execute the program code of the above steps in the image generation method.

[0374] Optionally, Figure 24 is a structural block diagram of a computer terminal according to an embodiment of the present application. As Figure 24 shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 2402, a memory 2404, and a transmission device 2406.

[0375] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image generation method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned image generation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and the above-mentioned remote memory can be connected to the computer terminal A through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0376] By adopting the embodiments of the present application, an image generation method is provided. If it is necessary to match the display object in a certain target image, the target image can be obtained, and the image processing model is used to analyze the style to which the display object in the target image belongs, and obtain the target number of matching images that match the style. The spatial position relationship between the display object and the matching object in the matching image can be determined, and then, according to the spatial position relationship, the corresponding matching image and the target image are rendered into the target scene image including the display scene. In the embodiments of the present application, by introducing an image processing model for automated style analysis, rapid retrieval of matching images, automatic determination of spatial position relationships, and optimized image rendering and synthesis technologies, it is possible to generate a large number of stylized, personalized, and high-quality matching images in a short time, greatly promoting the digital process of the matching design industry, improving the user experience and business operation efficiency, thus achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0377] Those of ordinary skill in the art can understand that Figure 24 the structure shown is only schematic, and the computer terminal A can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (abbreviated as MID), a PAD and other terminal devices. Figure 24 It does not limit the structure of the above computer terminal A. For example, the computer terminal A may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 24 in the figure, or have a different configuration from that shown Figure 24 in the figure.

[0378] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disk, etc.

[0379] An embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium can be used to store the program code executed by the image generation method provided in the first embodiment above.

[0380] Optionally, in this embodiment, the above computer-readable storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.

[0381] Optionally, in this embodiment, the computer-readable storage medium is configured to store the program code for executing the steps in the above web resource loading method.

[0382] An embodiment of the present application can provide an electronic device, and the electronic device can include a memory and a processor.

[0383] Figure 25 is a block diagram of an electronic device for an image generation method according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0384] As Figure 25 shown, the device 2500 includes a computing unit 2501, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 2502 or the computer program loaded from the storage unit 2508 into the random access memory (RAM) 2503. In the RAM 2503, various programs and data required for the operation of the device 2500 can also be stored. The computing unit 2501, the ROM 2502, and the RAM 2503 are connected to each other through a bus 2504. The input / output (I / O) interface 2505 is also connected to the bus 2504.

[0385] Multiple components in device 2500 are connected to I / O interface 2505, including: an input unit 2506, such as a keyboard, a mouse, etc.; an output unit 2504, such as various types of displays, speakers, etc.; a storage unit 2508, such as a magnetic disk, an optical disc, etc.; and a communication unit 2509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 2509 allows device 2500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0386] The computing unit 2501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 2501 include but are not limited to a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various dedicated Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 2501 executes the various methods and processes described above, such as the method for generating an image. For example, in some embodiments, the method for generating an image can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 2508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 2500 via the ROM 2502 and / or the communication unit 2509. When the computer program is loaded into the RAM 2503 and executed by the computing unit 2501, one or more steps of the method for generating an image described above can be executed. Alternatively, in other embodiments, the computing unit 2501 can be configured to execute the method for generating an image in any other suitable manner (e.g., by means of firmware).

[0387] Embodiments of the present application also provide a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and when the computer program is executed by a processor, it implements the method for generating an image in the above embodiments of the present application.

[0388] According to an embodiment of the present application, a method for generating an image is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0389] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 26 It is a hardware block diagram of a computer terminal (or mobile device) for implementing an image generation method according to an embodiment of the present application. As Figure 26 shown, the computer terminal 260 (or mobile device) may include one or more processors 2602 (shown as 2602a, 2602b,..., 2602n in the figure) (the processor 2602 may include, but is not limited to, a processing device such as a microcontroller unit (MCU) or a field programmable gate array (FPGA)), a memory 2604 for storing data, and a transmission device 2606 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 26 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 260 may further include more or fewer components than Figure 26 shown in the figure, or have a different configuration from Figure 26 shown in the figure.

[0390] Figure 26 The shown hardware block diagram can be used not only as an exemplary block diagram of the above computer terminal 260 (or mobile device), but also as an exemplary block diagram of the above server. In an alternative embodiment, Figure 26 it shows in a block diagram an embodiment of using the above Figure 26 shown computer terminal 260 (or mobile device) as a computing node in the computing environment 2601.

[0391] Figure 27 It is a block diagram of the structure of a computing environment for an image generation method according to an embodiment of the present application. As Figure 27As shown, the computing environment 2701 includes multiple computing nodes (such as servers, shown as 2710-1, 2710-2, … in the figure) running on a distributed network. Each computing node includes local processing and memory resources, and end users 2702 can remotely run applications or store data in the computing environment 2701. Applications can be provided as multiple services 2720-1, 2720-2, 2720-3, and 2720-4 in the computing environment 2701, representing services "F", "G", "I", and "H" respectively.

[0392] End users 2702 can provide and access services through a web browser or other software applications on the client side. In some embodiments, the provisioning and / or requests of end users 2702 can be provided to the ingress gateway 2730. The ingress gateway 2730 can include a corresponding proxy to handle the provisioning and / or requests for services (one or more services provided in the computing environment 2701).

[0393] Services are provided or deployed according to various virtualization technologies supported by the computing environment 2701. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar means. VM-based virtualization can be achieved by initializing virtual machines to simulate real computers and execute programs and applications without directly accessing any actual hardware resources. While virtualizing machines with VMs, according to container-based virtualization, containers can be launched to virtualize the entire operating system so that multiple workloads can run on a single operating system instance.

[0394] In one embodiment of container-based virtualization, several containers of a service can be assembled into a Pod (e.g., Kubernetes Pod). For example, as Figure 27 shown, service 2720-2 can be equipped with one or more Pods 2740-1, 2740-2, …, 2740-N (collectively referred to as Pods). A Pod can include a proxy 2745 and one or more containers 2742-1, 2742-2, …, 2742-M (collectively referred to as containers). One or more containers in the Pod handle requests related to one or more corresponding functions of the service, and the proxy 2745 generally controls network functions related to the service, such as routing, load balancing, etc. Other services can also be equipped with Pods similar to this one.

[0395] During operation, executing user requests from end users 2702 may require invoking one or more services in the computing environment 2701, and executing one or more functions of a service may require invoking one or more functions of another service. As Figure 27As shown, service "F" 2720-1 receives a user request from end-user 2702 at ingress gateway 2730. Service "F" 2720-1 may invoke service "G" 2720-2, and service "G" 2720-2 may request service "I" 2720-3 to perform one or more functions.

[0396] The computing environment described above may be a cloud computing environment where the allocation of resources is managed by a cloud service provider, allowing the development of functions without the need to analyze implementation, adjust, or scale servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be split into groups of functions that can scale automatically and independently, rather than scaling a single hardware device to handle potential loads.

[0397] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system-on-a-chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. The various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0398] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. The above program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on a remote machine or server.

[0399] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0400] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD), monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0401] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0402] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server that incorporates blockchain.

[0403] It should be noted that the serial numbers of the embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0404] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0405] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0406] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0407] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0408] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories, random access memories, mobile hard disks, magnetic disks, or optical discs.

[0409] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and the above improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for generating an image, characterized in that, Including: Obtain a target image, where the image content of the target image includes at least one display object; Analyze the style feature information of the target image by using an image processing model to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one of the attribute features; Determine the spatial position relationship between the display object and the matching object; Render the matching image and the target image into a target scene image according to the spatial position relationship.

2. The method according to claim 1, wherein Analyze the style feature information of the target image by using an image processing model to obtain at least one matching image, including: Analyze the style feature information of the target image by using the image processing model, and determine a target candidate image that meets the style feature information of the target image from multiple candidate images, where the image content of the candidate image includes a candidate object; Use the image processing model to determine the target candidate image whose attribute feature matches the attribute feature of the display object as the matching image, where the matching object is the candidate object whose attribute feature matches the attribute feature of the display object.

3. The method according to claim 2, characterized in that Analyze the style feature information of the target image by using the image processing model, and determine a target candidate image that meets the style feature information of the target image from multiple candidate images, including: Use the image processing model to determine the similarity between the style feature information of the target image and the style feature information of the candidate image, where the style feature information of the candidate image is used to represent at least one attribute feature of the candidate object; Determine the candidate image with the similarity greater than the similarity threshold as the target candidate image.

4. The method according to claim 3, characterized in that, The image processing model includes a feature comparison model, and the style feature information of the target image includes a first style feature vector, which is used to represent the style semantic feature of the corresponding display object in the embedding space. Using the image processing model to determine the similarity between the style feature information of the target image and the style feature information of the candidate image includes: Obtain a second style feature vector of the candidate image, where the second style feature vector is used to represent the style semantic feature of the corresponding candidate object in the embedding space; Use the feature comparison model to determine the similarity between the first style feature vector and the second style feature vector.

5. The method according to claim 2, wherein The attribute features of the display object include at least one of the following: the color attribute, size attribute, and function attribute of the display object. Using the image processing model to determine the target candidate image whose attribute feature matches the attribute feature of the display object as the matching image, including: Using the image processing model, determine the target candidate image corresponding to a candidate object whose similarity of color attributes to the color attributes of the display object is greater than a similarity threshold, and / or whose size attributes match the size attributes of the display object, and / or whose function attributes match the function attributes of the display object, as the matching image.

6. The method according to claim 1, wherein Analyze the style feature information of the target image using an image processing model to obtain at least one matching image, including: Obtain foreground information from the target image, where the foreground information includes the display object; Use the image processing model to analyze the foreground information to obtain the style feature information of the target image; Use the image processing model to determine the matching image based on the style feature information of the target image.

7. The method according to claim 6, characterized in that, The image processing model includes a feature extraction model, and the style feature information of the target image includes a first style feature vector, which is used to represent the style semantic features of the corresponding display object in the embedding space. Using the image processing model to analyze the foreground information to obtain the style feature information of the target image includes: Use the feature extraction model to identify the style semantic features of the display object from the foreground information; Use the feature extraction model to map the style semantic features of the display object into the embedding space to obtain the first style feature vector.

8. The method according to claim 1, characterized in that, Analyze the style feature information of the target image using an image processing model to obtain at least one matching image, including: Determine a generation strategy corresponding to the target image based on the style feature information of the target image, where the generation strategy is used to represent the rule for comparing the target image and candidate images to generate the target number of matching images, and the image content of the candidate images includes candidate objects; According to the generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate images to obtain the target number of matching images, where the matching object is a candidate object whose style feature information matches the style of the target image.

9. The method according to claim 8, wherein The generation strategy includes a first generation strategy, which is used to represent the rule for comparing the target image and the candidate images, and for comparing the matching images and the candidate images to generate the target number of matching images. According to the generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate images to obtain the target number of matching images, including: According to the first generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate images to obtain a first number of matching images, where the first number is less than the target number; Control the image processing model to compare the style feature information of the first quantity of the collocation images with the style feature information of the candidate image, and obtain the second quantity of the collocation images, where the second quantity is less than the target quantity; Based on the first quantity of the collocation images and the second quantity of the collocation images, determine the target quantity of the collocation images.

10. The method according to claim 9, characterized in that, Based on the first quantity of the collocation images and the second quantity of the collocation images, determining the target quantity of the collocation images includes: In response to the sum of the first quantity and the second quantity being less than the target quantity, determine the second quantity as the first quantity, and return to execute the following steps until the sum of the first quantity and the second quantity is equal to the target quantity. Determine the first quantity of the collocation images and the second quantity of the collocation images as the target quantity of the collocation images: Control the image processing model to compare the style feature information of the first quantity of the collocation images with the style feature information of the candidate image, and obtain the second quantity of the collocation images.

11. The method according to claim 8, wherein The generation strategy includes a second generation strategy, and the second generation strategy is used to represent the rule for comparing the target image and the candidate image to generate the target quantity of the collocation images. According to the generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate image, and obtain the target quantity of the collocation images, including: According to the second generation strategy, control the image processing model to compare the style feature information of the target image with the style feature information of the candidate image, and obtain a target candidate image, where the similarity between the style feature information of the target candidate image and the style feature information of the target image is greater than the similarity threshold; Use the image processing model to determine the target candidate images whose attribute features of the target quantity match the attribute features of the display object as the target quantity of the collocation images.

12. The method according to claim 1, wherein The image content of the target scene image includes the display scene for displaying the display object and the collocation object. According to the spatial position relationship, rendering the collocation image and the target image into the target scene image includes: Obtain the three-dimensional object models corresponding to the target image and the collocation image respectively, and the three-dimensional scene model corresponding to the display scene; According to the spatial position relationship, render the three-dimensional object model into the three-dimensional scene model to obtain a rendered three-dimensional scene model; Based on the rendered three-dimensional scene model, determine the rendered target scene image.

13. The method according to claim 12, wherein According to the spatial position relationship, rendering the three-dimensional object model into the three-dimensional scene model to obtain a rendered three-dimensional scene model includes: In response to the image content of both the target image and the matching image including a target object, perform a deformation process on the three-dimensional object model corresponding to the target object, wherein, in any dimension in the three-dimensional space, the size of the target object is smaller than a size threshold; Adjust the spatial position relationship corresponding to the deformed three-dimensional object model to obtain an adjusted spatial position relationship, wherein the spatial position relationship corresponding to the deformed three-dimensional object model is the spatial position relationship between the three-dimensional object models corresponding to the display object and the matching object other than the three-dimensional object model corresponding to the target object in the three-dimensional object model; Render the deformed three-dimensional object model into the three-dimensional scene model according to the adjusted spatial position relationship.

14. The method according to any one of claims 1 to 13, characterized in that, The image content of the target scene image includes a display scene for displaying the display object and the matching object. Determining the spatial position relationship between the display object and the matching object includes: Obtain the attribute characteristics of both the display object and the matching object, and the layout strategy of the display scene, wherein the attribute characteristics at least include one of the following: the functional attributes of the display object and the matching object, and the size attributes of the display object and the matching object in the three-dimensional space. The layout strategy is used to represent the rules for laying out the display object and the matching object in the display scene that matches the style feature information of the display object; Based on the layout strategy, the attribute characteristics, and the layout information of the target scene image, determine the spatial position relationship, wherein the layout information is used to represent the structure and size of the display scene in the three-dimensional space.

15. A method for determining an image processing model, characterized in that Include: Obtain an image sample set, wherein the image sample set includes a target image sample set and a candidate image sample set. The image content of the target image sample includes at least one display object sample, and the image content of the candidate image sample set includes at least one candidate object sample; Use the image sample set to perform contrastive learning on a deep learning model to obtain an image processing model, wherein the image processing model is used to analyze the style feature information of the target image to obtain at least one matching image. The image content of the target image includes at least one display object, and the style feature information of the target image is used to represent the attribute characteristics of the display object. The image content of the matching image includes a matching object that matches at least one of the attribute characteristics. The matching image and the target image are used to be rendered into a target scene image based on the spatial position relationship between the display object and the matching object.

16. The method according to claim 15, wherein Use the image sample set to perform contrastive learning on a deep learning model to obtain an image processing model, including: Use the deep learning model to determine a matching image sample set from the candidate image sample set, wherein the image content of the matching image sample set includes at least one matching object sample, and the matching object sample is a candidate object sample whose attribute characteristics match the attribute characteristics of the display object sample; The deep learning model is subjected to contrastive learning using at least one of the display object samples, the corresponding collocation object samples, and the candidate object samples other than the collocation object samples among the at least one candidate object samples, to obtain the image processing model.

17. The method according to claim 16, wherein Using the deep learning model, a collocation image sample set is determined from the candidate image sample set, including: The deep learning model is used to analyze the image sample set to obtain a first style feature vector sample of the display object sample and a second style feature vector sample of the candidate object sample, where the first style feature vector sample is used to represent the style semantic features of the display object sample in the embedding space, and the second style feature vector sample is used to represent the style semantic features of the candidate object sample in the embedding space; The deep learning model is used to compare the first style feature vector sample and the second style feature vector sample to obtain a comparison result; The candidate image sample set corresponding to the candidate object sample whose comparison result is that the similarity between the first style feature vector sample and the second style feature vector sample is greater than the similarity threshold is determined as the collocation image sample set.

18. The method according to claim 15, wherein An image sample set is obtained, including: Using a search prompt, a first initial image sample set is obtained, where the search prompt is at least used to describe the features of the image content of the initial image samples in the first initial image sample set; Using the prompt information, the first initial image sample set is screened to obtain a screened first initial image sample set, where the prompt information is used to determine whether the initial image samples in the first initial image sample set can be used to train the deep learning model; At least one initial image sample with an image quality higher than the image quality threshold in the screened first initial image sample set is determined as the second initial image sample set; A foreground information sample set is extracted from the second initial image sample set, and the foreground information sample set is determined as the image sample set, where the image content of the foreground information sample set includes the display object sample and the candidate object sample.

19. A method for generating an image, characterized in that, Including: Identifying a product object image to be collocated on a product collocation platform, where the image content of the product object image includes at least one product object; Using a product collocation model to analyze the style feature information of the product object image to obtain at least one collocation image, where the style feature information of the product object image is used to represent at least one attribute feature of the product object, and the image content of the collocation image includes a collocation object matching at least one of the attribute features; Determining the spatial position relationship between the product object and the collocation object; Rendering the collocation image and the product object image to a target scene image according to the spatial position relationship; Returning the rendered target scene image to the product collocation platform.

20. A method for generating an image, characterized in that, Including: Responding to an input operation on the operation interface, a target image is displayed on the operation interface, where the image content of the target image includes at least one display object; In response to an image generation instruction acting on the operation interface, a rendered target scene image is displayed on the operation interface, where the rendered target scene image is obtained by rendering a matching image and the target image into the target scene image according to the spatial position relationship between the display object and the matching object, the matching image is obtained by analyzing the style feature information of the target image using an image processing model, the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one of the attribute features.

21. A method for generating an image, characterized in that, Comprising: By calling a first interface, a target image is obtained, where the first interface includes a first parameter, and the parameter value of the first parameter includes the target image, and the image content of the target image includes at least one display object; Using an image processing model to analyze the style feature information of the target image to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one of the attribute features; Determine the spatial position relationship between the display object and the matching object; According to the spatial position relationship, render the matching image and the target image into the target scene image; Output the rendered target scene image by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the rendered target scene image.

22. An image generation system, characterized in that, Comprising: A client for uploading a target image, where the image content of the target image includes at least one display object; A server for using an image processing model to analyze the style feature information of the target image to obtain at least one matching image, where the style feature information of the target image is used to represent at least one attribute feature of the display object, and the image content of the matching image includes a matching object that matches at least one of the attribute features; determining the spatial position relationship between the display object and the matching object; rendering the matching image and the target image into the target scene image according to the spatial position relationship; Wherein, the client is used to display the rendered target scene image.

23. An electronic device, characterized in that, Comprising: A memory storing an executable program; A processor connected to the memory through a bus for running the program, where when the program runs, it executes the method according to any one of claims 1 to 21.

24. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 21.

25. A computer program product, characterized in that, Comprising a computer program, where when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 21.

Citation Information

Cited By

  • Depth data automatic acquisition method and device, electronic equipment and storage medium

    CN121280504A