Image processing method and device, image display method and device, equipment and storage medium

Through a multi-stage quality inspection and evaluation mechanism, combined with automated objective inspection and user evaluation data, high-quality AI-generated catering images are screened out, solving the problem of unstable quality of AI-generated images and improving user experience and platform conversion rate.

CN120765355APending Publication Date: 2025-10-10RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511144569.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In existing technologies, the quality of AI-generated images of food dishes in the catering industry is difficult to guarantee stably, which affects their usability and credibility in e-commerce and advertising.

Method used

By introducing a multi-stage quality inspection and evaluation mechanism, combining automated objective quality inspection with user behavior data and subjective evaluation data, high-quality target product images are screened out.

Benefits of technology

It significantly improves the overall quality and usability of AI-generated product images, enhances the user visual experience, and improves product display effects and platform conversion rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765355A_ABST
    Figure CN120765355A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, an image display method and device, equipment and a storage medium. The image processing method is applied to the server, and comprises the following steps: obtaining candidate commodity images; evaluating the candidate commodity image according to a preset detection dimension to obtain detection information corresponding to the candidate commodity image; under the condition that the detection information indicates that the candidate commodity image meets a first quality condition corresponding to the preset detection dimension, obtaining user evaluation information corresponding to the candidate commodity image; under the condition that the user evaluation information indicates that the candidate commodity image meets a second quality condition, determining a target commodity image according to the candidate commodity image; and sending the target commodity image to the user terminal, so that the user terminal displays the target commodity image. By adopting the image processing method, the quality of the commodity image can be improved, and the visual experience of a user is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, an image display method, an apparatus, a device and a storage medium. Background Art

[0002] With the rapid development of AI-generated content (AIGC) technology, its application in the catering industry is becoming increasingly widespread, particularly in the generation of dish images. Compared to traditional manual photography methods, AI-generated imagery offers significant advantages in efficiency and cost, enabling rapid production of large-scale image content.

[0003] However, in related technologies, the quality of AI-generated images is still difficult to guarantee stably. Therefore, there is an urgent need to improve the quality of AI-generated images to enhance their usability and credibility in practical applications such as catering e-commerce and advertising communication. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, image display method, apparatus, device, and storage medium that can effectively improve the quality of product images and enhance the visual experience. The above technical solutions are as follows: In a first aspect, an embodiment of the present application provides an image processing method, applied to a server, comprising: Obtain candidate product images; Detect the candidate product image according to the preset detection dimension to obtain detection information corresponding to the candidate product image. The detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension; If the detection information indicates that the candidate product image meets a first quality condition corresponding to a preset detection dimension, obtaining user evaluation information corresponding to the candidate product image, where the first quality condition is that quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; If the user evaluation information indicates that the candidate product image meets a second quality condition, determining a target product image based on the candidate product image, where the second quality condition is used to measure the content performance of the candidate product image in user interaction; The target product image is sent to the user terminal so that the user terminal displays the target product image.

[0005] In one possible implementation, obtaining a candidate product image includes: Obtain preset product information, which includes the product name and the corresponding product demand weight; Determine whether the product demand weight is greater than the preset weight threshold; When the product demand weight is greater than a preset weight threshold, the product name is input into a first preset model, and the output image of the first preset model is determined as a candidate product image; Among them, the first preset model is a text-driven image generation model.

[0006] In a possible implementation, the preset product information also includes an initial product image; Obtaining candidate product images also includes: If the product demand weight is less than or equal to a preset weight threshold, inputting the product name and the initial product image into a second preset model, and determining the output image of the second preset model as a candidate product image; Among them, the second preset model is an image generation model based on joint input of images and text.

[0007] In a possible implementation, the first quality condition includes quality targets corresponding to each preset detection dimension, and the first quality condition is that the quality targets corresponding to each preset detection dimension are all satisfied; Detect the candidate product image according to the preset detection dimension and obtain the detection information corresponding to the candidate product image, including: Detect the candidate product images according to each preset detection dimension, and obtain the detection results corresponding to the candidate product images under each preset detection dimension; Based on the test results, determine whether the candidate product image meets the quality targets of each preset test dimension; If yes, generating detection information indicating that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension; If not, detection information is generated to indicate that the candidate product image does not meet the first quality condition corresponding to the preset detection dimension.

[0008] In a possible implementation, the preset detection dimension includes at least one of the following: a redundant information detection dimension, a clarity detection dimension, an illumination detection dimension, an integrity detection dimension, a copy detection dimension, a boundary line detection dimension, and a consistency detection dimension; The quality goal corresponding to the redundant information detection dimension is: the candidate product image does not contain redundant information; The quality goal corresponding to the clarity detection dimension is: the clarity of the candidate product image is greater than the preset clarity threshold; The quality goal corresponding to the illumination detection dimension is: the exposure of the candidate product image is within the preset exposure range; The quality target corresponding to the integrity detection dimension is: the integrity score of the image subject in the candidate product image is greater than the preset integrity score threshold; The quality goal for the copy detection dimension is: the candidate product image is not a copy image. A copy image is an image obtained by re-photographing an image displayed in an existing medium. The quality target corresponding to the boundary line detection dimension is: the image edge of the candidate product image does not contain boundary lines; The quality target corresponding to the consistency detection dimension is: the similarity score between the image subject in the candidate product image and the corresponding product name is greater than or equal to the preset similarity threshold.

[0009] In one possible implementation, after detecting the candidate product image according to a preset detection dimension and obtaining detection information corresponding to the candidate product image, the method further includes: When the detection information indicates that the candidate product image does not meet the first quality condition, if the candidate product image contains redundant information, the redundant information is eliminated, and based on the processed candidate product image, the step of detecting the candidate product image according to the preset detection dimension is returned to; or When the detection information indicates that the candidate product image does not meet the first quality condition, if the integrity score of the image body in the candidate product image is less than the preset integrity score threshold, the candidate product image is extended, and based on the processed candidate product image, the step of detecting the candidate product image according to the preset detection dimension is returned.

[0010] In one possible implementation, eliminating redundant information includes: Detecting redundant information included in the candidate product image and determining the redundant information location area; The redundant information location area is subjected to image erasure processing to generate processed candidate product images.

[0011] In a possible implementation, performing extension processing on the candidate product image includes: Extracting mask information corresponding to the image subject in the candidate product image; Determine the corresponding minimum bounding rectangle area based on the mask information; Determine the image area to be extended according to the minimum circumscribed rectangular area; The image region to be extended is extended to obtain a processed candidate product image.

[0012] In one possible implementation, the subjective evaluation data includes at least one of the following: image subject evaluation information, image background evaluation information, and image overall evaluation information; In the case where the subjective evaluation data includes image subject evaluation information, the second quality condition includes: a deformation score of the image subject in the candidate product image is less than a first deformation score threshold; In the case where the subjective evaluation data includes image background assessment information, the second quality condition includes: a deformation score of the image background in the candidate product image is less than a second deformation score threshold; When the subjective evaluation data includes overall image evaluation information, the second quality condition includes: the image naturalness score of the candidate product image is greater than a preset naturalness score threshold, the image composition score of the candidate product image is greater than a preset composition score threshold, and the image color balance score of the candidate product image is greater than a preset color balance score threshold.

[0013] In a possible implementation, the candidate product image is a dish image, and the image body of the candidate product image includes a dish area and a container area for loading the dish.

[0014] In one possible implementation, determining a target product image based on candidate product images includes: Get the encoding information to be added to the watermark; Extracting spectrum information of candidate product images; Embedding the coded information into the spectrum information to obtain processed spectrum information; The processed spectrum information is inversely transformed to generate the target product image with the watermark added.

[0015] In a second aspect, an embodiment of the present application provides an image display method, applied to a user terminal, comprising: Receive a target product image sent by a server, where the target product image is an image determined by the server based on the candidate product image when user evaluation information indicates that the candidate product image satisfies a second quality condition, where the second quality condition is used to measure the content performance of the candidate product image in user interaction; user evaluation information is information corresponding to the candidate product image obtained by the server when detection information indicates that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension, where the first quality condition is that all quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated from the user's behavioral data and / or subjective evaluation data; and detection information is information obtained by the server through detection of the candidate product image according to the preset detection dimension, where the detection information is used to characterize the objective quality performance of the candidate product image under the preset detection dimension; In response to a trigger operation on the product indicated by the target product image, the target product image is displayed.

[0016] In a third aspect, an embodiment of the present application provides an image processing system, comprising: a server and a user terminal; The server is configured to obtain candidate product images, detect the candidate product images according to preset detection dimensions, obtain detection information corresponding to the candidate product images, and, if the detection information indicates that the candidate product images satisfy a first quality condition corresponding to the preset detection dimensions, obtain user evaluation information corresponding to the candidate product images. The detection information is used to characterize the objective quality performance of the candidate product images under the preset detection dimensions. The first quality condition is that all quality targets corresponding to the preset detection dimensions are satisfied. The user evaluation information is information generated from user behavior data and / or subjective evaluation data. Furthermore, if the user evaluation information indicates that the candidate product images satisfy a second quality condition, a target product image is determined based on the candidate product images, and the target product image is sent to a user terminal. The second quality condition is used to measure the content performance of the candidate product images in user interaction. The user terminal is configured to receive the target product image sent by the server, and display the target product image in response to a triggering operation on the product indicated by the target product image.

[0017] In a fourth aspect, an embodiment of the present application provides an image processing device, applied to a server, comprising: A first acquisition module is used to acquire candidate product images; A detection module is used to detect the candidate product image according to a preset detection dimension and obtain detection information corresponding to the candidate product image. The detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension; a second acquisition module configured to acquire user evaluation information corresponding to the candidate product image if the detection information indicates that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension, wherein the first quality condition is that quality targets corresponding to each preset detection dimension are satisfied, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; a determination module configured to determine a target product image based on the candidate product images if the user evaluation information indicates that the candidate product images meet a second quality condition, the second quality condition being used to measure content performance of the candidate product images in user interaction; The sending module is used to send the target product image to the user terminal so that the user terminal displays the target product image.

[0018] In a fifth aspect, an embodiment of the present application provides an image display device, applied to a user terminal, comprising: A receiving module is configured to receive a target product image sent by a server, wherein the target product image is an image determined by the server based on the candidate product image when user evaluation information indicates that the candidate product image satisfies a second quality condition, the second quality condition being used to measure the content performance of the candidate product image in user interaction; user evaluation information is information corresponding to the candidate product image obtained by the server when detection information indicates that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension, the first quality condition being that quality targets corresponding to each preset detection dimension are all satisfied, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; detection information is information obtained by the server through detection of the candidate product image according to the preset detection dimension, the detection information being used to characterize the objective quality performance of the candidate product image under the preset detection dimension; The display module is configured to display the target product image in response to a triggering operation on the product indicated by the target product image.

[0019] In a sixth aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory; wherein the memory stores a computer program, and when the processor executes the computer program, the method steps provided in the first aspect or the second aspect of the embodiment of the present application are implemented.

[0020] In the seventh aspect, an embodiment of the present application provides a computer storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor and executing the method steps provided in the first aspect or the second aspect of the embodiment of the present application.

[0021] The above-mentioned image processing method, image display method, device, equipment and storage medium are applied to a server, and obtain a candidate product image and detect the candidate product image according to a preset detection dimension to obtain detection information corresponding to the candidate product image. The detection information is used to characterize the objective quality performance of the candidate product image under the preset detection dimension. Then, when the detection information indicates that the candidate product image meets the first quality condition corresponding to the preset detection dimension, user evaluation information corresponding to the candidate product image is obtained. The first quality condition is that the quality targets corresponding to each preset detection dimension are all met. The user evaluation information is information generated by the user's behavior data and / or subjective evaluation data. And when the user evaluation information indicates that the candidate product image meets the second quality condition, a target product image is determined based on the candidate product image. The second quality condition is used to measure the content performance of the candidate product image in user interaction. The target product image is then sent to the user terminal so that the user terminal displays the target product image. Therefore, by introducing a multi-stage quality detection and evaluation mechanism, combining automated objective quality detection with subjective user evaluation information combined with user behavior data and / or subjective evaluation data, the overall quality and usability of AI-generated product images can be effectively improved. Not only does this ensure that images meet the standards for pre-defined detection dimensions, but manual review also further selects higher-quality images of target products. Ultimately, these high-quality images are pushed to terminal devices for display, significantly enhancing the user's visual experience, increasing content credibility, and helping to improve product display and platform conversion rates. The technical solutions provided in this application can be applied to transactions and delivery services on instant e-commerce platforms, such as Taobao Flash Sales, Taoxianda, and Ele.me food delivery and retail. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 A schematic diagram of the architecture of an image processing system provided by an exemplary embodiment of the present application; Figure 2 A flowchart of an image processing method provided by an exemplary embodiment of the present application; Figure 3 A flowchart of another image processing method provided by an exemplary embodiment of the present application; Figure 4 A flowchart of another image processing method provided by an exemplary embodiment of the present application; Figure 5A flowchart of a method for automatically detecting candidate product images provided by an exemplary embodiment of the present application; Figure 6 A flowchart of an image extension method provided by an exemplary embodiment of the present application; Figure 7 A schematic diagram of an effect of extending an image subject provided by an exemplary embodiment of the present application; Figure 8 A flowchart of another image processing method provided by an exemplary embodiment of the present application; Figure 9 A schematic diagram of an image review dimension provided by an exemplary embodiment of the present application; Figure 10 A schematic diagram of a candidate product image including deformation provided by an exemplary embodiment of the present application; Figure 11 A flowchart of a watermark embedding method provided by an exemplary embodiment of the present application; Figure 12 A flowchart of an image processing method provided by an exemplary embodiment of the present application; Figure 13 A schematic structural diagram of an image processing device provided by an exemplary embodiment of the present application; Figure 14 A schematic structural diagram of an image display device provided by an exemplary embodiment of the present application; Figure 15 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0025] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0026] See Figure 1, which is a schematic diagram of the architecture of an image processing system provided by an exemplary embodiment of the present application. In particular, server 10 communicates with user terminal 20 via a network. A data storage system can store data that server 10 needs to process. The data storage system can be integrated with server 10 or placed on a cloud or other network server.

[0027] In some possible embodiments, the server 10 obtains a candidate product image; detects the candidate product image according to a preset detection dimension to obtain detection information corresponding to the candidate product image, and the detection information is used to characterize the objective quality performance of the candidate product image under the preset detection dimension; when the detection information indicates that the candidate product image meets a first quality condition corresponding to the preset detection dimension, obtains user evaluation information corresponding to the candidate product image, and the first quality condition is that the quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated by the user's behavior data and / or subjective evaluation data; when the user evaluation information indicates that the candidate product image meets a second quality condition, determine a target product image based on the candidate product image, and the second quality condition is used to measure the content performance of the candidate product image in user interaction; and then send the target product image to the user terminal 20, so that the user terminal 20 displays the target product image.

[0028] It is understood that the server 10 can be implemented as a standalone server or a server cluster consisting of multiple servers. The user terminal 20 can be, but is not limited to, various smartphones, tablet computers, personal computers, laptop computers, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can include smart watches, smart bracelets, head-mounted devices, etc.

[0029] Optionally, the user terminal 20 may include a terminal device used by an ordering user, such as a terminal device used to browse products, place orders, and receive product image display content; it may also include a terminal device used by a supply user (such as a merchant), such as a merchant-side terminal device used to upload product images and set product information.

[0030] In one embodiment, Figure 2 As shown, an image processing method is provided, which is applied to a server and includes the following steps: S201: Acquire candidate product images.

[0031] Among them, the image subject in the candidate product image may include products, for example, dishes, drinks, medicines and other products. The candidate product image is a product picture that has not been finally determined whether to be displayed to the user terminal, and can be used as the object of subsequent evaluation and screening.

[0032] Optionally, the candidate commodity image can be generated by a related AI model, or a plurality of candidate commodity images that can be displayed or recommended can be obtained from a related database or source (such as a commodity library of an e-commerce platform, an image acquisition system, etc.).

[0033] In one embodiment, the candidate commodity image is a dish image, and the image subject in the candidate commodity image includes a dish area and a vessel area for loading the dish.

[0034] S202: detecting the candidate commodity image according to a preset detection dimension to obtain detection information corresponding to the candidate commodity image, the detection information being used to represent an objective quality performance of the candidate commodity image under the preset detection dimension.

[0035] The preset detection dimension can be one or more. The preset detection dimension can be a pre-set evaluation dimension, such as sharpness, contrast, composition integrity, etc., which can be used for quality detection of the candidate image to obtain a preliminary detection result, i.e., the detection information.

[0036] S203: In a case where the detection information indicates that the candidate commodity image satisfies a first quality condition corresponding to the preset detection dimension, obtaining user evaluation information corresponding to the candidate commodity image, the first quality condition being that the quality targets corresponding to each preset detection dimension are all satisfied, and the user evaluation information being information generated by behavior data and / or subjective evaluation data of a user.

[0037] Optionally, the first quality condition can include a quality target corresponding to each preset detection dimension, and the evaluation dimension and the quality target correspond one-to-one, and the first quality condition is that the quality targets corresponding to each preset detection dimension are all satisfied. That is, if any quality target is not satisfied, it can be determined that the candidate commodity image does not satisfy the first quality condition corresponding to the preset detection dimension.

[0038] In one embodiment, the user can be a related reviewer, such as a professional person in charge of commodity image quality review, an operation person or a content management person, and the user can subjectively evaluate the candidate commodity image according to the second quality condition.

[0039] In another embodiment, the user can also be an ordering user, i.e., a real user in a target user group or a terminal purchaser. The ordering user can indirectly express the subjective acceptance of the candidate commodity image by the ordering user itself based on behavior information such as browsing, clicking, collecting, staying time, purchase conversion, etc.; or can subjectively evaluate the image through explicit scoring, complaint feedback, image preference selection, etc.; and then the data can be processed into user evaluation information for further determining whether the candidate commodity image satisfies the second quality condition.

[0040] Optionally, behavioral data can be objective behavioral information generated by users during their interaction with candidate product images, which can be used to indirectly reflect the user's interest in or acceptance of the image. Behavioral data may include, but is not limited to, click behavior (i.e., whether the user clicks to view the product containing the image), browsing time (how long the user stays on the page containing the image), scrolling behavior (whether the user swipes to view the image area or quickly skips), favorite / add-to-cart behavior (whether the user adds the product to favorites or a shopping cart), conversion behavior (whether the user ultimately places an order), and bounce rate / return rate (whether the user leaves the page or returns to the previous interface after viewing the image). This behavioral data can be automatically collected and generated by the system, for example, through embedded point detection, log analysis, and other methods, without the need for active user input.

[0041] It is understandable that although behavioral data is essentially objective data automatically collected by the system, it can still be used to indirectly reflect users' subjective feelings or preferences towards candidate product images, so as to model user preferences or image attractiveness from a data level.

[0042] Optionally, subjective evaluation data can be explicit user evaluation information based on a candidate product image's subjective judgment. This data can be obtained through manual input or active feedback, directly reflecting the user's subjective opinion on the image quality or visual effects. Subjective evaluation data may include, but is not limited to, rating information (e.g., user-assigned ratings of images), textual comments (user-entered comments such as "unclear image" or "good shooting angle"), tag selections (user-selected evaluation tags such as "reasonable composition"), complaint records (user complaints or reports regarding unsatisfactory images), and image preference feedback (e.g., selecting a favorite image from multiple images). Subjective evaluation data can be derived from explicit feedback from professional reviewers or end users. It is highly subjective and can reflect the image's display quality, acceptance, or potential issues at a perceptual level.

[0043] Specifically, the server can be connected to an external display or to other user terminals via a network. The server can display candidate product images and questions related to the candidate product images to users via the external display or user terminal, and collect users' responses to the questions via the external display or user terminal to form subjective evaluation data. The responses can be multiple-choice questions, ratings, text comments, or other structured data.

[0044] S204: When the user evaluation information indicates that the candidate product image meets a second quality condition, determining a target product image based on the candidate product image, where the second quality condition is used to measure content performance of the candidate product image in user interaction.

[0045] The second quality condition may include criteria set by the user based on actual needs or subjective judgment, such as the candidate product image's deformation, exposure, consistency with the product description, and brand image fit. If a candidate product image meets the second quality condition, it may be determined as the target product image for subsequent display or recommendation to the user terminal. If it does not meet the second quality condition, the candidate product image may be eliminated or returned to the candidate product image pool for further optimization or retesting.

[0046] S205: Send the target product image to the user terminal, so that the user terminal displays the target product image.

[0047] Optionally, the server may send the target product image to the user terminal according to a specific application scenario (such as an e-commerce platform, a restaurant ordering system, an advertising recommendation system), so that the user terminal can trigger the display of the target product image.

[0048] In one embodiment, the server may send the target product image to the user terminal upon receiving a request from the user terminal, or may actively push the target product image according to a recommendation algorithm.

[0049] The user terminal may be configured to receive a target product image sent by a server, and display the target product image in response to a triggering operation on the product indicated by the target product image.

[0050] The above-mentioned image processing method, image display method, device, equipment and storage medium are applied to a server, and obtain a candidate product image and detect the candidate product image according to a preset detection dimension to obtain detection information corresponding to the candidate product image, and the detection information is used to characterize the objective quality performance of the candidate product image under the preset detection dimension; then, when the detection information indicates that the candidate product image meets the first quality condition corresponding to the preset detection dimension, user evaluation information corresponding to the candidate product image is obtained, the first quality condition is that the quality targets corresponding to each preset detection dimension are all met, and the user evaluation information is information generated by the user's behavior data and / or subjective evaluation data; and when the user evaluation information indicates that the candidate product image meets the second quality condition, a target product image is determined based on the candidate product image, and the second quality condition is used to measure the content performance of the candidate product image in user interaction; thereby, the target product image is sent to the user terminal so that the user terminal displays the target product image. Therefore, by introducing a multi-stage quality evaluation mechanism, combining automated objective quality detection with subjective user evaluation information combined with user behavior data and / or subjective evaluation data, the overall quality and usability of AI-generated product images can be effectively improved. Not only does this ensure that images meet pre-defined inspection criteria, but manual review further selects higher-quality images of target products. Ultimately, these high-quality images are pushed to end-user devices for display, significantly enhancing the user experience, increasing content credibility, and helping to improve product display and platform conversion rates.

[0051] In one embodiment, Figure 3 As shown, another image processing method is provided, which is applied to a server and includes the following steps: S301: Acquire preset product information, which includes the product name and the corresponding product demand weight.

[0052] The commodity demand weight may be the degree of demand for the commodity indicated by the preset commodity information. A larger commodity demand weight indicates a greater degree of demand for the commodity, and a smaller commodity demand weight indicates a smaller degree of demand for the commodity.

[0053] Optionally, in a restaurant ordering scenario, the preset product information may be dish information or drink information. If the preset product information is dish information, the product name included in the preset product information is the dish name, and the corresponding product demand weight is the dish demand weight. The dish demand weight can be determined or adjusted based on factors such as dish sales, customer preferences, and seasonality.

[0054] S302: Determine whether the product demand weight is greater than a preset weight threshold.

[0055] The preset weight threshold is a pre-set reference value used to determine whether the demand for a product is high enough.

[0056] Optionally, relevant indicators that reflect demand heat can be selected, such as the product's historical sales volume, search popularity, number of user clicks and favorites, add-to-cart ratio, inventory turnover rate, etc.; then, each indicator is assigned a weight coefficient according to its importance in actual application, and then a weighted summation or multi-indicator scoring model is used to calculate the comprehensive demand weight of the product, and the result is mapped to a unified numerical range through normalization; finally, based on the capacity planning, operation strategy and historical data distribution in actual application, combined with demand percentile analysis, a preset weight threshold is determined, so that the preset weight threshold can not only filter out products with lower demand, but also retain products with potential conversion value.

[0057] For example, based on the platform transaction data of the past year, historical sales, search popularity, number of user clicks and favorites, add-to-cart ratio and inventory turnover rate can be assigned weight coefficients of 0.3, 0.25, 0.2, 0.15 and 0.1 respectively to calculate the comprehensive demand weight of each product, and normalize the result to the interval [0,1]. Assuming that after analyzing the distribution of the demand weights of each product, it is found that the demand weights of high-conversion products are mostly concentrated above 0.75, 0.75 can be set as the preset weight threshold, so that in subsequent screening, when the comprehensive demand weight of a product is greater than or equal to 0.75, it can be determined that the demand level of the product is high.

[0058] S303: When the product demand weight is greater than a preset weight threshold, the product name is input into a first preset model, and the output image of the first preset model is determined as a candidate product image; wherein the first preset model is a text-driven image generation model.

[0059] Optionally, the first preset model is a model that generates product images based on product names. This model can use natural language processing technology to understand product names and generate related images. For example, after inputting "Fish-flavored Shredded Pork", the model will output a corresponding dish image as a candidate product image.

[0060] S304: When the product demand weight is less than or equal to a preset weight threshold, the product name and the initial product image are input into a second preset model, and the output image of the second preset model is determined as a candidate product image; wherein the second preset model is an image generation model based on joint input of images and text.

[0061] Optionally, when the product demand weight is less than or equal to a preset weight threshold, another image generation model, namely a second preset model, is required. The model receives the product name and the initial product image as input and generates an improved product image as a candidate product image.

[0062] That is to say, the second preset model not only relies on the product name, but can also generate higher quality candidate product images based on the joint analysis of text and image information.

[0063] For example, in a restaurant ordering scenario, there might be a dish with low demand, and its initial product image might be low-definition. In this case, the second preset model combines the product name (dish name) and the original initial product image to generate a more relevant and beautiful candidate product image, thereby increasing the click-through rate of users on the restaurant ordering platform.

[0064] In one embodiment, using a restaurant ordering scenario as an example, the initial model can be trained based on a relevant dish dataset. A dual-link generation strategy is employed for dishes with different demand weights. For dishes with higher demand weights, a direct link generation strategy is implemented using a text-based graph (the first preset model), where the dish name is input and an image is generated. For less common dishes with lower demand weights, a link generation strategy combining images and text (the second preset model) is employed, where both the dish name and the initial dish image are input. This strategy helps the model generate candidate dish images that better meet actual needs.

[0065] S305: Detect the candidate product image according to the preset detection dimension to obtain detection information corresponding to the candidate product image, where the detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension.

[0066] Specifically, S305 is consistent with S202 and will not be repeated here.

[0067] S306: When the detection information indicates that the candidate product image meets the first quality condition corresponding to the preset detection dimension, obtain user evaluation information corresponding to the candidate product image, where the first quality condition is that the quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated by the user's behavioral data and / or subjective evaluation data.

[0068] Specifically, S306 is consistent with S203 and will not be repeated here.

[0069] S307: When the user evaluation information indicates that the candidate product image meets a second quality condition, determining a target product image based on the candidate product image, where the second quality condition is used to measure content performance of the candidate product image in user interaction.

[0070] Specifically, S307 is consistent with S204 and will not be repeated here.

[0071] S308: Send the target product image to the user terminal, so that the user terminal displays the target product image.

[0072] Specifically, S308 is consistent with S205 and will not be repeated here.

[0073] In the embodiment of the present application, by introducing product demand weights and adopting a dual-link image generation strategy, different image generation models can be flexibly selected according to the popularity of products (such as dishes) in actual application scenarios: for products with high demand, images are directly generated through the first preset model driven by text, which is highly efficient and fast; while for products with low demand or relatively uncommon products, more exquisite and realistic images are generated through the second preset model that integrates the product name and the original image, thereby improving the overall product display quality.

[0074] In one embodiment, Figure 4 As shown, another image processing method is provided, which is applied to a server and includes the following steps: S401: Acquire candidate product images.

[0075] Specifically, S401 is consistent with S201 and will not be repeated here.

[0076] S402: Detect the candidate product image according to each preset detection dimension to obtain the detection results corresponding to the candidate product image under each preset detection dimension.

[0077] The preset detection dimensions are used to perform multi-dimensional detection on the candidate product images, so as to detect the quality of the candidate product images in different aspects respectively.

[0078] In one embodiment, the preset detection dimensions include at least one of the following: a redundant information detection dimension, a clarity detection dimension, a lighting detection dimension, an integrity detection dimension, a copy detection dimension, a boundary line detection dimension, and a consistency detection dimension. The redundant information detection dimension is used to detect whether candidate product images contain redundant information (such as abnormal text). The clarity detection dimension is used to assess the clarity of candidate product images, retaining candidate product images that clearly display product details. The lighting detection dimension is used to detect the lighting conditions of candidate product images, retaining candidate product images with appropriate exposure. The integrity detection dimension is used to check the integrity of candidate product images, retaining candidate product images whose main body is not cropped, obscured, or missing. The copy detection dimension is used to check whether an image is a copy, that is, whether it is an image regenerated by photographing an existing image (such as an image displayed on another display device or an image printed on a medium such as paper), in order to eliminate low-quality, distorted copy images. The boundary line detection dimension is used to check whether candidate product images contain unnecessary boundary lines.

[0079] Optionally, when the preset detection dimension includes a redundant information detection dimension, discrete wavelet transform or discrete Fourier transform can be used to perform frequency domain analysis on the image to identify redundant information embedding features. Convolutional neural networks can also be trained to identify redundant information areas to determine whether the selected product image contains redundant information and obtain corresponding detection results.

[0080] Optionally, if the preset detection dimension includes a clarity detection dimension, the corresponding clarity can be determined by calculating the Laplace transform variance of the candidate product image, where a larger variance indicates a higher clarity of the candidate product image. Alternatively, a pre-trained image quality assessment model can be used to score the clarity of the candidate product image to obtain the corresponding clarity, thereby determining whether the clarity of the candidate product image is greater than a preset clarity threshold and obtaining the corresponding detection result. The clarity can be a value in the interval [0, 1], and the preset clarity threshold can be set to 0.8.

[0081] In one embodiment, the preset clarity threshold is a reference value used to determine whether a candidate product image meets the clarity standards required for practical applications. Specifically, a large number of product image samples with different clarity levels can be pre-collected and manually subjectively graded for clarity, with these scores used as benchmark labels. Subsequently, the Laplace transform variance of the product image samples is calculated, or a pre-trained image quality assessment model is used to generate clarity scores. The variance calculation results or model output results are then correlated with the manual scores to determine the numerical distribution characteristics and calculate the distribution range of clarity scores for high-quality images. Finally, based on the requirements for image detail fidelity, sharpness, and display effects in practical applications, a cutoff value that can effectively distinguish between clear and blurry images is selected as the preset clarity threshold.

[0082] For example, based on the image data of products that have been listed on the platform in the past year, samples that have been manually rated as "clear" or above can be used as a high-quality image set, and the distribution range of the clarity scores of the high-quality image set can be calculated. When the clarity score range is [0,1], assuming that the clarity score corresponding to the 80% percentile is 0.8, then 0.8 can be set as the preset clarity threshold, so that when the clarity score of the candidate product image is greater than or equal to 0.8, it can be determined that the candidate product image meets the actual application requirements in terms of detail fidelity, edge sharpness and overall display effect.

[0083] Optionally, when the preset detection dimension includes a light detection dimension, the exposure corresponding to the candidate product image can be determined by the grayscale mean method, the brightness standard deviation method or the brightness histogram analysis method, and then the exposure of the candidate product image can be determined to be within the preset exposure range, and the corresponding detection result can be obtained.

[0084] Optionally, in the case that the preset detection dimension includes the integrity detection dimension, whether the outline of the image subject is complete can be detected by a related model, the occlusion area is identified, and the completeness score of the image subject is determined, so as to determine whether the completeness score of the image subject is greater than a preset completeness score threshold, and then the corresponding detection result is obtained.

[0085] Optionally, the completeness score can be a quantitative numerical value for measuring the completeness of the image subject in the candidate commodity image in terms of shape, structure, visibility, and the like. The higher the numerical value of the completeness score, the more complete the image subject and the less missing or occlusion. The specific calculation process can include: first, identifying and extracting the pixel area of the image subject by using a target detection or image segmentation model; then, calculating the proportion of the complete area of the subject in the image and the proportion of the occluded or missing part by using contour analysis, shape matching, and occlusion detection methods, wherein the ratio of the number of subject pixels to the number of expected complete subject pixels can be used as a basic completeness score indicator; then, the difference between the current image subject and the reference form in terms of geometric structure, edge contour, and the like can be further compared; finally, the above indicators are weighted and fused and normalized to the interval [0, 1] to obtain the completeness score. When the completeness score is greater than a preset completeness score threshold, it can be determined that the image subject meets the quality requirements of complete presentation in vision.

[0086] The preset completeness score threshold is used to determine whether the subject of the candidate commodity image is complete in shape and structure. Specifically, based on a large number of commodity image samples with labeled completeness score levels, the subject outline can be extracted by using an image segmentation or target detection model, and the completeness score indicators such as the proportion of subject pixels, the proportion of missing areas, and the proportion of occluded areas can be calculated, and the results can be normalized to a unified interval. Then, by analyzing the completeness score distribution of high-completeness-score samples (artificially labeled as "complete" or above), and combining the requirements for the visibility and detail presentation of the image subject in actual applications, a threshold value that can effectively distinguish between complete and incomplete images can be selected, for example, the lower quantile value (such as 0.85) of the distribution of high-quality samples is taken as the preset completeness score threshold, so that when the completeness score is greater than the threshold value, the structure of the image subject has no missing or has less missing area, and meets the related display standards.

[0087] Optionally, in the case that the preset detection dimension includes the integrity detection dimension, whether the outline of the image subject is complete can be detected by a related model, the occlusion area is identified, and the completeness score of the image subject is determined, so as to determine whether the completeness score of the image subject is greater than a preset completeness score threshold, and then the corresponding detection result is obtained.

[0088] Optionally, when the preset detection dimension includes a boundary line detection dimension, the image edge of the candidate product image can be scanned and pixel brightness statistics can be performed to determine whether the image edge of the candidate product image contains a boundary line (such as a black or white boundary) and obtain the corresponding detection result.

[0089] S403: Based on the detection results, determine whether the candidate product image meets the quality targets of each preset detection dimension; if so, execute S404; if not, execute S405.

[0090] Among them, the quality target corresponding to the redundant information detection dimension is: the candidate product image does not contain redundant information; the quality target corresponding to the clarity detection dimension is: the clarity of the candidate product image is greater than the preset clarity threshold; the quality target corresponding to the lighting detection dimension is: the exposure of the candidate product image is within the preset exposure range; the quality target corresponding to the integrity detection dimension is: the integrity score of the image subject in the candidate product image is greater than the preset integrity score threshold; the quality target corresponding to the copy detection dimension is: the candidate product image is not a copy image, and the copy image is an image obtained by re-photographing the image displayed in the existing medium; the quality target corresponding to the boundary line detection dimension is: the image edge of the candidate product image does not contain a boundary line; the quality target corresponding to the consistency detection dimension is: the similarity score between the image subject in the candidate product image and the corresponding product name is greater than or equal to the preset similarity threshold.

[0091] The similarity score can be used to measure the degree of semantic match between the main body of the candidate product image and the corresponding product name. This score can range from 0 to 1 or 0 to 100, with larger scores indicating a higher degree of match. Optionally, a preset similarity threshold can be used to determine whether the semantic match between the main body of the candidate product image and the corresponding product name meets practical requirements. Specifically, based on a large sample of product images and names with annotated matches / mismatches, a multimodal feature alignment model can be used to extract the main image feature vector and the text feature vector. The cosine similarity between the main image feature vector and the text feature vector is calculated and normalized to the interval [0, 1] or [0, 100]. To determine the similarity threshold, the similarity score distribution of highly matching samples (manually labeled as "match" or above) can be statistically analyzed. Based on the practical requirements for balancing recognition accuracy and recall, a cutoff value that effectively distinguishes between highly matching and poorly matching samples can be selected. For example, the lower quantile of the distribution of highly matching samples (e.g., 0.85 or 85) can be used as the preset similarity threshold.

[0092] S404: Generate detection information indicating whether the candidate product image meets a first quality condition corresponding to a preset detection dimension.

[0093] That is, if the candidate product images all meet the quality targets corresponding to the preset detection dimensions, detection information is generated, which indicates that the candidate product images meet the first quality condition, indicating that the candidate product images have passed the preliminary automated detection.

[0094] Figure 5 This is a flow chart of a method for automatically detecting candidate product images provided by an exemplary embodiment of the present application. Figure 5 Taking the preset detection dimensions including redundant information detection dimension, clarity detection dimension, illumination detection dimension, integrity detection dimension, copy detection dimension, boundary line detection dimension and consistency detection dimension as an example, the candidate product image can be initially automatically detected based on the redundant information detection dimension, clarity detection dimension, illumination detection dimension, integrity detection dimension, copy detection dimension and boundary line detection dimension. When the detection results of each preset detection dimension meet the corresponding quality target, the candidate product image is then subjected to image-text consistency detection based on the consistency detection dimension to determine the similarity score between the image subject in the candidate product image and the corresponding product name, and then determine whether the similarity score is greater than or equal to the preset similarity threshold, and obtain the corresponding detection result. If the detection result corresponding to the consistency detection dimension indicates that the similarity score is less than the preset similarity threshold, it can be determined that the candidate product image does not meet the quality target of each preset detection dimension; if the detection result corresponding to the consistency detection dimension indicates that the similarity score is greater than or equal to the preset similarity threshold, it can be determined that the candidate product image meets the quality target of each preset detection dimension, determines that the candidate product image meets the first quality condition, and generates corresponding detection information.

[0095] Specifically, taking the processing of candidate dish images as an example, to evaluate the aforementioned consistency dimension, a comparative language-image pre-trained model, generated by training on a large amount of dish image data, can be used to calculate the similarity score between the dish name (product name) and the candidate dish image (candidate product image). The comparative language-image pre-trained model includes at least two encoders: an image encoder (ImageEncoder), which extracts features from image data, and a text encoder (TextEncoder), which extracts features from text data. Based on the features extracted by these two encoders, a similarity score can be calculated between the dish name and the candidate dish image. This similarity score reflects whether the candidate dish image satisfies the description of the dish name; a higher similarity score indicates a higher degree of match between the two.

[0096] In addition, for the evaluation of the above-mentioned consistency dimension, the main area of ​​the candidate product image can also be located and extracted, and the main area is input into the image feature extraction model to obtain an image feature vector; and the corresponding product name is input into the text encoding model to obtain a text feature vector. After model training, the image feature vector and the text feature vector can be aligned to the same semantic space; then, the cosine similarity algorithm is used to calculate the cosine value of the angle between the image feature vector and the text feature vector to quantify the degree of semantic matching between the image feature vector and the text feature vector. The cosine value of the angle is the similarity score, and a higher score indicates a higher degree of matching.

[0097] In the embodiments of the present application, by setting multiple preset detection dimensions and setting corresponding quality targets for each detection dimension, comprehensive quality testing of candidate product images is achieved, so that the candidate product images meet the standards in multiple dimensions such as clarity, lighting, redundant information, and boundaries. When the candidate product image meets the quality targets in all dimensions, it is determined to meet the first quality condition, thereby generating clear detection information. This refined quality control mechanism helps to improve the quality and professionalism of candidate product images, filter out low-quality content, and enhance the platform display effect and user experience.

[0098] S405: When the detection information indicates that the candidate product image meets the first quality condition corresponding to the preset detection dimension, obtain user evaluation information corresponding to the candidate product image, where the first quality condition is that the quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated by the user's behavioral data and / or subjective evaluation data.

[0099] Specifically, after the candidate product image meets the first quality condition, the candidate product image can be further manually reviewed to obtain subjective evaluation data determined by the reviewer (user); and / or, the behavioral data generated by the user during the interaction process is collected, and then the subjective evaluation data and / or behavioral data are integrated to generate corresponding user evaluation information.

[0100] S406: When the user evaluation information indicates that the candidate product image meets a second quality condition, determining a target product image based on the candidate product image, where the second quality condition is used to measure content performance of the candidate product image in user interaction.

[0101] Specifically, S406 is consistent with S204 and will not be repeated here.

[0102] S407: Send the target product image to the user terminal, so that the user terminal displays the target product image.

[0103] Specifically, S407 is consistent with S205 and will not be repeated here.

[0104] S408: Generate detection information indicating that the candidate product image does not meet a first quality condition corresponding to a preset detection dimension.

[0105] Optionally, if the detection information indicates that the candidate product image does not meet the first quality condition corresponding to the preset detection dimension, the candidate product image can be eliminated or further image optimization can be performed. The specific operation can be flexibly determined according to the actual scenario requirements.

[0106] S409: When the detection information indicates that the candidate product image does not meet the first quality condition, if the candidate product image contains redundant information, the redundant information is eliminated, and based on the processed candidate product image, the step of detecting the candidate product image according to the preset detection dimension is returned; or, when the detection information indicates that the candidate product image does not meet the first quality condition, if the integrity score of the image body in the candidate product image is less than the preset integrity score threshold, the candidate product image is extended, and based on the processed candidate product image, the step of detecting the candidate product image according to the preset detection dimension is returned.

[0107] It is worth noting that after eliminating the redundant information, the above-mentioned step of obtaining the user evaluation information corresponding to the candidate product image can be directly performed based on the processed candidate product image; similarly, after extending the candidate product image, the above-mentioned step of obtaining the user evaluation information corresponding to the candidate product image can also be directly performed based on the processed candidate product image. The specific operation can be flexibly determined according to the actual scenario requirements.

[0108] In one embodiment, image restoration or filling technology may be used to restore areas in the candidate product image that are blocked by redundant information, so that the processed candidate product image is cleaner and neater.

[0109] In the above embodiment, when the candidate product image does not meet the first quality condition, redundant information in the image is eliminated, or when the image body integrity score is insufficient, the image is extended, and the processed image is returned to the quality inspection process, so that the image meets the standard requirements in key dimensions, thereby achieving closed-loop quality optimization. This not only improves the intelligence level of image quality restoration, but also allows the user evaluation stage to be directly entered after processing, avoiding repeated operations and improving overall processing efficiency and system response speed. At the same time, this method is flexible and can select the processing path as needed according to the actual scenario, effectively balancing image quality control and processing efficiency.

[0110] Optionally, in S409, the redundant information is eliminated, including: detecting the redundant information included in the candidate commodity image, determining the redundant information position area; performing image erasing processing on the redundant information position area to generate the processed candidate commodity image.

[0111] Specifically, the redundant information included in the candidate commodity image can be detected by a related model based on deep learning, the redundant information position area is determined, and then the image erasing processing is performed on the redundant information position area to erase the redundant information, thereby obtaining the processed candidate commodity image.

[0112] In one embodiment, if the completeness score of the image subject in the candidate commodity image is less than a preset completeness score threshold, the candidate can be extended by a related image generation technology. The extension processing can refer to using image completion technology to infer the missing part by context, and restore the complete commodity image. For example, if a part of the commodity is cropped, the missing part of the image can be automatically generated, so that the candidate commodity image appears more complete.

[0113] In the above embodiments, the redundant information position in the candidate commodity image is accurately detected by a related model, and then the detected area is intelligently erased, which can effectively remove the redundant information in the image that affects the visual experience, and improve the neatness and professionalism of the image.

[0114] Optionally, in S409, the candidate commodity image is extended, including: extracting mask information corresponding to the image subject in the candidate commodity image; determining a minimum bounding rectangle area corresponding to the mask information; determining a to-be-extended image area according to the minimum bounding rectangle area; and performing extension processing on the to-be-extended image area to obtain a processed candidate commodity image.

[0115] Specifically, please refer to Figure 6 , Figure 6A flowchart of an image extension method provided for an exemplary embodiment of the present application, first, the candidate product image 601 can be processed using a relevant cutout model to extract the mask information (mask) corresponding to the image subject in the candidate product image 601. Then, the mask information can be geometrically analyzed to calculate and crop the minimum bounding rectangle area corresponding to the mask information, which can tightly surround the image subject area. Furthermore, the distribution of the image subject in the entire candidate product image 601 can be determined by analyzing the distance between the image subject and the upper, lower, left, and right boundaries within the minimum bounding rectangle area, and the size of the area that needs to be extended in each direction can be dynamically calculated based on a preset margin threshold to obtain the extended area; further, the model can be used to perform preliminary image restoration based on the extended area; finally, the extended area of ​​the candidate product image after the preliminary image restoration is completed based on the relevant extension generation model to obtain the processed candidate product image, that is, Figure 6 The candidate product images 602 in .

[0116] See also Figure 7 , Figure 7 This is a schematic diagram of the effect of extending an image body provided by an exemplary embodiment of the present application. In candidate product image 801, the vessel area is missing. This image has an image body integrity score less than a preset integrity score threshold. After the extension processing as described above, candidate product image 802 is obtained. The image body integrity score of the processed candidate product image is greater than or equal to the preset integrity score threshold.

[0117] In the above embodiment, by extracting mask information from the main image, determining the extended area to be completed based on the minimum bounding rectangle, combining it with the model for preliminary restoration, and then completing the extended area with high quality, the integrity and aesthetics of the main image are effectively improved. This not only repairs missing edges and completes important details such as product containers, but also enhances the visual quality and professionalism of the image on display, significantly improving the quality of candidate product images.

[0118] In other embodiments, in the above S202, the candidate product image is detected according to the preset detection dimension to obtain the detection information corresponding to the candidate product image, which may also include: performing a preprocessing operation on the candidate product image, wherein the preprocessing operation may include any one or more of format standardization, resolution normalization, color adjustment, noise suppression, etc.; and then inputting the preprocessed candidate product image into a preset image feature extraction model to extract the multi-level image semantic features of the preprocessed candidate product image. The multi-level image semantic features can be used to support quality judgment of multiple preset detection dimensions; then, for each preset detection dimension, a corresponding natural language prompt information is generated. (such as "whether the image is clear", "whether there is any product obstruction", etc.), and input the multi-level image semantic features and natural language prompt information into the preset multimodal quality assessment model together to obtain the quality score of each preset detection dimension output by the multimodal quality assessment model; further, the score of each preset detection dimension is compared with the corresponding preset quality target to determine whether the preset detection dimension meets the standard, and combined with the attention distribution or posterior distribution within the model, the corresponding score confidence value is output; the score, judgment result (whether it meets the standard), confidence information, optional abnormal mark or semantic explanation text of each preset detection dimension are aggregated to generate structured detection information.

[0119] In the embodiment of the present application, by introducing image preprocessing, multi-level feature extraction, natural language prompts and multimodal quality assessment models, it is possible to achieve intelligent evaluation of candidate product images under multiple preset detection dimensions, effectively improving the accuracy and stability of detection. In addition, the detection information not only includes the scores of each preset detection dimension, but also includes judgment marks, confidence information and semantic explanation text, which can provide better explanatory support while improving detection accuracy, helping users understand the source of image problems and assisting in their subsequent image optimization processing. Furthermore, through the natural language prompt mechanism, it is possible to flexibly expand the detection requirements under different product categories or quality rules, and has good scalability.

[0120] In one embodiment, Figure 8 As shown, another image processing method is provided, which is applied to a server and includes the following steps: S801: Acquire candidate product images.

[0121] Specifically, S801 is consistent with S201 and will not be repeated here.

[0122] S802: Evaluate the candidate product image according to a preset detection dimension to obtain detection information corresponding to the candidate product image, where the detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension.

[0123] Specifically, S802 is consistent with S202 and will not be repeated here.

[0124] S803: When the detection information indicates that the candidate product image meets the first quality condition corresponding to the preset detection dimension, obtain user evaluation information corresponding to the candidate product image, where the first quality condition is that the quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated by the user's behavioral data and / or subjective evaluation data.

[0125] In one embodiment, the subjective evaluation data includes at least one of the following: image subject evaluation information, image background evaluation information, and image overall evaluation information; in addition, the subjective evaluation data may also include supplementary review information.

[0126] Optionally, when the subjective evaluation data includes image subject assessment information, the second quality condition includes: a deformation score of the image subject in the candidate product image is less than a first deformation score threshold; when the subjective evaluation data includes image background assessment information, the second quality condition includes: a deformation score of the image background in the candidate product image is less than a second deformation score threshold. The deformation score may be a quantitative value used to measure the degree of distortion of the image subject in the candidate product image relative to a standard form in terms of shape, proportion, and structure. A larger deformation score indicates a more severe deformation. The specific calculation process may include: first, performing subject detection and segmentation on the candidate product image to extract the main area of ​​the image; then, obtaining the standard reference form of the corresponding product, which can be obtained from a high-quality template image of a similar product, a standard product photo or a three-dimensional model rendering; then, extracting geometric structural features (such as contour edges, aspect ratio, key point positions, local curvature distribution) and texture features from the main area of ​​the image and the standard reference form respectively, and calculating the difference between the main area of ​​the image and the standard reference form through shape matching and deformation analysis algorithms (such as affine / perspective transformation fitting based on key points, shape context matching, etc.); finally, normalizing the difference, and the resulting value is the deformation score, which can be taken within a uniform range.

[0127] Optionally, a first deformation score threshold and a second deformation score threshold can be used to determine whether the degree of deformation of the image subject and image background in the candidate product image meets quality requirements, respectively. Specifically, based on a large number of product image samples with annotated deformation levels, manually annotating levels such as "slight deformation" and "severe deformation" and combining them with the aforementioned deformation score calculation method, the deformation score distribution range of high-quality samples (labeled as "slight deformation" or "no deformation") can be statistically calculated. Furthermore, the cutoff values ​​can be determined based on different requirements for subject integrity and background naturalness in actual applications. For example, the requirements for the image subject can be more stringent, and the upper quantile of the deformation score distribution of high-quality subject samples (such as the 90th percentile) can be selected as the first deformation score threshold. Since background deformation has a relatively small impact on overall quality, the higher quantile of the deformation score distribution of high-quality background samples (such as the 95th percentile) can be selected as the second deformation score threshold. This ensures that the degree of deformation of both the image subject and image background is within an acceptable range, while balancing image quality and screening pass rate.

[0128] Optionally, when the subjective evaluation data includes overall image evaluation information, the second quality condition includes: the image naturalness score of the candidate product image is greater than a preset naturalness score threshold, the image composition score of the candidate product image is greater than a preset composition score threshold, and the image color balance score of the candidate product image is greater than a preset color balance score threshold.

[0129] The naturalness score is a quantitative value used to measure whether the overall appearance of a candidate product image presents a natural effect that conforms to the human eye's perception habits. A higher naturalness score indicates a more natural image and greater visual comfort. The naturalness score calculation process may include: first, performing a multi-dimensional image quality feature analysis on the candidate product image to extract multiple indicators, including color distribution characteristics (such as color saturation, hue distribution, and color contrast), brightness distribution characteristics (such as global brightness uniformity and local brightness gradient), and texture and detail characteristics (such as clarity, sharpness, and noise level). Subsequently, these multiple indicators are input into a trained naturalness assessment model. This naturalness assessment model can perform supervised learning based on a large-scale natural image sample and manually annotated scores, thereby learning the mapping relationship between different features and the human eye's subjective evaluation of naturalness. Finally, the naturalness assessment model outputs a predicted value, which is the naturalness score, which can take values ​​within a uniform range.

[0130] Optionally, with respect to the method for determining the preset naturalness score threshold, a trained picture naturalness assessment model can be used to analyze a large number of natural images and non-natural images with subjective evaluation labels of the human eye to obtain a picture naturalness score distribution within a unified range (such as 0 to 100); then, the score differences between natural images and non-natural images are statistically analyzed in the verification data set, and a score that can better distinguish the two types of images is selected as the preset naturalness score threshold according to actual needs, such as taking the low segment quantile of the natural image score distribution; in addition, the preset naturalness score threshold can also be fine-tuned and pre-updated according to the category of the image subject, shooting conditions or time period.

[0131] In addition, the image composition score can be a quantitative value used to measure the rationality of the candidate product image in terms of subject layout, framing ratio, and spatial distribution. It can be used to reflect whether the image conforms to relevant aesthetic principles or product display specifications. The specific calculation process may include: first, using object detection or semantic segmentation algorithms to locate the image subject area; then, based on the image subject area, extracting composition features, such as the proximity of the image subject's position in the image to the rule of thirds or the golden section, whether the proportion of the subject to the image area is reasonable, and the spatial separation between the subject and the background; then, inputting the above composition features into a composition quality assessment model trained based on large-scale manually scored samples to obtain a composition score, and normalizing the results to a uniform range to obtain the image composition score.

[0132] In one embodiment, to determine the preset composition score threshold, a composition quality assessment model can be used to analyze sample images with subjective human ratings to obtain a uniform distribution of composition scores. Then, in a validation dataset, the scores of images that meet aesthetic principles or display specifications are compared with those that do not. The score that effectively distinguishes the two categories is selected as the preset composition score threshold, for example, the low quantile of the distribution of qualified sample scores. Finally, the preset composition score threshold can be fine-tuned based on different product categories, shooting scenarios, or application requirements.

[0133] Optionally, the image color balance score can be a quantitative value used to measure the naturalness and harmony of the color distribution of the candidate product image, and to assess whether there is color cast, abnormal color temperature, or hue imbalance. The specific calculation process may include: first, converting the image to the hue-saturation-lightness color space and extracting color features such as the overall color histogram, average color temperature, and hue distribution to analyze color deviations, including the balance of each color channel and the consistency of the main and background colors; finally, inputting the color features into a color balance quality assessment model trained with manually annotated samples, and normalizing the output of the color balance quality assessment model to a uniform range to obtain the image color balance score.

[0134] In one embodiment, for the determination of the preset color balance score threshold, the sample images with human subjective scores can be analyzed by using the color balance quality evaluation model to obtain the color balance score distribution in a unified range; then the score difference between the qualified images with natural and coordinated colors and the unqualified images with obvious color deviation or color imbalance in the verification data set is compared, and the score that can effectively distinguish the two types of images is selected as the preset color balance score threshold, for example, the low quantile of the qualified sample score distribution. Finally, the preset color balance score threshold can also be appropriately adjusted according to different categories, shooting conditions or actual needs.

[0135] It should be noted that the problem corresponding to the candidate commodity image can be determined in advance to collect the user's answer to form the subjective evaluation data. The answer can be a multiple-choice question, a score, a text comment or other structured data.

[0136] Next, referring to Figure 9 , Figure 9 A schematic diagram of an image review dimension provided by an exemplary embodiment of the present application, still taking the candidate commodity image as the candidate dish image, Figure 9 The candidate dish image 901 can include the image subject (including the dish and the utensil), the image background, and the dish feature information. The problem corresponding to the candidate commodity image can include at least one of the following types of problems: the first type of problem corresponding to the image subject, the second type of problem corresponding to the image background, and the third type of problem corresponding to the image as a whole. In addition, other supplementary questions, i.e. the fourth type of problem, can also be included. The first type of problem is used for the user to answer to collect the image subject evaluation information determined by the user, the second type of problem is used for the user to answer to collect the image background evaluation information determined by the user, the third type of problem is used for the user to answer to collect the image overall evaluation information determined by the user, and the fourth type of problem is used for the user to answer to collect the supplementary review information determined by the user.

[0137] Specifically, the first category of questions may include at least one of the following: dish deformation (used to obtain dish deformation information, such as whether the dish is clear, whether the material and shape of the ingredients are realistic, whether the size and proportion of the ingredients are realistic, whether the ingredients and the utensils are fused and deteriorated, etc.), utensil deformation (used to obtain utensil deformation information), and subject decoration deformation (used to obtain subject decoration deformation information, such as whether the utensils are deformed, distorted, missing, and whether the texture of the utensils is realistic.). The second category of questions may include at least one of the following: background deformation (used to obtain background deformation information, such as whether the tabletop is deformed or distorted, whether the tabletop material is realistic, whether there are multiple materials spliced ​​together, whether the tabletop perspective is consistent with the perspective of the image subject and whether they are on the same horizontal plane, whether the image background is clear, whether there is excessive blur (not lens blur), whether there is a sense of splicing, and whether the subject is suspended in the background, etc.) and background decoration deformation (used to obtain background decoration deformation information, such as whether the shape and material of the background decorations and embellishments are realistic, whether there are obvious deformations, distortions, or missing, etc.). The third category of questions may include at least one of the following: image naturalness (used to obtain an image naturalness score, which reflects whether the image's lighting and shadows are realistic and natural), image composition (used to obtain an image composition score, which reflects whether the image composition is complete and whether it can be cropped or cut out), and color balance (used to obtain an image color balance score, which reflects whether the color vividness of the image subject and background is realistic, whether the overall vividness of the image is realistic, and whether the color coordination is reasonable). Specifically, image composition issues may include: whether the size of the image subject is within a preset safe area to avoid being too large or too small, and whether it can be cropped 1:1; whether the image subject is centered, fully prominent, and clearly defined; whether the image subject is placed horizontally, without tilt (long utensils can be placed at an angle); whether the shooting angle can highlight the content of the image subject; and whether the composition contains splicing elements.

[0138] It should be noted that the above-mentioned questions can be provided to the user in the form of multiple-choice questions, so that the user can select the options corresponding to each question based on the candidate product images viewed. For example, each question in the first and second categories can correspond to three options representing deformation scores, and the deformation scores are "undeformed", "generally deformed" and "severely deformed" from low to high. For the first category of questions, if the obtained dish deformation information, utensil deformation information and main body decoration deformation information do not include the option representing the most serious deformation score, that is, "severely deformed", it can be determined that the deformation score of the image subject in the candidate product image is less than the first deformation score threshold; if the obtained dish deformation information, utensil deformation information and main body decoration deformation information include the option representing the most serious deformation score, that is, "severely deformed", it can be determined that the deformation score of the image subject in the candidate product image is greater than or equal to the first deformation score threshold. For the second type of questions, if the obtained background deformation information and background decoration deformation information do not include the option with the most severe deformation score, i.e., "severe deformation," then the deformation score of the image background in the candidate product image can be determined to be less than the second deformation score threshold. If the obtained background deformation information and background decoration deformation information include the option with the most severe deformation score, i.e., "severe deformation," then the deformation score of the image background in the candidate product image can be determined to be greater than or equal to the second deformation score threshold. Each question in the third type of questions can have three corresponding score options: "1 point," "3 points," and "5 points." The preset naturalness score threshold, preset composition score threshold, and preset color balance score threshold can all be set to 1 point. Thus, if a "1 point" answer is given to each question, it indicates that the following conditions, including the second quality condition, are not met: the candidate product image's image naturalness score is greater than the preset naturalness score threshold, the candidate product image's image composition score is greater than the preset composition score threshold, and the candidate product image's image color balance score is greater than the preset color balance score threshold. Therefore, the candidate product image can be determined to not meet the second quality condition.

[0139] See also Figure 10 , Figure 10 A schematic diagram of deformed candidate product images provided for an exemplary embodiment of the present application, wherein the dish in candidate product image 1001 is deformed, that is, the image body is deformed; the image background in candidate product image 1002 is deformed; and the utensils in candidate product image 1003 are deformed, that is, the image body is deformed.

[0140] The above-described embodiment effectively enhances the accuracy of quality assessment and fine-grained control over candidate product images by introducing multi-dimensional evaluation information and combining it with an interactive review mechanism. By setting deformation thresholds and image scoring criteria, a comprehensive assessment of image authenticity and aesthetics is achieved, ensuring that the visual presentation of images meets platform requirements. Furthermore, the use of structured questions to guide user evaluations helps standardize the review process, improve review efficiency and consistency, and provide strong support for building a high-quality product image library.

[0141] S804: When the user evaluation information indicates that the candidate product image meets the second quality condition, obtain the coding information to be watermarked.

[0142] In some embodiments, when user evaluation information includes behavioral data, user behavioral data during the display of candidate product images can be statistically analyzed based on preset behavioral indicator rules. Specifically, multiple behavioral indicators can be constructed for this behavioral data, such as click-through rate, average dwell time, conversion rate, bounce rate, and interaction frequency, and corresponding weighting parameters can be set based on the importance of each behavioral indicator. Furthermore, a comprehensive score is generated for the behavioral data based on a preset machine learning scoring model, which is trained using historical image data and corresponding user behavioral feedback. The machine learning scoring model takes as input the standardized behavioral indicators and outputs a user acceptance score for the image. This user acceptance score can then be compared with a preset acceptance score threshold to determine whether the candidate product image meets the second quality condition. For example, if the user acceptance score is greater than the acceptance score threshold, the candidate product image is determined to meet the second quality condition. If the user acceptance score is less than the acceptance score threshold, the candidate product image is determined to not meet the second quality condition. Furthermore, outlier processing can be performed on the collected behavioral data. For example, when the behavioral data of an image deviates significantly from the statistical distribution of similar images (the click-through rate exceeds a specific threshold but the dwell time is less than the preset dwell time), anomaly detection and elimination of abnormal samples can be performed by setting upper and lower limits, or based on cluster analysis, to avoid evaluation distortion due to data anomalies.

[0143] The acceptance score threshold can be determined by comparing historical image user acceptance scores with actual user feedback. Specifically, existing behavioral data (click-through rate, dwell time, conversion rate, etc.) and corresponding historical user feedback can be used to calculate and store the user acceptance score for each historical image. This is then statistically analyzed against manually determined high-quality or high-conversion images to determine a cutoff point that effectively distinguishes high-quality from low-quality images. For example, the user acceptance score at the 80th percentile of the distribution can be used as the acceptance score threshold. The acceptance score threshold can also be fine-tuned based on recent actual data and performance testing.

[0144] In some embodiments, when the user evaluation information includes behavioral data and subjective evaluation data, the determination of the second quality condition can be based on a joint decision made based on the comprehensive judgment results of the two dimensions of behavioral data and subjective evaluation data, so as to take into account the user's actual interactive feedback and perceived aesthetic judgment.

[0145] Optionally, during the joint decision-making phase, if both the subjective evaluation data and the behavioral data satisfy the corresponding second quality condition, the user evaluation information may be determined to indicate that the candidate product image satisfies the second quality condition. Specifically, if the subjective evaluation data does not meet the quality condition, the user evaluation information may be directly determined to indicate that the candidate product image does not meet the second quality condition, and the behavioral data may be determined to meet the quality condition only after the subjective evaluation data meets the quality condition.

[0146] Optionally, the behavioral data and the subjective rating data can be combined according to weights to calculate a comprehensive score, which is then compared with a comprehensive threshold. When the comprehensive score is greater than the comprehensive threshold, it is determined that the user evaluation information indicates that the candidate product image meets the second quality condition. The comprehensive threshold can be determined by comparing and analyzing the comprehensive scores of historical product images with the true quality labels. Specifically, the behavioral score and the subjective score can be unified into the same numerical range first, and the comprehensive score can be calculated according to the set weights; then, in the historical data, the dividing point at which the comprehensive score can better distinguish between high-quality and low-quality images is determined, for example, the 80th percentile of the comprehensive score distribution is taken as the comprehensive threshold. In addition, the threshold can be fine-tuned according to actual needs.

[0147] In one embodiment, the watermark to be added can be a dark watermark (also known as a hidden watermark or digital watermark), which can be used for protection and content authentication. Dark watermarks in images do not affect the visual quality of the image and are difficult to remove manually, providing covert identification of ownership while maintaining image quality.

[0148] S805: Extracting spectrum information of the candidate product image.

[0149] It can be understood that the spectrum information of the candidate commodity image is extracted, which means that the candidate commodity image is converted from the spatial domain (pixel) to the frequency domain, so as to more stably and covertly embed the watermark.

[0150] S806: Embed the encoded information into the spectrum information to obtain processed spectrum information.

[0151] The encoded information is embedded into the spectrum information, which can be a small adjustment of the frequency coefficient at a specific position (such as the middle frequency region) of the spectrum information, and the encoded information is injected.

[0152] S807: Perform inverse transform operation on the processed spectrum information to generate the target commodity image with added watermark.

[0153] Optionally, the inverse transform operation can refer to converting the frequency domain data (spectrum information) back to the spatial domain (pixel), which can be the inverse operation of extracting the spectrum information of the candidate commodity image.

[0154] Referring to Figure 11 , Figure 11 The flowchart of the watermark embedding method provided by the exemplary embodiment of the present application is shown. The candidate commodity image 1101 is the original image without added watermark. After the spectrum extraction, watermark encryption, transform output and other processes, the candidate commodity image 1101 can be restored to the original image format, thereby obtaining a candidate commodity image 1102 with unchanged appearance but containing hidden watermark information.

[0155] S808: Send the target commodity image to the user terminal to display the target commodity image on the user terminal.

[0156] Specifically, S808 is consistent with S205, which will not be described here.

[0157] In the above embodiment, the steganographic watermark is embedded in the candidate commodity image after meeting the high-quality standard, which realizes the protection and content authentication of the image without affecting the visual effect of the image. By embedding the encoded information in the spectrum information of the image and using the frequency domain steganographic technology, the watermark has high concealment and tamper resistance, and is not easy to be removed or damaged by human beings, which effectively enhances the content protection ability of the target commodity image, prevents unauthorized use and illegal dissemination, protects the platform and creator rights, and does not affect the image quality, improves the user experience and the professionalism and credibility of the platform content.

[0158] The embodiments of the present application will be described in conjunction with the overall image processing method, please refer to Figure 12 , Figure 12A flowchart of an image processing method provided as an exemplary embodiment of the present application is provided. First, candidate product images 1201 can be automatically generated based on a model. Automated quality inspection can then be performed on candidate product images 1201 to filter out those that do not meet a first quality requirement. Furthermore, if specific issues exist with candidate product images 1201, targeted solutions can be implemented, such as eliminating redundant information or extending the main image, to ensure that the processed candidate product images 1201 meet the first quality requirement. After automated processing, candidate product images 1201 enter a manual review phase, where a multi-dimensional assessment is performed based on pre-defined evaluation criteria to determine whether candidate product images 1201 meet a second quality requirement. If candidate product images 1201 meet the second quality requirement, content protection processing can be performed on candidate product images 1201, namely, a dark watermark (specifically, using wavelet transform, Fourier transform, etc.) can be added to candidate product images 1201 to track image ownership and prevent theft. Ultimately, a high-quality candidate product image 1202 with a dark watermark is obtained.

[0159] The image processing method provided in this application automatically generates product images through AIGC. Combined with automatic quality inspection, image restoration (such as removing redundant information and subject extension), and manual multi-dimensional evaluation, it achieves efficient image screening and optimization. Once quality standards are met, a hidden dark watermark is further embedded to enhance image protection. This method improves image production efficiency and quality control capabilities, reduces image discard rates, and enhances the security of product images.

[0160] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0161] Based on the inventive concept of the above image processing method, Figure 13 As shown, the embodiment of the present application further provides an image processing device 1300 for implementing the above-mentioned image processing method. The image processing device 1300 includes: The first acquisition module 1301 is used to acquire candidate product images; Detection module 1302, configured to detect the candidate product image according to a preset detection dimension and obtain detection information corresponding to the candidate product image, where the detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension; A second acquisition module 1303 is configured to acquire user evaluation information corresponding to the candidate product image if the detection information indicates that the candidate product image meets a first quality condition corresponding to a preset detection dimension, wherein the first quality condition is that the quality targets corresponding to each preset detection dimension are met, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; Determining module 1304, configured to determine a target product image based on the candidate product images if the user evaluation information indicates that the candidate product images meet a second quality condition, where the second quality condition is used to measure the content performance of the candidate product images in user interaction; The sending module 1305 is configured to send the target product image to the user terminal so that the user terminal displays the target product image.

[0162] In one embodiment, the first acquisition module 1301 is specifically used to: obtain preset product information, which includes the product name and the corresponding product demand weight; determine whether the product demand weight is greater than the preset weight threshold; if the product demand weight is greater than the preset weight threshold, input the product name into the first preset model, and determine the output image of the first preset model as the candidate product image; wherein the first preset model is a text-driven image generation model.

[0163] In one embodiment, the preset product information also includes an initial product image; the first acquisition module 1301 is further specifically used to: when the product demand weight is less than or equal to a preset weight threshold, input the product name and the initial product image into the second preset model, and determine the output image of the second preset model as the candidate product image; wherein the second preset model is an image generation model based on joint input of images and text.

[0164] In one embodiment, the first quality condition includes quality targets corresponding to each preset detection dimension, and the first quality condition is that the quality targets corresponding to each preset detection dimension are met; the detection module 1302 is specifically used to: detect the candidate product image according to each preset detection dimension, and obtain the detection results corresponding to the candidate product image under each preset detection dimension; based on the detection results, determine whether the candidate product image meets the quality targets of each preset detection dimension; if so, generate detection information indicating that the candidate product image meets the first quality condition corresponding to the preset detection dimension; if not, generate detection information indicating that the candidate product image does not meet the first quality condition corresponding to the preset detection dimension.

[0165] In one embodiment, the preset detection dimensions include at least one of the following: redundant information detection dimension, clarity detection dimension, lighting detection dimension, integrity detection dimension, copy detection dimension, boundary line detection dimension and consistency detection dimension; wherein, the quality target corresponding to the redundant information detection dimension is: the candidate product image does not contain redundant information; the quality target corresponding to the clarity detection dimension is: the clarity of the candidate product image is greater than a preset clarity threshold; the quality target corresponding to the lighting detection dimension is: the exposure of the candidate product image is within a preset exposure range; the quality target corresponding to the integrity detection dimension is: the integrity score of the image subject in the candidate product image is greater than a preset integrity score threshold; the quality target corresponding to the copy detection dimension is: the candidate product image is not a copy image, and a copy image is an image obtained by re-shooting an image displayed in an existing medium; the quality target corresponding to the boundary line detection dimension is: the image edge of the candidate product image does not contain a boundary line; the quality target corresponding to the consistency detection dimension is: the similarity score between the image subject in the candidate product image and the corresponding product name is greater than or equal to the preset similarity threshold.

[0166] In one embodiment, the image processing apparatus further includes: a first processing module configured to, when the detection information indicates that the candidate product image does not meet the first quality condition, eliminate redundant information if the candidate product image contains redundant information, and return to executing the step of detecting the candidate product image according to the preset detection dimension based on the processed candidate product image; or The second processing module is configured to, when the detection information indicates that the candidate product image does not meet the first quality condition, perform extension processing on the candidate product image if the integrity score of the image body in the candidate product image is less than a preset integrity score threshold, and return to execute the step of detecting the candidate product image according to the preset detection dimension based on the processed candidate product image.

[0167] In one embodiment, the first processing module is specifically configured to: detect redundant information included in the candidate product image to determine the redundant information location area; and perform image erasing processing on the redundant information location area to generate a processed candidate product image.

[0168] In one embodiment, the second processing module is specifically used to: extract mask information corresponding to the image body in the candidate product image; determine the corresponding minimum bounding rectangle area based on the mask information; determine the image area to be extended based on the minimum bounding rectangle area; and perform extension processing on the image area to be extended to obtain the processed candidate product image.

[0169] In one embodiment, the subjective evaluation data includes at least one of the following: image subject evaluation information, image background evaluation information, and image overall evaluation information; when the subjective evaluation data includes image subject evaluation information, the second quality condition includes: the deformation score of the image subject in the candidate product image is less than a first deformation score threshold; when the subjective evaluation data includes image background evaluation information, the second quality condition includes: the deformation score of the image background in the candidate product image is less than a second deformation score threshold; when the subjective evaluation data includes image overall evaluation information, the second quality condition includes: the image picture naturalness score of the candidate product image is greater than a preset naturalness score threshold, and the picture composition score of the candidate product image is greater than a preset composition score threshold, and the image color balance score of the candidate product image is greater than a preset color balance score threshold.

[0170] In one embodiment, the candidate product image is a dish image, and the image body of the candidate product image includes a dish area and a container area for loading the dish.

[0171] In one embodiment, the determination module 1304 is specifically used to: obtain the coding information to be watermarked; extract the spectrum information of the candidate product image; embed the coding information into the spectrum information to obtain the processed spectrum information; and perform an inverse transformation operation on the processed spectrum information to generate the target product image after adding the watermark.

[0172] Each module in the image processing apparatus 1300 may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0173] Based on the inventive concept of the above image processing method, Figure 14 As shown, the embodiment of the present application further provides an image processing and display device 1400 for implementing the above-mentioned image display method. The image display device 1400 includes: Receiving module 1401 is configured to receive a target product image sent by a server, wherein the target product image is an image determined by the server based on the candidate product image when user evaluation information indicates that the candidate product image satisfies a second quality condition, where the second quality condition is used to measure the content performance of the candidate product image in user interaction; user evaluation information is information corresponding to the candidate product image obtained by the server when detection information indicates that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension, where the first quality condition is that all quality targets corresponding to each preset detection dimension are satisfied, and user evaluation information is information generated from user behavior data and / or subjective evaluation data; detection information is information obtained by the server through detection of the candidate product image according to the preset detection dimension, and the detection information is used to characterize the objective quality performance of the candidate product image under the preset detection dimension; The display module 1402 is configured to display the target product image in response to a triggering operation on the product indicated by the target product image.

[0174] The embodiment of the present application also provides an electronic device, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 15 As shown. The electronic device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the electronic device is used to store active configuration information. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal through a network connection. The processor of the electronic device executes a computer program to implement an image processing method or an image display method.

[0175] Those skilled in the art will understand that Figure 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0176] The present application also provides a computer storage medium having instructions stored therein that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of the above-described embodiments. If the components of the electronic device described above are implemented as software functional units and sold or used as independent products, they may be stored in the computer-readable storage medium described above.

[0177] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted via a computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. Available media may be magnetic media (eg, floppy disks, hard disks, tapes), optical media (eg, digital versatile discs (DVDs)), or semiconductor media (eg, solid-state drives (SSDs)).

[0178] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. The technical features of this embodiment and the implementation scheme can be combined in any manner unless they conflict.

[0179] The above embodiments are merely preferred embodiments of the present application and are not intended to limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements made to the technical solutions of the present application by ordinary technicians in this field should fall within the scope of protection determined by the claims.

[0180] The above described specific embodiments of the application. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the process depicted in the figures does not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

Claims

1. An image processing method, characterized in that: Applied to a server, the method includes: Obtain candidate product images; Detect the candidate product image according to a preset detection dimension to obtain detection information corresponding to the candidate product image, where the detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension; If the detection information indicates that the candidate product image satisfies a first quality condition corresponding to the preset detection dimension, obtaining user evaluation information corresponding to the candidate product image, wherein the first quality condition is that quality targets corresponding to each preset detection dimension are satisfied, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; determining a target product image based on the candidate product image if the user evaluation information indicates that the candidate product image satisfies a second quality condition, wherein the second quality condition is used to measure content performance of the candidate product image in user interaction; The target product image is sent to a user terminal, so that the user terminal displays the target product image.

2. The method according to claim 1, wherein The step of obtaining candidate product images includes: Obtaining preset product information, wherein the preset product information includes the product name and the corresponding product demand weight; Determining whether the commodity demand weight is greater than a preset weight threshold; If the product demand weight is greater than the preset weight threshold, inputting the product name into a first preset model, and determining an output image of the first preset model as a candidate product image; Among them, the first preset model is a text-driven image generation model.

3. The method according to claim 2, wherein The preset product information also includes an initial product image; The acquiring of candidate product images further includes: If the product demand weight is less than or equal to the preset weight threshold, inputting the product name and the initial product image into a second preset model, and determining the output image of the second preset model as a candidate product image; Among them, the second preset model is an image generation model based on joint input of images and texts.

4. The method according to claim 1, wherein The first quality condition includes quality targets corresponding to each preset detection dimension, and the first quality condition is that the quality targets corresponding to each preset detection dimension are all met; The detecting the candidate product image according to the preset detection dimension to obtain detection information corresponding to the candidate product image includes: Detect the candidate product image according to each preset detection dimension, and obtain the detection results corresponding to the candidate product image under each preset detection dimension; Determining whether the candidate product image meets the quality targets of each preset detection dimension based on the detection result; If yes, generating detection information indicating that the candidate product image satisfies the first quality condition corresponding to the preset detection dimension; If not, detection information is generated to indicate that the candidate product image does not meet the first quality condition corresponding to the preset detection dimension.

5. The method according to claim 4, wherein The preset detection dimension includes at least one of the following: redundant information detection dimension, clarity detection dimension, illumination detection dimension, integrity detection dimension, copy detection dimension, boundary line detection dimension and consistency detection dimension; The quality target corresponding to the redundant information detection dimension is: the candidate product image does not contain redundant information; The quality target corresponding to the clarity detection dimension is: the clarity of the candidate product image is greater than a preset clarity threshold; The quality target corresponding to the illumination detection dimension is: the exposure of the candidate product image is within a preset exposure range; The quality target corresponding to the integrity detection dimension is: the integrity score of the image subject in the candidate product image is greater than a preset integrity score threshold; The quality target corresponding to the copy detection dimension is: the candidate product image is not a copy image, and the copy image is an image obtained by re-photographing an image displayed in an existing medium; The quality target corresponding to the boundary line detection dimension is: the image edge of the candidate product image does not contain a boundary line; The quality target corresponding to the consistency detection dimension is: the similarity score between the image body in the candidate product image and the corresponding product name is greater than or equal to a preset similarity threshold.

6. The method according to claim 1, wherein After detecting the candidate product image according to the preset detection dimension and obtaining detection information corresponding to the candidate product image, the method further includes: In a case where the detection information indicates that the candidate product image does not meet the first quality condition, if the candidate product image contains redundant information, the redundant information is eliminated, and based on the processed candidate product image, the step of detecting the candidate product image according to the preset detection dimension is returned to; or In a case where the detection information indicates that the candidate product image does not meet the first quality condition, if the integrity score of the image body in the candidate product image is less than a preset integrity score threshold, the candidate product image is extended, and based on the processed candidate product image, the step of detecting the candidate product image according to the preset detection dimension is returned to be executed.

7. The method according to claim 6, wherein The eliminating process of the redundant information includes: Detecting the redundant information included in the candidate product image and determining a redundant information location area; An image erasing process is performed on the redundant information location area to generate a processed candidate product image.

8. The method according to claim 6, wherein The extending process on the candidate product image includes: Extracting mask information corresponding to the image subject in the candidate product image; Determine a corresponding minimum circumscribed rectangular area based on the mask information; Determining the image area to be extended according to the minimum circumscribed rectangular area; The image region to be extended is extended to obtain a processed candidate product image.

9. The method according to claim 2, wherein The subjective evaluation data includes at least one of the following: image subject evaluation information, image background evaluation information, and image overall evaluation information; In a case where the subjective evaluation data includes the image subject evaluation information, the second quality condition includes: a deformation score of the image subject in the candidate product image is less than a first deformation score threshold; In a case where the subjective evaluation data includes the image background assessment information, the second quality condition includes: a deformation score of the image background in the candidate product image is less than a second deformation score threshold; When the subjective evaluation data includes the overall image evaluation information, the second quality condition includes: the image naturalness score of the candidate product image is greater than a preset naturalness score threshold, the image composition score of the candidate product image is greater than a preset composition score threshold, and the image color balance score of the candidate product image is greater than a preset color balance score threshold.

10. The method according to claim 9, wherein The candidate product image is a dish image, and the image body of the candidate product image includes a dish area and a container area for loading the dish.

11. The method according to claim 1, wherein The determining of the target product image based on the candidate product images includes: Get the encoding information to be added to the watermark; extracting spectrum information of the candidate product image; Embedding the coding information into the spectrum information to obtain processed spectrum information; An inverse transformation operation is performed on the processed spectrum information to generate a target product image with a watermark added.

12. An image display method, characterized in that: Applied to a user terminal, the method includes: Receive a target product image sent by a server, wherein the target product image is an image determined by the server based on the candidate product image when user evaluation information indicates that the candidate product image satisfies a second quality condition, the second quality condition being used to measure the content performance of the candidate product image in user interaction; the user evaluation information is information corresponding to the candidate product image obtained by the server when detection information indicates that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension, the first quality condition being that quality targets corresponding to each preset detection dimension are all satisfied, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; the detection information is information obtained by the server through detection of the candidate product image according to the preset detection dimension, the detection information being used to characterize the objective quality performance of the candidate product image under the preset detection dimension; In response to a triggering operation on the product indicated by the target product image, the target product image is displayed.

13. An image processing device, characterized in that: Applied to a server, the device includes: A first acquisition module is used to acquire candidate product images; a detection module, configured to detect the candidate product image according to a preset detection dimension and obtain detection information corresponding to the candidate product image, wherein the detection information is used to represent the objective quality performance of the candidate product image under the preset detection dimension; a second acquisition module configured to acquire user evaluation information corresponding to the candidate product image if the detection information indicates that the candidate product image satisfies a first quality condition corresponding to the preset detection dimension, the first quality condition being that quality targets corresponding to each preset detection dimension are satisfied, and the user evaluation information being information generated from user behavior data and / or subjective evaluation data; a determination module configured to determine a target product image based on the candidate product images if the user evaluation information indicates that the candidate product images meet a second quality condition, wherein the second quality condition is used to measure content performance of the candidate product images in user interaction; The sending module is configured to send the target product image to a user terminal so that the user terminal displays the target product image.

14. An image display device, characterized in that: Applied to a user terminal, the device includes: A receiving module is configured to receive a target product image sent by a server, wherein the target product image is an image determined by the server based on the candidate product image when user evaluation information indicates that the candidate product image satisfies a second quality condition, the second quality condition being used to measure the content performance of the candidate product image in user interaction; the user evaluation information is information corresponding to the candidate product image obtained by the server when detection information indicates that the candidate product image satisfies a first quality condition corresponding to a preset detection dimension, the first quality condition being that quality targets corresponding to each preset detection dimension are all satisfied, and the user evaluation information is information generated from user behavior data and / or subjective evaluation data; the detection information is information obtained by the server through detection of the candidate product image according to the preset detection dimension, the detection information being used to characterize the objective quality performance of the candidate product image under the preset detection dimension; The display module is configured to display the target product image in response to a triggering operation on the product indicated by the target product image.

15. An electronic device, characterized in that: include: A processor and a memory; the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.

16. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method for screening images containing specific features from a plurality of images

    CN113592960A

  • Commodity image generation method and device, electronic equipment and storage medium

    CN117593083A

  • Image generation method and device, program product and storage medium

    CN118154727A

  • Image aesthetics evaluation method based on subjective and objective attribute fusion

    CN118155053A

  • Data processing method and electronic equipment

    CN119476222A