Automatic on-cloud training method and device for e-commerce image model

Through automated cloud training methods, the problem of time-consuming and labor-consuming manual operation in e-commerce image model training is solved, and the training cycle is shortened and iterative efficiency is improved, ensuring the rapid improvement of model effect.

CN120070618APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108219.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing e-commerce image model training methods require a large number of repetitive operations manually, resulting in too long training cycles and being unable to achieve rapid iteration and effect improvement.

Method used

It provides an automated cloud training method for e-commerce image model, establishing a connection with cloud services through the client, transmitting image data to the cloud server, receiving and storing training data, using the cloud server to perform model training, and evaluating the passing rate of generated images through the scoring model, and automatically saving and retraining the model.

Benefits of technology

It significantly shortens the model training cycle, improves iteration efficiency, ensures the traceability and standardization of experiments, and can quickly improve the model effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070618A_ABST
    Figure CN120070618A_ABST
Patent Text Reader

Abstract

The invention provides an automatic on-cloud training method and device for an e-commerce image model, and the method comprises the steps: building a connection between a client and a cloud server, and transmitting image data to the cloud server; the cloud server receives and stores training data to obtain the training data; the cloud server performs e-commerce image model training through the training data to obtain a trained e-commerce image model; generating a set number of generated images through the trained e-commerce image model; each generated image is scored through a scoring model, an image score is obtained, and when the image score reaches a set threshold score, the generated image is qualified; otherwise, determining that the product is unqualified; when the qualified rate of the generated image reaches a set percentage, storing the trained e-commerce image model; otherwise, re-training the e-commerce image model until the qualified rate of the generated image reaches a set percentage, storing the trained e-commerce image model, quickly generating the required model, and accurately performing model evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an automated cloud training method and device for an e-commerce image model. Background Art

[0002] Existing designers only need to input prompt words into the e-commerce image model to generate images to obtain creative inspiration. Each type of e-commerce image model needs to undergo a lot of training before it can be used. The existing training methods have the following problems: Due to the lack of automated tools, it is necessary to manually upload local data to the cloud, download cloud data to the local, convert data formats and manage storage. These repetitive manual operations are not only time-consuming and labor-intensive, but also prone to errors. At the same time, parameter tuning and model evaluation in the training process also require manual login to the server for intervention and judgment, and it is impossible to achieve rapid training iteration and evaluation of multiple servers in one configuration. When faced with massive image data of hundreds of thousands in e-commerce scenarios, this manual login to multiple services to download, train and evaluate one by one leads to a long model training cycle, which seriously restricts the rapid improvement of model effects and timely iteration and update of products. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide an automated cloud training method and device for e-commerce image models.

[0004] In a first aspect, the present invention provides an automated cloud training method for an e-commerce image model, comprising the following steps:

[0005] Step 1: The client establishes a connection with the cloud service and transmits the image data to the cloud server;

[0006] Step 2: The cloud server receives and stores the training data;

[0007] Step 3: The cloud server trains the e-commerce image model using the training data to obtain a trained e-commerce image model;

[0008] Step 4: Generate a set number of generated images using the trained e-commerce image model;

[0009] Step 5. Score each generated image through the scoring model to obtain an image score. When the image score reaches the set threshold score, the generated image is qualified; otherwise, it is unqualified; when the qualified rate of the generated image reaches the set percentage, the trained e-commerce image model is saved; otherwise, the e-commerce image model is retrained until the qualified rate of the generated image reaches the set percentage, and the trained e-commerce image model is saved.

[0010] Second aspect, the present invention provides an e-commerce image model automated cloud training device, including:

[0011] A connection establishment module, which establishes a connection between the client and the cloud service and transmits image data to the cloud server;

[0012] A receiving and storing module, which receives and stores the data by the cloud server to obtain training data;

[0013] A training model module, which trains the e-commerce image model by the cloud server using the training data to obtain a trained e-commerce image model;

[0014] An image generation module, which generates a set number of generated images by the trained e-commerce image model;

[0015] An evaluation model module, which scores each generated image through a scoring model to obtain an image score. When the score reaches a set threshold score, the generated image is qualified; otherwise, it is unqualified. When the qualified rate of the generated images reaches a set percentage, the trained e-commerce image model is saved; otherwise, the e-commerce image model is retrained until the qualified rate of the generated images reaches the set percentage, and the trained e-commerce image model is saved.

[0016] Third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0017] Fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0018] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0019] The present invention integrates the automated process, significantly improving the model iteration efficiency; compared with the traditional manual operation, the training cycle is shortened by more than 70%, and at the same time, the traceability and standardization of the experiment are ensured.

[0020] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0022] Figure 1This is the flowchart of the method in Embodiment 1 of the present invention;

[0023] Figure 2 This is the structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed implementation manners

[0024] By providing an automated cloud training method and device for e-commerce image models in the embodiments of this application,

[0025] The overall idea of the technical solution in the embodiments of this application is as follows:

[0026] I. System connection and data transmission

[0027] Establish a connection

[0028] Use paramiko.SSHClient() to establish an SSH secure connection with the cloud server. At the same time, configure OSS object storage access credentials through oss2.Auth and oss2.Bucket, supporting storage on Alibaba Cloud, Tencent Cloud, and Huawei Cloud. Establish a persistent connection with the cloud server according to the configuration table.

[0029] Manage multiple concurrent connections based on the connection pool mechanism to ensure the stability and efficiency of data transmission. Set up an automatic reconnection and exception handling mechanism to prevent connection interruptions caused by network fluctuations.

[0030] Data compression and transmission

[0031] In view of the large number of e-commerce images, select Pigz for multi-threaded compression, significantly shortening the compression time. For large datasets, the system supports data splitting with 50,000 as a folder, and uses multiprocessing to create multiple processes to traverse the million-scale dataset for parallel and rapid dataset splitting.

[0032] Subsequently, use multiprocessing to enable multiple processes to perform multi-threaded compression on the separated datasets using Pigz, and batch obtain dataset compressed package files to be uploaded. In the transmission scheme, use multi-threaded concurrent.futures to upload multiple compressed data to the cloud service storage system, making full use of network bandwidth.

[0033] While uploading, record a file mapping dictionary locally, such as {"dataset1.zip":"oss: / / dataset / dataset1.zip"}, to ensure the accuracy of data synchronization and facilitate subsequent traceability.

[0034] II. Remote data processing and monitoring

[0035] Remote data decompression

[0036] Use paramiko.SSHClient().exec_command() to directly manipulate the cloud server on the local machine to run a Python script for decompressing data packets. Inside the script, use subprocess.Popen to execute the data packet decompression command unpigz on the cloud, and support starting multiple processes for decompression using multiprocessing according to the number of data packets to be decompressed as needed.

[0037] Data Transfer Monitoring and Recording

[0038] The system implements monitoring and logging of the data transfer process, including transfer speed, progress, error messages, etc. Each step is recorded as a.pkl file and stored at a specified local location for convenient problem location and performance optimization.

[0039] III. E-commerce Image Model Training

[0040] Training Environment Setup

[0041] Based on Docker container technology, encapsulate the training environment, which can quickly start the training environment for the e-commerce image model. Support GPU training through Nvidia-Docker.

[0042] Use torch.distributed to achieve multi-GPU parallel training. Support data loading for the e-commerce image model through the DistributedDataParallel class, and then distribute it to different graphics cards for multi-graphics card parallel training.

[0043] Adopt Tensorboard to record metrics such as loss curves and learning rate changes during the training process. Regularly use torch.save to save checkpoints, support resuming training from breakpoints, and support starting training at any time after locally modifying the configuration.

[0044] Model Evaluation and Image Generation

[0045] After the training of the e-commerce image model is completed, the system will automatically trigger an evaluation script. This script is executed on a remote machine through paramiko.SSHClient().exec_command(), loads the trained model and runs the text-to-image script in batches to generate images with a resolution of 1024x1024.

[0046] The generated images are saved to a specified directory (such as generated_images / ) on the remote machine to prepare for subsequent aesthetic scoring. The evaluation script uses the trained model to generate images according to preset prompts (such as "fashion women's clothing", "outdoor sports shoes", etc.).

[0047] IV. Image Aesthetic Scoring and Result Management

[0048] Aesthetic scoring

[0049] After generation is completed, call the qwenvl2 tool to perform aesthetic scoring on each image. qwenvl2 evaluates the image quality from multiple dimensions such as color, composition, and clarity, and combines the prompt words to ensure the relevance of the scoring results to the e-commerce scenario.

[0050] The prompt words used are: "Please evaluate the overall quality of this picture. Please score from the following aspects: 1. Whether the color of the picture is coordinated and natural, and whether the saturation and brightness are appropriate; 2. Whether the composition is reasonable, the main body is prominent, and the picture is balanced; 3. Whether the image is clear and whether the details are rich. Please give a comprehensive score between 0 and 1, where 1 means very good quality and 0 means poor quality. At the same time, please briefly explain the reasons for the score."

[0051] Processing of scoring results

[0052] After the aesthetic scoring is completed, the system saves the score of each image as a CSV file (such as aesthetic_scores.csv), and the content includes the image path, prompt words, aesthetic score, and generation time.

[0053] Subsequently, upload the scoring result file to cloud storage (such as Alibaba Cloud OSS), and record the mapping dictionary locally. The system will automatically download all scoring files from cloud storage and summarize them into the local training result storage directory for convenient subsequent unified analysis and management of the evaluation results.

[0054] V. System Monitoring and Alarms

[0055] The system will monitor key metrics during the training and evaluation processes in real time, such as the loss curve, learning rate changes, GPU utilization, number of images generated, aesthetic score distribution, etc.

[0056] When an anomaly occurs (such as insufficient GPU memory, network interruption, scoring anomaly, etc.), the system will send alarm notifications through multiple API interfaces:

[0057] Send email alarms using smtplib.SMTP().

[0058] Send SMS alarms using the SendSms interface of the Alibaba Cloud SMS service SDK.

[0059] Send DingTalk group message alarms using the robot.send_markdown() interface of the DingTalk open platform.

[0060] The alarm content includes key information such as the anomaly type, occurrence time, and anomaly stack to ensure that problems can be discovered and resolved in a timely manner.

[0061] Embodiment 1

[0062] As Figure 1 shown, this embodiment provides an automated cloud training method for an e-commerce image model, including the following steps:

[0063] Step 1: The client establishes a connection with the cloud service and transmits image data to the cloud server;

[0064] Step 2: The cloud server receives and stores it to obtain training data;

[0065] Step 3: The cloud server trains the e-commerce image model with the training data to obtain a trained e-commerce image model;

[0066] Step 4: Generate a set number of generated images through the trained e-commerce image model;

[0067] Step 5: Score each generated image through a scoring model to obtain an image score. When the score reaches a set threshold score, the generated image is qualified; otherwise, it is unqualified. When the qualified rate of the generated images reaches a set percentage, save the trained e-commerce image model; otherwise, retrain the e-commerce image model until the qualified rate of the generated images reaches the set percentage, and save the trained e-commerce image model.

[0068] In this embodiment, preferably, Step 1 is specifically as follows: The client uses paramiko.SSHClient to establish an SSH secure connection with the cloud server. At the same time, it configures OSS object storage access credentials through oss2.Auth and oss2.Bucket, and establishes a persistent connection with the cloud server according to a preset configuration table;

[0069] Select Pigz to compress the image data in multiple threads to obtain at least one dataset compressed package file to be uploaded; use the multi-threaded concurrent.futures to upload multiple compressed data to the cloud server in parallel; while uploading, record a file mapping dictionary locally on the client, and the file mapping dictionary is used to ensure the accuracy of data synchronization and backtracking;

[0070] Step 2 specifically includes: using paramiko.SSHClient().exec_command() to manipulate the cloud server on the client to run a Python script for decompressing the data packet. The Python script uses subprocess.Popen to execute the data packet decompression command unpigz on the cloud server, and uses multiprocessing to start multiple processes for decompression according to the number of data packets to be decompressed; on the cloud server, monitor and log the data transmission process, record it as a.pkl file and store it at a specified location on the local client to obtain training data.

[0071] In this embodiment, preferably, Step 3 specifically includes: encapsulating the training environment based on Docker container technology, supporting GPU training through Nvidia-Docker, using torch.distributed to achieve multi-GPU parallel training, supporting data loading of the e-commerce image model through the DistributedDataParallel class, and then distributing it to different graphics cards for multi-graphics card parallel training; using Tensorboard to record the metric data during the training process, and regularly saving checkpoints using torch.save; inputting the training data into the e-commerce image model for model training, and after the training is completed, obtaining the trained e-commerce image model.

[0072] In this embodiment, preferably, Step 4 specifically includes: using paramiko.SSHClient().exec_command to load the trained e-commerce image model on the cloud server, and generating a set number of generated images through the trained e-commerce image model; storing the generated images at a specified location.

[0073] In this embodiment, preferably, the scoring model in Step 5 specifically is:

[0074] Step 61: Perform color evaluation, composition evaluation, and image aesthetics evaluation on the input image, and mark good color, bad color, good composition, bad composition, good-looking image, and bad-looking image with 0 or 1 to obtain a marked image;

[0075] The color evaluation specifically includes: color space conversion and channel extraction; using cv2.cvtColor to convert the input image from the RGB color space to the HSV color space, and respectively extracting three channels of hue, saturation, and value;

[0076] Histogram distribution feature calculation: For the three channels of hue, saturation, and value, respectively use cv2.calcHist to calculate the corresponding histogram distribution features; use np.percentile to obtain the statistical values of the set quantiles of each channel;

[0077] Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of value; when S_mean is within the range of [0.3, 0.7] and V_mean is within the range of [0.4, 0.8], and at the same time H_std is less than the set first threshold, then mark the color as good as 1 and the color as bad as 0; when S_mean < 0.3 or > 0.7, or V_mean < 0.4 or > 0.8, or H_std is greater than or equal to the first set threshold, then mark the color as good as 0 and the color as bad as 1;

[0078] The composition evaluation is specifically as follows: Use cv2.Laplacian to calculate the sharpness score sharpness_score of the image;

[0079] Use the cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate the edge_strength;

[0080] Based on cv2.moments, calculate the centroid position of the image to obtain the symmetry_score;

[0081] Divide the image into a 3x3 grid, and calculate the feature response intensity golden_ratio_score at the golden section point position;

[0082] When sharpness_score > 0.7 and edge_strength is greater than the second set threshold, and at the same time symmetry_score > 0.6 and golden_ratio_score > 0.5, then mark the composition as good as 1 and the composition as bad as 0; when sharpness_score ≤ 0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score ≤ 0.6 or golden_ratio_score ≤ 0.5, then mark the composition as good as 0 and the composition as bad as 1;

[0083] The evaluation of the image aesthetics is specifically as follows:

[0084] The evaluation of the image looking good is specifically as follows: Use cv2.Laplacian to calculate the sharpness score model_sharpness_score of the image;

[0085] Use the QwenVL2 vision model to score the aesthetics of the image, set the prompt words, and obtain the qwen_beauty_score;

[0086] When the model_sharpness_score > 0.7 and the qwen_beauty_score > 0.7, mark the image as good-looking (1) if it is good-looking and as not good-looking (0) if it is not; in other cases, mark the image as not good-looking (0) if it is good-looking and as good-looking (1) if it is not;

[0087] Step 62: Build the evaluation model architecture. The evaluation model includes the aesthetic evaluation small model base and the aesthetic evaluation large model base:

[0088] Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the aesthetic evaluation small model base;

[0089] Use torch.load to load the CLIP model as the aesthetic evaluation large model base;

[0090] Delete the fully connected layer in the output part of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain model outputs of the same dimension;

[0091] Finally, the outputs are concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, corresponding to the probabilities of good color, bad color, good composition, bad composition, good-looking image of the product, and not good-looking image of the product respectively;

[0092] Model training: Input the labeled images into the evaluation model in sequence, and perform fine-tuning training on the evaluation model architecture: Use the torch.optim.Adam optimizer, and use the MSE loss function to calculate the error between the predicted score and the set value; when the error is less than the set threshold, stop training to obtain the trained evaluation model;

[0093] Step 63: Input the images to be screened into the trained evaluation model to obtain the probability values of six dimensions;

[0094] First, perform normalization using np.normalize: Set the weights of the probability of good color and the probability of bad color to 0.8, the weights of the probability of good composition and the probability of bad composition to 1.0, and the weights of the probability of good-looking image and the probability of not good-looking image to 1.2;

[0095] Use np.average to achieve weighted average to obtain the final evaluation score. The evaluation score = (probability value of good color × 0.8 + probability value of good composition × 1.0 + probability value of good-looking image of the product × 1.2) ÷ 3, and limit the score within the range of 0 to 1 through np.clip, which is the image score.

[0096] Based on the same inventive concept, this application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.

[0097] Embodiment 2

[0098] As Figure 2 shown, in this embodiment, an e-commerce image model automated cloud training apparatus is provided, including:

[0099] A connection establishment module, where the client establishes a connection with the cloud service and transmits image data to the cloud server;

[0100] A receiving and storing module, where the cloud server receives and stores to obtain training data;

[0101] A training model module, where the cloud server trains an e-commerce image model through the training data to obtain a trained e-commerce image model;

[0102] An image generation module, which generates a set number of generated images through the trained e-commerce image model;

[0103] An evaluation model module, which scores each generated image through a scoring model to obtain an image score. When the image score reaches a set threshold score, the generated image is qualified; otherwise, it is unqualified. When the pass rate of the generated images reaches a set percentage, the trained e-commerce image model is saved; otherwise, the e-commerce image model is retrained until the pass rate of the generated images reaches the set percentage, and the trained e-commerce image model is saved.

[0104] In this embodiment, preferably, the connection establishment module is specifically: the client uses paramiko.SSHClient to establish an SSH secure connection with the cloud server, and at the same time configures OSS object storage access credentials through oss2.Auth and oss2.Bucket. According to a preset configuration table, the client establishes a persistent connection with the cloud server;

[0105] Select Pigz for multi-threading to compress the image data to obtain at least one dataset compressed package file to be uploaded; use multi-threading concurrent.futures to upload multiple compressed data to the cloud server in parallel; while uploading, record a file mapping dictionary locally on the client, and the file mapping dictionary is used to ensure the accuracy of data synchronization and backtracking;

[0106] The receiving and storing module is specifically as follows: Use paramiko.SSHClient().exec_command() to manipulate the cloud server on the client side to run a Python script for decompressing data packets. The Python script uses subprocess.Popen to execute the data packet decompression command unpigz on the cloud server, and uses multiprocessing to start multiple processes for decompression according to the number of data packets to be decompressed; Monitor and log the data transmission process on the cloud server, record it as a.pkl file and store it at a specified location on the client side locally to obtain training data.

[0107] In this embodiment, preferably, the training model module is specifically as follows: Package the training environment based on Docker container technology, support GPU training through Nvidia-Docker, use torch.distributed to achieve multi-GPU parallel training, support data loading of the e-commerce image model through the DistributedDataParallel class, and then distribute it to different graphics cards for multi-graphics card parallel training; Use Tensorboard to record the metric data during training, and regularly use torch.save to save checkpoints; Input the training data into the e-commerce image model for model training, and after the training is completed, obtain the trained e-commerce image model.

[0108] In this embodiment, preferably, the training model module is specifically as follows: Load the trained e-commerce image model on the cloud server through paramiko.SSHClient().exec_command, and generate a set number of generated images through the trained e-commerce image model; Store the generated images at a specified location.

[0109] In this embodiment, preferably, the scoring model in the evaluation model module is specifically as follows:

[0110] A marking unit that performs color evaluation, composition evaluation, and image aesthetics evaluation on the input image, and marks good color, bad color, good composition, bad composition, good-looking image, and bad-looking image with 0 or 1 to obtain a marked image;

[0111] The color evaluation is specifically as follows: Color space conversion and channel extraction; Use cv2.cvtColor to convert the input image from the RGB color space to the HSV color space, and extract the three channels of hue, saturation, and value respectively;

[0112] Histogram distribution feature calculation: For the three channels of hue, saturation, and value, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; Use np.percentile to obtain the statistical values of the set quantiles for each channel;

[0113] Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of lightness; when S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and at the same time H_std is less than the set first threshold, then mark the color as good as 1 and the color as bad as 0; when S_mean < 0.3 or > 0.7, or V_mean < 0.4 or > 0.8, or H_std is greater than or equal to the first set threshold, then mark the color as good as 0 and the color as bad as 1;

[0114] The composition evaluation is specifically as follows: Use cv2.Laplacian to calculate the sharpness score sharpness_score of the image;

[0115] Use the cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate the edge_strength;

[0116] Based on cv2.moments, calculate the centroid position of the image to obtain the symmetry_score;

[0117] Divide the image into a 3x3 grid, and calculate the feature response intensity golden_ratio_score at the position of the golden section point;

[0118] When sharpness_score > 0.7 and edge_strength is greater than the second set threshold, and at the same time symmetry_score > 0.6 and golden_ratio_score > 0.5, then mark the composition as good as 1 and the composition as bad as 0; when sharpness_score ≤ 0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score ≤ 0.6 or golden_ratio_score ≤ 0.5, then mark the composition as good as 0 and the composition as bad as 1;

[0119] The evaluation of the image aesthetics is specifically as follows:

[0120] The evaluation of the image looking good is specifically as follows: Use cv2.Laplacian to calculate the sharpness score model_sharpness_score of the image;

[0121] Use the QwenVL2 vision model to score the aesthetics of the image, set the prompt words, and obtain the qwen_beauty_score;

[0122] When the model_sharpness_score > 0.7 and the qwen_beauty_score > 0.7, mark the image as good-looking (1) if it is, and as not good-looking (0) if not; in other cases, mark the image as not good-looking (0) if it is, and as good-looking (1) if not;

[0123] A training model unit for evaluating the model architecture construction. The evaluation model includes a base of an aesthetic evaluation small model and a base of an aesthetic evaluation large model:

[0124] Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the base of the aesthetic evaluation small model;

[0125] Use torch.load to load the CLIP model as the base of the aesthetic evaluation large model;

[0126] Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use an adaptive pooling layer to obtain model outputs of the same dimension;

[0127] Finally, the outputs are concatenated into the same output using the concat function, and a fully connected layer with an output dimension of 6 is used to obtain the final result. The output dimension includes 6 dimensions, corresponding to the probabilities of good color, bad color, good composition, bad composition, good-looking image commodity, and not good-looking image commodity respectively;

[0128] Model training: Input the labeled images into the evaluation model in sequence to perform fine-tuning training on the evaluation model architecture: Use the torch.optim.Adam optimizer and calculate the error between the predicted score and the set value using the MSE loss function; when the error is less than the set threshold, stop training to obtain the trained evaluation model;

[0129] A screening unit that inputs the images to be screened into the trained evaluation model to obtain probability values in six dimensions;

[0130] First, perform normalization using np.normalize: Set the weights of the probability of good color and the probability of bad color to 0.8, the weights of the probability of good composition and the probability of bad composition to 1.0, and the weights of the probability of good-looking image and the probability of not good-looking image to 1.2;

[0131] Use np.average to achieve weighted average to obtain the final evaluation score. The evaluation score = (probability value of good color × 0.8 + probability value of good composition × 1.0 + probability value of good-looking image commodity × 1.2) ÷ 3, and limit the score within the range of 0 to 1 through np.clip, which is the image score.

[0132] Since the device described in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method described in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated herein. Any device adopted for the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0133] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as detailed in the third embodiment.

[0134] Embodiment Three

[0135] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in the first embodiment can be realized.

[0136] Since the electronic device described in this embodiment is the device adopted for implementing the method in the first embodiment of this application, based on the method described in the first embodiment of this application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, how this electronic device realizes the method in the embodiments of this application will not be introduced in detail herein. Any device adopted by those skilled in the art for implementing the method in the embodiments of this application falls within the scope of protection of this application.

[0137] Based on the same inventive concept, this application provides a storage medium corresponding to the first embodiment, as detailed in the fourth embodiment.

[0138] Embodiment Four

[0139] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in the first embodiment can be realized.

[0140] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0141] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0144] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should all be covered by the scope protected by the claims of the present invention.

Claims

1. An automated cloud training method for e-commerce image models, characterized by: The steps include: Step 1: The client establishes a connection with the cloud service and transmits the image data to the cloud server; Step 2: The cloud server receives and stores the training data; Step 3: The cloud server trains the e-commerce image model using the training data to obtain a trained e-commerce image model; Step 4: Generate a set number of generated images using the trained e-commerce image model; Step 5. Score each generated image through the scoring model to obtain an image score. When the image score reaches the set threshold score, the generated image is qualified; otherwise, it is unqualified; when the qualified rate of the generated image reaches the set percentage, the trained e-commerce image model is saved; otherwise, the e-commerce image model is retrained until the qualified rate of the generated image reaches the set percentage, and the trained e-commerce image model is saved.

2. According to claim 1, the method for automatic cloud training of e-commerce image models is characterized by: The step 1 is specifically as follows: the client uses paramiko.SSHClient to establish an SSH secure connection with the cloud server, and configures the OSS object storage access credentials through oss2.Auth and oss2.Bucket. According to the preset configuration table, the client establishes a persistent connection with the cloud server; Select Pigz to perform multi-threading to compress the image data and obtain at least one data set compression package file to be uploaded; use multi-threaded concurrent.futures to upload multiple compressed data to the cloud server in parallel; while uploading, record a file mapping dictionary locally on the client, and the file mapping dictionary is used to ensure the accuracy of data synchronization and backtracking; The step 2 is specifically as follows: paramiko.SSHClient().exec_command() is used to manipulate the cloud server on the client to run a Python script for decompressing data packets, wherein the Python script uses subprocess.Popen to execute the data packet decompression command unpigz on the cloud server, and multiprocessing is used to start multiple processes for decompression according to the number of data packets to be decompressed; The cloud server records the monitoring and logs of the data transmission process and stores them as .pkl files in the specified location of the client to obtain training data.

3. The method for automated cloud training of e-commerce image models according to claim 1, characterized in that: The step 3 is specifically as follows: encapsulating the training environment based on Docker container technology, supporting GPU training through Nvidia-Docker, using torch.distributed to implement multi-GPU parallel training, supporting data loading of the e-commerce image model through the DistributedDataParallel class, and then distributing it to different graphics cards for multi-graphics card parallel training; using Tensorboard to record indicator data during the training process, and regularly using torch.save to save checkpoints; inputting the training data into the e-commerce image model for model training, and obtaining the trained e-commerce image model after the training is completed.

4. The method for automated cloud training of e-commerce image models according to claim 1, characterized in that: The step 4 is specifically as follows: loading the trained e-commerce image model on the cloud server through paramiko.SSHClient().exec_command, and generating a set number of generated images through the trained e-commerce image model; and storing the generated images to a set location.

5. The method for automated cloud training of e-commerce image models according to claim 1, characterized in that: The scoring model in step 5 is specifically: Step 61, perform color evaluation, composition evaluation and image aesthetic evaluation on the input image, and use 0 or 1 to mark good color, bad color, good composition, bad composition, good image and bad image to obtain a marked image; Color evaluation is specifically as follows: color space conversion and channel extraction; using cv2.cvtColor to convert the input image from RGB color space to HSV color space, and extracting the three channels of hue, saturation and brightness respectively; Histogram distribution feature calculation: For the three channels of hue, saturation and brightness, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; use np.percentile to obtain the statistical value of the set quantile of each channel; Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of brightness; when S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and H_std is less than the set first threshold, then the good color is recorded as 1, and the bad color is recorded as 0; when S_mean<0.3 or>0.7, or V_mean<0.4 or>0.8, or H_std is greater than or equal to the first set threshold, then the good color is recorded as 0, and the bad color is recorded as 1; The composition evaluation is as follows: cv2.Laplacian is used to calculate the sharpness score of the image; Use cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate edge_strength; Calculate the image centroid position based on cv2.moments and get the symmetry_score; Divide the image into a 3x3 grid and calculate the feature response intensity golden_ratio_score at the golden section point; When sharpness_score>0.7 and edge_strength is greater than the second set threshold, and symmetry_score>0.6 and golden_ratio_score>0.5, good composition is recorded as 1, and bad composition is recorded as 0; when sharpness_score≤0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score≤0.6 or golden_ratio_score≤0.5, good composition is recorded as 0, and bad composition is recorded as 1; The image aesthetic evaluation is as follows: The specific evaluation of image beauty is as follows: use cv2.Laplacian to calculate the image sharpness score model_sharpness_score; Use the QwenVL2 visual model to score the beauty of the image, set the prompt word, and obtain qwen_beauty_score; When model_sharpness_score>0.7 and qwen_beauty_score>0.7, the image is recorded as 1 if it looks good and 0 if it looks bad. In other cases, the image is recorded as 0 if it looks good and 1 if it looks bad. Step 62: Building the evaluation model framework. The evaluation model includes a small aesthetic evaluation model base and a large aesthetic evaluation model base: Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the base of the aesthetic evaluation model; Use torch.load to load the CLIP model as the basis of the aesthetic evaluation model; Delete the fully connected layers in the output part of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain the model output of the same dimension; The final output is concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, which correspond to the probability values ​​of good color, bad color, good composition, bad composition, good image product, and bad image product. Model training: Input the labeled images into the evaluation model in sequence, and fine-tune the evaluation model architecture: Use the torch.optim.Adam optimizer and the MSE loss function to calculate the error between the predicted score and the set value; when the error is less than the set threshold, stop training and obtain the trained evaluation model; Step 63: Input the pictures to be screened into the trained evaluation model to obtain the probability values ​​of the six dimensions; First, use np.normalize to perform normalization: set the weights of the probability of good color and bad color to 0.8, the weights of the probability of good composition and bad composition to 1.0, and the weights of the probability of good image and bad image to 1.2; Use np.average to achieve weighted averaging to get the final evaluation score, evaluation score = (probability value of good color × 0.8 + probability value of good composition × 1.0 + probability value of good-looking image product × 1.2) ÷ 3, and use np.clip to limit the score to the range of 0 to 1, which is the image score.

6. An automated cloud training device for e-commerce image models, characterized by: include: Establish a connection module, the client establishes a connection with the cloud service and transmits the image data to the cloud server; The receiving storage module receives and stores the training data on the cloud server; Training model module: The cloud server trains the e-commerce image model through training data to obtain the trained e-commerce image model; The generated image module generates a set number of generated images through the trained e-commerce image model; The evaluation model module scores each generated image through the scoring model to obtain an image score. When the image score reaches a set threshold score, the generated image is qualified; otherwise, it is unqualified; when the qualified rate of the generated image reaches a set percentage, the trained e-commerce image model is saved; otherwise, the e-commerce image model is retrained until the qualified rate of the generated image reaches a set percentage, and the trained e-commerce image model is saved.

7. The automatic cloud training device for e-commerce image models according to claim 6 is characterized by: The connection establishment module is specifically as follows: the client uses paramiko.SSHClient to establish an SSH secure connection with the cloud server, and configures the OSS object storage access credentials through oss2.Auth and oss2.Bucket. According to the preset configuration table, the client establishes a persistent connection with the cloud server; Select Pigz to perform multi-threading to compress the image data and obtain at least one data set compression package file to be uploaded; use multi-threaded concurrent.futures to upload multiple compressed data to the cloud server in parallel; while uploading, record a file mapping dictionary locally on the client, and the file mapping dictionary is used to ensure the accuracy of data synchronization and backtracking; The receiving storage module specifically comprises: controlling the cloud server to run a Python script for decompressing data packets on the client through paramiko.SSHClient().exec_command(), wherein the Python script uses subprocess.Popen to execute the data packet decompression command unpigz on the cloud server, and uses multiprocessing to start multiple processes for decompression according to the number of data packets to be decompressed; The cloud server records the monitoring and logs of the data transmission process and stores them as .pkl files in the specified location of the client to obtain training data.

8. The automatic cloud training device for e-commerce image models according to claim 6 is characterized by: The training model module specifically includes: encapsulating the training environment based on Docker container technology, supporting GPU training through Nvidia-Docker, using torch.distributed to implement multi-GPU parallel training, supporting data loading of the e-commerce image model through the DistributedDataParallel class, and then distributing it to different graphics cards for multi-graphics card parallel training; using Tensorboard to record indicator data during training, and regularly using torch.save to save checkpoints; inputting the training data into the e-commerce image model for model training, and obtaining the trained e-commerce image model after the training is completed.

9. The automatic cloud training device for e-commerce image model according to claim 6, characterized in that: The training model module specifically includes: loading the trained e-commerce image model on the cloud server through paramiko.SSHClient().exec_command, and generating a set number of generated images through the trained e-commerce image model; and storing the generated images to a set location.

10. The automatic cloud training device for e-commerce image model according to claim 6, characterized in that: The scoring model in the evaluation model module is specifically: A marking unit performs color evaluation, composition evaluation and image aesthetic evaluation on the input image, and marks good color, bad color, good composition, bad composition, good image and bad image with 0 or 1 to obtain a marked image; Color evaluation is specifically as follows: color space conversion and channel extraction; using cv2.cvtColor to convert the input image from RGB color space to HSV color space, and extracting the three channels of hue, saturation and brightness respectively; Histogram distribution feature calculation: For the three channels of hue, saturation and brightness, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; use np.percentile to obtain the statistical value of the set quantile of each channel; Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of brightness; when S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and H_std is less than the set first threshold, then the good color is recorded as 1, and the bad color is recorded as 0; when S_mean<0.3 or>0.7, or V_mean<0.4 or>0.8, or H_std is greater than or equal to the first set threshold, then the good color is recorded as 0, and the bad color is recorded as 1; The composition evaluation is as follows: cv2.Laplacian is used to calculate the sharpness score of the image; Use cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate edge_strength; Calculate the image centroid position based on cv2.moments and get the symmetry_score; Divide the image into a 3x3 grid and calculate the feature response intensity golden_ratio_score at the golden section point; When sharpness_score>0.7 and edge_strength is greater than the second set threshold, and symmetry_score>0.6 and golden_ratio_score>0.5, good composition is recorded as 1, and bad composition is recorded as 0; when sharpness_score≤0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score≤0.6 or golden_ratio_score≤0.5, good composition is recorded as 0, and bad composition is recorded as 1; The image aesthetic evaluation is as follows: The specific evaluation of image beauty is as follows: use cv2.Laplacian to calculate the image sharpness score model_sharpness_score; Use the QwenVL2 visual model to score the beauty of the image, set the prompt word, and obtain qwen_beauty_score; When model_sharpness_score>0.7 and qwen_beauty_score>0.7, the image is recorded as 1 if it looks good and 0 if it looks bad. In other cases, the image is recorded as 0 if it looks good and 1 if it looks bad. Training model unit, evaluation model architecture construction, the evaluation model includes aesthetic evaluation small model base and aesthetic evaluation large model base: Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the base of the aesthetic evaluation model; Use torch.load to load the CLIP model as the basis of the aesthetic evaluation model; Delete the fully connected layers in the output part of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain the model output of the same dimension; The final output is concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, which correspond to the probability values ​​of good color, bad color, good composition, bad composition, good image product, and bad image product. Model training: Input the labeled images into the evaluation model in sequence, and fine-tune the evaluation model architecture: Use the torch.optim.Adam optimizer and the MSE loss function to calculate the error between the predicted score and the set value; when the error is less than the set threshold, stop training and obtain the trained evaluation model; The screening unit inputs the images to be screened into the trained evaluation model to obtain the probability values ​​of six dimensions; First, use np.normalize to perform normalization: set the weights of the probability of good color and bad color to 0.8, the weights of the probability of good composition and bad composition to 1.0, and the weights of the probability of good image and bad image to 1.2; Use np.average to achieve weighted averaging to get the final evaluation score, evaluation score = (probability value of good color × 0.8 + probability value of good composition × 1.0 + probability value of good-looking image product × 1.2) ÷ 3, and use np.clip to limit the score to the range of 0 to 1, which is the image score.