Content generation and graph selection method, system and device, storage medium and program product
By performing visual and perceptual screening on candidate images, combined with image preprocessing and large model technology, the problem of low efficiency and high cost in content creation on social networking service platforms has been solved, achieving efficient and low-cost generation of high-quality content.
Patent Information
- Application Number
- CN202510914197.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, content creation for sharing to social networking service platforms is inefficient and labor-intensive, with cumbersome manual creation processes leading to creative burnout and increased costs.
By acquiring candidate images, the first technique is used for visual quality screening, followed by a second technique for image perception screening, ultimately generating content to be published. The first technique may include image preprocessing, image quality assessment, and text detection, while the second technique may be a large-scale model technique for aesthetic and sentiment perception screening.
It improves content creation efficiency, reduces labor costs, and enhances image and generation quality through multi-dimensional screening, resulting in more attractive content.
Smart Images

Figure CN120994860A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet, and particularly relates to a content generation and image selection method, system, device, storage medium and program product. BACKGROUND
[0002] With the continuous development of Internet technology, a social networking service (SNS) platform gradually becomes a commonly used tool for people. For example, a user can publish created information and the like to the SNS platform, so that friends or other users can know the user's dynamic and achieve the purpose of communication; or a merchant user can share content (such as copywriting, images, videos and the like) created on an e-commerce platform to the SNS platform, to help the merchant user attract consumers on the SNS platform by using the shared content.
[0003] At present, the content shared to the SNS platform is usually created by an artificial way. However, the artificial creation way has the problems of low creation efficiency and high artificial cost. SUMMARY
[0004] Aspects of the present application provide a content generation and image selection method, system, device, storage medium and program product to solve or improve the problems in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a content generation method. The method comprises: obtaining at least one candidate image; screening the at least one candidate image by using a first technology to obtain at least one tentative image meeting a first screening condition; screening the at least one tentative image by using a second technology to obtain at least one target image meeting a second screening condition; and generating to-be-published content based on the at least one target image; wherein the first technology and the second technology screen images from different dimensions.
[0006] In a second aspect, an embodiment of the present application provides a content generation method. The method comprises: obtaining at least one candidate image of a commodity; screening the at least one candidate image by using a first technology to obtain at least one tentative image meeting a first screening condition; screening the at least one tentative image by using a second technology to obtain at least one target image meeting a second screening condition; and generating to-be-published content for the commodity based on the at least one target image; wherein the first technology and the second technology screen images from different dimensions.
[0007] In a third aspect, the embodiments of the present application provide a content generation method. The method comprises: in response to a first operation of a user, sending at least one candidate image input by the user to a server or triggering the server to obtain at least one candidate image related to a target object specified by the user; receiving to-be-posted content returned by the server for the at least one candidate image; and displaying the to-be-posted content; wherein the to-be-posted content is generated based on at least one target image, and the at least one target image is selected from the at least one candidate image by using two different technologies; the two different technologies select the at least one candidate image from different dimensions.
[0008] In a fourth aspect, the embodiments of the present application provide a picture selection method. The method comprises: obtaining at least one candidate image; performing image quality selection of at least one visual dimension on the at least one candidate image to obtain at least one to-be-determined image meeting an image quality requirement; performing image perception aspect selection of at least one dimension on the at least one to-be-determined image to obtain at least one target image meeting an image perception aspect requirement; and outputting the at least one target image.
[0009] In a fifth aspect, the embodiments of the present application provide a service system. The service system comprises a client and a server. The client is configured to implement the steps in the content generation method provided in the third aspect; and the server is configured to implement the steps in the content generation method provided in the first or second aspect.
[0010] In a sixth aspect, the embodiments of the present application provide a service system. The service system comprises a client and a server. The server is configured to implement the steps in the picture selection method provided in the fourth aspect; and the client is configured to obtain at least one target image output by the server.
[0011] In a seventh aspect, the embodiments of the present application provide an electronic device. The electronic device comprises a memory and a processor. The memory is configured to store a program; and the processor is coupled to the memory and is configured to execute the program stored in the memory to implement the steps in the content generation method provided in the first, second or third aspect, or the steps in the picture selection method provided in the fourth aspect.
[0012] In an eighth aspect, the embodiments of the present application provide a computer readable storage medium. The computer readable storage medium stores a computer program or instructions. When the computer program or instructions are executed, the steps in the content generation method provided in the first, second or third aspect, or the steps in the picture selection method provided in the fourth aspect are implemented.
[0013] In an eighth aspect, an embodiment of the present application provides a computer program product, comprising a computer program which, when executed by a computer, causes the computer to perform the steps of the content generation method provided in the first aspect, the second aspect or the third aspect, or the steps of the image selection method provided in the fourth aspect.
[0014] In the embodiment of the present application, at least one candidate image is obtained, the at least one candidate image is screened by using a first technique, and the at least one candidate image is screened again by using a second technique to screen at least one target image, and then the to-be-published content is generated based on the at least one target image. It can be seen that the technical solution provided in the embodiment of the present application can automatically generate the to-be-published content based on the screened target image, improve the creation efficiency of the published content, and is conducive to reducing the labor cost. In addition, the embodiment of the present application adopts a twice screening process, and the two techniques screen the image from different dimensions, which is helpful to improve the screening quality of the image, that is, the screened image is of higher quality, and thus the generation quality of the to-be-published content is improved.
[0015] Another embodiment of the present application provides an image selection method, which performs image quality screening in at least one visual dimension on at least one candidate image, and then performs screening in at least one dimension of image perception; so that the at least one target image screened is better in image quality and image perception, and the target image is better in quality in multiple dimensions. It can be seen that the technical solution provided in the embodiment improves the screening quality of the image. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, and the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:
[0017] Figure 1 A flowchart of a content generation method provided in an embodiment of the present application is shown in FIG. 1;
[0018] Figure 2 An interface in a content generation process provided in an embodiment of the present application is shown in FIG. 2;
[0019] Figure 3 Another interface in a content generation process provided in an embodiment of the present application is shown in FIG. 3;
[0020] Figure 4 A flowchart of another content generation method provided in an embodiment of the present application is shown in FIG. 4;
[0021] Figure 5 A flowchart of still another content generation method provided in an embodiment of the present application is shown in FIG. 5;
[0022] Figure 6A flowchart of a picture selection method provided by an embodiment of the present application is shown in FIG. 1.
[0023] Figure 7 A structural diagram of a service system provided by an embodiment of the present application is shown in FIG. 4.
[0024] Figure 8 A structural diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be described below in conjunction with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0026] It should be noted that, in the case where an embodiment of the present application involves user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiment of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use, and processing of the relevant data need to comply with relevant laws, regulations, and standards of the country and region, and provide corresponding operation portals for the user to choose authorization or rejection. In addition, the various models (including but not limited to language models or large models) involved in the present application comply with relevant legal and standard regulations.
[0027] In addition, it should be noted that, in the case where an embodiment of the present application involves user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiment of the present application include but are not limited to: touch operations, gesture operations, voice operations, head movement operations, eye movement operations, and various modes of interaction operations; wherein the touch operations include but are not limited to: click operations, double-click operations, long-press operations, sliding operations, pinch operations, or mouse hovering operations, etc. The sliding operation includes but is not limited to: straight-line sliding, curve sliding, etc.
[0028] Furthermore, it should be noted that, in the case where an embodiment of the present application involves switching between a first interface and a second interface, the switching mode involved in the embodiment of the present application includes but is not limited to: directly switching from the first interface to the second interface, switching from the first interface to the task interface and completing the corresponding task operation on the task interface, and then switching to the second interface.
[0029] Firstly, the words involved in the embodiments of the present application are explained. It can be understood that the explanation is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation on the embodiments of the present application.
[0030] SNS refers to a service for promoting interaction, sharing information and establishing social relationships between users through an Internet platform. In some scenarios, the SNS platform can become an important tool for people to communicate and disseminate information, and users can publish the created information to the SNS platform so that friends or other users can understand their dynamics and achieve the purpose of communication. In other scenarios, the SNS platform can also become a key channel for brand shaping, promotion and flow, and the merchant user can bind the social media account of the SNS platform on the e-commerce platform, so that the merchant user can share the content (such as copywriting, images, videos, etc.) created on the e-commerce platform to the SNS platform through the bound social media account, helping the merchant user to attract consumers on the SNS platform by using the shared content.
[0031] Artificial intelligence (AI) is a branch of computer science that aims to simulate human intelligent behavior by machines, including learning, reasoning, perception, decision-making, etc. AI technologies include machine learning (ML), deep learning (DL), computer vision (CV), natural language processing (NLP), multi-modal AI (combining visual, language, speech, and other multi-dimensional data), AI large models, etc.
[0032] The preset model, commonly known as the large model, refers to a machine learning model with appropriate parameters and strong computing power, which can process more data and complete various complex tasks such as natural language processing, computer vision, speech recognition, etc. It is an AI model. The preset model can be implemented by a large language model (LLM) or a multi-modal large model (MLM), etc. The preset model can implement generative tasks such as text-to-text, text-to-image, text-to-video, etc. Of course, in addition to the large language model, there are many other types of large models, which are not limited by the present application. The first preset model, the second preset model, the third preset model, the fourth preset model and the fifth preset model mentioned below can be the large model introduced here, but the structure, parameter quantity and computing power of each model can be different.
[0033] Deep learning: refers to a machine learning method that belongs to a subfield of machine learning, which learns the feature representation of data by simulating the hierarchical structure of the human brain neural network (deep neural network), the core feature is to use multiple layers of nonlinear transformation (such as convolutional layer, recurrent layer, attention mechanism, etc.) to automatically extract and learn the complex feature hierarchy in the data, and is widely used in image recognition, speech recognition, natural language processing and other fields. Deep learning models such as CNN (convolutional neural network), RNN (recurrent neural network), Transformer model, GAN (generative adversarial network) and so on. The "large model (preset model)" mentioned in the above is a scaled expansion of deep learning, and the large model is based on the deep learning architecture, but through increasing the parameter quantity, data and computing power to realize qualitative change. Deep learning models are specific (such as CNN for images, RNN for time series data), which need to design features or adjust network structure manually; while large models learn general representations (such as language, multi-modal understanding, etc.) through pre-training.
[0034] Computer vision: is a subfield of artificial intelligence, which is committed to enabling computers to understand, analyze and extract information from digital images or videos (such as object detection, image segmentation, etc.). The "deep learning" mentioned in the above is a basic technology that supports CV and large models. Large models are a scaled expansion of deep learning, and some large models are designed for CV. Computer vision relies on both deep learning models and gradually introduces large model paradigms.
[0035] OpenCV (open source computer vision library): is a cross-platform computer vision and machine learning library that aims to provide efficient and easy-to-use tools to solve various problems in image and video processing.
[0036] PIQE (perception based image quality evaluator): is a perception-based non-reference image quality evaluation index, which uses the block structure and noise features of the image to calculate the quality score of the image.
[0037] OCR (optical character recognition): is a technology that uses optical and computer technologies to convert text information in images into machine-readable text.
[0038] Python Imaging Library (PIL): is an image processing standard library on the Python (computer programming language) platform. In PIL, many image processing related tasks can be performed, including but not limited to: image archiving, image display, and image processing, etc. In the image archiving task, PIL can perform image archiving and batch processing tasks, and can use PIL to create thumbnails, convert image formats, print images, etc. In the image display task, PIL supports various graphical user interface (GUI) framework interfaces and can be used for image display. In the image processing task, PIL includes basic image processing functions that can process pixels, use various convolution kernels for filtering, and also perform color space conversion.
[0039] PIQE (Perceptual Image Quality Evaluator, perceptual image quality evaluator) aims to quantify perceptual quality issues in images by detecting local texture changes and noise levels to obtain a preliminary assessment of perceptual quality. PIQE provides a score on overall perceptual quality, focusing on visual clarity and comfort.
[0040] Wavelet transform technology: can be regarded as a technology that decomposes images into detail and approximation information, which analyzes images through different frequency components. The PyWavelets (an open source Python package for performing wavelet transform) library can implement various wavelet transforms, including discrete wavelet transform (DWT) and continuous wavelet transform (CWT).
[0041] Post, originally referring to a short text written on paper (such as an invitation, a writing sample), later extended to network content, referring to text, image or video content published on the network.
[0042] Currently, the content shared on the SNS platform can be called a "post". A post can include at least one of the following: text, image, and video. Posts are usually created manually. However, this manual creation method faces the problems of low creation efficiency and creative exhaustion, and the manual creation process is complicated and requires the cooperation of multiple tools. The cost of manual creation and material procurement is also rising, which will also lead to the problem of high labor costs.
[0043] To solve the above at least one technical problem, the embodiments of the present application provide the following embodiments. The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the drawings.
[0044] Figure 1 A flowchart of a content generation method provided by an embodiment of the present application is shown. The method can be executed by a server, or by a client, or by a part of the steps being executed by a server and the remaining steps being executed by a client. As shown in the figure, the content generation method can specifically include the following steps: Figure 1
[0045] S101, obtaining at least one candidate image.
[0046] S102, screening the at least one candidate image by using a first technique to obtain at least one pending image meeting a first screening condition;
[0047] S103, screening the at least one pending image by using a second technique to obtain at least one target image meeting a second screening condition;
[0048] S104, generating to-be-published content based on the at least one target image.
[0049] The first technique and the second technique screen images from different dimensions.
[0050] The embodiment of the present application obtains at least one candidate image, screens the at least one candidate image by using a first technique, screens again by using a second technique, and screens out at least one target image. Then, to-be-published content is generated based on the at least one target image. It can be seen that the technical solution provided by the embodiment of the present application can automatically generate to-be-published content based on the screened target image, improves the creation efficiency of the published content, and is conducive to reducing the labor cost. In addition, the embodiment of the present application adopts a twice-screening process, and the two techniques screen images from different dimensions, which helps to improve the screening quality, that is, the screened images are of higher quality, and thus the generation quality of the to-be-published content is improved.
[0051] The at least one candidate image can be an image uploaded by the user through the client, or a photographed photo, or a screenshot obtained picture, or a video frame in a recorded video, etc. Alternatively, the at least one candidate image is autonomously acquired from a picture page or a picture library designated by the user. For example, in an e-commerce scenario, a merchant user sends a content generation request for a product A to be published to the server through the client, and the server acquires an image of the product A on the e-commerce platform based on the identification of the product A as the candidate image after receiving the request. More specifically, the image is acquired from the product detail page of the product A in the merchant's e-commerce store as the candidate image. The product detail page can include pictures, videos, etc.; these pictures can include but are not limited to: pictures of the product A, usage instruction pictures, and third-party detection and authentication pictures, etc. The video in the product detail page can be but is not limited to: product commentary video, product live video, product usage method video, etc. Further, the at least one candidate image can also be acquired from the comment area, such as the picture, video, etc. containing the product A acquired from the comment area of the product A. It should be noted that the information (such as pictures and videos, etc.) in the comment area is allowed by the commenters when acquired.
[0052] In a scenario where the user needs to generate content to be published, the user can operate on the client, so that the user interface of the client can display a content generation page. The user can add at least one candidate image in the content generation page. In an e-commerce scenario, for example, to generate promotion content to be published for a product A, these candidate images are all related to a target object. The target object is the product A. In a social scenario, when the user uploads multiple candidate images, the candidate images can be unrelated, such as pictures taken at different tourist attractions. Then, the "content generation control" is touched, and the client sends the first image related to the target object to the server in response to the user's touch operation on the "content generation control", or the client triggers a local picture screening function to perform the above S102 step.
[0053] The content to be published refers to various forms of SNS content that can be shared and spread to an SNS platform later, including text, images, videos, etc.
[0054] In different application scenarios, the implementation form of the target object is different. For example, in an e-commerce scenario, the target object can be a product, which can be food, clothing, electronic products, books, etc.; in a news scenario, the target object can be news; in a short video scenario, the target object can be a theme object for generating short video content, such as a person, an animal, a car, etc. It can be understood that the implementation form of the target object of the embodiments of the present application is not limited to the above-mentioned forms, which will not be described one by one here.
[0055] In addition to adding at least one candidate image in the content generation page, the user can also input the identification of the product A in the content generation page, so that the server obtains at least one candidate image of the product A based on the identification of the product A. Alternatively, in another application scenario, the user inputs the name of a person in the content generation page, so that the client obtains the portrait of the person from a local album or from the cloud.
[0056] In the above S102, the screening of the at least one candidate image can be screening in terms of image quality. For example, images with poor clarity, poor image composition, high noise, etc. are excluded, and images with good clarity, good image composition, and low noise are screened out. In addition to screening in terms of image quality, it can also include detection of text in the image, such as detecting whether the image contains artificially embedded text (such as watermarks), detecting whether the text in the image meets the set specifications (such as not containing prohibited words, not containing personal information, not containing links and URLs (uniform resource locator), etc.).
[0057] That is, in the present embodiment, the first technology performs image quality screening in at least one visual dimension on the at least one candidate image, wherein the image quality screening in at least one visual dimension includes but is not limited to at least one of the following: screening in the image clarity dimension, screening in the image noise dimension, screening in the image composition dimension, screening in the image brightness dimension, and screening in the image contrast dimension.
[0058] The first technology can not include a large model technology. The first technology can be an independent technology or a fusion technology of multiple technologies. For example, the first technology is a fusion technology of multiple technologies, which can include but is not limited to at least one of the following:
[0059] Image preprocessing technology for converting images to grayscale images, adjusting the size of images, and denoising images, such as OpenCV;
[0060] At least one image quality evaluation technology, such as at least one of PIQE (perceptual image quality evaluator), NIQE (natural image quality evaluator, a no-reference image quality evaluation algorithm), and PyWavelets (wavelet transform technology).
[0061] PIQE is a hybrid solution combining image processing technology and machine learning algorithms. NIQE is a statistical modeling method based on prior knowledge.
[0062] If the first technology is an independent technology, the first technology can be one of the at least one image quality evaluation technology.
[0063] Further, the first technology can further include but is not limited to: an OCR technology, an artificial embedded text (such as a watermark) detection technology, a non-compliant text detection technology, and the like. That is, the first technology in the embodiment also performs text recognition and detection on the at least one candidate image to filter out images without artificial embedded text and with compliant image text.
[0064] The OCR technology is used to extract text from the image. The non-compliant text detection technology applies compliance rules to filter the extracted text to detect whether the text contains, but is not limited to, personal information, connections and URLs, prohibited words, and the like. The artificial embedded text detection technology can include but is not limited to: a frequency domain analysis technology (such as Python+OpenCV tools), a steganalysis tool, and the like, which are not specifically limited in the embodiment.
[0065] The first technology can include an AI technology other than a large model, such as a machine learning algorithm (such as SVM, which is a subset of AI technology) or a deep learning model. For example, the artificial embedded text detection technology uses a deep learning model.
[0066] The filtering of the at least one pending image in S103 can be filtering of the image in terms of aesthetics and emotions. That is, the second technology in the embodiment filters the at least one pending image in at least one dimension of image perception, wherein the filtering in at least one dimension of image perception includes but is not limited to at least one of: filtering in terms of image aesthetic perception and / or filtering in terms of image emotional perception.
[0067] The second technology includes a large model technology. For example, the preset model mentioned in the foregoing can be a multi-modal large model. Accordingly, step S103 in the embodiment can be specifically:
[0068] The at least one pending image is filtered using the first preset model to obtain at least one target image filtered by the model.
[0069] Among the content published on the SNS platform, images are an essential element. At present, some SNS platforms support image content, and even the content published on some SNS platforms mainly focuses on images. Therefore, in order to carefully select an image suitable for publishing on an SNS platform from a plurality of candidate images, the embodiments of the present application filter a better target image from at least one candidate image with the aid of a plurality of technologies.
[0070] In practical application, the large model (i.e., the preset model in this paper) has some shortcomings: first, in image watermarking and text recognition, there are misjudgment conditions in the multi-modal large model mentioned above, so in image-related recognition tasks, the large model cannot be completely relied on; second, in image quality evaluation, the large model also has the problem of unstable results before and after scoring.
[0071] If only the large model is used to screen the target image from the at least one candidate image, the above two problems will occur. Therefore, in the quality scoring of images, the recognition of non-emotional and aesthetic content such as image watermarking and text, the embodiments of the present application use non-AI technology or use computer vision technology and deep learning fusion technology to realize the whole process of image quality detection, watermark detection, and text compliance detection, and to screen at least one to-be-determined image that meets the corresponding requirements in terms of clarity and text content, and then use the large model to further screen the target image that meets the requirements of image perception from the at least one to-be-determined image.
[0072] Among them, the artificially embedded text is different from the natural text. The artificially embedded text refers to the text content artificially added to the image through post-processing means, such as characters, watermarks, subtitles, etc. artificially added to the image; while the natural text refers to the text content originally existing in the natural scene, such as road signs, book texts, etc.
[0073] The illegal text refers to the text used to represent illegal content. Specifically, the illegal text includes but is not limited to false propaganda text, uncivilized language, personal information, links or URLs, and prohibited words, etc.
[0074] Based on the above, in an optional implementation, the step S102 "adopting the first screening of the at least one candidate image" can include:
[0075] S1021, preprocessing the candidate image;
[0076] Among them, the preprocessing includes at least one of the following: converting the candidate image into a grayscale image, adjusting the candidate image to a set size, and denoising the candidate image;
[0077] S1022, quantifying the quality of the candidate image by detecting the texture change and noise level of the candidate image, to obtain a first score;
[0078] S1023, using a wavelet transform library to perform multi-scale wavelet decomposition on the candidate image, decomposing the candidate image into subbands of different scales and frequencies; and evaluating the quality of the candidate image by calculating the energy of the subbands to obtain a quantized second score;
[0079] S1024, obtaining a quality score of the candidate image according to the first score and the second score;
[0080] S1025, in a case where the quality score is greater than or equal to a first threshold, the image quality of the candidate image meets the requirement.
[0081] In the above S1021, the first image can be preprocessed by using OpenCV to obtain a preprocessed image. Specifically, the first image can be converted into a grayscale image by using OpenCV to simplify the processing complexity of the image and focus on the brightness information; then the size of the grayscale image is uniformly adjusted to a standard size to ensure the consistency of the processing scale, for example, the standard size can be 256 pixels*256 pixels; and the grayscale image adjusted to the standard size is subjected to denoising processing by using a Gaussian blur method or the like to reduce the noise of the image and improve the edge definition, thereby obtaining the preprocessed image.
[0082] In the above S1022, the local texture variation and noise level of the preprocessed image can be detected by using PIQE to obtain a preliminary evaluation of the perceived image quality. PIQE focuses on visual clarity and comfort. The value of the PIQE result is adjusted to a range of 0 to 1, thereby obtaining a PIQE score (i.e., the first score). Of course, in addition to using PIOE, other image quality detection tools mentioned above, such as NIQE, can also be used, and the present embodiment does not make specific limitations on this.
[0083] In the above S1023, the preprocessed image can be subjected to wavelet transform processing by using a PyWavelets library to decompose the image into subbands of different scales and frequencies; the energy of the subbands is calculated to evaluate the details and edge definition of the preprocessed image, thereby obtaining a wavelet analysis result; and the value of the wavelet analysis result is adjusted to a range of 0 to 1, thereby obtaining a wavelet analysis score (i.e., the second score).
[0084] One implementation of the above S1024 is to perform weighted summation on the first score and the second score to obtain the quality score of the candidate image. The value of the quality score of the candidate image is in a range of 0 to 1, and the closer the score is to 1, the better the image quality of the candidate image, and the closer the score is to 0, the worse the image quality of the candidate image.
[0085] For example, the weight corresponding to the first score and the weight corresponding to the second score can both be 50%, so as to balance the perceived quality and the detail features, i.e., the quality score of the candidate image = 50%*first score + 50%*second score.
[0086] Therefore, the image is comprehensively evaluated in terms of sharpness, blur detection, noise analysis, and brightness / contrast evaluation, and a quality score in the range of 0 to 1 is obtained. Moreover, images with low resolution, too dark / too bright, or blur can be filtered out by the quality score.
[0087] Specifically, the embodiment of the present application can eliminate candidate images with a quality score less than a first threshold, and retain candidate images with a quality score greater than or equal to the first threshold as pending images, so that the pending images screened have high definition, reasonable composition, and no obvious noise or blur area, i.e., images with a quality score greater than or equal to the first threshold have image quality meeting the preset requirements.
[0088] For example, in actual application, after multiple manual audits according to the quality score of the image, it is considered that images with a quality score of 0.6 or higher have good definition, reasonable composition, less noise, and no obvious blur area, and therefore the first threshold can be 0.6. It should be noted that the first threshold can be pre-set in the server, and the specific value of the first threshold can be set according to actual conditions, and is not limited to the above 0.6. The embodiment of the present application does not limit the specific value of the first threshold.
[0089] In a specific embodiment, the first technique includes a first deep learning model that can be used to detect artificial embedded text in the candidate image. For example, the first deep learning model is trained in advance using data in a first training sample set. The data in the training data set includes sample images and labels indicating whether the sample images contain artificial embedded text. The candidate image is input into the trained first deep model, and the first deep model outputs a probability that the candidate image contains artificial embedded text. For example, the probability value is between 0 and 1, and the closer the probability value is to 1, the more likely the candidate image contains artificial embedded text, and the closer the probability value is to 0, the less likely the candidate image contains artificial embedded text.
[0090] Therefore, the probability value that the candidate image contains artificial embedded text can be used to filter out images containing artificial embedded text. Specifically, the embodiment of the present application can eliminate candidate images with a probability value greater than a second threshold, so that the pending images screened do not contain artificial embedded text, i.e., images with a probability value less than or equal to the second threshold do not contain artificial embedded text. For example, in actual application, images with a probability value less than or equal to 0.001 do not contain artificial embedded text, i.e., the second threshold can be 0.001.
[0091] It should be noted that the second threshold value can be pre-set in the server, and the specific value of the second threshold value can be set according to actual conditions, and is not limited to 0.001. The specific value of the second threshold value is not limited in the embodiments of the present application.
[0092] After the text content is extracted from the candidate image, a preset rule can be used to determine whether the text content contains illegal text, so as to eliminate the first image whose text content contains illegal text, so that the selected pending image does not contain illegal text.
[0093] Specifically, the preset rule can include a virtual propaganda text detection rule, an uncivil language detection rule, a personal information detection rule, a link and URL detection rule, a prohibited word detection rule, and the like. Of course, a text library can be pre-set in the specific implementation, such as a false propaganda word library, an uncivil language word library, a prohibited word library, and the like.
[0094] Based on the above content, in the embodiments of the present application, the first kind of technology can include multiple function modules, such as an image quality screening module, an artificial text embedding detection module, and a text auditing module. Among them, the image quality screening module is realized based on the image preprocessing technology and at least one image quality evaluation technology mentioned above. The artificial text embedding detection module is realized based on the artificial embedded text detection technology mentioned above. The text auditing module is realized based on the OCR technology and the illegal text detection technology. When performing the S102 step of the embodiments of the present application, the detection engine will sequentially execute the above three modules, such as inputting at least one candidate image into the image quality screening module; then inputting the candidate image meeting the image quality requirement output by the image quality screening module into the artificial text embedding detection module; finally inputting the candidate text with the probability of artificial text being less than the second threshold value output by the artificial text embedding detection module into the text auditing module. Of course, the execution order of the above three modules can also be changed, such as the detection engine first calling the artificial text embedding detection module, then calling the text auditing module, and finally calling the image quality screening module. Because the processing logic of the image quality screening module is complex and consumes computing power, the detection engine first calls the artificial text embedding detection module or the text auditing module, which can quickly screen out a part of the candidate images with watermarks or containing illegal text; in this way, the number of candidate images processed by the image quality screening module will be less when the image quality screening module is called.
[0095] In the embodiments of the present application, the to-be-determined images screened from the candidate images have met the basic publishing standards in terms of clarity and text content. However, in the competitive social media environment, it is not enough to meet these standards. In order to make the to-be-published content stand out on the SNS platform, the image also needs to have strong user appeal, which comes from the theme, scene and emotion conveyed by the image, and can win the recognition of users in the aesthetic level and trigger emotional resonance.
[0096] Therefore, in the present embodiment, a second technique is also used to screen at least one to-be-determined image. The second technique can be a large model technique. That is, a first preset model is used to screen the at least one to-be-determined image, and at least one target image screened by the model is obtained. The first preset model can screen the to-be-determined image in terms of aesthetics and image emotional perception. For example, some multi-modal large models play a key role in image aesthetics and emotion detection, and their advantage lies in cross-media understanding ability, which can deeply excavate the connotation of the image, capture the core points and context clues. This all-round data analysis capability breaks the boundaries of traditional image cognition and realizes a leap from surface description to deep understanding. Therefore, combining the capabilities of artificial experience and large models, it can be concluded that in the aesthetic and emotional detection of image content, the large model can comprehensively evaluate the following three aspects:
[0097] The first aspect is visual expressiveness. The visual performance of the image is the first impression presented to the user, which mainly includes color and style, background design, and subject focus.
[0098] The second aspect is emotion and brand communication. Emotion can be conveyed through images and brand consistency can be maintained, which mainly includes emotional resonance, brand consistency, and fashion trends.
[0099] The third aspect is user interaction and practicality. Through images, interaction with users can be promoted and practical value can be provided, which mainly includes relevance and applicability, interaction potential, and format adaptation.
[0100] In order to improve the screening accuracy of the first preset model, a first prompt word for guiding screening can be input when inputting the at least one to-be-determined image to the first preset model. That is, the first preset model screens the at least one to-be-determined image based on the first prompt word, and outputs at least one target image that passes the screening.
[0101] The first prompt word can guide the first preset model to screen images that meet the set aesthetic and emotional requirements. Specifically, the "determining the first prompt word for guiding screening" can include:
[0102] S105, acquire first information, the first information comprising at least one of: to-be-published platform information, a subject of to-be-published content, related information of a target object to which a candidate image belongs, and user input information, the to-be-published platform information comprising at least one of: characteristics of users on a to-be-published platform, platform requirements for to-be-published content, and a style to which the to-be-published platform belongs;
[0103] S106, determine the first prompt word based on the first information.
[0104] The characteristics of users on a to-be-published platform can include but are not limited to: user group characteristics (such as youth, high interactivity, geographical distribution, extensive interests and hobbies, high male proportion, or high female proportion, etc.), and user group common habits (such as a higher proportion of watching short videos, etc.). Some platforms (such as social platforms and short video platforms) have requirements for the content published on the platform, such as requirements for content length (such as maximum number of characters, video file size, etc.), requirements for the field involved in the content, requirements for the format of the content file, etc. In specific implementation, the platform requirements can be acquired from the network side or the service end corresponding to the platform of the to-be-published content. The style to which the to-be-published platform belongs can include but is not limited to: visual and interactive style, content tone, and user interaction mode, etc. Different platforms have their own styles in the above listed several items. For example, some platforms belong to minimalism in visual and interactive style (such as emphasizing function priority and clean interface, etc.); some platforms belong to high saturation and fragmentation style in visual and interactive style (such as emphasizing the use of pictures / videos to attract attention, bright colors, etc.). For example, some platforms belong to entertainment and relaxation style in content tone (such as mostly short, fast, and efficient / talent videos, etc.); some platforms belong to professional and vertical style in content tone (such as emphasizing industry insights or knowledge sharing, and strict content structure, etc.). For example, some platforms belong to strong social chain style in user interaction mode (such as sharing content based on acquaintance relationship, etc.); some platforms belong to weak social chain style in user interaction mode (such as connecting strangers based on topics / interests, etc.).
[0105] In different scenarios, the target object to which the candidate image belongs can be different. For example, in an e-commerce scenario, the target object to which the candidate image belongs can be a commodity or a service (such as a logistics service or a promotion activity). If the target object is a commodity, the related information of the target object to which the candidate image belongs can include but is not limited to: commodity name, commodity detail information, commodity related comment information, commodity belonging brand information (such as brand style, brand story, etc.). If the target object is a promotion activity, the related information of the target object to which the candidate image belongs can include but is not limited to: activity time, activity rules, activity target group, activity theme, etc. In a social scenario, the target object to which the candidate image belongs can be: scenery, person, food, article, etc. If the target object is an article, assuming that the article is a bag shared by a user, the related information of the target object to which the candidate image belongs can include but is not limited to: the article belonging brand information, the article introduction information, the evaluation information for the article, etc.
[0106] Of course, the user can input some information by himself / herself to increase some user's own filtering image ideas. The user input information can be one or more keywords, or a longer text sentence, or a set of instructions. The specific input prompt content depends on the actual needs of the user. In this case, when the user adds the candidate image related to the target object in the content generation page, the user can also input some information in the content generation page, so that the client can send the first information containing the user input information to the server when sending the candidate image related to the target object to the server. Alternatively, the first prompt word can also be one or more keywords pre-set in the server. Of course, in specific implementation, the first prompt word can also be determined locally in the client. That is, the client determines the first prompt word based on the first information.
[0107] The above-mentioned "topic of the to-be-published content" can be input by the user himself / herself, that is, the user input information can contain the "topic of the to-be-published content". Alternatively, the "topic of the to-be-published content" can be determined by the server through recognizing and analyzing at least one candidate image. For example, after obtaining at least one candidate image, the server determines the topic of the to-be-published content by recognizing the content of the at least one candidate image and understanding the image content.
[0108] It needs to be supplemented here that the prompt word (Prompt) refers to the instruction or question provided to the large model, which is used to guide the large model to generate a specific type of answer, content or perform a task. The quality of the prompt word directly affects the output result. In the present embodiment, the first prompt word determined based on the first information can be multiple, or the determined first prompt word contains one or more words, phrases or sentences.
[0109] Based on the above, the first prompt word generated based on the "to-be-published platform information" in the first information is used to guide the first preset model to screen out target images whose visual expressiveness, user interaction and practicality (i.e., the first aspect and the third aspect mentioned above) conform to the style of the to-be-published platform. The first prompt word generated based on the "topic of to-be-published content", "related information of the target object to which the candidate image belongs" and / or "user input information" in the first information is used to guide the first preset model to screen out target images whose emotions and brand communication conform to user needs.
[0110] For example, the embodiment of the present application inputs at least one pending image and the first prompt word into the first preset model, and the output result of the first preset model can include "yes" and "no". If the output result of the first preset model is "yes", it means that the aesthetic and emotional detection of the pending image is passed, and the pending image can be used as a target image. If the output result of the first preset model is "no", it means that the aesthetic and emotional detection of the pending image is not passed. The embodiment of the present application can eliminate the pending image with the output result of "no", so as to determine the pending image with the output result of "yes" as a target image.
[0111] The technical solution provided by the present application can generate to-be-published content, which can be but is not limited to the following several types:
[0112] The first type is that the to-be-published content only contains at least one target image.
[0113] The second type is that the to-be-published content contains at least one target image and adaptive text information.
[0114] The third type is that the to-be-published content contains a carousel video generated based on at least one target image.
[0115] The fourth type is that the to-be-published content contains a carousel video and adaptive text information.
[0116] The fifth type is that the to-be-published content contains video information generated based on at least one target image.
[0117] The sixth type is that the to-be-published content contains video information and adaptive text information.
[0118] In the above-mentioned multiple cases, the adaptive text can be generated by the first preset model based on at least one target image. For example, a multi-modal large model can recognize the image content in at least one target image, understand and mine the connotation of the image, and generate text information.
[0119] In the first case, the at least one target image obtained by the screening can be directly used as the to-be-published content. If the target object is multiple, the multiple target images can be sorted randomly or based on some rules. For example, the multiple target images are multiple images of a product A. The multiple target images include a front image of the product, a side image of the product, a back image of the product, a detail image of the product, a brand image to which the product belongs, and the like. The corresponding rule can be: first the product image, then the product detail image, and then the brand image to which the product belongs. Further, for the multiple product images, the rule can further include: first the front image, then the side image, and then the back image. The image sorting rule is not specifically limited in the embodiment, and can be determined according to actual scene requirements.
[0120] In the second case, the to-be-published content can include the at least one target image and adapted text information. In the to-be-published content, the at least one target image and the adapted text information are associated with each other, and the text information can be displayed in different functional areas. In some embodiments, the text information can be displayed in the target image, for example, the text information is located at the left side of the target image, or at the upper right or lower right of the target image, and the like. In other embodiments, the text information can be displayed outside the target image, for example, the text information is located at a lower position or a right position outside the target image, and the like.
[0121] In the third case, an image processing algorithm can be used to process the at least one target image to generate a carousel video. For example, the image processing algorithm can be, but is not limited to, OpenCV and PIL for image processing in Python. For example, the at least one target object is processed by OpenCV and PIL for image processing in Python to obtain a carousel video.
[0122] The carousel image video includes multiple images displayed in sequence at a specified frequency in a specified area. For example, the carousel video can be a carousel video with a left sliding effect. Taking a carousel video including image one, image three, and image four as an example, after image one is displayed for a preset time length, image one slides to the left and exits, and image three slides to the left and enters, so as to switch to image three for display. After image three is displayed for a preset time length, image three slides to the left and exits, and image four slides to the left and enters, so as to switch to image four for display.
[0123] It can be seen that the technical solution provided by the embodiment of the application can quickly screen out target images that meet the technical standards and are attractive to users, and then use an image processing algorithm to process the target images to generate a carousel video, which can help users quickly obtain rich product information. The carousel video can be adapted to a short video platform for delivery, and the embodiment of the application can quickly and batch produce and deliver videos based on the screened target images.
[0124] Further, the method provided by the embodiment further includes: adding background music to the carousel video, and / or adding an instruction text in a preset frame interval of the carousel video; wherein the instruction text is used to instruct a user to enter the first website.
[0125] The first website can be a website where the to-be-published content is generated, or a website specified by a publisher (such as a social website, a video website, and the like). Taking a target object as a commodity on an e-commerce platform as an example, the first website can be a website corresponding to the e-commerce platform.
[0126] Taking a target object as a commodity on an e-commerce platform as an example, the instruction text can be a link address of the first website, so as to guide a consumer to enter the first website; or the instruction text can also be a link address of a commodity detail page of the commodity, so that the consumer can open the commodity detail page of the commodity based on the link address to perform an order placing, a shopping cart adding, and the like.
[0127] In a fourth case, the generated carousel video and text information adapted to the carousel video can be published together. The carousel video and the text information can be associated and displayed on an interface of a publishing platform. The text information can be displayed below, above, left, or right of the carousel video.
[0128] In a fifth case, the following manner can be adopted, that is, step S104, “generating to-be-published content based on the at least one target image” includes:
[0129] S1041, obtaining target object information, the target image being an image of the target object;
[0130] S1042, obtaining a second prompt word;
[0131] S1043, inputting the at least one target image, the target object information, and the second prompt word into a second preset model, and generating video information by using the second preset model; or S1043’, inputting the at least one target image, the target object information, and the second prompt word into a third preset model, generating a video script by using the third preset model, inputting the video script and the at least one target image into a fourth preset model, and generating video information by using the fourth preset model.
[0132] In the above S1041, taking an e-commerce scenario as an example, if the target object is a commodity, the target object information can include commodity information. The commodity information can include but is not limited to a commodity title and commodity attribute information, wherein the commodity attribute includes a commodity price, a commodity category, a commodity brand, a commodity model, a commodity color, and a commodity material, and the like.
[0133] The target object information can be input by the user in the content generation page, or can be obtained from the network side or a database based on a target object identifier carried in a content publishing request sent by the user.
[0134] In S1043, the second prompt word can be input as text of the second preset model to guide the second preset model to generate the video. The second prompt word can be preset by the system or input by the user. For example, the video is requested to be generated by the publisher to promote the product A, and the publisher can input "please use at least one target image to make a promotional video of product A". In the case of system preset, the technical personnel can set it in advance according to some data or experience. In the case of user input information, the user can input the second prompt word in the content generation page when adding at least one candidate image related to the target object in the content generation page, so that the client can send the second prompt word to the server at the same time when sending the at least one candidate image related to the target object to the server. Alternatively, the second prompt word can also be one or more keywords preset in the server, and the server can automatically obtain the second prompt word when generating the video information by using the second preset model.
[0135] For the scheme of S1043', the difference from S1043 is that two preset models are used, such as a third preset model is used to generate a video script, and a fourth preset model is used to generate video information. The two preset models selected in S1043' have a smaller parameter amount than the second preset model, so the model computing power requirement is not very high, and the cost of the provider of AI service (i.e. the provider of content generation in this embodiment) will not be too high, and it is helpful to improve the professionalism of each model. The third preset model in S1043' is designed and trained for video script generation; and the fourth preset model is designed and trained for video information generation.
[0136] In addition, the second preset model and the first preset model can be the same model or two different models.
[0137] The video script is a kind of information that can indicate the video content and the video structure of a video. Of course, the video script can also indicate other information, which is not limited in this embodiment. The video script can also be understood as a design idea for the to-be-generated video. Through creating the video script, the video content and the video structure contained in the to-be-generated video can be intuitively and clearly obtained, so as to determine the materials (such as target images) required by the to-be-generated video according to the video script, and process these materials according to the video script to generate video information meeting the requirements of the video script, thereby improving the video generation efficiency.
[0138] The generated video information can be a short video with a time length not exceeding a set time length. The set time length is not specifically limited in this embodiment, and can be determined based on the requirements of the corresponding to-be-published platform.
[0139] In the sixth case, the generated video information is published together with the video-adapted text information. In the to-be-published content, the generated video information and the adapted text information are associated with each other, and the text information can be displayed in different functional areas. In some embodiments, the text information can be embedded in the video information for display, such as being embedded in a frame or multiple frames in the video information, or being segmented to display a text block on all frames of the video information or on frames with an equal interval and a fixed number of frames. In other embodiments, the text information can be displayed outside the area where the video information is located.
[0140] In addition to the generation of text information based on at least one target image mentioned above, the embodiments of the present application also provide a scheme for generating to-be-published text based on collected text material information, as a supplement to the text information generated above, to enrich the content of the text and improve the quality of the text. The to-be-published text generated based on the text material information can be published together with the to-be-published content generated in the embodiments. Specifically, the method provided in the embodiments can further include the following steps:
[0141] S107, acquiring text material information;
[0142] S108, processing the text material information by using a fifth preset model to generate to-be-published text;
[0143] The text material information includes at least one of the following: to-be-published platform information, a text theme, related information of a target object to which a candidate image belongs, and user input information. The to-be-published platform information includes at least one of the following: characteristics of a user on a to-be-published platform, platform requirements for to-be-published content, and a style to which the to-be-published platform belongs. It should be noted that the examples of each item in the text material information can be referred to the content above, and will not be repeated here.
[0144] Taking a target object as a commodity on an e-commerce platform as an example, the text theme includes but is not limited to: recommending high-quality products, pushing products according to real-time hot trends, various marketing activities, and promotion activities, etc.
[0145] The fifth preset model can be the same as the first preset model mentioned above, or can be a different model, and the embodiments do not specifically limit this.
[0146] It can be seen that the embodiment of the application can use a large model to create a script, which can not only quickly generate a large number of scripts, but also optimize creative expression through algorithms, providing a new solution for script production that needs to be published on an SNS platform, breaking through the efficiency bottleneck and creative limitations of traditional script creation, and providing a more efficient, more accurate, and more creative solution for generating content to be published.
[0147] It should be noted here that the to-be-published text generated by the fifth preset model can be fused with the "adapted text information generated based on at least one target image" described above. The "adapted text information generated based on at least one target image" focuses on the image content itself; the to-be-published text generated here focuses on the target object, user demand, to-be-published platform, and the like. The to-be-published text can be simply spliced with the "adapted text information generated based on at least one target image" described above; or the fifth preset model can use its natural language processing capability to arrange and adjust the sentences of the two, making them more smooth and reasonable.
[0148] Further, the method provided by the embodiment can further include: sending the generated to-be-published content to the client for the user to review. The user can modify, store, and the like, the text information in the to-be-published content. After confirming the to-be-published content, the user can trigger "publish" through the client for the to-be-published content.
[0149] The technical solutions provided by the embodiments of the application will be described below in conjunction with specific application scenarios.
[0150] Scenario one, taking a target object as a commodity on an e-commerce platform as an example, a simple description of the scheme corresponding to "generating a to-be-published text based on script material information" is given. The script material information includes at least one of the following: to-be-published platform information, script theme, related information of the target object to which the candidate image belongs, and user input information. The specific content of each item of information can be referred to the description above, which will not be repeated here. In addition, it should be noted here that the related information of the commodity to which the candidate image belongs can also include commodity characteristic information.
[0151] For example, the commodity characteristics refer to various internal and external characteristics possessed by the commodity, which together determine the performance of the commodity in the market and the acceptance of the commodity by consumers. Commodity characteristics can be divided into many aspects, such as quality characteristics, functional characteristics, appearance characteristics, and brand characteristics. Quality characteristics can include durability, reliability, safety, and the like of the commodity; functional characteristics refer to what needs of consumers can be met by the commodity and the degree of satisfaction; appearance characteristics refer to the appearance design, color, size, and the like of the commodity, which can have a certain influence on the purchase decision of consumers; a well-known brand can bring consumers a sense of trust and quality assurance, thereby increasing market demand.
[0152] For another example, the copy material information includes user input information, such as promotion keywords (e.g., “limited-time discount”), real-time hotspots, and brand tone, etc. The brand tone is a market impression formed based on the external performance of the brand, and the brand tone includes: brand core value definition interpretation, brand value appeal, brand catchphrase, and brand story, etc. The brand tone can be understood as an impression memory conveyed by the brand to consumers, which can be a brand name, logo, color, slogan, packaging, spokesperson, etc.
[0153] Among the information in the copy material information other than the user input information, part of the information can be input by the user (e.g., the merchant user) in the content generation page, and another part of the information can be autonomously acquired by the server. For example, the to-be-published platform information can be autonomously acquired by the server.
[0154] Referring to Figure 2 , the merchant user can log in to a merchant backstage page provided by an e-commerce platform through the client, and perform a touch operation on the “SNS intelligent sharing” control in the merchant backstage page, so that the user interface of the client can display a content generation page.
[0155] Subsequently, the merchant user can add information related to the product in the content generation page, such as a first image, etc. For example, as shown in (a) of Figure 2 , the content generation page can include an “add image” control 201, and the merchant user can perform a touch operation on the “add image” control 201 in the content generation page to upload a first image related to the product. For example, the uploaded first image related to the product can include image one, image two, image three, image four, and image five, etc.
[0156] In addition, the merchant user can also input copy material information related to the product in the content generation page. For example, as shown in (a) of Figure 2 , the content generation page can further include a copy theme input box 202 and a product characteristic input box 203, etc. The merchant user can input a copy theme through the copy theme input box 202, and input a product characteristic through the product characteristic input box 203.
[0157] For example, the first image uploaded by the merchant user in the content generation page is a product image containing “sunscreen”, the copy theme input by the merchant user in the content generation page is “XXX festival promotion activity”, the input product characteristic is “good sunscreen effect”, and the input marketing strategy is “limited-time discount, 8.8 times”, etc.
[0158] For example, as shown in (a) of Figure 2As shown in (a), the content generation page may also include a content generation control 204. After the merchant user completes the input of product-related information on the content generation page, the merchant user can perform a touch operation on the content generation control 204. The client responds to the merchant user's touch operation on the content generation control 204 and sends candidate images related to the product and product-related copywriting material information to the server.
[0159] The server can filter candidate images (e.g., using two filtering techniques) to obtain the target image. The server also generates target text (i.e., the text to be published) related to the product based on the copywriting material information. This ensures that the content to be published includes both the target image and the target text. The server can then send the content to be published to the client for display on the client's interface.
[0160] In this way, after the server generates the content to be published, including the target image and target text, and sends the content to the client, the client's user interface can then access it from... Figure 2 The content generation page shown in (a) will redirect to the page shown in Figure 1. Figure 2 The content publishing page shown in (b) is shown in the image.
[0161] like Figure 2 As shown in (b), the content publishing page may include content to be published 205, which includes a target image and target text. For example, the target images include Image 1, Image 3, and Image 4, and the target text is "XXX Carnival, summer sun protection essential, reduced to 8.8% off", and the target text is located in the upper right corner of the target image.
[0162] And, as Figure 2 As shown in (b), the content publishing page may also include a text zoom-in control 206, a text zoom-out control 207, a font selection control 208, and a text color selection control 209, etc. Merchants can use these controls to adjust the target text in the content to be published 205.
[0163] The text enlargement control 206 is used to enlarge the text of the target text in the content to be published 205; the text shrinking control 207 is used to shrink the text of the target text in the content to be published 205; the font selection control 208 is used to instruct the merchant user to select the corresponding font to adjust the font of the target text in the content to be published 205 to the selected font, such as adjusting the font of the target text in the content to be published 205 to KaiTi or SongTi, etc.; the text color selection control 209 is used to instruct the merchant user to select the corresponding text color to adjust the text color of the target text in the content to be published 205 to the selected text color, such as adjusting the text color of the target text in the content to be published 205 to red or blue, etc.
[0164] In other embodiments, other controls can also be set on the content publishing page to adjust the target image in the to-be-published content 205, such as changing the contrast of the target image, adding a filter to the target image, and adding a watermark to the target image, etc.
[0165] In this way, by providing the editing controls on the content publishing page, the to-be-published content published to the SNS platform is more in line with the user's needs, and the user's experience is improved.
[0166] In addition, as shown in (b) of Figure 2 The content publishing page can also include a publishing control 210, a draft saving control 211, and a return control 212, etc.
[0167] If the merchant user wants to publish the to-be-published content to the SNS platform, the merchant user can perform a touch operation on the publishing control 210, and the client responds to the touch operation of the merchant user on the publishing control 210 to perform a publishing operation on the to-be-published content. If the merchant user does not want to publish the to-be-published content to the SNS platform temporarily, but wants to save the generated to-be-published content, the merchant user can perform a touch operation on the draft saving control 211, and the client responds to the touch operation of the merchant user on the draft saving control 211 to store the to-be-published content to the draft box. If the merchant user is not satisfied with the generated to-be-published content, the merchant user can perform a touch operation on the return control 212, and the client responds to the touch operation of the merchant user on the return control 212 to return to the content generation page as shown in (a) of Figure 2 In this way, the merchant user can re-upload the first image or input the text material information.
[0168] Scenario two, taking the target object as a commodity on an e-commerce platform as an example, the generated target video (such as the carousel video or video information described above) and the adapted text information are described. Taking the target object as a commodity on an e-commerce platform as an example, in the scenario where the merchant user needs to generate to-be-published content related to the commodity, the merchant user can log in to a merchant back-end page provided by an e-commerce platform through the client, and perform a touch operation on the "SNS intelligent sharing" control in the merchant back-end page, so that the user interface of the client can display a content generation page.
[0169] After that, the merchant user can add information related to the commodity in the content generation page, such as a first image containing the commodity, etc. As shown in (a) of Figure 3As shown in (a), the content generation page may include an "Add Image" control 201. Merchant users can touch the "Add Image" control 201 on the content generation page to upload candidate images related to the product. For example, the uploaded candidate images related to the product may include Image 6, Image 7, Image 8, Image 9, Image 10, and Image 11, etc.
[0170] In addition, merchants can input product-related copywriting information on the content generation page, such as the copywriting theme, product characteristics, marketing strategies, social media platform information to be published, and compliance rules. Figure 3 As shown in (a), the content generation page may also include a copywriting theme input box 202 and a product feature input box 203, etc. Merchants can input copywriting themes through the copywriting theme input box 202 and input product features through the product feature input box 203.
[0171] Optionally, merchants can also enter product information, such as product title and product attributes, on the content generation page.
[0172] like Figure 3 As shown in (a), the content generation page may further include a first generation control 301 and a second generation control 302. The first generation control 301 is used to generate content to be published, including the target image, and the second generation control 302 is used to generate content to be published, including the target video.
[0173] If a merchant wants to generate content to be published, including a target image, they input product-related information on the content generation page. Then, the merchant can touch the first generation control 301. The client responds to this touch by sending candidate images and related text to the server. The server can then generate the content to be published, including the target image and target text (i.e., the adapted text information), based on these candidate images and text. Finally, the server sends the content to the client, allowing the client's user interface to display the content, including the target image and target text.
[0174] If a merchant wants to generate content to be published, including a target video, after entering product-related information on the content generation page, the merchant can touch the second generation control 302. The client responds to the touch operation by sending product-related candidate images and text materials to the server. The server can then generate the content to be published, including the target video and target text, based on the product-related candidate images and text materials. The server can then send the content to be published to the client, allowing the client's user interface to view the content. Figure 3 The content generation page shown in (a) will redirect to the page shown in Figure 1. Figure 3 The content publishing page shown in (b) is shown in the image.
[0175] like Figure 3 As shown in (b), the content publishing page may include content to be published, which includes the target video 303 and the target text, with the target text located below and outside the area where the target video is located. Furthermore, the content publishing page may also include a publishing control 210, a draft control 211, and a return control 212, etc.
[0176] Understandably, in combination Figure 2 and Figure 3 As shown, embodiments of this application can generate content to be published that includes a target image based on an uploaded first image, or they can generate content to be published that includes a target video based on an uploaded first image. One implementation is as follows: Figure 3 As shown in (a), a first generation control 301 and a second generation control 302 can be set on the content generation page. The first generation control 301 is used to generate content to be published, including the target image, and the second generation control 302 is used to generate content to be published, including the target video. Another implementation method is as follows... Figure 2 As shown in (a), a content generation control 204 can be set in the content generation page. The client can respond to the touch operation of the content generation control 204 and display a prompt window. The prompt window includes prompt information, a confirmation control, and a rejection control. The prompt information can be "Generate content to be published including the target image?" When the user touches the confirmation control, the server can generate content to be published including the target image. When the user touches the rejection control, the server can generate content to be published including the target video.
[0177] In one alternative implementation, the method further includes sending the content to be published to a client to display the content on the client interface.
[0178] For example, the content to be published displayed on the client can be like this: Figure 2the to-be-published content indicated by (b) in FIG. 1B, or Figure 3 the to-be-published content indicated by (b) in FIG. 1B.
[0179] In some embodiments, after the content publishing page displays the to-be-published content, the client can also edit the target text, target image, or target video in the to-be-published content in response to the user's editing operation.
[0180] After that, if the user wants to publish the to-be-published content to the SNS platform, the user can perform a touch operation on the publishing control 210. The client responds to the user's touch operation on the publishing control 210 and performs a publishing operation on the to-be-published content to publish the to-be-published content to the second website. The second website can be a website corresponding to the SNS platform.
[0181] In summary, in the embodiments of the present application, the second image with image quality meeting the preset requirements and without artificially embedded text or illegal text can be selected from the first image, and the content processing model is used to select the target image with user attraction from the second image, so that the selected target image not only meets the technical standards but also has user attraction, providing a more comprehensive, efficient, and creative solution for generating the to-be-published content. This not only improves the creation efficiency but also significantly reduces the labor cost.
[0182] In the embodiments of the present application, target images, target texts, and target videos can be generated. In the e-commerce scenario, by generating the to-be-published content that meets the technical standards and has user attraction and publishing the to-be-published content to the second website, the users on the SNS platform can be attracted through the to-be-published content, and high-quality growth of traffic can be achieved.
[0183] It can be understood that Figure 2 and Figure 3 The user interface shown in the above is only an example for better understanding the technical solutions of the embodiments of the present application, and is not the only limitation on the embodiments. In actual applications, the form of the user interface can be set as needed, which is not limited herein.
[0184] Figure 4 Another flowchart of a content generation method provided by the embodiments of the present application is shown in FIG. 4. The method can be executed by a server, or by a client, or a part of the steps is executed by a server and the remaining steps are executed by a client. For example, Figure 4 The content generation method can specifically include the following steps:
[0185] S401, obtaining at least one candidate image of a product;
[0186] S402, screen the at least one candidate image by using a first technique to obtain at least one pending image meeting a first screening condition;
[0187] S403, screen the at least one pending image by using a second technique to obtain at least one target image meeting a second screening condition;
[0188] S404, generate the to-be-published content for the commodity based on the at least one target image;
[0189] The first technique and the second technique screen images from different dimensions.
[0190] The method provided in this embodiment is limited to screening of commodity images and generating to-be-published content of commodities. The specific implementation of each step is the same as or similar to the corresponding step in the above embodiment, and the specific content can be referred to the above description.
[0191] In S401, the at least one candidate image of the commodity can include but is not limited to at least one of the following:
[0192] At least one picture in a commodity detail page of the commodity is obtained as the at least one candidate image.
[0193] A commodity video in the commodity detail page of the commodity is obtained, at least one video frame in the commodity video is extracted as the at least one candidate image.
[0194] At least one picture input by a user for the commodity is obtained as the at least one candidate image.
[0195] Further, the method provided in this embodiment can further include the following steps:
[0196] S405, obtain text information related to the commodity;
[0197] When the to-be-published content for the commodity is generated based on the at least one target image, it includes:
[0198] S401', generate the to-be-published content for the commodity based on the text information and the at least one target image;
[0199] The text information can include but is not limited to at least one of the following: a commodity name, a commodity description text in a commodity detail page, comment information for the commodity, brand information of a brand to which the commodity belongs, and a category to which the commodity belongs.
[0200] Figure 5A flowchart of a content generation method provided by another embodiment of the present application is shown. The execution subject of the method provided by the present embodiment can be a client. Specifically, referring to Figure 5 , the method comprises:
[0201] S501, in response to a first operation of a user, sending at least one candidate image input by the user to a server or triggering the server to obtain at least one candidate image related to a target object specified by the user;
[0202] S502, receiving content to be published returned by the server for the at least one candidate image;
[0203] S503, displaying the content to be published;
[0204] The content to be published is generated based on at least one target image, and the at least one target image is selected from the at least one candidate image by using two different technologies; the two different technologies select the at least one candidate image from different dimensions.
[0205] The content of each step described above and the generation scheme of the content to be published can be referred to the above, and will not be repeated here.
[0206] Further, the present embodiment can further comprise the following steps: in response to a second operation of the user, publishing the content to be published to a second website or editing the content to be published.
[0207] Referring to Figure 6 , based on the technical scheme provided by each embodiment of the present application, another embodiment of the present application provides a picture selection method. The method can be executed by a server, or executed by a client, or a part of the steps is executed by a server and the remaining steps is executed by a client. As shown in Figure 6 , the content generation method can specifically comprise the following steps:
[0208] S601, obtaining at least one candidate image;
[0209] S602, performing image quality screening of at least one visual dimension on the at least one candidate image to obtain at least one to-be-determined image meeting the image quality requirement;
[0210] S603, performing screening of at least one image perception aspect on the at least one to-be-determined image to obtain at least one target image meeting the image perception aspect requirement;
[0211] S604, outputting the at least one target image.
[0212] Further, the S603 "perform screening on the at least one pending image in at least one dimension of image perception aspect to obtain at least one target image meeting the requirement of the image perception aspect" can include:
[0213] obtain second information, the second information including at least one of user input information and information about a target object to which the candidate image belongs;
[0214] determine a third prompt word based on the second information;
[0215] input the at least one pending image and the third prompt word into a first preset model, and perform screening on the at least one pending image in at least one dimension of image perception aspect based on the third prompt word by the first preset model to obtain at least one target image screened by the model;
[0216] wherein the screening in at least one dimension of image perception aspect includes screening in an image aesthetic perception aspect and / or screening in an image emotional perception aspect.
[0217] The specific implementation of each step is the same as or similar to the corresponding step in the above embodiment, and the specific content can be referred to the above, which will not be repeated here.
[0218] Figure 7 A structural schematic diagram of a content generation system provided by an embodiment of the present application is shown in FIG. 7. Figure 7 As shown in the figure, the content generation system can include a client 701 and a server 702.
[0219] The client 701 and the server 702 are connected through a network, and the client 701 can interact with the server 702 through the network to receive or send messages, etc. The network provides a communication link medium between the client 701 and the server 702, and the network can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0220] The client 701 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a vehicle-mounted device, a smart wearable device or an internet of things (IOT) device, etc. For the convenience of understanding, Figure 7 the smart phone, the notebook computer or the desktop computer will be mainly exemplarily described.
[0221] A server-side 702 error can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0222] In some embodiments, the functions of client 701 and server 702 are as follows:
[0223] Client 701 is used to implement the steps in the corresponding method embodiments described above, such as... Figure 5 The steps in the method embodiments shown, or those described above Figure 1 or Figure 4 Some steps in the method embodiments are provided. Server 702 is used to implement the above. Figure 1 or Figure 4 Provide some or all of the steps in the method embodiments.
[0224] or
[0225] Server-side 702 is used to achieve the above. Figure 6 The steps in the provided method embodiment. Client 701 is used to acquire at least one target image output by the server. Or, as described above. Figure 6 In the provided method embodiments, some steps are implemented by the server 702, and other steps are implemented by the client 701.
[0226] The specific implementation of the client, server, and interaction process between them provided in this application embodiment can be found in the corresponding content of the above embodiments, and will not be repeated here.
[0227] One embodiment of this application provides a content generation apparatus. The content generation apparatus may include: an acquisition module, a first filtering module, a second filtering module, and a generation module. The acquisition module is used to acquire at least one candidate image. The first filtering module is used to filter the at least one candidate image using a first technique to obtain at least one undetermined image that meets the first filtering criteria. The second filtering module is used to filter the at least one undetermined image using a second technique to obtain at least one target image that meets the second filtering criteria. The generation module is used to generate content to be published based on the at least one target image. The first and second techniques filter the images from different dimensions.
[0228] Further, at least one of the first technology and the second technology comprises artificial intelligence technology. If the second technology is artificial intelligence technology, the second screening module is configured to screen the at least one pending image by using a first preset model to obtain at least one target image screened by the model.
[0229] Further, the first technology is configured to perform image quality screening of at least one visual dimension on the at least one candidate image, wherein the image quality screening of the at least one visual dimension comprises at least one of the following: screening in an image sharpness dimension, screening in an image noise dimension, screening in an image composition dimension, screening in an image brightness dimension, and screening in an image contrast dimension.
[0230] The second technology is configured to perform image perception aspect screening of at least one dimension on the at least one pending image, wherein the image perception aspect screening of the at least one dimension comprises at least one of the following: image aesthetic perception aspect screening and / or image emotional perception aspect screening.
[0231] Further, the first technology is further configured to perform text recognition and detection on the at least one candidate image to screen out images without artificially embedded text and with compliant image text.
[0232] Further, when the first screening module screens the candidate image by using the first technology, the first screening module is specifically configured to: perform preprocessing on the candidate image, wherein the preprocessing comprises at least one of the following: converting the candidate image into a grayscale image, adjusting the candidate image to a set size, and performing denoising processing on the candidate image; quantifying the quality of the candidate image by detecting texture changes and noise levels of the candidate image to obtain a first score; performing multi-scale wavelet decomposition on the candidate image by using a wavelet transform library to decompose the candidate image into subbands of different scales and frequencies; evaluating the quality of the candidate image by calculating the energy of the subbands to obtain a quantized second score; obtaining a quality score of the candidate image according to the first score and the second score; and in a case where the quality score is greater than or equal to a first threshold value, the image quality of the candidate image meets the requirements.
[0233] Further, when the second screening module screens the at least one pending image by using the first preset model, the second screening module is specifically configured to: determine a first prompt word for guided screening; and the first preset model screens the at least one pending image based on the first prompt word to output at least one target image that passes the screening.
[0234] Further, the second screening module, in determining the first prompt word for guiding the screening, is specifically configured to: acquire first information, and determine the first prompt word based on the first information. The first information includes at least one of the following: to-be-published platform information, a subject of the to-be-published content, related information of a target object to which the candidate image belongs, and user input information. The to-be-published platform information includes at least one of the following: characteristics of users on the to-be-published platform, platform requirements for the to-be-published content, and a style to which the to-be-published platform belongs.
[0235] Further, the generation module, in generating the to-be-published content based on the at least one target image, is specifically configured to: process the at least one target image by using an image processing algorithm to generate a carousel video. The to-be-published content includes the carousel video.
[0236] Further, the generation module is further configured to add background music to the carousel video, and / or add an instruction text in a preset frame interval of the carousel video. The instruction text is used to instruct a user to enter a first website.
[0237] Further, the generation module, in generating the to-be-published content based on the at least one target image, is specifically configured to: acquire target object information, the target image being an image of a target object; acquire a second prompt word; input the at least one target image, the target object information, and the second prompt word into a second preset model, and generate video information by using the second preset model; or input the at least one target image, the target object information, and the second prompt word into a third preset model, generate a video script by using the third preset model, and input the video script and the at least one target image into a fourth preset model, and generate video information by using the fourth preset model. The to-be-published content includes the video information.
[0238] Further, the generation module, in generating the to-be-published content based on the at least one target image, is specifically configured to: perform image content recognition on the at least one target image; and generate text information based on a recognition result. The to-be-published content further includes the text information.
[0239] Further, the generation module is further configured to: acquire copy material information; and process the copy material information by using a fifth preset model to generate a to-be-published text. The copy material information includes at least one of the following: to-be-published platform information, a copy subject, related information of a target object to which a candidate image belongs, and user input information. The to-be-published platform information includes at least one of the following: characteristics of users on the to-be-published platform, platform requirements for the to-be-published content, and a style to which the to-be-published platform belongs.
[0240] Another embodiment of the present application provides a content generation module. The content generation module comprises an acquisition module, a first screening module, a second screening module and a generation module. The acquisition module is configured to acquire at least one candidate image of a commodity. The first screening module is configured to screen the at least one candidate image by using a first technique to obtain at least one pending image meeting a first screening condition. The second screening module is configured to screen the at least one pending image by using a second technique to obtain at least one target image meeting a second screening condition. The generation module is configured to generate, based on the at least one target image, to-be-published content for the commodity. The first technique and the second technique screen images from different dimensions.
[0241] Further, when acquiring the at least one candidate image of the commodity, the acquisition module is specifically configured to:
[0242] acquire at least one picture in a commodity detail page of the commodity as the at least one candidate image; and / or
[0243] acquire a commodity video of the commodity, extract at least one video frame in the commodity video as the at least one candidate image; and / or
[0244] acquire at least one picture input by a user for the commodity as the at least one candidate image.
[0245] Further, the generation module is further configured to acquire text information related to the commodity, generate, based on the text information and the at least one target image, to-be-published content for the commodity, wherein the text information comprises at least one of the following: a commodity name, a commodity description text in a commodity detail page, comment information for the commodity, brand information of a brand to which the commodity belongs, and a commodity category to which the commodity belongs.
[0246] Yet another embodiment of the present application provides a content generation module. The content generation module comprises a sending module, a receiving module and a display module. The sending module is configured to, in response to a first operation of a user, send, to a server, at least one candidate image input by the user or trigger the server to acquire at least one candidate image related to a target object specified by the user. The receiving module is configured to receive to-be-published content returned by the server for the at least one candidate image. The display module is configured to display the to-be-published content. The to-be-published content is generated based on at least one target image, and the at least one target image is screened from the at least one candidate image by using two different techniques. The two different techniques screen the at least one candidate image from different dimensions.
[0247] Further, the sending module is further configured to, in response to a second operation of the user, publish the to-be-published content to a second website or edit the to-be-published content.
[0248] An embodiment of the present application further provides a picture selecting device. The picture selecting device comprises an obtaining module, a first screening module, a second screening module and an output module. The obtaining module is configured to obtain at least one candidate image. The first screening module is configured to perform image quality screening of at least one visual dimension on the at least one candidate image to obtain at least one to-be-determined image meeting an image quality requirement. The second screening module is configured to perform screening of at least one dimension on the at least one to-be-determined image in terms of image perception to obtain at least one target image meeting a requirement in terms of image perception. The output module is configured to output the at least one target image.
[0249] When the second screening module performs screening of at least one dimension on the at least one to-be-determined image in terms of image perception to obtain at least one target image meeting a requirement in terms of image perception, the second screening module is specifically configured to: obtain second information, the second information comprising at least one of the following: user input information, and related information of a target object to which the candidate image belongs; determine a third prompt word based on the second information; input the at least one to-be-determined image and the third prompt word into a first preset model, and perform screening of at least one dimension on the at least one to-be-determined image in terms of image perception by the first preset model based on the third prompt word to obtain at least one target image screened by the model. The screening of at least one dimension in terms of image perception comprises screening in terms of image aesthetic perception and / or screening in terms of image emotional perception.
[0250] It should be noted that each of the device embodiments described above can implement the technical solutions described in the corresponding method embodiments described above, and the implementation principles and technical effects will not be described again. For the specific manner of executing operations of each module or unit in each device provided in the above embodiments, refer to the related content in the corresponding method embodiments described above.
[0251] Figure 8 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. In a possible implementation manner, the structure of the content generating device described above can be implemented as an electronic device, which can be a server, or a client. As shown in FIG. 1, the electronic device can include a memory 1001 and a processor 1002. Figure 8
[0252] The memory 1001 is configured to store computer programs and can be configured to store other various data to support operations on the electronic device. Examples of these data include instructions of any application program or method for operating on the electronic device, data structures, contact data, phonebook data, messages, images, videos, etc.
[0253] The processor 1002 is coupled to the memory 1001 and configured to execute the computer program stored in the memory 1001 to perform the steps of the content generation method or the picture selection method.
[0254] Further, as shown in Figure 8 , the electronic device further includes a communication component 1003, a power supply component 1004, a display 1005, an audio component 1006, and other components. Figure 8 Some components are only schematically shown in the working node, and it does not mean that the electronic device only includes Figure 8 the components shown. In addition, Figure 8 the components in the dashed box are optional components, not mandatory components, and the specific product form of the working node can be determined. The working node of the present embodiment can be implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone or an Internet of Things device, or a server device such as a conventional server, a cloud server or a server array. If the working node of the present embodiment is implemented as a terminal device such as a desktop computer, a notebook computer or a smart phone, it can include Figure 8 the components in the dashed box; if the working node of the present embodiment is implemented as a server device such as a conventional server, a cloud server or a server array, it can not include Figure 8 the components in the dashed box.
[0255] The memory 1001 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0256] Accordingly, the embodiments of the present application also provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. The computer readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of the computer readable storage medium include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium.
[0257] Accordingly, the embodiments of the present application also provide a computer program product, which includes a computer program or instructions, when the computer program or instructions are executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. It should be understood that each of the above-mentioned method processes or a combination of multiple processes can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable data processing devices can be implemented as a device for implementing the corresponding functions in the above-mentioned method embodiments.
[0258] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0259] The above merely provides an example of the present application, and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall fall into the scope of claims of the present application.
Claims
1. A content generation method, characterized in that, The method includes: Obtain at least one candidate image; The first technique is used to filter the at least one candidate image to obtain at least one undetermined image that meets the first filtering condition; The at least one undetermined image is filtered using a second technique to obtain at least one target image that meets the second filtering conditions; Based on the at least one target image, generate content to be published; The first and second techniques filter images from different dimensions.
2. The method according to claim 1, characterized in that, At least one of the first and second technologies includes artificial intelligence technology; If the second technology is artificial intelligence technology, then the second technology is used to filter the at least one undetermined image to obtain at least one target image that meets the second filtering condition, including: The at least one undetermined image is filtered using a first preset model to obtain at least one target image filtered out by the model.
3. The method according to claim 1, characterized in that, The first technique involves performing image quality screening on the at least one candidate image in at least one visual dimension, wherein the image quality screening in the at least one visual dimension includes at least one of the following: screening in the image sharpness dimension, screening in the image noise dimension, screening in the image composition dimension, screening in the image brightness dimension, and screening in the image contrast dimension. The second technique involves filtering the at least one image to be determined in at least one dimension of image perception, wherein the at least one dimension of image perception filtering includes at least one of the following: image aesthetic perception filtering and / or image emotion perception filtering.
4. The method according to claim 3, characterized in that, The first technique also performs text recognition and detection on the at least one candidate image to screen out images without artificially embedded text and whose text is compliant.
5. The method according to claim 3, characterized in that, The first technique is used to filter candidate images, including: The candidate image is preprocessed, and the preprocessing includes at least one of the following: converting the candidate image to grayscale, adjusting the candidate image to a set size, and denoising the candidate image; The quality of the candidate image is quantified by detecting texture changes and noise levels, and a first score is obtained. The candidate image is decomposed into sub-bands of different scales and frequencies using a wavelet transform library; the quality of the candidate image is evaluated by calculating the energy of the sub-bands to obtain a quantized second score. The quality score of the candidate image is obtained based on the first score and the second score; If the quality score is greater than or equal to the first threshold, the image quality of the candidate image meets the requirements.
6. The method according to claim 2 or 3, characterized in that, The at least one undetermined image is filtered using a first preset model, including: Determine the primary prompt word for the guided filtering; The first preset model filters the at least one pending image based on the first prompt word and outputs at least one target image that has passed the filter.
7. The method according to claim 6, characterized in that, Determine the primary prompt words for the guided filtering, including: Obtain first information, which includes at least one of the following: information about the platform to be published, the theme of the content to be published, information about the target object to which the candidate image belongs, and user input information. The information about the platform to be published includes at least one of the following: characteristics of the user on the platform to be published, platform requirements for the content to be published, and the style of the platform to be published. Based on the first information, the first prompt word is determined.
8. The method according to any one of claims 1 to 5, characterized in that, Based on the at least one target image, generate content to be published, including: The at least one target image is processed using an image processing algorithm to generate a carousel video; The content to be published includes the carousel video.
9. The method according to claim 8, characterized in that, The method further includes: Add background music to the carousel video, and / or add instruction text within a preset frame range of the carousel video; wherein the instruction text is used to instruct the user to enter the first website.
10. The method according to any one of claims 1 to 5, characterized in that, Based on the at least one target image, generate content to be published, including: Obtain target object information; the target image is the image of the target object. Get the second clue word; The at least one target image, the target object information, and the second prompt word are input into a second preset model, and video information is generated using the second preset model; or, the at least one target image, the target object information, and the second prompt word are input into a third preset model, and a video script is generated using the third preset model, and the video script and the at least one target image are input into a fourth preset model, and video information is generated using the fourth preset model. The content to be published includes the video information.
11. The method according to claim 10, characterized in that, Based on the at least one target image, generate content to be published, including: Perform image content recognition on the at least one target image; Based on the recognition results, generate text information; The content to be published also includes the text information.
12. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain copywriting material information; The text material information is processed using the fifth preset model to generate the text to be published; The copywriting material information includes at least one of the following: information about the platform to be published, the copywriting theme, relevant information about the target object to which the candidate image belongs, and user input information. The information about the platform to be published includes at least one of the following: the characteristics of the users on the platform to be published, the platform requirements for the content to be published, and the style of the platform to be published.
13. A content generation method, characterized in that, include: Obtain at least one candidate image of the product; The first technique is used to filter the at least one candidate image to obtain at least one undetermined image that meets the first filtering condition; The at least one undetermined image is filtered using a second technique to obtain at least one target image that meets the second filtering conditions; Based on the at least one target image, generate content to be published for the product; The first and second techniques filter images from different dimensions.
14. The method according to claim 13, characterized in that, Obtain at least one candidate image of the product, including: Obtain at least one image from the product details page of the product as the at least one candidate image; and / or Obtain the product video from the product details page of the product, and extract at least one video frame from the product video as the at least one candidate image; and / or At least one image input by the user for the product is obtained as the at least one candidate image.
15. The content generation method according to claim 13 or 14, characterized in that, Also includes: Obtain text information related to the product; When generating content to be published for the product based on the at least one target image, the process includes: Based on the text information and the at least one target image, generate content to be published for the product; The text information includes at least one of the following: product name, product description text on the product details page, comments about the product, brand information of the brand to which the product belongs, and the product category.
16. A content generation method, characterized in that, The method includes: In response to the user's first action, send at least one candidate image input by the user to the server or trigger the server to retrieve at least one candidate image related to the target object specified by the user; Receive the content to be published returned by the server for the at least one candidate image; Display the content to be published; The content to be published is generated based on at least one target image, which is selected from at least one candidate image using two different techniques; the two different techniques filter at least one candidate image from different dimensions.
17. The method according to claim 16, characterized in that, The method further includes: In response to the user's second action, the content to be published is published to a second website or the content to be published is edited.
18. A method for selecting images, characterized in that, include: Obtain at least one candidate image; The at least one candidate image is subjected to image quality screening in at least one visual dimension to obtain at least one undetermined image that meets the image quality requirements; The at least one undetermined image is filtered in at least one dimension of image perception to obtain at least one target image that meets the requirements of image perception. Output the at least one target image.
19. The method according to claim 18, characterized in that, Performing image perception filtering on the at least one undetermined image in at least one dimension to obtain at least one target image that meets the image perception requirements includes: Obtain second information, which includes at least one of the following: user input information and information related to the target object to which the candidate image belongs; Based on the second information, the third prompt word is determined; The at least one undetermined image and the third prompt word are input into a first preset model. The first preset model performs image perception filtering on the at least one undetermined image based on the third prompt word to obtain at least one target image filtered by the model. Among them, at least one dimension of image perception screening includes: screening based on image aesthetic perception and / or screening based on image emotional perception.
20. A service system, characterized in that, include: A client, used to implement the steps in the content generation method as described in claim 16 or 17; A server-side component for implementing the steps in the content generation method as described in any one of claims 1 to 15.
21. A service system, characterized in that, include: The server is used to implement the steps in the image selection method as described in claim 18 or 19; The client is used to obtain at least one target image output by the server.
22. An electronic device, characterized in that, include: Memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program stored in the memory to implement the steps in the content generation method as described in any one of claims 1 to 17, or the steps in the image selection method as described in claim 18 or 19.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed, implement the steps in the content generation method as described in any one of claims 1 to 17, or the steps in the image selection method as described in claim 18 or 19.
24. A computer program product, characterized in that, Includes a computer program that, when run, causes a computer to perform the steps in the content generation method as described in any one of claims 1 to 17, or the steps in the image selection method as described in claim 18 or 19.