Image processing method and device
By calculating the importance index and regional proportion of self-media cover images, the screenshot that best represents the target image is selected as the cover image, which solves the problems of low efficiency and low quality in generating self-media cover images and improves user interactivity.
Patent Information
- Application Number
- CN202111301128.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-04
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-11-04
AI Technical Summary
In existing technologies, the generation efficiency and quality of cover images for self-media are relatively low, and it is impossible to quantitatively evaluate the effects and differences of different cropping results.
By identifying at least two screenshots corresponding to the target image in the target information stream, calculating the importance index and region proportion of the first object in each screenshot, comprehensively scoring the correlation between the screenshots, and selecting the target screenshot that best represents the target image as the cover image.
It improved the efficiency and quality of cover image generation, thereby increasing user click-through rates and engagement time with the target information feed.
Smart Images

Figure CN114332195B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method and device. Background Art
[0002] We-media differs from information dissemination led by professional media organizations. It is an information dissemination activity led by the general public, providing individuals with a means of producing, accumulating, sharing, and disseminating information, while maintaining both privacy and public transparency. The cover image is the face of a we-media platform, attracting attention. The quality of the cover image and the message it conveys significantly influence users' interest in viewing.
[0003] In related technologies, when self-media content is published, some cropping rules are manually defined and the original image is screenshotted to obtain the cover image. During the screenshot process, it is impossible to quantitatively evaluate the effects and differences of different cropping results; thus, the generation efficiency and quality of the cover image are low.
[0004] Therefore, it is necessary to provide an image processing method and device to improve the generation efficiency and quality of cover images. Summary of the Invention
[0005] The present application provides an image processing method and device, which can improve the generation efficiency and quality of cover images.
[0006] In one aspect, the present application provides an image processing method, comprising:
[0007] Determine at least two screenshots corresponding to the target image in the target information flow, wherein the at least two screenshots include respective corresponding first objects;
[0008] Determining an importance index result of the first object in each screenshot based on an area ratio in the target image of a second object corresponding to the first object in each screenshot; the first object is part or all of the second object; and the second object is an object in the target image;
[0009] Determining an area ratio of the first object in each screenshot based on a preset area of the first object in each screenshot and an overlapping area of each screenshot; the preset area being an area occupied by a second object corresponding to the first object in the target image;
[0010] Determining a comprehensive score for each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot; the comprehensive score of each screenshot represents the degree of association between each screenshot and the target image;
[0011] Based on the comprehensive scores of the screenshots corresponding to the target image, a target screenshot corresponding to the target image is determined; the target screenshot is used to determine the cover image of the target information flow.
[0012] Another aspect provides an image processing device, the device comprising:
[0013] A screenshot determining module, configured to determine at least two screenshots corresponding to a target image in a target information stream, wherein the at least two screenshots include respective corresponding first objects;
[0014] an importance index result determination module, configured to determine an importance index result of the first object in each screenshot based on an area ratio of a second object corresponding to the first object in the target image; the first object being part or all of the second object; and the second object being an object in the target image;
[0015] an area ratio result determination module, configured to determine an area ratio result of the first object in each screenshot based on a preset area of the first object in each screenshot and an overlapping area of each screenshot; the preset area being an area occupied by a second object corresponding to the first object in the target image;
[0016] a comprehensive score determination module, configured to determine a comprehensive score for each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot; the comprehensive score of each screenshot represents the degree of association between each screenshot and the target image;
[0017] A target screenshot determination module is used to determine a target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image; the target screenshot is used to determine the cover image of the target information flow.
[0018] On the other hand, an image processing device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the image processing method described above.
[0019] On the other hand, a computer storage medium is provided, wherein the computer storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the image processing method described above.
[0020] Another aspect provides a computer program product or computer program, the computer program product or computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image processing method described above.
[0021] The image processing method and device provided in this application have the following technical effects:
[0022] The present application first determines at least two screenshots corresponding to a target image in a target information stream, and then determines an importance index result of the first object in each screenshot based on the area ratio of the second object corresponding to the first object in each screenshot in the target image; the first object is part or all of the second object; and the second object is an object in the target image; thereby, the importance of the first object in the original image in each screenshot can be determined; based on the overlapping area between a preset area of the first object in each screenshot and each screenshot, the area ratio result of the first object in each screenshot is determined; the preset area is the area occupied by the second object corresponding to the first object in the target image; thereby, the completeness of the first object relative to the corresponding second object in each screenshot can be determined; based on the importance index result and area ratio result of the first object in each screenshot, a comprehensive score of each screenshot corresponding to the target image is determined; based on the comprehensive score of each screenshot corresponding to the target image, a target screenshot corresponding to the target image is determined; thereby, a target screenshot that can well represent the target image can be determined. The target screenshot can be used to determine the cover image of the target information stream, thereby improving the efficiency and quality of cover image generation, and facilitating increasing users' click-through rate and consumption time on the target information stream. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 is a schematic diagram of an image processing system provided in an embodiment of the present application;
[0025] Figure 2 This is a flowchart of an image processing method provided by an embodiment of the present application;
[0026] Figure 3This is a flowchart of a method for determining an importance index result of a first target object in a candidate screenshot provided by an embodiment of the present application;
[0027] Figure 4 This is a flowchart of another method for determining the importance index result of the first target object in the candidate screenshot provided by an embodiment of the present application;
[0028] Figure 5 This is a flowchart of a method for determining an area ratio result of a first object in each screenshot provided by an embodiment of the present application;
[0029] Figure 6 1 is a flow chart of a method for determining a comprehensive score of a candidate screenshot provided in an embodiment of the present application;
[0030] Figure 7 is a flowchart of another method for determining a comprehensive score of a candidate screenshot provided in an embodiment of the present application;
[0031] Figure 8 The screenshots are obtained using the existing preset rule screenshot method, the manual standard screenshot method and the image processing method of the present application respectively;
[0032] Figure 9 This is the target image example 1 provided in the embodiment of the present application;
[0033] Figure 10 is a target screenshot corresponding to the target image example 1 provided in the embodiment of the present application;
[0034] Figure 11 This is the second target image example provided in the embodiment of the present application;
[0035] Figure 12 This is a target screenshot corresponding to the target image example 2 provided in the embodiment of the present application;
[0036] Figure 13 This is the target image example 3 provided in the embodiment of the present application;
[0037] Figure 14 This is a target screenshot corresponding to target image example 3 provided in the embodiment of the present application;
[0038] Figure 15 is a structural diagram of an image processing device provided in an embodiment of the present application;
[0039] Figure 16 is a structural diagram of an image processing system provided in an embodiment of the present application;
[0040] Figure 17 This is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The following explains the terms in the embodiments of this application:
[0042] Feeds: A news source (also translated as information flow, source material, feed, information provider, feed, summary, source, news subscription, web source) is a data format through which websites disseminate the latest information to users, typically arranged in a timeline format. The timeline is the most primitive, intuitive, and basic form of a feed. A prerequisite for users to subscribe to a website is that the website provides a news source. The aggregation of feeds is called aggregation, and the software used for aggregation is called an aggregator.
[0043] MCN (Multi-Channel Network), a product form of multi-channel network, is a new operating model of the internet celebrity economy.
[0044] OCR: (Optical Character Recognition) refers to the process in which an electronic device examines characters printed on paper, determines their shape by detecting dark and light patterns, and then uses character recognition methods to translate the shape into computer text.
[0045] Short video: also known as short video, is a form of Internet content dissemination, generally referring to video content with a duration of less than 5 minutes that is disseminated on new Internet media; with the popularization of mobile terminals and the acceleration of the Internet, short, flat and fast high-traffic content has gradually gained the favor of major platforms, fans and capital.
[0046] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0047] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0048] See also Figure 1 , Figure 1 is a schematic diagram of an image processing system provided in an embodiment of the present application, such as Figure 1 As shown, the image processing system may include at least a server 01 and a client 02 .
[0049] Specifically, in an embodiment of the present application, the server 01 may include an independently operated server, a distributed server, or a server cluster consisting of multiple servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 01 may include a network communication unit, a processor, and a memory, etc. Specifically, the server 01 can be used to determine at least two screenshots corresponding to a target image in a target information flow, and based on the area ratio of the second object corresponding to the first object in each screenshot in the target image, determine the importance index result of the first object in each screenshot; and based on the overlapping area of the preset area of the first object in each screenshot and each screenshot, determine the area ratio result of the first object in each screenshot; and based on the importance index result and area ratio result of the first object in each screenshot, determine the comprehensive score of each screenshot corresponding to the target image; and based on the comprehensive score of each screenshot corresponding to the target image, determine the target screenshot corresponding to the target image.
[0050] Specifically, in the embodiments of the present application, the client 02 may include a physical device such as a smartphone, desktop computer, tablet computer, laptop computer, digital assistant, smart wearable device, smart speaker, in-vehicle terminal, smart TV, etc. It may also include software running on the physical device, such as a web page provided by a service provider to a user, or an application provided by the service provider to a user. Specifically, the client 02 may be used to send a target information stream to the server 01 and receive a target screenshot sent by the server 01.
[0051] The following describes an image processing method of this application. Figure 2 It is a flow chart of an image processing method provided by an embodiment of the present application. This specification provides the method operation steps as described in the embodiment or flow chart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many steps and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, in a parallel processor or multi-threaded processing environment).
[0052] Specific examples Figure 2 As shown, the method may include:
[0053] S201: Determine at least two screenshots corresponding to a target image in a target information flow, wherein the at least two screenshots include respective corresponding first objects.
[0054] In the embodiments of the present application, the target information stream may include professionally generated content or user-generated content such as videos (long or short), live broadcasts, and graphics. There may be one or more target images. When the target information stream is a video, the video may be subjected to frame extraction to obtain the target image. During frame extraction, key frames in the video may be extracted, or the video may be uniformly sampled to obtain multiple target images.
[0055] In the embodiment of the present application, the at least two screenshots corresponding to the target image in the target information stream may include:
[0056] S20101: Determine a target image in a target information flow.
[0057] In the embodiment of the present application, the target image may include a second object, and the second object may be one or more.
[0058] In the embodiment of the present application, determining the target image in the target information stream may include:
[0059] Obtain at least two candidate images in the target information stream;
[0060] Based on the image quality index, the at least two candidate images are filtered to obtain a target image.
[0061] In an embodiment of the present application, image quality indicators may include but are not limited to image clarity, image aesthetics, whether it is pornographic or vulgar, etc., so as to filter out images that do not meet the image quality indicators from the candidate images and obtain target images with good clarity, high aesthetics, and non-pornographic or vulgarity.
[0062] In an embodiment of the present application, the target image can be stored based on a preset database; and the identification information of each target image can be set; for example, the unique mark Rowkey of the content where the image is located can be used as the identification information, and the Rowkey consists of a 16-bit length string in a fixed format, and its content includes the area ID; it can be determined according to the type of the target information stream (for example, pictures, small videos, and short videos correspond to different area IDs respectively), or according to a timestamp, a random number (such as the source field serial number of the self-media account owner, the MCN partner organization, or the crawled content). The serial number ID of the target image in the target information stream can also be used as the identification information. For example, the identification information can be the image serial number encoding obtained by the order in which the target image appears in the picture and text content or by extracting the frame.
[0063] S20103: Determine a target size of the screenshot based on the attribute information of the second object in the target image.
[0064] In an embodiment of the present application, the attribute information may be the coordinates corresponding to the outline of the second object, and the target size may be determined based on the coordinates corresponding to the outline of the second object; that is, the screenshot may be ensured to include one or more second objects in the target image. When the target image includes multiple second objects, the type of the second object may be determined; and the target size may be determined based on the attribute information of the second object of the target type. The type of the second object may include a person, an animal (e.g., a cat head, a dog head), a building, text, and lace. The target type is an important type, for example, it may include a person, an animal, etc.
[0065] In this embodiment of the present application, the target size of the screenshot can be determined using the target frame attributes of the second object (mainly the coordinates of the detection frame and the category to which it belongs) and point attributes (mainly the detection results of human key points, including coordinates, target category, and point category information). Different target images can correspond to different target sizes.
[0066] S20105: Based on the target size, determine at least two screenshots corresponding to the target image.
[0067] In the embodiments of this application, the most important point of screenshot is to ensure the integrity of the image's meaning under the constraint of the target size. The screenshot corresponding to the target image can be determined based on the target size; the screenshot process can be carried out according to the following preset rules:
[0068] (1) The larger the screenshot area, the better; (2) The more targets included, the better; (3) The included targets should be as complete as possible; (4) The integrity of the screenshot of the human face should be preserved as much as possible; (5) The upper and lower text areas should be considered as a whole and either included or not; (6) If there is an animal, the animal's head should be preserved as much as possible; (7) Targets with a large proportion are important targets and need to be protected; (8) The central area of non-human animal targets is more important than the edge area; (9) The upper border has a higher priority than the lower border text; (10) The border area should be either symmetrical or flush with the bottom of the screenshot area and the lower border division. At the same time, the OCR information in the border should be preserved as much as possible, and the loss problem is not big, but the OCR should not be truncated to keep the text box content intact; (11) The spliced image should either be a complete screenshot or focus on the sub-image with more clues. The clues can be objects in the image, for example, the first image in the screenshot and the second image in the target image. The sub-image with more clues can be the sub-image with a large number of objects.
[0069] In an embodiment of the present application, a screenshot may be determined based on the detection result of the second object in the target image. The detection content may include:
[0070] (1) Face detection and human body detection (2) General object detection (3) Text detection (4) Watermark detection (5) Lace detection (6) Mosaic detection (7) Splice image segmentation, etc.
[0071] It can also be expanded based on the clues needed for cropping. In order to perform intelligent cropping, the image element area detection clue attributes used here include: (1) human target coordinates and key point coordinates; (2) non-human target coordinates (such as cars, animals, etc.); (3) lace coordinates; (4) image size; (5) OCR text content and coordinates; (6) long image dividing line; (7) human face and its key point coordinates.
[0072] In an embodiment of the present application, faces in target images can be detected using a face detection model. For images with people, people are usually the key areas of the image that require attention. Specifically, the five key features of the face can be used to mark the face, thereby training a face detection model. At the same time, a self-supervised face encoder can be used to solve the detection of difficult faces. For non-face objects in the target image, including buildings, animals, human bodies, cars, etc., the latest YOLOv5 model can be used to detect and obtain the detection box coordinates. YOLO is a fast and compact open source object detection model. Compared with other networks, it has stronger performance at the same size and better stability. It is the first end-to-end neural network that can predict the category and bounding box of an object.
[0073] The YOLO network mainly consists of three main components.
[0074] 1) Backbone: A convolutional neural network that aggregates and forms image features at different image granularities.
[0075] 2) Neck: A series of network layers that mix and combine image features and pass them to the prediction layer.
[0076] 3) Head: Predict image features, generate bounding boxes and predict categories.
[0077] YOLOV5 passes each batch of training data through the data loader and enhances the training data at the same time. The data loader performs three types of data augmentation: scaling, color space adjustment, and mosaic enhancement.
[0078] S203: Determine an importance index result of the first object in each screenshot based on the area ratio of the second object corresponding to the first object in the target image; the first object is part or all of the second object; and the second object is an object in the target image.
[0079] In the embodiment of the present application, the importance index result characterizes the degree of influence of the first object in the screenshot on the screenshot being determined as the target screenshot; the probability of the screenshot being the target screenshot can be preliminarily determined based on the importance index result of the first object in the screenshot; if the screenshot includes one first object, the importance index result of the first object in the screenshot characterizes the probability of the screenshot being determined as the target screenshot; if the screenshot includes multiple first objects, the average of the importance index results of the multiple first objects in the screenshot can be used as the probability of the screenshot being determined as the target screenshot. Specifically, if the numerical value corresponding to the importance index result is large, it preliminarily indicates that the probability of the screenshot being determined as the target screenshot is large; if the numerical value corresponding to the importance index result is small, it preliminarily indicates that the probability of the screenshot being determined as the target screenshot is small.
[0080] In the embodiment of the present application, the target images are a preset number, the preset number is at least two, and the at least two screenshots include candidate screenshots, such as Figure 3 As shown, the above determination of the importance index result of the first object in each screenshot based on the area ratio of the second object corresponding to the first object in each screenshot in the target image includes:
[0081] S2031: Determine the area ratio of the second target object corresponding to the first target object in the candidate screenshot in the candidate target image; the candidate screenshot is the screenshot corresponding to the candidate target image;
[0082] In the embodiment of the present application, each screenshot may include different types of first objects, and each target image may include different types of second objects, wherein the first target object and the second target object are objects of the same type.
[0083] S2033: Determine an area ratio interval corresponding to the first target object based on the area ratio corresponding to the first target object in the candidate screenshot;
[0084] In an embodiment of the present application, multiple area ratio intervals can be preset, such as: less than 1%, 1-2%, 2-3%, 3-5%, 5-8%, 8-12%, 12-20%, 20-30%, 30-50%, and above 50%. If the second target object corresponding to the first target object in the candidate screenshot has an area ratio of 10% in the candidate target image, then its corresponding area ratio interval is 8-12%.
[0085] S2035: Based on the first area ratio of the second object in each target image, determine the target image whose first area ratio is within the area ratio range as a first screening image;
[0086] S2037: Determine a first number of second objects in the first screening image;
[0087] In the embodiment of the present application, the number of first screening images within the above area ratio interval, that is, the number of first objects in all target images, can be determined based on the area ratio of the second object in each target image.
[0088] S2039: Determine an importance index result of the first target object in the candidate screenshot based on the first number corresponding to the first target object in the candidate screenshot and the preset number.
[0089] In the embodiment of the present application, each of the above target images includes at least two second objects, such as Figure 4As shown, the above-mentioned determination of the importance index result of the first target object in the candidate screenshot based on the first number corresponding to the first target object in the candidate screenshot and the preset number includes:
[0090] S20391: Based on the second area ratio of each second object in each target image, determine the target image whose second area ratio is within the area ratio range as the second screening image;
[0091] S20393: Determine a second number of second objects in the second screening image;
[0092] In the embodiment of the present application, the total number of second objects in each second screening image may be calculated, and the second number is the number of all second objects in each second screening image.
[0093] S20395: Determine the target image including the first target object as a third screening image, and determine a third number of the third screening images;
[0094] In the embodiment of the present application, the third screening image is the number of target images including the first target object.
[0095] S20397: Determine an importance index result of the first target object in the candidate screenshot based on the first quantity, the second quantity, the third quantity, and the preset quantity corresponding to the first target object in the candidate screenshot.
[0096] In the embodiment of the present application, the calculation formula of the importance index result O of the first target object is as follows:
[0097] O=first quantity / second quantity*log(preset quantity / (third quantity+1)).
[0098] In some embodiments, determining the importance index result of the first target object in the candidate screenshot based on the first number, the second number, the third number, and the preset number corresponding to the first target object in the candidate screenshot may include:
[0099] Determining a weight coefficient of the first target object based on the category of the target information flow;
[0100] In some embodiments, a weight coefficient for the first target object can be set based on the category of the target information stream. For example, if the target information stream is entertainment-related, the first target object can be determined as a person, and a higher weight coefficient can be assigned to the person. If the target information stream is sports-related, the first target object can be determined as football, basketball, etc., and a higher weight coefficient can be assigned to balls. The specific value of the weight coefficient can be set based on actual conditions and can generally be set to a value greater than 1.
[0101] Obtaining a first updated quantity based on the first quantity and the weight coefficient;
[0102] In some embodiments, the product of the first quantity and the weight coefficient may be used as the first update quantity.
[0103] Obtaining a second updated quantity based on the second quantity and the first updated quantity;
[0104] In some embodiments, the second quantity may be adjusted according to the first updated quantity, and the calculation formula is as follows:
[0105] Second updated quantity=second quantity−first quantity+first updated quantity.
[0106] Based on the first update quantity, the second update quantity, the third quantity and the preset quantity corresponding to the first target object in the candidate screenshot, an importance index result of the first target object in the candidate screenshot is determined.
[0107] In some embodiments, the calculation formula of the importance index result O of the first target object is as follows:
[0108] O=first update quantity / second update quantity*log(preset quantity / (third quantity+1)).
[0109] In some embodiments, there may be multiple candidate screenshots. The determining of the importance index result of the first target object in the candidate screenshots based on the first number, the second number, the third number, and the preset number corresponding to the first target object in the candidate screenshots may include:
[0110] Obtain multiple candidate screenshots within a preset time period;
[0111] In some embodiments, in a specific application scenario, the start time and end time of the candidate screenshots may be determined to obtain a preset time period.
[0112] Based on the first object in each candidate screenshot, a target candidate screenshot including the first target object is determined.
[0113] In some embodiments, the target candidate screenshots may be determined based on the type of the first object in the candidate screenshots, and candidate screenshots that do not include the first target object may be screened out.
[0114] Based on the first quantity, the second quantity, the third quantity and the preset quantity corresponding to the first target object in the candidate screenshot, an importance index result of the first target object in the candidate screenshot is determined.
[0115] In some embodiments, in specific application scenarios, for example, in a scenario where an athlete is diving, the cover image can usually be set to a diving image; by obtaining the start time and end time of the athlete's diving, a target screenshot including the diving athlete is determined, thereby improving the screening efficiency of the cover image of the target information flow.
[0116] In the embodiment of the present application, the importance index result can be an importance index value (Object Importance, ObjImp). The relative area of the second object in the image can be divided into multiple area ratio intervals (subdivided into: less than 1%, 1-2%, 2-3%, 3-5%, 5-8%, 8-12%, 12-20%, 20-30%, 30-50%, and more than 50%, a total of 10 levels). The importance ObjImp of targets of different sizes and types (similar to the TFIDF concept, counting the number of occurrences and the number of selections) is calculated as follows:
[0117] ObjImp = (Src: number of times a specific target with a given area percentage is selected / sum of the number of times all selected targets with a given area percentage are selected) * log (total number of images / (number of images containing a specific target + 1));
[0118] Among them, Src (source) is the target image corresponding to the screenshot, and the target is the object. The "number of specific selected targets with a given area ratio" is exemplified as follows: first calculate the area ratio of the second object corresponding to the first target object in the screenshot in the target image, for example, 21%; then determine the given area ratio interval as 20-30%, and calculate the number of occurrences of the first target object in this interval; and the number of first objects in all target images whose corresponding area ratio is in this interval; the total number of images is the total number of target images corresponding to a target information flow, and the number of images containing specific targets is the number of target images including the first object.
[0119] In the embodiment of the present application, the importance of each first object can be determined by the importance index result, thereby further determining the comprehensive score of the screenshot corresponding to the first object.
[0120] S205: Determine the area ratio of the first object in each screenshot based on the preset area of the first object in each screenshot and the overlapping area of each screenshot; the preset area is the area occupied by the second object corresponding to the first object in the target image.
[0121] In the embodiments of this application, Figure 5 As shown, the above-mentioned determination of the area ratio of the first object in each screenshot based on the preset area of the first object in each screenshot and the overlapping area of each screenshot includes:
[0122] S2051: Determine a preset area of the first object in each of the screenshots and a screenshot area corresponding to each of the screenshots;
[0123] S2053: Determine an overlapping area and a merged area between the preset area and the screenshot area;
[0124] In the embodiment of the present application, the overlapping area is the intersection of the preset area and the corresponding screenshot area, and the merged area is the union of the preset area and the corresponding screenshot area.
[0125] S2055: Determine a first area corresponding to the overlapping area and a second area corresponding to the merged area;
[0126] S2057: Utilize the ratio of the first area to the second area as the area ratio of the first object in each of the screenshots.
[0127] In an embodiment of the present application, the area ratio result iou (INTERSECTION OF UNION) obtained based on the screenshot area and the preset area of the first object can be subdivided into multiple intervals, for example, including: less than 10%, 10-20%, 20-30%, 30-40%, 40-50%, and more than 50%.
[0128] In the embodiment of the present application, the first object may be of various types, for example, a person, text, an animal head, lace, etc.;
[0129] Among them, the area ratio results corresponding to the person can be divided into 10 levels of person key point loss information statistical intervals: less than 1%, 1-2%, 2-3%, 3-5%, 5-8%, 8-12%, 12-20%, 20-30%, 30-50%, and 50%. According to the human body detection results, if there are no human key points, the area ratio result is calculated according to the area of the human target area. If it is more precise, it is to see how many human key points exist. If all exist, it is a complete human body. If there are only some key points, it may be that the person itself is incomplete or blocked by other subjects. Considering the amount of calculation here, it can be processed directly according to the area of the human target area, and the human key points can be ignored.
[0130] The area percentage corresponding to the text is a statistical calculation of the degree of text damage in different image positions (top, middle, and bottom) (subdivided into six levels: less than 10%, 10-20%, 20-30%, 30-40%, 40-50%, and more than 50%). The area percentage corresponding to the lace can be obtained based on the positional relationship between the lace and the edge of the screenshot (the conflict between the four edges of the top, bottom, left, and right, similar for the left and right).
[0131] S207: Determine a comprehensive score of each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot; the comprehensive score of each screenshot represents the degree of association between each screenshot and the target image.
[0132] In the embodiments of this application, Figure 6 As shown, the candidate screenshots include at least two first target objects. The comprehensive scores of the screenshots corresponding to the target image are determined based on the importance index results and area ratio results of the first objects in each screenshot, including:
[0133] S2071: Determine the importance index result and area proportion result of each first target object in the candidate screenshot;
[0134] S2073: Determine the comprehensive score of each first target object in the candidate screenshot by multiplying the importance index result and the area proportion result of each first target object in the candidate screenshot;
[0135] S2075: Calculating an average of the comprehensive scores of the first target objects in the candidate screenshots based on the comprehensive scores of the first target objects in the candidate screenshots;
[0136] S2077: Using the average of the comprehensive scores as the comprehensive score of the candidate screenshots.
[0137] In this embodiment of the present application, the comprehensive score (score_obj) of the first target object in the candidate screenshot is calculated as follows:
[0138] score_obj={(z1*c)+(z2*p)+(z3*t)+(z4*h)+(z5*e)} / sum z
[0139] Among them, z1, z2, ..., z5 are ObjImp of different types of first objects, sum z is the sum of the number of z1, z2, ..., z5; c, p, t, h, e are the area proportion results corresponding to each target z respectively.
[0140] In an embodiment of the present application, a comprehensive score for each first target object in the candidate screenshot can be determined based on the product of the importance index result and the area proportion result of each first target object in the candidate screenshot, thereby facilitating the rapid and accurate determination of the target screenshot corresponding to each target image, and the determined target screenshot has the greatest degree of correlation with the corresponding target image and the best image quality.
[0141] In the embodiments of this application, Figure 7 As shown, the above-mentioned comprehensive score average is used as the comprehensive score of the above-mentioned candidate screenshots, including:
[0142] S20791: Obtain the size of the display area and the size of the candidate screenshots;
[0143] In the embodiment of the present application, the display area may be an area corresponding to the cover image, for example, including but not limited to a display interface corresponding to the target information flow, or an interface customized for different applications. Different display areas have different sizes.
[0144] S20793: Determine a normalization coefficient based on the size of the display area and the size of the candidate screenshot;
[0145] S20795: Based on the normalization coefficient, the comprehensive score mean is normalized to obtain the comprehensive score of the candidate screenshot.
[0146] In an embodiment of the present application, the area ratio of the screenshot frame to the maximum proportional screenshot frame can be calculated as a normalization coefficient. The screenshot calculated above is based on the edge of each target detection, and the actual frame obtained will be smaller than the final screenshot frame. This is equivalent to expanding to the nearest output actual screenshot frame size. For example, there is only one person in the target frame. According to the screenshot importance formula, the importance of capturing this complete person is definitely very high and there is no conflict. However, the size of this person does not necessarily correspond exactly to a specific size of the final screenshot frame, and the person that meets the final specific screenshot frame itself can be large or small. This is equivalent to normalizing the size of each candidate screenshot to the size of the display area.
[0147] In the embodiment of the present application, a cover image with good quality can be determined quickly and accurately according to different application scenarios and different display interfaces.
[0148] S209: Determine a target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image; the target screenshot is used to determine the cover image of the target information flow.
[0149] In the embodiment of the present application, the target screenshot corresponding to the target image is determined based on the comprehensive scores of the screenshots corresponding to the target image, including:
[0150] Sorting the screenshots based on the comprehensive scores of the screenshots corresponding to the target image to obtain a first sorting result;
[0151] Based on the first sorting result, a target screenshot corresponding to the target image is determined.
[0152] In the embodiment of the present application, the screenshots may be sorted from high to low according to the comprehensive scores, and then the screenshot at the top of the sort may be determined as the target screenshot of the target image.
[0153] In the embodiment of the present application, the comprehensive performance of different screenshots can be quantified to determine the target screenshot with the greatest correlation with the target image.
[0154] In the embodiment of the present application, the target screenshot corresponding to the target image is determined based on the comprehensive scores of the screenshots corresponding to the target image, including:
[0155] Determining a target screenshot corresponding to each target image based on the comprehensive scores of the screenshots corresponding to each target image;
[0156] After determining the target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image, the method further includes:
[0157] Based on the target screenshots corresponding to the target images, the cover image of the target information flow is determined.
[0158] In the embodiment of the present application, determining the cover image of the target information stream based on the target screenshots corresponding to the target images includes:
[0159] Obtaining subject information of the target information flow;
[0160] Based on the relevance of each target screenshot to the above-mentioned subject information, sorting the target screenshots to obtain a second sorting result;
[0161] Based on the second sorting result, the cover image of the target information flow is determined.
[0162] In an embodiment of the present application, the target screenshot with the highest relevance can be determined based on its relevance to the target information flow's subject information, and used as the cover image. This improves the quality of the cover image and the strength of the subject's expression, increasing the user's click-through rate and consumption time for the target information flow. It also improves the level of automated and intelligent processing of information flow images in various scenarios, resolving the problem of the lack of flexibility of rigid cropping rules, effectively lowering the threshold for self-media authors to generate cover images for display in various specifications and content scenarios, and improving the efficiency and quality of cover image generation.
[0163] In some embodiments, determining the cover image of the target information stream includes:
[0164] Obtaining cover template information, where the cover template information is used to indicate a display style of an object in a cover image;
[0165] In some embodiments, the display style of an object may include the typesetting style of an image or text, and may also include additional text for the object or set the object to have a dynamic display effect.
[0166] Generate a cover image based on the above cover template information and the target screenshot.
[0167] In some embodiments, the target screenshot with the highest correlation with the target information flow theme information among multiple target screenshots can be determined as the target cover screenshot, and each object in the target cover screenshot can be displayed according to the cover template information; thereby enriching the display effect and style of articles or images in the information flow, thereby optimizing the user's reading or viewing experience.
[0168] In a specific embodiment, Figure 8 As shown, the images obtained by using the existing preset rule screenshot method, the manual standard screenshot method and the image processing method of the present application are shown in FIG. Figure 8 In the figure, target screenshot a is a screenshot obtained using preset rules; target screenshot b is a screenshot obtained by the method of the present application; target screenshot c is a manually annotated screenshot; in comparison, it is obvious that the target screenshot obtained by the method of the present application has the best effect, which not only includes all objects in the target image, but also the area of the objects in the target screenshot accounts for the largest proportion. It can be seen that target screenshot b is the best choice for the cover image.
[0169] In a specific embodiment, Figure 9 As shown, Figure 9 The target image example 1 provided in the embodiment of the present application is obtained by using the image processing method of the present application. Figure 10 The target screenshot shown, Figure 9 The image includes two objects, buildings and vehicles. The image corresponds to a scene of a vehicle video, and the focus object is the vehicle. Therefore, the target screenshot only contains one object, the vehicle.
[0170] In a specific embodiment, Figure 11 As shown, Figure 11 The target image example 2 provided in the embodiment of the present application is obtained by using the image processing method of the present application. Figure 12 The target screenshot shown, Figure 11 The object is an animal, and the corresponding target screenshot needs to be displayed on a display device of a target size. However, not all parts of the animal's body can be displayed on the display device. Therefore, a target screenshot of the animal's head image that matches the size of the display device is obtained.
[0171] In a specific embodiment, Figure 13 As shown, Figure 13 The target image example 3 provided in the embodiment of the present application is obtained by using the image processing method of the present application. Figure 14 The target screenshot shown, Figure 13 The objects are two characters c and d, where character c only shows the upper body and character d is a full body image; in order to adapt to the size of the display device corresponding to the target screenshot, the target screenshot Figure 14 Only a part of the original image is captured, including only the upper body image of person d and person c in the original image.
[0172] Specifically, in an embodiment of the present application, a histogram statistics can be used. A histogram, also known as a quality distribution diagram, is a statistical report diagram that represents the distribution of data by a series of longitudinal stripes or line segments of varying heights. The horizontal axis is generally used to represent the data type, and the vertical axis represents the distribution. In order to construct a histogram, the first step is to segment the range of values, that is, to divide the entire range of values into a series of intervals, and then calculate how many values are in each interval. These values are usually specified as continuous, non-overlapping variable intervals. The intervals must be adjacent and are usually (but not necessarily) of equal size. The histogram interval here is the binning rule mentioned above, and statistics is to determine the gear position of the clue object in the image (the following are some examples, and more objects can be used in the same way):
[0173] (1) obj_statistics represents the histogram statistics of all objects with different area ratios in the image. There are 70 types of objects, including 65 general objects and 5 special objects such as faces, cat heads, dog heads, text, and lace. Each type of object has a unit value on the histogram, and the subsequent statistics are performed in a similar way.
[0174] (2) obj_sel_statistics represents the histogram statistics of objects (iou>0) contained in the screenshot frame with different area ratios in the image. There are 70 types of objects, including 65 general objects and 5 special objects such as human faces, cat heads, dog heads, text, and lace.
[0175] (3) obj_block_conflicts represents the statistics of different conflict areas when different object frames (text and borders are considered separately) conflict with the screenshot frame. The conflict area here refers to the overlapping area between the two.
[0176] (4) text_conflicts represents the histogram statistics of different iou (intersection of union) intervals when the text (position in the image) conflicts with the screenshot frame.
[0177] (5) edge_conflicts represents the histogram statistics of different iou intervals when the lace (four different edges) conflicts with the screenshot frame.
[0178] (6) person_node_conflicts represents the histogram statistics of key points of objects with different area ratios when a person or face conflicts with the screenshot frame.
[0179] (7) person_node_discards represents the histogram statistics of key points of objects with different area ratios that are lost when a person or face conflicts with the screenshot frame.
[0180] The most appropriate target screenshot output frame is obtained by scoring the results of the above strategies, thereby realizing an intelligent screenshot solution based on the comprehensive scoring of image area clues.
[0181] It can be seen from the technical solution provided by the above embodiments of the present application that the embodiments of the present application first determine at least two screenshots corresponding to the target image in the target information flow, and then determine the importance index result of the first object in each screenshot based on the area ratio of the second object corresponding to the first object in each screenshot in the target image; the first object is part or all of the second object; the second object is the object in the target image; thereby, the importance of the first object in each screenshot in the original image can be determined; based on the preset area of the first object in each screenshot and the overlapping area of each screenshot, the area ratio result of the first object in each screenshot is determined; the preset area is the The area occupied by the second object corresponding to the first object in the target image; thereby, the degree of completeness of the first object relative to the corresponding second object in each screenshot can be determined; based on the importance index result and the area proportion result of the first object in each screenshot, the comprehensive score of each screenshot corresponding to the target image is determined; based on the comprehensive score of each screenshot corresponding to the target image, the target screenshot corresponding to the target image is determined; thereby, a target screenshot that can well represent the target image can be determined, and the target screenshot can be used to determine the cover image of the target information flow, thereby improving the generation efficiency and quality of the cover image, and facilitating improving the user's click rate and consumption time on the target information flow.
[0182] The present application also provides an image processing device, such as Figure 15 As shown, the device includes:
[0183] A screenshot determining module 1510 is configured to determine at least two screenshots corresponding to a target image in a target information stream, wherein the at least two screenshots include respective corresponding first objects;
[0184] Importance index result determination module 1520, configured to determine an importance index result of the first object in each screenshot based on an area ratio of a second object corresponding to the first object in the target image; the first object being part or all of the second object; and the second object being an object in the target image;
[0185] An area ratio determination module 1530 is configured to determine an area ratio of the first object in each screenshot based on an overlapping area between a preset area of the first object in each screenshot and each screenshot; the preset area being an area occupied by a second object corresponding to the first object in the target image;
[0186] Comprehensive score determination module 1540, configured to determine a comprehensive score for each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot; the comprehensive score of each screenshot represents the degree of association between each screenshot and the target image;
[0187] The target screenshot determination module 1550 is used to determine the target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image; the target screenshot is used to determine the cover image of the target information flow.
[0188] In some embodiments, the target images are a preset number, the preset number is at least two, the at least two screenshots include candidate screenshots, and the importance index result determination module may include:
[0189] an area ratio determining unit, configured to determine an area ratio of a second target object corresponding to the first target object in the candidate screenshot in the candidate target image; the candidate screenshot is a screenshot corresponding to the candidate target image;
[0190] an area ratio interval determining unit, configured to determine an area ratio interval corresponding to the first target object based on the area ratio corresponding to the first target object in the candidate screenshot;
[0191] a first screening image determining unit, configured to determine, based on a first area ratio of the second object in each target image, a target image with a first area ratio within the area ratio interval as a first screening image;
[0192] a first quantity determining unit, configured to determine a first quantity of second objects in the first screening image;
[0193] The importance index result determining unit is configured to determine the importance index result of the first target object in the candidate screenshot based on the first quantity corresponding to the first target object in the candidate screenshot and the preset quantity.
[0194] In some embodiments, each target image includes at least two second objects, and the importance index result determining unit may include:
[0195] a second screening image determining subunit, configured to determine, based on the second area ratios of the second objects in each target image, target images whose second area ratios are within the area ratio interval as second screening images;
[0196] a second quantity determining subunit, configured to determine a second quantity of second objects in the second screening image;
[0197] a third number determining subunit, configured to determine the target image including the first target object as a third screening image, and determine a third number of the third screening images;
[0198] The importance index result determining subunit is configured to determine the importance index result of the first target object in the candidate screenshot based on the first quantity, the second quantity, the third quantity, and the preset quantity corresponding to the first target object in the candidate screenshot.
[0199] In some embodiments, the target screenshot determination module may include:
[0200] The target screenshot determining unit is configured to determine a target screenshot corresponding to each target image based on the comprehensive scores of the screenshots corresponding to each target image.
[0201] In some embodiments, the apparatus may further include:
[0202] The cover image determination module is used to determine the cover image of the target information flow based on the target screenshots corresponding to each target image.
[0203] In some embodiments, the area proportion result determination module may include:
[0204] a screenshot area determining unit, configured to determine a preset area of the first object in each screenshot and a screenshot area corresponding to each screenshot;
[0205] a merged area determining unit, configured to determine an overlapping area between the preset area and the screenshot area and a merged area;
[0206] a second area determining unit, configured to determine a first area corresponding to the overlapping area and a second area corresponding to the merged area;
[0207] The area ratio result determining unit is configured to use the ratio of the first area to the second area as the area ratio result of the first object in each screenshot.
[0208] In some embodiments, the candidate screenshots include at least two first target objects, and the comprehensive score determination module may include:
[0209] A first target object information determination unit, configured to determine an importance index result and an area proportion result of each first target object in the candidate screenshot;
[0210] a comprehensive score determination unit, configured to determine the comprehensive score of each first target object in the candidate screenshot by multiplying the importance index result and the area proportion result of each first target object in the candidate screenshot;
[0211] a comprehensive score average calculation unit, configured to calculate an average of the comprehensive scores of the first target objects in the candidate screenshots based on the comprehensive score of each first target object in the candidate screenshots;
[0212] A comprehensive score determination unit is configured to use the comprehensive score mean as the comprehensive score of the candidate screenshots.
[0213] In some embodiments, the comprehensive score determination unit may include:
[0214] A candidate screenshot size acquisition subunit is used to acquire the size of the display area and the size of the candidate screenshot;
[0215] a normalization coefficient determination subunit, configured to determine a normalization coefficient based on the size of the display area and the size of the candidate screenshot;
[0216] The comprehensive score determination subunit of the candidate screenshots is used to normalize the comprehensive score mean based on the normalization coefficient to obtain the comprehensive score of the candidate screenshots.
[0217] In some embodiments, the target screenshot determination module may include:
[0218] a screenshot sorting unit, configured to sort the screenshots corresponding to the target image based on comprehensive scores thereof to obtain a first sorting result;
[0219] A target screenshot determining unit is configured to determine a target screenshot corresponding to the target image based on the first sorting result.
[0220] In some embodiments, the cover image determination module may include:
[0221] A topic information acquisition unit, configured to acquire topic information of the target information flow;
[0222] a sorting unit, configured to sort the target screenshots based on the relevance between each target screenshot and the subject information to obtain a second sorting result;
[0223] A cover image determination unit is used to determine the cover image of the target information flow based on the above second sorting result.
[0224] The device and method embodiments in the device embodiments are based on the same inventive concept.
[0225] The present application also provides an image processing system. Figure 16 The system includes: a server 1600, a content production end 1620 and a content consumption end 1630. The server 1600 includes a content distribution export server 1601, a content database 1602, a content deduplication service 1603, a regional clue scoring and intelligent screenshot service 1604, a user feedback or reporting interface service 1605, a manual review service 1606, a scheduling center service 1607, an upstream and downstream content interface server 1608, an image region clue feature library 1609, an image region clue extraction service 1610, an image region clue model 1611, an image sample library 1612, a download file system 1613 and a content storage service 1614.
[0226] like Figure 16 As shown, the server 1600 is electrically connected to the content production end 1620 based on the uplink and downlink content interface service 1608, and the server 1600 is electrically connected to the content consumption end 1630 based on the user feedback or reporting interface service 1605, the content distribution export service 1601 or the content storage service 1614.
[0227] In the server 1600, the content distribution export service 1601 is electrically connected to the content database 1602, the regional clue scoring and intelligent screenshot service 1604 is electrically connected to the user feedback or reporting interface service 1605, the manual review service 1606 is electrically connected to the dispatch center service 1607, the user feedback or reporting interface service 1605, and the content database 1602 respectively, the uplink and downlink content interface server 1608 is electrically connected to the content database 1602 and the dispatch center service 1607 respectively; the dispatch center service 1607 is electrically connected to the content deduplication service 1606. 603 is electrically connected; the image area clue feature library 1609 is electrically connected to the area clue scoring and intelligent screenshot service 1604 and the image area clue extraction service 1610 respectively, the image area clue extraction service 1610 is electrically connected to the electrically connected image area clue model 1611, the image area clue model 1611 is electrically connected to the image sample library 1612, the image sample library 1612 is electrically connected to the user feedback or reporting interface service 1605 and the download file system 1613 respectively, and the download file system 1613 is electrically connected to the content storage service 1614.
[0228] The content production terminal 1620 (Professional Generated Content, PGC) provides local or filmed video content, self-written media articles, or photo albums through the mobile terminal or back-end interface API system. Authors can choose to actively upload cover images for the corresponding content, which are the main sources of content distribution. By communicating with the upstream and downstream content interface services, the upload server interface address is first obtained, and then the local file is uploaded. During the shooting process, the local video content can be matched with music, filter templates, and video beautification functions.
[0229] The content consumer (1630) (User Generated Content, UGC) communicates with the content distribution server to obtain index information for the corresponding content. For videos, it then communicates with the video storage server to download the corresponding streaming media file and play it through a local player. For images and text, it typically communicates directly with a CDN service deployed at the edge. Consumption data is viewed through a feed stream. Low-quality image content on the consumer side is directly reported and reported, and feedback is provided. The user interface directly connects to the manual review system for confirmation and review, ultimately serving as sample data for the low-quality image filtering feature model.
[0230] The content production end 1620 and the content consumption end 1630 report the user's browsing behavior data, reading speed, completion rate, reading time, freeze, loading time, playback clicks, etc. during the upload and download process to the server 1600.
[0231] The content database 1602 is used to store the content meta-information of the target information stream, which may include file size, cover image link, bit rate, file format, title, release time, author, video file size, video format, whether it is original or first release, etc.
[0232] (1) The core database of content. The metadata of all content published by producers is stored in this business database. The focus is on the metadata of the content itself, such as file size, cover image link, bit rate, file format, title, release time, author, video file size, video format, whether it is original or first release, and the classification of content during the manual review process (including first-level, second-level, and third-level classification and label information, such as a content about XX mobile phone, the first-level classification is technology, the second-level classification is smart phone, the third-level classification is domestic mobile phone, and the label information is XX);
[0233] (2) During the manual review process, the information in the content database will be read, and the results and status of the manual review will also be sent back to the content database
[0234] (3) The dispatch center service mainly processes content through machine processing and manual review. The core of machine processing is various quality judgments such as low-quality filtering, content labels such as classification, label information, and content deduplication. Their results will be written into the content database, and completely duplicate content will not be manually processed again.
[0235] The user feedback or reporting interface service 1605 is used to receive user feedback or reports of images with quality issues (such as vulgar images, incomplete images, etc.), and then call the manual review service 1606 to review such images. After review and confirmation, the images are stored in the image sample library 1612. Specifically:
[0236] (1) Receive consumption flow reports from the content review end and the content consumption end;
[0237] (2) Conduct statistical mining and analysis on the reported flow, focusing on feedback and reports related to image quality. After manual review, write them into the image sample library as the sample data source for image clue modeling (for example, mosaics, watermarks, OCR text, etc. all need to correspond to sample data that conforms to the business scenario). In this way, the samples obtained are more targeted and conform to the distribution of business problems.
[0238] The image region clue feature library 1609 is specifically used for:
[0239] (1) Provide storage and query services for image region clues required for region clue scoring and intelligent screenshot services;
[0240] (2) Typical clue features include: (a) face detection and human body detection (b) general object detection (c) text detection (d) watermark detection (e) lace detection (f) mosaic detection (g) mosaic image segmentation, etc.
[0241] Image region clue extraction service 1610 is used to train an image region clue model library 1611 based on the image region clue feature library 1609. For example, by downloading a video file, extracting frames, and obtaining images from graphic content, an image sample library 1612 is generated. The constructed image region clue model is service-oriented, creating a callable service to generate image region clues and store them in the image region clue library.
[0242] The image region clue model library 1611 is used to automatically extract various types of objects (including people, animals, watermarks, lace, mosaics, etc.) in images based on training based on the image sample library 1612. For example, based on the face detection model, facial features are extracted from the target image.
[0243] (1) Models corresponding to various features described in the image region clue library, such as face detection and basic models corresponding to text, mosaic, watermark detection, etc. These models are stored in the feature model library;
[0244] (2) To avoid model degradation, these module libraries are updated regularly, usually on a weekly basis.
[0245] The download file system 1613 is specifically used for:
[0246] (1) Download and obtain the original video content from the content storage server and control the download speed and progress. It is usually composed of a group of parallel servers with related task scheduling and distribution clusters;
[0247] (2) The downloaded file calls the frame extraction service to obtain the necessary video file key frames from the video source file as the data source for the subsequent construction of the image clue extraction service.
[0248] The regional clue scoring and intelligent screenshot service 1604 is used to score multiple preset screenshots based on the image regional clue feature library 1609, so as to determine the target screenshot of the target image and use it as a candidate cover image.
[0249] The image sample library 1612 is constructed based on the user feedback or reporting interface service 1605 and the download file system 1613, and is used to store multiple target images and cover images with unqualified quality reported by users. Specifically:
[0250] (1) Performing primary processing of video file features on the files downloaded from the video content storage service by downloading the file system - extracting video frames, including key frames and evenly extracted frames, which are used as image samples;
[0251] (2) Images with quality issues reported and reported by content consumers for manual review are used as image sample data for the quality filtering and classification model.
[0252] The dispatch center service 1607 can receive target information streams from the upstream and downstream content interface services 1608 and obtain multiple target images corresponding to the target information streams. It can also obtain content metadata of the target information streams from the content database 1602. The dispatch center service 1607 can send multiple target images or target information streams to the regional clue scoring and intelligent screenshot service 1604, the manual review service 1606, and the content deduplication service 1603. The dispatch center service 1607 can control the order and priority of scheduling and send the target information streams and cover images that have passed the manual review service 1606 to the content distribution export service 1601 for content recommendation. Specifically:
[0253] (1) Responsible for the entire scheduling process of video and graphic content flow, receiving the content through the upstream and downstream content interface servers, and then obtaining the content metadata from the content metadata database;
[0254] (2) As the actual dispatch controller of the image, text, and video links, it dispatches regional clue scoring and intelligent screenshot services to process the corresponding content according to the type of content and completes intelligent screenshots;
[0255] (3) Dispatching manual review systems and machine processing systems to control the order and priority of dispatch;
[0256] (4) After the manual review system content is enabled, it is directly displayed to the terminal content consumers through the content export distribution service (usually a recommendation engine or search engine or operation), that is, the content index information obtained by the consumer end.
[0257] The manual review service 1606 is used to manually review the target information flow or target image, and to review the unqualified cover image fed back by the content consumption end 1630. Specifically,
[0258] (1) It is usually a web system that receives the results of machine filtering on the link, manually confirms and reviews the results, and writes the review results into the content information metadata database for record. At the same time, the actual effect of the machine filtering model can be evaluated online through the results of manual review here;
[0259] (2) Report the source of the manual review task, review results, review start and end time, and other detailed review processes to the statistics server.
[0260] The content deduplication service 1603 is used to check for duplicates in a target information stream or target image.
[0261] The content storage service 1614 is used to store the source files of the target information stream. The source files of the target information stream can be obtained from the content storage service 1614 and sent to the content consumption end 1630. Specifically:
[0262] (1) It is usually a group of storage servers that are widely distributed and accessed nearby. There are also CDN (Content Delivery Network) acceleration servers on the periphery for distributed cache acceleration, which saves the video and image content uploaded by content producers through the upstream and downstream content interface servers;
[0263] (2) After obtaining the content index information, the terminal consumer can also directly access the video content storage server to download the corresponding content;
[0264] (3) In addition to being a data source for external services, it also serves as a data source for internal services, allowing the download file system to obtain raw video data for related processing. The paths of internal and external data sources are usually deployed separately to avoid mutual influence.
[0265] The content distribution export service 1601 is used to distribute the target information stream according to the target cover image and push the target information stream to the content consumption end 1630.
[0266] The uplink and downlink content interface service 1608 is used to receive the target information stream uploaded by the content production terminal 1620 and forward the target information stream to the content storage service 1614, the scheduling center service 1607 and the content database 1602 for processing. Specifically, it includes:
[0267] (1) Communicate directly with the content production end. The content submitted from the front end is usually the title, publisher, summary, cover image, and release time of the content; or the video shot directly enters the server end through the server and stores the file in the video content storage service;
[0268] (2) Writing metadata of the video content, such as video file size, cover image link, bit rate, file format, title, release time, author, etc., into the content database;
[0269] (3) Submit the uploaded files and content metadata to the dispatch center service for subsequent content processing and circulation.
[0270] The system provided in the above embodiment can execute the method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects of executing the method. For technical details not fully described in the above embodiment, please refer to an image processing method provided in any embodiment of the present application.
[0271] An embodiment of the present application provides an image processing device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the image processing method provided in the above method embodiment.
[0272] An embodiment of the present application also provides a computer storage medium, which can be set in a terminal to store at least one instruction or at least one program related to an image processing method in a method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the image processing method provided by the above method embodiment.
[0273] Embodiments of the present application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to implement the image processing method provided in the above method embodiment.
[0274] Optionally, in an embodiment of the present application, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0275] The memory described in the embodiment of the present application can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.
[0276] The image processing method provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 17 This is a hardware structure diagram of a server of an image processing method provided in an embodiment of the present application. Figure 17As shown, the server 1700 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 1710 (the central processing unit 1710 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1730 for storing data, and one or more storage media 1720 (such as one or more mass storage devices) for storing application programs 1723 or data 1722. Among them, the memory 1730 and the storage medium 1720 can be temporary storage or permanent storage. The program stored in the storage medium 1720 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1710 can be configured to communicate with the storage medium 1720 to execute a series of instruction operations in the storage medium 1720 on the server 1700. The server 1700 may also include one or more power supplies 1760, one or more wired or wireless network interfaces 1750, one or more input and output interfaces 1740, and / or one or more operating systems 1721, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0277] The input / output interface 1740 can be used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communication provider of the server 1700. In one embodiment, the input / output interface 1740 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In another embodiment, the input / output interface 1740 can be a radio frequency (RF) module for wirelessly communicating with the Internet.
[0278] It can be understood by those skilled in the art that Figure 17 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 17 More or fewer components than shown, or with Figure 17 Different configurations shown.
[0279] It can be seen from the embodiments of the image processing method, device, equipment or storage medium provided by the above-mentioned present application that the present application first determines at least two screenshots corresponding to the target image in the target information flow, and then determines the importance index result of the first object in each screenshot based on the area ratio of the second object corresponding to the first object in each screenshot in the target image; the first object is part or all of the second object; the second object is an object in the target image; thereby, the importance of the first object in each screenshot in the original image can be determined; based on the preset area of the first object in each screenshot and the overlapping area of each screenshot, the area ratio result of the first object in each screenshot is determined; the preset The area is the area occupied by the second object corresponding to the first object in the target image; thereby, the degree of completeness of the first object relative to the corresponding second object in each screenshot can be determined; based on the importance index result and the area proportion result of the first object in each screenshot, the comprehensive score of each screenshot corresponding to the target image is determined; based on the comprehensive score of each screenshot corresponding to the target image, the target screenshot corresponding to the target image is determined; thereby, a target screenshot that can well represent the target image can be determined, and the target screenshot can be used to determine the cover image of the target information flow, thereby improving the generation efficiency and quality of the cover image, and facilitating improving the user's click rate and consumption time on the target information flow.
[0280] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0281] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant portions, refer to the descriptions of the method embodiments.
[0282] Those skilled in the art will understand that all or part of the steps of implementing the above embodiments may be accomplished by hardware, or by a program instructing the relevant hardware to accomplish the steps. The program may be stored in a computer storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0283] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Determine at least two screenshots corresponding to the target image in the target information flow, wherein the at least two screenshots include respective corresponding first objects; Determining an importance index result of the first object in each screenshot based on an area ratio in the target image of a second object corresponding to the first object in each screenshot; the first object being part or all of the second object; and the second object being an object in the target image; the importance index result representing a degree of influence of the first object in each screenshot on determining each screenshot as a target screenshot; Determining an area ratio of the first object in each screenshot based on an overlapping area between a preset area of the first object in each screenshot and each screenshot; the preset area being an area occupied by a second object corresponding to the first object in the target image; the area ratio representing a degree of completeness of the first object relative to the corresponding second object in each screenshot; Determining a comprehensive score for each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot; the comprehensive score of each screenshot represents the degree of association between each screenshot and the target image; Based on the comprehensive scores of the screenshots corresponding to the target image, a target screenshot corresponding to the target image is determined; the target screenshot is used to determine the cover image of the target information flow.
2. The method according to claim 1, characterized in that The target images are a preset number, which is at least two, and the at least two screenshots include candidate screenshots. Determining the importance index result of the first object in each screenshot based on the area ratio of the second object corresponding to the first object in each screenshot in the target image includes: Determine an area ratio of a second target object corresponding to the first target object in the candidate screenshot in the candidate target image; the candidate screenshot is a screenshot corresponding to the candidate target image; Determining an area ratio interval corresponding to the first target object based on an area ratio corresponding to the first target object in the candidate screenshot; Based on the first area ratio of the second object in each target image, determining the target images whose first area ratio is within the area ratio interval as first screening images; determining a first number of second objects in the first screening image; Based on the first number corresponding to the first target object in the candidate screenshot and the preset number, an importance index result of the first target object in the candidate screenshot is determined.
3. The method according to claim 2, characterized in that Each target image includes at least two second objects, and determining the importance index result of the first target object in the candidate screenshot based on the first number corresponding to the first target object in the candidate screenshot and the preset number includes: Based on the second area ratio of each second object in each target image, determining the target images whose second area ratios are within the area ratio interval as second screening images; determining a second number of second objects in the second screening image; determining a target image including the first target object as a third screening image, and determining a third number of the third screening images; Based on the first quantity, the second quantity, the third quantity and the preset quantity corresponding to the first target object in the candidate screenshot, an importance index result of the first target object in the candidate screenshot is determined.
4. The method according to claim 2, characterized in that The determining a target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image includes: Determining a target screenshot corresponding to each target image based on the comprehensive scores of the screenshots corresponding to each target image; After determining the target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image, the method further includes: Based on the target screenshots corresponding to the target images, a cover image of the target information flow is determined.
5. The method according to claim 1, wherein The determining of an area ratio of the first object in each screenshot based on an overlapping area between a preset area of the first object in each screenshot and each screenshot includes: Determining a preset area of the first object in each screenshot and a screenshot area corresponding to each screenshot; Determine an overlapping area and a merged area between the preset area and the screenshot area; Determine a first area corresponding to the overlapping area and a second area corresponding to the merged area; The ratio of the first area to the second area is used as the area ratio result of the first object in each screenshot.
6. The method according to claim 2, characterized in that The candidate screenshots include at least two first target objects, and determining the comprehensive score of each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot includes: Determine the importance index result and area proportion result of each first target object in the candidate screenshot; Determine the comprehensive score of each first target object in the candidate screenshot by multiplying the importance index result and the area proportion result of each first target object in the candidate screenshot; Calculating an average of the comprehensive scores of the first target objects in the candidate screenshots based on the comprehensive scores of each first target object in the candidate screenshots; The average of the comprehensive scores is used as the comprehensive score of the candidate screenshots.
7. The method according to claim 6, characterized in that The taking the comprehensive score mean as the comprehensive score of the candidate screenshots includes: Obtain the size of the display area and the size of the candidate screenshot; Determining a normalization coefficient based on the size of the display area and the size of the candidate screenshot; Based on the normalization coefficient, the comprehensive score mean is normalized to obtain the comprehensive score of the candidate screenshot.
8. The method according to claim 1, characterized in that The determining a target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image includes: sorting the screenshots based on the comprehensive scores of the screenshots corresponding to the target image to obtain a first sorting result; Based on the first sorting result, a target screenshot corresponding to the target image is determined.
9. The method according to claim 4, characterized in that The determining of the cover image of the target information stream based on the target screenshots corresponding to the target images includes: Acquire subject information of the target information flow; sorting the target screenshots based on the relevance between each target screenshot and the subject information to obtain a second sorting result; Based on the second sorting result, a cover image of the target information flow is determined.
10. An image processing device, characterized in that: The device comprises: A screenshot determining module, configured to determine at least two screenshots corresponding to a target image in a target information stream, wherein the at least two screenshots include respective corresponding first objects; an importance index result determination module, configured to determine an importance index result of the first object in each screenshot based on a proportion of an area of a second object corresponding to the first object in the target image; the first object being part or all of the second object; the second object being an object in the target image; the importance index result representing a degree of influence of the first object in each screenshot on the determination of each screenshot as a target screenshot; an area ratio result determination module, configured to determine an area ratio result of the first object in each screenshot based on an overlapping area between a preset area of the first object in each screenshot and each screenshot; the preset area being the area occupied by a second object corresponding to the first object in the target image; the area ratio result representing a degree of completeness of the first object relative to the corresponding second object in each screenshot; a comprehensive score determination module, configured to determine a comprehensive score for each screenshot corresponding to the target image based on the importance index result and the area ratio result of the first object in each screenshot; the comprehensive score of each screenshot represents the degree of association between each screenshot and the target image; A target screenshot determination module is used to determine a target screenshot corresponding to the target image based on the comprehensive scores of the screenshots corresponding to the target image; the target screenshot is used to determine the cover image of the target information flow.
Citation Information
Patent Citations
Cover image determination method, device and equipment
CN110602554A
Cover image acquisition method and device
CN113254696A