A method and matching device for generating a thumbnail based on information flow

By using natural language processing and computer vision to match text and image semantics, the method enhances thumbnail generation in information flows, improving user engagement through relevant and high-quality thumbnails.

CN114220115BActive Publication Date: 2025-07-15BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010915689.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-03
Publication Date
2025-07-15
Estimated Expiration
2040-12-27

AI Technical Summary

Technical Problem

In the prior art, the image and text in the information flow have low matching ability, resulting in a low visual quality of the generated thumbnails, affecting the user's browsing experience.

Method used

Text semantics are obtained through natural language processing technology, and the image semantics of candidate images are processed in combination with computer vision technology, matching images with text semantics, and generating thumbnails that meet preset display sizes and categories.

Benefits of technology

Improve the matching and visual quality of thumbnails and text, increase the user's attractiveness to information flow, and improve the browsing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114220115B_ABST
    Figure CN114220115B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method and related device for generating a thumbnail based on an information stream. The method includes: obtaining the text semantics of the text in the information stream by using natural language processing technology; processing multiple candidate images corresponding to the text by using computer vision technology to obtain the image semantics of each candidate image; matching the image semantics of each candidate image with the text semantics of the text, and determining at least one candidate image as the target image from the multiple candidate images; and generating a thumbnail by processing the target image by using computer vision technology based on the category of the target image and a preset display size. It can be seen that the target image is determined by matching the text semantics of the text with the image semantics of each candidate image, and the generated thumbnail has a relatively high matching degree with the text in the information stream, improving the matching between the thumbnail and the text; and when generating the thumbnail, the category of the image and the preset display size are considered, and different thumbnail methods are adopted to improve the visual quality of the thumbnail.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of data processing, and in particular, to a method for generating a thumbnail based on information flow and a matching device. Background Art

[0002] With the rapid development of information technology, more and more information is presented to users in the form of images. Images can display information that is difficult to express in text, and compared with text, they have more display advantages and can attract users' attention more. The sizes of the images corresponding to the text in the information flow are inconsistent. Directly displaying the original images causes the entire display layout of the information flow to be chaotic; if the size of the original image is too large and occupies a large display layout, it also causes too many images to be displayed in the display layout of the information flow, affecting the text display volume.

[0003] Before the information flow is displayed, it is particularly important to crop or scale the images corresponding to the text to obtain appropriate regional images as thumbnails, so that the entire display layout of the information flow is neatly typeset and the number of images displayed in the display layout is appropriate. Generally, the text in the information flow corresponds to multiple images. In the prior art, the method for generating thumbnails based on the information flow actually generates thumbnails for multiple images separately, so as to display the multiple thumbnails corresponding to the text in the information flow.

[0004] However, through research, it is found that there are images with low matching degree with the text among the multiple images corresponding to the text in the information flow. When using the method for generating thumbnails in the prior art, the phenomenon of non - matching between the thumbnail and the text is likely to occur when displaying the information flow; and the method for generating thumbnails in the prior art uses a unified thumbnail method for all images, and the visual quality of the generated thumbnails is relatively low, resulting in insufficient attraction of the displayed information flow to users, thus seriously affecting the user experience of browsing the information flow. Summary of the Invention

[0005] In view of this, there is an urgent need to provide a method for generating a thumbnail based on information flow and a matching device, so as to display the generated thumbnail when displaying the information flow, which can improve the matching degree between the thumbnail and the text, increase the attraction to users, and thus improve the user experience of browsing the information flow.

[0006] In a first aspect, an embodiment of the present invention provides a method for generating a thumbnail based on information flow, and the method includes:

[0007] Perform natural language processing on the text in the information flow to obtain the text semantics of the text;

[0008] Perform computer vision processing on multiple candidate images corresponding to the text to obtain the image semantics of each candidate image;

[0009] Match the image semantics of each candidate image with the text semantics of the text, and determine at least one candidate image from the multiple candidate images as the target image;

[0010] Based on the category of the target image and a preset display size, perform computer vision processing on the target image to generate a thumbnail.

[0011] Optionally, the matching the image semantics of each candidate image with the text semantics of the text, and determining at least one candidate image from the multiple candidate images as the target image includes:

[0012] Match the image semantics of each candidate image with the text semantics of the text, and obtain the matching degree between the image semantics of each candidate image and the text semantics of the text;

[0013] Determine the top N candidate images as the target images from the multiple candidate images according to the matching degree from high to low; N is a positive integer less than the total number of the multiple candidate images.

[0014] Optionally, the performing computer vision processing on the target image based on the category of the target image and a preset display size to generate a thumbnail is specifically:

[0015] If the category of the target image is a general category, perform computer vision processing on the target image based on the positions of the key elements in the target image and the preset display size to generate the thumbnail; or,

[0016] If the category of the target image is an icon category, perform computer vision processing on the icon area in the target image based on the background color of the target image and the preset display size to generate the thumbnail; or,

[0017] If the category of the target image is a portrait category, perform computer vision processing on the face and neck and shoulder areas in the target image based on the edge color of the target image and the preset display size to generate the thumbnail; or,

[0018] If the category of the target image is a text / chart category, perform computer vision processing on the title area of the text / chart in the target image based on the preset display size to generate the thumbnail.

[0019] Optionally, the performing computer vision processing on the target image based on the positions of the key elements in the target image and the preset display size to generate the thumbnail includes:

[0020] Based on the positions of the key elements in the target image, determine candidate regions in the target image;

[0021] Performing computer vision processing on the candidate region based on the preset display size to generate the thumbnail.

[0022] Optionally, the performing computer vision processing on the candidate region based on the preset display size to generate the thumbnail specifically includes:

[0023] When the candidate region meets the preset display size, performing computer vision processing on the candidate region according to the preset display size to generate the thumbnail; or,

[0024] When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is greater than the preset distance, deleting the region between key elements in the candidate region to obtain an updated candidate region; performing computer vision processing on the updated candidate region according to the preset display size to generate the thumbnail; or,

[0025] When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is less than or equal to the preset distance or there is only one key element in the candidate region, performing computer vision processing on the candidate region based on the background color of the target image and the preset display size to generate the thumbnail.

[0026] Optionally, before performing computer vision processing on the multiple candidate images corresponding to the text, it further includes:

[0027] Obtaining multiple images corresponding to the text in the information stream;

[0028] Performing computer vision processing on the multiple images corresponding to the text, and filtering sensitive images from the multiple images to obtain multiple candidate images corresponding to the text.

[0029] Optionally, the sensitive images include quality-sensitive images, category-sensitive images, and / or content-sensitive images; the quality-sensitive images are specifically images with image quality lower than the preset image quality, the category-sensitive images are specifically images whose image category belongs to the preset sensitive image category, and the content-sensitive images are specifically images with image content sensitivity higher than the preset image content sensitivity.

[0030] In a second aspect, an embodiment of the present invention provides a device for generating thumbnails based on an information stream, and the device includes:

[0031] A first obtaining unit, configured to perform natural language processing on the text in the information stream to obtain the text semantics of the text;

[0032] A second obtaining unit, configured to perform computer vision processing on the multiple candidate images corresponding to the text to obtain the image semantics of each candidate image;

[0033] A determination unit, configured to match the image semantics of each of the candidate images with the text semantics of the text, and determine at least one candidate image from the multiple candidate images as the target image;

[0034] A generation unit, configured to perform computer vision processing on the target image based on the category of the target image and a preset display size to generate a thumbnail.

[0035] Optionally, the determination unit includes:

[0036] An obtaining subunit, configured to match the image semantics of each of the candidate images with the text semantics of the text, and obtain the matching degree between the image semantics of each of the candidate images and the text semantics of the text;

[0037] A first determination subunit, configured to determine the first N candidate images as the target images from the multiple candidate images in descending order of the matching degree; N is a positive integer less than the total number of the multiple candidate images.

[0038] Optionally, the generation unit is specifically configured to:

[0039] If the category of the target image is a general category, perform computer vision processing on the target image based on the positions of the respective key elements in the target image and the preset display size to generate the thumbnail; or,

[0040] If the category of the target image is an icon category, perform computer vision processing on the icon area in the target image based on the background color of the target image and the preset display size to generate the thumbnail; or,

[0041] If the category of the target image is a portrait category, perform computer vision processing on the face and neck area in the target image based on the edge color of the target image and the preset display size to generate the thumbnail; or,

[0042] If the category of the target image is a text / chart category, perform computer vision processing on the title area of the text / chart in the target image based on the preset display size to generate the thumbnail.

[0043] Optionally, when the generation unit is specifically configured to perform computer vision processing on the target image based on the positions of the respective key elements in the target image and the preset display size if the category of the target image is a general category, the generation unit includes:

[0044] A second determination subunit, configured to determine a candidate area in the target image based on the positions of the respective key elements in the target image;

[0045] A generating subunit, configured to perform computer vision processing on the candidate region based on the preset display size to generate the thumbnail.

[0046] Optionally, the generating subunit is specifically configured to:

[0047] When the candidate region meets the preset display size, perform computer vision processing on the candidate region according to the preset display size to generate the thumbnail; or,

[0048] When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is greater than a preset distance, delete the region between key elements in the candidate region to obtain an updated candidate region; perform computer vision processing on the updated candidate region according to the preset display size to generate the thumbnail; or,

[0049] When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is less than or equal to the preset distance or there is only one key element in the candidate region, perform computer vision processing on the candidate region based on the background color of the target image and the preset display size to generate the thumbnail.

[0050] Optionally, the apparatus further includes:

[0051] An obtaining unit, configured to obtain a plurality of images corresponding to the text in the information flow;

[0052] A third obtaining unit, configured to perform computer vision processing on the plurality of images corresponding to the text, and filter out sensitive images from the plurality of images to obtain a plurality of candidate images corresponding to the text.

[0053] Optionally, the sensitive images include quality-sensitive images, category-sensitive images, and / or content-sensitive images; the quality-sensitive images are specifically images with an image quality lower than a preset image quality, the category-sensitive images are specifically images whose image category belongs to a preset sensitive image category, and the content-sensitive images are specifically images with an image content sensitivity higher than a preset image content sensitivity.

[0054] In a third aspect, an embodiment of the present invention provides an apparatus for generating a thumbnail based on an information flow. The apparatus includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations:

[0055] Perform natural language processing on the text in the information flow to obtain the text semantics of the text;

[0056] Perform computer vision processing on multiple candidate images corresponding to the text to obtain the image semantics of each candidate image;

[0057] Match the image semantics of each candidate image with the text semantics of the text, and determine at least one candidate image as the target image from the multiple candidate images;

[0058] Based on the category and preset display size of the target image, perform computer vision processing on the target image to generate a thumbnail.

[0059] In a fourth aspect, an embodiment of the present invention provides a machine-readable medium, on which instructions are stored. When executed by one or more processors, the instructions cause the device to execute the method for generating a thumbnail based on an information stream according to any one of the above first aspects.

[0060] Compared with the prior art,

[0061] Adopting the technical solution of the embodiment of the present invention, the text semantics of the text in the information stream are obtained by using natural language processing technology; multiple candidate images corresponding to the text are processed by using computer vision technology to obtain the image semantics of each candidate image; the image semantics of each candidate image are matched with the text semantics of the text, and at least one candidate image is determined as the target image from the multiple candidate images; based on the category and preset display size of the target image, the target image is processed by using computer vision technology to generate a thumbnail. It can be seen that the target image is determined by obtaining the text semantics of the text and the image semantics of each candidate image, and matching the image semantics of each candidate image with the text semantics of the text. The generated thumbnail has a high matching degree with the text in the information stream. When displaying the information stream, displaying the thumbnail can improve the matching between the thumbnail and the text; and when generating the thumbnail, considering the category and preset display size of the image, different thumbnail methods are adopted to improve the visual quality of the thumbnail, increase the attractiveness of the thumbnail to users, and thus improve the user experience of browsing the information stream. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a schematic diagram of the system framework involved in an application scenario in an embodiment of the present invention;

[0064] Figure 2Schematic flowchart of a method for generating a thumbnail based on information flow provided by an embodiment of the present invention;

[0065] Figure 3 Schematic diagram of generating a thumbnail when the candidate region in a target image of a general category meets the preset display size and all key elements in the candidate region are complete, provided by an embodiment of the present invention;

[0066] Figure 4 Schematic diagram of generating a thumbnail when the candidate region in a target image of a general category does not meet the preset display size and the intermediate distance in the candidate region is greater than the preset distance, provided by an embodiment of the present invention;

[0067] Figure 5 Schematic diagram of generating a thumbnail when the candidate region in a target image of a general category does not meet the preset display size and there is only one key element in the candidate region, provided by an embodiment of the present invention;

[0068] Figure 6 Schematic diagram of generating a thumbnail for a target image of an icon category, provided by an embodiment of the present invention;

[0069] Figure 7 Schematic diagram of generating a thumbnail for a target image of a portrait category, provided by an embodiment of the present invention;

[0070] Figure 8 Schematic diagram of generating a thumbnail for a target image of a text / chart category, provided by an embodiment of the present invention;

[0071] Figure 9 Schematic flowchart of another method for generating a thumbnail based on information flow provided by an embodiment of the present invention;

[0072] Figure 10 Schematic structural diagram of a device for generating a thumbnail based on information flow provided by an embodiment of the present invention;

[0073] Figure 11 Schematic structural diagram of a device for generating a thumbnail based on information flow provided by an embodiment of the present invention;

[0074] Figure 12 Schematic structural diagram of a server provided by an embodiment of the present invention. Detailed implementation manners

[0075] To enable those skilled in the art to better understand the solutions of the embodiments of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the embodiments of the present invention.

[0076] At present, in order to avoid the direct display of the original images of multiple images corresponding to the text when displaying the information stream, the inconsistent image sizes may cause chaos in the entire display layout of the information stream; or the overly large size of the original images may result in too many images being displayed in the display layout of the information stream, affecting the text display amount. Before displaying the information stream, the images corresponding to the text are cropped or scaled to obtain appropriate regional images as thumbnails for subsequent display of the information stream when the corresponding text is displayed. In the prior art, the method of generating thumbnails based on the information stream actually generates thumbnails for multiple images separately so that multiple thumbnails can be displayed corresponding to the text in the information stream. However, among the multiple images corresponding to the text in the information stream, there are images with low matching degree with the text. Using the above prior art method, it is easy to have the phenomenon that the thumbnail does not match the text when displaying the information stream. Moreover, all the images in the above prior art method adopt a unified thumbnail method, and the visual quality of the generated thumbnails is relatively low, resulting in insufficient attraction of the displayed information stream to users, thus seriously affecting the user experience of browsing the information stream.

[0077] To solve this problem, in the embodiments of the present invention, the text semantics of the text in the information stream is obtained by using natural language processing technology; multiple candidate images corresponding to the text are processed by using computer vision technology to obtain the image semantics of each candidate image; the image semantics of each candidate image is matched with the text semantics of the text, and at least one candidate image is determined as the target image from the multiple candidate images; based on the category and preset display size of the target image, the computer vision technology is used to process the target image to generate a thumbnail. It can be seen that the target image is determined by obtaining the text semantics of the text and the image semantics of each candidate image, and matching the image semantics of each candidate image with the text semantics of the text. The generated thumbnail has a high matching degree with the text in the information stream. Displaying this thumbnail when displaying the information stream can improve the matching degree between the thumbnail and the text; and when generating the thumbnail, considering the category and preset display size of the image, different thumbnail methods are adopted to improve the visual quality of the thumbnail, increase the attraction of the thumbnail to users, and thus improve the user experience of browsing the information stream.

[0078] For example, one of the scenarios of the embodiments of the present invention can be applied to, such as Figure 1In the scenario shown, the scenario includes a terminal 101 and a processor 102. The terminal 101 sends an information stream including text and multiple images corresponding thereto to the processor 102. The processor 102 generates thumbnails based on the information stream using an implementation method of an embodiment of the present invention and returns them to the terminal 101, so that the terminal 101 displays thumbnails corresponding to the text when displaying the information stream.

[0079] It can be understood that in the above application scenario, although the action description of the implementation method provided by the embodiment of the present invention is executed by the processor 102, the embodiment of the present invention is not limited in terms of the execution subject, as long as the actions disclosed in the implementation method of this application are executed.

[0080] It can be understood that the above scenario is only an example scenario provided by the embodiment of the present invention, and the embodiment of the present invention is not limited to this scenario.

[0081] The specific implementation of the method for generating thumbnails based on information flow and the matching device in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings by way of embodiments.

[0082] Exemplary method

[0083] See also Figure 2 , shows a flow chart of a method for generating thumbnails based on information flow in an embodiment of the present invention. In this embodiment, the method may include the following steps:

[0084] Step 201: Perform natural language processing on the text in the information flow to obtain the text semantics of the text.

[0085] It should be noted that the information stream includes text and multiple images corresponding to the text. Since there are likely to be images with low matching with the text among the multiple images corresponding to the text, if thumbnails are directly generated for each image and multiple thumbnails are displayed corresponding to the text, it is easy for the thumbnails to not match the text, resulting in insufficient appeal of the displayed information stream to the user, thus seriously affecting the user's browsing experience of the information stream. Therefore, in the embodiment of the present invention, it is necessary to first consider the matching of each image and the text, avoid generating thumbnails for images that do not match the text and displaying them corresponding to the text; then it is necessary to use natural language processing technology to obtain the text semantics of the text in the information stream.

[0086] Specifically, natural language processing can be, for example, a TextRank algorithm or a Lexrank algorithm, etc. The TextRank algorithm or the Lexrank algorithm can automatically calculate the weight of each word in the text to extract keywords in the text as the text semantics of the text. It should be noted that the natural language processing in the embodiment of the present invention includes but is not limited to the TextRank algorithm or the Lexrank algorithm.

[0087] Step 202: Perform computer vision processing on multiple candidate images corresponding to the text to obtain the image semantics of each candidate image.

[0088] It should be noted that in the embodiments of the present invention, in order to consider the matching degree between each image and the text and avoid generating thumbnails for images that do not match the text and displaying them corresponding to the text, it is also necessary to process multiple candidate images corresponding to the text by using computer vision technology to obtain the image semantics of each candidate image. The image semantics of the candidate image specifically may refer to the categories of various key elements in the candidate image.

[0089] Specifically, the computer vision processing may be, for example, image key element recognition or image key element detection, etc. It should be noted that the computer vision processing in the embodiments of the present invention includes but is not limited to image key element recognition or image key element detection.

[0090] It should also be noted that multiple images corresponding to the text in the information stream can be directly used as multiple candidate images corresponding to the text in Step 202; or a part of the multiple images corresponding to the text in the information stream can be used as multiple candidate images corresponding to the text in Step 202. For a detailed description, reference can be made to the following embodiments.

[0091] It should be further noted that in the embodiments of the present invention, the execution order of Step 201 and Step 202 is not limited. Step 201 can be executed first and then Step 202, or Step 202 can be executed first and then Step 201, or Step 201 and Step 202 can be executed simultaneously.

[0092] Step 203: Match the image semantics of each candidate image and the text semantics of the text, and determine at least one candidate image as the target image from the multiple candidate images.

[0093] It should be noted that in order to avoid generating thumbnails for images that do not match the text and displaying them corresponding to the text, in the embodiments of the present invention, it is necessary to match the image semantics of each candidate image and the text semantics of the text, and determine at least one candidate image with a relatively high matching degree with the text from the multiple candidate images as the target image.

[0094] Specifically, by matching the image semantics of each candidate image with the text semantics of the text, the matching degree between the image semantics in each candidate image and the text semantics of the text can actually be obtained. After the above matching operations are completed for multiple candidate images, corresponding multiple matching degrees can be obtained. According to the order from high to low of the multiple matching degrees, the corresponding multiple candidate images can be sorted. The candidate image ranked higher has a higher matching degree with the text. Selecting the top N candidate images after sorting can be used as the target images, and N must be a positive integer less than the total number of multiple candidate images. Therefore, in an optional implementation manner of the embodiment of the present invention, the step 203 may include the following steps, for example:

[0095] Step A: Match the image semantics of each candidate image with the text semantics of the text to obtain the matching degree between the image semantics of each candidate image and the text semantics of the text;

[0096] Step B: Determine the top N candidate images from the multiple candidate images as the target images according to the order from high to low of the matching degree; N is a positive integer less than the total number of the multiple candidate images.

[0097] As an example, sort the multiple candidate images according to the order from high to low of the matching degree between the image semantics in each candidate image and the text semantics of the text, and determine the first candidate image after sorting as the target image, and this target image has the highest matching degree with the text. As another example, the total number of multiple candidate images is 6. Sort the multiple candidate images according to the order from high to low of the matching degree between the image semantics in each candidate image and the text semantics of the text, and determine the top 2 candidate images after sorting as the target images. At this time, the number of target images is 2.

[0098] Step 204: Based on the category and preset display size of the target image, perform computer vision processing on the target image to generate a thumbnail.

[0099] It should be noted that after the target image is determined in step 203, a thumbnail of the target image needs to be generated so that the corresponding text of the thumbnail can be displayed when presenting the information stream. If the method of generating thumbnails in the prior art is adopted and a unified thumbnail method is used for all images, the visual quality of the generated thumbnails is relatively low, resulting in insufficient attraction of the presented information stream to users, thus seriously affecting the user experience of browsing the information stream. Therefore, in the embodiments of the present invention, considering that the categories of different images may be different and the characteristics of images in different categories are different, different thumbnail methods need to be adopted when generating thumbnails for different categories of images; and the thumbnails generated by the images need to conform to the display size of a unified display device, denoted as the preset display size; after the target image is determined in step 203, during the process of generating a thumbnail of the target image, both the category of the target image and the preset display size need to be considered, and the corresponding thumbnail method is used to perform computer vision processing on the target image to generate a thumbnail of the target image.

[0100] In an alternative implementation manner of the embodiments of the present invention, the category of the target image may be a general category, an icon category, a portrait category, and a text / chart category. The following takes the target images of the above four different categories as examples in turn to respectively and specifically describe four specific implementation manners corresponding to step 204:

[0101] For the first specific implementation manner of step 204, when generating a thumbnail for a target image of the general category, the key is to focus on each key element area in the target image, and basically do not need to pay attention to other areas, that is, the thumbnail needs to include each key element area in the target image; therefore, in the process of using computer vision technology to process the target image to generate a thumbnail, it is necessary to consider the positions of each key element in the target image and the preset display size. That is, in an alternative implementation manner of the embodiments of the present invention, step 204 is specifically, for example: if the category of the target image is the general category, based on the positions of each key element in the target image and the preset display size, perform computer vision processing on the target image to generate the thumbnail.

[0102] Specifically, first, the positions of each key element in the target image should be considered, and a candidate area including each key element is determined in the target image. The principles for determining the candidate area include but are not limited to: as much as possible to ensure that each key element in the target image exists, as much as possible to ensure that each key element is complete, as much as possible to ensure that each key element is not deformed, as much as possible to ensure that each key element is clear, etc.; then, considering the preset display size, use computer vision technology to process the candidate area to generate a thumbnail. Therefore, in an alternative implementation manner of the embodiments of the present invention, step 204 may include the following steps, for example:

[0103] Step C: Determine the candidate regions in the target image based on the positions of the key elements in the target image;

[0104] Step D: Perform computer vision processing on the candidate regions based on the preset display size to generate the thumbnail.

[0105] Among them, whether the candidate regions determined in Step C meet the preset display size and the situations of the key elements in the candidate regions determine the specific implementation manner of Step D to perform computer vision processing on the candidate regions based on the preset display size to generate the thumbnail. The following respectively elaborates on three specific implementation manners corresponding to Step D in detail according to whether the candidate regions meet the preset display size and different situations of the key elements in the candidate regions.

[0106] The first specific implementation manner of Step D: When the candidate regions meet the preset display size, based on the above principles for determining the candidate regions, all the key elements in the candidate regions are complete, indicating that the candidate regions can be directly used to generate the thumbnail; then directly process the candidate regions using computer vision technology according to the preset display size to generate the thumbnail; for example, as Figure 3 shown in the schematic diagram of generating a thumbnail when the candidate regions in a general category of target images meet the preset display size and all the key elements in the candidate regions are complete. Therefore, in an optional implementation manner of the embodiments of the present invention, Step D may specifically be, for example: When the candidate regions meet the preset display size, perform computer vision processing on the candidate regions according to the preset display size to generate the thumbnail.

[0107] The second specific implementation manner of Step D: When the candidate regions do not meet the preset display size, if the distance between any two adjacent key elements among all the key elements in the candidate regions is large at this time, it indicates that the effect of directly using the candidate regions to generate the thumbnail is poor, then it can be considered to first delete the region between the two adjacent key elements with a large distance in the candidate regions to obtain updated candidate regions, so that the updated candidate regions are as close as possible to the preset display size; then process the updated candidate regions using computer vision technology according to the preset display size to generate the thumbnail, which can improve the quality of the thumbnail; for example, as Figure 4 shown in the schematic diagram of generating a thumbnail when the candidate regions in a general category of target images do not meet the preset display size and the distance in the middle of the candidate regions is greater than the preset distance. Therefore, in an optional implementation manner of the embodiments of the present invention, Step D may specifically be, for example: When the candidate regions do not meet the preset display size, if the distance between the key elements in the candidate regions is greater than the preset distance, delete the region between the key elements in the candidate regions to obtain updated candidate regions; perform computer vision processing on the updated candidate regions according to the preset display size to generate the thumbnail.

[0108] The third specific implementation manner of step D: When the candidate region does not meet the preset display size, if the distance between any two adjacent key elements among all the key elements in the candidate region is small at this time or there is only one key element in the candidate region, if only considering the preset display size and using computer vision technology to process the candidate region to generate a thumbnail, it may cause the key elements in the thumbnail to be incomplete. Then, the background color of the target image can be obtained. During the process of using computer vision technology to process the candidate region to generate a thumbnail, on the basis of the region that is smaller than the preset display size and contains all the key elements completely, it is filled with the background color of the target image so that the generated thumbnail meets the preset display size. For example, as Figure 5 shown in the schematic diagram of generating a thumbnail when the candidate region in a general category of target image does not meet the preset display size and there is only one key element in the candidate region. Therefore, in an optional implementation manner of the embodiment of the present invention, step D may specifically be: When the candidate region does not meet the preset display size, if the distance between the key elements in the candidate region is less than or equal to the preset distance or there is only one key element in the candidate region, based on the background color of the target image and the preset display size, perform computer vision processing on the candidate region to generate the thumbnail.

[0109] The second specific implementation manner of step 204: Since when generating a thumbnail for a target image of the icon category, the complete icon needs to be focused on, that is, the thumbnail needs to include the icon region in the target image. Therefore, it is necessary to use computer vision technology to process the icon region in the target image to generate a thumbnail. And since the non-icon region in the target image of the icon category is very small, if only considering the preset display size and using computer vision technology to process the icon region in the target image to generate a thumbnail, it may cause the icon in the thumbnail to be incomplete. Then, the background color of the target image can be obtained. During the process of using computer vision technology to process the icon region in the target image to generate a thumbnail, on the basis of the region that is smaller than the preset display size and contains the complete icon, it is filled with the background color of the target image so that the generated thumbnail meets the preset display size. For example, as Figure 6 shown in the schematic diagram of generating a thumbnail for a target image of the icon category. That is, in an optional implementation manner of the embodiment of the present invention, step 204 may specifically be: If the category of the target image is the icon category, based on the background color of the target image and the preset display size, perform computer vision processing on the icon region in the target image to generate the thumbnail.

[0110] The third specific implementation manner of step 204: When generating a thumbnail for a target image of the portrait category, the key areas to focus on are the face and the shoulders and neck near it. That is, the thumbnail needs to include the face and the shoulders and neck areas in the target image. Therefore, computer vision technology is required to process the face and shoulders and neck areas in the target image to generate the thumbnail. If only considering the preset display size and using computer vision technology to process the face and shoulders and neck areas to generate the thumbnail, it may result in an incomplete face and the shoulders and neck near it in the thumbnail. Moreover, for a target image of the portrait category, the edge color is emphasized, and the edge color of the target image also needs to be obtained. During the process of using computer vision technology to process the face and shoulders and neck areas in the target image to generate the thumbnail, based on the area that contains the complete face and the shoulders and neck near it and is smaller than the preset display size, it is stitched with the edge color of the target image, so that the generated thumbnail meets the preset display size. For example, as Figure 7 shown in the schematic diagram of generating a thumbnail for a target image of a portrait category. That is, in an alternative implementation manner of the embodiment of the present invention, step 204 is specifically: If the category of the target image is the portrait category, based on the edge color of the target image and the preset display size, perform computer vision processing on the face and shoulders and neck areas in the target image to generate the thumbnail.

[0111] The fourth specific implementation manner of step 204: When generating a thumbnail for a target image of the text / chart category, the key area to focus on is the title of the text / chart. That is, the thumbnail needs to include the title area of the text / chart in the target image. Moreover, the title area of the text / chart in the target image of the text / chart category is generally small. Therefore, generating the thumbnail by using computer vision technology to process the title area of the text / chart in the target image according to the preset display size is sufficient. For example, as Figure 8 shown in the schematic diagram of generating a thumbnail for a target image of a text / chart category. That is, in an alternative implementation manner of the embodiment of the present invention, step 204 is specifically: If the category of the target image is the text / chart category, based on the preset display size, perform computer vision processing on the title area of the text / chart in the target image to generate the thumbnail.

[0112] Through various embodiments provided in this embodiment, natural language processing technology is used to obtain the text semantics of the text in the information flow; computer vision technology is used to process multiple candidate images corresponding to the text to obtain the image semantics of each candidate image; the image semantics of each candidate image is matched with the text semantics of the text, and at least one candidate image is determined as the target image from multiple candidate images; based on the category and preset display size of the target image, computer vision technology is used to process the target image to generate a thumbnail. It can be seen that the target image is determined by obtaining the text semantics of the text and the image semantics of each candidate image, and matching the image semantics of each candidate image with the text semantics of the text. The generated thumbnail has a high matching degree with the text in the information flow. When displaying the information flow, displaying this thumbnail can improve the matching of the thumbnail and the text; and when generating the thumbnail, considering the category of the image and the preset display size, different thumbnail methods are adopted to improve the visual quality of the thumbnail and increase the attractiveness of the thumbnail to users, thereby improving the user experience of browsing the information flow.

[0113] It should be noted that among the multiple images corresponding to the text in the information flow, there may be images that are not suitable for display to users, which are called sensitive images; therefore, corresponding to the description of the multiple candidate images corresponding to the text in step 202, before performing step 202, it is also necessary to use computer vision technology to process the multiple images corresponding to the text in the information flow to filter out the sensitive images in the multiple images corresponding to the text in the information flow, and use the filtered multiple images as the multiple candidate images corresponding to the text. The following combines with the attached Figure 9 On the basis of the above embodiment, through another embodiment, the specific implementation manner of another method and matching device for generating thumbnails based on information flow in the embodiments of the present invention will be described in detail.

[0114] See Figure 9 , which shows a schematic flowchart of another method for generating thumbnails based on information flow in the embodiments of the present invention. In this embodiment, the method may include the following steps, for example:

[0115] Step 901: Perform natural language processing on the text in the information flow to obtain the text semantics of the text.

[0116] It should be noted that step 901 in this embodiment is the same as step 201 in the above embodiment. For the specific implementation manner, please refer to the above embodiment and will not be elaborated here.

[0117] Step 902: Obtain multiple images corresponding to the text in the information flow.

[0118] Step 903: Perform computer vision processing on the multiple images corresponding to the text, and filter out sensitive images in the multiple images to obtain the multiple candidate images corresponding to the text.

[0119] It should be noted that sensitive images refer to images that are not suitable for display to users. Generally, whether an image is a sensitive image can be measured from three dimensions: the quality, category, and content of the image. That is, in an alternative implementation manner of the embodiments of the present invention, the sensitive images include quality-sensitive images, category-sensitive images, and / or content-sensitive images; the quality-sensitive image is specifically an image whose image quality is lower than a preset image quality, the category-sensitive image is specifically an image whose image category belongs to a preset sensitive image category, and the content-sensitive image is specifically an image whose image content sensitivity is higher than a preset image content sensitivity. By using computer vision technology to process multiple images corresponding to the text, it is possible to determine whether each image is a sensitive image. If the image is a sensitive image, it needs to be filtered from the multiple images corresponding to the text, and the remaining multiple images after filtering the multiple images corresponding to the text are the multiple candidate images corresponding to the text.

[0120] As an example, the quality-sensitive images include images with a resolution lower than a preset image resolution and / or images with a noise higher than a preset image noise, etc.; the category-sensitive images include advertisement images, pornographic images, horror images, and / or violent images, etc.; the content-sensitive images include images containing pornographic, horror, or violent element content and / or images containing disgusting element content, etc.

[0121] Step 904: Perform computer vision processing on the multiple candidate images corresponding to the text to obtain the image semantics of each candidate image.

[0122] It should also be noted that in the embodiments of the present invention, the execution order of step 901 and steps 902 - 904 is not limited. It is possible to execute step 901 first and then steps 902 - 904, or execute steps 902 - 904 first and then step 901, or execute step 901 and steps 902 - 904 simultaneously.

[0123] Step 905: Match the image semantics of each candidate image with the text semantics of the text, and determine at least one candidate image as the target image from the multiple candidate images.

[0124] Step 906: Based on the category and preset display size of the target image, perform computer vision processing on the target image to generate a thumbnail.

[0125] It should be noted that steps 904 - 906 in this embodiment are the same as steps 202 - 204 in the above embodiment. For the specific implementation manner, refer to the above embodiment and will not be elaborated here.

[0126] Through various implementation manners provided by this embodiment, the text semantics of the text in the information flow is obtained by using natural language processing technology; multiple images corresponding to the text in the information flow are obtained for computer vision processing, and sensitive images in the multiple images are filtered to obtain multiple candidate images corresponding to the text; the multiple candidate images corresponding to the text are processed by using computer vision technology to obtain the image semantics of each candidate image; the image semantics of each candidate image is matched with the text semantics of the text, and at least one candidate image is determined as the target image from the multiple candidate images; based on the category of the target image and the preset display size, the target image is processed by using computer vision technology to generate a thumbnail. It can be seen that filtering sensitive images in the multiple images to obtain multiple candidate images can exclude images that are not suitable for display to users and ensure that the candidate images are suitable for display; the target image is determined by obtaining the text semantics of the text and the image semantics of each candidate image, and matching the image semantics of each candidate image with the text semantics of the text. The generated thumbnail has a high matching degree with the text in the information flow. Displaying the thumbnail when displaying the information flow can improve the matching of the thumbnail with the text; and when generating the thumbnail, considering the category of the image and the preset display size, different thumbnail methods are adopted to improve the visual quality of the thumbnail and increase the attraction of the thumbnail to users, thereby improving the user experience of browsing the information flow.

[0127] Exemplary device

[0128] See Figure 10 , which shows the structural schematic diagram of a device for generating a thumbnail based on an information flow in an embodiment of the present invention. In this embodiment, the device may specifically include, for example:

[0129] The first obtaining unit 1001 is configured to perform natural language processing on the text in the information flow to obtain the text semantics of the text;

[0130] The second obtaining unit 1002 is configured to perform computer vision processing on the multiple candidate images corresponding to the text to obtain the image semantics of each candidate image;

[0131] The determining unit 1003 is configured to match the image semantics of each candidate image with the text semantics of the text, and determine at least one candidate image as the target image from the multiple candidate images;

[0132] The generating unit 1004 is configured to perform computer vision processing on the target image based on the category of the target image and the preset display size to generate a thumbnail.

[0133] In an optional implementation manner of an embodiment of the present invention, the determining unit 1003 includes:

[0134] An acquisition subunit, configured to match the image semantics of each candidate image with the text semantics of the text, and obtain the matching degree between the image semantics of each candidate image and the text semantics of the text;

[0135] A first determination subunit, configured to determine the top N candidate images from the multiple candidate images as the target images according to the matching degree from high to low; N is a positive integer less than the total number of the multiple candidate images.

[0136] In an alternative implementation manner of the embodiment of the present invention, the generating unit 1004 is specifically configured to:

[0137] If the category of the target image is a general category, perform computer vision processing on the target image based on the positions of the key elements in the target image and the preset display size to generate the thumbnail; or,

[0138] If the category of the target image is an icon category, perform computer vision processing on the icon area in the target image based on the background color of the target image and the preset display size to generate the thumbnail; or,

[0139] If the category of the target image is a portrait category, perform computer vision processing on the face and neck and shoulder areas in the target image based on the edge color of the target image and the preset display size to generate the thumbnail; or,

[0140] If the category of the target image is a text / chart category, perform computer vision processing on the title area of the text / chart in the target image based on the preset display size to generate the thumbnail.

[0141] In an alternative implementation manner of the embodiment of the present invention, when the generating unit 1004 is specifically configured to, if the category of the target image is a general category, perform computer vision processing on the target image based on the positions of the key elements in the target image and the preset display size to generate the thumbnail, the generating unit 1004 includes:

[0142] A second determination subunit, configured to determine a candidate area in the target image based on the positions of the key elements in the target image;

[0143] A generating subunit, configured to perform computer vision processing on the candidate area based on the preset display size to generate the thumbnail.

[0144] In an alternative implementation manner of the embodiment of the present invention, the generating subunit is specifically configured to:

[0145] When the candidate region meets the preset display size, perform computer vision processing on the candidate region according to the preset display size to generate the thumbnail; or,

[0146] When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is greater than the preset distance, delete the region between the key elements in the candidate region to obtain an updated candidate region; perform computer vision processing on the updated candidate region according to the preset display size to generate the thumbnail; or,

[0147] When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is less than or equal to the preset distance or there is only one key element in the candidate region, perform computer vision processing on the candidate region based on the background color of the target image and the preset display size to generate the thumbnail.

[0148] In an optional implementation manner of the embodiment of the present invention, the device further includes:

[0149] An acquisition unit, configured to acquire multiple images corresponding to the text in the information stream;

[0150] A third obtaining unit, configured to perform computer vision processing on the multiple images corresponding to the text, and filter out sensitive images from the multiple images to obtain multiple candidate images corresponding to the text.

[0151] In an optional implementation manner of the embodiment of the present invention, the sensitive images include quality-sensitive images, category-sensitive images, and / or content-sensitive images; the quality-sensitive images are specifically images with an image quality lower than a preset image quality, the category-sensitive images are specifically images whose image categories belong to a preset sensitive image category, and the content-sensitive images are specifically images with an image content sensitivity higher than a preset image content sensitivity.

[0152] Through various embodiments provided in this embodiment, the text semantics of the text in the information stream are obtained by using natural language processing technology; a plurality of candidate images corresponding to the text are processed by using computer vision technology to obtain the image semantics of each candidate image; the image semantics of each candidate image are matched with the text semantics of the text, and at least one candidate image is determined as the target image from the plurality of candidate images; based on the category of the target image and a preset display size, the target image is processed by using computer vision technology to generate a thumbnail. It can be seen that the target image is determined by obtaining the text semantics of the text and the image semantics of each candidate image, and matching the image semantics of each candidate image with the text semantics of the text. The generated thumbnail has a high matching degree with the text in the information stream. When displaying the information stream, displaying the thumbnail can improve the matching between the thumbnail and the text; and when generating the thumbnail, considering the category of the image and the preset display size, different thumbnail methods are adopted to improve the visual quality of the thumbnail and increase the attraction of the thumbnail to users, thereby improving the user experience of browsing the information stream.

[0153] Figure 11 FIG. 4 is a block diagram of an apparatus 1100 for generating a thumbnail based on an information stream according to an exemplary embodiment. For example, the apparatus 1100 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0154] Referring to Figure 11 , the apparatus 1100 may include one or more of the following components: a processing component 1102, a memory 1104, a power component 1106, a multimedia component 1108, an audio component 1110, an input / output (I / O) interface 1112, a sensor component 1114, and a communication component 1116.

[0155] The processing component 1102 generally controls the overall operation of the apparatus 1100, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 1102 may include one or more processors 1120 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1102 may include one or more modules to facilitate the interaction between the processing component 1102 and other components. For example, the processing component 1102 may include a multimedia module to facilitate the interaction between the multimedia component 1108 and the processing component 1102.

[0156] The memory 1104 is configured to store various types of data to support the operation of the device 1100. Examples of such data include instructions for any application or method operating on the device 1100, contact data, phone book data, messages, pictures, videos, and the like. The memory 1104 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0157] The power supply component 1106 provides power to various components of the device 1100. The power supply component 1106 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 1100.

[0158] The multimedia component 1108 includes a screen that provides an output interface between the device 1100 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1108 includes a front camera and / or a rear camera. When the device 1100 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0159] The audio component 1110 is configured to output and / or input audio signals. For example, the audio component 1110 includes a microphone (MIC) that is configured to receive external audio signals when the device 1100 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1104 or transmitted via the communication component 1116. In some embodiments, the audio component 1110 further includes a speaker for outputting audio signals.

[0160] The I / O interface 1112 provides an interface between the processing component 1102 and the peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, and the like. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0161] The sensor assembly 1114 includes one or more sensors for providing a status assessment of various aspects of the device 1100. For example, the sensor assembly 1114 can detect the on / off state of the device 1100, the relative positioning of components, such as the display and keypad of the device 1100. The sensor assembly 1114 can also detect a change in the position of the device 1100 or a component of the device 1100, the presence or absence of user contact with the device 1100, the orientation or acceleration / deceleration of the device 1100, and a change in the temperature of the device 1100. The sensor assembly 1114 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1114 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1114 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0162] The communication component 1116 is configured to facilitate communication between the device 1100 and other devices in a wired or wireless manner. The device 1100 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1116 receives a broadcast signal or broadcast matching information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1116 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0163] In an exemplary embodiment, the device 1100 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0164] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1104 including instructions, and the above instructions can be executed by the processor 1120 of the device 1100 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0165] A non - transitory computer - readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute a method for generating a thumbnail based on information flow. The method includes:

[0166] Performing natural language processing on the text in the information flow to obtain the text semantics of the text;

[0167] Performing computer vision processing on multiple candidate images corresponding to the text to obtain the categories and positions of key elements in each candidate image;

[0168] Matching the categories of key elements in each candidate image and the text semantics of the text, and determining at least one candidate image as the target image from the multiple candidate images;

[0169] Based on the category of the target image and a preset display size, performing computer vision processing on the target image to generate a thumbnail.

[0170] Figure 12 It is a schematic structural diagram of a server in an embodiment of the present invention. The server 1200 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1222 (for example, one or more processors) and a memory 1232, and one or more storage media 1230 (for example, one or more mass storage devices) for storing application programs 1242 or data 1244. Among them, the memory 1232 and the storage media 1230 can be transient storage or persistent storage. The program stored in the storage media 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1222 can be configured to communicate with the storage media 1230 and execute a series of instruction operations in the storage media 1230 on the server 1200.

[0171] The server 1200 may further include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, one or more keyboards 1256, and / or one or more operating systems 1241, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0172] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the corresponding parts, reference can be made to the description in the method part.

[0173] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0174] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including an..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0175] The above is only the preferred embodiments of the embodiments of the present invention, and does not impose any form of limitation on the embodiments of the present invention. Although the embodiments of the present invention have been disclosed above in the preferred embodiments, they are not intended to limit the embodiments of the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solutions of the embodiments of the present invention, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the embodiments of the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the embodiments of the present invention without departing from the content of the technical solutions of the embodiments of the present invention still fall within the scope of the protection of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a thumbnail based on information flow, characterized in that Including: Performing natural language processing on the text in the information flow to obtain the text semantics of the text; Performing computer vision processing on multiple candidate images corresponding to the text to obtain the image semantics of each candidate image; Matching the image semantics of each candidate image with the text semantics of the text, and determining at least one candidate image as the target image from the multiple candidate images; Performing computer vision processing on the target image based on the category of the target image and a preset display size to generate a thumbnail; Among them, the performing computer vision processing on the target image based on the category of the target image and a preset display size to generate a thumbnail specifically includes: If the category of the target image is a general category, performing computer vision processing on the target image based on the positions of key elements in the target image and the preset display size to generate the thumbnail; or, If the category of the target image is an icon category, performing computer vision processing on the icon area in the target image based on the background color of the target image and the preset display size to generate the thumbnail; or, If the category of the target image is a portrait category, performing computer vision processing on the face and neck area in the target image based on the edge color of the target image and the preset display size to generate the thumbnail; or, If the category of the target image is a text / chart category, performing computer vision processing on the title area of the text / chart in the target image based on the preset display size to generate the thumbnail.

2. The method according to claim 1, characterized in that, The matching the image semantics of each candidate image with the text semantics of the text, and determining at least one candidate image as the target image from the multiple candidate images includes: Matching the image semantics of each candidate image with the text semantics of the text to obtain the matching degree between the image semantics of each candidate image and the text semantics of the text; Determining the top N candidate images as the target images from the multiple candidate images in descending order of the matching degree; N is a positive integer less than the total number of the multiple candidate images.

3. The method according to claim 1, characterized in that, The performing computer vision processing on the target image based on the positions of key elements in the target image and the preset display size to generate the thumbnail includes: Determining a candidate area in the target image based on the positions of key elements in the target image; Performing computer vision processing on the candidate area based on the preset display size to generate the thumbnail.

4. The method according to claim 3, wherein The performing computer vision processing on the candidate area based on the preset display size to generate the thumbnail specifically includes: When the candidate area meets the preset display size, performing computer vision processing on the candidate area according to the preset display size to generate the thumbnail; Or, When the candidate area does not meet the preset display size, if the distance between key elements in the candidate area is greater than a preset distance, deleting the area between key elements in the candidate area to obtain an updated candidate area; Performing computer vision processing on the updated candidate area according to the preset display size to generate the thumbnail; Or, When the candidate region does not meet the preset display size, if the distance between key elements in the candidate region is less than or equal to the preset distance or there is only one key element in the candidate region, computer vision processing is performed on the candidate region based on the background color of the target image and the preset display size to generate the thumbnail.

5. The method according to claim 1, wherein Before performing computer vision processing on the multiple candidate images corresponding to the text, it further includes: Obtaining multiple images corresponding to the text in the information stream; Performing computer vision processing on the multiple images corresponding to the text, and filtering out sensitive images from the multiple images to obtain multiple candidate images corresponding to the text.

6. The method according to claim 5, wherein The sensitive images include quality-sensitive images, category-sensitive images, and / or content-sensitive images; the quality-sensitive image is specifically an image with an image quality lower than the preset image quality, the category-sensitive image is specifically an image whose image category belongs to the preset sensitive image category, and the content-sensitive image is specifically an image with an image content sensitivity higher than the preset image content sensitivity.

7. An apparatus for generating a thumbnail based on information flow, characterized in that It includes: The first obtaining unit is used to perform natural language processing on the text in the information stream to obtain the text semantics of the text; The second obtaining unit is used to perform computer vision processing on the multiple candidate images corresponding to the text to obtain the image semantics of each candidate image; The determining unit is used to match the image semantics of each candidate image with the text semantics of the text, and determine at least one candidate image as the target image from the multiple candidate images; The generating unit is used to perform computer vision processing on the target image based on the category of the target image and the preset display size to generate a thumbnail; Among them, the generating unit is specifically used for: If the category of the target image is a general category, performing computer vision processing on the target image based on the positions of the key elements in the target image and the preset display size to generate the thumbnail; or, If the category of the target image is an icon category, performing computer vision processing on the icon area in the target image based on the background color of the target image and the preset display size to generate the thumbnail; or, If the category of the target image is a portrait category, performing computer vision processing on the face and neck area in the target image based on the edge color of the target image and the preset display size to generate the thumbnail; or, If the category of the target image is a text / chart category, performing computer vision processing on the title area of the text / chart in the target image based on the preset display size to generate the thumbnail.

8. The device according to claim 7, characterized in that, The determining unit includes: The obtaining subunit is used to match the image semantics of each candidate image with the text semantics of the text to obtain the matching degree between the image semantics of each candidate image and the text semantics of the text; The first determining subunit is used to determine the first N candidate images as the target images from the multiple candidate images in descending order of the matching degree; N is a positive integer less than the total number of the multiple candidate images.

9. The device according to claim 7, characterized in that, When the generation unit is specifically configured to perform computer vision processing on the target image to generate the thumbnail based on the positions of the key elements in the target image and the preset display size if the category of the target image is a general category, the generation unit includes: A second determination subunit, configured to determine a candidate region in the target image based on the positions of the key elements in the target image; A generation subunit, configured to perform computer vision processing on the candidate region based on the preset display size to generate the thumbnail.

10. The device according to claim 9, characterized in that, The generation subunit is specifically configured to: When the candidate region meets the preset display size, perform computer vision processing on the candidate region according to the preset display size to generate the thumbnail; Or, When the candidate region does not meet the preset display size, if the distance between the key elements in the candidate region is greater than a preset distance, delete the region between the key elements in the candidate region to obtain an updated candidate region; Perform computer vision processing on the updated candidate region according to the preset display size to generate the thumbnail; Or, When the candidate region does not meet the preset display size, if the distance between the key elements in the candidate region is less than or equal to the preset distance or there is only one key element in the candidate region, perform computer vision processing on the candidate region based on the background color of the target image and the preset display size to generate the thumbnail.

11. The device according to claim 7, characterized in that, The apparatus further includes: An acquisition unit, configured to acquire a plurality of images corresponding to the text in the information stream; A third obtaining unit, configured to perform computer vision processing on the plurality of images corresponding to the text, and filter out sensitive images from the plurality of images to obtain a plurality of candidate images corresponding to the text.

12. The device according to claim 11, wherein The sensitive images include quality-sensitive images, category-sensitive images, and / or content-sensitive images; the quality-sensitive image is specifically an image with an image quality lower than a preset image quality, the category-sensitive image is specifically an image whose image category belongs to a preset sensitive image category, and the content-sensitive image is specifically an image with an image content sensitivity higher than a preset image content sensitivity.

13. An apparatus for generating a thumbnail based on an information stream, characterized in that, Including a memory, and more than one program, where the more than one program is stored in the memory and is configured to be executed by more than one processor. The more than one program includes instructions for performing the following operations: Perform natural language processing on the text in the information stream to obtain the text semantics of the text; Perform computer vision processing on the plurality of candidate images corresponding to the text to obtain the image semantics of each candidate image; Match the image semantics of each candidate image with the text semantics of the text, and determine at least one candidate image as the target image from the plurality of candidate images; Perform computer vision processing on the target image based on the category of the target image and the preset display size to generate a thumbnail; Wherein, performing computer vision processing on the target image based on the category of the target image and the preset display size to generate a thumbnail is specifically: If the category of the target image is a general category, computer vision processing is performed on the target image based on the positions of the key elements in the target image and the preset display size to generate the thumbnail; or, If the category of the target image is an icon category, computer vision processing is performed on the icon area in the target image based on the background color of the target image and the preset display size to generate the thumbnail; or, If the category of the target image is a portrait category, computer vision processing is performed on the face and neck area in the target image based on the edge color of the target image and the preset display size to generate the thumbnail; or, If the category of the target image is a text / chart category, computer vision processing is performed on the title area of the text / chart in the target image based on the preset display size to generate the thumbnail.

14. A machine-readable medium having instructions stored thereon that, when executed by one or more processors, cause the apparatus to perform the method for generating a thumbnail based on an information stream as recited in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Content based sensitive web page identification method

    CN101055621A

  • Method for generating thumbnail image and electronic device thereof

    US20140104477A1