Image-guided video thumbnail generation for e-commerce applications
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EBAY INC
- Filing Date
- 2022-12-07
- Publication Date
- 2026-08-07
AI Technical Summary
手动生成缩略图对卖家来说是一项耗时的任务
[0006]根据本公开,通过使用在线购物市场中的机器学习模型基于视频的内容自动生成缩略图来解决上述和其他问题。本公开涉及在电子商务系统中从视频自动生成缩略图图像。具体地,所公开的技术确定用于交易的物品的图像与描述该物品的视频的一个或多个视频帧(即,帧图像)之间的加权视觉相似度。
Smart Images

Figure CN116437122B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims priority to U.S. Patent Application No. 17 / 545,497, filed December 8, 2021, entitled “IMAGE GUIDED VIDEO THUMBNAIL GENERATION FOR E-COMMERCE APPLICATIONS”, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] This application relates to e-commerce, and more specifically, to a method and system for automatically generating thumbnail images. Background Technology
[0004] In e-commerce platforms, using videos to describe items in an item listing helps buyers purposefully understand the items by watching them in motion, in actual use. When not playing videos, some e-commerce platforms display text titles and / or thumbnails (e.g., images) representing the video. Some e-commerce systems prompt sellers to create and upload thumbnails representing the video along with it. Other e-commerce systems automatically generate thumbnails using thumbnails from predetermined video frames related to the item. Still others automatically select and extract video frames at predetermined locations within the video (e.g., the first video frame, a video frame from the longest segment of the video) as thumbnail images. Some systems evaluate the visual quality of video frames in a video and select video frames with quality levels exceeding a threshold as candidate thumbnails. In practice, viewers often do not find automatically generated thumbnails sufficient to represent the item associated with the video. While automatically generated thumbnails can represent the video, the video frames extracted from the video do not necessarily represent the item a potential buyer would expect to see. Using accurate thumbnails representing items in e-commerce systems improves the efficiency and effectiveness of online shopping marketplaces because more sellers and buyers use e-commerce systems with confidence. Manually generating thumbnails is a time-consuming task for sellers. Therefore, developing a technology that can better meet their needs while minimizing compromises would be desirable.
[0005] It is with regard to these and other general considerations that the various aspects disclosed herein have been made. While relatively specific issues may be discussed, it should be understood that the examples are not limited to addressing specific problems identified in the context of this disclosure or elsewhere. Summary of the Invention
[0006] According to this disclosure, the aforementioned and other problems are addressed by automatically generating thumbnails based on video content using machine learning models in online shopping marketplaces. This disclosure relates to automatically generating thumbnail images from videos in e-commerce systems. Specifically, the disclosed technique determines a weighted visual similarity between an image of an item for sale and one or more video frames (i.e., frame images) of a video describing that item.
[0007] This disclosure relates to systems and methods for automatically generating thumbnails, at least according to the examples provided in the following sections. Specifically, this disclosure relates to a computer-implemented method for automatically generating thumbnail images representing videos associated with a list of items in an electronic marketplace. The method includes: retrieving a video associated with a list of items, wherein the video comprises a plurality of ordered frame images; retrieving item images associated with the list of items; determining, from the plurality of ordered frame images, at least a first score for a first frame image and a second score for a second frame image, wherein the scores are based at least on the similarity between the item images and the frame images. The method further includes: determining a first weight associated with the first frame image and a second weight associated with the second frame image, wherein the first weight is based at least on the ordered position of the first frame image; updating the first score based on the first weight and updating the second score based on the second weight; selecting a thumbnail image from either the first frame image or the second frame image, wherein the selection is based on the updated first score and the updated second score; and automatically adding the thumbnail image to the video.
[0008] This disclosure also relates to a system for automatically generating thumbnail images representing videos associated with items. The system includes: a processor; and a memory storing computer-executable instructions that, when executed by the processor, cause the system to: retrieve a video associated with a list of items, wherein the video comprises a plurality of ordered frame images; retrieve item images associated with the list of items; determine, from the plurality of ordered frame images, at least a first score for a first frame image and a second score for a second frame image, wherein the scores are based at least on the similarity between the item images and the frame images; determine a first weight associated with the first frame image and a second weight associated with the second frame image, wherein the first weight is based at least on the ordered position of the first frame image; update the first score based on the first weight and update the second score based on the second weight; select a thumbnail image from either the first frame image or the second frame image, wherein the selection is based on the updated first score and the updated second score; and automatically add the thumbnail image to the video.
[0009] This disclosure also relates to a computer-implemented method for automatically generating thumbnail images representing videos associated with a list of items in an electronic marketplace. The method includes: retrieving a video associated with an item, wherein the video comprises a plurality of frame images; retrieving a set of item images, wherein the set of item images comprises item images associated with the item, and wherein the position of an item image corresponds to the position of the item image relative to other item images published on a list of items in the electronic marketplace; extracting video frames from the video based on a sampling rule; determining a score for the video frame, wherein the score represents the degree of similarity between the video frame and each item image in the set of item images; determining a weighted score for the video frame based on the score, wherein the weighted score includes weights associated with the positions of the item images in the set of item images; selecting the content of the video frame as a thumbnail image of the video based on the weighted score; and publishing the thumbnail image to represent the video.
[0010] The summary is provided to introduce, in a simplified form, the selection of concepts further described below in the detailed description. The summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be set forth in part in the description which follows, and will be apparent in part from that description, or may be learned by practice of this disclosure. Attached Figure Description
[0011] The following figures illustrate examples of non-restrictive and non-exhaustive properties.
[0012] Figure 1 An overview of an example system for automatically generating thumbnail images of items according to various aspects of this disclosure is shown.
[0013] Figure 2 Examples of determining similarity scores according to various aspects of this disclosure are shown.
[0014] Figure 3 An example of a data structure for similarity score data of videos according to various aspects of this disclosure is shown.
[0015] Figure 4 An example of a data structure for weighted visual similarity scores of videos according to various aspects of this disclosure is shown.
[0016] Figure 5 Examples of methods for automatically generating thumbnails according to various aspects of this disclosure are shown.
[0017] Figure 6 This is a block diagram illustrating example physical components of a computing device that can be used to practice various aspects of this disclosure. Detailed Implementation
[0018] Various aspects of this disclosure are described more fully below with reference to the accompanying drawings, which are taken from a part of this disclosure and illustrate specific example aspects. However, different aspects of this disclosure may be implemented in many different ways and should not be construed as limited to the aspects set forth herein; rather, these aspects are provided so that this disclosure will be comprehensive and complete, and will fully convey the scope of these aspects to those skilled in the art. The aspects may be practiced as methods, systems, or apparatuses. Thus, aspects may take the form of hardware implementations, entirely software implementations, or implementations combining software and hardware aspects. Therefore, the following detailed description should not be considered limiting.
[0019] Online shopping systems (such as e-commerce systems and online auction systems) rely on sellers to prepare content describing items for posting on online shopping sites. This content typically includes a text description of the item, one or more images depicting the item, a video describing the item, and a thumbnail that briefly describes the video. Some traditional e-commerce systems automatically generate video thumbnails by selecting video frames at predetermined locations within the video.
[0020] For example, a seller can upload a video describing a pair of shoes for sale. The video could begin with a few seconds of title screen, followed by an overview image of the shoes, a series of close-up images of different parts of the shoes (e.g., the sole), a textual description of the upper's features, an image of the shoes in a landscape setting, and a list of the shoes' features before the video ends. Sellers can also upload a collection of images associated with the shoes.
[0021] Some traditional systems select a video frame at a predetermined location in the video as the thumbnail. Examples of predetermined locations may include the first video frame or segment, a specific video frame that enters the video at a predetermined time (e.g., three seconds, which typically includes the video's title screen), or the last video frame.
[0022] As discussed in more detail below, this disclosure relates to automatically generating thumbnails of videos describing items. Specifically, the disclosed technique uses a weighted similarity score between images used in an item list and sample video frames from the video. A similarity score generator can use a machine learning model (i.e., a similarity model) to determine the degree of similarity between images in an image set and each sampled video frame. A weighted score generator generates a weighted sum of the similarity scores of the sampled video frames compared to the individual images in the image set. Furthermore, the disclosed technique weights the frequency of occurrence of the sample video frames in the video.
[0023] Figure 1An overview of an example system 100 for automatically generating thumbnails of videos according to various aspects of this disclosure is shown. System 100 represents a system for determining similarity scores of individual sampled video frames of a video using a similarity model. In some aspects, the similarity score indicates the level of similarity between the content of the sampled video frame and an image in a collection of images published in an item listing on an online shopping site.
[0024] System 100 includes a client device 102, an application server 110, an online shopping server 120, a network 150, and an image data server 160. The client device 102 communicates with the application server 110, which includes one or more sets of instructions executed on the client device 102 as an application. The application server 110 includes an online shopping application 112 (i.e., an application). The one or more sets of instructions in the application server 110 can provide an interactive user interface (not shown) through an interactive interface 104.
[0025] Online shopping server 120 includes a video receiver 124, an item image receiver, a visual similarity determiner 128, a similarity model 130, a weighted score determiner 132, a weighting rule 134, a thumbnail determiner 136, a shopping store server 140, and an item database 142. Client device 102 connects to application server 110 via network 150 to execute an application including user interaction through interactive interface 104. Application server 110 interacts with client device 102 and online shopping server 120 via network 150 to perform online shopping as a seller or buyer of items.
[0026] Client device 102 is a general-purpose computer device that provides user input capabilities (e.g., online shopping via network 150 through interactive interface 104). In some aspects, client device 102 may optionally receive user input from sellers of items. Sellers upload information about items being sold in an online shopping marketplace (e.g., an e-marketplace). Information about the items includes image data of the items, a brief description of the items, price information, quantity information, etc. For example, interactive interface 104 may present a graphical user interface associated with a web browser. In some aspects, client device 102 may communicate with application server 110 via network 150.
[0027] Application server 110 is a server that enables sellers (who can post items for sale) and buyers (who purchase items) to interactively use system 100 on client device 102. Application server 110 may include applications, including online shopping application 112. Online shopping application 112 can provide the presentation of items for users to purchase.
[0028] In various aspects, the online shopping application 112 can connect to the video receiver 124 of the thumbnail generator 122 in the online shopping server 120 to publish information about items for sale on an online shopping site (not shown). In some aspects, the video receiver 124 can prompt the user to upload item images in a sequential order, where the item image that best represents the item is uploaded before other item images. Information about the item may include its name, a brief description of the item, its quantity, price, and one or more images describing the item, as well as a video describing the item. Additionally or alternatively, information about the item may include category information. For example, the item may include a pair of shoes. The one or more images may include photographs of the shoes from different angles, a table listing product features, close-ups of product information with product codes, and the item's serial number. When the online shopping server 120 successfully receives information about the item, the online shopping application 112 can receive confirmation from the online shopping server 120.
[0029] Online shopping server 120 refers to the application / system used to generate thumbnails for videos and publish those thumbnails on online shopping sites. Specifically, thumbnail generator 122 uses a similarity model to match the content of sampled video frames with images associated with the item. The thumbnail generator determines the video thumbnail from a set of video frames based on a weighted similarity score.
[0030] The video receiver 124 receives videos of items from the online shopping application 112 used by the seller via an interactive interface 104 on the client device 102. Alternatively, the video receiver 124 may receive the video from the item database 142 when the seller has already uploaded the video and stored it on the online shopping server 120.
[0031] The item image receiver 126 receives a set of images associated with an item from the online shopping application 112 used by the seller via an interactive interface 104 on the client device 102. Alternatively or additionally, when the seller has already uploaded the image set and stored it on the online shopping server 120, the item image receiver 126 receives the item image set from the item database 142.
[0032] The visual similarity determiner 128 extracts a set of video frames from the video by sampling. For example, the visual similarity determiner 128 extracts video frames at predetermined time intervals. In some other examples, the visual similarity determiner 128 extracts video frames that follow changes in the video content exceeding a predetermined threshold. The visual similarity determiner 128 generates pairs between each sampled video frame and each image in the image set. For example, when there are ten sampled video frames and five images in the image set, the visual similarity determiner 128 generates fifty pairs of video frames and images.
[0033] The visual similarity determiner 128 can use the similarity model 130 to determine the similarity score for each pair. In various respects, the similarity model 130 can be a machine learning model used to predict the similarity score. Similarity indicates the probability that the content of a video frame is similar to that of an image. For example, the similarity score can be represented in a range from zero to one hundred, where zero indicates difference and one hundred indicates similarity.
[0034] The weighted score determiner 132 determines a weighted similarity score (i.e., a weighted score) for a pair of sampled video frames and images based on their similarity scores. In each respect, each image may correspond to a different weight. In each respect, the weighted score determiner 132 determines the weighted similarity score based on weighting rule 134. For example, an image of an item that appears before other images may correspond to more weight. In an item description page on an online shopping site, an item image that appears earlier in the list of item images is more likely to represent the item because these images are more likely to attract the viewer's attention than other images that appear later in the list. In another example, the content of a video frame that appears more frequently than the content of other video frames is more likely to represent the item, based on the assumption that images representing an item will appear more frequently than other images in the video. In yet another example, video frames closer to the beginning of the video may have a higher weight than video frames closer to the end of the video. In each respect, the weighted score determiner 132 determines the weighted similarity score based on a weighted sum of the similarity scores associated with the sampled video frames. The weighted score determiner 132 can also incorporate weights associated with the number of times the content of the sampled video frame appears in other video frames in the video (i.e., the number of occurrences). For example, the content of a video frame that appears more frequently than the content of other video frames in the video can carry more weight. Additionally or alternatively, the weights can be based on the degree to which the frame image indicates a specific characteristic of an object. Characteristics can include the shape, color, and / or visual content of the frame image.
[0035] Thumbnail determiner 136 determines the thumbnail of the video. In all aspects, thumbnail determiner 136 selects the sampled video frame with the highest weighted similarity score from the sampled video frames as the video thumbnail. Thumbnail determiner 136 stores the video thumbnail in item database 142. Shopping store server 140 can publish the video thumbnail on an online shopping site.
[0036] As will be understood, Figure 1 The various methods, devices, applications, features, etc., described are not intended to limit System 100 to being performed by the specific applications and features described. Therefore, additional controller configurations may be used to practice the methods and systems described herein, and / or the features and applications described may be excluded without departing from the methods and systems disclosed herein.
[0037] Figure 2 An example of a visual similarity determiner 202 according to various aspects of this disclosure is shown. Figure 2 In this embodiment, system 200 illustrates a visual similarity determiner 202 using a similarity model 214. The visual similarity determiner 202 includes a video frame sampler 204 and a similarity matrix generator 206. In each respect, the visual similarity determiner 202 receives a set of object images 210 and a video 212 as input and generates a similarity score matrix 216.
[0038] Video frame sampler 204 extracts sample video frames from the video frames of video 212 using a predetermined sampling rule 208. For example, the predetermined sampling rule 208 may include sampling video frames according to a predetermined number of video frames (e.g., every ten video frames) and / or time intervals. Based on the sampling, video frame sampler 204 generates a set of sampled video frames.
[0039] Similarity matrix generator 206 generates a similarity score matrix 216 by generating similarity scores for pairs of video frames and images. The similarity score matrix includes the similarity score between each sampled video frame in the sampled video frames and each item image in the set of item images 210. In each respect, similarity matrix generator 206 uses a similarity model 214 to determine the similarity score for each pair of video frames and item images. Similarity model 214 can be a machine learning model trained to predict the similarity between two sets of image data containing information associated with items used for transactions on an online shopping site. For example, the similarity score can be represented as a value between zero and one hundred, where zero indicates that the two sets of image data are different, and one hundred indicates that the two sets of image data are the same.
[0040] In various aspects, the similarity score matrix 216 may include sampled video frames as rows and item images as columns. Therefore, the similarity score matrix 216 includes similarity scores associated with pairs of video frames and item images. In various aspects, a weighted score determiner (e.g., such as...) Figure 1 The weighted score determiner 132 shown generates a weighted similarity score matrix by determining the weighted sum of the similarity scores (e.g., as shown in the figure). Figure 3 (Example of weighted similarity score matrix 302 shown).
[0041] Figure 3 An exemplary data structure for a weighted similarity score matrix according to various aspects of this disclosure is shown. Figure 3An example 300 depicting a weighted similarity score matrix 302 is shown. The weighted similarity score matrix 302 includes weights 322 associated with individual item images and scores associated with individual sampled video frames 304 in the rows. The weighted similarity score matrix 302 also includes a first image 308 (i.e., the first item image), a second image 310 (i.e., the second item image), a third image 312 (i.e., the third item image), a fourth image 314 (i.e., the fourth item image), and a fifth image 316 (i.e., the fifth item image). The weighted similarity score matrix 302 also includes the number of times the video frame appears in the video 318 and weighted similarity scores 320 in the respective columns. The exemplary data is based on a frame sampling frequency 324 of every ten video frames. The weighted similarity score matrix 302 can specify a threshold (e.g., 10,000 points) for determining the sampled video frames as thumbnails. In various respects, the threshold can be a predetermined threshold.
[0042] In all respects, this disclosed technique uses a higher multiplier as the weight value for item images that appear earlier in the list of item images. On item description pages of online shopping sites, item images that appear earlier in the list of item images tend to be more representative of the item than those that appear later in the list. For example, the first item image may depict a panoramic photograph of an image with a clear background. The second item image may depict an image with a landscape background. Item images near the end of the list may depict textual descriptions of the item's characteristics (e.g., such as...). Figure 4 (The list of item images 402 shown). Therefore, the row labeled weight 322 indicates x10 for the first image 308 (i.e., the weighted similarity score is calculated by multiplying the similarity score by ten), x5 for the second image 310 (i.e., multiplied by five), x3 for the third image 312, x2 for the fourth image 314, and x1 for the fifth image 316 (i.e., unweighted). For example, when the similarity score between the video frame and the first item image is 70, the weighted similarity score is 70 x 10 = 700. The occurrence count 318 corresponds to the number of times the content of the sampled video frame appears in the video. The weighted similarity score 320 indicates the weighted sum of the similarity scores associated with the sampled video frames.
[0043] For example, the first sampled video frame indicates a similarity score of zero for the first image 308, the second image 310, the third image 312, and the fourth image 314. The first sampled video frame has a similarity score of 30 with the fifth image 316. The first sampled video frame appears three times in the video. The second sampled video frame indicates the following similarity scores with the corresponding five images: 70, 100 (i.e., the second sampled video frame is the same as the second image 310), 60, 10, 0. The second sampled image appears ten times in the video. The third sampled video frame indicates the following similarity scores with the corresponding five images: 100 (i.e., the third sampled video frame is the same as the first image 308), 70, 80, 20, 0. The third sampled image appears seven times in the video. The 106th sampled video frame indicates the following similarity scores with the corresponding five images: 0, 10, 0, 0, 100 (i.e., the 106th sampled video frame is the same as the fifth image 316). The third sampled image appears ten times in the video.
[0044] In all aspects, the weighted similarity score 320 of the sampled video frames is a weighted sum multiplied by the number of times they appear in the video. For example, the weighted similarity score 320 of the first sampled video frame is (0+0+0+0+10)x3 = 90. The weighted similarity score of the second sampled video frame is (700+500+180+20+0)x10 = 14000. The weighted similarity score of the third sampled video frame is (1000+350+240+40+0)x7 = 11410. The weighted similarity score of the 106th sampled video frame is (0+50+0+0+100)x10 = 1500. The weighted similarity score of the second sampled video frame is greater than that of the other sampled video frames. Therefore, the disclosed technique allows the second sampled video frame to be selected as the thumbnail of the video.
[0045] Figure 4 Examples of lists of item images and sequences of video frames are shown in accordance with various aspects of this disclosure. Figure 4 An example 400 is shown, comprising an ordered list of item images 402, video frames 420 in a time series, and thumbnails automatically generated using the disclosed techniques.
[0046] As an example, the ordered list of item images 402 is arranged in order of prominent representation of the items. The first item image 410 (an overview photograph of the shoes) is likely the most representative of shoes as an item for sale on an online shopping site. The ordered list of item images 402 also includes, in descending order, the second item image 412 (an overview photograph of the shoes with a background landscape), the third item image 414 (a vertical photograph of the shoes), the fourth item image 416 (the sole of the shoe), and the fifth item image 418 (a photograph indicating a characteristic description of the shoe). In all respects, the ordered list of item images 402 corresponds to five item images (i.e., the first image 308, the second image 310, the third image 312, the fourth image 314, and the fifth image 316), as... Figure 3 As shown.
[0047] Video frame 420 includes an ordered list of sampled video frames from a video sampler (e.g., such as...). Figure 2 The video frame sampler 204 of the visual similarity determiner 202 shown samples from the video of the item (i.e., the shoes). In all respects, as indicated by the directional timeline 450, there is an ordered list in the time sequence of the video frames in the video. The first sampled video frame 422 indicates the title scene of the video, which displays the title of the video (i.e., "INTRODUCING NEW SHOES"). The first sampled video frame 422 has no corresponding item image in item image 402. The second sampled video frame 424 indicates an overview of the shoes with a background landscape. The content of the second sampled video frame 424 is the same as the second item image 412. The third sampled video frame 426 indicates an overview of the shoes with a clear background. The content of the third sampled video frame 426 is the same as the first item image 410. The fourth sampled video frame 428 represents a combination of a vertical photograph of the shoe and the sole. The content of the fourth sampled video frame 428 combines the third item image 414 and the fourth item image 416 from the sampled video frames. The first hundred and six sampled video frame 430 indicates a photograph describing the features of the shoe. The content of the 106th sampled video frame 430 is identical to that of the fifth item image 418. In all respects, the first sampled video frame 422, the second sampled video frame 424, and the third sampled video frame 426 correspond to the corresponding sampled video frames "first," "second," and "third," as shown below. Figure 3 As shown.
[0048] Thumbnail 432 indicates a thumbnail automatically generated by the thumbnail generator based on the weighted similarity scores of each video frame determined by the thumbnail determiner. In all respects, this thumbnail generator corresponds to thumbnail generator 122. The thumbnail determiner corresponds to, as... Figure 1 The thumbnail determiner 136 is shown. Due to the highest weighted similarity score (i.e., as shown) Figure 3With a weighted similarity score of 320 (shown as 14,000), the thumbnail determiner has determined the second sampled video frame 424 as the video thumbnail 432.
[0049] Figure 5 This is an example of a method for automatically generating thumbnails for videos according to various aspects of this disclosure. The general order of operations in method 500 is as follows: Figure 5 As shown in the diagram. Typically, method 500 begins with start operation 502 and ends with end operation 520. Method 500 may include more or fewer steps, or may be combined with... Figure 5 The different arrangements of the steps shown illustrate this. Method 500 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 500 can be executed by gates or circuits associated with a processor, ASIC, FPGA, SOC, or other hardware device. References will be incorporated herein by reference. Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 6 Method 500 is described by describing the system, components, devices, modules, software, data structures, data characteristic representations, signaling diagrams, and methods.
[0050] After initiating operation 502, method 500 begins with receiving operation 504, which receives a video describing the item from an item list (e.g., an item's description page) on an online shopping site. In some aspects, sellers use an online shopping application as sellers and upload information about the item to an online shopping server. This information may include the item's name, a brief description, the item's category, one or more item images representing the item, and a video describing the item. For example, the video may include... Figure 4 The video frame 420 shown, as well as other video frames not sampled in this disclosure.
[0051] The receiving operation 506 receives an ordered list of item images. As described above, the seller uses an online shopping application as a seller and uploads an ordered list of item images. In various respects, the ordered list can be sorted in descending order of the degree to which the item images represent the items.
[0052] Extraction operation 508 extracts multiple video frames from the video using sampling rules. For example, extraction operation 508 can extract frames based on sampling rules (such as...). Figure 2 The predetermined sampling rule 208 shows that video frames are extracted every ten video frames.
[0053] Operation 510 determines the similarity score of the sampled video frame. The similarity score indicates the degree of similarity between two image data sets. Specifically, the similarity score of the sampled video frame indicates the degree of similarity between the sampled video frame and an ordered list of item images. For example, the degree of similarity can be expressed by setting zero as different and one hundred as the same.
[0054] Operation 512 determines the weighted similarity score for each sampled video frame. For example, the weighted similarity score can use multiple weights as multipliers for the similarity score. These multiple weights may include, but are not limited to, different weight values corresponding to the object image (e.g., such as...). Figure 3 The weights shown are 322) and the number of times each item image appears.
[0055] Ranking operation 514 ranks each sampled video frame based on a weighted similarity score. In all respects, sampled video frames with higher weighted similarity values can rank higher than other sampled video frames with lower weighted similarity values.
[0056] Selection operation 516 selects a sampled video frame as the video thumbnail. Specifically, selection operation 516 selects the highest-ranking sampled video frame among other sampled video frames. In various aspects, selection operation 516 can also generate a thumbnail based on the content of the selected video frame according to the thumbnail specification of the video by resizing and / or cropping the content of the selected video frame. In various aspects, the thumbnail may be identical to an item image in the item image set. Additionally or alternatively, the thumbnail may indicate a degree of similarity to the item image, which is greater than a predetermined threshold.
[0057] Storage operation 518 stores the thumbnail as an attribute of an item in an item database for e-commerce transactions at an online shopping site. The stored thumbnail can be published on the online shopping site as an image representing a video. The thumbnail is based on one of the sampled video frames from the video. End operation 520 follows the storage operation.
[0058] Figure 6 A simplified block diagram of an apparatus that can be used to practice various aspects of this disclosure is shown. One or more of the embodiments described herein can be implemented in operating environment 600. This is only one example of a suitable computing environment and is not intended to imply any limitation on the scope of functionality or use. Other well-known computing systems, environments, and / or configurations that may be suitable for use include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumer electronic devices such as smartphones, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the aforementioned systems or devices, etc.
[0059] In its most basic configuration, the operating environment 600 typically includes at least one processing unit 602 and memory 604. Depending on the exact configuration and type of computing device, the memory 604 (for instructions that automatically generate thumbnails as described herein) can be volatile (e.g., RAM), non-volatile (e.g., ROM, flash memory, etc.), or some combination of both. This most basic configuration in... Figure 6 The image is shown by dashed line 606. Furthermore, the operating environment 600 may also include storage devices (removable storage device 608, and / or non-removable storage device 610), including but not limited to disks, optical discs, or magnetic tapes. Similarly, the operating environment 600 may also have input devices 614 such as a keyboard, mouse, pen, voice input, onboard sensors, etc., and / or output devices 616 such as a display, speaker, printer, motor, etc. The environment may also include one or more communication connections 612, such as LAN, WAN, near-field communication networks, point-to-point, etc.
[0060] Operating environment 600 typically includes at least some form of computer-readable medium. Computer-readable medium can be any available medium accessible by at least one processing unit 602 or other devices including the operating environment. By way of example and not limitation, computer-readable medium can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Universal Optical Disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other tangible, non-transitory medium that can be used to store desired information. Computer storage media does not include communication media. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0061] Communication media embody computer-readable instructions, data structures, program modules, or other data in the form of modulated data signals (such as carrier waves or other transmission mechanisms), and include any information transmission medium. The term "modulated data signal" refers to a signal configured or altered in a manner that encodes information in the signal, which has one or more characteristics. By way of example and not limitation, communication media include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media.
[0062] Operating environment 600 can be a single computer operating in a network environment using logical connections to one or more remote computers. Remote computers can be personal computers, servers, routers, network PCs, peer-to-peer devices, or other public network nodes, and typically include many or all of the aforementioned elements as well as other unmentioned elements. Logical connections can include any method supported by available communication media. This type of network environment is common in offices, enterprise-wide computer networks, intranets, and the Internet.
[0063] The descriptions and illustrations of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the claimed disclosure as the best mode. The claimed disclosure should not be construed as limited to any aspect, such as, or the details provided in this application. Whether shown and described in combination or separately, various features (structural and methodological) are intended to be selectively included or omitted to produce embodiments with a particular set of features. Having provided the descriptions and illustrations of this application, those skilled in the art can contemplate variations, modifications, and alternatives falling within the spirit of the broader aspects of the overall inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.
[0064] The descriptions and illustrations of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the claimed disclosure as the best mode. The claimed disclosure should not be construed as limited to any aspect, such as, or the details provided in this application. Whether shown and described in combination or separately, various features (structural and methodological) are intended to be selectively included or omitted to produce embodiments with a particular set of features. Having provided the descriptions and illustrations of this application, those skilled in the art can contemplate variations, modifications, and alternatives falling within the spirit of the broader aspects of the overall inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.
[0065] This disclosure relates to systems and methods for automatically generating thumbnails, at least according to the examples provided in the following sections. This disclosure relates to a computer-implemented method for automatically generating thumbnail images representing videos associated with a list of items in an electronic marketplace. The method includes: retrieving a video associated with a list of items, wherein the video comprises a plurality of ordered frame images; retrieving item images associated with the list of items; determining, from the plurality of ordered frame images, at least a first score for a first frame image and a second score for a second frame image, wherein the scores are based at least on the similarity between the item images and the frame images; determining a first weight associated with the first frame image and a second weight associated with the second frame image, wherein the first weight is based at least on the ordered position of the first frame image; updating the first score based on the first weight and updating the second score based on the second weight; selecting a thumbnail image from either the first frame image or the second frame image, wherein the selection is based on the updated first score and the updated second score; and automatically adding the thumbnail image to the video. The similarity includes the similarity between the item image and the first frame image. The first weight is based at least on the ordered position of the first frame image in the video. The first weight is based at least on the number of times the content of the first frame image appears in other frames of the video. The first weight is based at least on the degree to which the first frame image indicates the characteristics of an item, wherein the characteristics of the item include at least one of the following: shape, color, or visual content. The method further includes: extracting the first frame image and a second frame image from the video based on a predetermined sampling rule. The method includes: receiving another item image associated with a list of items, wherein the other item image is different from the item image; updating a first score based on the similarity between the other item image and the first frame image; and updating the first score based on a third weight, wherein the third weight corresponds to the other item image. The method further includes: prompting the user to upload one or more item images associated with an item in descending order of relevance to the item.
[0066] Another aspect of this technology relates to a system for automatically generating thumbnail images representing videos associated with items. The system includes: a processor; and a memory storing computer-executable instructions that, when executed by the processor, cause the system to: retrieve a video associated with a list of items, wherein the video comprises a plurality of ordered frame images; retrieve item images associated with the list of items; determine, from the plurality of ordered frame images, at least a first score for a first frame image and a second score for a second frame image, wherein the scores are based at least on the similarity between the item images and the frame images; determine a first weight associated with the first frame image and a second weight associated with the second frame image, wherein the first weight is based at least on the ordered position of the first frame image; update the first score based on the first weight and update the second score based on the second weight; select a thumbnail image from either the first frame image or the second frame image, wherein the selection is based on the updated first score and the updated second score; and automatically add the thumbnail image to the video. The first score indicates the similarity between the item image and the first frame image. The first weight is based at least on the ordered position of the first frame image in the video. The first weight is based at least on the number of times the content of the first frame image appears in other frames of the video. The first weight is based at least on the degree to which the first frame image indicates the characteristics of an item, wherein the characteristics of the item include at least one of the following: shape, color, or visual content. The computer-executable instructions, when executed, further cause the system to extract the first frame image and the second frame image from the video based on a predetermined sampling rule. The computer-executable instructions, when executed, further cause the system to receive another item image associated with the list of items, wherein the other item image is different from the first item image; update the first score based on the similarity between the other item image and the first frame image; and update the first score based on a third weight, wherein the third weight corresponds to the other item image. The computer-executable instructions, when executed, further cause the system to prompt the user to upload one or more item images associated with the item in descending order of relevance to the item.
[0067] In other aspects, the technology relates to a computer-implemented method for automatically generating thumbnail images representing videos associated with a list of items in an electronic marketplace. The method includes: retrieving a video associated with an item, wherein the video comprises a plurality of frame images; retrieving a set of item images, wherein the set of item images comprises item images associated with the item, and wherein the position of an item image corresponds to the position of the item image relative to other item images published on a list of items in the electronic marketplace; extracting video frames from the video based on a sampling rule; determining a score for the video frame, wherein the score represents the degree of similarity between the video frame and each item image in the set of item images; determining a weighted score for the video frame based on the score, wherein the weighted score includes weights associated with the positions of the item images in the set of item images; selecting the content of the video frame as a thumbnail image of the video based on the weighted score; and publishing the thumbnail image to represent the video. The sampling rule includes: extracting video frames in the video at predetermined intervals. The method includes: updating the weighted score of the video frame based on another weight associated with the position of the video frame within the plurality of video frames. The similarity between the thumbnail image and at least one item image in the set of item images is greater than a predetermined threshold.
[0068] Any one of the above-mentioned aspects in combination with any other aspect of the above-mentioned aspects. Any one of the one or more aspects described herein.
Claims
1. A computer-implemented method for automatically generating thumbnail images, the thumbnail images representing videos associated with a list of items in an electronic marketplace, the method comprising: Retrieve a video associated with the list of items, wherein the video comprises a plurality of ordered frame images; Retrieve item images associated with the list of items; A first score for a first frame image and a second score for a second frame image are determined from the plurality of ordered frame images, wherein the scores are based at least on the degree of similarity between the item image and the frame image; A first weight associated with the first frame image and a second weight associated with the second frame image are determined, wherein the first weight is based at least on the ordered position of the first frame image in the video; The first score is updated based on the first weight, and the second score is updated based on the second weight; The thumbnail image is selected from either the first frame image or the second frame image, wherein the selection is based on an updated first score and an updated second score; and The thumbnail image is automatically added to the video.
2. The computer-implemented method according to claim 1, wherein, The similarity includes the similarity between the item image and the first frame image.
3. The computer-implemented method according to claim 1, wherein, The first weight is based at least on the number of times the content of the first frame appears in other frames of the video.
4. The computer-implemented method according to claim 1, wherein, The first weight is based at least on the degree to which the first frame image indicates the characteristics of the item, wherein the characteristics of the item include at least one of the following: shape, color, or visual content.
5. The computer-implemented method according to claim 1, further comprising: Based on predetermined sampling rules, the first frame image and the second frame image are extracted from the video.
6. The computer-implemented method according to claim 1, the method comprising: Receive another item image associated with the list of items, wherein the other item image is different from the item image; The first score is updated based on the similarity between the other item image and the first frame image; and The first score is updated based on a third weight, wherein the third weight corresponds to the other item image.
7. The computer-implemented method according to claim 1, further comprising: The user is prompted to upload one or more images of items associated with the item in descending order of their relevance to the item.
8. A system for automatically generating thumbnail images, the thumbnail images representing videos associated with items, the system comprising: processor; as well as The memory stores computer-executable instructions that, when executed by the processor, cause the system to: Retrieve a video associated with the list of items, wherein the video comprises a plurality of ordered frame images; Retrieve item images associated with the list of items; A first score for a first frame image and a second score for a second frame image are determined from the plurality of ordered frame images, wherein the scores are based at least on the degree of similarity between the item image and the frame image; A first weight associated with the first frame image and a second weight associated with the second frame image are determined, wherein the first weight is based at least on the ordered position of the first frame image in the video; The first score is updated based on the first weight, and the second score is updated based on the second weight; The thumbnail image is selected from either the first frame image or the second frame image, wherein the selection is based on an updated first score and an updated second score; and The thumbnail image is automatically added to the video.
9. The system according to claim 8, wherein, The first weight is based at least on the number of times the content of the first frame appears in other frames of the video.
10. The system according to claim 8, wherein, The first weight is based at least on the degree to which the first frame image indicates the characteristics of the item, wherein the characteristics of the item include at least one of the following: shape, color, or visual content.
11. The system of claim 8, wherein the computer-executable instructions, when executed, further cause the system to: Based on predetermined sampling rules, the first frame image and the second frame image are extracted from the video.
12. The system of claim 8, wherein the computer-executable instructions, when executed, further cause the system to: Receive another item image associated with the list of items, wherein, The other item image is different from the item image; The first score is updated based on the similarity between the other item image and the first frame image; and The first score is updated based on a third weight, wherein the third weight corresponds to the other item image.
13. The system of claim 8, wherein the computer-executable instructions, when executed, further cause the system to: The user is prompted to upload one or more images of items associated with the item in descending order of their relevance to the item.
14. A computer-implemented method for automatically generating thumbnail images, the thumbnail images representing videos associated with a list of items in an electronic marketplace, the method comprising: Retrieve a video associated with the item, wherein the video comprises multiple frames. Retrieve a set of item images, wherein the set of item images includes item images associated with the item, and wherein the position of the item image corresponds to the position of the item image relative to other item images published on the list of the item in the electronic marketplace; Video frames are extracted from the video based on sampling rules; Determine a score for the video frame, wherein the score represents the degree of similarity between the video frame and each item image in the set of item images; A weighted score for the video frame is determined based on the score, wherein the weighted score is based on weights associated with the positions of the object images in the set of object images; The content of the video frame is selected as the thumbnail image of the video based on the weighted score; and The thumbnail image is published to represent the video.
15. The computer-implemented method according to claim 14, in, The sampling rules include: extracting video frames from the video at predetermined intervals.
16. The computer-implemented method according to claim 14, the method comprising: The weighted score of the video frame is updated based on another weight associated with its position among multiple video frames in the video.
17. The computer-implemented method according to claim 14, wherein, The similarity between the thumbnail image and at least one item image in the set of item images is greater than a predetermined threshold.
Citation Information
Patent Citations
Generating moving thumbnails for videos
CN108780654A
Video thumbnail recommendation method fusing visual semantic information
CN111680190A