Method, device and system for processing visual material, computer terminal

By receiving a set of visual materials, selecting target shots based on product characteristics and perceptual persuasiveness, and sorting the shots using film and television production principles, the problem of high cost and time consumption caused by manual processing is solved, and efficient automated video generation is achieved.

CN114677190BActive Publication Date: 2025-12-30ALIBABA GROUP HOLDING LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202011554704.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-24
Publication Date
2025-12-30
Estimated Expiration
2040-12-24

AI Technical Summary

Technical Problem

In existing technologies, visual material processing relies on manual processing, which results in high costs and long processing times, making it impossible to generate videos efficiently.

Method used

By receiving a set of visual materials, the system filters target shots based on the semantic and perceptual persuasiveness of product features, and uses film and television production principles to sort the shots, generating videos that recommend products.

Benefits of technology

It improves the efficiency and quality of video generation, reduces costs, enhances the viewing experience and perceptual persuasiveness, and enables automated video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677190B_ABST
    Figure CN114677190B_ABST
Patent Text Reader

Abstract

The application discloses a kind of visual material processing method, device and system, computer terminal.Therein, the method includes: receiving visual material set, wherein visual material has included the product feature associated with product to be recommended;Visual material in visual material set is formed candidate lens set in the way of material sequence;Based on the semantics of product feature and the perceptual persuasion of visual material, multiple target lenses are filtered from candidate lens set;Video for recommending product is generated by lens sorting of multiple target lenses based on sorting factor, wherein sorting factor includes at least one of the following: semantic distance between target lenses, significant area ratio of product in target lens and similarity between target lenses;Output video.The application solves the technical problem that visual material is manually processed by manpower in the related art, resulting in higher cost and longer time-consuming.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the Internet field, and more specifically, to a method, apparatus and system for processing visual materials, and a computer terminal. Background Technology

[0002] In recent years, video has become a mainstream way to attract consumer attention. On e-commerce platforms, using video as a promotional tool is a viable method to increase product sharing rates and sales. One of the core steps in video production is generating a sequence of visual materials; however, this step is currently performed by experienced directors, and the entire process is costly and time-consuming.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method, apparatus, system, and computer terminal for processing visual materials, in order to at least solve the technical problem in the related art that the manual processing of visual materials by human labor results in high costs and long time consumption.

[0005] According to one aspect of the embodiments of this application, a method for processing visual materials is provided, comprising: receiving a set of visual materials, wherein each visual material contains product features associated with a product to be recommended; assembling the visual materials in the set into a candidate shot set in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots; and outputting the video.

[0006] According to another aspect of the embodiments of this application, a method for processing visual materials is also provided, comprising: acquiring a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set of visual materials into a candidate shot set in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; and ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of the product in the target shot, and similarity between target shots.

[0007] According to another aspect of the embodiments of this application, a method for processing visual materials is also provided, comprising: obtaining a set of visual materials by calling a first interface, wherein the first interface includes: a first parameter, the parameter value of the first parameter being a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling a candidate shot set from the visual materials in the set of visual materials in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots; and outputting a video by calling a second interface, wherein the second interface includes: a second parameter, the parameter value of the second parameter being a video.

[0008] According to another aspect of the embodiments of this application, a visual material processing apparatus is also provided, comprising: a receiving module for receiving a set of visual materials, wherein each visual material contains product features associated with a product to be recommended; a combining module for combining the visual materials in the set of visual materials into a candidate shot set in the form of a material sequence; a filtering module for filtering multiple target shots from the candidate shot set based on the semantics of product features and the perceptual persuasiveness of the visual materials; a sorting module for sorting the multiple target shots based on sorting factors to generate a video for recommending products, wherein the sorting factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots; and an output module for outputting the video.

[0009] According to another aspect of the embodiments of this application, a visual material processing apparatus is also provided, comprising: an acquisition module for acquiring a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; a combination module for assembling the visual materials in the set of visual materials into a candidate shot set in the form of a material sequence; a filtering module for filtering multiple target shots from the candidate shot set based on the semantics of product features and the perceptual persuasiveness of visual materials; and a ranking module for ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots.

[0010] According to another aspect of the embodiments of this application, a visual material processing apparatus is also provided, comprising: a first calling module, configured to obtain a visual material set by calling a first interface, wherein the first interface includes: a first parameter, the parameter value of the first parameter being a visual material set, wherein the visual materials all contain product features associated with the product to be recommended; a combination module, configured to combine the visual materials in the visual material set into a candidate shot set in the form of a material sequence; a filtering module, configured to filter multiple target shots from the candidate shot set based on the semantics of product features and the perceptual persuasiveness of visual materials; a sorting module, configured to sort the multiple target shots based on sorting factors to generate a video for recommending products, wherein the sorting factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots; and a second calling module, configured to output a video by calling a second interface, wherein the second interface includes: a second parameter, the parameter value of the second parameter being a video.

[0011] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the computer-readable storage medium is located to perform the above-described visual material processing method.

[0012] According to another aspect of the embodiments of this application, a computer terminal is also provided, including: a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the above-described method for processing visual materials when it runs.

[0013] According to another aspect of the embodiments of this application, a visual material processing system is also provided, including: a processor; and a memory connected to the processor, for providing the processor with instructions to perform the following processing steps: receiving a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set of visual materials into a candidate shot set in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; sorting the multiple target shots to generate a video for recommending the product; and outputting the video.

[0014] In this embodiment, after receiving a set of visual materials, the visual materials in the set can be arranged into a candidate shot set in the form of a material sequence. Further, based on the semantics of product features and the perceptual persuasiveness of the visual materials, multiple target shots are selected from the candidate shot set. These target shots are then ranked based on ranking factors to generate a video for recommending the product, which is then output to the user for viewing, achieving the purpose of video production. It is noteworthy that different ranking factors can be determined based on film and television production principles, integrating film and television production knowledge into the candidate shot set selection and shot ranking process. Based on the extraction of original video visual and structural information, editing techniques are modeled as optimization sub-modules, thereby enhancing the logical flow, improving the viewing experience and perceptual persuasiveness, and more effectively promoting the product's technical effects. This solves the technical problem in related technologies where visual materials are manually processed, resulting in high costs and long processing times. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0016] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for processing visual materials according to an embodiment of this application;

[0017] Figure 2 This is a flowchart of a method for processing visual materials according to an embodiment of this application;

[0018] Figure 3 This is a schematic diagram of an optional interactive interface according to an embodiment of this application;

[0019] Figure 4 This is a flowchart of an optional method for generating a sequence of visual materials according to an embodiment of this application;

[0020] Figure 5 This is a flowchart of another method for processing visual materials according to an embodiment of this application;

[0021] Figure 6 This is a flowchart of another method for processing visual materials according to an embodiment of this application;

[0022] Figure 7 This is a schematic diagram of a visual material processing apparatus according to an embodiment of this application;

[0023] Figure 8 This is a schematic diagram of another visual material processing apparatus according to an embodiment of this application;

[0024] Figure 9 This is a schematic diagram of another visual material processing apparatus according to an embodiment of this application;

[0025] Figure 10 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0029] Visual Material Sequence Generation (VMS) refers to taking some visual materials (images and videos) as input, selecting and sorting them to generate a sequence of materials, which will be used to generate the final video.

[0030] Film-making principles can refer to the rules of thumb for conveying a narrative through editing, such as using close-ups to emphasize a product.

[0031] Recursive clustering: Visual materials in a collection can be recursively divided into specific classes. Visual materials in the same class have similar characteristics, while visual materials in different classes have different characteristics.

[0032] The Wundt curve describes the psychological response patterns of users to information.

[0033] Significant region: This can refer to the area in an image that attracts attention or is relatively important, such as the area that a user first focuses on when viewing an image.

[0034] Nearest Neighbor Search refers to finding the element with the smallest distance to a given query term within a defined distance metric and a search space.

[0035] To generate sequences of visual materials, the following methods can be used: the first method is to skip certain images by sampling to sort a set of images; the second method is to calculate image differences and encourage the sequence of video clips to follow a "general plot of the story"; the third method can use RNNs and submodule optimizations to construct storylines for generating videos of events such as travel and parties.

[0036] However, the above solutions cannot guarantee logical flow or appropriate graphic discontinuity, resulting in a poor viewing experience. To address these issues, this application provides the following technical solution:

[0037] Example 1

[0038] According to an embodiment of this application, a method for processing visual materials is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0039] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a method for processing visual materials is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0040] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0041] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the visual material processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned visual material processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0042] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0043] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0044] It should be noted here that, in some optional embodiments, the above... Figure 1The computer device (or mobile device) shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance and is intended to illustrate the types of components that may exist in the aforementioned computer device (or mobile device).

[0045] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for processing the visual materials shown. Figure 2 This is a flowchart of a method for processing visual materials according to an embodiment of this application. Figure 2 As shown, the method includes the following steps:

[0046] Step S202: Receive a set of visual materials, wherein each visual material contains product features associated with the product to be recommended.

[0047] The products mentioned above can be items that need to be recommended to users. For example, in an e-commerce shopping scenario, products can be goods sold by different merchants, such as clothing, skincare products, cosmetics, and home appliances. Product features can be characteristics that characterize the product's own structure, such as appearance, quality, material, function, trademark, and packaging, but are not limited to these.

[0048] The aforementioned visual materials can be videos, images, etc., taken by merchants in e-commerce shopping scenarios, but are not limited to these.

[0049] In one optional embodiment, the entity executing the visual material processing method can be a client installed on a user's mobile terminal or computer terminal. To conserve computing resources on the mobile terminal or computer terminal and improve processing efficiency, the entity executing the visual material processing method can be a server, such as a cloud server, allowing users to upload collections of visual materials via a client installed on their mobile terminal or computer terminal.

[0050] In another alternative embodiment, an interactive interface can be provided to the user, such as... Figure 3 As shown, when a user needs to create a video, the user can click the "Upload Materials" button in the material upload area to select the visual materials to be uploaded, and then click the "Generate Video" button to confirm. In this way, the client can obtain the set of visual materials and process them; or, the client can directly upload the set of visual materials to the server via the network for processing.

[0051] Step S204: The visual materials in the visual material set are arranged into a candidate shot set in the form of a material sequence.

[0052] The candidate shots in the candidate shot set in the above steps can be subsequences of candidate footage used to generate the final shot sequence for video.

[0053] In one optional embodiment, after obtaining the set of visual materials, all visual materials can be grouped, different categories of visual materials can be divided into different groups, and the visual materials in the same group can be sorted in the form of material sequence to obtain a set of candidate shots.

[0054] Step S206: Based on the semantics of product features and the perceptual persuasiveness of visual materials, multiple target shots are selected from the candidate shot set.

[0055] The semantics of the product features in the above steps can refer to the specific meaning of the product features, which can be obtained through semantic recognition of the product features. Different visual materials contain different product features, and the degree of correlation between different product features and products varies. In order to ensure that the final generated recommended product videos are more attractive to users, in this embodiment of the application, it is necessary to ensure that the selected visual materials have a high degree of correlation with the products.

[0056] The perceived persuasiveness in the above steps refers to the likelihood that the recommended product in the visual materials is acceptable to the user. The higher the perceived persuasiveness, the more likely the user is to accept the product recommendation. To ensure that the final generated recommended product has high perceived persuasiveness, it is necessary to ensure that the selected visual materials have high perceived persuasiveness in this application.

[0057] In one optional embodiment, after generating a candidate shot set, visual materials with high relevance to the recommended product and high perceptual persuasiveness can be selected based on the semantics and perceptual persuasiveness of product features in all visual materials in the set, thereby obtaining multiple target shots.

[0058] Step S208: Rank multiple target shots based on ranking factors to generate a video for product recommendation. The ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots.

[0059] The aforementioned sorting factors can be based on principles of filmmaking, which can include the following three principles: a gradual increase in shot distance, logical storytelling, and shot discontinuity. For the first principle, "a gradual increase in shot distance," the goal of shot sorting is to ensure the visual storyline begins with a long shot, gradually decreasing in distance until the final close-up. Therefore, two factors can be considered: the salient area ratio (SSR) and the semantic distance between the shot and the product title, where the salient area ratio gradually increases and the semantic distance gradually decreases. For the second principle, "logical storytelling," it is necessary to ensure that shots from one scene category are displayed before shots from another scene category are shown. For the third principle, "shot discontinuity," it is necessary to reduce the similarity between two adjacent shots.

[0060] The video used to recommend the product in the above steps can refer to the product's promotional video. For example, in an e-commerce shopping scenario, the video can be a short product video.

[0061] In one alternative embodiment, in order to improve the viewing experience of the final video and obtain higher perceived persuasiveness, the principles of filmmaking can be used to determine the ranking factors, and the multiple selected target shots can be ranked using the ranking factors to generate the final visual material sequence, thereby generating the final promotional video.

[0062] Step S210: Output video.

[0063] In an alternative embodiment, where the visual material processing method is performed by the client, the video used for recommending products can be displayed as follows: Figure 3 The video display area shown is for user convenience. When the visual material processing is performed by the server, the server can return the video used to recommend products to the client via the network, which the client then displays, as shown in the image. Figure 3 The video display area shown is for user convenience.

[0064] The solution provided by the above embodiments of the present invention, after receiving a set of visual materials, can organize the visual materials in the set into a candidate shot set in the form of a material sequence. Further, based on the semantics of product features and the perceptual persuasiveness of the visual materials, multiple target shots are selected from the candidate shot set. These target shots are then ranked based on ranking factors to generate a video for recommending the product, which is then output to the user for viewing, achieving the purpose of video production. It is noteworthy that different ranking factors can be determined based on film and television production principles, integrating film and television production knowledge into the candidate shot set selection and shot ranking process. Based on the extraction of original video visual and structural information, editing techniques are modeled as optimization sub-modules, thereby enhancing the logical flow, improving the viewing experience and perceptual persuasiveness, and more effectively promoting the product's technical effects. This solves the technical problem in related technologies where visual materials are manually processed, resulting in high costs and long processing times.

[0065] In the above embodiments of this application, the visual materials in the visual material set are arranged into a candidate shot set in the form of material sequences, including: using a scene detection model to analyze the visual materials in the visual material set and obtain the scene category of the visual materials in the visual material set, wherein a convolutional neural network model is used to train the sample materials to generate the scene detection model; based on the scene category of the visual materials in the visual material set, the visual material set is recursively clustered to obtain multiple material sequences of different categories; and a candidate shot set composed of material sequences is obtained.

[0066] The scene detection model described above can be obtained by pre-training a convolutional neural network, which can predict the scene category in visual materials.

[0067] The above-mentioned scene categories can be specific scenes of shooting products in visual materials, such as model display, overall product display, product detail display, etc. Product details can refer to the display of a certain part of the product (such as cuffs, collars, etc.).

[0068] Optionally, the visual materials contained in the above material sequences may have the same scene type, and the visual appearance similarity of the visual materials may exceed a threshold; different material sequences may have different scene types.

[0069] The aforementioned threshold can be a similarity threshold determined in advance based on the actual processing accuracy and processing speed requirements. If the similarity is greater than the threshold, it indicates that the visual materials have similar visual appearances; if the similarity is less than the threshold, it indicates that the visual materials have different visual appearances.

[0070] In one alternative embodiment, a scene detection model can be used to process the visual material set, estimate the scene category of the visual materials, and then recursively cluster the visual materials based on the scene category. This ensures that visual materials clustered into the same category not only belong to the same scene category but also have similar visual appearances, thereby realizing the logic within the shot. Furthermore, for each cluster, an iterative nearest neighbor retrieval method can be used to select visual materials from each cluster, and the selected visual materials are sorted into candidate shots, thus obtaining a candidate shot set.

[0071] In the above embodiments of this application, obtaining a candidate shot set composed of a material sequence includes: randomly selecting the first visual material in the material sequence; selecting the next visual material with the highest similarity to the first visual material from the candidate materials, placing it as an adjacent material in the material sequence where the first visual material is located, and adjacent to the playback position of the first visual material; performing iterative selection on each visual material in the material sequence to select the next adjacent visual material, and outputting a candidate shot set.

[0072] The aforementioned alternative materials can be visual materials that have not been selected within the same cluster.

[0073] In an optional embodiment, during the process of using the iterative nearest neighbor retrieval method, for each cluster, the first visual material in the candidate shots can be randomly selected, and based on the nearest neighbor retrieval algorithm, the most similar unselected neighbor is selected from the candidate materials as the next in the sequence. The above method is then iteratively executed until all mutually selected shots are filtered out, thereby obtaining a set of candidate shots to ensure the logic within the set of candidate shots.

[0074] In the above embodiments of this application, multiple target shots are selected from a set of candidate shots based on the semantics of product features and the perceptual persuasiveness of visual materials. This includes: obtaining multiple visual materials contained in the candidate shots in the set of candidate shots; obtaining the semantic distance between each visual material based on the semantics of product features in each visual material; processing each visual material based on the Wundt curve to obtain the perceptual persuasiveness of each visual material; and selecting multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0075] The semantic distance mentioned above can be the Euclidean distance calculated based on the semantics of product features, but it is not limited to this.

[0076] The Wundt curve mentioned above is a curve obtained through prior learning.

[0077] In the selection process of target shots, to ensure a strong relationship between the shot and the product and high perceived persuasiveness, while also considering the balance between scene categories, the three factors of semantic distance (SED), perceived persuasiveness, and scene category can be comprehensively considered for shot selection. In an optional embodiment, for candidate shots in the candidate shot set, the semantic distance between each visual material can be calculated based on the semantics of each visual material and the product title, and the perceived persuasiveness of each visual material can be calculated according to the Wundt curve. Target shots are then selected based on semantic distance, perceived persuasiveness, and scene category.

[0078] In the above embodiments of this application, multiple target shots are obtained by filtering based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material. This includes: obtaining the weighted sum of the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material to obtain the score of each shot in the candidate shot set; and selecting the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter multiple target shots.

[0079] In an alternative embodiment, to comprehensively consider the three factors of semantic distance, perceived persuasiveness, and scene category, the following evaluation function can be used to calculate the score for each shot. :

[0080] ,

[0081] in, Indicates the lens semantic distance, Indicates the lens Perceptual persuasiveness Indicates the lens Scene categories, and These are the weight values ​​for semantic distance and perceived persuasiveness, respectively. The weight values ​​for different factors can be determined based on the importance of different factors.

[0082] Then, by sorting the sub-modes, the highest-scoring shot is selected from all candidate shot sets as the target shot each time, until the selected target shots reach the pre-set upper limit.

[0083] In the above embodiments of this application, sorting multiple target shots based on sorting factors to generate a video for recommending products includes: sorting multiple target shots based on sorting factors to generate a target sequence; and generating a video for recommending products according to the shots sorted according to the target sequence.

[0084] For the three ranking factors mentioned above, if the changes in product features in the visual materials do not conform to the aforementioned trends, then these changes can be penalized using an objective function. Based on the three film and television production principles mentioned above, the following objective function can be determined:

[0085] ,

[0086] ,

[0087] in, Indicates the significance ratio. Indicates semantic distance. Indicates the lens and the lens similarity, Indicates the scene category. The weight values ​​representing the significance ratio, The weight values ​​represent semantic distance. The weight value represents the similarity. Optionally, different ranking factors can have different weight values. The weight value can be used to determine the result of the shot ranking. In the embodiments of this application, the weight values ​​of different factors can be determined according to the importance of different factors in determining the result of the shot ranking.

[0088] In one alternative embodiment, the discontinuity of graphics is considered when the logical flow is satisfied, thereby achieving appropriate visual stimulation. The target sequence can be determined by searching all possible permutations through the above objective function, thereby achieving the purpose of generating a sequence of visual materials, and then generating promotional videos for recommended products.

[0089] The following is combined Figure 4 A preferred embodiment of this application will be described in detail using an e-commerce shopping scenario as an example. Figure 4 As shown, this method can be executed by either a client or a server. In this embodiment, server execution is used as an example. The method can be executed in three stages: Shot Composition, Shot Selection, and Shot Plotting. The specific execution steps for each stage are as follows:

[0090] Step S41: Input a set of visual materials, which contains a set of video clips and images with a user-specified duration.

[0091] Step S42: Shots are composed based on the input set of visual materials. These shots can be used as candidate subsequences to form the final sequence.

[0092] Optionally, the visual materials can be grouped and sorted into candidate shots, which can be further divided into two sub-steps: scene detection and recursive clustering.

[0093] Step S421: In the scene detection sub-step, the scene category of the visual material can be estimated using a scene detection model;

[0094] In step S422, in the recursive clustering sub-step, recursive clustering can be performed on the visual materials, and iterative nearest neighbor retrieval can be performed. For each cluster, the first visual material in the subsequence is randomly selected, and then the most similar unselected neighbor is selected from the candidate materials as the next one in the sequence, thereby obtaining a set of candidate shots.

[0095] Step S43: Select a shot based on perceived persuasiveness and semantic distance of the product.

[0096] Optionally, perceived persuasiveness can be calculated using a learnable Wundt curve, and the semantic distance between the shot and the product title can be calculated based on the semantics of the visual material. Then, the evaluation function comprehensively considers three factors: semantic distance, perceived persuasiveness, and scene category. The submodal sorting is used to select the shot with the highest score from the candidate shots each time until the number of shots reaches the set upper limit.

[0097] Step S44: Based on three film and television production principles, a visual storyline sequence is generated by considering semantic distance (SED), significant region ratio (SRR), and similarity of selected shots (SIM).

[0098] Optionally, an objective function can be constructed based on semantic distance (SED), significant region ratio (SRR), and similarity (SIM) of the selected shots, and the selected shots can be sorted based on the objective function to obtain the final output visual storyline sequence.

[0099] Through the above steps, this application proposes a method for generating visual material sequences using film and television production principles. By integrating film and television production knowledge, the principles of film production are incorporated into automatically generated visual storylines to enhance logical flow, viewing experience, and perceptual persuasiveness. Based on the extraction of original video visual and structural information, editing techniques are modeled as optimization sub-modules, thereby more effectively promoting e-commerce products.

[0100] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0102] Example 2

[0103] According to an embodiment of this application, a method for processing visual materials is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0104] Figure 5 This is a flowchart of another method for processing visual materials according to an embodiment of this application. For example... Figure 5 As shown, the method includes the following steps:

[0105] Step S502: Obtain a set of visual materials, wherein each visual material contains product features associated with the product to be recommended.

[0106] The products mentioned above can be items that need to be recommended to users. For example, in an e-commerce shopping scenario, products can be goods sold by different merchants, such as clothing, skincare products, cosmetics, and home appliances. Product features can be characteristics that characterize the product's own structure, such as appearance, quality, material, function, trademark, and packaging, but are not limited to these.

[0107] The aforementioned visual materials can be videos, images, etc., taken by merchants in e-commerce shopping scenarios, but are not limited to these.

[0108] Step S504: The visual materials in the visual material set are arranged into a candidate shot set in the form of a material sequence.

[0109] The candidate shots in the candidate shot set in the above steps can be subsequences of candidate footage used to generate the final shot sequence for video.

[0110] Step S506: Based on the semantics of product features and the perceptual persuasiveness of visual materials, multiple target shots are selected from the candidate shot set.

[0111] The semantics of the product features in the above steps can refer to the specific meaning of the product features, which can be obtained through semantic recognition of the product features. Different visual materials contain different product features, and the degree of correlation between different product features and products varies. In order to ensure that the final generated recommended product videos are more attractive to users, in this embodiment of the application, it is necessary to ensure that the selected visual materials have a high degree of correlation with the products.

[0112] The perceived persuasiveness in the above steps refers to the likelihood that the recommended product in the visual materials is acceptable to the user. The higher the perceived persuasiveness, the more likely the user is to accept the product recommendation. To ensure that the final generated recommended product has high perceived persuasiveness, it is necessary to ensure that the selected visual materials have high perceived persuasiveness in this application.

[0113] Step S508: Rank multiple target shots based on ranking factors to generate a video for recommending products. The ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots.

[0114] The video used to recommend the product in the above steps can refer to the product's promotional video. For example, in an e-commerce shopping scenario, the video can be a short product video.

[0115] In the above embodiments of this application, the visual materials in the visual material set are arranged into a candidate shot set in the form of material sequences, including: using a scene detection model to analyze the visual materials in the visual material set and obtain the scene category of the visual materials in the visual material set, wherein a convolutional neural network model is used to train the sample materials to generate the scene detection model; based on the scene category of the visual materials in the visual material set, the visual material set is recursively clustered to obtain multiple material sequences of different categories; and a candidate shot set composed of material sequences is obtained.

[0116] The scene detection model described above can be obtained by pre-training a convolutional neural network, which can predict the scene category in visual materials.

[0117] The above-mentioned scene categories can be specific scenes of shooting products in visual materials, such as model display, overall product display, product detail display, etc. Product details can refer to the display of a certain part of the product (such as cuffs, collars, etc.).

[0118] Optionally, the visual materials contained in the above material sequences may have the same scene type, and the visual appearance similarity of the visual materials may exceed a threshold; different material sequences may have different scene types.

[0119] The aforementioned threshold can be a similarity threshold determined in advance based on the actual processing accuracy and processing speed requirements. If the similarity is greater than the threshold, it indicates that the visual materials have similar visual appearances; if the similarity is less than the threshold, it indicates that the visual materials have different visual appearances.

[0120] In the above embodiments of this application, obtaining a candidate shot set composed of a material sequence includes: randomly selecting the first visual material in the material sequence; selecting the next visual material with the highest similarity to the first visual material from the candidate materials, placing it as an adjacent material in the material sequence where the first visual material is located, and adjacent to the playback position of the first visual material; performing iterative selection on each visual material in the material sequence to select the next adjacent visual material, and outputting a candidate shot set.

[0121] The aforementioned alternative materials can be visual materials that have not been selected within the same cluster.

[0122] In the above embodiments of this application, multiple target shots are selected from a set of candidate shots based on the semantics of product features and the perceptual persuasiveness of visual materials. This includes: obtaining multiple visual materials contained in the candidate shots in the set of candidate shots; obtaining the semantic distance between each visual material based on the semantics of product features in each visual material; processing each visual material based on the Wundt curve to obtain the perceptual persuasiveness of each visual material; and selecting multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0123] The semantic distance mentioned above can be the Euclidean distance calculated based on the semantics of product features, but it is not limited to this.

[0124] The Wundt curve mentioned above is a curve obtained through prior learning.

[0125] In the above embodiments of this application, multiple target shots are obtained by filtering based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material. This includes: obtaining the weighted sum of the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material to obtain the score of each shot in the candidate shot set; and selecting the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter multiple target shots.

[0126] In the above embodiments of this application, sorting multiple target shots based on sorting factors to generate a video for recommending products includes: sorting multiple target shots based on sorting factors to generate a target sequence; and generating a video for recommending products according to the shots sorted according to the target sequence.

[0127] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0128] Example 3

[0129] According to an embodiment of this application, a method for processing visual materials is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0130] Figure 6 This is a flowchart illustrating another method for processing visual materials according to an embodiment of this application. For example... Figure 6 As shown, the method includes the following steps:

[0131] Step S602: Obtain a set of visual materials by calling the first interface, wherein the first interface includes: a first parameter, the value of which is a set of visual materials, and each visual material contains product features associated with the product to be recommended.

[0132] The first interface in the above steps can be an interface for data interaction between the client and the server. The client can pass the visual material collection to the interface function as two parameters of the interface function to achieve the purpose of uploading the visual material collection to the server.

[0133] The products mentioned above can be items that need to be recommended to users. For example, in an e-commerce shopping scenario, products can be goods sold by different merchants, such as clothing, skincare products, cosmetics, and home appliances. Product features can be characteristics that characterize the product's own structure, such as appearance, quality, material, function, trademark, and packaging, but are not limited to these.

[0134] The aforementioned visual materials can be videos, images, etc., taken by merchants in e-commerce shopping scenarios, but are not limited to these.

[0135] Step S604: The visual materials in the visual material set are arranged into a candidate shot set in the form of a material sequence.

[0136] The candidate shots in the candidate shot set in the above steps can be subsequences of candidate footage used to generate the final shot sequence for video.

[0137] Step S606: Based on the semantic and visual persuasiveness of the product features, multiple target shots are selected from the candidate shot set.

[0138] The semantics of the product features in the above steps can refer to the specific meaning of the product features, which can be obtained through semantic recognition of the product features. Different visual materials contain different product features, and the degree of correlation between different product features and products varies. In order to ensure that the final generated recommended product videos are more attractive to users, in this embodiment of the application, it is necessary to ensure that the selected visual materials have a high degree of correlation with the products.

[0139] The perceived persuasiveness in the above steps refers to the likelihood that the recommended product in the visual materials is acceptable to the user. The higher the perceived persuasiveness, the more likely the user is to accept the product recommendation. To ensure that the final generated recommended product has high perceived persuasiveness, it is necessary to ensure that the selected visual materials have high perceived persuasiveness in this application.

[0140] Step S608: Rank multiple target shots based on ranking factors to generate a video for product recommendation. The ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots.

[0141] The video used to recommend the product in the above steps can refer to the product's promotional video. For example, in an e-commerce shopping scenario, the video can be a short product video.

[0142] Step S610: Output video by calling the second interface, wherein the second interface includes a second parameter, the value of which is video.

[0143] The second interface in the above steps can be an interface for data interaction between the server and the client. The server can pass the video used for recommending products into the interface function as a parameter of the interface function, so as to achieve the purpose of sending the video used for recommending products to the client.

[0144] In the above embodiments of this application, the visual materials in the visual material set are arranged into a candidate shot set in the form of material sequences, including: using a scene detection model to analyze the visual materials in the visual material set and obtain the scene category of the visual materials in the visual material set, wherein a convolutional neural network model is used to train the sample materials to generate the scene detection model; based on the scene category of the visual materials in the visual material set, the visual material set is recursively clustered to obtain multiple material sequences of different categories; and a candidate shot set composed of material sequences is obtained.

[0145] The scene detection model described above can be obtained by pre-training a convolutional neural network, which can predict the scene category in visual materials.

[0146] The above-mentioned scene categories can be specific scenes of shooting products in visual materials, such as model display, overall product display, product detail display, etc. Product details can refer to the display of a certain part of the product (such as cuffs, collars, etc.).

[0147] Optionally, the visual materials contained in the above material sequences may have the same scene type, and the visual appearance similarity of the visual materials may exceed a threshold; different material sequences may have different scene types.

[0148] The aforementioned threshold can be a similarity threshold determined in advance based on the actual processing accuracy and processing speed requirements. If the similarity is greater than the threshold, it indicates that the visual materials have similar visual appearances; if the similarity is less than the threshold, it indicates that the visual materials have different visual appearances.

[0149] In the above embodiments of this application, obtaining a candidate shot set composed of a material sequence includes: randomly selecting the first visual material in the material sequence; selecting the next visual material with the highest similarity to the first visual material from the candidate materials, placing it as an adjacent material in the material sequence where the first visual material is located, and adjacent to the playback position of the first visual material; performing iterative selection on each visual material in the material sequence to select the next adjacent visual material, and outputting a candidate shot set.

[0150] The aforementioned alternative materials can be visual materials that have not been selected within the same cluster.

[0151] In the above embodiments of this application, multiple target shots are selected from a set of candidate shots based on the semantics of product features and the perceptual persuasiveness of visual materials. This includes: obtaining multiple visual materials contained in the candidate shots in the set of candidate shots; obtaining the semantic distance between each visual material based on the semantics of product features in each visual material; processing each visual material based on the Wundt curve to obtain the perceptual persuasiveness of each visual material; and selecting multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0152] The semantic distance mentioned above can be the Euclidean distance calculated based on the semantics of product features, but it is not limited to this.

[0153] The Wundt curve mentioned above is a curve obtained through prior learning.

[0154] In the above embodiments of this application, multiple target shots are obtained by filtering based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material. This includes: obtaining the weighted sum of the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material to obtain the score of each shot in the candidate shot set; and selecting the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter multiple target shots.

[0155] In the above embodiments of this application, sorting multiple target shots based on sorting factors to generate a video for recommending products includes: sorting multiple target shots based on sorting factors to generate a target sequence; and generating a video for recommending products according to the shots sorted according to the target sequence.

[0156] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0157] Example 4

[0158] According to embodiments of this application, a visual material processing apparatus for implementing the above-described visual material processing method is also provided, such as... Figure 7 As shown, the device 700 includes: a receiving module 702, a combining module 704, a filtering module 706, a sorting module 708, and an output module 710.

[0159] The receiving module 702 receives a set of visual materials, each containing product features associated with the product to be recommended; the combining module 704 combines the visual materials in the set into a candidate shot set as a sequence; the filtering module 706 filters multiple target shots from the candidate shot set based on the semantics of product features and the perceptual persuasiveness of the visual materials; the ranking module 708 ranks the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots; and the output module 710 outputs the video.

[0160] It should be noted that the receiving module 702, combining module 704, filtering module 706, sorting module 708, and output module 710 mentioned above correspond to steps S202 to S210 in Embodiment 1. The five modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0161] In the above embodiments of this application, the combination module includes: an analysis unit, a clustering unit, and a combination unit.

[0162] The analysis unit is used to analyze the visual materials in the visual material set using a scene detection model to obtain the scene category of the visual materials in the visual material set. The scene detection model is generated by training the sample materials using a convolutional neural network model. The clustering unit is used to recursively cluster the visual material set based on the scene category of the visual materials in the visual material set, and the clustering results in multiple material sequences of different categories. The combination unit is used to obtain a candidate shot set composed of material sequences.

[0163] In the above embodiments of this application, the combination unit includes: a first selection subunit, a second selection subunit, and an output subunit.

[0164] The first selection subunit is used to randomly select the first visual material in the sequence of materials; the second selection subunit is used to select the next visual material with the highest similarity to the first visual material from the candidate materials, and place it as an adjacent material in the sequence of materials where the first visual material is located, and adjacent to the playback position of the first visual material; the output subunit is used to perform iterative selection of the next adjacent visual material for each visual material in the sequence of materials, and output the candidate shot set.

[0165] In the above embodiments of this application, the filtering module includes: a first acquisition unit, a second acquisition unit, a third acquisition unit, and a filtering unit.

[0166] The first acquisition unit is used to acquire multiple visual materials contained in the candidate shots in the candidate shot set; the second acquisition unit is used to acquire the semantic distance between each visual material based on the semantics of the product features in each visual material; the third acquisition unit is used to process each visual material based on the Wundt curve to acquire the perceptual persuasiveness of each visual material; and the filtering unit is used to filter out multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0167] In the above embodiments of this application, the filtering unit includes: an acquisition subunit and a filtering subunit.

[0168] The acquisition subunit is used to acquire the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the weighted sum of the scene categories of each visual material to obtain the score of each shot in the candidate shot set; the filtering subunit is used to select the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter and obtain multiple target shots.

[0169] In the above embodiments of this application, the sorting module includes a sorting unit and a generation unit.

[0170] The sorting unit is used to sort multiple target shots based on sorting factors to generate a target sequence; the generation unit is used to generate a video for recommending products based on the shots sorted according to the target sequence.

[0171] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0172] Example 5

[0173] According to embodiments of this application, a visual material processing apparatus for implementing the above-described visual material processing method is also provided, such as... Figure 8 As shown, the device 800 includes: an acquisition module 802, a combination module 804, a filtering module 806, and a sorting module 808.

[0174] The acquisition module 802 is used to acquire a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; the combination module 804 is used to combine the visual materials in the set of visual materials into a candidate shot set in the form of a material sequence; the filtering module 806 is used to filter multiple target shots from the candidate shot set based on the semantics of product features and the perceptual persuasiveness of visual materials; and the ranking module 808 is used to rank the multiple target shots based on ranking factors to generate a video for recommending products, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots.

[0175] It should be noted that the acquisition module 802, combination module 804, filtering module 806, and sorting module 808 mentioned above correspond to steps S502 to S508 in Embodiment 2. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run on the computer terminal 10 provided in Embodiment 1.

[0176] In the above embodiments of this application, the combination module includes: an analysis unit, a clustering unit, and a combination unit.

[0177] The analysis unit is used to analyze the visual materials in the visual material set using a scene detection model to obtain the scene category of the visual materials in the visual material set. The scene detection model is generated by training the sample materials using a convolutional neural network model. The clustering unit is used to recursively cluster the visual material set based on the scene category of the visual materials in the visual material set, and the clustering results in multiple material sequences of different categories. The combination unit is used to obtain a candidate shot set composed of material sequences.

[0178] In the above embodiments of this application, the combination unit includes: a first selection subunit, a second selection subunit, and an output subunit.

[0179] The first selection subunit is used to randomly select the first visual material in the sequence of materials; the second selection subunit is used to select the next visual material with the highest similarity to the first visual material from the candidate materials, and place it as an adjacent material in the sequence of materials where the first visual material is located, and adjacent to the playback position of the first visual material; the output subunit is used to perform iterative selection of the next adjacent visual material for each visual material in the sequence of materials, and output the candidate shot set.

[0180] In the above embodiments of this application, the filtering module includes: a first acquisition unit, a second acquisition unit, a third acquisition unit, and a filtering unit.

[0181] The first acquisition unit is used to acquire multiple visual materials contained in the candidate shots in the candidate shot set; the second acquisition unit is used to acquire the semantic distance between each visual material based on the semantics of the product features in each visual material; the third acquisition unit is used to process each visual material based on the Wundt curve to acquire the perceptual persuasiveness of each visual material; and the filtering unit is used to filter out multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0182] In the above embodiments of this application, the filtering unit includes: an acquisition subunit and a filtering subunit.

[0183] The acquisition subunit is used to acquire the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the weighted sum of the scene categories of each visual material to obtain the score of each shot in the candidate shot set; the filtering subunit is used to select the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter and obtain multiple target shots.

[0184] In the above embodiments of this application, the sorting module includes a sorting unit and a generation unit.

[0185] The sorting unit is used to sort multiple target shots based on sorting factors to generate a target sequence; the generation unit is used to generate a video for recommending products based on the shots sorted according to the target sequence.

[0186] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0187] Example 6

[0188] According to embodiments of this application, a visual material processing apparatus for implementing the above-described visual material processing method is also provided, such as... Figure 9 As shown, the device 900 includes: a first calling module 902, a combination module 904, a filtering module 906, a sorting module 908, and a second calling module 910.

[0189] The system comprises the following components: a first calling module 902, which obtains a set of visual materials by calling a first interface, wherein the first interface includes a first parameter, the value of which is a set of visual materials, each containing product features associated with the product to be recommended; a combination module 904, which assembles the visual materials in the set into a candidate shot set as a sequence of materials; a filtering module 906, which filters multiple target shots from the candidate shot set based on the semantics of product features and the perceptual persuasiveness of the visual materials; a ranking module 908, which ranks the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of products in target shots, and similarity between target shots; and a second calling module 910, which outputs a video by calling a second interface, wherein the second interface includes a second parameter, the value of which is a video.

[0190] It should be noted that the first calling module 902, the combination module 904, the filtering module 906, the sorting module 908, and the second calling module 910 mentioned above correspond to steps S602 to S610 in Embodiment 3. The five modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0191] In the above embodiments of this application, the combination module includes: an analysis unit, a clustering unit, and a combination unit.

[0192] The analysis unit is used to analyze the visual materials in the visual material set using a scene detection model to obtain the scene category of the visual materials in the visual material set. The scene detection model is generated by training the sample materials using a convolutional neural network model. The clustering unit is used to recursively cluster the visual material set based on the scene category of the visual materials in the visual material set, and the clustering results in multiple material sequences of different categories. The combination unit is used to obtain a candidate shot set composed of material sequences.

[0193] In the above embodiments of this application, the combination unit includes: a first selection subunit, a second selection subunit, and an output subunit.

[0194] The first selection subunit is used to randomly select the first visual material in the sequence of materials; the second selection subunit is used to select the next visual material with the highest similarity to the first visual material from the candidate materials, and place it as an adjacent material in the sequence of materials where the first visual material is located, and adjacent to the playback position of the first visual material; the output subunit is used to perform iterative selection of the next adjacent visual material for each visual material in the sequence of materials, and output the candidate shot set.

[0195] In the above embodiments of this application, the filtering module includes: a first acquisition unit, a second acquisition unit, a third acquisition unit, and a filtering unit.

[0196] The first acquisition unit is used to acquire multiple visual materials contained in the candidate shots in the candidate shot set; the second acquisition unit is used to acquire the semantic distance between each visual material based on the semantics of the product features in each visual material; the third acquisition unit is used to process each visual material based on the Wundt curve to acquire the perceptual persuasiveness of each visual material; and the filtering unit is used to filter out multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0197] In the above embodiments of this application, the filtering unit includes: an acquisition subunit and a filtering subunit.

[0198] The acquisition subunit is used to acquire the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the weighted sum of the scene categories of each visual material to obtain the score of each shot in the candidate shot set; the filtering subunit is used to select the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter and obtain multiple target shots.

[0199] In the above embodiments of this application, the sorting module includes a sorting unit and a generation unit.

[0200] The sorting unit is used to sort multiple target shots based on sorting factors, and the target sequence generation unit is used to generate a video for recommending products by sorting the shots according to the target sequence.

[0201] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0202] Example 7

[0203] According to an embodiment of this application, a visual material processing system is also provided, comprising:

[0204] Processor; and

[0205] The memory, connected to the processor, provides instructions to the processor to perform the following processing steps: receiving a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set as a sequence of materials; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; sorting the multiple target shots to generate a video for recommending the product; and outputting the video.

[0206] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0207] Example 8

[0208] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.

[0209] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0210] In this embodiment, the computer terminal described above can execute the program code for the following steps in the visual material processing method: receiving a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; sorting the multiple target shots to generate a video for recommending the product; and outputting the video.

[0211] Optionally, Figure 10 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 10 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1002 and memory 1004.

[0212] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the visual material processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned visual material processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0213] The processor can invoke information and applications stored in memory via a transmission device to perform the following steps: receiving a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set as a sequence of materials; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of the product in the target shot, and similarity between target shots; and outputting the video.

[0214] Optionally, the processor may also execute program code for the following steps: using a scene detection model to analyze the visual materials in the visual material set, and obtaining the scene category of the visual materials in the visual material set, wherein a convolutional neural network model is used to train the sample materials to generate the scene detection model; based on the scene category of the visual materials in the visual material set, recursively clustering the visual material set to obtain multiple material sequences of different categories; and obtaining a candidate shot set composed of material sequences.

[0215] Optionally, the processor may also execute program code that performs the following steps: randomly selects the first visual material from the material sequence; selects the next visual material with the highest similarity to the first visual material from the candidate materials, places it as an adjacent material in the material sequence where the first visual material is located, and is adjacent to the playback position of the first visual material; performs iterative selection on each visual material in the material sequence to select the next adjacent visual material, and outputs a set of candidate shots.

[0216] Optionally, the processor may also execute program code that performs the following steps: obtaining multiple visual materials contained in the candidate lenses in the candidate lens set; obtaining the semantic distance between each visual material based on the semantics of the product features in each visual material; processing each visual material based on the Wundt curve to obtain the perceptual persuasiveness of each visual material; and selecting multiple target lenses based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0217] Optionally, the processor may also execute program code that performs the following steps: obtains the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the weighted sum of the scene categories of each visual material, to obtain the score of each shot in the candidate shot set; and selects the shot with the highest score from the candidate shot set each time by using a sub-modal sorting method to filter and obtain multiple target shots.

[0218] Optionally, the processor may also execute program code that performs the following steps: sorting multiple target shots based on sorting factors to generate a target sequence; and generating a video for recommending products based on the shots sorted according to the target sequence.

[0219] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set in the form of a sequence of materials; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of the product in the target shot, and similarity between target shots.

[0220] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring a set of visual materials by invoking a first interface, wherein the first interface includes a first parameter, the value of which is the set of visual materials, each containing product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set as a sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, the salient area ratio of the product in the target shots, and the similarity between target shots; and outputting a video by invoking a second interface, wherein the second interface includes a second parameter, the value of which is the video.

[0221] This application provides a scheme for generating a sequence of visual materials. By integrating film and television production knowledge into the process of candidate shot selection and shot sorting, and modeling editing techniques as optimization sub-modules based on the extraction of original video visual and structural information, the scheme enhances the logical flow, improves the viewing experience and perceptual persuasiveness, and more effectively promotes the technical effects of the product. This solves the technical problem in related technologies where visual materials are manually processed, resulting in high costs and long processing times.

[0222] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Figure 10 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more... Figure 10 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 10 The different configurations shown.

[0223] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0224] Example 9

[0225] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the visual material processing method provided in the above embodiments.

[0226] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0227] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of the product in the target shot, and similarity between target shots; and outputting the video.

[0228] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: using a scene detection model to analyze the visual materials in the visual material set, obtaining the scene categories of the visual materials in the visual material set, wherein a convolutional neural network model is used to train the sample materials to generate the scene detection model; based on the scene categories of the visual materials in the visual material set, recursively clustering the visual material set to obtain multiple material sequences of different categories; and obtaining a candidate shot set composed of material sequences.

[0229] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: randomly selecting the first visual material from the material sequence; selecting the next visual material with the highest similarity to the first visual material from the candidate materials, placing it as an adjacent material in the material sequence where the first visual material is located, and adjacent to the playback position of the first visual material; performing iterative selection on each visual material in the material sequence to find the next adjacent visual material, and outputting a set of candidate shots.

[0230] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: obtaining multiple visual materials contained in the candidate shots in the candidate shot set; obtaining the semantic distance between each visual material based on the semantics of the product features in each visual material; processing each visual material based on the Wundt curve to obtain the perceptual persuasiveness of each visual material; and selecting multiple target shots based on the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the scene category of each visual material.

[0231] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: obtaining the semantic distance between each visual material, the perceptual persuasiveness of each visual material, and the weighted sum of the scene categories of each visual material to obtain the score of each shot in the candidate shot set; and selecting the shot with the highest score from the candidate shot set each time by using a submodal sorting method to filter and obtain multiple target shots.

[0232] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: sorting multiple target shots based on sorting factors to generate a target sequence; and generating a video for recommending products based on the shots sorted according to the target sequence.

[0233] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a set of visual materials, wherein each visual material contains product features associated with the product to be recommended; assembling the visual materials in the set into a candidate shot set in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of the product in the target shot, and similarity between target shots.

[0234] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a set of visual materials by calling a first interface, wherein the first interface includes: a first parameter, the value of which is a set of visual materials, all of which contain product features associated with the product to be recommended; assembling a candidate shot set from the visual materials in the set of visual materials in the form of a material sequence; selecting multiple target shots from the candidate shot set based on the semantics of the product features and the perceptual persuasiveness of the visual materials; ranking the multiple target shots based on ranking factors to generate a video for recommending the product, wherein the ranking factors include at least one of the following: semantic distance between target shots, salient area ratio of the product in the target shot, and similarity between target shots; and outputting a video by calling a second interface, wherein the second interface includes: a second parameter, the value of which is a video.

[0235] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0236] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0237] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0238] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0239] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0240] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0241] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method of processing visual material, characterized by, The method comprises the following steps: receiving a set of visual materials, wherein each of the visual materials contains a product feature associated with a product to be recommended; composing a set of candidate shots in the form of sequences of visual materials from the set of visual materials, wherein the sequences of visual materials are determined by analyzing scene categories of the visual materials in the set of visual materials by using a scene detection model, the visual materials contained in the sequences of visual materials have the same scene category, and the visual appearance similarity of the visual materials exceeds a threshold value, the threshold value is determined according to processing accuracy and processing speed, and the scene detection model is generated by training sample materials by using a convolutional neural network model; selecting a plurality of target shots from the set of candidate shots based on semantics of product features and perceptual persuasiveness of visual materials; performing shot sorting on the plurality of target shots based on sorting factors to generate a video for recommending the product, wherein the sorting factors include at least one of the following: semantic distance between the target shots, ratio of a prominent area of the product in the target shots, and similarity between the target shots; outputting the video; the step of selecting a plurality of target shots from the set of candidate shots based on semantics of product features and perceptual persuasiveness of visual materials comprises the following steps: obtaining semantic distances between a plurality of the visual materials based on semantics of the product features and product titles of corresponding products of the product features; selecting the plurality of target shots based on the semantic distances between the plurality of the visual materials, the perceptual persuasiveness of the plurality of the visual materials, and scene categories of the plurality of the visual materials.

2. The method of claim 1, wherein, the step of composing a set of candidate shots in the form of sequences of visual materials from the set of visual materials comprises the following steps: analyzing the visual materials in the set of visual materials by using a scene detection model to obtain scene categories of the visual materials in the set of visual materials; recursively clustering the set of visual materials based on the scene categories of the visual materials in the set of visual materials to obtain a plurality of sequences of visual materials of different categories; obtaining the set of candidate shots composed in the form of the sequences of visual materials.

3. The method of claim 2, wherein, The scene types of different sequences of visual materials are different.

4. The method of claim 3, wherein, the step of obtaining the set of candidate shots composed in the form of the sequences of visual materials comprises the following steps: randomly selecting a first visual material in a sequence from the sequences of visual materials; selecting a next visual material with the highest similarity to the first visual material from alternative visual materials as an adjacent visual material and placing the adjacent visual material in the sequence of visual materials in which the first visual material is located and adjacent to a playing position of the first visual material; performing iterative selection of an adjacent next visual material for each visual material in the sequence of visual materials to output the set of candidate shots.

5. The method of claim 1, wherein, the step of selecting a plurality of target shots from the set of candidate shots based on semantics of product features and perceptual persuasiveness of visual materials further comprises the following steps: obtaining a plurality of visual materials contained in a candidate shot in the set of candidate shots; processing each of the plurality of visual materials by using a von Frey curve to obtain perceptual persuasiveness of the each of the visual materials.

6. The method of claim 5, wherein, The filtering of the multiple target shots based on the semantic distance between each visual material, the perceived persuasiveness of each visual material, and the scene category of each visual material comprises: Obtaining a weighted sum of the semantic distance between each visual material, the perceived persuasiveness of each visual material, and the scene category of each visual material to obtain a score of each shot in the candidate shot set; Selecting a shot with the highest score from the candidate shot set each time in a way of secondary mode sorting to filter the multiple target shots.

7. The method of claim 1, wherein, The shot sorting of the multiple target shots based on the sorting factors to generate a video for recommending the product comprises: The shot sorting of the multiple target shots based on the sorting factors to generate a target sequence; Generating the video for recommending the product according to the shots sorted in the target sequence.

8. The method of claim 7, wherein, Different sorting factors have different weight values, and the weight values are used to determine the result of the shot sorting.

9. A visual material processing apparatus, characterized by comprising: Comprise: The receiving module is configured to receive a visual material set, wherein each visual material in the visual material set comprises a product feature associated with a product to be recommended; The combination module is configured to combine the visual materials in the visual material set into a candidate shot set in a material sequence, wherein the material sequence is determined by analyzing the scene category of the visual materials obtained by using a scene detection model, the visual materials included in the material sequence have the same scene category, and the visual appearance similarity of the visual materials exceeds a threshold value, the threshold value is determined according to processing accuracy and processing speed, and the scene detection model is generated by training sample materials using a convolutional neural network model; The filtering module is configured to filter multiple target shots from the candidate shot set based on the semantics of the product features and the perceived persuasiveness of the visual materials; The sorting module is configured to sort the multiple target shots based on sorting factors to generate a video for recommending the product, wherein the sorting factors comprise at least one of the following: a semantic distance between the target shots, a significant area ratio of the product in the target shots, and a similarity between the target shots; The output module is configured to output the video. The filtering module is further configured to obtain a semantic distance between multiple visual materials based on the semantics of the product features and the product title of the corresponding product, and filter the multiple target shots based on the semantic distance between the multiple visual materials, the perceived persuasiveness of the multiple visual materials, and the scene category of the multiple visual materials.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program controls the device where the computer-readable storage medium is located to execute the visual material processing method of any one of claims 1 to 8 when the program is running.

11. A computer terminal, characterized in that Comprise: A memory and a processor, wherein the processor is configured to run a program stored in the memory, and the program executes the visual material processing method of any one of claims 1 to 8 when the program is running.

12. A system for processing visual material, characterized by Comprise: A processor; and ​ The memory is connected with the processor and is used to provide the processor with instructions for processing the following processing steps: receiving a set of visual materials, wherein each of the visual materials contains a product feature associated with a product to be recommended; composing a set of candidate shots in the form of a sequence of visual materials from the set of visual materials, wherein the sequence of visual materials is determined by analyzing the scene categories of the visual materials obtained by using a scene detection model, the visual materials contained in the sequence of visual materials have the same scene category, and the visual appearance similarity of the visual materials exceeds a threshold value, the threshold value is determined according to the processing accuracy and the processing speed, and the scene detection model is generated by training sample materials using a convolutional neural network model; selecting a plurality of target shots from the set of candidate shots based on the semantics of the product features and the perceived persuasiveness of the visual materials; performing shot sorting on the plurality of target shots to generate a video for recommending the product; outputting the video; wherein the selecting the plurality of target shots from the set of candidate shots based on the semantics of the product features and the perceived persuasiveness of the visual materials comprises: obtaining the semantic distance between a plurality of the visual materials based on the semantics of the product features and the product title of the corresponding product; and selecting the plurality of target shots based on the semantic distance between a plurality of the visual materials, the perceived persuasiveness of a plurality of the visual materials, and the scene categories of a plurality of the visual materials.