Video generation method, electronic equipment, chip system and storage medium
By receiving instructions on time and event elements, obtaining memory nodes, and generating videos using templates, the problem of lack of narrative linearity and coherence in existing videos is solved, thus improving the user experience.
Patent Information
- Application Number
- CN202411050087.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2044-07-31
AI Technical Summary
Existing electronic devices lack narrative linearity and coherence when generating videos, resulting in a poor user experience.
By receiving instructions on time and event elements, memory nodes are obtained and videos are generated using templates. Media materials that match the target tags are selected, and memory nodes are clustered to form a storyline.
The generated videos have a clear narrative line and logical structure, enhancing the user experience.
Smart Images

Figure CN121509700A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic equipment technology, and in particular to a video generation method, electronic equipment, chip system, and storage medium. Background Technology
[0002] Many electronic devices are equipped with cameras, allowing users to take photos and videos. These devices also have communication functions, enabling users to download images or videos from the internet or receive images or videos sent by other electronic devices.
[0003] Users can trigger electronic devices to generate videos by selecting media materials that match their intentions based on voice commands. However, videos generated in this way tend to be disorganized and result in a poor user experience. Summary of the Invention
[0004] This application provides a control method for an electronic device, an electronic device, a chip system, and a storage medium, which can generate videos with a storyline to improve the user experience.
[0005] To achieve the above objectives, the first aspect of this application adopts the following technical solution:
[0006] The first aspect of this application provides a video generation method, including:
[0007] Receive a first instruction, which includes a time element and an event element;
[0008] Based on the time element in the first instruction, a memory node with a start time within the range of the time element is obtained. The memory node is generated by one or more related media materials. The memory node includes a start time and an event. The event is related to the content of the multiple media materials that generated the memory node.
[0009] Based on the event elements in the first instruction, a finished product template is obtained, wherein the finished product template includes multiple themes and multiple tags corresponding to each theme;
[0010] For the first memory node, media materials whose image content contains the target tag of the first memory node are selected from the media materials that generated the first memory node. The first memory node is any memory node within the time element range, and the target tag of the first memory node is determined by the event of the first memory node, the theme and tag in the final template.
[0011] A video is generated based on media material whose selected image content contains the target tag of the first memory node.
[0012] In this application, when selecting media materials for generating videos, on the one hand, they are obtained based on memory nodes, which are generated from related media materials, so the final video has story relevance; on the other hand, memory nodes for generating videos are obtained through finished product templates; and finished product templates are pre-set templates related to the event elements in the instructions, so the final video has a story that matches the finished product template; furthermore, the media materials selected for each memory node are also media materials corresponding to tags related to the theme of that memory node, making the story relevance of the resulting video stronger and more organized.
[0013] As one implementation of this application, the events of the memory node include one or more themes in the piece template;
[0014] Before selecting media materials whose image content contains the target tag of the first memory node from the media materials from which the first memory node was generated, the method further includes:
[0015] Based on the events of the first memory node, the topic of the first memory node is obtained;
[0016] The multiple tags corresponding to the theme of the first memory node in the finished template are used as the target tags of the first memory node.
[0017] In this application, since memory nodes are generated from multiple related media materials, a memory node may include media materials from multiple events; in order to make the specific storyline of the final media material coherent, a memory node retains a theme.
[0018] As another implementation of this application, obtaining the topic of the first memory node based on the events of the first memory node includes:
[0019] When there is only one event in the first memory node, the event of the first memory node is determined as the topic of the first memory node;
[0020] When there are multiple events for the first memory node, the event with the highest priority among the multiple themes set in the image template is determined as the theme of the first memory node.
[0021] In this application, the theme of the memory node is matched as closely as possible to the theme in the finished video template, so that the final video matches the user's needs.
[0022] As another implementation of this application, obtaining the topic of the first memory node based on the events of the first memory node includes:
[0023] Generate a storyline from multiple memory nodes related to the first memory node;
[0024] The events that generate memory nodes in the same storyline are clustered to obtain the theme of the storyline, and the theme of each memory node in the storyline is the theme of the storyline.
[0025] This application can also generate storylines based on memory nodes, thereby enabling the clustering of multiple memory nodes with the same theme, so that the theme of the final video is more in line with the user's intent.
[0026] As another implementation of this application, the step of clustering events that generate memory nodes of the same storyline to obtain the theme of the storyline includes:
[0027] Based on the priority of multiple themes set in the finished template, the event with the highest priority is selected from the events that generate the memory nodes of the same storyline as the theme of the storyline.
[0028] As another implementation of this application, the step of generating a video based on media material containing the target tag of the first memory node in the selected image content includes:
[0029] The selected media materials containing the target tag of the first memory node are filtered to obtain the filtered media materials;
[0030] Generate a video based on the selected media materials.
[0031] As another implementation of this application, the selected media materials whose image content contains the target tag of the first memory node are filtered to obtain the filtered media materials, which also include:
[0032] The selected media materials whose image content contains the target tag of the first memory node are subjected to a first filtering process to obtain the media materials after the first filtering process. The first filtering process includes at least one of the following: retaining one of two or more media materials whose image content repetition is greater than a first threshold, removing media materials whose image content contains preset content, and removing media materials whose image clarity is less than a second threshold.
[0033] In this application, in order to ensure the quality of the final product, some media materials that do not produce good results may be deleted.
[0034] As another implementation of this application, the step of filtering media materials whose selected image content contains the target tag of the first memory node to obtain filtered media materials further includes:
[0035] Divide the time element range into multiple sub-time ranges;
[0036] For the media materials after the first screening process, among the memory nodes with the same theme within the same sub-time range, the media materials with one memory node are retained, and the media materials with other memory nodes are deleted to obtain the media materials after the second screening process. The other memory nodes are the memory nodes other than the retained memory nodes among the memory nodes with the same theme within the same sub-time range.
[0037] To avoid the final product being too focused on one theme, only one memory node with the same theme will be retained.
[0038] As another implementation of this application, the memory node further includes: popular tags; the media material for retaining one memory node from memory nodes with the same theme within the same sub-time range includes:
[0039] Obtain the first number of first tags contained in the popular tags of each memory node with the same theme, wherein the same theme is the first theme and the first tags are the tags of the first theme in the template;
[0040] Sort by the first number of the first tags contained in the popular tags of each memory node with the same topic;
[0041] If the two largest first quantities are not equal, then the media material of the memory node corresponding to the largest first quantity is retained;
[0042] If the two largest first quantities are equal, then the memory nodes with the largest and equal first quantities are retained;
[0043] Calculate the median score of the most recently retained memory node, where the median score of the memory node is the median score of the media material after the first filtering process that generated the memory node;
[0044] Sort the most recently retained memory nodes by their median scores;
[0045] If the two largest median ratings are not equal, then the media material corresponding to the memory node with the largest median rating is retained;
[0046] If the two largest median scores are equal, then the memory nodes with the largest and equal median scores are retained.
[0047] Calculate the average score of the most recently retained memory nodes, where the average score of the memory nodes is the average score of the media materials after the first filtering process that generated the memory nodes;
[0048] Sort the most recently retained memory nodes by their average scores;
[0049] If the two largest average ratings are not equal, then the media material corresponding to the memory node with the largest average rating will be retained.
[0050] If the two largest average scores are equal, then the memory nodes with the largest and equal average scores are retained.
[0051] Get the total number of media materials after the first filtering process of the latest retained memory node;
[0052] Sort the total number of media materials after the first screening process of the latest retained memory nodes;
[0053] If the total number of the two largest media assets is not equal, then the media asset of the memory node corresponding to the largest total number of media assets is retained;
[0054] If the two largest media assets have the same total number, then the memory nodes with the largest and equal total number of media assets will be retained.
[0055] Select one memory node from the latest retained memory nodes to retain, and obtain the media material after the first filtering of the retained memory nodes.
[0056] In this application, one of the memory nodes is selected and retained in the following order: number of tags contained; median rating, average rating, total number of media materials contained, start time, etc.
[0057] As another implementation of this application, the step of filtering media materials whose selected image content contains the target tag of the first memory node to obtain filtered media materials further includes:
[0058] The media materials after the second screening process are subjected to time hashing to obtain the number of media materials in each sub-time range;
[0059] Based on the preset total number of media materials for the generated video, and the proportion of the number of media materials in each sub-time range to the number of media materials after the second filtering process, the number of media materials to be extracted from each sub-time range is obtained.
[0060] Based on the number of media materials to be extracted from each sub-time range, media materials are extracted from the media materials in each sub-time range to obtain the media materials after the third filtering process.
[0061] In this application, in order to make the media materials more dispersed within the time range required by the user for the final cut, a certain number of media materials can be extracted from each sub-time period.
[0062] As another implementation of this application, the step of extracting media materials from the media materials in each sub-time range based on the number of media materials to be extracted from each sub-time range to obtain the media materials after the third filtering process includes:
[0063] Determine the maximum number of topics to extract from the media materials in each sub-time range based on the number of sub-time ranges with media materials.
[0064] For the first sub-time range, media materials are extracted from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, resulting in filtered media materials. The first sub-time range is any sub-time range within the time element range.
[0065] When a sub-time range contains memory nodes for multiple themes, in order to avoid the videos becoming disorganized, a certain number of memory nodes for each theme can be extracted from each month.
[0066] As another implementation of this application, the step of extracting media materials from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, to obtain the filtered media materials, includes:
[0067] Determine whether the number of topics in the first sub-time range is greater than 1;
[0068] If the number of topics in the first sub-time range is not greater than 1, then media materials are extracted from the first sub-time range according to the number of media materials to be extracted from the first sub-time range.
[0069] If the number of topics in the first sub-time range is greater than 1, then the topic with the highest priority is selected from the topics in the first sub-time range.
[0070] Determine whether the number of media materials corresponding to the highest priority topic is greater than or equal to the number of media materials to be extracted from the first sub-time range;
[0071] If the number of media materials corresponding to the highest priority theme is greater than or equal to the number of media materials to be extracted from the first sub-time range, then media materials are extracted from the media materials corresponding to the highest priority theme according to the number of media materials to be extracted from the first sub-time range.
[0072] If the number of media materials corresponding to the highest priority topic is less than the number of media materials to be extracted from the first sub-time range, then all media materials corresponding to the highest priority topic will be extracted.
[0073] If the maximum number of topics is equal to 1, then media materials are extracted from the first sub-time range according to topic priority.
[0074] If the maximum number of topics is greater than 1, then the second highest priority topic is obtained from the topics in the first sub-time range; and media materials are extracted from the media materials corresponding to the second highest priority topic until the media materials corresponding to the p-th highest priority topic are extracted, where p represents the maximum number of topics; after the media materials corresponding to the p-th highest priority topic are extracted, the media materials are extracted from the first sub-time range according to topic priority.
[0075] In this application, media materials are extracted from various topics in order of priority to generate videos that meet the user's intent.
[0076] As another implementation of this application, after extracting media materials from the first sub-time range according to topic priority, the difference between the number of media materials to be extracted from the first sub-time range and the number of media materials already extracted is obtained.
[0077] After obtaining the difference for each sub-time range within the time element range, the sum of the differences for each sub-time range within the time element range is calculated;
[0078] From the media materials that were not extracted within the time element range, continue to extract media materials of the sum of the differences based on the media material scores.
[0079] If the number of extracted media materials is insufficient to produce a complete film, it is necessary to continue extracting from the unextracted media materials based on the scores, in order to complete the film.
[0080] As another implementation of this application, the step of extracting media materials from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, to obtain the filtered media materials, includes:
[0081] Determine whether the number of topics in the memory nodes within the first sub-time range is greater than the maximum number of topics;
[0082] If the number of topics in the memory nodes within the first sub-time range is greater than the maximum number of topics, then the maximum number of topics is selected according to the topic priority set in the final template.
[0083] The number of media materials to be extracted from each theme is obtained by taking into account the proportion of the number of media materials corresponding to each selected theme to the total number of media materials corresponding to all selected themes, and the number of media materials to be extracted from the first sub-time range.
[0084] Based on the number of media materials extracted from each theme, media materials are extracted from the first sub-time period.
[0085] If the number of topics in the memory nodes within the first sub-time range is less than or equal to the maximum number of topics, then the number of media materials extracted from each topic is obtained based on the ratio of the number of media materials corresponding to each topic within the first sub-time range to the total number of media materials within the first sub-time range, and the number of media materials to be extracted from the first sub-time range.
[0086] Media materials are extracted from the first sub-time frame based on the number of media materials extracted from each theme.
[0087] In this application, media materials can also be extracted proportionally based on the quantity of media materials under each theme to generate videos.
[0088] As another implementation of this application, the scoring of the media material includes at least one of the following dimensions: portrait dimension, picture quality dimension, aspect ratio dimension, and user behavior dimension;
[0089] The portrait dimension is related to the proportion of the face area in the media material to the entire frame;
[0090] The image quality dimension is related to the clarity of the media material and the jitter when the media material is video;
[0091] The aspect ratio is related to the percentage of the filled area relative to the standard ratio when the aspect ratio of the media material is filled to the standard ratio;
[0092] The user behavior dimension is related to the number of times the media material is collected, shared, and viewed by users.
[0093] In this application, media materials can be scored from multiple dimensions, resulting in a better final product.
[0094] In a second aspect, an electronic device is provided, including a processor for calling a computer program stored in a memory to implement the method of any one of the first aspects of this application.
[0095] Thirdly, a chip system is provided, including a processor coupled to a memory, wherein the processor executes a computer program stored in the memory to cause an electronic device to implement the method of any one of the first aspects of this application.
[0096] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when computer instructions are executed on an electronic device, causes the electronic device to implement the method of any one of the first aspects of this application.
[0097] Fifthly, embodiments of this application provide a computer program product that, when run on a device, causes the electronic device to execute the method of any one of the first aspects of this application.
[0098] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0099] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0100] Figure 2 This application provides a schematic diagram of an interface for waking up a voice assistant.
[0101] Figure 3 A schematic diagram of an interface for generating video via a voice assistant, provided as an embodiment of this application;
[0102] Figure 4 A timing diagram for obtaining media materials for generating video based on user voice, provided in an embodiment of this application;
[0103] Figure 5 Provided for the embodiments of this application Figure 4 A flowchart illustrating the process of extracting media materials in S28;
[0104] Figure 6 Provided for the embodiments of this application Figure 5 A flowchart illustrating the process of deleting memory nodes for duplicate event topics within the same month in section S102;
[0105] Figure 7 A schematic diagram illustrating the event themes and quantities of media materials distributed across different months, as provided in an embodiment of this application;
[0106] Figure 8 A flowchart illustrating a strategy for extracting media materials from each month, provided as an embodiment of this application;
[0107] Figure 9 A flowchart illustrating the process of determining the number of event topics to be extracted when extracting media materials each month, as provided in an embodiment of this application;
[0108] Figure 10This application provides a timing diagram for generating and displaying video based on extracted media materials. Detailed Implementation
[0109] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limiting purposes, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details.
[0110] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0111] It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more; "and / or" describes the relationship between the associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.
[0112] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," "fourth," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0113] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0114] This application provides a video generation method that can be applied to electronic devices. These electronic devices can be tablets, mobile phones, wearable devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. This application does not limit the specific type of electronic device.
[0115] Figure 1 A schematic diagram of an electronic device is shown. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0116] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0117] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0118] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0119] Internal memory 121 can be used to store computer executable program code, including instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as image playback). Touch sensor 180K, also called a "touch panel," can be disposed on display screen 194. Touch sensor 180K and display screen 194 together form a touch screen, also called a "touch screen." Touch sensor 180K is used to detect touch operations applied to or near it. Touch sensor can transmit the detected touch operation to application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be disposed on the surface of electronic device 100, in a different location than display screen 194.
[0120] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0121] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized display, a microLED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0122] This application does not specifically limit the structure of the execution entity of a video generation method. As long as the code recording the video generation method of this application is executed, communication can be performed according to the video generation method provided in this application. For example, the execution entity of the video generation method provided in this application can be a functional module in an electronic device capable of calling and executing a program, or a communication device applied in an electronic device, such as a chip.
[0123] Media stored in electronic devices include images and videos. Images can be photos taken by the user using the device's camera, images received from other electronic devices, or screenshots of content displayed on the device. Videos can be videos taken by the user using the device's camera, videos received from other electronic devices, or screen recordings of content displayed on the device.
[0124] Electronic devices can generate a memory node based on one or more related media materials. The media materials that generate the memory node have relationships, such as causal relationships or chronological order. For example, a memory node can be generated based on time and location, connecting media materials located in the same location within a few days. In practical applications, multiple conditions can be set; for example, the electronic device might only execute the memory node generation step if it is at night, the screen is off, or a certain number of media materials have not yet generated a memory node.
[0125] As an example of generating memory nodes, imagine a user attending a wedding in a certain location and planning to stay for three days. On the first day, they attend the wedding; on the second day, they visit a local art exhibition; and on the third day, they visit the local beach. Photos or videos taken on the first day, the second day, and the third day can all generate a memory node. Each memory node records the following information: Memory Node ID, start time, end time, city, event, and toptags.
[0126] Table 1. Memory nodes generated based on media materials
[0127]
[0128] Referring to Table 1, an example of multiple memory nodes generated from media footage is provided. The memory node ID serves as the identifier for this memory node, distinguishing it from other memory nodes. The start time is the shooting time of the earliest-shot media footage among those that generated this memory node. The end time is the shooting time of the latest-shot media footage among those that generated this memory node. The city is the shooting location of the media footage that generated this memory node.
[0129] Among them, event and toptags can be obtained by semantic analysis of the media materials that generate this memory node. event represents the major category of the content of these media materials, and toptags represent the minor categories of the content of these media materials. The minor category can be understood as a more detailed branch under the major category.
[0130] In practical implementation, multiple major categories and multiple subcategories under each major category need to be pre-defined. A semantic model is used to identify these major and subcategories of media materials. Thus, when generating memory nodes based on these media materials, the event is obtained based on the major category of the media material generating each memory node, and toptags are generated based on the subcategories of the media material generating each memory node. The media material is the media material that generated this memory node. In practical applications, this information can be a unique identifier for each media material, such as its name, content, or hash value.
[0131] Furthermore, in practical applications, storylines can be generated based on multiple memory nodes. For example, in the example above, three memory nodes generated in a certain location over three days can generate another storyline. This storyline can have a single event theme. Multiple conditions can also be set when generating a storyline, such as the location being the same or similar, the time being within a certain duration (e.g., within 3 days, within 7 days), and memory nodes containing the same or similar events or toptags.
[0132] Electronic devices can generate memory nodes based on stored media materials at preset time intervals; they can also generate memory nodes based on stored media materials that have not yet generated memory nodes at preset time intervals.
[0133] Of course, media footage that has already generated a memory node in the previous period (e.g., media footage shot on the last day of the previous period) and media footage that has not yet generated a memory node in the current period (e.g., media footage shot on the first day of the current period) may also generate a memory node. Therefore, in each period, the media footage corresponding to the latest start time of the memory node in the previous period can also be used as media footage that has not generated a memory node, along with the newly added media footage that has not yet generated a memory node in the current period, to form a new set of media footage. One or more memory nodes can be generated based on this new set of media footage, overwriting the memory node with the latest start time of the previous period.
[0134] The above-described method for generating memory nodes is for illustrative purposes only. This application does not impose any restrictions on the time, method, or rules for generating memory nodes.
[0135] Electronic devices can generate videos by selecting media materials that match the user's intent from the media materials stored in the device based on the user's voice information. Due to the time, location, and event relevance of memory nodes, media materials that match the user's intent can be obtained based on memory nodes, and thus a video with a storyline and in accordance with the user's intent can be generated based on these media materials.
[0136] Reference Figure 2 The image shown is a schematic diagram of a scenario for waking up a voice assistant, provided in an embodiment of this application.
[0137] Reference Figure 2 In (a), the electronic device display interface 201 can be any interface displayed by the electronic device, such as the system desktop, an interface of any application, etc. This application embodiment uses the system desktop as an example.
[0138] Reference Figure 2 In (b) of the above, during the display interface 201 of the electronic device, if the user says the wake-up word "Hello yoyo", the electronic device detects the voice information and determines that the voice information contains the wake-up word, then displays interface 202. Interface 202 displays the voice assistant window 2021 based on interface 201. The voice assistant window 2021 is used to remind the user that the voice assistant has been activated, and the user can continue to say relevant voice commands.
[0139] Reference Figure 3 The image shown is a schematic diagram of an interface for collecting media materials and generating videos based on user voice commands, provided in an embodiment of this application.
[0140] Reference Figure 3 In (a), the user continues with the voice command "Help me generate a personal annual summary video for 2023." After detecting the voice information, the electronic device, on the one hand, obtains the corresponding text information based on the voice information and displays the text information 2031 corresponding to the voice information in the voice assistant window 2021. The text information 2031 is "Help me generate a personal annual summary video for 2023." On the other hand, it obtains semantic information based on the voice information, searches for media materials related to the 2023 personal annual summary in the media material library based on the semantic information, and displays the searched media materials 2032 in the voice assistant window 2021. While displaying the media materials 2032, the voice assistant window 2021 also displays a control 2033, which is used to generate a video based on the searched media materials. Figure 3 The media material shown in (a) can be obtained by searching based on relevant information (e.g., event and toptags information) in the memory nodes in the above embodiments.
[0141] Figure 3 The content in the voice assistant window in interface 203 and interface 202 shown in (a) is different.
[0142] Reference Figure 3 In (b), the user can click on the control 2033 or say "Generate video" to trigger the electronic device to generate a video based on the displayed media material and display the content 2051 in the voice assistant window 2021.
[0143] Content 2051 includes the cover image of the generated video and the duration of the video. Figure 3 The content in the voice assistant window in interfaces 204, 203, and 202 is different.
[0144] Of course, the voice assistant in the above example is just one example. In actual applications, the voice assistant can be any application with voice recognition capabilities.
[0145] To better understand the implementation process of the above scenario, the following will explain... Figure 4 This describes the process of extracting media materials based on memory nodes. This process corresponds to the above. Figures 2 to 3 The interface shown in (a) is shown in the image.
[0146] S01, the voice assistant detected the user's voice "Hello, yoyo".
[0147] In this embodiment of the application, when the electronic device displays any interface, the user says the wake word "Hello, yoyo" for the voice assistant application. "Hello, yoyo" is an example of a wake word for a voice assistant application; in practical applications, other wake words may also be used.
[0148] S02, the voice assistant determines that the detected user voice is the wake word and displays the voice assistant window.
[0149] In this embodiment, after detecting a user's voice, the voice assistant can determine whether the current user's voice is the voice assistant's wake-up word through semantic analysis and other recognition methods. If it is the voice assistant's wake-up word, the voice assistant window is displayed.
[0150] S11, the voice assistant detected the user's voice message "Help me generate a personal annual summary video for 2023".
[0151] In this embodiment of the application, after the user wakes up the voice assistant, the user can continue to output the voice information "Help me generate my personal annual summary video for 2023".
[0152] In practical applications, other related videos can also be generated, such as personal growth videos, children's growth videos, and personal annual travel videos. Correspondingly, users need to speak audio related to the video to be generated.
[0153] As an example, if a user says, "Generate a personal travel video for me in 2023," the electronic device will generate a travel video of the user in 2023. If a user says, "Generate a personal growth video for my daughter," the electronic device will generate a growth video of the user's daughter. Other examples are not listed here.
[0154] Additionally, if the user's spoken voice message is "Help me generate a personal annual summary video," meaning the user's voice message does not explicitly specify an event, then the time element can be one year prior to the current time. For example, if the current year is 2024, then the time element would be 2023. Alternatively, the time element can be set to 12 months prior to the current month. For example, if the current year is July 2024, then the time element would be from June 2023 to June 2024. Of course, the user can also explicitly specify a personal summary video from a specific year and month to a specific year and month. This application does not impose any restrictions on this.
[0155] S12: After detecting the user's voice, the voice assistant generates text information based on the user's voice.
[0156] S13, the voice assistant inputs text information into the large language model.
[0157] S14. After receiving text information, the large language model identifies the semantics of the text information and obtains semantic elements.
[0158] In this embodiment of the application, semantic elements may include time elements, person elements, location elements, event elements, etc.
[0159] As an example, the time element is set to 2023, the person element is set to the local user, the location element is empty, and the event element is "Annual Summary". In practical applications, some element values can be empty.
[0160] In practical applications, recognizing the semantics of text information can yield either semantic elements or user-defined semantics. Of course, both can also be obtained simultaneously. This application takes obtaining semantic elements as an example.
[0161] S15, the large language model sends the obtained semantic elements to the voice assistant.
[0162] S16, after receiving the semantic elements, the voice assistant sends a material search request to the media processing platform, which carries the semantic elements.
[0163] S17, the media processing platform determines whether there is preset content in the semantic elements.
[0164] In this embodiment, "summary" can be set as a preset content for whitelist judgment. When the preset content exists, the video generation process based on memory nodes provided in this embodiment can be executed.
[0165] S18, when the media processing platform determines that the semantic elements carry preset content, it sends a memory node search request to the learning platform. The memory node search request carries a time element (e.g., 2023).
[0166] S19, the learning platform searches for memory nodes whose start time is within 2023 from all memory nodes based on the time element in the memory node search request.
[0167] S20, the learning platform sends the memory nodes found within 2023 to the media processing platform.
[0168] S21, the media processing platform determines the event topic for each memory node based on the events in the memory nodes within 2023 and the topic priority in the template.
[0169] In this embodiment of the application, when there are multiple templates in the media processing platform, the media processing platform can determine the currently used personal annual summary template based on semantic elements.
[0170] The personal annual summary template includes categories, themes, and tags. These themes have different priorities.
[0171] Table 2 Personal Annual Summary Template
[0172]
[0173] In Table 2, the priority of topics decreases from top to bottom. The topics and tags in the personal annual summary template shown in Table 2 were obtained by the user based on the major and minor categories in the semantic model of the above embodiments.
[0174] As an example, developers first set some categories, and then set the priority of these categories; developers can select a portion of the content from the major and minor categories that the semantic model can recognize as the themes and tags in the personal annual summary template; after selecting the themes and tags under each category, the personal annual summary template is obtained.
[0175] In practical applications, users can generate other templates in the same way as described above, such as personal annual travel templates, children's growth video templates, etc. This application will not list them all.
[0176] In addition, as shown in Table 1, each memory node's event includes one or more items; based on the priority of the topics in the personal annual summary template, the item with the highest priority from the various items of the event in the memory node can be selected as the event topic of that memory node.
[0177] As another example of obtaining the event theme for each memory node, as mentioned earlier, multiple adjacent memory nodes can also be grouped into a storyline based on certain principles (e.g., same location or proximity). For example, memory nodes 5 through 10 can be grouped into a storyline, and the event content in memory nodes 5 through 10 can be clustered into an event theme. When the events in memory nodes 5 through 10 are aggregated, the event with the highest priority (determined from the template) is taken as the event theme for each memory node in the storyline.
[0178] For example, query the contents of the `event` field in all memory nodes that generate a storyline, and determine if it contains the following: "Food", "Travel", "Scenery", and "Sports". Then, based on the priority of themes in the pre-set template, determine the highest priority play as the event theme.
[0179] S22, After determining the event topic of each memory node, select the tags belonging to the event topic in the template from the toptags of the memory node, and use them as the tags of the memory node.
[0180] It should be noted that the event topic in the template has multiple tags, while the toptags of the memory node are generated by the semantic model based on the image content. This means that the toptags of the memory node may only contain one or more (or may not contain) tags under a specific event topic. Of course, the memory node itself may include multiple events, so the toptags of the memory node may contain tags other than those under a specific event topic.
[0181] For example, a memory node includes eventA and eventB, and toptags1, toptags2, toptags3, and toptags4. Toptags1 belongs to the toptags under eventA, while toptags2, toptags3, and toptags4 belong to the toptags under eventB. If the event topic of the memory node is determined to be eventB, then the determined tags for this memory node are toptags2, toptags3, and toptags4 from the multiple tags under the topic (eventB) in the template. The determined toptags for this memory node do not include toptags1 under eventA.
[0182] As another example, when this memory node and other memory nodes generate a storyline, and the determined event theme is eventC, then this memory node does not include eventC, and the toptags of this memory node may not include multiple tags under eventC, so this memory node has no tags.
[0183] S23, the media processing platform sends a material search request to the material search module. The material search request carries the memory nodes within 2023 and the tag of each memory node.
[0184] In practical implementation, memory nodes with tags can be sent.
[0185] S24, after receiving a material search request, the material search module uses the clip model to search for media materials whose image content contains any tag of the memory node from among the multiple media materials that generated the memory node.
[0186] In this embodiment, the image content (text content) of each media material that generates a memory node can be obtained through the clip model; the media material whose image content contains the tag of the memory node is used as the media material for generating the video.
[0187] The clip model can learn the correspondence between images and text. By inputting media material into the clip model, at least one text that can represent the image content in the media material can be obtained. In other words, the clip model can obtain the text corresponding to each media material.
[0188] S25, the media search module sends the found media materials and the memory nodes where these media materials are located to the media processing platform.
[0189] S26. After receiving media materials, the media processing platform conducts an initial screening.
[0190] In this embodiment of the application, the preliminary screening operation includes deduplication, denegation, and removal of low-quality media materials.
[0191] Deduplication means removing duplicate media materials. For example, if the content similarity between two photos is greater than 95%, one of the photos can be deleted from the media material set.
[0192] "Denegative removal" refers to the deletion of media materials containing specific negative content. For example, some media materials contain artificially labeled tags, some contain content with inauspicious connotations (such as graves), and some contain meaningless content such as keyboards or computer screens. These types of media materials can be deleted.
[0193] Removing low-quality content means deleting photos or videos with a resolution below a certain threshold from the media archive. For example, you can calculate the blur value of the media archive and delete those with a blur value higher than the threshold.
[0194] Of course, in practical applications, more conditions can be set for initial screening. For example, some media materials that do not conform to a specific aspect ratio can be deleted, while media materials that conform to a specific aspect ratio (e.g., 16:9 or 9:16) can be kept.
[0195] S27, The media processing platform calculates the score of the media materials after the initial screening.
[0196] In this embodiment of the application, the media materials after initial screening can be scored from multiple dimensions, specifically from the following dimensions:
[0197] (1) From the perspective of the human face, the position score is determined based on the distance of the human face from the center of the image. The smaller the distance, the higher the score, and the farther the distance, the lower the score. The proportion score is determined based on the proportion of the human face to the entire image. If the proportion is too large or too small, the score is lower. The closer the proportion is to a certain value or range, the higher the score. The model determines the aesthetic score based on the composition of the human face. The position score, proportion score, and aesthetic score are weighted and summed to obtain the human face dimension score.
[0198] For media materials that do not contain human figures, this score can be set as a base score. This base score can be a fixed value or the average score of the human figure dimension of media materials that contain human figures.
[0199] (2) From the perspective of image quality, if the media material is an image, the image sharpness is calculated, and a sharpness score is determined based on the image sharpness. The higher the sharpness, the higher the score; the lower the sharpness, the lower the score. The composition score of the image also needs to be calculated (which can be calculated using a model). Finally, the image quality score is obtained by weighting the sharpness score and the composition score. If the media material is a video, the video sharpness is calculated, and a sharpness score is determined based on the video sharpness. The higher the sharpness, the higher the score; the lower the sharpness, the lower the score. The video jitter also needs to be calculated, and a jitter score is obtained based on the jitter. The greater the jitter, the lower the score; the smaller the jitter, the higher the score. Finally, the image quality score is obtained by weighting the sharpness score and the jitter score.
[0200] (3) From the aspect ratio perspective, if media materials that do not conform to a specific aspect ratio (e.g., 16:9 or 9:16) are not deleted in the initial screening, then the aspect ratio score of these media materials is calculated in this step.
[0201] The standard aspect ratio can be set to 16:9 (length:width). If the aspect ratio of the media material is greater than 16:9, the length of the media material will be set to the standard aspect ratio, and the aspect ratio of the media material will be proportionally reduced or enlarged; the percentage of the media material's aspect ratio to the standard aspect ratio will be calculated. If the aspect ratio of the media material is less than 16:9, the width of the media material will be set to the standard aspect ratio, and the aspect ratio of the media material will be proportionally reduced or enlarged; the percentage of the media material's aspect ratio to the standard aspect ratio will be calculated; if the aspect ratio of the media material is equal to 16:9, then the media material's aspect ratio will occupy 100% of the standard aspect ratio.
[0202] The aspect ratio score can be set to full when the media footage occupies 100% of the standard aspect ratio. The higher the percentage of the media footage in the standard aspect ratio, the higher the aspect ratio score. The lower the percentage of the media footage in the standard aspect ratio, the lower the aspect ratio score. Of course, when the percentage of the media footage in the standard aspect ratio is less than a certain value, this score can be set to 0.
[0203] (4) From the user behavior dimension, a user behavior score is obtained based on the number of times media materials are collected, shared, and viewed. For example, the base score can be set to 0 points, with the first score added for media materials that have been collected, the second score added for media materials that have been shared once (the more shares, the more points are added), and the third score added for media materials that have been viewed once (the more views, the more points are added). An upper limit can be set for the score of this dimension, for example, the score of this dimension cannot exceed 20 points.
[0204] Finally, the scores of each media material across the four dimensions are weighted and summed to obtain the total score for that media material.
[0205] S28, the media processing platform extracts a portion of media materials from the initially screened materials based on the memory node where the media materials are located and the scores of the media materials, and generates a finished product ID for the extracted media materials.
[0206] In this embodiment, media materials can be extracted according to the number extracted each month. Of course, in practical applications, step S28 can also be executed according to... Figure 5 The extraction is performed in the manner shown, that is, S28 includes steps S101 to S104.
[0207] S101, based on the start time of the memory node where the media material is located, perform time hashing on the media material.
[0208] In this step, the media material is first scattered to various months, that is, the month in which the media material is located is determined according to the start time of the memory node in which the media material is located.
[0209] S102, delete the memory nodes of repeated event themes within the same month, so that media materials with the same event theme within the same month have one memory node.
[0210] Delete media materials corresponding to memory nodes with the same event theme in the same month, so that media materials with the same event theme in the same month have only one memory node.
[0211] In practical applications, when generating videos from annual summaries, it's possible to retain only one memory node within a month that shares the same event theme. When there are at least two memory nodes within a month sharing the same event theme, they can be arranged according to... Figure 6 The method shown selects one of the memory nodes as the memory node for generating the video.
[0212] A1, retrieve the number of tag types under the same event topic contained in each memory node with the same event topic;
[0213] This application uses memory node 1 and memory node 2 as examples of memory nodes with the same event theme. In practical applications, there may be more memory nodes with the same event theme. The event theme of memory node 1 is tourism, and the event theme of memory node 2 is also tourism. Memory node 1 contains m types of tags under the tourism event theme, and memory node 2 contains n types of tags under the tourism event theme.
[0214] As an example, memory node 1 contains three tags under the theme of travel events in the life summary template: scenery, performance, and animals; memory node 2 contains five tags under the theme of travel events in the life summary template: luxury cars and cruises, amusement parks, scenery, pets, and food.
[0215] A2, determine whether the number of tag types under the event topic contained in memory node 1 is equal to the number of tag types under the event topic contained in memory node 2.
[0216] A3. If the number of tag types under the event theme contained in memory node 1 is not equal to the number of tag types under the event theme contained in memory node 2, the memory node containing the most tag types under the event theme shall be used as the memory node for generating the video.
[0217] Of course, if there are 3 or more memory nodes under the same event theme, and the two memory nodes with the most tag types under the same event theme have different numbers of memory nodes, then the memory node with the most tag types under the same event theme will be used as the memory node for generating the video.
[0218] A4. If the number of tag types under the same topic contained in memory node 1 is equal to the number of tag types under the same topic contained in memory node 2, calculate the median rating of the media material that generated memory node 1 and the median rating of the media material that generated memory node 2.
[0219] If the two memory nodes with the most tag types under the event topic have the same number of tags, then calculate the median score corresponding to the multiple memory nodes with the most tags and the same number of tags.
[0220] As an example, the number of labels for the top three memory nodes with the largest number of labels may be as shown in Table 1.
[0221] Table 3 determines the execution process based on the number of tags in the memory nodes.
[0222]
[0223] In subsequent embodiments, although two memory nodes (memory node 1 and memory node 2) are used as examples, the execution process when multiple memory nodes exist can be determined based on the principles described above. Further examples will not be provided hereafter.
[0224] As mentioned earlier, media assets can be rated from multiple perspectives, such as portrait dimensions, image quality dimensions, aspect ratio dimensions, and user behavior dimensions. Multiple media assets for a single memory node correspond to multiple ratings, which can be sorted from high to low (or low to high). When the number of media assets generating a memory node is odd, the middle rating in the sorted list is selected as the median; when the number of media assets generating a memory node is even, the average of the two middle ratings in the sorted list is selected as the median.
[0225] A5. Determine whether the median rating of the median material that generated memory node 1 is equal to the median rating of the median material that generated memory node 2.
[0226] A6. If the median score of the media material used to generate memory node 1 is not equal to the median score of the media material used to generate memory node 2, the memory node with the highest median score shall be used as the memory node for generating the video.
[0227] A7. If the median rating of the median material generated for memory node 1 is equal to the median rating of the median material generated for memory node 2, calculate the average rating of the median material generated for memory node 1 and the average rating of the median material generated for memory node 2.
[0228] A8. Determine whether the average rating of the media material that generates memory node 1 is equal to the average rating of the media material that generates memory node 2.
[0229] A9. If the average score of the media material used to generate memory node 1 is not equal to the average score of the media material used to generate memory node 2, the memory node with the highest average score will be used as the memory node for generating the video.
[0230] A10, if the average score of the media material generated for memory node 1 is equal to the average score of the media material generated for memory node 2, obtain the total number of filtered materials in the media material generated for memory node 1 and the total number of filtered materials in the media material generated for memory node 2.
[0231] A11, determine whether the total number of filtered materials in the media materials of generated memory node 1 is equal to the total number of filtered materials in the media materials of generated memory node 2.
[0232] A12, if the total number of materials filtered from the media materials used to generate memory node 1 is not equal to the total number of materials filtered from the media materials used to generate memory node 2, the memory node with the highest total number of filtered materials shall be used as the memory node for generating the video.
[0233] A13, if the total number of materials filtered from the media materials used to generate memory node 1 is equal to the total number of materials filtered from the media materials used to generate memory node 2, obtain the start time of memory node 1 and the start time of memory node 2.
[0234] A14 uses the earliest start time memory node as the memory node for generating the video.
[0235] In this way, each month of the year can contain only one memory node for the same event theme as the basis for generating a video. Of course, each month of the year can include multiple memory nodes corresponding to different event themes.
[0236] S103. The number of media materials extracted from the media materials of the i-th month is obtained by multiplying the proportion of the number of media materials belonging to the i-th month to the total number of media materials in the media material set by the number of media materials in a block.
[0237] After steps A1 to A10, the resulting media material set contains media materials for the same event theme within the same month that contain only one memory node.
[0238] Going forward, we can continue to select media materials from the media materials corresponding to each month's memory nodes (which have already undergone chip model screening and initial screening) as the final media materials for the finished product, in the following manner:
[0239] The proportion of media materials belonging to the i-th month to the total number of media materials in the media material set (the total number of media materials in the media material set after step A10) is multiplied by the number of media materials in a block, and the result is rounded to obtain the number of media materials extracted from the media materials of the i-th month.
[0240] After determining the number of media materials to be extracted from the i-th month, the media materials are extracted by referring to the total score of each media material.
[0241] Reference Figure 7 As an example of calculating the number of media materials extracted from the media materials of the i-th month, there are 120 media materials filtered from the memory node. These 120 media materials involve 10 event themes and are distributed in January, February, May, June and July.
[0242] There are 30 media clips in January, covering two event themes. We need to select 30 clips from 120. First, we can calculate the proportion of January's media clips to the total (30 / 120). Then, multiply this proportion by the number of complete media clips (30 / 120*30), which is 7.5 clips. Rounding the result to 8, we get 8 clips. Therefore, we need to select 8 media clips from January.
[0243] There are 20 media clips for February, covering one event theme. We need to select 30 clips from these 120. First, we can calculate the proportion of February's clips to the total (20 / 120). Then, multiply this proportion by the number of complete clips (20 / 120*30), which equals 5. Rounding the result to 5, we have 5 media clips to be selected from February.
[0244] There are 40 media clips in May, covering 2 event themes. We need to select 30 clips from 120. First, we can calculate the proportion of May's media clips to the total (40 / 120). Then, multiply this proportion by the number of complete media clips (40 / 120*30), which equals 10. Rounding the result to 10, we can select 10 media clips from May.
[0245] There are 20 media clips in June, covering 2 event themes. We need to select 30 clips from 120. First, we can calculate the proportion of June's media clips to the total (20 / 120). Then, multiply this proportion by the number of complete media clips (20 / 120*30), which equals 5. Rounding the result to 5, we need to select 5 media clips from June.
[0246] There are 10 media clips for July, covering one event theme. We need to select 30 clips from 120 clips. First, we can calculate the proportion of July's clips to the total (10 / 120). Then, multiply this proportion by the number of complete clips (10 / 120*30), which equals 2.5 clips. Rounding the result to 3, we need to select 3 media clips from July.
[0247] S104. Select media materials based on the total score of each media material.
[0248] As an example of media material extracted as a reference for the overall score of various media materials.
[0249] This application provides two extraction strategies to extract a certain number of media materials from the media materials of the corresponding month.
[0250] Reference Figure 8 Strategy 1:
[0251] First, the electronic devices are configured with the priority of each event topic. As an example of priority, the order is: birthday > family gathering > travel > graduation ceremony > outing > wedding > exercise.
[0252] B1, determine whether the number of event topics in the i-th month is greater than 1.
[0253] B2. If the number of event topics in the i-th month is less than or equal to 1, then extract the corresponding number of media materials from the media materials of the i-th month.
[0254] When selecting media materials, priority should be given to those with high total scores. For example, a certain number of media materials can be selected based on the total score from high to low.
[0255] B3. If the number of event topics in the i-th month is greater than 1, then determine the event topic with the highest priority from the event topics in the i-th month.
[0256] B4. Determine if the number of materials corresponding to the highest priority event theme is greater than the number to be extracted.
[0257] B5. If the number of media materials corresponding to the highest priority event theme is greater than or equal to the number to be extracted, then extract the corresponding number of media materials from the media materials corresponding to the highest priority event theme.
[0258] When selecting media materials, priority should be given to those with high total scores. For example, a certain number of media materials can be selected based on the total score from high to low.
[0259] B6. If the number of media materials corresponding to a high-priority event topic is less than the number to be extracted, then all media materials will be extracted from the media materials corresponding to the high-priority event topic.
[0260] B7 records the first difference in media materials for the i-th month (the number of materials to be extracted in the current month minus the number of materials already extracted).
[0261] In practical applications, after step B7, media materials can continue to be extracted from the media materials corresponding to the next highest priority event topics. Specifically, after B7, the following steps are also included:
[0262] C1, determine the next highest priority event topic from the topics of the i-th month.
[0263] C2 determines whether the number of materials corresponding to the second highest priority event topic is greater than the first difference in the i-th month.
[0264] C3. If the number of materials corresponding to the second-highest priority event theme is greater than or equal to the first difference in the i-th month, then extract the media materials of the first difference from the media materials corresponding to the second-highest priority event theme.
[0265] During the selection process, media materials with high overall scores are prioritized.
[0266] C4. If the number of materials corresponding to the second-highest priority event theme is less than the first difference in the i-th month, then all media materials are extracted from the media materials corresponding to the second-highest priority event theme.
[0267] C5 records the second difference in media materials for the i-th month (the number to be extracted minus the number already extracted).
[0268] Of course, in practical applications, the number of event topics for media materials that can be extracted from a month can be set. For example, the process shown in B1 to B7 allows media materials for one event topic to be extracted from a month; after B7, the execution flow continues from C1 to C5, which allows media materials for two event topics to be extracted from a month. The process from C1 to C5 is similar to the process from B3 to B7.
[0269] Similarly, when the number of event topics for media materials extracted from a month is greater than the number allowed, the number of materials corresponding to the second-highest priority event topic can be determined, and so on. This is equivalent to looping through B3 to B7 or C1 to C5, only with the priority decreasing by one level with each loop. Examples will not be provided in this application.
[0270] Regardless of the method used, if it is finally determined that there is a difference in the i-th month (the number to be extracted minus the number already extracted), then the differences of each month are summed to obtain the total difference.
[0271] From the remaining unselected media materials throughout the year, continue selecting the remaining media materials according to the priority of the event themes, based on the total difference in media materials.
[0272] As mentioned above, the number of event themes for which media materials can be extracted from a month may be one or more. This application embodiment also sets the following rules to determine the maximum number of event themes of media materials that can be extracted from the current month, as detailed below. Figure 9 The flowchart shown is shown.
[0273] D1 determines whether the number of months to which the media material belongs is greater than or equal to 6, that is, to determine how many months have media material.
[0274] D2. If the number of months to which the media material belongs is greater than or equal to 6, then each month is allowed to select media material from the highest priority event theme at most.
[0275] D3. If the number of months to which the media material belongs is less than 6, continue to determine whether the number of months to which the media material belongs is greater than or equal to 3.
[0276] D4. If the number of months to which the media material belongs is less than 6 and greater than or equal to 3, then each month is allowed to select media material from the two highest priority event themes at most.
[0277] D5. If the number of months to which the media material belongs is less than 3, then determine whether the number of months to which the media material belongs is equal to 2.
[0278] D6. If the number of months to which the media material belongs is equal to 2, then each month is allowed to select media material from the three highest priority event themes.
[0279] D7. If the number of months to which the media material belongs is not equal to 2 (equal to 1), then each month is allowed to select media material from the four highest priority event themes.
[0280] In practical applications, media materials can also be selected using another strategy. Strategy Two:
[0281] Based on the ratio between the number of media materials corresponding to different event themes in the i-th month, extract media materials corresponding to each event theme.
[0282] The media materials for the i-th month involve n event themes, and the number of media materials corresponding to each of the n event themes is i1:i2:i3...:in. When j media materials need to be extracted from the i-th month, the number of materials extracted from the media materials corresponding to each theme is obtained as follows: i1 / (i1+i2+i3+...+in)*j is: the number of media materials extracted from the media materials corresponding to the first event theme.
[0283] i2 / (i1+i2+i3+……+in)*j is the number of media materials extracted from the media materials corresponding to the second event theme.
[0284] ...
[0285] in / (i1+i2+i3+……+in)*j represents the number of media materials extracted from the media materials corresponding to the nth event theme.
[0286] Of course, the calculated number of media materials can be rounded down. After calculating the number of media materials selected for each event theme in the i-th month according to Strategy 2, they can be extracted from high to low based on the total score of the media materials under that event theme.
[0287] Strategy Two can extract media materials from each event theme within the current month. Alternatively, it can be configured to allow media materials to be extracted from several event themes. This application embodiment also sets the following rules for Strategy Two to determine the maximum number of event themes from which media materials can be extracted within the current month.
[0288] E1 determines whether the number of months to which the media material belongs is greater than or equal to 6, that is, it determines how many months have media material.
[0289] E2. If the number of months to which the media material belongs is greater than or equal to 6, then the media material will be selected from the highest priority event theme each month.
[0290] E3. If the number of months to which the media material belongs is less than 6, continue to determine whether the number of months to which the media material belongs is greater than or equal to 3.
[0291] E4. If the number of months to which the media material belongs is less than 6 and greater than or equal to 3, then media material will be selected from the two highest priority event themes each month.
[0292] E5. If the number of months to which the media material belongs is less than 3, then determine whether the number of months to which the media material belongs is equal to 2.
[0293] E6. If the number of months to which the media material belongs is equal to 2, then media material will be selected from the three highest priority event themes each month.
[0294] E7. If the number of months to which the media material belongs is not equal to 2 (equal to 1), then media material will be selected from the four highest priority event themes each month.
[0295] As an example, if the media material for the i-th month involves n event topics;
[0296] If n is less than or equal to the number of allowed event topics (k), then the media materials corresponding to each event topic are extracted according to the ratio between the number of media materials corresponding to different event topics (n event topics) as described above.
[0297] If n is greater than the allowed number of event topics (k), then determine whether the number of media materials corresponding to the k highest priority event topics is greater than or equal to the number of media materials to be extracted in the i-th month; if it is greater than or equal to, then extract media materials from the k event topics according to the ratio between the number of media materials corresponding to the k highest priority event topics; if it is less than, then select all media materials corresponding to the k highest priority event topics and calculate the difference in the i-th month.
[0298] In this way, after obtaining the extracted media materials, the difference for each month is summed to obtain the total difference; from the remaining unselected media materials throughout the year, media materials of the total difference are selected according to the priority of the event theme.
[0299] In practical applications, media materials can be extracted from each month according to Strategy 1, or according to Strategy 2.
[0300] After extracting media materials from each month and generating a finished product ID from the extracted media materials, perform the following steps:
[0301] S29, the media processing platform sends the extracted media clip card and final clip ID to the voice assistant. This final clip ID is associated with the extracted media clip.
[0302] S30: After receiving the extracted media material card and finished product ID, the voice assistant displays the media material card and finished product control on the voice assistant interface.
[0303] The final video control can be the same as the control for "Generate Video". See the attached document for details. Figure 3 The interface shown in (a) is shown in the image.
[0304] Reference Figure 10 The image shows the process of generating a video based on extracted media materials in an embodiment of this application.
[0305] S31, the voice assistant receives a voice message instructing the generation of a video or receives a click operation on the voice assistant interface to generate a video control.
[0306] S32, the voice assistant sends a finished film request to the media processing center, which carries the finished film ID.
[0307] S33, after receiving the finished film request, the media processing platform retrieves the media material associated with the finished film ID based on the finished film ID.
[0308] S34, the media processing platform generates the final film theme based on the acquired media materials.
[0309] S35, the media processing platform obtains copy and other materials that match the theme of the finished film.
[0310] S36, the media processing platform sends a final cut request to the light editing service. The final cut request includes media materials, the final cut theme, and the script.
[0311] S37, after receiving a finished film request, the light editing service matches the finished film template and music according to the finished film theme.
[0312] S38, a lightweight editing service, generates videos based on media footage, templates, text, and music.
[0313] S39, the Light Editing service sends the percentage of the finished video to the media processing platform during the video generation process.
[0314] After receiving the percentage of completed videos, the media processing platform in S40 sends the completed video card to the voice assistant.
[0315] S41, after the voice assistant receives the completed card, it displays the animation and the completed card in the voice assistant window.
[0316] S42, after generating the video, the light editing service sends a message to the media processing platform indicating that the final cut is complete. This message includes information such as the video cover and video duration.
[0317] S43: After receiving the information that the finished film is complete, the media processing platform sends the finished film ID, video cover, and video duration to the voice assistant.
[0318] S44: After receiving the finished video ID, video cover, and video duration, the voice assistant displays the video cover and video duration in the voice assistant window.
[0319] In this way, it can be displayed in the voice assistant interface. Figure 3 The interface shown in (b) includes a video cover and video duration. Users can click the play control on the video cover to play the video.
[0320] This application also provides a video generation method, including:
[0321] Receive a first instruction, which includes a time element and an event element;
[0322] Based on the time element in the first instruction, a memory node with a start time within the range of the time element is obtained. The memory node is generated by one or more related media materials. The memory node includes a start time and an event. The event is related to the content of the multiple media materials that generated the memory node.
[0323] Based on the event elements in the first instruction, a finished product template is obtained, wherein the finished product template includes multiple themes and multiple tags corresponding to each theme;
[0324] For the first memory node, media materials whose image content contains the target tag of the first memory node are selected from the media materials that generated the first memory node. The first memory node is any memory node within the time element range, and the target tag of the first memory node is determined by the event of the first memory node, the theme and tag in the final template.
[0325] A video is generated based on media material whose selected image content contains the target tag of the first memory node.
[0326] In this embodiment, the first instruction can be the user's voice instruction in the above embodiment: "Help me generate a personal annual summary video for 2023".
[0327] In another embodiment of this application, the events of the memory node include one or more topics in the slab template;
[0328] Before selecting media materials whose image content contains the target tag of the first memory node from the media materials from which the first memory node was generated, the method further includes:
[0329] Based on the events of the first memory node, the topic of the first memory node (i.e., the event topic in the above embodiment) is obtained;
[0330] The multiple tags corresponding to the theme of the first memory node in the finished template are used as the target tags of the first memory node.
[0331] As another embodiment of this application, obtaining the topic of the first memory node based on the events of the first memory node includes:
[0332] When there is only one event in the first memory node, the event of the first memory node is determined as the topic of the first memory node;
[0333] When there are multiple events for the first memory node, the event with the highest priority among the multiple themes set in the image template is determined as the theme of the first memory node.
[0334] As another embodiment of this application, obtaining the topic of the first memory node based on the events of the first memory node includes:
[0335] Generate a storyline from multiple memory nodes related to the first memory node;
[0336] The events that generate memory nodes in the same storyline are clustered to obtain the theme of the storyline, and the theme of each memory node in the storyline is the theme of the storyline.
[0337] As another embodiment of this application, the step of clustering events that generate memory nodes of the same storyline to obtain the theme of the storyline includes:
[0338] Based on the priority of multiple themes set in the finished template, the event with the highest priority is selected from the events that generate the memory nodes of the same storyline as the theme of the storyline.
[0339] As another embodiment of this application, the step of generating a video based on media material containing the target tag of the first memory node in the selected image content includes:
[0340] The selected media materials containing the target tag of the first memory node are filtered to obtain the filtered media materials;
[0341] Generate a video based on the selected media materials.
[0342] This step can be achieved using the clip model.
[0343] In another embodiment of this application, media materials whose selected image content contains the target tag of the first memory node are filtered to obtain filtered media materials, which further include:
[0344] The selected media materials whose image content contains the target tag of the first memory node are subjected to a first filtering process to obtain the media materials after the first filtering process. The first filtering process includes at least one of the following: retaining one of two or more media materials whose image content repetition is greater than a first threshold, removing media materials whose image content contains preset content, and removing media materials whose image clarity is less than a second threshold.
[0345] The first screening process is the initial screening described in the above embodiments.
[0346] As another embodiment of this application, the step of filtering media materials whose selected image content contains the target tag of the first memory node to obtain filtered media materials further includes:
[0347] Divide the time element range into multiple sub-time ranges;
[0348] For the media materials after the first screening process, among the memory nodes with the same theme within the same sub-time range, the media materials with one memory node are retained, and the media materials with other memory nodes are deleted to obtain the media materials after the second screening process. The other memory nodes are the memory nodes other than the retained memory nodes among the memory nodes with the same theme within the same sub-time range.
[0349] In the case where the time element range is one year, the sub-time range can be one month; of course, in practical applications, there may be other ways of dividing the time.
[0350] The second filtering process can be to retain one memory node for the same event theme in a month, as described in the above embodiment.
[0351] In another embodiment of this application, the memory node further includes: popular tags; the media material for retaining one memory node from memory nodes with the same theme within the same sub-time range includes:
[0352] Obtain the first number of first tags contained in the popular tags of each memory node with the same theme, wherein the same theme is the first theme and the first tags are the tags of the first theme in the template;
[0353] Sort by the first number of the first tags contained in the popular tags of each memory node with the same topic;
[0354] If the two largest first quantities are not equal, then the media material of the memory node corresponding to the largest first quantity is retained;
[0355] If the two largest first quantities are equal, then the memory nodes with the largest and equal first quantities are retained;
[0356] Calculate the median score of the most recently retained memory node, where the median score of the memory node is the median score of the media material after the first filtering process that generated the memory node;
[0357] Sort the most recently retained memory nodes by their median scores;
[0358] If the two largest median ratings are not equal, then the media material corresponding to the memory node with the largest median rating is retained;
[0359] If the two largest median scores are equal, then the memory nodes with the largest and equal median scores are retained.
[0360] Calculate the average score of the most recently retained memory nodes, where the average score of the memory nodes is the average score of the media materials after the first filtering process that generated the memory nodes;
[0361] Sort the most recently retained memory nodes by their average scores;
[0362] If the two largest average ratings are not equal, then the media material corresponding to the memory node with the largest average rating will be retained.
[0363] If the two largest average scores are equal, then the memory nodes with the largest and equal average scores are retained.
[0364] Get the total number of media materials after the first filtering process of the latest retained memory node;
[0365] Sort the total number of media materials after the first screening process of the latest retained memory nodes;
[0366] If the total number of the two largest media assets is not equal, then the media asset of the memory node corresponding to the largest total number of media assets is retained;
[0367] If the two largest media assets have the same total number, then the memory nodes with the largest and equal total number of media assets will be retained.
[0368] Select one memory node from the latest retained memory nodes to retain, and obtain the media material after the first filtering of the retained memory nodes.
[0369] In this embodiment, the retained memory nodes are determined in the following order: the number of tags included; the median rating, the average rating, the total number of media materials included, and the start time. In practical applications, the retained memory nodes can also be determined in other orders of the above parameters. Of course, the retained memory nodes can also be determined using fewer or more parameters than those mentioned above.
[0370] As another embodiment of this application, the step of filtering media materials whose selected image content contains the target tag of the first memory node to obtain filtered media materials further includes:
[0371] The media materials after the second screening process are subjected to time hashing to obtain the number of media materials in each sub-time range;
[0372] Based on the preset total number of media materials for the generated video, and the proportion of the number of media materials in each sub-time range to the number of media materials after the second filtering process, the number of media materials to be extracted from each sub-time range is obtained.
[0373] Based on the number of media materials to be extracted from each sub-time range, media materials are extracted from the media materials in each sub-time range to obtain the media materials after the third filtering process.
[0374] In the above embodiments, the three screening processes can be performed in the above order or in other orders. For example, the second screening process can be performed first, followed by the first screening process, and finally the third screening process. Of course, in practical applications, only one or more of the screening processes can be performed, and this application does not limit this.
[0375] As another embodiment of this application, the step of extracting media materials from the media materials in each sub-time range based on the number of media materials to be extracted from each sub-time range to obtain the media materials after the third filtering process includes:
[0376] Determine the maximum number of topics to extract from the media materials in each sub-time range based on the number of sub-time ranges with media materials.
[0377] For the first sub-time range, media materials are extracted from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, resulting in filtered media materials. The first sub-time range is any sub-time range within the time element range.
[0378] As another embodiment of this application, the step of extracting media materials from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, to obtain the filtered media materials, includes:
[0379] Determine whether the number of topics in the first sub-time range is greater than 1;
[0380] If the number of topics in the first sub-time range is not greater than 1, then media materials are extracted from the first sub-time range according to the number of media materials to be extracted from the first sub-time range.
[0381] If the number of topics in the first sub-time range is greater than 1, then the topic with the highest priority is selected from the topics in the first sub-time range.
[0382] Determine whether the number of media materials corresponding to the highest priority topic is greater than or equal to the number of media materials to be extracted from the first sub-time range;
[0383] If the number of media materials corresponding to the highest priority theme is greater than or equal to the number of media materials to be extracted from the first sub-time range, then media materials are extracted from the media materials corresponding to the highest priority theme according to the number of media materials to be extracted from the first sub-time range.
[0384] If the number of media materials corresponding to the highest priority topic is less than the number of media materials to be extracted from the first sub-time range, then all media materials corresponding to the highest priority topic will be extracted.
[0385] If the maximum number of topics is equal to 1, then media materials are extracted from the first sub-time range according to topic priority.
[0386] If the maximum number of topics is greater than 1, then the second highest priority topic is obtained from the topics in the first sub-time range; and media materials are extracted from the media materials corresponding to the second highest priority topic until the media materials corresponding to the p-th highest priority topic are extracted, where p represents the maximum number of topics; after the media materials corresponding to the p-th highest priority topic are extracted, the media materials are extracted from the first sub-time range according to topic priority.
[0387] As another embodiment of this application, after extracting media materials from the first sub-time range according to topic priority, the difference between the number of media materials to be extracted from the first sub-time range and the number of media materials currently extracted is obtained;
[0388] After obtaining the difference for each sub-time range within the time element range, the sum of the differences for each sub-time range within the time element range is calculated;
[0389] From the media materials that were not extracted within the time element range, continue to extract media materials of the sum of the differences based on the media material scores.
[0390] As another embodiment of this application, the step of extracting media materials from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, to obtain the filtered media materials, includes:
[0391] Determine whether the number of topics in the memory nodes within the first sub-time range is greater than the maximum number of topics;
[0392] If the number of topics in the memory nodes within the first sub-time range is greater than the maximum number of topics, then the maximum number of topics is selected according to the topic priority set in the final template.
[0393] The number of media materials to be extracted from each theme is obtained by taking into account the proportion of the number of media materials corresponding to each selected theme to the total number of media materials corresponding to all selected themes, and the number of media materials to be extracted from the first sub-time range.
[0394] Based on the number of media materials extracted from each theme, media materials are extracted from the first sub-time period.
[0395] If the number of topics in the memory nodes within the first sub-time range is less than or equal to the maximum number of topics, then the number of media materials extracted from each topic is obtained based on the ratio of the number of media materials corresponding to each topic within the first sub-time range to the total number of media materials within the first sub-time range, and the number of media materials to be extracted from the first sub-time range.
[0396] Media materials are extracted from the first sub-time frame based on the number of media materials extracted from each theme.
[0397] As another embodiment of this application, the scoring of the media material includes at least one of the following dimensions: portrait dimension, picture quality dimension, aspect ratio dimension, and user behavior dimension;
[0398] The portrait dimension is related to the proportion of the face area in the media material to the entire frame;
[0399] The image quality dimension is related to the clarity of the media material and the jitter when the media material is video;
[0400] The aspect ratio is related to the percentage of the filled area relative to the standard ratio when the aspect ratio of the media material is filled to the standard ratio;
[0401] The user behavior dimension is related to the number of times the media material is collected, shared, and viewed by users.
[0402] This embodiment can be specifically referred to in the above embodiment for the process of calculating the total score of media materials.
[0403] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0404] This application also provides a computer-readable storage medium storing a computer program that, when run on an electronic device, can implement the steps in the above-described method embodiments.
[0405] This application also provides a computer program product that, when run on an electronic device or a wireless router, enables the electronic device to perform the steps described in the various method embodiments above.
[0406] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the first device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0407] This application also provides a chip, which includes a processor coupled to a memory. The processor calls a computer program stored in the memory to implement the steps of any method embodiment of this application. The chip can be a single chip or a chip module composed of multiple chips.
[0408] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0409] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0410] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A video generation method, characterized in that, include: Receive a first instruction, which includes a time element and an event element; Based on the time element in the first instruction, a memory node with a start time within the range of the time element is obtained. The memory node is generated by one or more related media materials. The memory node includes a start time and an event. The event is related to the content of the multiple media materials that generated the memory node. Based on the event elements in the first instruction, a finished product template is obtained, wherein the finished product template includes multiple themes and multiple tags corresponding to each theme; For the first memory node, media materials whose image content contains the target tag of the first memory node are selected from the media materials that generated the first memory node. The first memory node is any memory node within the time element range, and the target tag of the first memory node is determined by the event of the first memory node, the theme and tag in the final template. A video is generated based on media material whose selected image content contains the target tag of the first memory node.
2. The method as described in claim 1, characterized in that, The events of the memory node include one or more themes in the block template; Before selecting media materials whose image content contains the target tag of the first memory node from the media materials from which the first memory node was generated, the method further includes: Based on the events of the first memory node, the topic of the first memory node is obtained; The multiple tags corresponding to the theme of the first memory node in the finished template are used as the target tags of the first memory node.
3. The method as described in claim 2, characterized in that, The step of obtaining the topic of the first memory node based on the events of the first memory node includes: When there is only one event in the first memory node, the event of the first memory node is determined as the topic of the first memory node; When there are multiple events for the first memory node, the event with the highest priority among the multiple themes set in the image template is determined as the theme of the first memory node.
4. The method as described in claim 2, characterized in that, The step of obtaining the topic of the first memory node based on the events of the first memory node includes: Generate a storyline from multiple memory nodes related to the first memory node; The events that generate memory nodes in the same storyline are clustered to obtain the theme of the storyline, and the theme of each memory node in the storyline is the theme of the storyline.
5. The method as described in claim 4, characterized in that, The process of clustering events that generate memory nodes for the same storyline to obtain the theme of the storyline includes: Based on the priority of multiple themes set in the finished template, the event with the highest priority is selected from the events that generate the memory nodes of the same storyline as the theme of the storyline.
6. The method according to any one of claims 1 to 5, characterized in that, The step of generating a video based on media material containing the target tag of the first memory node in the selected image content includes: The selected media materials containing the target tag of the first memory node are filtered to obtain the filtered media materials; Generate a video based on the selected media materials.
7. The method as described in claim 6, characterized in that, The selected media materials containing the target tag of the first memory node are filtered to obtain the filtered media materials, which also include: The selected media materials whose image content contains the target tag of the first memory node are subjected to a first filtering process to obtain the media materials after the first filtering process. The first filtering process includes at least one of the following: retaining one of two or more media materials whose image content repetition is greater than a first threshold, removing media materials whose image content contains preset content, and removing media materials whose image clarity is less than a second threshold.
8. The method as described in claim 7, characterized in that, The step of filtering media materials whose selected image content contains the target tag of the first memory node to obtain filtered media materials also includes: Divide the time element range into multiple sub-time ranges; For the media materials after the first screening process, among the memory nodes with the same theme within the same sub-time range, the media materials with one memory node are retained, and the media materials with other memory nodes are deleted to obtain the media materials after the second screening process. The other memory nodes are the memory nodes other than the retained memory nodes among the memory nodes with the same theme within the same sub-time range.
9. The method as described in claim 8, characterized in that, The memory node also includes: popular tags. The media material that retains one memory node from memory nodes with the same theme within the same sub-time range includes: Obtain the first number of first tags contained in the popular tags of each memory node with the same theme, wherein the same theme is the first theme and the first tags are the tags of the first theme in the template; Sort by the first number of the first tags contained in the popular tags of each memory node with the same topic; If the two largest first quantities are not equal, then the media material of the memory node corresponding to the largest first quantity is retained; If the two largest first quantities are equal, then the memory nodes with the largest and equal first quantities are retained; Calculate the median score of the most recently retained memory node, where the median score of the memory node is the median score of the media material after the first filtering process that generated the memory node; Sort the most recently retained memory nodes by their median scores; If the two largest median ratings are not equal, then the media material corresponding to the memory node with the largest median rating is retained; If the two largest median scores are equal, then the memory nodes with the largest and equal median scores are retained. Calculate the average score of the most recently retained memory nodes, where the average score of the memory nodes is the average score of the media materials after the first filtering process that generated the memory nodes; Sort the most recently retained memory nodes by their average scores; If the two largest average ratings are not equal, then the media material corresponding to the memory node with the largest average rating will be retained. If the two largest average scores are equal, then the memory nodes with the largest and equal average scores are retained. Get the total number of media materials after the first filtering process of the latest retained memory node; Sort the total number of media materials after the first screening process of the latest retained memory nodes; If the total number of the two largest media assets is not equal, then the media asset of the memory node corresponding to the largest total number of media assets is retained; If the two largest media assets have the same total number, then the memory nodes with the largest and equal total number of media assets will be retained. Select one memory node from the latest retained memory nodes to retain, and obtain the media material after the first filtering of the retained memory nodes.
10. The method as described in claim 8 or 9, characterized in that, The step of filtering media materials whose selected image content contains the target tag of the first memory node to obtain filtered media materials also includes: The media materials after the second screening process are subjected to time hashing to obtain the number of media materials in each sub-time range; Based on the preset total number of media materials for the generated video, and the proportion of the number of media materials in each sub-time range to the number of media materials after the second filtering process, the number of media materials to be extracted from each sub-time range is obtained. Based on the number of media materials to be extracted from each sub-time range, media materials are extracted from the media materials in each sub-time range to obtain the media materials after the third filtering process.
11. The method as described in claim 10, characterized in that, The process involves extracting media materials from the media materials in each sub-time range based on the number of media materials to be extracted from each sub-time range, resulting in the media materials after the third filtering process, including: Determine the maximum number of topics to extract from the media materials in each sub-time range based on the number of sub-time ranges with media materials. For the first sub-time range, media materials are extracted from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, resulting in filtered media materials. The first sub-time range is any sub-time range within the time element range.
12. The method as described in claim 11, characterized in that, The step of extracting media materials from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, to obtain the filtered media materials, includes: Determine whether the number of topics in the first sub-time range is greater than 1; If the number of topics in the first sub-time range is not greater than 1, then media materials are extracted from the first sub-time range according to the number of media materials to be extracted from the first sub-time range. If the number of topics in the first sub-time range is greater than 1, then the topic with the highest priority is selected from the topics in the first sub-time range. Determine whether the number of media materials corresponding to the highest priority topic is greater than or equal to the number of media materials to be extracted from the first sub-time range; If the number of media materials corresponding to the highest priority theme is greater than or equal to the number of media materials to be extracted from the first sub-time range, then media materials are extracted from the media materials corresponding to the highest priority theme according to the number of media materials to be extracted from the first sub-time range. If the number of media materials corresponding to the highest priority topic is less than the number of media materials to be extracted from the first sub-time range, then all media materials corresponding to the highest priority topic will be extracted. If the maximum number of topics is equal to 1, then media materials are extracted from the first sub-time range according to topic priority. If the maximum number of topics is greater than 1, then the second highest priority topic is obtained from the topics in the first sub-time range; and media materials are extracted from the media materials corresponding to the second highest priority topic until the media materials corresponding to the p-th highest priority topic are extracted, where p represents the maximum number of topics; after the media materials corresponding to the p-th highest priority topic are extracted, the media materials are extracted from the first sub-time range according to topic priority.
13. The method as described in claim 12, characterized in that, After extracting media materials from the first sub-time range according to thematic priority, obtain the difference between the number of media materials to be extracted from the first sub-time range and the number of media materials already extracted. After obtaining the difference for each sub-time range within the time element range, the sum of the differences for each sub-time range within the time element range is calculated; From the media materials that were not extracted within the time element range, continue to extract media materials of the sum of the differences based on the media material scores.
14. The method as described in claim 11, characterized in that, The step of extracting media materials from the first sub-time range based on the number of media materials to be extracted from the first sub-time range and the maximum number of topics, to obtain the filtered media materials, includes: Determine whether the number of topics in the memory nodes within the first sub-time range is greater than the maximum number of topics; If the number of topics in the memory nodes within the first sub-time range is greater than the maximum number of topics, then the maximum number of topics is selected according to the topic priority set in the final template. The number of media materials to be extracted from each theme is obtained by taking into account the proportion of the number of media materials corresponding to each selected theme to the total number of media materials corresponding to all selected themes, and the number of media materials to be extracted from the first sub-time range. Based on the number of media materials extracted from each theme, media materials are extracted from the first sub-time period. If the number of topics in the memory nodes within the first sub-time range is less than or equal to the maximum number of topics, then the number of media materials extracted from each topic is obtained based on the ratio of the number of media materials corresponding to each topic within the first sub-time range to the total number of media materials within the first sub-time range, and the number of media materials to be extracted from the first sub-time range. Media materials are extracted from the first sub-time frame based on the number of media materials extracted from each theme.
15. The method as described in claim 9 or 13, characterized in that, The scoring of the media materials includes at least one of the following dimensions: portrait dimension, picture quality dimension, aspect ratio dimension, and user behavior dimension; The portrait dimension is related to the proportion of the face area in the media material to the entire frame; The image quality dimension is related to the clarity of the media material and the jitter when the media material is video; The aspect ratio is related to the percentage of the filled area relative to the standard ratio when the aspect ratio of the media material is filled to the standard ratio; The user behavior dimension is related to the number of times the media material is collected, shared, and viewed by users.
16. An electronic device, characterized in that, It includes one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store a computer program, which, when executed by the one or more processors, causes the electronic device to perform the method as described in any one of claims 1-15.
17. A chip system applied to an electronic device, the chip system comprising one or more processors, characterized in that, The processor is used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-15.
18. A computer-readable storage medium comprising a computer program, characterized in that, When the computer program is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-15.
Citation Information
Patent Citations
Apparatus and method for acceleration data structure refit
CN112085827A
Video backtracking method and device, computer equipment and storage medium
CN112788270A
Optical amplifier gain adjusting method and device based on machine learning
CN118074806A
Learning method, learning material marking language and learning machine
CN1991933A