Method, system and device for generating pet highlight video and computer equipment

Through adjustments based on pet video scene annotation and evaluation standards and combined with video template matching methods, the problem of generating high-quality pet short videos in the existing technology is solved, and a fast, flexible and efficient video generation effect is achieved.

CN120050446APending Publication Date: 2025-05-27HANGZHOU TUYA INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510143870.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

It is difficult to quickly generate high-quality pet short videos in the prior art, and video generation depends on the ability of terminal devices and is costly. The training of its own algorithms is limited to fixed feature dimensions, and the generated videos are too similar and the content is messy.

Method used

By quickly adjusting the evaluation criteria of video clips based on different pet video scenes, selecting video clips with high evaluation marks, and matching them with their corresponding video templates, and finally generating high-quality pet highlight videos.

Benefits of technology

It realizes the rapid generation of high-quality pet highlight videos, avoiding the limitations of video generation relying on terminal devices, reducing costs, and improving the flexibility and viewing of videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050446A_ABST
    Figure CN120050446A_ABST
Patent Text Reader

Abstract

The invention provides a method, system and device for generating a pet highlight video and computer equipment, and relates to the field of pet short video generation. Acquiring a pet activity video clip; performing scene labeling on the video clip according to the video content in the video clip, obtaining a corresponding evaluation standard according to the scene labeling, and performing evaluation labeling on the video clip according to the evaluation standard; according to the behaviors of the pet in the video clip, carrying out behavior marking on the video clip; matching a video template from a plurality of preset video templates according to the behavior annotation, and determining a to-be-rendered fragment set of the pet activity according to the evaluation annotation; and rendering the video clips in the pet activity to-be-rendered clip set into a pet highlight video by using the matched video template. According to the method, the evaluation standards of the video clips can be quickly adjusted based on different scene labels, the video clips with high evaluation labels are selected and matched with the corresponding video templates, and the high-quality pet highlight video is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of short video generation, and particularly to a method, system, device and computer equipment for generating pet highlight videos. Background Art

[0002] With the accelerating pace of modern life, keeping pets has become a choice for many people, and pets are becoming increasingly important in families. However, it is difficult for pet owners to understand the situation of their pets at home when they go to work because making a high-quality pet short video requires much more time and resources than imagined. Pet short videos require high-quality shooting equipment and post-editing tools. Moreover, due to the unpredictable behavior of pets, it is necessary to wait for a long time to shoot video clips with ideal effects, which conflicts with the limited energy of most people.

[0003] There are obvious deficiencies in the related technologies. It is necessary to pre-determine a shooting script in advance, and then send the requirements for video clips required by the script to the terminal device. The terminal device shoots and uploads the video clips according to the requirements, and then assembles the clips in the background to generate a video. The shot video clips are highly dependent on the capabilities of the terminal device, are not very reproducible and scalable, and have high costs. Or it is possible to perform feature extraction and matching of pet video clips based on a proprietary algorithm, and then edit and splice the selected video clips to generate a pet video. However, the training of the proprietary algorithm is based on fixed feature dimensions and data, and the selection of video clips has certain limitations, is not flexible enough, the generated videos will be too similar, and the content of the videos is relatively messy. Summary of the Invention

[0004] To solve the deficiencies of the prior art, the purpose of the present application is to provide a method, system, device and computer equipment for generating pet highlight videos. This method can quickly adjust the evaluation criteria for pet video clips based on different pet video scene annotations, select video clips with high evaluation annotations from all video clips, match them with the corresponding video templates, and finally generate high-quality pet highlight videos.

[0005] To achieve the above purpose, the present application adopts the following technical solutions:

[0006] In the first aspect, the present application provides a method for generating a pet highlight video, the method comprising:

[0007] Obtain pet activity video clips;

[0008] According to the video content in the pet activity video clips, perform scene annotation on the pet activity video clips, obtain corresponding evaluation criteria according to the scene annotation, and perform evaluation annotation on the pet activity video clips according to the evaluation criteria;

[0009] Perform behavior annotation on pet activity video clips according to the behaviors of the pets in the pet activity video clips; wherein, each pet activity video clip obtains one or more behavior annotations;

[0010] Match video templates from multiple pre-set video templates according to the behavior annotations, and determine a set of pet activity segments to be rendered according to the evaluation annotations; wherein, the set of pet activity segments to be rendered includes one or more pet activity video clips;

[0011] Render the pet activity video clips in the set of pet activity segments to be rendered into a pet highlight video by using the matched video templates.

[0012] In one embodiment, matching video templates from multiple pre-set video templates according to the behavior annotations includes:

[0013] If the set time is reached and / or the number of obtained pet activity video clips is greater than the number threshold, determine the target behavior annotation according to the behavior annotations of each pet activity video clip; wherein, the target behavior annotation is the behavior annotation that appears the most times in multiple pet activity video clips;

[0014] Match video templates from multiple pre-set video templates according to the target behavior annotation.

[0015] In one embodiment, determining the set of pet activity segments to be rendered according to the evaluation annotations includes:

[0016] Determine the pet activity video clips marked with the target behavior annotation from the pet activity video clips, and define them as candidate pet activity video clips;

[0017] According to the evaluation annotations of the candidate pet activity video clips, screen out a target number of pet activity video clips from the candidate pet activity video clips, and the target number of pet activity video clips form the set of pet activity segments to be rendered; wherein, the pre-set video templates include the number of video slots that the video template can support for rendering, and the target number of pet activity video clips in the set of pet activity segments to be rendered is equal to the number of video slots of the matched video template.

[0018] In one embodiment, matching video templates from multiple pre-set video templates according to the behavior annotations, and determining the set of pet activity segments to be rendered according to the evaluation annotations includes:

[0019] Screen the pet activity video clips that have been evaluated and annotated. If the evaluation annotations of the pet activity video clips meet the set conditions, the pet activity video clips form the set of pet activity segments to be rendered;

[0020] If there is a template that supports rendering a single pet activity video clip among multiple pre-set video templates, match the video template from the pre-set video templates that support rendering a single pet activity video clip according to the behavior annotation of the pet activity video clip.

[0021] In one embodiment, the method further includes: performing highlight start and end annotation on the pet activity video clip according to the video content in the pet activity video clip; wherein, any pet activity video clip in the pet activity video clip set to be rendered obtains highlight start and end annotation.

[0022] The method further includes: setting a negative scene annotation list, and excluding the pet activity video clips marked with the scene annotations on the negative list from the scope of the pet activity video clip set to be rendered.

[0023] The method further includes: when the result of the highlight start and end annotation cannot be matched to the video template, excluding the corresponding pet activity video clip from the scope of the pet activity video clip set to be rendered.

[0024] In one embodiment, the method further includes: updating one or more of the evaluation annotation, scene annotation, and highlight start and end annotation.

[0025] Among them, updating one or more of the evaluation annotation, scene annotation, and highlight start and end annotation includes:

[0026] Providing the generated pet highlight video to the user, and providing a user interface to receive one or more feedbacks from the user on the pet highlight video, and integrating the feedback into one or more of the evaluation annotation, scene annotation, highlight start and end annotation, and video template matching.

[0027] In one embodiment, rendering the pet activity video clips in the pet activity video clip set to be rendered into a pet highlight video by using the matched video template includes:

[0028] Pre-setting multiple video slots in the video template, and integrating multiple pet activity video clips with the same or related behavior annotations in the pet activity video clip set to be rendered into the multiple video slots; wherein, one or more of a background music slot, a copywriting slot, and a transition effect slot are also set in the video template.

[0029] In one embodiment, performing scene annotation on the pet activity video clip according to the video content in the pet activity video clip includes:

[0030] Obtaining multiple candidate scenes.

[0031] Identify the video content in the video clip of pet activities, and analyze the corresponding scene of the video clip of pet activities from multiple candidate scenes according to the video content for scene annotation.

[0032] In a second aspect, the present application provides a system for generating a pet highlight video, the system comprising:

[0033] A terminal configured to collect video clips of pet activities;

[0034] A cloud configured to generate a pet highlight video by using the method for generating a pet highlight video in the first aspect.

[0035] In a third aspect, the present application provides an apparatus for generating a pet highlight video, the apparatus comprising:

[0036] An acquisition unit for acquiring video clips of pet activities;

[0037] An evaluation annotation unit for performing scene annotation on the video clip of pet activities according to the video content in the video clip of pet activities, obtaining corresponding evaluation criteria according to the scene annotation, and performing evaluation annotation on the video clip of pet activities according to the evaluation criteria;

[0038] A behavior annotation unit for performing behavior annotation on the video clip of pet activities according to the behavior of the pet in the video clip of pet activities; wherein, one or more behavior annotations are obtained for each video clip of pet activities;

[0039] A template matching unit for matching video templates from multiple pre-set video templates according to the behavior annotation, and determining a set of pet activity segments to be rendered according to the evaluation annotation; wherein, the set of pet activity segments to be rendered includes one or more video clips of pet activities;

[0040] A rendering unit for rendering the video clips of pet activities in the set of pet activity segments to be rendered into a pet highlight video by using the matched video templates.

[0041] In a fourth aspect, the present application provides a computer device, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method for generating a pet highlight video in the first aspect when executing the computer program.

[0042] The above method for generating a pet highlight video obtains video clips of pet activities from the cloud, performs scene annotation based on the video content, obtains corresponding evaluation criteria and performs evaluation annotation; the cloud performs behavior annotation on pet behaviors, and one or more behavior annotations can be obtained for each clip. The cloud matches a suitable video template based on the behavior annotation and determines a set of pet activity clips to be rendered according to the evaluation annotation. The cloud uses the matched video template to render the clips in the set of clips to be rendered into a pet highlight video. This method can quickly adjust the evaluation criteria of pet video clips based on different pet video scene annotations, select video clips with high evaluation annotations from all video clips, match them with the corresponding video templates, and finally generate high-quality pet highlight videos. Description of the Drawings

[0043] Figure 1 is a flowchart of a method for generating a pet highlight video in an embodiment;

[0044] Figure 2 is a flowchart of matching a video template from multiple pre-set video templates according to behavior annotation in an embodiment;

[0045] Figure 3 is a flowchart of determining a set of pet activity clips to be rendered according to evaluation annotation in an embodiment;

[0046] Figure 4 is a flowchart of determining a set of pet activity clips to be rendered in an embodiment;

[0047] Figure 5 is a flowchart of performing scene annotation on pet activity video clips in an embodiment;

[0048] Figure 6 is a flowchart of generating a pet highlight video based on multiple pet activity video clips in an embodiment;

[0049] Figure 7 is a diagram of a device for generating a pet highlight video in an embodiment;

[0050] Figure 8 is a structural diagram of a computer device in an embodiment. Detailed Embodiments

[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs.

[0052] In one embodiment, as Figure 1 shown, a method for generating a highlight video of a pet is provided. The method includes the following steps:

[0053] Step 101: Obtain video clips of pet activities;

[0054] There are two different ways to obtain video clips of pet activities from the cloud:

[0055] The first way is that for terminals supporting pet detection, during the operation of such terminals, if the activity of a pet is detected, the shooting function of such terminals is activated to shoot and record video clips of the pet's activities. The terminal can directly upload the shot video clips of the pet's activities to the cloud.

[0056] The second way is that for terminals only equipped with motion detection functions, as long as such terminals detect that an object moves, regardless of whether the moving object is a pet, the shooting function of such terminals will be activated, and the shot video clips will be uploaded to the cloud. Further, since such terminals cannot distinguish whether there is a pet in the captured picture, the video clips uploaded to the cloud may contain irrelevant content. To ensure that the subsequent processed video clips are all related to pet activities, the cloud needs to perform pet detection and filtering processing on the uploaded video clips. It should be noted that the cloud can use a variety of detection technologies for pet detection and filtering processing, such as local small model processing (such as yolo / mmdetection / self-developed models, etc.), large model multi-modal detection, or cloud provider detection services, and select a suitable detection method according to actual cost and efficiency requirements to screen out video clips containing pet activities and eliminate video content without pet activities, so that the video clips are all pet activity video clips containing pets.

[0057] Step 102: Perform scene annotation on the video clips of pet activities according to the video content in the video clips of pet activities, obtain corresponding evaluation criteria according to the scene annotation, and perform evaluation annotation on the video clips of pet activities according to the evaluation criteria;

[0058] The cloud performs recognition and analysis on the video content in the obtained video clips of pet activities to identify the scene where the pet activities are located, such as indoor rest, outdoor play, feeding scene, etc. Further, the cloud can perform scene annotation according to different scenes and obtain evaluation criteria corresponding to the scene annotation based on the scene annotation.

[0059] It should be noted that different scenarios may have different evaluation criteria. For example, in the indoor rest scenario, more attention may be paid to the comfort of the environment and the relaxation state of the pet; in the outdoor play scenario, the activity level of the pet and its interaction with the environment are emphasized. The cloud evaluates and annotates the pet activity video clips according to the obtained evaluation criteria. Among them, the evaluation annotation can be a quantified score, that is, the cloud can score the pet activity video clips according to the evaluation criteria corresponding to the scenario.

[0060] Step 103: Perform behavior annotation on the pet activity video clip according to the behavior of the pet in the pet activity video clip; among them, each pet activity video clip obtains one or more behavior annotations;

[0061] When processing the pet activity video clip, the cloud can perform behavior annotation based on the behavior shown by the pet in the pet activity video clip. Exemplarily, if the cloud detects that the behavior of the pet in the pet activity video clip is eating, it can be annotated as "eating behavior"; if the pet is interacting and playing with a toy, it can be annotated as "playing behavior"; if the pet is in a rest and sleep state, it can be annotated as "rest behavior".

[0062] It should be noted that the pet in the pet activity video clip may show multiple behaviors. For example, if it plays first and then eats, then this pet activity video clip can obtain two annotations: "playing behavior" and "eating behavior".

[0063] Furthermore, the basis for annotating pet behavior can be flexible. On the one hand, pet behavior can be pre-set common pet behaviors, such as "eating", "playing", "sleeping", etc., and the pet activity video clip is annotated according to different behaviors. On the other hand, the annotation of pet behavior can also be derived by the cloud. The cloud can analyze the video content based on data and algorithms to automatically summarize the pet's behavior and perform behavior annotation.

[0064] Step 104: Match the video template from multiple pre-set video templates according to the behavior annotation, and determine the set of pet activity segments to be rendered according to the evaluation annotation; among them, the set of pet activity segments to be rendered includes one or more pet activity video clips;

[0065] Based on the behavior annotation of the pet activity video clip, the cloud can find the video template that matches the behavior annotation of the pet activity video clip from multiple pre-set video templates. For example, if the behavior annotation of this pet activity video clip is "playing behavior", then the cloud can select the template designed for pet playing from multiple pre-set video templates.

[0066] It should be noted that the video template can be a preset framework for generating pet highlight videos. The video template can include video slots, specifying the number, layout, and duration of pet activity video clips that can be placed; there is a background music slot to match suitable music to create an atmosphere; there may also be a copywriting slot, a transition effect slot, etc. Matching the template according to the behavior annotation can make the video style adapt to the pet behavior.

[0067] Similarly, the cloud can select pet activity video clips whose evaluation annotations meet the requirements based on the evaluation annotations of the video clips, and then determine the set of pet activity video clips to be rendered. The set of pet activity video clips to be rendered can contain only one high-quality pet activity video clip. For example, in a certain clip, the behavior of the pet is extremely interesting and the overall quality of the video is high, and a single clip is sufficient to meet the requirements; it can also contain multiple video clips. For example, multiple clips respectively show different interesting behaviors of the pet, and multiple clips together constitute the set of video clips to be rendered.

[0068] It should be noted that in the process of generating pet highlight videos, the scene annotation, behavior annotation, evaluation annotation, and matching of video templates for pet activity video clips can be completed by the cloud based on a large model. Among them, the large model can be an artificial intelligence model built based on a deep learning architecture. Through unsupervised or supervised learning of a large amount of multi-source data, for example, on the basis of architectures such as Transformer, it is trained to contain hundreds of millions of parameters. When generating pet highlight videos, the large model can learn pet-related video data and analyze elements such as the environment, actions, and light in it. When performing scene annotation, the large model can judge whether it is an indoor or outdoor scene based on the environmental characteristics in the video; when performing behavior annotation, the large model can identify the pet's behavior by capturing the pet's limb movements and movement trajectories; when performing evaluation annotation, the large model can score and evaluate the quality of pet activity video clips according to the evaluation criteria corresponding to the scene annotation; when matching video templates, the large model can comprehensively consider the pet's behavior and the styles of each template, select a suitable template, and generate wonderful pet highlight videos.

[0069] Step 105: Use the matched video template to render the pet activity video clips in the set of pet activity video clips to be rendered into pet highlight videos.

[0070] Based on the previous steps, the cloud has completed operations such as scene annotation, behavior annotation, and evaluation annotation on the pet activity video clips, determined the set of pet activity video clips to be rendered, and matched a suitable video template. Further, the cloud can use the matched video template to render the pet activity video clips in the set of pet activity video clips to be rendered. Among them, the video template usually includes specific elements such as picture layout, editing rhythm, special effects, and music. The cloud can combine one or more pet activity video clips together according to the rules of the template, add corresponding special effects and music, and integrate the pet activity video clips into a complete pet highlight video.

[0071] In this embodiment, the cloud obtains pet activity video clips, performs scene annotation according to the video content, obtains corresponding evaluation criteria and performs evaluation annotation; the cloud performs behavior annotation on pet behaviors, and each clip can obtain one or more behavior annotations. The cloud matches a suitable video template based on the behavior annotation and determines the set of pet activity clips to be rendered according to the evaluation annotation. The cloud uses the matched video template to render the clips in the set of clips to be rendered into a pet highlight video. This method can quickly adjust the evaluation criteria of pet video clips based on different pet video scene annotations, select video clips with high evaluation annotations from all video clips, match them with their corresponding video templates, and finally generate high-quality pet highlight videos. Through the combination of the large and small models in the cloud, it is possible to obtain the ability to detect pet clips online without relying on the local pet detection algorithm on the terminal device.

[0072] In one embodiment, as Figure 2 shown, matching a video template from multiple pre-set video templates according to the behavior annotation includes the following steps:

[0073] Step 201: If the set time is reached and / or the number of obtained pet activity video clips is greater than the number threshold, determine the target behavior annotation according to the behavior annotations of each pet activity video clip; wherein, the target behavior annotation is the behavior annotation that appears the most times among multiple pet activity video clips;

[0074] When the preset specific time is reached and / or the number of obtained pet activity video clips exceeds the set number threshold, the cloud can analyze according to the behavior annotations of each pet activity video clip. The behavior annotation of each pet activity video clip can record the specific behaviors of the pet in each pet activity video clip, such as "playing", "eating", "sleeping", etc.

[0075] The cloud can count the behavior annotations of all segments, select the behavior annotation that appears most frequently among multiple pet activity video segments, and determine it as the target behavior annotation. For example, among the multiple pet activity video segments obtained, the behavior annotation of "playing" appears 15 times, "eating" appears 8 times, and "sleeping" appears 5 times, then the behavior annotation of "playing" is determined as the target behavior annotation.

[0076] Step 202: Match a video template from multiple pre-set video templates according to the target behavior annotation.

[0077] Based on the target behavior annotation determined in step 201, the cloud can screen and match in a library of multiple pre-set video templates. Among them, each pre-set video template has a specific style, rhythm, and behavioral characteristics suitable for display. For example, there can be a template designed specifically to show the "playing" behavior of pets, which contains lively and cheerful music to highlight the vitality during play; there can also be a template suitable for the "eating" behavior, which can be a warm color tone and a slow rhythm, focusing on showing the details of eating.

[0078] When "playing" is determined as the target behavior annotation, the cloud searches in the template library for the video template that best matches this target behavior annotation, that is, the cloud finds a template that can better present the "playing" behavior in terms of design, so as to be used subsequently to produce a pet video showing this target behavior, making the style and rhythm of the video match the behavior of the pet, thus bringing a more immersive and ornamental video experience to the audience.

[0079] In this embodiment, a trigger condition is set. When the set time is reached or the number of pet activity video segments obtained is greater than the quantity threshold, the behavior annotation that appears most frequently is found based on the behavior annotations of each segment as the target behavior annotation, and a matching operation is performed among multiple pre-set video templates. This method ensures the material basis for video generation through the trigger condition. Matching the template according to the target behavior annotation can make the highlight video match the most common behavior of the pet, show daily typical scenes, and enhance the video's ornamental value.

[0080] In one embodiment, as Figure 3 shown, according to the evaluation annotation, determining the set of pet activity video segments to be rendered includes the following steps:

[0081] Step 301: Determine the pet activity video segments marked with the target behavior annotation from the pet activity video segments, and define them as candidate pet activity video segments;

[0082] The cloud can perform screening operations on the obtained pet activity video clips. The cloud can traverse all the obtained pet activity video clips based on the determined target behavior annotation, and define the pet activity video clips with the target behavior annotation as candidate pet activity video clips.

[0083] Exemplarily, if the target behavior annotation is "playing", the cloud can select all the pet activity video clips with the behavior annotation of "playing" as candidate pet activity video clips.

[0084] Step 302: According to the evaluation annotation of the candidate pet activity video clips, screen out a target number of pet activity video clips from the candidate pet activity video clips. The pet activity video clips with the target number form a set of pet activity video clips to be rendered; wherein, the preset video template includes the number of video slots that the video template can support for rendering, and the target number of pet activity video clips in the set of pet activity video clips to be rendered is equal to the number of video slots of the matched video template.

[0085] The cloud can perform screening based on the evaluation annotation of the candidate pet activity video clips. Among them, the evaluation annotation can include multiple evaluation dimensions of the pet activity video clips, such as the evaluation of background quality, pet behavior performance, exposure quality, etc. The cloud can use the evaluation annotation to evaluate and sort the candidate pet activity video clips.

[0086] Furthermore, the cloud selects a target number of video clips from the candidate clips, and the target number can be determined according to the number of video slots that the matched video template can support for rendering. For example, assuming that the matched video template has 5 video slots, then 5 pet activity video clips need to be screened out from the candidate pet activity video clips.

[0087] Through the comprehensive evaluation of the evaluation annotation of the candidate pet activity video clips, the cloud can select high-quality and outstanding clips and form a set of pet activity video clips to be rendered with them. The video clips in the set of pet activity video clips to be rendered will be exactly matched with the matched video template, and the number of clips in the set of pet activity video clips to be rendered is equal to the number of video slots supported by the video template, thus providing accurate quantity and high-quality materials for the subsequent rendering of the pet highlight video using this video template, and ensuring that each video slot can be filled with a suitable video clip during the rendering process.

[0088] In this embodiment, segments marked with target behavior annotations are found from multiple pet activity video segments, and they are defined as candidate pet activity video segments. According to the evaluation annotations of the candidate pet activity video segments, segments that match the number of video slots in the video template are selected to form a set of pet activity video segments to be rendered. This method can screen out targeted and high-quality video segments, ensuring that the final pet highlight video not only showcases typical pet behaviors but also has excellent quality, enhancing the viewing experience and meeting user needs.

[0089] In one embodiment, as Figure 4 shown, a video template is matched from multiple pre-set video templates according to the behavior annotation, and according to the evaluation annotation, a set of pet activity video segments to be rendered is determined, including the following steps:

[0090] Step 401: Screen the pet activity video segments with evaluation annotations. If there are pet activity video segments whose evaluation annotations meet the set conditions, then these pet activity video segments form a set of pet activity video segments to be rendered;

[0091] If there are pet activity video segments whose evaluation annotations meet the set conditions, then these pet activity video segments form a set of pet activity video segments to be rendered. Among them, the set conditions can be a clear screening criterion, which can be determined according to the requirements for generating a pet highlight video. If the evaluation annotation of a pet activity video segment meets the set conditions, then this pet activity video segment will be selected and screened into the set of pet activity video segments to be rendered.

[0092] Exemplarily, if the set condition is that the evaluation annotation of the pet activity video segment needs to be above 80 points, then the cloud traverses all pet activity video segments with evaluation annotations, adds the pet activity video segments that meet the condition of the evaluation annotation above 80 points to the set of pet activity video segments to be rendered, and if the evaluation annotation of the pet activity video segment is below 80 points, it cannot be added to the set of pet activity video segments to be rendered.

[0093] Step 402: If there is a template in the multiple pre-set video templates that supports rendering a single pet activity video segment, then match a video template from the pre-set video templates that support rendering a single pet activity video segment according to the behavior annotation of this pet activity video segment.

[0094] The cloud checks multiple pre-set video templates with the aim of finding the templates that support rendering a single pet activity video clip. It should be noted that different video templates may have different functions and scopes of application. Some templates may be more suitable for processing the splicing and combination of multiple video clips, while some are specifically designed for rendering a single pet activity video clip. When the cloud discovers that there is a template that supports rendering a single pet activity video clip, the cloud can perform further matching operations based on the behavior annotation of the current pet activity video clip.

[0095] The behavior annotation can reflect the specific behavior of the pet in the video clip, such as playing, eating, sleeping, or running. The cloud can use the behavior annotation to screen among the video templates that support rendering a single pet activity video clip to find a suitable video template. For example, if the behavior annotation of the pet activity video clip is "playing", the cloud can search for a video template that shows the pet playing scene among the video templates that support single clip rendering.

[0096] It should be noted that the cloud will also determine whether the pet activity video clip with this behavior annotation has recently been matched with a template that supports rendering a single pet activity video clip. If so, the pet activity video clip with this behavior annotation cannot be matched with the video template that supports single clip rendering; if there is no template that supports rendering a single pet activity video clip, the pet activity video clip with this behavior annotation also cannot be matched with the video template.

[0097] Furthermore, the cloud will set exclusive video templates related to holidays for some special holidays, and the exclusive video templates need to be launched at a specified time. When selecting video templates in the same situation, the cloud will first screen the templates with specific launch times. For example, if it is the Spring Festival recently, it will first match the video templates launched during the Spring Festival to enhance the festive atmosphere of the video; if there is no video template with a relevant launch time currently, then select ordinary video templates for assembly.

[0098] In this embodiment, the pet activity video clips with evaluation annotations are screened, and the clips that meet the set conditions are grouped into a set of clips to be rendered; if there are templates that support single clip rendering, templates will be matched in such templates according to the behavior annotations of the corresponding clips. This method ensures the quality of the set to be rendered by screening the clips with evaluation annotations that meet the conditions, and uses behavior annotations to match suitable single clip templates, making the generated video more wonderful and more targeted.

[0099] In one embodiment, the method further includes: performing highlight start and end annotations on the pet activity video clip according to the video content in the pet activity video clip; among them, any pet activity video clip in the set of pet activity clips to be rendered obtains highlight start and end annotations.

[0100] The cloud will perform highlight start and end annotations on the video content in the pet activity video clips. Among them, the highlight can refer to the most wonderful and prominent part in the pet activity video clip. For example, if a pet completes a difficult jumping action in the video, the process from the moment of preparing to jump to the moment of standing firm after landing can be marked as a highlight period. In this way, each pet activity video clip in the set of pet activity video clips to be rendered can obtain corresponding highlight start and end annotations. It should be noted that the cloud can use a large model to perform highlight start and end annotations on the clips of pet activity videos. The large model can learn a large amount of pet behavior data and master the wonderful moment patterns corresponding to different behaviors. When processing pet activity videos, the large model can analyze features such as pet movements and expressions. When a highly ornamental moment appears, it marks the start of the highlight; when the wonderful behavior ends, it marks the end time to complete the annotation.

[0101] Furthermore, during subsequent video rendering and synthesis, the cloud can highlight the wonderful moments of pet activities according to the highlight start and end annotations, making the generated pet highlight video focused and enhancing the viewing and attractiveness of the video.

[0102] In one embodiment, the method further includes: setting a negative scene annotation list and excluding the pet activity video clips marked with the scene annotations on the negative list from the scope of the set of pet activity video clips to be rendered.

[0103] The cloud can set a negative scene annotation list, which can include various scene annotations that are not suitable to appear in the pet highlight video. For example, the negative scene annotation list can include scene annotations such as no babies and no humans can appear.

[0104] When the cloud processes pet activity video clips, it will check each video clip to see if it is marked with the scene annotations on the negative scene annotation list. If the cloud checks that a certain pet activity video clip has the scene annotations in the negative scene list, then this clip will be excluded from the scope of the set of pet activity video clips to be rendered.

[0105] It should be noted that due to the limitations of different business requirements, the cloud will have a certain screening for pet activity video clips. The purpose is to ensure that the video clips in the set of pet activity video clips to be rendered all meet the business requirements and avoid negative scenes that affect the user viewing experience from appearing in the generated pet highlight video.

[0106] In one embodiment, the method further includes: when the result of the highlight start and end annotation cannot be matched to the video template, excluding the corresponding pet activity video clip from the scope of the set of pet activity video clips to be rendered.

[0107] After performing highlight start and end annotations on pet activity video clips in the cloud, the cloud can attempt to match the results of the highlight start and end annotations with video templates. Video templates have their own characteristics and structures. For example, some templates are designed to showcase the continuous actions of pets, while others focus on highlighting a certain wonderful moment of a pet. These characteristics determine the scope and form of the highlight start and end annotations that different templates can match.

[0108] Exemplarily, if the results of the highlight start and end annotations cannot be successfully matched with any existing video template, it means that the pet activity video clip corresponding to the structure of the highlight start and end annotations cannot generate a pet highlight video according to the template.

[0109] Furthermore, to ensure the smoothness and overall effect of the finally generated pet highlight video, the cloud can exclude the pet activity video clip corresponding to the results of the highlight start and end annotations from the set of pet activity video clips to be rendered. This avoids unmatched clips from disrupting the coherence and aesthetics of the video, ensuring that each clip in the set of pet activity video clips to be rendered can fit the selected video template, so as to facilitate subsequent video rendering work.

[0110] In one embodiment, the method further includes: updating one or more annotation methods among the evaluation annotation, scene annotation, and highlight start and end annotations; wherein, updating one or more annotation methods among the evaluation annotation, scene annotation, and highlight start and end annotations includes: providing the generated pet highlight video to the user, and providing a user interface to receive one or more feedbacks from the user on the pet highlight video, and integrating the feedback into one or more annotation methods or matching methods among the evaluation annotation, scene annotation, highlight start and end annotation, and video template matching.

[0111] If the user obtains the pet highlight video sent by the cloud through the APP on the mobile phone, there is a user interface in the APP for receiving user feedback. For example, the user points out that the pet in the video faces the camera and looks cute, but the duration is short, and hopes to see more similar clips. And the user also feedbacks that they don't like the picture of the pet in the cage in the video, and hopes that the pet can move around in a large indoor area.

[0112] After collecting these feedbacks, the cloud can update the annotation method and matching strategy. Increase the annotation weight of the behavior that the pet faces the camera and looks cute, and at the same time adjust the highlight start and end annotations to extend the capture duration of similar clips; the cloud can also mark the scene of the pet in the cage as a low priority to reduce the probability of such clips appearing in future video generation.

[0113] Furthermore, the cloud can also adjust the video template matching strategy according to the user's preference for video style, and preferentially select templates with lively background music and dynamic transition effects.

[0114] In this embodiment, through a dynamic update mechanism based on user feedback, the cloud can continuously optimize the labeling and matching rules to generate pet highlight videos that better meet the user's personalized needs, thereby improving user satisfaction and the attractiveness of the video content.

[0115] In one embodiment, using the matched video template to render the pet activity video clips in the set of pet activity clips to be rendered into a pet highlight video includes:

[0116] Multiple video slots are pre-set in the video template, and multiple pet activity video clips with the same or related behavior annotations in the set of pet activity clips to be rendered are integrated into the multiple video slots; wherein the video template is also provided with one or more of a background music slot, a text slot and a transition effect slot.

[0117] Specifically, multiple video slots are pre-set in the video template, and the video slots can be used to place video clips in the set of pet activity clips to be rendered. The cloud can integrate multiple pet activity video clips with the same or related behavior labels in the set of pet activity clips to be rendered into the video slots. For example, if there are multiple pet activity video clips with the behavior label of "playing", these pet activity video clips can be arranged in the video slots in the corresponding video template, or several clips with the behavior label of "eating" will be placed in the corresponding video slot of another accessory template.

[0118] After the pet activity clip set to be rendered matches the appropriate video template, the cloud renders the pet activity video clips in the pet activity clip set to be rendered. The pet activity video clips are placed in the corresponding video slots in order according to the requirements set by the video template. Combined with the elements of the video template, the pet activity video clips are processed and finally a complete pet highlight video is generated.

[0119] Furthermore, slots for other elements are also set in the video template, such as one or more of the background music slot, text slot and transition effect slot. It should be noted that during the rendering process, suitable background music can be added to the background music slot to make the atmosphere of the video more vivid and interesting; the text slot can be used to add some descriptive text, such as an explanation of the pet's behavior or praise for the pet; the transition effect slot can add various transition effects, such as fade in and fade out, rotation switching, etc., to enhance the coherence and viewing experience of the video.

[0120] In this embodiment, the method combines pet activity video clips with different slot elements to render the video clips in the set of pet activity clips to be rendered into a complete and attractive pet highlight video, giving users a better viewing experience.

[0121] In one embodiment, as Figure 5 shown, according to the video content in the pet activity video clip, scene annotation is performed on the pet activity video clip, including the following steps:

[0122] Step 501: Obtain multiple candidate scenes;

[0123] The cloud can obtain candidate scenes in various ways. Exemplarily, the cloud can analyze common scene elements in a large amount of existing pet activity video data, such as the indoor living room, outdoor lawn, pet park, etc. where the pet is active, and surrounding environmental items such as food bowls, toys, pet beds, etc. The cloud combines and generalizes different scene elements to obtain multiple possible scenes and thus obtains multiple candidate scenes. The cloud can also use image recognition technology to analyze the image features in the pet activity video clip and identify the scene type involved in the image features as candidate scenes.

[0124] Step 502: Identify the video content in the pet activity video clip, and analyze the scene corresponding to the pet activity video clip from multiple candidate scenes according to the video content for scene annotation.

[0125] Specifically, the cloud can identify the video content in the pet activity video. The video content can be the behavior actions of the pet in the video frame, the surrounding environment, and whether there are other pets or people in the video, etc. After the cloud identifies the pet activity video clip, it obtains the corresponding video content and filters and analyzes the most suitable scene for the pet activity video clip from the multiple candidate scenes obtained in step 501. For example, if it is identified that the pet in the video is eating in an indoor environment with a food bowl and a water bowl, comparing multiple candidate scenes, this video clip can be labeled as an "indoor eating scene" to complete the scene annotation.

[0126] It should be noted that the cloud can also directly perform custom scene reasoning on the pet activity video clip, that is, the cloud performs video content recognition on the pet activity video clip. For the recognized video content, the cloud performs custom scene reasoning on it, directly infers the scene corresponding to the pet activity video clip, and performs scene annotation.

[0127] In this embodiment, when performing scene annotation on the pet activity video clip, multiple candidate scenes are obtained, the video content in the video clip is identified, and the corresponding scene is analyzed from the candidate scenes based on this content to complete the annotation. This method provides multiple possible choices for annotation by pre-obtaining candidate scenes, and the accurate recognition and analysis of video content can ensure that the scene annotation fits the actual situation, thereby providing accurate scene information for subsequent video processing.

[0128] In one embodiment, as Figure 6As shown in the figure, the specific steps for generating a daily pet highlight video based on multiple pet activity video clips are as follows:

[0129] Step 601: The cloud obtains pet activity video clips.

[0130] Step 602: The cloud performs scene annotation on the scenes of the obtained pet activity video clips.

[0131] Step 603: The cloud obtains corresponding evaluation criteria based on the scene annotation and performs evaluation annotation according to the evaluation criteria.

[0132] Step 604: The cloud performs behavior annotation on the pet activity video clips according to the behaviors of the pets in the pet activity video clips, and performs highlight start and end annotation on the pet activity video clips according to the video content in the pet activity video clips.

[0133] Step 605: If the set time is reached and / or the number of obtained pet activity video clips is greater than the quantity threshold, the cloud determines the target behavior annotation according to the behavior annotations of each pet activity video clip.

[0134] Step 606: The cloud determines the pet activity video clips marked with the target behavior annotation in the pet activity video clips and defines them as candidate pet activity video clips.

[0135] Step 607: The cloud performs video template matching based on the target behavior annotation and preferentially selects a video template with a matching delivery time according to the current time; if there is no list of templates with a matching delivery time, a common video template is selected.

[0136] Step 608: The cloud selects multiple candidate pet activity video clips with high scores to form a set of pet activity video clips to be rendered, and the number of video clips in the set of pet activity video clips to be rendered matches the number of video slots of the video template.

[0137] Step 609: The cloud uses the matched video template to render the pet activity video clips in the set of pet activity video clips to be rendered into a daily pet highlight video.

[0138] It should be noted that the positions of Step 606 and Step 607 can be interchanged, that is, there is no sequential relationship between the two.

[0139] Based on the same concept, the present application also provides a system for generating a pet highlight video, and the system includes:

[0140] A terminal configured to collect pet activity video clips;

[0141] A cloud configured to generate a pet highlight video by using the method for generating a pet highlight video.

[0142] Specifically, the system for generating pet highlight videos includes two parts: a terminal and a cloud. Among them, the terminal mainly collects video clips of pet activities. The terminal can be various devices with video shooting functions. During the daily activities of the pet, the terminal can timely record various behavior performances of the pet and accumulate rich video materials. The cloud is the core processing part of the system. The cloud can process the video clips of pet activities uploaded by the terminal by using the method for generating pet highlight videos. The process of generating pet highlight videos mainly involves screening, analyzing, and annotating the video clips of pet activities, matching video templates according to specific rules, and rendering appropriate video clips into wonderful pet highlight videos. Through the cooperation of the terminal and the cloud, a complete process from material collection to highlight video generation is realized, providing users with a convenient and efficient pet highlight video production experience. It should be noted that the system can also include a user terminal, and the user can receive the pet highlight videos generated by the cloud and can provide feedback and suggestions on the videos.

[0143] Based on the same inventive concept, an embodiment of the present application provides a device for generating pet highlight videos. The implementation solution provided by the device for solving problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of generating pet highlight videos provided below can refer to the limitations on the method for generating pet highlight videos in the above text, and will not be repeated here.

[0144] In one embodiment, as Figure 7 shown, an embodiment of the present application further provides a device for generating pet highlight videos. The device includes:

[0145] An acquisition unit 701, configured to acquire video clips of pet activities;

[0146] An evaluation and annotation unit 702, configured to perform scene annotation on the video clips of pet activities according to the video content in the video clips of pet activities, obtain corresponding evaluation criteria according to the scene annotation, and perform evaluation and annotation on the video clips of pet activities according to the evaluation criteria;

[0147] A behavior annotation unit 703, configured to perform behavior annotation on the video clips of pet activities according to the behaviors of the pets in the video clips of pet activities; wherein, each video clip of pet activities obtains one or more behavior annotations;

[0148] A template matching unit 704, configured to match video templates from a plurality of pre-set video templates according to the behavior annotations, and determine a set of pet activity segments to be rendered according to the evaluation annotations; wherein, the set of pet activity segments to be rendered includes one or more video clips of pet activities;

[0149] A rendering unit 705 is configured to render the pet activity video segments in the pet activity video segment set to be rendered into a pet highlight video by using the matched video template.

[0150] In one embodiment, the behavior annotation unit 703 matches a video template from multiple pre-set video templates according to the behavior annotation. Specifically, when the set time is reached and / or the number of pet activity video segments obtained is greater than the number threshold, the target behavior annotation is determined according to the behavior annotations of each pet activity video segment, where the target behavior annotation is the behavior annotation that appears the most times among multiple pet activity video segments; and a video template is matched from multiple pre-set video templates according to the target behavior annotation.

[0151] In one embodiment, the template matching unit 704 determines the pet activity video segment set to be rendered according to the evaluation annotation. Specifically, the pet activity video segments marked with the target behavior annotation are determined from the pet activity video segments and defined as candidate pet activity video segments; according to the evaluation annotations of the candidate pet activity video segments, a target number of pet activity video segments are selected from the candidate pet activity video segments, and the target number of pet activity video segments form the pet activity video segment set to be rendered. The pre-set video template includes the number of video slots that the video template can support for rendering, and the target number of pet activity video segments in the pet activity video segment set to be rendered is equal to the number of video slots of the matched video template.

[0152] In one embodiment, the template matching unit 704 matches a video template from multiple pre-set video templates according to the behavior annotation and determines the pet activity video segment set to be rendered according to the evaluation annotation. Specifically, the pet activity video segments that have been evaluated and annotated are screened. If the evaluation annotation of a pet activity video segment meets the set conditions, the pet activity video segment forms the pet activity video segment set to be rendered. If there is a template among the multiple pre-set video templates that supports rendering a single pet activity video segment, a video template is matched from the pre-set video templates that support rendering a single pet activity video segment according to the behavior annotation of the pet activity video segment.

[0153] In one embodiment, the evaluation annotation unit 702 performs highlight start and end annotations on the pet activity video segments according to the video content in the pet activity video segments. Any pet activity video segment in the pet activity video segment set to be rendered is given a highlight start and end annotation. A negative scene annotation list is set, and the pet activity video segments marked with the scene annotations on the negative list are excluded from the scope of the pet activity video segment set to be rendered. When the result of the highlight start and end annotation cannot be matched to a video template, the corresponding pet activity video segment is excluded from the scope of the pet activity video segment set to be rendered.

[0154] In one embodiment, the evaluation annotation unit 702 updates one or more of the evaluation annotation, scene annotation, and highlight start and end annotation; wherein, updating one or more of the evaluation annotation, scene annotation, and highlight start and end annotation includes: providing the generated pet highlight video to the user, and providing a user interface to receive one or more feedbacks from the user on the pet highlight video, and integrating the feedbacks into one or more of the annotation methods or matching methods of the evaluation annotation, scene annotation, highlight start and end annotation, and video template matching.

[0155] In one embodiment, the template matching unit 704 renders the pet activity video segments in the pet activity segments to be rendered set into a pet highlight video by using the matched video template, specifically for: presetting a plurality of video slots in the video template, and integrating a plurality of pet activity video segments with the same or related behavior annotations in the pet activity segments to be rendered set into the plurality of video slots; wherein, one or more of a background music slot, a copywriting slot, and a transition effect slot are also set in the video template.

[0156] In one embodiment, the evaluation annotation unit 702 performs scene annotation on the pet activity video segments according to the video content in the pet activity video segments, specifically for: obtaining a plurality of candidate scenes;

[0157] identifying the video content in the pet activity video segments, and analyzing the scene corresponding to the pet activity video segments from the plurality of candidate scenes according to the video content to perform scene annotation.

[0158] Based on the same concept, the present application further provides a computer device, including a memory and a processor. The computer device may be a terminal, and its internal structure diagram may be as Figure 8 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. When the computer program is executed by the processor, a method for generating a pet highlight video is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0159] Those skilled in the art can understand, Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0160] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0161] The above embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A method for generating a pet highlight video, characterized in that: The method comprises: Get video clips of pet activities; According to the video content in the pet activity video clip, the pet activity video clip is marked with a scene, corresponding evaluation criteria are obtained according to the scene marking, and the pet activity video clip is marked with an evaluation according to the evaluation criteria; According to the behavior of the pet in the pet activity video clip, the pet activity video clip is labeled with a behavior; wherein each of the pet activity video clips obtains one or more behavior labels; Matching a video template from a plurality of pre-set video templates according to the behavior annotation, and determining a set of pet activity segments to be rendered according to the evaluation annotation; wherein the set of pet activity segments to be rendered includes one or more of the pet activity video segments; The matched video template is used to render the pet activity video segments in the set of pet activity segments to be rendered into pet highlight videos.

2. The method for generating a pet highlight video according to claim 1, characterized in that: Matching a video template from a plurality of preset video templates according to the behavior annotation includes: If the set time is reached and / or the number of acquired pet activity video clips is greater than the number threshold, a target behavior label is determined based on the behavior labels of each pet activity video clip; wherein the target behavior label is the behavior label that appears most frequently in the multiple pet activity video clips; A video template is matched from a plurality of preset video templates according to the target behavior annotation.

3. The method for generating a pet highlight video according to claim 2, characterized in that: According to the evaluation annotation, a set of pet activity segments to be rendered is determined, including: Determine, from the pet activity video clips, pet activity video clips annotated with the target behavior annotation, and define them as candidate pet activity video clips; According to the evaluation annotations of the candidate pet activity video clips, a target number of pet activity video clips are screened out from the candidate pet activity video clips, and the target number of pet activity video clips constitute the set of pet activity clips to be rendered; wherein the pre-set video template includes the number of video slots that the video template can support for rendering, and the target number of pet activity video clips in the set of pet activity clips to be rendered is equal to the number of video slots of the matched video template.

4. The method for generating a pet highlight video according to claim 1, characterized in that: Matching a video template from a plurality of preset video templates according to the behavior annotation, and determining a set of pet activity segments to be rendered according to the evaluation annotation, including: Screening the pet activity video clips that have been evaluated and marked, if there is a pet activity video clip whose evaluation and marking meets the set conditions, then the pet activity video clip constitutes the pet activity clip set to be rendered; If there is a template supporting rendering of a single pet activity video clip among the multiple preset video templates, a video template is matched from the preset video templates supporting rendering of a single pet activity video clip according to the behavior annotation of the pet activity video clip.

5. The method for generating a pet highlight video according to claim 1, characterized in that: The method further includes: marking the start and end points of highlights on the pet activity video clips according to the video content in the pet activity video clips; wherein any of the pet activity video clips in the set of pet activity clips to be rendered obtains the highlight start and end markings; The method further includes: setting a negative scene annotation list, excluding the pet activity video segments annotated with scene annotations on the negative list from the range of the pet activity segment set to be rendered; The method further includes: when the result of highlight start and end marking cannot be matched to the video template, excluding the corresponding pet activity video segment from the range of the pet activity segment set to be rendered.

6. The method for generating a pet highlight video according to claim 5, characterized in that: The method further comprises: updating one or more annotation modes among the evaluation annotation, the scene annotation and the highlight start and end annotation; Wherein, updating one or more of the evaluation annotation, scene annotation and highlight start and end annotation includes: The generated pet highlight video is provided to the user, and a user interface is provided to receive one or more feedbacks from the user on the pet highlight video, and the feedback is integrated into one or more annotation methods or matching methods including evaluation annotation, scene annotation, highlight start and end annotation, and video template matching.

7. The method for generating a pet highlight video according to any one of claims 1 to 6, characterized in that: Rendering the pet activity video segments in the set of pet activity segments to be rendered into pet highlight videos using the matched video template includes: Multiple video slots are pre-set in the video template, and multiple pet activity video clips with the same or related behavior labels in the set of pet activity clips to be rendered are integrated into the multiple video slots; wherein, the video template is also provided with one or more of a background music slot, a text slot and a transition effect slot.

8. The method for generating a pet highlight video according to any one of claims 1 to 6, characterized in that: According to the video content in the pet activity video clip, the pet activity video clip is subjected to scene labeling, including: Obtain multiple candidate scenes; The video content in the pet activity video clip is identified, and the scene corresponding to the pet activity video clip is analyzed from the plurality of candidate scenes according to the video content to perform scene annotation.

9. A system for generating a pet highlight video, characterized in that: The system comprises: A terminal configured to collect video clips of pet activities; The cloud is configured to generate a pet highlight video using the method described in any one of claims 1 to 8.

10. A device for generating a pet highlight video, characterized in that: The device comprises: An acquisition unit, used for acquiring video clips of pet activities; an evaluation and annotation unit, configured to perform scene annotation on the pet activity video clip according to the video content in the pet activity video clip, obtain corresponding evaluation criteria according to the scene annotation, and perform evaluation and annotation on the pet activity video clip according to the evaluation criteria; A behavior annotation unit, configured to perform behavior annotation on the pet activity video clips according to the behavior of the pet in the pet activity video clips; wherein each of the pet activity video clips obtains one or more behavior annotations; A template matching unit, configured to match a video template from a plurality of pre-set video templates according to the behavior annotation, and determine a set of pet activity segments to be rendered according to the evaluation annotation; wherein the set of pet activity segments to be rendered includes one or more of the pet activity video segments; The rendering unit is used to render the pet activity video segments in the set of pet activity segments to be rendered into pet highlight videos by using the matched video template.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for generating a pet highlight video according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Video production control method and system based on cloud platform

    CN120916024A

  • A video production control method and system based on a cloud platform

    CN120916024B

  • A pet highlight video generation system, method

    CN122534303A