Animal narrative video generation method and device, computer equipment and medium

By using narrative structure templates and control strategies, pet behavior events are automatically identified and filtered to generate videos with narrative logic, solving the problem of lack of narrative in existing videos and achieving high-quality automated video generation.

CN121509784APending Publication Date: 2026-02-10深圳市灵智无界科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511790079.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies cannot automatically identify and extract pet behavioral events with narrative value, resulting in generated video content that lacks narrative and entertainment value, and has a low degree of automation.

Method used

By introducing narrative structure templates and corresponding control strategies, behavioral events of the target animal are identified from video streams from multiple cameras. Valid behavioral events that satisfy logical relationships are selected, and the duration of the episodes is set according to the narrative logic weight. The best perspective and visual effects are selected, and the episodes are spliced ​​sequentially to generate a narrative video.

Benefits of technology

It achieves fully automated generation from raw video streams to high-quality narrative videos. The generated videos have professional narrative structures and are visually appealing, highlighting the key points and pacing of the videos, thus improving the quality of pet video recordings and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509784A_ABST
    Figure CN121509784A_ABST
Patent Text Reader

Abstract

The invention relates to an animal narrative video generation method and device, computer equipment and a medium, and the method comprises the steps: recognizing different behavior events of a target animal from video streams of a plurality of cameras associated with the same spatial region identifier; based on a logic control strategy in the narrative structure template, screening out a plurality of effective behavior events meeting a corresponding logic relationship from different behavior events; endowing each effective behavior event with a corresponding narrative logic weight according to the logic relationship, and correspondingly setting plot duration; on the basis of a lens control strategy in the narrative structure template, selecting a fragment source matched with a view angle from a plurality of video streams in an event time period for each effective behavior event, and extracting continuous corresponding plot fragments from the fragment source; and according to a packaging control strategy in the narrative structure template, performing time sequence splicing on each plot video clip according to a narrative sequence determined by a logical relationship to generate a narrative video. According to the invention, the video with the narrative capability can be automatically generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for generating animal narrative videos, as well as computer equipment and media. Background Technology

[0002] With the increasing popularity of pet ownership, the demand for pet owners to record and share their pets' daily lives continues to grow. Currently, using smart devices in the home environment to record pet videos has become a common practice. Traditional technical solutions mainly rely on two modes: one is that users actively use mobile devices such as smartphones to manually record; the other is to record continuously through multiple pre-deployed home surveillance cameras, which can then be reviewed and viewed by the user afterward.

[0003] However, the aforementioned traditional technical solutions have inherent limitations in automating and story-based recording of pets' daily behaviors. Firstly, from a technical perspective, both manual filming and fixed monitoring are essentially passive recording or simple capture of video streams. Manual filming relies heavily on the user's immediate reaction and operation, making it difficult to continuously capture the unpredictable and fleeting moments of a pet, resulting in a high miss rate. While fixed monitoring can achieve uninterrupted recording, the generated video content is essentially a long, continuous time-stream data, lacking semantic understanding and structured processing of the video content itself.

[0004] Based on the aforementioned passive recording technology, a series of technical problems have arisen. The most critical issue is that none of these solutions can automatically identify and extract independent behavioral events with narrative value from the raw video stream. Because they cannot understand the semantics of behavioral units such as "peeping," "jumping," and "playing," nor can they recognize the possible logical connections between these behavioral events, the resulting video material can only be a linear arrangement or mechanical editing based on a simple chronological order. This simple splicing based on timestamps cannot organically organize discrete behavioral fragments into a coherent narrative structure with a clear beginning, development, climax, and conclusion, making it difficult to generate truly story-driven videos. Therefore, existing technical solutions consistently face technical bottlenecks such as low automation, lengthy and tedious video content, and a final product lacking narrative and entertainment value. Summary of the Invention

[0005] The primary objective of this application is to address at least one of the aforementioned problems by providing a method, apparatus, computer device, and medium for generating animal narrative videos.

[0006] To achieve the various objectives of this application, the following technical solution is adopted: A method for generating animal narrative videos, provided for one of the purposes of this application, includes the following steps: Identify different behavioral events of a target animal from video streams from multiple cameras associated with the same spatial region identifier; Based on the logical control strategy in the narrative structure template, multiple valid behavioral events that satisfy the corresponding logical relationships are selected from the different behavioral events; Each valid behavioral event is assigned a corresponding narrative logic weight based on the logical relationship, and a corresponding plot duration is set for each narrative logic weight. Based on the camera control strategy in the narrative structure template, for each valid behavioral event, a segment source with a matching perspective is selected from multiple video streams within its event time period, and a plot segment with a continuous corresponding plot duration is extracted from it. According to the encapsulation control strategy in the narrative structure template, the various plot video segments are spliced ​​together in chronological order according to the narrative sequence determined by the logical relationship to generate a narrative video.

[0007] An animal narrative video generation apparatus, proposed to meet one of the purposes of this application, includes: The video recognition module is configured to identify different behavioral events of a target animal from video streams from multiple cameras associated with the same spatial region identifier; The logic control module is configured to filter out multiple valid behavior events that satisfy the corresponding logical relationships from the different behavior events based on the logic control strategy in the narrative structure template. The duration control module is configured to assign a corresponding narrative logic weight to each valid behavioral event according to the logical relationship, and set a corresponding plot duration for the narrative logic weight. The camera control module is set to use a camera control strategy based on the narrative structure template to select a perspective-matching segment source from multiple video streams within the event time period for each valid behavioral event, and extract a plot segment that lasts for the corresponding plot duration. The encapsulation control module is configured to sequentially splice together various plot video segments according to the narrative order determined by the logical relationship, based on the encapsulation control strategy in the narrative structure template, to generate a narrative video.

[0008] On the one hand, a computer device provided for one of the purposes of this application includes a processor and a memory, wherein the processor invokes and runs a computer program in the memory to perform the steps of the animal narrative video generation method.

[0009] In another aspect, a computer-readable storage medium is provided to suit another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the described animal narrative video generation method, which, when invoked by a computer, performs the steps included in the corresponding method.

[0010] Compared to traditional technologies, this application fundamentally changes the traditional paradigm of passively capturing and simply splicing pet videos by introducing narrative structure templates and corresponding control strategies. Its most significant benefit lies in achieving automated and intelligent generation from raw video streams to finished videos with narrative logic. Specifically, by filtering identified behavioral events through logical control strategies, it can automatically identify effective behavioral events with narrative potential and construct a storyline based on their inherent logical relationships. This effectively solves the problem of lengthy and unfocused videos caused by traditional technologies that can only arrange segments chronologically, resulting in videos with a narrative structure and watchability similar to those of a human director.

[0011] Furthermore, by assigning narrative logic weights to each valid behavioral event and setting the episode duration accordingly, this application can automatically determine the level of detail in different parts of the video. High-scoring, compelling moments are given longer display time, while secondary or transitional content is shortened accordingly, thus automatically creating a balanced pacing in the video, highlighting the climax and core content of the story, and avoiding the egalitarianism and blandness of traditional narrative videos. In addition, based on the camera control strategy, it automatically selects source clips with matching perspectives for different behavioral events, mimicking the multi-camera collaboration approach in professional filming, ensuring that each narrative unit is presented from the best angle. Finally, the encapsulation control strategy splices and packages the selected high-quality shots sequentially according to the narrative order determined by logical relationships, smoothly organizing them into a complete narrative video file.

[0012] Therefore, this application transforms the original time-consuming and labor-intensive video creation process, which relies on professional skills, into a fully automated and highly efficient intelligent processing, greatly improving the quality of pet video recordings and user experience, and completely overcoming many technical defects revealed in the background technology. Attached Figure Description

[0013] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating a typical embodiment of the animal narrative video generation method of this application; Figure 2 This is a schematic diagram of the animal narrative video generation device of this application; Figure 3 This is a schematic diagram of the structure of a computer device used in this application. Detailed Implementation

[0014] The animal narrative video generation method provided in this application can be implemented as a computer program, and is particularly suitable for deployment in home network scenarios with smart home environments. In this typical application scenario, multiple network cameras are pre-deployed in different functional areas of the residence, such as the living room, dining room, balcony, and corridor—spaces where pets are frequently active. These cameras are connected to a computer device acting as a local computing node via the home LAN. This node can be a smart central control device, a home gateway, a network-attached storage device, or a dedicated edge computing box, which is responsible for receiving and processing real-time video streams from each camera. Alternatively, it can be implemented on a cloud platform, with a cloud server undertaking the core computing tasks. The technical solution of this application automatically creates short videos with a cinematic narrative quality that condense the essence of a pet's daily activities through intelligent analysis of multiple video streams, thereby meeting the owner's needs for recording and sharing.

[0015] Before delving into the specific implementation of this application, it is necessary to provide a unified explanation of several fundamental concepts that run through this application.

[0016] The narrative structure template referred to in this application integrates multiple control strategies, including logic control, camera control, and encapsulation control. These strategies work together to guide the generation of narrative videos. For ease of understanding, a typical template for pet daily life, such as the "comedic mishap" template, is provided. In this template, the logic control strategy predefines a series of logical relationship rules to identify and filter sequences of behavioral events with narrative value. These rules include, but are not limited to, the following categories: first, causal chain relationships, used to identify a series of actions with clear intent and coherence, such as "spying on a target," "sneaking closer," and "leaping to pounce"; second, dramatic conflict relationships, used to capture moments that create suspense or turning points, such as "successfully stealing food" followed by "being discovered by the owner"; and third, contrast humor relationships, suitable for linking behaviors that create strong contrasts, such as "excitedly chasing" followed by "slipping awkwardly." Through these rules, complete micro-story clues, such as "attempted theft" or "failed chase," can be constructed from a large number of raw behaviors.

[0017] The camera control strategy is responsible for planning the optimal visual presentation for the behavioral events selected by the logic control strategy. This strategy predefines shooting rules that match different behavior types, covering aspects such as perspective, framing, and even camera movement. For example, for the behavior of "lurking and approaching," the strategy might prioritize a low-angle tracking shot to simulate the pet's subjective perspective and create tension; for the climactic behavior of "missing and falling," it mandates a close-up shot combined with slow motion to maximize the comedic impact; and for the subsequent "awkward grooming," a static side close-up can be specified to capture its subtle "feigned composure" expression. The strategy further sets quality scoring criteria for each perspective, including composition, sharpness, and lighting, ensuring the automatic selection of the best shot.

[0018] The encapsulation control strategy is ultimately responsible for assembling the selected shot clips into the final video product, ensuring the rhythm and emotional impact of the finished film. This strategy first arranges the clips according to the narrative order determined by logical relationships, such as constructing a storyline in the order of "lurking-attacking-failure-embarrassment." Based on this, the strategy predefines transition rules that match the narrative rhythm, such as using gentle fade-in / fade-out effects when changing perspectives, and employing rapid jump cuts or wipes when scenes or emotions undergo dramatic shifts. Simultaneously, the strategy can also include audio matching rules to automatically select a light and humorous background music based on the overall tone of the story, such as for the aforementioned "comedic mishap" story, aligning the editing points with the music's beat to achieve a synchronized audio-visual effect.

[0019] As can be seen from the examples above, the narrative structure template essentially transforms professional video narrative knowledge into a set of calculable and executable rules. This set of rules, combined with the control methods of this application, enables computer devices to understand the content of the action, plan the shooting techniques, and perform post-editing, just like a human director, thereby automatically generating high-quality storytelling videos.

[0020] The following detailed description of each step of the method in this application will gradually reveal how to utilize the above-mentioned architecture and concepts to achieve the entire process of automatically and intelligently generating animal narrative videos.

[0021] Please see Figure 1 In some embodiments, the animal narrative video generation method of this application can be implemented as an application running on a computer device, and the method includes: Step S3100: Identify different behavioral events of the target animal from video streams of multiple cameras associated with the same spatial region identifier; Multiple cameras in the same smart home network are typically pre-configured by the user and associated with the same spatial area identifier. This spatial area identifier can be represented by various data, such as the user's spatial map identifier, network identifier, group identifier, user identity identifier, etc. By establishing a mapping relationship between these spatial area identifiers and each camera, it can be known that the video streams of the corresponding multiple cameras are deployed in the same spatial environment. Therefore, the video streams generated by these cameras can be collected centrally, and the continuous, unstructured video data can be transformed into discrete behavioral units with semantic tags, so as to generate the narrative video required by the cost application in the future.

[0022] In one embodiment, behavioral events can be analyzed from video streams in real time, meaning that video frames are continuously processed by edge computing devices or cloud servers as the video stream is generated. In another embodiment, a post-processing approach can be used, where multiple video streams are first synchronized and stored in a distributed storage system, and then computing nodes perform batch analysis on the stored video files. Regardless of the timing processing method used, behavioral event identification relies on computer vision and deep learning technologies.

[0023] Specifically, the identification process can include multiple sub-steps such as object detection, individual identification, and behavior classification. First, the target animal needs to be located and segmented from the video frames. In one implementation, a deep learning-based object detection model, such as YOLO or Faster R-CNN series models, can be used to extract frames from the video stream, detect and output bounding boxes containing the target animal frame by frame. For multi-pet households or scenes with multiple animals, image segmentation models like Unet can be used for animal instance segmentation to distinguish different individuals. Subsequently, a discriminative feature representation needs to be generated for each identified animal individual. This can be achieved by extracting the target animal's appearance features, for example, by using a convolutional neural network to encode the segmented animal image region to generate a high-dimensional feature vector. Furthermore, other features can be fused, including but not limited to skeletal keypoint features to describe the animal's posture, and motion features calculated using optical flow to describe its movement patterns. Any one or any combination of these features constitutes the animal's identity, providing a basis for cross-camera tracking and behavior attribution.

[0024] After successfully detecting and identifying individual animals, the next step is to identify their specific behaviors. Behavior recognition is a typical temporal classification task. In one embodiment, a temporal behavior recognition model, such as the SlowFast network or the X3D model, can be used. This type of model analyzes a continuous sequence of image frames, such as 16 consecutive frames, using a sliding window approach. It comprehensively examines the appearance and motion changes between frames and finally outputs the behavior type and its confidence level corresponding to the behavioral events that occurred within the time period corresponding to the image frame sequence. Each identified behavioral event is marked with a start timestamp, and its end timestamp can also be further marked to define a complete behavioral event. In addition to behavior type and time information, the spatial location of the behavioral event can be estimated or recorded. For example, through camera calibration and coordinate mapping techniques, the 2D position in the image can be converted into 3D coordinates in the actual physical space. This position information can be used to assist in subsequent shot selection strategies.

[0025] Step S3200: Based on the logical control strategy in the narrative structure template, select multiple valid behavioral events that satisfy the corresponding logical relationships from the different behavioral events; Based on the logical control strategy in the narrative structure template, multiple valid behavioral events that satisfy the corresponding logical relationships can be selected from the different behavioral events identified in the preceding steps. In this way, the discrete, time-sequential sequence of behavioral events can be organized into a story-like plot chain according to the preset narrative logic.

[0026] Logical control strategies are essentially a predefined set of logical relationship rules, which define the types of narrative associations allowed between events of different behavioral types. In one embodiment, these logical relationship types can include causal chains to identify a series of behaviors with cause-and-effect connections, such as a pet's complete hunting intention chain from spying on a target to stealthily approaching and then pouncing. In another embodiment, dramatic conflict relationships can be included to capture unexpected turns in behavioral development, such as a successful theft being abruptly interrupted by the owner's discovery. Furthermore, contrast humor relationships can be included to associate consecutive behaviors with stark contrasts in emotion or outcome, such as an exhilarating chase followed immediately by a clumsy fall.

[0027] In the specific screening process, firstly, all identified behavioral events are arranged chronologically to form a linear sequence of behavioral events. Next, this sequence is matched against predefined logical relationship rules in the narrative structure template. This matching process aims to find rule combinations from the rule base that can explain or fit the current arrangement pattern of the behavioral event sequence. For example, it checks whether there exists a subset of consecutive events conforming to the causal chain rule "Action A—Action B—Action C," or whether there exists a pair of adjacent events conforming to the dramatic conflict rule "Positive Action—Negative Turning Point Action." If so, the behavioral event sequence has achieved a match with the corresponding logical relationship rule.

[0028] By matching, one or more target logical relationship rules applicable to the current context are determined. Based on this, events in the sequence of behavioral events whose behavior types conform to the determined target logical relationship rules are marked as valid behavioral events. These valid behavioral events are considered narrative units that can contribute to the overall narrative structure. Conversely, events whose behavior types cannot be explained by any logical relationship rules and are narratively isolated are filtered out by the system and do not enter the subsequent video generation process.

[0029] Step S3300: Assign a corresponding narrative logic weight to each valid behavioral event according to the logical relationship, and set a corresponding plot duration for the narrative logic weight; Based on the logical relationships carried by the target logical relationship rules determined in the previous step, each selected valid behavioral event can be assigned a corresponding narrative logical weight. Based on this narrative logical weight, the plot length of each event in the final video can be set so as to quantify the importance of each behavioral event in the overall narrative and allocate resources accordingly, thereby controlling the rhythm and focus of the final video.

[0030] The calculation of narrative logic weights can be based on a multi-factor weighted evaluation system. In one embodiment, each behavioral event is assigned a corresponding basic weight in advance based on its behavioral type, and this basic weight can be predefined by the narrative structure template. For example, rarer or visually impactful behaviors, such as a failed jump, can be assigned a higher basic weight; while everyday behaviors, such as quietly grooming, correspond to a lower basic weight.

[0031] Building upon this foundation, logical relationships serve as the key factor for dynamically adjusting weights. The position and role of the action / event within the narrative chain constructed by logical relationships are examined. Pre-defined weight adjustment factors are matched according to different roles, and the narrative logical weight is obtained by superimposing the base weight with the weight adjustment factors. In one implementation, weight adjustment rules can also be set. For example, the weight of an action at the end of a causal chain, serving as the climax of an event, can receive a significant boost; similarly, the weight of an action that triggers a dramatic conflict or turning point can also be increased accordingly. Through this mechanism, the narrative logical weight is no longer a static value but reflects the relative importance of the event within a specific story context.

[0032] After obtaining the narrative logic weight of each valid behavioral event, the next step is to map it to a specific episode duration. This can be implemented based on a preset mapping relationship. A typical mapping relationship is configured such that the higher the narrative logic weight, the longer the episode duration is allocated. For example, events with extremely high weights can be allocated several seconds of duration and can be rendered with slow-motion effects; while events with lower weights can be allocated only one or two seconds of duration as transitional segments in the narrative. This mapping relationship can be linear or a non-linear piecewise function, the purpose of which is to ensure that the narrative rhythm of the video is well-paced and highlights key points.

[0033] In another embodiment, the plot duration can be allocated based on a preset total duration using preset narrative logic weights. Specifically, the sum of the narrative logic weights of all valid behavioral events can be used as the denominator, the narrative logic weight of each valid behavioral event can be used as the numerator, the ratio of the numerator to the denominator can be used as the normalized weight of the valid behavioral event, and multiplied by the total duration to obtain the corresponding plot duration.

[0034] It can be seen that by transforming abstract narrative logic relationships into quantifiable weights, and then into specific video duration allocations to obtain plot durations, intelligent rhythm control of video content is achieved.

[0035] Step S3400: Based on the camera control strategy in the narrative structure template, select the perspective-matching segment source from multiple video streams within the event time period for each valid behavioral event, and extract the plot segment with the corresponding plot duration from it. Based on the shot control strategy in the narrative structure template, the source of the segment with the matching perspective can be selected from multiple video streams within the event period for each effective action event, and the plot segment with the corresponding plot duration can be extracted from it. This imitates the shot selection process of a professional director and selects the best visual presentation for each narrative unit.

[0036] The primary function of the camera control strategy is to associate the semantics of effective behavioral events with the optimal shooting angle. Therefore, firstly, based on predefined rules within the strategy, one or more target angles are determined for each type of behavior in an effective behavioral event. For example, for dynamic behaviors like playing and running, the strategy might prioritize side-following shots or low-angle shots to emphasize dynamism and power; while for static behaviors like sleeping or eating, a close-up or overhead shot could be specified to capture details and expressions. These rules form the fundamental basis for camera selection.

[0037] After identifying the target perspective of the behavioral event, it is necessary to filter candidate video streams that can provide the target perspective from the video streams of all available cameras within the time period corresponding to the event. In one embodiment, this can be achieved through physical location calibration of the cameras and scene space modeling, thereby determining the field of view and shooting angle of each camera. In another embodiment, the shooting angle of image frames in the video stream can be determined by directly obtaining the image control parameters provided in advance by each camera for its corresponding video stream.

[0038] Subsequently, the selected candidate video streams can be evaluated for picture quality to select the optimal segment source. The evaluation system can include multiple quantitative scoring dimensions. These dimensions include, but are not limited to, composition scoring, which assesses whether the position of the subject in the frame conforms to aesthetic rules; sharpness scoring, which judges whether the image is accurately focused and clear in detail; perspective scoring, which measures the fit between the current image and the target perspective; and lighting scoring, which checks whether the image is properly exposed and without obvious overexposure or underexposure. In one implementation, video image frames within the event period can be sampled, and the above-mentioned scores can be calculated for each frame. Then, a comprehensive picture quality score can be calculated for each candidate video stream using a weighted average or other data fusion method.

[0039] Finally, the candidate video stream with the highest overall image quality score is selected as the source of the segment for this behavioral event. After selecting the source, a video segment of the corresponding length is extracted from the source according to the plot duration set for the event. For example, if the plot duration of an event is 5 seconds, then 5 seconds of video content is extracted from the selected optimal camera video stream within the event period, starting from the pre-marked start timestamp of the behavioral event, as the final plot segment used for narrative splicing.

[0040] Step S3500: According to the encapsulation control strategy in the narrative structure template, the various plot video segments are spliced ​​sequentially according to the narrative order determined by the logical relationship to generate a narrative video.

[0041] Based on the encapsulation control strategy in the narrative structure template, the various plot video segments can be spliced ​​together in chronological order according to the narrative sequence determined by the logical relationship revealed above, generating the final narrative video file. This allows the independent plot segments obtained in the pre-processing stage to be combined into a coherent, smooth, and impactful complete video work through professional post-production techniques.

[0042] The encapsulation control strategy first guides how to arrange the plot segments corresponding to each effective action event. Specifically, the plot segments corresponding to each effective action event can be organized according to the narrative order determined by logical relationships, rather than a simple chronological order. For example, for a complete story chain containing cause, development, climax, and ending, the segments will be arranged according to this logic, thus forming an initial editing sequence with narrative tension.

[0043] Subsequently, predefined transition rules from the encapsulation control strategy are applied. These rules specify the transition effects to be used in different contexts. In one embodiment, a smooth fade-in / fade-out effect can be used to achieve a smooth transition when switching between different camera angles of the same action. In another embodiment, when the narrative sequence indicates a drastic change in scene or emotion, such as switching from warm play to tense confrontation, a more impactful transition such as a rapid jump cut or wipe can be used. Appropriate transition effects can be automatically matched and added based on the attributes of adjacent plot segments (such as action type or the type of logical relationship rules between them).

[0044] Simultaneously, the audio matching rules in the encapsulation control strategy can be applied synchronously. These rules in the encapsulation control strategy indicate the overall tone presented by analyzing the entire narrative sequence, such as cheerful, heartwarming, suspenseful, or humorous. Based on this tone, matching background music and sound effects are automatically selected from a pre-built audio library. In one implementation, audio synchronization processing can also be performed, aligning video clip points with the music beat to achieve a synchronized audio-visual effect, thereby greatly enhancing the video's visual appeal and professionalism.

[0045] After combining visual and auditory elements, the initial edit sequence with overlaid transition effects and audio information is uniformly rendered and encoded, combining all elements into a standard video file format, and outputting the final narrative video, which can be played and displayed by users at any time.

[0046] As can be seen from the above embodiments, this application, by introducing a narrative structure template, integrates the originally discrete technical modules into an organic and intelligent creative system, possessing multiple technical advantages, including but not limited to: First, this application achieves fully automated generation from raw video streams to high-quality narrative videos. Traditional technologies can only provide lengthy recordings or rely on tedious manual editing, while this application automatically filters effective behavioral events and constructs storylines through logical control strategies, intelligently selects the best perspective through camera control strategies, and automatically completes editing and packaging through encapsulation control strategies, completely liberating users from complex technical operations and achieving an unprecedented level of automation.

[0047] Secondly, based on automation, this application brings about a qualitative leap in narrative ability. Videos generated by traditional technologies are essentially simple chronological lists, lacking storytelling and artistic appeal. This application, through the application of narrative structure templates, can understand the semantic connections between events, such as cause and effect, and transitions, thus enabling the construction of a story structure with introduction, development, climax, and conclusion, much like a director. Simultaneously, by dynamically allocating plot durations through narrative logic weights, it can highlight the climax of the story and control the narrative pace. The resulting video is no longer a mere chronicle, but a miniature documentary with professional narrative pacing and emotional expressiveness, greatly enhancing the content's viewing value.

[0048] Furthermore, another significant advantage of this application lies in its intelligent and precise processing, which is reflected on multiple levels: at the content analysis level, it comprehensively utilizes advanced AI models such as object detection, feature extraction, and temporal behavior recognition to ensure the accuracy of behavior recognition; at the shot selection level, it not only considers perspective matching but also introduces a multi-dimensional image quality evaluation system to ensure that each segment originates from the best image quality; at the resource allocation level, it quantifies narrative importance into weights and maps them to plot duration, achieving precise allocation of computational and attentional resources. This comprehensive intelligent processing guarantees the superiority of the final output in terms of content value and visual quality.

[0049] Finally, this application demonstrates strong flexibility and scalability. The narrative structure template, as a rule base, allows for pre-definition or post-production expansion of content such as logical relationship types, shot preferences, and transition rules, based on different pet types, user preferences, or video styles such as comedic or heartwarming. This means that the same technical architecture can adapt to diverse creative needs, generating video works of different styles. Furthermore, this application supports both real-time processing and post-processing batch processing, and can be deployed on local edge computing devices or run in the cloud. This flexibility in implementation allows it to easily adapt to different application scenarios and hardware conditions.

[0050] Based on any embodiment of the method in this application, different behavioral events of a target animal are identified from video streams of multiple cameras associated with the same spatial region identifier, including: Step S3110: Detect the video streams from the multiple cameras, identify the behavioral events of the target animal, and determine the behavior type, start timestamp, and spatial location of each behavioral event; The video streams continuously generated by multiple cameras in this application can be analyzed in real time. By analyzing the image information in each video stream in real time, the behavioral events of the target animal can be identified so as to obtain the key attributes of each behavioral event, including but not limited to the behavior type, start timestamp and its location in space.

[0051] In one embodiment, a preset object detection model is first used to process the image frames obtained after frame extraction from the video stream, identifying the bounding boxes belonging to the target animal. The object detection model can be a deep learning-based model, such as the YOLO series or Faster R-CNN series, which can efficiently locate the target animal's region in each image frame. Subsequently, based on the identified bounding boxes, the images within the boxes are extracted from the original image frames, and a preset image segmentation model, such as U-Net or Mask R-CNN, is further used to perform fine instance segmentation of the images within the boxes, obtaining pixel-level precision animal image regions. This step is particularly important for families with multiple pets, as it effectively distinguishes different individual animals.

[0052] After obtaining the segmented animal image, it is necessary to extract a feature vector to identify the animal. The extracted image features include, but are not limited to, the following categories: first, appearance features, such as features extracted through a pre-trained convolutional neural network that characterize the animal's coat color, texture, and body contour; second, skeletal keypoint features, which are the coordinates of joints such as the animal's head, limbs, and tail extracted through a keypoint detection model, and then the posture features calculated; and third, motion features, which describe the animal's direction and speed of movement, calculated through optical flow or inter-frame difference methods. These features can be used individually or fused together to form a feature vector that uniquely identifies the animal.

[0053] Based on object detection and feature extraction, a pre-defined temporal behavior recognition model is used to identify specific behavior types. This model analyzes a continuous sequence of image frames in a video stream using a sliding window approach. For example, the model can select 16 consecutive frames as an analysis unit. The feature vectors corresponding to each image frame are arranged chronologically to form a feature vector sequence, which serves as the input to the temporal behavior recognition model. This model, such as a SlowFast network or an X3D model, analyzes the temporal changes of features in this sequence to output the behavior type corresponding to the behavioral event that occurred within that time period, such as jumping, eating, or playing, and records the start timestamp of the behavioral event.

[0054] Furthermore, the spatial location of the behavioral event can be estimated. In one implementation, by pre-calibrating the cameras and establishing a mapping between the image coordinate system and the actual physical space, the position of the target animal in the image can be converted into three-dimensional coordinates within the home environment. In another implementation, the spatial position of the animal can be calculated more accurately by combining the perspectives of multiple cameras and using triangulation or multi-view geometry principles. This positional information is of significant reference value for selecting the optimal shooting angle in subsequent lens control strategies.

[0055] Step S3120: After identifying the behavioral event, generate a perspective adjustment instruction based on the preferred camera angle specified in the camera control strategy in the preset narrative structure template according to the behavior type, and drive the corresponding camera to enter a continuous tracking shooting state following the location where the behavioral event occurs according to the instruction. The camera's shooting state can be proactively adjusted according to the type of behavior in order to achieve optimized capture of behavioral events. Specifically, this can be achieved based on the camera control strategy in the narrative structure template, which predefines the preferred camera angles for different types of behavior.

[0056] Specifically, by querying the camera control strategy, one or more preferred camera angles associated with the currently identified behavior type can be determined. For example, for dynamic behaviors such as playing and running, the strategy can specify a side-view follow-up shot as the preferred camera angle to emphasize dynamism; for static behaviors such as sleeping and resting, a frontal close-up shot can be specified to capture details. Using these preferred camera angles as the target angles, specific angle adjustment instructions are generated. In one embodiment, these instructions may include parameters such as the gimbal rotation angle and / or optical zoom, designed to drive the camera to adjust its orientation and focal length so that its field of view continuously covers the target animal from the target angle, always pointing at the location where the behavioral event occurs.

[0057] After generating the perspective adjustment command, it is sent to the corresponding camera. The camera then enters a continuous tracking shooting state based on the command. In this state, the camera is not stationary but enters a dynamic working mode. In one embodiment, the camera can maintain lock on the target animal in subsequent frames based on a lightweight target tracking algorithm, and fine-tune the gimbal and zoom in real time to ensure the target is always in the ideal position in the frame and maintains the required perspective. In another embodiment, for a fixed-position camera, the video streams from multiple cameras can be combined, and the video stream from the camera that best provides the target's perspective can be selected as the primary signal source to achieve logical tracking.

[0058] Step S3130: When the behavior event is determined to have ended, exit the tracking and shooting state and record the termination timestamp. The start timestamp and the termination timestamp define the corresponding event time period.

[0059] In the process of identifying behavioral events, the ending state of the behavioral event can also be identified, and the event time period corresponding to the behavioral event can be finally defined, while releasing the system resources occupied for tracking the event.

[0060] The determination of the end of a behavioral event can be based on various conditions. In one embodiment, it can directly rely on the output of a temporal behavior recognition model. When the model analyzes subsequent consecutive image frame sequences and no longer outputs the same behavior type as the behavioral event, or outputs a special marker indicating the end of the behavior, the behavioral event can be determined to have ended. In another embodiment, a timeout threshold for continuous inactivity can be set. For example, if no valid action matching a preset behavior library is detected on the target animal within a preset time window, such as 2 seconds, the current behavioral event can be determined to have ended. Furthermore, for certain behaviors with a clear termination state, such as an animal leaving its food bowl at the end of a feeding behavior, detecting the target animal leaving the specific area where the behavior occurred can be used as an end signal.

[0061] Once the behavioral event is determined to have ended, a corresponding control command is generated, driving the corresponding camera in continuous tracking mode to exit that state. This means the camera stops actively adjusting its position, such as gimbal rotation and optical zoom, to track this specific event, and returns to standby, a preset fixed viewing angle, or a state awaiting the next behavioral event trigger command. This helps reduce system power consumption and avoids ineffective tracking of subsequent unrelated scenes.

[0062] Upon exiting the tracking and recording state, the termination timestamp of the event can be recorded. This termination timestamp, together with the start timestamp recorded in step S3110, defines a complete event time period. This event time period precisely defines the time range in which the specific behavior occurs in the original video stream, theoretically defining a key video segment and providing an indispensable time index for accurately extracting the corresponding plot segment from the video stream subsequently.

[0063] Through the above embodiments, this application integrates real-time behavioral event recognition and dynamic perspective tracking control, realizing a paradigm shift from passive monitoring to active narrative capture. Its advantage lies in constructing an intelligent shooting system with closed-loop feedback: by triggering a lens adjustment command based on narrative rules the instant a behavioral event is recognized, the camera can be driven into a continuous tracking state, rather than relying solely on the camera's traditional motion detection function, ensuring the capture of high-quality original material from the best perspective; at the same time, by accurately determining the end point of the behavioral event and automatically exiting the tracking state, not only is resource utilization efficiency optimized, but more importantly, each behavioral event is precisely encapsulated into a standardized narrative unit with clear start and end timestamps, clear spatial location, and corresponding optimal lens perspective. This provides a structured input with accurate timing, optimized perspective, and complete content for subsequent video generation based on logical relationships, thereby improving the coherence and professionalism of the final narrative video from the source.

[0064] Based on any embodiment of the method in this application, and based on the logical control strategy in the narrative structure template, multiple valid behavioral events that satisfy the corresponding logical relationships are selected from the different behavioral events, including: Step S3210: Arrange the identified behavioral events in chronological order to form a behavioral event sequence; The identified, time-discrete behavioral events can be arranged in chronological order to form a linear sequence of behavioral events. This transforms the original, unordered set of events into a standardized data structure that facilitates temporal relationship analysis by computer programs.

[0065] In one embodiment, the starting timestamp recorded when each behavioral event is identified is read, and the events are sorted according to the numerical value of these timestamps. This sorting method based on absolute timestamps can accurately reflect the true time sequence of events.

[0066] In another embodiment, the relative temporal relationship between behavioral events can be considered. For example, when processing video streams from different cameras, even if time synchronization has been performed, the order of some events may be very close. In this case, in addition to the start timestamp, the duration or end timestamp of the event can be introduced as an auxiliary criterion to ensure the accuracy of the sequence, for example, arranging the events that end first before the events that end later.

[0067] Accordingly, individual behavioral events, such as a pet eating, jumping, and playing at different times, are linked together into a continuous sequence of behavioral events unfolding chronologically. This sequence provides direct input for subsequent steps using logical control strategies for analysis, allowing for the examination of potential relationships between these events, much like reading a story outline that unfolds over time.

[0068] Step S3220: Match the sequence of behavioral events with multiple logical relationship rules contained in the predefined logical control strategy in the narrative structure template to determine the applicable target logical relationship rule. The logical relationship rule is used to define the types of logical relationships that are allowed between behavioral events of different behavioral types. By matching the sorted sequence of behavioral events with predefined logical relationship rules in the narrative structure template, narrative patterns can be identified within the sequence, thereby determining one or more target logical relationship rules applicable to the current sequence. This process is essentially a crucial step in assigning narrative meaning to an unordered sequence of events.

[0069] Logical relationship rules are a component of logical control strategies. These rules define, in a machine-readable form, the permissible narrative associations between events of different behavior types. Each rule contains a conditional part and a conclusion part. The conditional part specifies what sequence of behavior types can trigger the rule, while the conclusion part defines the type of logical relationship represented by this sequence. For example, a rule might state that if a sequence of behavior events consecutively contains behavior types A, B, and C, then this set of events constitutes a causal chain.

[0070] To facilitate understanding, several typical logical relationship rules and their data content can be specifically described. A common rule is the causal chain relationship rule. This type of rule is used to identify a series of behavioral events with a cause-and-effect relationship. For example, a predefined causal chain rule can be specifically described as follows: when the three behavioral event sequences of spying on a target, stealthily approaching, and leaping to attack appear sequentially, it should be determined that a complete hunting intent chain is formed, i.e., a causal chain relationship. The data form of this rule in the template can be an ordered list of behavioral types [spying, stealth, leaping], associated with the causal chain logical relationship type identifier.

[0071] Another important rule is the dramatic conflict relationship rule. This type of rule focuses on capturing unexpected twists in the development of events to create narrative tension. A practical example can be described as follows: when a sequence of behavioral events first involves a successful act of stealing food, followed immediately by an act of being discovered by the owner, the latter should be considered to constitute dramatic conflict to the former. The data for this rule can be represented as a binary pair (successful stealing, discovered by the owner), associated with the type of dramatic conflict relationship.

[0072] In addition, contrast humor relationship rules can be included. These rules are used to associate consecutive behaviors that present a stark contrast in emotion or outcome to create a comedic effect. A practical example could be defined as follows: when a type of behavior involving an enthusiastic chase is immediately followed by a type of behavior involving a clumsy slip, a contrast humor relationship should be considered formed. The data for this rule can also be a binary pair (enthusiastic chase, clumsy slip), corresponding to the contrast humor relationship type.

[0073] The matching process can be implemented using pattern matching algorithms. In one embodiment, the behavior types in a sequence of behavioral events can be extracted sequentially to form a behavior type string or list. This sequence is then compared with the behavior type patterns defined by each rule in the rule base. When a subsequence in the sequence completely matches the pattern of a rule, the rule is considered a match and thus identified as the target logical relationship rule. In another embodiment, the matching may not be precise, allowing for certain intervals or intermediate events. The rule with the highest degree of matching is determined by calculating sequence similarity or using more complex temporal pattern recognition algorithms.

[0074] Step S3230: Mark the behavior events in the behavior event sequence whose behavior type conforms to the target logical relationship rule as valid behavior events, and filter out isolated behavior events whose behavior type does not conform to any logical relationship rule.

[0075] The sequence of behavioral events is traversed to identify those whose behavioral types participate in a pattern defined by any matched target logical relationship rule. For example, if the target rule is a causal chain rule with the pattern [spying, stealth, attack], then the three events in the behavioral event sequence with behavioral types of peeping, stealth, and attacking are marked as valid behavioral events because their behavioral types conform to the rule pattern. These valid behavioral events are considered to constitute the core material of the narrative fragments.

[0076] Events in a sequence of behavioral events whose behavior types fail to match any of the matched target logical relationship rules are classified as isolated behavioral events by the system. In one embodiment, these isolated events can be simply removed from the candidate set to be processed or ignored. In another embodiment, a specific identifier, such as "unused" or "isolated," can be added to them, and their metadata can be stored in a separate log for subsequent analysis, while ensuring that they do not enter the subsequent video generation process.

[0077] Through the above embodiments, this application constructs a sequence of behavioral events and intelligently matches and filters them with a predefined logical relationship rule base. This enables the automatic identification and extraction of core events with narrative value from massive discrete behaviors. Its technical advantage lies in transforming a linear event stream based on timestamps into a narrative material library based on semantic logic. Through precise pattern matching algorithms, such as causal chains, dramatic conflict, and contrast humor rules, it effectively distinguishes between narrative-related events and isolated events unrelated to the narrative. This not only ensures the coherence and logic of the final video storyline but, more importantly, establishes a structured bridge from raw behavioral data to narrative video. This allows subsequent weight allocation, shot selection, and video editing to be based on high-quality narrative units, fundamentally improving the narrative quality and viewing value of the generated video at the content selection level.

[0078] Based on any embodiment of the method in this application, each valid behavioral event is assigned a corresponding narrative logic weight according to the logical relationship, and a corresponding plot duration is set for each narrative logic weight, including: Step S3310: For each valid behavioral event, calculate its narrative logic weight based on the basic weight corresponding to its behavioral type in the narrative structure template and the weight adjustment rules corresponding to the logical relationship type with the preceding and following behavioral events. The narrative structure template predetermines a base weight for each behavior type. This base weight reflects the inherent narrative value or visual appeal of the behavior type, detached from its specific context. Base weights are typically stored numerically, forming a mapping table between behavior types and base weights. For example, in a specific weight configuration, rare and dynamic behavior types like a leaping tackle can be assigned a higher base weight, such as 90 points; common, everyday behavior types like quietly grooming themselves are assigned a lower base weight, such as 50 points; and behavior types with potential comedic effect, such as awkwardly slipping, can have a base weight set at 85 points. These base weights serve as the benchmark for subsequent adjustments.

[0079] Building upon the basic weights, the weights are further dynamically adjusted based on the role and position of the valid behavioral event within the narrative chain constructed from logical relationships, applying weight adjustment rules. These rules are closely tied to the type of logical relationship; essentially, they define the weight increments or decrements that events at different sequence positions should receive under a specific logical relationship. In one embodiment, the weight adjustment rule can be represented as a list of adjustment values ​​or an adjustment function.

[0080] For example, for a causal chain relationship, a weight adjustment rule can be defined: in this chain, the weight of the starting event remains unchanged, the weight of the intermediate transitional events increases by 10 points, and the weight of the ending event, which serves as the climax, increases by 20 points. Specifically, in the causal chain of voyeurism-stealth-attack, if the attack is determined to be the end point of the chain, its narrative logic weight will receive an additional 20-point causal chain climax bonus on top of its base weight of 90 points, ultimately reaching 110 points.

[0081] For dramatic conflict relationships, the following weighting adjustment rules can be defined: events that trigger conflict turning points receive a higher weighting bonus. For example, in a scene where someone is discovered by their owner after successfully stealing food, this act of being discovered serves as the trigger for the conflict, and its weighting adjustment rule can be set to increase by 25 points. Assuming its base weight is 70 points, the adjusted narrative logic weight would be 95 points.

[0082] For humorous relationships based on contrast, the weighting adjustment rule can be defined as follows: the subsequent event that creates the contrast effect receives a significant bonus. For example, if an embarrassing slip occurs after an exhilarating chase, this slip serves as a manifestation of the contrast effect, and its weighting adjustment rule can be set to add 15 points. Combined with its base weight of 85 points, its narrative logic weight can reach 100 points.

[0083] In a more complex embodiment, the weight adjustment rule may not be a fixed value, but a function calculated based on the specific attributes of the behavioral event in the sequence (such as the time interval with the preceding behavioral event), thereby achieving more refined adjustments. By combining the basic weight and the applicable weight adjustment rule, the narrative logic weight calculated for each valid behavioral event becomes a comprehensive quantitative indicator that can simultaneously reflect its own value and the value of the narrative context, providing a precise basis for subsequent resource allocation.

[0084] Step S3320: Based on the narrative logic weight calculated for each valid behavioral event and combined with the preset mapping relationship, set the plot duration for each valid behavioral event in the final narrative video. The mapping relationship is configured such that the higher the narrative logic weight, the longer the allocated plot duration.

[0085] Based on the narrative logic weight calculated for each valid behavioral event, and combined with the preset mapping relationship, the plot duration of each valid behavioral event in the final narrative video is set. This mapping relationship is configured such that the higher the narrative logic weight, the longer the plot duration is allocated. The purpose is to transform the quantitative assessment of narrative importance into specific video time resource allocation, thereby controlling the rhythm and expressiveness of the final video.

[0086] The predefined mapping relationship defines the conversion rules from the numerical domain of narrative logic weights to the duration domain of plot events. In one embodiment, this mapping relationship can be a simple linear function. For example, a base duration, such as 2 seconds, can be set, and it can be stipulated that for every 10 points increase in narrative logic weight, the plot duration increases by 0.5 seconds. Under this rule, the plot duration of a valid behavioral event with a narrative logic weight of 70 points can be calculated as 2 seconds + (70 / 10) * 0.5 seconds = 2 + 3.5 = 5.5 seconds. An event with a weight as high as 110 points can have a plot duration of 2 seconds + (110 / 10) * 0.5 seconds = 2 + 5.5 = 7.5 seconds.

[0087] In another embodiment, the mapping relationship can take the form of a piecewise function to more finely control the duration growth curve of different weight intervals. For example, the following rules can be set: when the narrative logic weight is below 60 points, the plot duration is fixed at a transition duration of 1 second; when the weight is between 60 and 100 points, the plot duration increases linearly from 1 second to 5 seconds; when the weight is greater than 100 points, the plot duration enters a high growth interval, for example, starting from 5 seconds, for every 10 points increase in weight, the duration increases by 1 second, so that the climax can be fully displayed.

[0088] The mapping relationship can also be expressed as a non-linear function, such as an exponential or logarithmic relationship, so that a small change in weight can cause a significant change in duration within a specific interval, or that the growth of duration tends to level off during periods with extremely high weights. Furthermore, the mapping relationship can incorporate a total duration constraint mechanism. In one embodiment, the preliminary total duration of all valid behavioral events can be calculated based on their weights and the mapping relationship. If this total exceeds a preset upper limit for the total video duration, the allocated duration of each event is compressed proportionally to ensure that the final video length is appropriate.

[0089] By applying mapping relationships, each valid behavioral event is assigned a precise episode duration. For example, a dramatic conflict event with a calculated narrative logic weight of 95 points (such as being discovered by the owner) might be allocated 5 seconds according to linear mapping rules; while a causal chain climax event with a weight as high as 110 points (such as a leaping attack) might be allocated 7.5 seconds and could be rendered with slow-motion effects. This ensures that the video's narrative resources are reasonably allocated to the most important plot points, thus creating a well-paced rhythm in the final narrative video.

[0090] The above embodiments achieve significant technical advantages in narrative resource allocation by introducing a dynamic calculation mechanism for narrative logic weights and an intelligent mapping mechanism for plot duration. They transform abstract storytelling principles into quantifiable resource allocation algorithms. By combining basic weights with dynamic adjustment rules based on logical relationships, the narrative logic weights assigned to each behavioral event accurately reflect its relative importance within the overall narrative structure. Furthermore, by utilizing preset mapping relationships, weight values ​​are linearly or non-linearly converted into specific plot durations, establishing a standardized conversion channel from narrative value to time resources. This quantitative management mechanism ensures that high-weight events receive sufficient display time, while secondary events are reasonably compressed, fundamentally solving the problem of a flat, monotonous pace in traditional video generation. Ultimately, it can automatically generate narrative videos with a good balance of tension and release, highlighting key points. While maintaining automated production efficiency, it achieves a near-professional level of pacing control, significantly enhancing the artistic expression and viewing value of the generated content.

[0091] Based on any embodiment of the method in this application, and based on the camera control strategy in the narrative structure template, for each valid behavioral event, a segment source with a matching perspective is selected from multiple video streams within its event time period, and a plot segment with a continuous corresponding plot duration is extracted from it, including: Step S3410: Based on the predefined preferred camera angles that match different behavior types in the camera control strategy, determine one or more target camera angles corresponding to the behavior type for each valid behavior event; A shot control strategy can include a shot rule base that defines the mapping between behavior types and preferred shot angles. This mapping can be represented as a data lookup table, where each behavior type is associated with one or more recommended angle identifiers. Angle identifiers can be described using standardized terminology, such as frontal close-up, side medium shot, low-angle shot, high-angle shot, and tracking shot. Each angle identifier corresponds to a specific composition and emotional expression intention.

[0092] To facilitate understanding, we can specifically explain the preferred camera angles and their narrative intentions mapped to several typical behavior types. In one embodiment, for highly dynamic behavior types such as playing and running, the rules can be mapped to a side-following camera angle and a low-angle shooting angle. The side-following camera angle can clearly capture the animal's movement trajectory and limb extension posture, emphasizing dynamism; the low-angle shooting angle can mimic the animal's subjective perspective, enhancing the impact and immersion of the image. Therefore, when the behavior type of the effective behavioral event is playing and running, the target camera angles determined for it can include side-following cameras and low-angle shooting angles.

[0093] In another embodiment, for static behavior types such as sleeping or eating, the rules can be mapped to a frontal close-up view and a top-down view. A frontal close-up view focuses on the animal's facial expressions and details, conveying a sense of tranquility or focus; a top-down view shows the animal's relationship with its environment, creating an observer's perspective. Therefore, for effective behavioral events corresponding to the behavior type of sleeping, the target perspectives can be determined as a frontal close-up and a top-down view.

[0094] For suspenseful behaviors such as peeping or stealth, rules can be mapped to occluded composition or tilted camera angles. Occluded composition can mimic the visual effect of peeping through a crack in a door or behind furniture, increasing mystery and anticipation; tilted camera angles can disrupt the balance of the image, creating an unsettling atmosphere. Therefore, the target perspectives for such behaviors can include occluded composition and tilted camera angles.

[0095] In a more complex embodiment, the mapping relationship may not be a simple one-to-one or one-to-many relationship, but may be conditional. For example, the rule could be defined as follows: when the playful running behavior occurs in a spacious area, a side-view shooting angle is preferred; while when the behavior occurs in a narrow corridor, a frontal close-up angle is recommended to avoid spatial limitations. In this case, when determining the target perspective, contextual information such as the location of the behavioral event in space can also be used for comprehensive judgment.

[0096] Step S3420: During the event period corresponding to the effective behavior event, the video streams from multiple cameras are evaluated for image quality, and the video stream from the camera that can provide the target viewpoint and has the best image quality is selected as the segment source of the effective behavior event. Within the event period corresponding to the effective behavioral event, the video streams from multiple cameras are evaluated for image quality. The camera video stream that provides the target viewpoint and has the best image quality is selected as the segment source for the effective behavioral event. This achieves intelligent selection of the best segment that simultaneously meets the viewpoint and image quality requirements from multiple selectable video sources.

[0097] First, candidate video streams that can provide the target viewpoint need to be selected. In one embodiment, this can be achieved by querying the deployment metadata of the cameras. When each camera registers in the system, its physical location, orientation, focal length, and other parameters can be recorded. By calculating the geometric relationship between the location of the target behavior event in space and the camera, it can be determined whether the camera's field of view can cover the event location and shoot in an direction that matches the target viewpoint. For example, for events requiring a low-angle upward shooting perspective, only cameras installed at a low position with their lenses tilted upward can provide this viewpoint. In another embodiment, computer vision technology can be used to analyze the video stream in real time, identifying scene landmarks in the image or estimating the camera's pose to determine whether the current image matches the target viewpoint.

[0098] After identifying the set of candidate video streams that can provide the target perspective, they need to be evaluated for image quality. Image quality evaluation aims to quantify the overall visual quality of the footage captured by each video stream within the event period. The evaluation can be based on a pre-defined system that includes multiple scoring dimensions. These scoring dimensions aim to measure the usability and aesthetics of the image from different perspectives. In one embodiment, the evaluation system may include a composition score to assess whether the position and proportion of the target subject in the image conform to aesthetic rules, such as the rule of thirds; a sharpness score to determine whether the image is accurately focused, detailed, and free of motion blur; a perspective score to measure the degree to which the current image matches the target perspective, such as whether it is the required frontal close-up; and a lighting score to check whether the image is properly exposed, the colors are normal, and there is no overexposure or underexposure. Each dimension can independently calculate a sub-score.

[0099] To obtain an overall evaluation of each candidate video stream, the sub-scores from the aforementioned multiple scoring dimensions need to be fused to obtain a comprehensive picture quality score. In one embodiment, fusion can be achieved through a weighted average, where each scoring dimension is assigned a weight, and then a weighted sum is calculated. The weight allocation can reflect the differences in importance between different dimensions; for example, sharpness can be assigned the highest weight, followed by composition. In another embodiment, fusion can also employ a more complex multi-objective decision algorithm, or introduce a pre-trained scoring model that takes each score as input and directly outputs a comprehensive score.

[0100] Finally, the overall picture quality scores of all candidate video streams are compared, and the video stream with the highest score is selected as the source of the valid behavioral event. This ensures that the source material used to generate the narrative video not only conforms to the narrative intent in terms of perspective, but also achieves optimal visual quality.

[0101] Step S3430: Extract a video segment from the selected segment source that corresponds to the plot duration set for the valid behavioral event, and use it as the final plot segment.

[0102] Based on the event time period recorded during the identification phase of a valid behavioral event, particularly the start timestamp, the system locates the event within the selected optimal camera video stream, i.e., the segment source. In one embodiment, the starting point of the interception is the start timestamp of the behavioral event. The intercepted length strictly adheres to the episode duration set for the valid behavioral event. For example, if the start timestamp of a valid behavioral event is T1, and its assigned episode duration is 5 seconds, then the system precisely intercepts a 5-second video data block from the segment source starting from time point T1.

[0103] In another embodiment, the starting point for the cut can be fine-tuned based on the specific characteristics of the behavioral event. For example, for some valid behavioral events, the actual visual climax may occur slightly later than the starting point of behavior recognition. Typical starting offsets for various behavior types can be pre-stored. During the cut, the starting timestamp T1 is added with a positive offset (e.g., +0.5 seconds) as the actual cut starting point to ensure that the most expressive moment is captured. The length of the cut remains strictly maintained at the set episode duration.

[0104] The resulting video clips are the final plot segments. Each plot segment is a self-contained, high-quality video unit with precise temporal boundaries. It carries specific behavioral content, is generated from a video source with optimal viewpoint and image quality, and its duration precisely reflects the importance and weight of that behavior within the overall narrative. All plot segments corresponding to valid behavioral events constitute the direct raw material for subsequent encapsulation control strategies to splice them sequentially.

[0105] The above embodiments upgrade the traditional passive recording from a fixed perspective to active creation through multi-dimensional dynamic optimization. First, a predefined lens rule library transforms the narrative intent of different behavioral types, such as dynamism, tranquility, and suspense, into specific preferred lens perspectives, such as side tracking shots, close-ups, and tilted shots, achieving semantic matching of shooting techniques. Next, a comprehensive evaluation of perspective matching and image quality is simultaneously performed on multiple candidate video streams during the event's occurrence, ensuring that the selected clips achieve optimal artistic expression and technical quality. Finally, standardized, high-quality plot segments are generated by precisely extracting durations strictly corresponding to the importance of the plot. This end-to-end intelligent processing fundamentally guarantees that every frame of the final narrative video possesses both narrative expressiveness and visual aesthetics, achieving a qualitative leap from having available materials to having high-quality selections, significantly improving the professionalism and viewing value of the generated video.

[0106] Based on any embodiment of the method in this application, the image quality of video streams from multiple cameras is evaluated, and the camera video stream that can provide the target viewpoint and has the best image quality is selected, including: Step S3421: For multiple candidate camera video streams that can provide the target viewpoint during the event period, determine multiple scores for image frames belonging to the event period, wherein the multiple scores include at least two of the following scoring dimensions: composition score, sharpness score, viewpoint score, and illumination score. In the image quality assessment process, the first step is to identify multiple scores for image frames belonging to the target viewpoint from multiple candidate camera video streams that can provide the target viewpoint within the effective event period. These scoring dimensions include at least two of the following: composition score, sharpness score, viewpoint score, and lighting score. The composition score assesses whether the position and proportion of the target subject in the image conforms to aesthetic rules. For example, it quantifies the score by detecting whether the subject, i.e., the target animal, is located at the intersection of the rule of thirds lines. The score range can be set from 0 to 100. The sharpness score evaluates the image detail by calculating the image gradient or spectral energy. For example, it uses the Laplacian variance method; a higher variance value indicates better sharpness, and it can be normalized to 0-100. The viewpoint score measures the degree of fit between the actual image and the target viewpoint. For example, for a required low-angle upward shot, the score is calculated by the difference between the camera's elevation angle and the ideal angle; a smaller difference results in a higher score. The lighting score assesses the exposure level by analyzing the image histogram to avoid overexposure or underexposure. For example, it scores based on the proportion of pixel values ​​distributed within the normal exposure range.

[0107] Step S3422: Fuse multiple scores for each candidate camera video stream within the event period to obtain a comprehensive image quality score; When fusing multiple scores for each candidate camera video stream within an event period, a weighted average method can be used to obtain the overall image quality score. In one embodiment, different weights are assigned to each scoring dimension, such as a sharpness weight of 0.4, a composition weight of 0.3, a viewing angle weight of 0.2, and a lighting weight of 0.1. The calculation formula is: Overall Score = Sharpness Score × 0.4 + Composition Score × 0.3 + Viewing Angle Score × 0.2 + Lighting Score × 0.1. In another embodiment, a fuzzy logic-based fusion algorithm can be used. First, the scores for each dimension are converted into fuzzy sets, and then the overall score is calculated using preset inference rules. This method can better handle the uncertainty of the scoring criteria. Yet another embodiment uses a machine learning model for fusion. A regression model is trained using a large number of labeled high-quality image samples, and the scores for each dimension are used as feature inputs to directly predict the overall quality score.

[0108] Step S3423: Compare the overall image quality scores of all candidate camera video streams, and select the video stream with the highest overall image quality score as the segment source with the best image quality.

[0109] When comparing the overall image quality scores of all candidate camera video streams, a sorting algorithm can be used to select the video stream with the highest score. In one implementation, the overall scores of each video stream are directly compared, and the video stream with the highest score is selected. In another implementation, when multiple video streams have similar scores (e.g., a difference of less than 5 points), the video stream with the higher viewpoint score can be prioritized to ensure viewpoint matching priority. For special cases where overall scores are the same, other factors can be considered in the decision-making process, such as selecting the video stream with lower network latency to ensure transmission stability, or selecting the camera video stream with a lower historical failure rate to improve system reliability. Through this multi-dimensional evaluation and optimization selection, it is ensured that the final selected video source meets the optimal standards both technically and artistically.

[0110] The above embodiments, by constructing a multi-dimensional and quantifiable image quality assessment and fusion decision-making mechanism, achieve a technological leap from subjective experience-based judgment to objective data-driven decision-making at the video source selection level. Its advantage lies in establishing a scientific standard for optimal segment selection. First, candidate video streams are comprehensively and quantitatively evaluated through multiple specialized scoring dimensions such as composition, sharpness, perspective, and lighting, ensuring accurate measurement of key indicators from aesthetic rules, detail performance, perspective fit to exposure quality. Then, data fusion methods such as weighted averaging, fuzzy logic, or machine learning models are used to synthesize the multi-dimensional scores into a unified comprehensive image quality score, effectively solving the trade-off problem of various indicators in multi-objective decision-making. Finally, by introducing a hierarchical decision-making mechanism, it is ensured that the final selected segment source is not only optimal in visual quality but also achieves a balance in system robustness and real-time performance. This structured evaluation system transforms the traditional video editing process, which relies on manual experience for selection, into an automated and reusable standardized process. It not only guarantees the visual quality of every frame in narrative videos but also significantly improves the efficiency and consistency of the material selection process, laying a solid foundation for generating high-quality narrative videos.

[0111] Based on any embodiment of the method in this application, according to the encapsulation control strategy in the narrative structure template, the various plot video segments are sequentially spliced ​​according to the narrative order determined by the logical relationship to generate a narrative video, including: Step S3510: Arrange the plot segments corresponding to each effective behavioral event according to the narrative order determined by the logical relationship to form an initial editing sequence; Guided by the encapsulation control strategy, the process of synthesizing individual plot video clips into the final narrative video begins by arranging these clips according to a narrative order determined by logical relationships. This narrative order is not a simple timeline arrangement, but a storyline constructed based on the inherent connections between behavioral events. For example, for the selected valid sequence of behavioral events—peeping, stealth, failed attack, and awkward grooming—logical relationships determine that it constitutes a complete narrative chain of "comedic blunders." Therefore, when arranging the initial clip sequence, the order is: peeping clip, stealth clip, failed attack clip, and awkward grooming clip, thus forming a complete story structure with cause, development, climax, and ending, forming the corresponding initial clip sequence, rather than mechanically splicing them according to their actual absolute timestamps.

[0112] Step S3520: According to the predefined transition rules in the encapsulation control strategy that match different behavior types or logical relationship types, add corresponding transition effects between adjacent plot segments; After the initial editing sequence is formed, corresponding transition effects are added between adjacent plot segments according to predefined transition rules in the encapsulation control strategy. Transition rules can be bound to behavior types or logical relationship types to enhance narrative fluency and emotional impact. In one embodiment, the transition rule library contains transition rules corresponding to various mapping relationships. For example, for adjacent segments with a causal chain relationship (such as a stealth segment and a failed pounce segment), the rule can specify the use of soft fade-in / fade-out effects to achieve a natural scene transition. For segments with dramatic conflict or contrasting humor (such as a failed pounce segment and an awkward grooming segment that follows), the rule can specify the use of fast jump cuts or wipe effects to highlight the dramatic effect of the transition or the comedic contrast. In another embodiment, transition rules can consider the behavior types of the preceding and following segments. For example, when switching from a dynamic play segment to a static sleeping segment, a fade-to-black transition effect can be used to suggest the passage of time or the calming of emotions.

[0113] Step S3530: Based on the overall narrative tone presented by the narrative order, match the corresponding background music and sound effects from the audio library to construct audio information, and synchronize it with the initial editing sequence; Synchronously or asynchronously with step S3520, audio information is constructed by matching corresponding background music and sound effects from the audio library based on the overall narrative tone presented by the entire narrative sequence. To this end, the emotional tone of the narrative sequence is first analyzed. This emotional tone can be pre-associated with corresponding logical relationship rules for retrieval. For example, the overall tone of the aforementioned "comedic mistake" sequence is lighthearted and humorous. Based on this tone, matching metadata tags are retrieved from the audio library, and a background music piece with a brisk rhythm and humorous melody is selected. Simultaneously, sound effects are matched for key behavioral points, such as adding a comical slip sound effect at the moment of a failed tackle, or adding a humorous sound effect expressing helplessness during an awkward grooming moment. In one embodiment, audio synchronization also includes aligning video editing points with the music beat, i.e., performing "beat-matching" processing. For example, the occurrence of transition effects or key actions is precisely mapped to the downbeat of the music, thereby greatly enhancing the rhythm and visual appeal of the sound-image combination.

[0114] Step S3540: Render and encode the initial clip sequence with overlaid transition effects and audio information to output the final generated narrative video.

[0115] Finally, the initial edit sequence, overlaid with transition effects and audio information, is rendered and encoded, merging all elements into a single, unified video file. The rendering engine composites video clips, transition effects, background music, and sound tracks frame by frame. During the encoding stage, a standard video encoding format, such as H.264 or H.265, is selected, and appropriate bitrate, resolution, and frame rate parameters are set to balance file size and video quality. The output is the final narrative video file, such as a 30-second MP4 file with a complete storyline, professional transitions, and synchronized background music, which users can directly play and share.

[0116] The above embodiments, through the refined implementation of encapsulation control strategies, achieve a deep integration of narrative, artistry, and technology in the final stage of video compositing, elevating the high-quality materials generated in the preceding steps into a professionally caliber narrative video. Its advantage lies in constructing a fully automated intelligent post-production pipeline. This complete encapsulation process transforms traditionally experience-based tasks relying on professional editors, such as pacing, transition selection, and audio-visual coordination, into rule-based automated processing. This not only ensures a high degree of consistency in narrative structure, visual smoothness, and audio coordination in the output video but also maximizes production efficiency while achieving a professional artistic effect close to that of human editing, ultimately achieving a dual breakthrough in quality and efficiency for automated video generation.

[0117] Please see Figure 2This application provides an animal narrative video generation device, which is a functional embodiment of the animal narrative video generation method of this application. The device includes a video recognition module 3100, a logic control module 3200, a duration control module 3300, a lens control module 3400, and an encapsulation control module 3500. The video recognition module 3100 is configured to identify different behavioral events of a target animal from video streams from multiple cameras associated with the same spatial region identifier. The logic control module 3200 is configured to filter out multiple events that satisfy corresponding logical relationships from the different behavioral events based on a logic control strategy in a narrative structure template. Effective behavioral events; the duration control module 3300 is configured to assign a corresponding narrative logic weight to each effective behavioral event according to the logical relationship, and set a corresponding plot duration for the narrative logic weight; the shot control module 3400 is configured to select a perspective-matching segment source from multiple video streams within its event time period for each effective behavioral event based on the shot control strategy in the narrative structure template, and extract a plot segment with a continuous plot duration; the encapsulation control module 3500 is configured to splice the various plot video segments in a temporal sequence according to the narrative order determined by the logical relationship according to the encapsulation control strategy in the narrative structure template, to generate a narrative video.

[0118] Based on any embodiment of the device in this application, the video recognition module 3100 includes: a detection and recognition module, configured to detect the video streams of the plurality of cameras, identify behavioral events of the target animal, and determine the behavior type, start timestamp, and spatial location of each behavioral event; a perspective tracking module, configured to, after identifying a behavioral event, generate a perspective adjustment instruction based on the preferred lens perspective specified in the lens control strategy in a preset narrative structure template according to its behavior type, and drive the corresponding camera to enter a continuous tracking shooting state following the location of the behavioral event according to the instruction; and an end marker module, configured to exit the tracking shooting state when it is determined that the behavioral event has ended, and record the end timestamp, with the start timestamp and end timestamp defining the corresponding event time period.

[0119] Based on any embodiment of the device in this application, the logic control module 3200 includes: a sequence construction module, configured to arrange the identified behavioral events in chronological order to form a behavioral event sequence; a rule matching module, configured to match the behavioral event sequence with multiple logical relationship rules included in the predefined logic control strategy in the narrative structure template to determine the applicable target logical relationship rule, wherein the logical relationship rule is used to define the allowed logical relationship types between behavioral events of different behavioral types; and an event filtering module, configured to mark behavioral events in the behavioral event sequence whose behavioral types conform to the target logical relationship rule as valid behavioral events, and filter out isolated behavioral events whose behavioral types do not conform to any logical relationship rule.

[0120] Based on any embodiment of the device in this application, the duration control module 3300 includes: a weight determination module, configured to calculate the narrative logic weight of each valid behavioral event based on the basic weight corresponding to its behavioral type in the narrative structure template and the weight adjustment rules corresponding to the logical relationship type with the preceding and following behavioral events; and a duration setting module, configured to set the plot duration of each valid behavioral event in the final narrative video based on the narrative logic weight calculated for each valid behavioral event and in combination with a preset mapping relationship, wherein the mapping relationship is configured such that the higher the narrative logic weight, the longer the allocated plot duration.

[0121] Based on any embodiment of the device in this application, the lens control module 3400 includes: a viewpoint selection module, configured to determine one or more target viewpoints corresponding to the behavior type of each effective behavior event according to the preferred lens viewpoints predefined in the lens control strategy and matched with different behavior types; an evaluation and selection module, configured to evaluate the image quality of video streams from multiple cameras within the event time period corresponding to the effective behavior event, and select the camera video stream that can provide the target viewpoint and has the best image quality as the segment source of the effective behavior event; and a plot extraction module, configured to extract a video segment from the selected segment source that corresponds to the plot duration set for the effective behavior event, as the final plot segment.

[0122] Based on any embodiment of the apparatus in this application, the evaluation and selection module includes: an image scoring module, configured to determine multiple scores for image frames belonging to the event period among multiple candidate camera video streams that can provide the target viewpoint during the event period, wherein the multiple scores include at least two scoring dimensions among composition score, sharpness score, viewpoint score, and illumination score; a scoring determination module, configured to fuse the multiple scores of each candidate camera video stream during the event period to obtain a comprehensive image quality score; and a source selection module, configured to compare the comprehensive image quality scores of all candidate camera video streams and select the video stream with the highest comprehensive image quality score as the source of the segment with the best image quality.

[0123] Based on any embodiment of the device in this application, the encapsulation control module 3500 includes: a narrative arrangement module, configured to arrange plot segments corresponding to each effective behavioral event according to the narrative order determined by the logical relationship, forming an initial editing sequence; a transition processing module, configured to add corresponding transition effects between adjacent plot segments according to the predefined transition rules in the encapsulation control strategy that match different behavioral types or logical relationship types; an audio processing module, configured to match corresponding background music and sound effects from an audio library to construct audio information according to the overall narrative tone presented by the narrative order, and synchronize it with the initial editing sequence; and an encoding output module, configured to render and encode the initial editing sequence superimposed with transition effects and audio information, and output the finally generated narrative video.

[0124] To address the aforementioned technical problems, embodiments of this application also provide a computer device implementation. For example... Figure 3 The diagram shows the internal structure of a computer device. This computer device includes a processor, a computer-readable storage medium, a memory, a network interface, and various communication components connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When executed by the processor, the computer-readable instructions enable the processor to implement an animal narrative video generation method. The processor provides computational and control capabilities, supporting the operation of the entire computer device. The memory stores computer-readable instructions, which, when executed by the processor, enable the processor to execute the animal narrative video generation method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 3The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0125] In this embodiment, the processor is used to execute... Figure 2 The system contains the specific functions of each module and its sub-modules, and the memory stores the program code and various data required to execute the aforementioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the animal narrative video generation device of this application, and the server can call the server's program code and data to execute the functions of all sub-modules.

[0126] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the animal narrative video generation method of any embodiment of this application.

[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0128] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in the prior art that are similar to those in the open-source operations, methods, and processes of this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.

[0129] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating animal narrative videos, characterized in that, Includes the following steps: Identify different behavioral events of a target animal from video streams from multiple cameras associated with the same spatial region identifier; Based on the logical control strategy in the narrative structure template, multiple valid behavioral events that satisfy the corresponding logical relationships are selected from the different behavioral events; Each valid behavioral event is assigned a corresponding narrative logic weight based on the logical relationship, and a corresponding plot duration is set for each narrative logic weight. Based on the camera control strategy in the narrative structure template, for each valid behavioral event, a segment source with a matching perspective is selected from multiple video streams within its event time period, and a plot segment with a continuous corresponding plot duration is extracted from it. According to the encapsulation control strategy in the narrative structure template, the various plot video segments are spliced ​​together in chronological order according to the narrative sequence determined by the logical relationship to generate a narrative video.

2. The method for generating animal narrative videos according to claim 1, characterized in that, Identify different behavioral events of the target animal from video streams from multiple cameras associated with the same spatial region identifier, including: The video streams from the multiple cameras are detected to identify behavioral events of the target animal, and the behavior type, start timestamp, and spatial location of each behavioral event are determined. After identifying a behavioral event, a perspective adjustment instruction is generated based on the preferred camera angle specified in the camera control strategy within the preset narrative structure template according to the behavior type. Based on this instruction, the corresponding camera is driven to enter a continuous tracking shooting state that follows the location where the behavioral event occurs. When the behavior event is determined to have ended, the tracking and shooting state is exited, and the termination timestamp is recorded. The corresponding event time period is defined by the start timestamp and the termination timestamp.

3. The method for generating animal narrative videos according to claim 1, characterized in that, Based on the logical control strategy in the narrative structure template, multiple valid behavioral events that satisfy the corresponding logical relationships are selected from the different behavioral events, including: The identified behavioral events are arranged in chronological order to form a behavioral event sequence; The sequence of behavioral events is matched with multiple logical relationship rules contained in the predefined logical control strategy in the narrative structure template to determine the applicable target logical relationship rule. The logical relationship rule is used to define the types of logical relationships that are allowed between behavioral events of different behavioral types. In the sequence of behavioral events, behavioral events whose behavior types conform to the target logical relationship rules are marked as valid behavioral events, and isolated behavioral events whose behavior types do not conform to any logical relationship rules are filtered out.

4. The method for generating animal narrative videos according to claim 1, characterized in that, Each valid behavioral event is assigned a corresponding narrative logic weight based on the aforementioned logical relationship, and a corresponding plot duration is set for each narrative logic weight, including: For each valid behavioral event, its narrative logical weight is calculated based on the basic weight corresponding to its behavioral type in the narrative structure template and the weight adjustment rules corresponding to the logical relationship type with the preceding and following behavioral events. Based on the narrative logic weight calculated for each valid behavioral event, and combined with a preset mapping relationship, the plot duration of each valid behavioral event in the final narrative video is set. The mapping relationship is configured such that the higher the narrative logic weight, the longer the plot duration is allocated.

5. The method for generating animal narrative videos according to claim 1, characterized in that, Based on the camera control strategy in the narrative structure template, for each valid behavioral event, a segment source with a matching perspective is selected from multiple video streams within its event time period, and a plot segment with a continuous corresponding plot duration is extracted from it, including: Based on the predefined preferred camera angles that match different behavior types in the camera control strategy, one or more target camera angles corresponding to the behavior type are determined for each valid behavior event; Within the event time period corresponding to the effective behavior event, the video streams from multiple cameras are evaluated for image quality, and the video stream from the camera that can provide the target viewpoint and has the best image quality is selected as the segment source of the effective behavior event. From the selected source of segments, extract video segments that correspond to the duration of the plot set for the valid behavioral event, and use them as the final plot segments.

6. The method for generating animal narrative videos according to claim 5, characterized in that, The video streams from multiple cameras are evaluated for image quality, and the camera video stream that provides the target viewpoint and has the best image quality is selected, including: For multiple candidate camera video streams that can provide the target viewpoint during the event period, determine multiple scores for image frames belonging to the event period, the multiple scores including at least two of the following scoring dimensions: composition score, sharpness score, viewpoint score, and illumination score; Multiple scores for each candidate camera video stream within the event period are fused to obtain a comprehensive image quality score. Compare the overall image quality scores of all candidate camera video streams, and select the video stream with the highest overall image quality score as the source of the segment with the best image quality.

7. The method for generating animal narrative videos according to any one of claims 1 to 6, characterized in that, Based on the encapsulation control strategy in the narrative structure template, the various plot video segments are sequentially spliced ​​according to the narrative order determined by the logical relationship to generate a narrative video, including: Arrange the plot segments corresponding to each effective behavioral event according to the narrative order determined by the aforementioned logical relationship to form an initial editing sequence; According to the predefined transition rules in the encapsulation control strategy that match different behavior types or logical relationship types, corresponding transition effects are added between adjacent plot segments. Based on the overall narrative tone presented by the narrative order, corresponding background music and sound effects are matched from the audio library to construct audio information, and synchronized with the initial editing sequence. The initial edit sequence, which incorporates transition effects and audio information, is rendered and encoded to output the final narrative video.

8. An animal narrative video generation device, characterized in that, include: The video recognition module is configured to identify different behavioral events of a target animal from video streams from multiple cameras associated with the same spatial region identifier; The logic control module is configured to filter out multiple valid behavior events that satisfy the corresponding logical relationships from the different behavior events based on the logic control strategy in the narrative structure template. The duration control module is configured to assign a corresponding narrative logic weight to each valid behavioral event according to the logical relationship, and set a corresponding plot duration for the narrative logic weight. The camera control module is set to use a camera control strategy based on the narrative structure template to select a perspective-matching segment source from multiple video streams within the event time period for each valid behavioral event, and extract a plot segment that lasts for the corresponding plot duration. The encapsulation control module is configured to sequentially splice together various plot video segments according to the narrative order determined by the logical relationship, based on the encapsulation control strategy in the narrative structure template, to generate a narrative video.

9. A computer device comprising a processor and a memory, characterized in that, The processor invokes and runs a computer program in the memory to perform the steps of the animal narrative video generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, performs the steps included in the corresponding method.