A pet neck camera Vlog generation method, device, equipment and medium
By collecting multimodal data through a pet neck camera and performing intelligent analysis, an automated pet vlog is generated, solving the problems of data fragmentation and lack of emotional understanding, and achieving efficient emotion recognition and automatic content generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN UASCENT TECH CO LTD
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing pet monitoring equipment cannot simultaneously collect and process multimodal data, resulting in data fragmentation, lack of emotional understanding, and reliance on manual content generation, making it impossible to achieve intelligent emotional understanding and automatic Vlog generation.
Multimodal data (visual, auditory, positional, and motion) is collected using a pet neck camera. Through preprocessing, multimodal feature extraction and fusion analysis, combined with a sentiment analysis model and a generative pre-trained transformer, pet vlogs are generated.
It achieves accurate identification of pet emotional states and automatic Vlog generation, enhancing the usability of pet data and user emotional engagement, saving more than 95% of manual editing time, and achieving an emotion recognition accuracy rate of 89.3%.
Smart Images

Figure CN121334466B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a method, apparatus, device, and medium for generating pet neck camera Vlogs. Background Technology
[0002] With the rapid development of the pet economy, users' demand for emotional companionship and intelligent management of their pets has surged. However, existing pet monitoring devices have many technical pain points, making it difficult to meet users' needs for efficient use of pet data and emotional recording, as follows:
[0003] (1) Data fragmentation problem: Existing pet monitoring equipment can only collect single-dimensional data (such as only video or only GPS), resulting in one-sided information and lack of correlation. It is impossible to fully understand the pet's behavior and status. Users need to manually sort out a large amount of fragmented information. The core reason is the lack of multimodal data fusion technology, which makes it impossible to establish cross-modal correlation analysis.
[0004] (2) Lack of emotional understanding: Traditional devices cannot understand the emotional state and health status of pets and can only passively record them, which makes it impossible for users to discover abnormal situations of pets in a timely manner and lacks emotional communication experience. Technically, this is due to the lack of pet emotion recognition algorithms based on multimodal information.
[0005] (3) Content generation relies on manual labor: Users need to manually select, edit and produce pet videos, which is time-consuming and laborious. This not only reduces the user experience, but also causes a large number of valuable pets to be wasted instantly. The fundamental problem is the lack of intelligent content understanding, theme extraction and automatic editing technology.
[0006] In summary, existing technologies cannot simultaneously acquire and process multimodal data (visual, auditory, positional, and motion) and achieve intelligent emotion understanding and automatic vlog generation. Therefore, there is an urgent need for a method, device, equipment, and medium for generating vlogs using a pet neck camera to address the shortcomings of existing technologies. Summary of the Invention
[0007] The purpose of this invention is to propose a method, device, equipment, and medium for generating pet neck camera vlogs, in order to solve the problems of data fragmentation, lack of emotional understanding, and reliance on manual content generation in existing pet monitoring equipment, and to achieve the ability to simultaneously collect multimodal data of vision, hearing, location, and motion, and to complete intelligent emotional understanding and automatic vlog generation.
[0008] Firstly, to achieve the above objectives, the present invention provides a method for generating pet neck-mounted camera Vlogs, comprising the following steps:
[0009] S1. Use a pet neck camera to acquire multimodal data of the pet;
[0010] S2. Preprocess the pet multimodal data to obtain preprocessed pet multimodal data;
[0011] S3. Perform multimodal feature extraction and fusion analysis on the preprocessed pet multimodal data to obtain the analysis results of pet multimodal fusion features;
[0012] S4. Based on the analysis results of the pet multimodal fusion features, generate the pet neck camera Vlog.
[0013] Optionally, S1, using a pet neck camera, acquire multimodal data about the pet, including:
[0014] Data was collected using a pet neck camera to obtain initial multimodal data of the pet;
[0015] Determine whether the timestamps corresponding to the initial pet multimodal data are consistent. If they are, then obtain the initial pet multimodal data as the pet multimodal data; otherwise, perform the first operation.
[0016] The first operation is as follows: the initial pet multimodal data is time-synchronized using a time synchronization mechanism to obtain pet multimodal data, which includes pet video data, pet audio data, pet GPS data, and pet IMU data.
[0017] Optionally, S2, preprocessing the pet multimodal data to obtain preprocessed pet multimodal data includes:
[0018] The pet video data and the pet audio data are compressed to obtain compressed pet video data and compressed pet audio data respectively;
[0019] The pet GPS data is processed to obtain the processed pet GPS data;
[0020] The Kalman filter algorithm is used to filter noise from the pet IMU data to obtain filtered pet IMU data.
[0021] The compressed pet video data, the compressed pet audio data, the normalized pet GPS data, and the filtered pet IMU data are integrated to obtain initial preprocessed pet multimodal data;
[0022] Anomaly detection is performed based on the initial preprocessed pet multimodal data to obtain the preprocessed pet multimodal data.
[0023] Optionally, S3, perform multimodal feature extraction and fusion analysis on the preprocessed pet multimodal data to obtain the analysis results of the pet multimodal fusion features, including:
[0024] Multimodal features are extracted using the preprocessed pet multimodal data to obtain pet multimodal features, which include pet visual features, pet audio features, pet location features, and pet movement features.
[0025] Cross-modal fusion is performed on the pet multimodal features to obtain pet multimodal fused features;
[0026] The analysis results of the pet multimodal fusion features are obtained by analyzing and processing the pet multimodal fusion features.
[0027] Optionally, cross-modal fusion is performed on the pet multimodal features to obtain pet multimodal fused features, including:
[0028] The pet multimodal features are used for feature mapping to obtain pet multimodal features in the same dimensional space;
[0029] The pet multimodal features in the same dimensional space are concatenated to obtain a pet multimodal feature matrix;
[0030] Based on the pet multimodal feature matrix, a multimodal attention mechanism is used to obtain the weight matrix of the pet multimodal features;
[0031] The pet multimodal features in the same dimensional space are weighted and fused according to the weight matrix of the pet multimodal features to obtain the pet multimodal fused features.
[0032] Optionally, the analysis results of the pet multimodal fusion features are obtained by performing analysis and processing based on the pet multimodal fusion features, including:
[0033] Based on the aforementioned pet multimodal fusion features, an emotion analysis model is used to obtain information on the intensity of pet emotions.
[0034] Based on the pet multimodal fusion features, event segmentation and clustering are performed to obtain pet event segmentation results;
[0035] Based on the pet event segmentation results, KMeans clustering and topic modeling are used to obtain pet event topics;
[0036] Highlight extraction is performed using the pet's emotional intensity information to obtain pet highlight segments;
[0037] The analysis results of obtaining the pet's emotional intensity information, the pet's event theme, and the pet's highlight segments as pet multimodal fusion features are obtained.
[0038] Optionally, S4, based on the analysis results of the pet's multimodal fusion features, generate pet neck camera Vlog generation results, including:
[0039] By combining the analysis results of the pet multimodal fusion features with the mapping relationship between emotion and music parameters, emotional matching background music is obtained;
[0040] Based on the analysis results of the pet multimodal fusion features, subtitle content is obtained through a generative pre-trained transformer model;
[0041] Based on the analysis results of the pet's multimodal fusion features, the background music matched with the emotion, and the subtitle content, the pet neck camera Vlog generation results are obtained.
[0042] Secondly, to achieve the above objectives, the present invention provides a pet neck-hanging camera Vlog generation device, comprising: a hardware acquisition module, an edge processing module, a cloud AI module, and an application service module;
[0043] The hardware acquisition module is used to acquire multimodal data of pets using a pet neck camera;
[0044] The edge processing module is used to preprocess the pet multimodal data to obtain preprocessed pet multimodal data.
[0045] The cloud-based AI module is used to extract and fuse the preprocessed pet multimodal data to obtain the analysis results of the pet multimodal fusion features.
[0046] The application service module is used to generate pet neck camera Vlog results based on the analysis results of the pet multimodal fusion features.
[0047] Thirdly, to achieve the above objectives, the present invention provides an electronic device, comprising: one or more processors; and a storage device having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation of the first aspect.
[0048] Fourthly, to achieve the above objectives, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any implementation of the first aspect.
[0049] Compared with the closest existing technology, the present invention has the following advantages:
[0050] This invention addresses the data fragmentation problem of existing devices by utilizing multimodal data such as video, audio, GPS, and IMU. It achieves deep fusion of four modalities—visual, auditory, location, and motion—and comprehensive acquisition of pet-related information, transforming fragmented data into structured narratives. This enhances the usability of pet-related data and prevents valuable moments from being wasted. Furthermore, this invention leverages multimodal AI-based pet emotion recognition technology to accurately identify pet emotional states, achieving an accuracy rate of 89.3%. This fills a gap in understanding pet emotions, enabling real-time monitoring of abnormal pet emotions and behaviors and providing early warnings, thus assisting pet owners. Pet health management; This invention relies on a fully automated generation process from material collection to finished Vlog, saving more than 95% of manual time. It eliminates the need for manual editing by users, automatically generating meaningful pet story Vlogs, enhancing user emotional engagement and strengthening the emotional connection between humans and pets. This invention adopts an architecture combining lightweight edge processing and deep cloud analysis, balancing real-time performance with intelligent requirements. Edge latency is <100ms, and cloud analysis accuracy is >85%. This not only enhances user emotional engagement and assists in pet health management but also fills the gap in AI analysis and content generation for smart pet devices, effectively enhancing market competitiveness. Attached Figure Description
[0051] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0052] Figure 1 This is a flowchart of a method for generating pet neck-mounted camera Vlogs according to an embodiment of the present invention;
[0053] Figure 2 This is a schematic diagram of the structure of a pet neck-hanging camera Vlog generation device according to an embodiment of the present invention;
[0054] Figure 3 This is a schematic diagram of the structure of the electronic device proposed in an embodiment of the present invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0056] The terminology used in the embodiments section of this invention is for the purpose of explaining specific embodiments of the invention only, and is not intended to limit the invention.
[0057] like Figure 1 As shown, this embodiment of the invention provides a method for generating pet neck camera Vlogs, including:
[0058] S1. Use a pet neck camera to acquire multimodal data of the pet;
[0059] Multimodal data acquisition is achieved through the hardware acquisition function of the pet neck camera. The acquired multimodal data includes video data, audio data, GPS data, and inertial measurement unit (IMU) data. At the same time, a unified timestamp system is used to perform time synchronization processing on all modal data to ensure that the data is strictly aligned in time with a synchronization accuracy of milliseconds, providing a time consistency foundation for subsequent data processing.
[0060] S2. Preprocess the pet multimodal data to obtain preprocessed pet multimodal data;
[0061] The collected multimodal data underwent preliminary processing and optimization. Specifically, this included compressing video data using the H.265 encoding standard and audio data using AAC encoding to reduce storage and transmission costs. Kalman filtering was used to filter sensor data such as IMUs to reduce noise interference. The TensorFlow Lite framework was also used to perform anomaly detection on the data, focusing on identifying obvious equipment malfunctions such as black screens and electrical noises. Finally, the preprocessed pet multimodal data was output to prepare for subsequent feature extraction.
[0062] S3. Perform multimodal feature extraction and fusion analysis on the preprocessed pet multimodal data to obtain the analysis results of pet multimodal fusion features;
[0063] Feature extraction and fusion operations were performed on preprocessed pet multimodal data. Multimodal feature extraction employed ResNet-50 for visual features, WavLM for audio features, geocoding for location features, and behavioral pattern recognition for motion features. Subsequently, a cross-modal fusion process projected each modality's features onto a 768-dimensional unified semantic space through a linear layer, concatenating them into a (4,768) matrix. An attention mechanism was then used to calculate a (4,4) weight matrix, and after weighted aggregation, a fused feature vector integrating visual, audio, location, and motion information was obtained. Finally, intelligent analysis was performed based on the fused features, outputting multimodal fusion feature analysis results including sentiment tags, event structure, and highlight segments, providing a decision-making basis for Vlog generation.
[0064] S4. Based on the analysis results of the pet multimodal fusion features, generate the pet neck camera Vlog generation result;
[0065] This step focuses on "structured narrative" and "user experience." First, music parameters are matched based on sentiment analysis results. Then, combined with information such as pet actions, environment, and emotions, engaging subtitles are generated using a GPT model. Highlights, matching music, and subtitles are integrated into a Vlog according to narrative logic (e.g., "beginning-development-ending"), completing the transformation from data to finished content. Finally, the Vlog can be pushed to users via the app, allowing them to view and share it in real time.
[0066] In summary, steps S1 to S4 achieve intelligent processing of multimodal data across the entire chain. This not only solves the pain points of traditional devices such as data fragmentation, lack of emotional understanding, and reliance on manual content generation, but also ensures processing efficiency and accuracy through an edge-cloud collaborative architecture. It can automatically generate emotionally resonant Vlogs with narrative logic, significantly saving users' manual editing time. At the same time, it enables intelligent monitoring of pet status, enhancing the emotional connection between humans and pets and the value of data utilization.
[0067] As one possible implementation, in the above embodiments, step S1 may specifically include the following steps:
[0068] S1-1. Use a pet neck camera to collect data and obtain initial pet multimodal data;
[0069] The initial multimodal data collection of pets is completed using the hardware acquisition layer of a pet neck-mounted camera. The hardware used is a detachable neck-mounted thumb-stabilized action camera, which integrates core components such as a 1080P high-definition wide-angle night vision camera, microphone, 4G+GPS module, and motion sensor, and can collect initial data related to pets in real time.
[0070] (1) Video data: A high-definition wide-angle night vision camera with a resolution of 1920×1080@30fps was used, which clearly recorded the pet's visual behavior and the scene it was in.
[0071] (2) Audio data: The microphone captures the pet's barking, movement sounds and ambient sounds in a 48kHz sampling rate and 16-bit bit depth, i.e., 48kHz / 16-bit stereo format.
[0072] (3) GPS data: The real-time location information of the pet is recorded by a 4G+GPS module with an update frequency of 1Hz and an accuracy of ±3 meters;
[0073] (4) IMU data: IMU data is collected using a 9-axis motion sensor at a sampling frequency of 100Hz, which can provide information such as the motion and attitude of the device.
[0074] S1-2. Determine whether the timestamps corresponding to the initial pet multimodal data are consistent. If they are, obtain the initial pet multimodal data as the pet multimodal data. Otherwise, execute S1-3.
[0075] The initial pet multimodal data is validated in terms of time dimension by comparing the timestamps of video, audio, GPS, and IMU data one by one. If the timestamps of the four types of data match perfectly, it indicates that the time dimension of the data has been naturally aligned during the acquisition process, and the initial pet multimodal data is directly identified as pet multimodal data that can be used for subsequent processing. If there are discrepancies or inconsistencies in the timestamps, the time synchronization mechanism is triggered.
[0076] S1-3. The initial pet multimodal data is synchronized using a time synchronization mechanism to obtain pet multimodal data;
[0077] When initial data timestamps are found to be inconsistent, a time synchronization mechanism is employed. This mechanism introduces a unified timestamp system to calibrate and align video, audio, GPS, and IMU data in the time dimension. This ensures that the obtained pet video, audio, GPS, and IMU data from different modalities are strictly matched in time, with synchronization accuracy reaching the millisecond level. This process effectively avoids information misalignment caused by time differences in the acquisition of different modalities, ensuring the temporal correlation and accuracy of data in subsequent multimodal feature extraction and cross-modal fusion processes. Ultimately, it yields pet multimodal data with good temporal consistency, providing a reliable data foundation for subsequent in-depth analysis.
[0078] In summary, steps S1-1 to S1-3 not only cover the core status and environmental information of the pet through multi-dimensional data collection, but also avoid information misalignment caused by time differences in the collection of different modal data through high-precision time synchronization. This effectively solves the pain point of data fragmentation in existing equipment. At the same time, it provides accurate and complete data support in the time dimension for subsequent intelligent analysis links such as multi-modal feature extraction, cross-modal fusion, emotion recognition, and event segmentation. This directly ensures the reliability of the subsequent pet Vlog generation in terms of content accuracy (such as emotion judgment and event segmentation) and narrative logic.
[0079] As one possible implementation, in the above embodiments, step S2 may specifically include the following steps:
[0080] S2-1. Compress the pet video data and the pet audio data to obtain the compressed pet video data and compressed pet audio data respectively;
[0081] The collected pet video and audio data were compressed using corresponding encoding standards. Pet video data was encoded using H.265, significantly reducing data size while maintaining 1080P high-definition quality, thus reducing subsequent transmission and storage pressure. Pet audio data was encoded using Advanced Audio Coding (AAC), compressing the audio data size while preserving key sound information related to pet behavior, such as barking and ambient sounds during play. The final results were compressed pet video and audio data that met the requirements for transmission and analysis.
[0082] S2-2. The pet GPS data is processed to obtain the processed pet GPS data;
[0083] For pet GPS data collected at a 1Hz update frequency with an accuracy of ±3 meters, a data normalization process was performed. This included standardizing the data format by converting the raw GPS signal into a standard latitude and longitude coordinate format; removing invalid data and filtering out abnormal coordinate values caused by signal interference; associating time stamps by accurately binding each set of GPS location data with a unified timestamp of multimodal data to ensure the temporal correlation between location information and video, audio, and IMU data; and supplementing data attribute tags to indicate auxiliary information such as signal strength at the time of collection. The final result was normalized pet GPS data with a unified format, complete information, and time alignment.
[0084] S2-3. Use the Kalman filter algorithm to filter noise from the pet IMU data to obtain the filtered pet IMU data;
[0085] Noise filtering was performed on pet IMU data acquired at a sampling rate of 100Hz. Since motion sensors are susceptible to interference from pet movement and electromagnetic interference during data acquisition, random noise and drift errors are introduced into the data. Therefore, a Kalman filter algorithm was employed. By establishing state and observation equations for the pet IMU data and combining prior estimates with real-time observations, the sensor output data was dynamically corrected to suppress noise interference and smooth data fluctuations. After filtering, the accuracy and stability of the pet IMU data were significantly improved, ultimately yielding filtered pet IMU data that accurately reflects the pet's movement state, providing a reliable data foundation for subsequent motion feature extraction.
[0086] S2-4. Integrate the compressed pet video data, the compressed pet audio data, the normalized pet GPS data, and the filtered pet IMU data to obtain the initial preprocessed pet multimodal data;
[0087] Using a unified timestamp as the core index, compressed pet video data, compressed pet audio data, normalized pet GPS data, and filtered pet IMU data are integrated. Following a timeline, video clips, audio clips, GPS coordinates, and IMU motion data from the same moment are linked one-to-one, constructing a structured data set linked by time. This eliminates fragmentation of data across different modalities, ensuring strict alignment of different data types along the time dimension. Ultimately, this yields initial pre-processed pet multimodal data, providing a complete data foundation for subsequent intelligent analysis.
[0088] S2-5. Perform anomaly detection based on the initial preprocessed pet multimodal data to obtain preprocessed pet multimodal data;
[0089] Using pre-processed pet multimodal data as the detection target, anomaly detection was conducted using the TensorFlow Lite lightweight AI framework. During the detection process, pre-defined rules for judging device malfunctions were implemented. For video data, anomalies such as black screens, blurry images, and stuttering were detected; for audio data, anomalies such as electrical noise, no sound, and sound distortion were detected; and for GPS and IMU data, anomalies such as missing data and data mutations were detected. Data segments containing anomalies were removed, retaining only complete and valid data segments for each modality. This resulted in high-quality pre-processed pet multimodal data suitable for subsequent analysis, preventing anomalies from affecting the accuracy of feature extraction, fusion analysis, and Vlog generation.
[0090] In summary, steps S2-1 to S2-5 not only achieve lightweight processing of multimodal data (reducing transmission and storage costs), standardization and regularization (eliminating data format and time deviations), and quality screening (ensuring data validity), but also provide a high-quality and highly reliable data foundation for subsequent feature extraction, cross-modal fusion analysis, and automatic Vlog generation, effectively improving the efficiency and accuracy of the overall Vlog generation process.
[0091] As one possible implementation, in the above embodiments, step S3 may specifically include the following steps:
[0092] S3-1. Use the preprocessed pet multimodal data to extract multimodal features and obtain pet multimodal features;
[0093] Using preprocessed pet multimodal data (including video, audio, GPS, and IMU data) as input, targeted feature extraction is performed via cloud-based AI, as detailed below:
[0094] (1) Pet visual features: The ResNet-50 model is used to extract visual features. ResNet-50 is a classic deep convolutional neural network that is widely used in computer vision tasks such as image classification and feature extraction. It can extract representative pet visual features from images (video stream related image frames, etc.).
[0095] (2) Pet audio features: Audio features are extracted using the WavLM model. WavLM is a model for audio processing that can extract features that characterize audio content from audio stream data for subsequent audio-related analysis.
[0096] (3) Pet location features: Location features are extracted through geocoding, which can convert location-related data such as GPS into geographically meaningful features to represent the location information of objects or devices.
[0097] (4) Pet movement characteristics; use behavior pattern recognition methods to extract movement characteristics, where the behavior pattern recognition method is based on IMU and other data to identify different behavior patterns, and then extract movement-related features to reflect the movement state of objects or equipment.
[0098] S3-2. Perform cross-modal fusion on the pet multimodal features to obtain pet multimodal fused features;
[0099] For the extracted pet multimodal features, each modal feature is first projected onto a unified semantic space through a linear layer to eliminate the dimensional differences between different modal features. Then, the projected modal features are concatenated to form a feature matrix, and an attention mechanism is used to calculate a cross-modal weight matrix, which can accurately reflect the importance of each modal feature in the fusion process. Finally, the modal features are weighted and aggregated according to the weight matrix to generate pet multimodal fusion features, thereby achieving spatiotemporal alignment and deep fusion of different modal features.
[0100] S3-3. Analyze and process the pet multimodal fusion features to obtain the analysis results of the pet multimodal fusion features;
[0101] Intelligent analysis is conducted based on the multimodal fusion features of pets. First, an emotion recognition model is used to calculate the scores of seven types of pet emotions, such as happiness and sadness, and their corresponding seven-dimensional quantization vectors. Second, event segmentation and clustering are performed. The L2 norm ratio of adjacent time points of each modality feature is calculated first, and the comprehensive change rate is calculated according to weights and combined with dynamic thresholds to complete event segmentation. Then, KMeans clustering and LDA models are used to mine potential event themes. Finally, highlight extraction is performed. Keyframe scores are calculated based on indicators such as dynamic region proportion and emotion intensity. At the same time, highlight segments are selected based on a three-act structure divided according to emotion changes and duration. Finally, the fusion feature analysis results containing pet emotional state, event theme, and highlight segments are obtained.
[0102] In summary, steps S3-1 to S3-3, based on preprocessed pet multimodal data, first extract pet visual, audio, location, and motion features using models such as ResNet-50 and WavLM, along with geocoding and behavioral pattern recognition technologies. Then, cross-modal fusion and alignment are achieved through unified dimensional projection and attention mechanisms, generating fused features that incorporate information from all modalities. Finally, pet emotion recognition, event segmentation and clustering, and highlight extraction analysis are performed based on these fused features. This process effectively overcomes the limitations of fragmented single-modal data, achieving deep integration and precise correlation of multi-dimensional information. It not only improves the accuracy of pet emotional state recognition and event theme mining but also provides high-quality, structured analytical data for subsequent automatic pet vlog generation, significantly reducing manual intervention costs. Furthermore, it provides comprehensive intelligent analytical support for pet behavior and health monitoring.
[0103] As one possible implementation, in the above embodiments, step S3-2 may specifically include the following steps:
[0104] S3-2-1. Use the pet multimodal features to perform feature mapping to obtain pet multimodal features in the same dimensional space;
[0105] Because the representations and semantics of features from different modalities (visual, audio, location, motion) vary greatly in their original spatial representations, feature alignment is necessary to map them into a unified semantic space. This is achieved by using a modality projection layer to project each modality onto a space of the same dimension, allowing features from different modalities to be fused within the same dimension. This solves the problem of dimensional mismatch in the original features of different modalities and aligns multimodal features in a unified semantic space. Specifically:
[0106] (1) Pet visual features: The 2048-dimensional pet visual features V(t) extracted by the ResNet-50 model are processed through a linear layer. V Projected into a 768-dimensional pet visual feature vector Z v The calculation formula is: Z v =Linear V (V(t));
[0107] (2) Pet audio features: The 1536-dimensional pet audio features A(t) extracted by the WavLM model are projected onto the 768-dimensional pet audio features Z through a linear layer. a The calculation formula is: Z a =Linear A (A(t));
[0108] (3) Pet location features: The 256-dimensional pet location features G(t) obtained through geocoding are processed through a linear layer. G Projecting the pet's positional features Z in 768 dimensions g The calculation formula is: Z g =Linear G (G(t));
[0109] (4) Pet movement features: The 512-dimensional pet movement features M(t) extracted through behavior pattern recognition are processed through a linear layer. M Projecting the 768-dimensional pet motion features Z m The calculation formula is: Z m =Linear M (M(t));
[0110] After projection, the originally multidimensional visual, audio, positional, and motion features are all unified into a 768-dimensional vector (Z). v Z a Z g Z m This lays the foundation for subsequent cross-modal fusion.
[0111] S3-2-2, The pet multimodal features in the same dimensional space are concatenated to obtain the pet multimodal feature matrix;
[0112] This step aims to achieve the structured integration of single-modal features of the same dimension and to mine temporal and spatial correlations through Transformer encoding, providing a feature carrier rich in spatiotemporal semantics for accurate calculation of cross-modal weights. First, based on the 768-dimensional multimodal feature vectors obtained after feature mapping, they are structurally concatenated according to a preset modal priority order (e.g., visual → audio → location → motion). Each single-modal feature vector is treated as a row, and the four vectors are stacked sequentially to form a multimodal feature matrix with a basic shape of (4, 768). This matrix initially integrates the semantic information of each modality, but it has not yet captured the dynamic changes of data over time or the correlation patterns under different spatial scenarios. To solve this problem, the Transformer encoding mechanism is introduced. The Transformer's self-attention mechanism excels at handling long-distance dependencies in sequential data and can effectively mine the correlations of multimodal data over time and space. Specifically:
[0113] After obtaining four 768-dimensional feature vectors, the features are concatenated according to the preset modality priority order. That is, each single-modality feature vector is used as a row, and the four 768-dimensional vectors are stacked in sequence to form a (4,768) pet multimodal feature matrix. The rows of the matrix correspond to different modality types, and the columns correspond to the feature dimensions under a unified semantic space. This matrix completely preserves the feature information of each modality, and organizes the scattered single-modality features into a structured data form that is convenient for attention mechanism analysis of cross-modal correlation.
[0114] The formula for calculating feature splicing is as follows:
[0115] Zconcat=Concat[(Z v Z a Z g Z m )]
[0116] Where Zconcat represents the pet multimodal feature matrix obtained after concatenation, and Concat is the feature concatenation operation.
[0117] S3-2-3. Based on the pet multimodal feature matrix, a multimodal attention mechanism is used to obtain the weight matrix of the pet multimodal features;
[0118] This step assigns cross-modal weights to features from different modalities using an attention mechanism. This highlights modal features that are more important to the current task, making the fusion process more targeted. Using a (4,768) pet multimodal feature matrix as input, a self-attention computation framework is employed for processing. The specific steps are as follows:
[0119] (1) Perform a linear transformation on each modal eigenvector in the matrix to generate the corresponding query (Q), key (K), and value (V) vectors. The calculation formula is as follows:
[0120] Q=Zconcat∙W Q K=Zconcat∙W K V=Zconcat∙W V
[0121] Among them, W Q W K W V is a learnable parameter matrix.
[0122] (2) Calculate the similarity score between different modal features by performing a dot product operation between the Q and K vectors, and then scale (divide by) the result. , The feature dimension is normalized using the Softmax function to obtain the attention weight matrix between each modality, which is the output (4,4) pet multimodal feature weight matrix. The element in the i-th row and j-th column of the matrix represents the contribution weight of the j-th modal feature to the i-th modal feature. The larger the weight value, the higher the importance of the corresponding modality in the fusion process.
[0123] The formula for calculating the attention weight matrix is:
[0124]
[0125] Among them, Attention weights Here, is the attention weight matrix, Softmax is the normalization function, and T is the transpose. For feature dimensions.
[0126] S3-2-4. Based on the weight matrix of the pet multimodal features, perform weighted fusion of the pet multimodal features in the same dimensional space to obtain the pet multimodal fusion features;
[0127] This step weights and aggregates features from each modality based on attention weights to generate a fusion feature that integrates key information from multiple modalities. First, the attention weight matrix is... weights Multiplying with the V matrix, the core features (V) of each modality are weighted and summed according to their attention levels to obtain the feature Z of each modality after cross-modal correlation. attended (Still a 4×768 matrix), the formula is:
[0128] Z attended =Attention weights ∙V
[0129] Then, a learnable weight vector α=[α0,α1,α2,α3] (initial value [0.4,0.3,0.2,0.1]) is introduced for Z. attended The four rows of feature vectors are weighted and summed to obtain the final 768-dimensional multimodal fusion feature Z. aligned The calculation formula is:
[0130] Z aligned =α0∙Z attended [0]+α1∙Z attended [1]+α2∙Z attended [2]+α3∙Z attended [3]
[0131] After the above processing, a 768-dimensional fusion feature vector is finally obtained. This vector is the pet multimodal fusion feature, which integrates information from multiple modalities such as vision, audio, location, and motion. It reflects the differences in contribution of different modalities through attention weights, and can accurately reflect the comprehensive semantic information of pet behavior, state, and environment, providing high-quality feature input for subsequent tasks such as emotion recognition and event analysis.
[0132] In summary, steps S3-2-1 to S3-2-4 achieve deep fusion of four modalities: visual, auditory, positional, and motor. This breaks through the limitations of single-modal data and solves the problem of data fragmentation in existing devices, making information acquisition more comprehensive. By fusing features, multimodal emotional cues can be integrated simultaneously, providing comprehensive data support for subsequent pet emotion recognition and helping to achieve an emotion recognition accuracy of 89.3%. The fusion process dynamically adjusts modal weights through an attention mechanism, improving feature targeting and providing a high-quality data foundation for deep intelligent analysis in the cloud, further ensuring that the cloud analysis accuracy is >85%. At the same time, it lays the core technical foundation for saving manual time in the subsequent automatic generation of coherent pet vlogs.
[0133] As one possible implementation, in the above embodiments, step S3-3 may specifically include the following steps:
[0134] S3-3-1. Based on the pet's multimodal fusion features, use an emotion analysis model to obtain information on the intensity of the pet's emotions;
[0135] The pet multimodal fusion feature is a 768-dimensional aligned fusion feature Z obtained after cross-modal fusion alignment. aligned The data is then fed into a sentiment analysis model (such as a classifier) to calculate scores for seven sentiment categories. In this way, the model can simultaneously utilize emotional cues from visual, audio, location, and motion (such as facial expressions, tone of voice, scene location, and amplitude of movement), thereby improving the accuracy of sentiment recognition.
[0136] After inputting the pet's multimodal fusion features into the sentiment analysis model, seven types of pet emotional states can be identified: happiness, sadness, anxiety, anger, fear, tension, and relaxation. Each emotional state is quantitatively represented by a seven-dimensional vector [activity level, sadness level, anxiety level, aggression level, fear level, tension level, and relaxation level]. The expression of the emotion in the seven dimensions is quantified by the numerical values (between 0 and 1) of different dimensions. For example, the vector corresponding to "happy" is [0.8,0.2,0.1,0.0,0.0,0.1,0.9], the vector corresponding to "sad" is [0.1,0.9,0.2,0.3,0.1,0.8,0.2], the vector corresponding to "anxious" is [0.2,0.6,0.9,0.2,0.5,0.7,0.3], the vector corresponding to "angry" is [0.3,0.4,0.7,0.9,0.1,0.6,0.1], the vector corresponding to "fear" is [0.1,0.5,0.8,0.2,0.9,0.9,0.2], the vector corresponding to "nervous" is [0.4,0.3,0.6,0.1,0.4,0.8,0.3], and the vector corresponding to "relaxed" is [0.6,0.1,0.1,0.0,0.1,0.2,0.9]. Simultaneously, emotional intensity information is calculated, which is derived using the following formula:
[0137] base score score :base score =max(emotion probs )×100 (Contribution of the highest probability emotion);
[0138] Confidence reward bonus confidence bonus =entropy penalty (emotion probs (The more concentrated the sentiment classification results, the higher the reward).
[0139] Stability reward bonus stability bonus =temporal stability (emotion history )×5 (The longer the emotional duration, the higher the reward);
[0140] Emotional intensity score: Score =min(base score +confidence bonus +stability bonus (100) (The maximum total score is 100).
[0141] Where max is the maximum value function, and emotion probs The probability distribution values output by the sentiment analysis model, entropy penalty For entropy penalty function, temporal stability Let `emotion` be a time-stability function. history This is a historical dataset of pet emotional states.
[0142] The rating not only considers the most likely emotional base at present score It also incorporates the model's confidence. bonus (For example, the more concentrated the sentiment classification results, the more points are awarded) and the time stability of sentiment. bonus (For example, if the "happiness" lasts longer, a bonus will be awarded), making the rating more reliable.
[0143] In addition, emotional information can be obtained in another way: First, for the four types of multimodal data—visual, audio, motion, and location—dedicated emotion networks (such as EmotionNet) can be used respectively. V First, extract the emotional features of each modality. Second, calculate the dynamic weight β of each modality using Softmax, giving higher weights to modalities that better reflect the current emotion (such as visual expressions when happy). Finally, sum the weighted features to obtain the fused emotional feature E. fused The accuracy is improved by integrating multimodal information; finally, the fused emotional features are input into the sentiment analysis model to obtain emotional intensity information of the same dimension.
[0144] S3-3-2. Perform event segmentation and clustering based on the pet multimodal fusion features to obtain pet event segmentation results;
[0145] Based on the multimodal fusion features of pets and the associated individual modal features (visual, audio, location, motion), the rate of change of each individual modality is first calculated using a formula to quantify the severity of the individual modal change (the larger the value, the more obvious the change):
[0146] The rate of change of visual features ΔV(t): ΔV(t) = ||V(t) - V(t-1)||2 / ||V(t-1)||2 (representing visual changes such as scene switching);
[0147] The rate of change of audio features ΔA(t): ΔA(t) = ||A(t) - A(t-1)||2 / ||A(t-1)||2 (characterizing abrupt changes in sound);
[0148] The rate of change of location characteristic ΔG(t): ΔG(t) = ||G(t) - G(t-1)||2 / ||G(t-1)||2 (characterizing a sudden move to a new location);
[0149] Rate of change of motion characteristic ΔM(t): ΔM(t) = ||M(t) - M(t-1)||2 / ||M(t-1)||2 (characterizing sudden cessation of motion, etc.);
[0150] Wherein, V(t) is the visual feature at time t, V(t-1) is the visual feature at time t-1, A(t) is the audio feature at time t, A(t-1) is the audio feature at time t-1, G(t) is the positional feature at time t, G(t-1) is the positional feature at time t-1, M(t) is the motion feature at time t, and M(t-1) is the motion feature at time t-1.
[0151] The aforementioned rate of change is calculated using the L2 norm ratio of features at adjacent time points, quantifying the severity of single-mode changes; a larger value indicates more severe single-mode changes. Subsequently, the comprehensive rate of change Δtotal(t) = w0×ΔV(t) + w1×ΔA(t) + w2×ΔG(t) + w3×ΔM(t) is calculated according to weights w = [w0, w1, w2, w3] = [0.4, 0.3, 0.2, 0.1], and a dynamic threshold is set. adaptive =0.15×(1+std(Δtotal(t),window=60)) (considering fluctuations within 60 frames), where std is the standard deviation, used to quantify the fluctuation of the overall rate of change Δtotal(t) within the sliding window, and window is the size of the sliding window; when the overall rate of change Δtotal(t) exceeds this dynamic threshold, that moment is determined as a "change point", that is, the boundary of event segmentation, thereby segmenting continuous pet behavior data into multiple independent events, such as switching from the "playing" event to the "eating" event), and finally obtaining the pet event segmentation result.
[0152] S3-3-3: Based on the pet event segmentation results, KMeans clustering and topic model are used to obtain pet event topics;
[0153] First, a "vocabulary" is constructed based on the multimodal features corresponding to the event segmentation results: The KMeans clustering algorithm transforms continuous multimodal features into discrete "vocabularies" (analogous to text words). For example, visual features are clustered into 1000 "visual words," and audio features into 500 "audio words." Motion and position features are clustered similarly, resulting in a total vocabulary of 2000 words, making non-textual data suitable for LDA model processing. Subsequently, the LDA topic model is used, and the topic-word distribution β is estimated using the Gibbs sampling method. k Document-Topic Distribution θ d :
[0154] Topic-word distribution β k : Represents the "vocabulary" composition of theme k, such as "grass + laughter + running" corresponding to the theme "outdoor play";
[0155] Document-Topic Distribution θ d : Represents the probability that a certain segmented event belongs to each topic;
[0156] The above process uncovers potential themes in pet behavior (such as "outdoor play", "indoor modification", "eating"), ultimately yielding the theme of pet events.
[0157] S3-3-4. Use the pet's emotional intensity information to extract highlights and obtain pet highlight segments;
[0158] Highlight fragment extraction is not simply a matter of selecting high-resolution keyframes, but rather a combination of two strategies: first, using keyframes... score First, select high-quality single-frame footage (strong dynamics, prominent subject, and rich emotion); second, use a three-act structure to ensure the extracted clips cover "beginning-development-ending," avoiding highlight editing becoming a mere collection of fragmented footage; ultimately, achieve the extraction of highlight clips that are both captivating and narratively complete, providing support for the core content presentation of the Vlog. The specific steps are as follows:
[0159] S3-3-4-1. Utilize the pet emotional intensity information to perform keyframe calculations, obtain the comprehensive score of each frame corresponding to the pet emotional intensity information, and filter out high-value frames.
[0160] Each frame is scored across four dimensions: dynamism, subject prominence, emotional intensity, and scene complexity. Weights reflect the importance of each dimension, with high-scoring keyframes corresponding to the core scenes of highlight sequences. The overall score for each frame is calculated using the keyframe function. score Calculated using the following formula:
[0161] keyframe score =(0.4×motion area_ratio +0.3×pet attention_score +0.2×emotion intensity (t)+0.1×scene complexity_score )
[0162] Among them, motion area_ratio The percentage of the moving area represents the richness of movement in the video. For example, when a pet is running or playing, a larger moving area / more varied movements result in a higher score and the highest weight (40%), because dynamic images are more likely to attract viewers. attention_score Pet prominence, or how clearly the pet is focused in the frame, is scored based on how prominent / clear the pet is in close-up shots. This score accounts for 30% of the overall score. Ensure that highlight scenes emphasize the main subject (e.g., the pet). (emotion) intensity(t) represents the emotional intensity, i.e., the intensity of the emotion corresponding to the current frame. For example, the higher the score for the "happy" emotional intensity, the higher the score for this dimension, with a weight of 20%, to ensure a scene full of emotion; scene complexity_score Scene complexity refers to the richness of environmental elements in the scene. The richer the scene information, the higher the score. For example, a park scene scores higher than a blank wall scene. It has a weight of 10% to balance the amount of information in the scene.
[0163] The overall score for each frame (keyframe) score The higher the score, the more likely the frame is to be a highlight, such as a close-up shot of a pet running excitedly. Subsequent frames containing high-scoring frames will be prioritized as highlights.
[0164] S3-3-4-2. Based on the comprehensive score of each frame corresponding to the pet's emotional intensity information, perform plot structure analysis to obtain the pet's highlight segments and ensure narrative logic;
[0165] The video clips are divided according to classic narrative structures to ensure that highlight edits conform to the story's logic. For example, the opening retains scene introductions, while the development section emphasizes emotional conflict. Through plot structure analysis, specifically dividing the clips into a "three-act structure," the highlight segments are ensured to not only contain exciting visuals but also form a complete storyline.
[0166] Start and end points setup end setup end =min(peaks[0],25% duration), that is, take the smaller value between the first emotional peak position peaks[0] and 25% of the total video length to ensure that the beginning (introduction part) is not too long and ends before the first emotional climax;
[0167] Ending starting point resolution start resolution start =max(valleys[-1],75% duration), which means taking the larger value between the last emotional trough position valleys[-1] and 75% of the total video length, to ensure that the ending (the closing part) includes an emotionally stable phase and does not start too early.
[0168] The event segmentation results are divided into a three-act structure story. structure It consists of three parts: "beginning (introduction of the scene), development (core of emotional fluctuations), and ending (emotional conclusion)," and the calculation formula is as follows:
[0169] story structure ={setup,conflict,resolution}.
[0170] The role of the three-act structure:
[0171] Setup: Extract basic information such as scene introduction and the pet's first appearance to let the audience understand the background.
[0172] Conflict: This includes the segments with the greatest emotional fluctuations and the most concentrated events, such as pets playing or encountering small challenges. It is the core source of highlight content.
[0173] Resolution: Select a scene where emotions have calmed down and the event has come to an end, such as a pet resting, to ensure the integrity of the story.
[0174] Ultimately, by combining high-resolution keyframes and a three-act structure, segments that "contain core images that are emotionally intense and dynamically rich, and also cover the complete narrative logic" were selected, avoiding the fragmented accumulation of highlight images, thus obtaining the highlight segments of the pet.
[0175] S3-3-5, The analysis results of obtaining the pet's emotional intensity information, the pet's event theme, and the pet's highlight segments as pet multimodal fusion features;
[0176] The outputs of the above steps are summarized as follows: First, pet emotional intensity information, including the pet's emotional state (one of seven types of emotions) and the corresponding seven-dimensional quantitative vector and emotional intensity score (0-100 points); second, pet event theme, such as "outdoor play" and "indoor rest"; third, pet highlight clips, which are core clips selected by combining keyframe scores and a three-act structure. These three together constitute the analysis results of pet multimodal fusion features, which can provide core basis for subsequent intelligent captioning and music matching, automatic generation of pet vlogs, and realize the transformation from fragmented multimodal data to structured analysis results.
[0177] In summary, event segmentation and topic clustering determine the narrative structure of a Vlog, highlight extraction filters core content, and ultimately together support the automatic generation of coherent Vlogs from multimodal data. This achieves the core logic of highlight segment extraction, and through keyframe scoring and plot structure analysis, the most expressive content is selected from the video while ensuring the integrity of the narrative.
[0178] As one possible implementation, in the above embodiments, step S4 may specifically include the following steps:
[0179] S4-1. Using the analysis results of the pet multimodal fusion features and the mapping relationship between emotion and music parameters, obtain the background music that matches the emotion.
[0180] The analysis results of pet multimodal fusion features include pet emotional states (such as happiness, relaxation, sadness, etc., totaling 7 categories) and corresponding emotional intensity information. Based on this, a fixed mapping relationship between emotions and music parameters is first established, such as "happiness" corresponding to fast-paced, major-key music, to ensure that the background music matches the video's emotion. The fixed mapping relationship between emotions and music parameters is then established. emotionmap The calculation formula is as follows:
[0181] music emotionmap ={Happy: {tempo:120-140,key:major,energy:high},Relaxed: {tempo:60-80,key:major,energy:low},Sad: {tempo:50-70,key:minor,energy:low}}
[0182] This mapping relationship clearly defines the tempo, key, and energy corresponding to different emotions. Specifically, "happy" emotion corresponds to a fast tempo, major key, and high energy of 120-140, "relaxed" emotion corresponds to a slow tempo, major key, and low energy of 60-80, and "sad" emotion corresponds to a tempo, minor key, and low energy of 50-70. This mapping ensures that the style of the background music is highly consistent with the pet's current emotional state.
[0183] Meanwhile, emotional intensity information, such as a "happy" emotional intensity score of 90, can be used to further fine-tune music parameters, such as appropriately increasing the music energy value when the emotional intensity is high, and finally matching and filtering background music content that is perfectly matched with the pet's emotions from the music library.
[0184] S4-2. Based on the analysis results of the pet multimodal fusion features, subtitle content is obtained through a generative pre-trained transformer model;
[0185] The analysis results of pet multimodal fusion features cover pet behavior, environment, and emotion. These three types of information are used as input prompts and substituted into a preset caption generation template. template The template explicitly requires "generating humorous captions of ≤20 characters based on pet actions, environmental objects, and emotional states." Subsequently, a Generative Pre-trained Transformer (GPT) model is used to perform semantic understanding and content generation on the input prompts. For example, when the input is "pet action: chasing butterflies, environmental objects: butterflies in the park, emotional state: happy," the model can generate captions such as "The puppy is chasing butterflies and is so happy it's flying!" which fits the scene and is also humorous, ensuring that the captions can accurately describe the content of the scene and convey the pet's emotions.
[0186] S4-3. Based on the analysis results of the pet multimodal fusion features, the background music of the emotion matching and the subtitle content, obtain the pet neck camera Vlog generation result;
[0187] The analysis results of pet multimodal fusion features provide a core narrative framework and content foundation for Vlogs. Among them, the event segmentation results determine the logical order of events in the Vlog (such as "departure - outdoor play - rest"), the event theme (such as "outdoor play") clarifies the core style positioning of the Vlog, and the highlight clips are the core visual content of the Vlog. The emotionally matched background music is synchronized with the highlight clips along the timeline (such as the "happy" style background music corresponding to the highlight clip of "outdoor play"). Then, the subtitles generated by the GPT model are accurately superimposed on the corresponding frame, such as superimposing fun subtitles on the scene of the pet chasing butterflies. Finally, the integration forms a coherent pet neck camera Vlog with "complete narrative logic, emotionally appropriate background music, and scene-based fun subtitles", realizing the fully automated generation from multimodal data to finished Vlog without human intervention.
[0188] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a pet neck-hanging camera Vlog generation device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0189] like Figure 2 As shown, a pet neck camera Vlog generation device in this embodiment includes: a hardware acquisition module, an edge processing module, a cloud AI module, and an application service module;
[0190] The hardware acquisition module is used to acquire multimodal data of pets using a pet neck camera;
[0191] This module uses a detachable neck-mounted pet thumb-sized action camera as its core hardware. It features optical image stabilization (using a suspended lens structure to compensate for camera vibration), electronic image stabilization (a digital image stabilization algorithm based on gyroscope data), adaptive adjustment (dynamically adjusting stabilization intensity according to the pet's movement), autofocus (real-time focus adjustment based on target detection), motion prediction (predicting the pet's movement trajectory using a motion sensor), and variable focal length (intelligently adjusting the camera's focal length to adapt to different shooting distances). The hardware configuration includes a 1080P high-definition wide-angle night vision camera, microphone, 4G+GPS module, motion sensor, and rechargeable battery, enabling real-time acquisition of multimodal data such as video, audio, location, and motion. Simultaneously, the module is equipped with a time synchronization unit, employing a unified timestamp system to achieve data time synchronization, ensuring strict time alignment of video, audio, GPS, and IMU data with millisecond-level synchronization accuracy, providing raw data support for the entire system.
[0192] The edge processing module is used to preprocess the pet multimodal data to obtain preprocessed pet multimodal data.
[0193] This module performs lightweight AI inference and undertakes the preprocessing of multimodal data. It organizes the multimodal data acquired by the hardware acquisition module, performs video and audio compression to reduce the data volume, and uses the Kalman filter algorithm to denoise the sensor data to improve data accuracy. At the same time, it uses the TensorFlow Lite framework to detect device malfunctions such as black screens and electrical noises. After completing the data preprocessing, the data is transmitted to the cloud AI module to achieve preliminary data optimization and screening.
[0194] The cloud-based AI module is used to extract and fuse the preprocessed pet multimodal data to obtain the analysis results of the pet multimodal fusion features.
[0195] This module performs deep intelligent analysis and content generation. First, the multimodal feature extraction unit extracts visual, audio, location, and motion features based on the preprocessed pet multimodal data output by the edge processing module, using ResNet-50, WavLM, geocoding, and behavioral pattern recognition technologies. Then, the cross-modal fusion unit unifies the features of each modality into a 768-dimensional vector through a modal projection layer, and then uses a multimodal attention mechanism to calculate cross-modal weights and perform weighted fusion of features to generate a 768-dimensional fused feature that integrates multimodal information. Finally, relying on the intelligent analysis unit, it completes the quantitative scoring and feature vector generation of seven types of pet emotions such as happiness and sadness, event segmentation based on comprehensive change rate and dynamic threshold, KMeans+LDA topic clustering, and extraction of highlight segments combined with weighted scoring and a three-act structure. The final output is a multimodal fusion feature analysis result containing emotional state, event theme, and highlight material, providing core decision-making basis for Vlog generation.
[0196] The application service module is used to generate pet neck camera Vlog results based on the analysis results of the pet multimodal fusion features;
[0197] This module serves as a bridge connecting cloud-based analytics results with user needs, offering services such as automatic Vlog generation, user interaction, and content management. For Vlog generation, based on cloud-based analytics, the module automatically matches background music parameters to match the emotions conveyed (e.g., "happy" is paired with high-energy major key music), generates engaging subtitles of ≤20 characters using a GPT model, and integrates highlight segments, background music, and subtitles according to a three-act narrative logic to create a complete pet Vlog. Regarding user services, the module allows users to view multimodal data captured by the camera in real-time via a mobile app, engage in two-way communication with their pets, and provides Vlog content management functions (viewing, secondary editing, and sharing to social media platforms), achieving a closed loop from intelligent generation to user experience and fully meeting users' needs for recording and interacting with their pets' emotions.
[0198] In summary, this device adopts a four-layer architecture design: hardware acquisition + edge preprocessing + cloud AI analysis + application services. Its core technological effects are reflected in three aspects: First, it solves the data fragmentation problem of traditional devices by achieving comprehensive acquisition of pet information through multimodal synchronous acquisition and deep fusion; second, it fills the gap in emotional understanding, with an emotion recognition model based on fusion features achieving an accuracy rate of 89.3%, accurately capturing the emotional state of pets; and third, it eliminates reliance on manual labor, achieving fully automated generation from raw materials to finished vlogs, saving over 95% of manual editing time. Simultaneously, the edge-cloud collaborative architecture between modules balances real-time performance and intelligence, ensuring both edge data processing latency of <100ms and cloud analysis accuracy exceeding 85%, effectively enhancing the emotional connection between humans and pets, supporting pet health management, and maximizing the value of pet-related data.
[0199] In this embodiment, the specific processing of a pet neck-hanging camera Vlog generation device and its resulting technical effects can be referred to separately. Figure 1 The relevant descriptions of steps S1, S2, S3 and S4 in the corresponding embodiments will not be repeated here.
[0200] It should be noted that the implementation details and technical effects of each module and unit in the device provided in the embodiments of this disclosure can be referred to the description of other embodiments in this disclosure, and will not be repeated here.
[0201] The following is for reference. Figure 3 It shows a schematic diagram of the structure of a computer system 500 suitable for implementing the electronic device of the present disclosure. Figure 3 The computer system 500 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0202] like Figure 3As shown, the computer system 500 may include a processing device 501 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM 502) or a program loaded from a storage device 508 into a random access memory (RAM 503). The RAM 503 also stores various programs and data required for the operation of the computer system 500. The processing device 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output interface (I / O interface 505) is also connected to the bus 504.
[0203] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows computer system 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 A computer system 500 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0204] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0205] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0206] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0207] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 1 The illustrated embodiments and their alternative implementations demonstrate a method for generating pet neck camera vlogs.
[0208] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0210] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the unit itself; for example, an acquisition module can also be described as "acquiring preset prompts, including modality fusion prompts, attention mechanism prompts, and / or time-related prompts."
[0211] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method for generating pet neck-mounted camera Vlogs, characterized in that, include: S1. Use a pet neck camera to acquire multimodal data of the pet; S2. Preprocess the pet multimodal data to obtain preprocessed pet multimodal data; S3. Perform multimodal feature extraction and fusion analysis on the preprocessed pet multimodal data to obtain the analysis results of the pet multimodal fusion features, including: Multimodal features are extracted using the preprocessed pet multimodal data to obtain pet multimodal features, which include pet visual features, pet audio features, pet location features, and pet movement features. Cross-modal fusion is performed on the pet multimodal features to obtain pet multimodal fused features; The analysis results of the pet multimodal fusion features are obtained by analyzing and processing the pet multimodal fusion features, including: Based on the aforementioned pet multimodal fusion features, an emotion analysis model is used to obtain information on the intensity of pet emotions. Based on the pet multimodal fusion features, event segmentation and clustering are performed to obtain pet event segmentation results; Based on the pet event segmentation results, KMeans clustering and topic modeling are used to obtain pet event topics; Highlight extraction is performed using the pet's emotional intensity information to obtain highlight segments of the pet, including: The pet's emotional intensity information is used to perform keyframe calculations to obtain a comprehensive score for each frame corresponding to the pet's emotional intensity information. Based on the comprehensive score of each frame corresponding to the pet's emotional intensity information, the plot structure is analyzed to obtain the pet's highlight segments; The analysis results of obtaining the pet's emotional intensity information, the pet's event theme, and the pet's highlight segments are used as the pet's multimodal fusion features; S4. Based on the analysis results of the pet multimodal fusion features, generate the pet neck camera Vlog.
2. The method for generating pet neck camera Vlogs according to claim 1, characterized in that, S1. Using a pet neck camera, acquire multimodal data about the pet, including: Data was collected using a pet neck camera to obtain initial multimodal data of the pet; Determine whether the timestamps corresponding to the initial pet multimodal data are consistent. If they are, then obtain the initial pet multimodal data as the pet multimodal data; otherwise, perform the first operation. The first operation is as follows: the initial pet multimodal data is time-synchronized using a time synchronization mechanism to obtain pet multimodal data, which includes pet video data, pet audio data, pet GPS data, and pet IMU data.
3. The method for generating pet neck camera Vlogs according to claim 2, characterized in that, S2. Preprocess the pet multimodal data to obtain preprocessed pet multimodal data, including: The pet video data and the pet audio data are compressed to obtain compressed pet video data and compressed pet audio data respectively; The pet GPS data is processed to obtain the processed pet GPS data; The Kalman filter algorithm is used to filter noise from the pet IMU data to obtain filtered pet IMU data. The compressed pet video data, the compressed pet audio data, the normalized pet GPS data, and the filtered pet IMU data are integrated to obtain initial preprocessed pet multimodal data; Anomaly detection is performed based on the initial preprocessed pet multimodal data to obtain the preprocessed pet multimodal data.
4. The method for generating pet neck camera Vlogs according to claim 1, characterized in that, Cross-modal fusion of the pet multimodal features to obtain pet multimodal fused features includes: The pet multimodal features are used for feature mapping to obtain pet multimodal features in the same dimensional space; The pet multimodal features in the same dimensional space are concatenated to obtain a pet multimodal feature matrix; Based on the pet multimodal feature matrix, a multimodal attention mechanism is used to obtain the weight matrix of the pet multimodal features; The pet multimodal features in the same dimensional space are weighted and fused according to the weight matrix of the pet multimodal features to obtain the pet multimodal fused features.
5. The method for generating pet neck-mounted camera Vlogs according to claim 1, characterized in that, S4. Based on the analysis results of the pet multimodal fusion features, generate the pet neck camera Vlog generation results, including: By combining the analysis results of the pet multimodal fusion features with the mapping relationship between emotion and music parameters, emotional matching background music is obtained; Based on the analysis results of the pet multimodal fusion features, subtitle content is obtained through a generative pre-trained transformer model; Based on the analysis results of the pet's multimodal fusion features, the background music matched with the emotion, and the subtitle content, the pet neck camera Vlog generation results are obtained.
6. A pet neck-mounted camera Vlog generation device, employing the method described in any one of claims 1-5, characterized in that, include: Hardware acquisition module, edge processing module, cloud AI module and application service module; The hardware acquisition module is used to acquire multimodal data of pets using a pet neck camera; The edge processing module is used to preprocess the pet multimodal data to obtain preprocessed pet multimodal data. The cloud-based AI module is used to extract and fuse the preprocessed pet multimodal data to obtain the analysis results of the pet multimodal fusion features. The application service module is used to generate pet neck camera Vlog results based on the analysis results of the pet multimodal fusion features.
7. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Automatic editing method and system for pet video
CN119364111A
Deep learning-based pet dog emotion recognition method and system
CN120708251A