Method, system and medium for generating wonderful pet videos
By extracting feature data from each frame of pet videos and evaluating their excitement using a three-step method, efficient and accurate pet exciting videos are generated. This solves the problems of inaccurate excitement judgment and complex processes in existing technologies, and improves pet owners' sense of happiness and video generation efficiency.
Patent Information
- Application Number
- CN202310187158.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-03-01
AI Technical Summary
When generating wonderful pet videos, the existing technology has inaccurate judgment of the degree of wonderfulness and a complicated process, resulting in a large workload and a long time consumption, and is unable to efficiently generate high-quality wonderful pet videos.
By extracting feature data from each frame of the video, a three-step method is used to evaluate the excitement, including calculating the numerical value, average value and verification of the feature data of each frame, combining it with the confidence score, editing and synthesizing the wonderful pet video, and selecting the presentation method based on the application score.
It improves the accuracy and generation efficiency of wonderful pet videos, enhances the happiness of pet owners, reduces resource waste, and improves the viewing experience.
Smart Images

Figure CN116366994B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pet videos, and more specifically, relates to a method, system and medium for generating wonderful pet videos. Background Art
[0002] With economic development, the pace of urban life has become increasingly fast, and interpersonal communication has become increasingly less frequent. Consequently, more and more people are keeping pets for personal companionship. Many pet owners enjoy sharing videos of their pets, showcasing their pets' adorable moments. However, since pet owners are often unable to record videos at home due to work or other reasons, they often miss out on many of their pets' fun moments. To address this issue, some pet owners install cameras at home to monitor their pets' behavior. However, interesting and engaging pet videos are not always produced, and the monitored data often contains a large amount of mundane video content. Sorting out the highlights from this mundane video is labor-intensive and time-consuming, and highlights are easily overlooked, with uncertain accuracy and quality.
[0003] Corresponding improvements have also been made to address the above-mentioned issues. For example, Chinese patent application number CN202111123623.4, published on December 10, 2021, discloses a pet video recording and automatic editing method based on an AI algorithm, including a device side, a backend, cloud storage, a CV cloud computing server, and an APP side. The device side is connected to the backend and cloud storage, the backend is connected to the cloud storage, the CV cloud computing server, and the APP side, the cloud storage is connected to the CV cloud computing server, and the CV cloud computing server is connected to the APP side. The backend includes a file management system and business logic. The method includes the following steps: S1: the device collects information, takes photos or records videos; S2: information and photos or videos are recorded and backed up; S3: picture or video processing; S4: video distribution; S5: system optimization. The shortcomings of this patent are that: although it can help pet owners eliminate unnecessary video information, the accuracy of the entire wonderful video is not high, and the specific steps of how to generate the wonderful video are not disclosed.
[0004] Another example is Chinese patent application number CN202011338142.0, published on March 5, 2021. The patent discloses a method for automatically generating a video collection based on content analysis, including: pre-screening the original video according to preset screening rules to obtain multiple original video clips; using the KTS algorithm to divide the video content after the pre-screening into multiple continuous clips; using the fscn algorithm to analyze the video highlights of each continuous clip to obtain multiple candidate wonderful continuous clips; scoring and assigning weights to the image quality, face detection and analysis, and age of each candidate wonderful continuous clip respectively, and combining the video content pornography identification results to give a final score to each candidate wonderful continuous clip; screening out multiple final wonderful continuous clips based on the final scoring results; adding special effects and transition effects at the junction of each final wonderful continuous clip to generate a video collection. The shortcomings of this patent are that the whole process is complicated and time-consuming. Summary of the Invention
[0005] 1. Problems to be solved
[0006] To address the current issues of limited coverage of pet highlight videos and inaccurate highlight assessment, the present invention provides a method, system, and medium for generating pet highlight videos. The highlight assessment process in this invention avoids the large errors and susceptibility to interference that can occur with a single assessment, thereby improving accuracy. The entire method is simple, enabling accurate assessment of pet highlight videos, thereby enhancing pet owners' happiness. It also enables rapid assessment without the need for complex steps or models for training or recognition, thus improving efficiency. The system of this invention is simple in composition and stable in operation.
[0007] 2. Technical solution
[0008] To solve the above problems, the present invention adopts the following technical solutions.
[0009] A method for generating a wonderful pet video comprises the following steps:
[0010] S1: Obtain feature data of each frame in the video;
[0011] S2: Evaluate the excitement of the feature data. Specifically, the excitement evaluation includes first calculating the value of each frame of feature data, then calculating the average value of all frame feature data values, and finally verifying the average value to obtain an excitement evaluation result.
[0012] S3: If the requirements for excitement are met, each frame meeting the requirements for excitement is automatically edited and synthesized to generate an exciting pet video.
[0013] Furthermore, the step S2 specifically includes the following steps:
[0014] S21: Calculate the value of the feature data of each frame: S = (100% + A + B + C) * Z, where A is the total label coefficient of each frame, B is the human-pet interaction coefficient, C is the image quality coefficient, and Z is the confidence score;
[0015] S22: Calculate the average value of all frames: S ’ =(S1+S2+S3+……+S N ) / N, where N is the number of all frames;
[0016] S23: Verify the average value: When the average value is greater than or equal to the set threshold, it is considered passed and proceeds to the subsequent steps; if it is less than the set threshold, it is considered failed and does not proceed to the subsequent steps.
[0017] Furthermore, step S23 also includes establishing a scoring countermeasure process, that is, when the verification is passed, the confidence score of the feature data of each frame is judged. When the confidence score is less than the threshold, it is deemed to have failed and no further steps are entered; when the verification fails, the confidence score of the feature data of each frame is judged. When the confidence score is greater than or equal to the threshold, it is deemed to have passed and subsequent steps are entered.
[0018] Furthermore, step S1 includes pre-analyzing each frame of the picture to obtain feature data in the picture: total label coefficient, human-pet interaction coefficient, picture quality coefficient and confidence; at the same time, the pre-analysis also includes determining whether there is misidentification and / or non-identification phenomenon in the picture.
[0019] Furthermore, the specific operations for generating the wonderful pet video in step S3 are as follows:
[0020] S31: Cropping: Cropping is performed at double speed from the original image to the moment the pet appears in each frame, and the image after the pet appears remains unchanged;
[0021] S32: Splicing: Splicing the previous and next frames. When the interval between the previous and next frames is less than or equal to a threshold, the images are filled in to fill the gap between the previous and next frames. When the interval between the previous and next frames is greater than the threshold, the previous and next frames are directly spliced.
[0022] S33: Generate cover: Select and display the cover of the video.
[0023] Furthermore, the step S33 specifically includes the following steps:
[0024] S331: Selecting a frame with the highest value in the video as the cover material;
[0025] S332: Pets are the center of the picture;
[0026] S333: Cutting out the portion of the video without the pet according to the aspect ratio of the original video;
[0027] S334: The cut-out image is used as the video cover;
[0028] S335: Determine the boundary value of the video cover.
[0029] Furthermore, the method further includes step S4: performing application rating on the generated wonderful pet videos, and generating different presentation modes of the wonderful pet videos according to different application rating results.
[0030] Furthermore, the presentation methods include: basic presentation: only displaying the generated wonderful pet videos;
[0031] Key presentation: Display not only the generated pet videos but also the copywriting that matches them;
[0032] Highlight presentation plus immediate push notification: Not only will the generated pet highlights video and the copywriting that matches the pet highlights video be displayed, but push notification will also be sent immediately;
[0033] Do not present: Do not display the video, but save it.
[0034] A system using any of the above methods for generating wonderful pet videos, comprising:
[0035] Data acquisition unit: used to obtain feature data of each frame in the video;
[0036] Wonderfulness evaluation unit: used to evaluate the wonderfulness of feature data;
[0037] Display unit: used to generate and display wonderful pet videos.
[0038] A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the method for generating a wonderful pet video described in any one of the above items is implemented.
[0039] 3. Beneficial effects
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] (1) The present invention extracts feature data from each frame of the video to ensure the comprehensiveness of the data, providing an accurate and comprehensive data basis for subsequent steps; then the feature data is evaluated for brilliance, and a three-step method is adopted for the evaluation of brilliance, first calculating the value of the feature data of a single frame, then calculating the average value of the feature data values of all frames, and finally verifying the average value. The whole process can effectively ensure the accuracy of the brilliance evaluation, avoid the problem of large errors and susceptibility to interference caused by a single evaluation, and improve the accuracy of the judgment of the brilliance video; then the corresponding pet brilliance video is generated according to the result of the brilliance evaluation; the whole method process is simple, which can not only realize the accurate judgment of the pet brilliance video, thereby improving the happiness of the pet owner, but also realize rapid judgment without introducing complex steps or models for training or recognition, thereby improving work efficiency;
[0042] (2) When verifying the average value, the present invention not only judges whether the average value is greater than or less than a set threshold, but also adds a scoring countermeasure process. The establishment of this process can eliminate the errors caused by the basic evaluation, avoid the occurrence of misscreening or failure to consider special circumstances due to the basic evaluation, and further increase the accuracy of the entire process. At the same time, when acquiring feature data, it also includes judging whether there is a bad case, that is, misidentification and / or non-identification phenomenon in each frame, thereby strengthening the recognition ability of pets, improving the accuracy and capture rate, and ensuring the authenticity and accuracy of the data at the source.
[0043] (3) In the process of generating wonderful pet videos, the present invention can play redundant, lengthy and boring video images at double speed in the cropping stage, so as to avoid losing interest due to long waiting time; in the splicing stage, the interval content can be supplemented in the interval switching stage, so as to avoid the connection between the front and the back being too abrupt and jumpy; in the cover generation stage, the frame with the highest value is selected as the cover material, which can attract attention; the entire generation stage can improve the integrity and focus of the wonderful video, avoid other boring video images taking up too much time and space, and improve the viewing experience of the wonderful video;
[0044] (4) The present invention also includes applying a rating to a wonderful pet video after it is generated, so that the video can be presented in different ways according to different rating results, making the presentation more accurate and the push more reasonable, thereby avoiding the poor user experience caused by the use of the same push method and the cost waste caused by the unreasonable application of resources, thereby further improving the pet owner's viewing experience of the wonderful pet video. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0046] The present invention is further described below with reference to specific embodiments and accompanying drawings.
[0047] Example 1
[0048] like Figure 1 As shown, a method for generating a wonderful pet video includes the following steps:
[0049] S1: Obtain feature data of each frame in the video; specifically, the feature data in the step includes pre-analysis of each frame, analyzing whether there is a corresponding label in the picture, whether there is a human-pet interaction scene, and picture quality indicators, so as to obtain the feature data in the picture: total label coefficient A, human-pet interaction coefficient B, picture quality coefficient C and confidence Z; at the same time, the pre-analysis also includes judging whether there is misidentification and / or non-recognition phenomenon in the picture, wherein misidentification means that the algorithm mistakenly thinks "plush toys", "hair" and other items are pets, and generates a wonderful video to show to the user, that is, the accuracy is problematic; non-recognition means that there is obviously a pet but the algorithm The probability of not recognizing or recognizing a pet is low, that is, the capture rate and accuracy are problematic. When a non-real "pet" such as a "plush toy" or "hair" is recognized in the video, "the pet" is misidentified in the Badcase. If there are multiple confidence levels, that is, multiple pets in the video, the confidence levels and recognition fields of all frames of "the pet" are discarded, and the confidence levels and recognition fields of the remaining pets are retained. If there is only the confidence level of "the pet" in the video, that is, there is only one pet, all confidence levels, recognition fields of "the pet", and the video are discarded. For example, if three pets are recognized in the picture, with confidence levels Z1, Z2, and Z3 respectively. Among them, Z3 is a plush toy, which is a misidentification in the Badcase. All confidence levels and recognition fields related to Z3 are discarded, and all confidence levels and recognition fields related to Z1 and Z2 are retained. One pet is recognized in the picture with confidence level Z1. Z1 is a plush toy, which is a misidentification in the Badcase. All confidence and recognition fields associated with Z3 are discarded, and the video is discarded. The purpose here is to automatically avoid non-pet items, strengthen pet recognition capabilities, and improve accuracy and capture rate.
[0050] Tags refer to the labels on the screen, which are divided into the following four categories:
[0051] (1) Pet identification tag: It can identify the specific pet and match the pet appearing in the video with the identified name;
[0052] (2) Pet behavior tag: specifically identifies the pet's actions, such as running, and can automatically identify various pet behaviors and match them with corresponding tags. This embodiment takes the following example: in the case of a single pet: if the edge smart camera, motion detection, gravity sensor, and distance sensor are all triggered, the tag is "eating"; if the X, Y, W, and H values all change at a very fast rate, the tag is "running"; if the Y, W, and H values change from small to large and then stop for 3 seconds without changing, and the gravity sensor senses that the weight in the weighing bowl is less than 5g, The label is "approaching the camera and staying"; if the Y, W, and H values change from large to small, the label is "leaving"; the rightward horizontal axis is the X-axis, and the downward vertical axis is the Y-axis. The X and Y values can be used to locate the pet's position, and the W and H values are the width and height of the pet. In the case of multiple pets: if the W and H values of one pet are close to those of another pet and the time period of the value change is relatively consistent, the label is "playful"; to further facilitate calculation, pet behavior labels are divided into three levels: high-quality behavior labels; general behavior labels; and ordinary behavior labels.
[0053] (3) Pet quantity label: A label that counts the number of pets and can accurately identify the number of pets in the video. The number of pets displayed will be displayed.
[0054] (4) Scenario tags: The combination of scenarios and pet behavior tags or multiple pet behavior tags constitutes scenario tags, making them more story-like. Specifically, the scenarios in the scenario tags may include: less than a period of time from the scheduled food delivery; 0:00-6:00 sleep time; time after eating; then the corresponding scenario tags may be: less than 15 minutes from the scheduled food delivery + approaching the camera and staying, it is judged as "waiting for dinner"; the time is 00:00-04:00 + eating, it is judged as "eating a midnight snack"; after eating + leaving, it is judged as "full", etc. This embodiment will not give detailed examples. In order to further facilitate calculation, the scenario tags are divided into three levels: high-quality scenario tags; general scenario tags and ordinary scenario tags;
[0055] Human-pet interaction refers to manual user-issued commands, meaning the device executes non-timed or non-quantitative commands, and then the pet triggers the generation of a wonderful video within a certain period of time (e.g., 30 seconds). This behavior requires both the user issuing the command and the pet triggering the generation of the video within the specified time. This behavior is considered a valid interaction, as both the user and the pet respond to each other, and this behavior has unique significance. Human-pet interaction scenarios include: real-time voice; voice playback; one-piece food delivery; taking photos; recording videos, etc. Human-pet interaction scenario recognition logic: The above behaviors occur within 30 seconds before the video starts.
[0056] The image quality index is based on computer vision's judgment of the pet's position in the picture, the pet's clarity, and the brightness of the picture, to analyze whether the picture looks good. The factors that affect the image quality include the pet's position in the picture, the pet's clarity, and the brightness of the picture. Therefore, to facilitate subsequent calculations, the image quality is converted into three levels: high quality, normal quality, and low quality. Different levels correspond to different image quality coefficients.
[0057] S2: Evaluate the excitement of the feature data. The excitement evaluation specifically includes first calculating the value of each frame of feature data, then calculating the average value of all frame feature data values, and finally verifying the average value to obtain the excitement evaluation result. Specifically, the steps include:
[0058] S21: Calculate the value of each frame's feature data: S = (100% + A + B + C) * Z, where A is the total label coefficient of each frame, B is the human-pet interaction coefficient, C is the image quality coefficient, and Z is the confidence score. The confidence score is the pet's credibility. Here, third-party software can be used to identify each frame and obtain a value for the pet's credibility. All coefficients are percentages, and the specific percentage value of each coefficient can be determined according to actual conditions.
[0059] The total label coefficient, in this formula, only contains a single label:
[0060] A=A ’ , that is, the total label coefficient is equal to the label coefficient: (1) When there is a label with a pet identity label, A is 5%; if there is a label with a pet behavior label: if the label type is a high-quality behavior label, A is (5% * the credibility of the label); if the label type is a general behavior label, A is (3% * the credibility of the label; if the label type is an ordinary behavior label), A is (2% * the credibility of the label); if there is a label credibility, it is calculated. If there is no label credibility, it defaults to 100%; (2) If there is a label with a pet quantity label: if the number of pets is only one, then A 1%; if there are two pets, A is 2%; if there are more than two pets, A is 4%; (3) If there is a scenario tag: if the tag type is a high-quality scenario tag, A is (3% * the credibility of the tag); if the tag type is a general scenario tag, A is (2% * the credibility of the tag); if the tag type is an ordinary scenario tag, A is (1% * the credibility of the tag). If there is a tag credibility, it will be calculated. If there is no tag credibility, it will default to 100%. It is explained here that the tag credibility is generally defaulted to 100%;
[0061] In the case of multiple tags:
[0062] A=A1+A2+A3+……+Am , m is the total number of labels, that is, the total label coefficient is equal to the sum of the A values of the multiple labels contained;
[0063] Human-pet interaction coefficient. In this formula, if there is any human-pet interaction behavior, B is 15%;
[0064] Picture quality coefficient. In this formula, if it is high quality, C is 10%; if it is normal quality, C is 5%; if it is low quality, C is 0%;
[0065] Confidence score. In this formula, if there is only a single pet's confidence score in this frame, then Z = Z ‘ ; If there are multiple pet confidence scores in this frame, Z is the highest score among the multiple confidence scores;
[0066] S22: Calculate the average value of all frames and use this average value as the basic score to enter the subsequent steps: S ’ =(S1+S2+S3+……+S N ) / N, where N is the number of all frames;
[0067] S23: Verify the average value: when the average value is greater than or equal to the set threshold, it is considered to have passed and the subsequent steps are carried out; if it is less than the set threshold, it is considered to have failed and the subsequent steps are not carried out; the selection of the threshold can be determined according to the actual situation, for example, the threshold in this embodiment is 75; this step is mainly to further screen the videos that have passed the basic scoring and further improve the accuracy; further, this step also includes establishing a scoring countermeasure process, that is, when the verification is passed, the confidence score of the feature data of each frame is judged, and when the confidence score is less than the threshold, it is considered to have failed and the subsequent steps are not carried out; when the verification fails, the confidence score of the feature data of each frame is judged, and when the confidence score is greater than or equal to the threshold, it is considered to have passed and the subsequent steps are not carried out; a specific example is as follows: when the feature data of each frame is based on When the basic score is greater than or equal to 75, but the number of Z scores < 70 points accounts for more than 50%, the discard process is entered; when the basic score of each frame feature data is less than 75 points, but in each frame feature data, there are at least three Z scores ≥ 90 points that are not in the same 1 second, and the interval between each second is less than 4 seconds, it is still considered to have passed; because the establishment of the scoring countermeasure process can eliminate the errors caused by the basic evaluation, it can avoid the occurrence of misscreening or failure to consider special circumstances due to the basic evaluation, because the total label coefficient, human-pet interaction coefficient and picture quality coefficient of each frame before and after each frame will not fluctuate greatly, and the credibility of the pet may appear in the previous frame, not in the next frame, and then appear in a jump fluctuation, so the confidence score parameter is selected for scoring countermeasure judgment, which further increases the accuracy of the entire process;
[0068] S3: If the wonderfulness requirement is met, each frame that meets the wonderfulness requirement is automatically edited and synthesized to generate a wonderful pet video. The specific operations for generating a wonderful pet video are as follows:
[0069] S31: Cropping: Each frame is cropped at a double speed from the original image to the moment the pet appears, and the image after the pet appears remains unchanged; this step removes redundant, lengthy and boring video images, and there is no need to wait for a long time before the pet appears; this step is divided into two situations: a: If there is an execution instruction before the pet appears: the video start time is one second before the device starts to execute the instruction; the gap between the execution and the appearance of the pet is cropped at the original speed to the set time (set to 3s in this embodiment); the image after the pet appears continues to be retained without any change; b: If there is no execution instruction before the pet appears: the video is cropped at the original speed to the moment the pet appears; the image after the pet appears continues to be retained without any change, to avoid losing interest due to waiting time;
[0070] S32: Splicing: Splicing the front and back frames to generate a wonderful video with a complete story; when the interval between the front and back frames is less than or equal to the threshold, the images of the interval between the front and back frames are filled in during splicing; the interval between the front and back frames refers to the interval between the end time of the current video and the start time of the next video. If the interval is less than or equal to the threshold (set to 3s in this embodiment), all videos that meet the time interval ≤ 3s are spliced until the end time of the last video and the start time of the next video are > 3s, and the video content with the middle interval ≤ 3s is filled in during splicing to avoid the front and back connection being too abrupt and jumpy; when the interval between the front and back frames is greater than the threshold, the front and back frames are directly spliced; at the same time, the total duration of the spliced video is determined in this step: if the total duration does not exceed 60 seconds, a "wonderful video" is directly generated; if the total duration exceeds 60 seconds, the splicing is canceled and "wonderful videos" are generated separately; after splicing is completed, the video can also be edited, that is, the generated wonderful video is flowered and the text is directly reflected in the video;
[0071] S33: Generate cover: Select and display the cover of the video to improve the video conversion rate and attract user clicks; specifically, the steps include:
[0072] S331: Selecting a frame with the highest value in the video as the cover material;
[0073] S332: Pets are the center of the picture;
[0074] S333: Cutting out the portion of the video without the pet according to the aspect ratio of the original video;
[0075] S334: The cut-out image is used as the video cover;
[0076] S335: Determine the boundary value of the video cover: the cut-out image is enlarged to a maximum set value (the set value in this embodiment is 2).
[0077] The present invention extracts feature data from each frame of the video to ensure the comprehensiveness of the data and provide an accurate and comprehensive data basis for subsequent steps; then the feature data is evaluated for brilliance, and a three-step method is adopted when performing the brilliance evaluation: first, the value of the feature data of a single frame is calculated, then the average value of the feature data values of all frames is calculated, and finally the average value is verified. The whole process can effectively ensure the accuracy of the brilliance evaluation, avoid the problem of large errors and susceptibility to interference caused by a single evaluation, and improve the accuracy of the brilliance video judgment; then, the corresponding pet brilliance video is generated according to the brilliance evaluation result; the whole method process is simple, which can not only realize the accurate judgment of the pet brilliance video, thereby improving the happiness of the pet owner; but also realize rapid judgment without introducing complex steps or models for training or recognition, thereby improving work efficiency. At the same time, it is explained that the specific numerical values in this embodiment do not play a limiting role, and they can generally be determined according to actual usage and on-site aspects.
[0078] Example 2
[0079] Basically the same as Example 1, this embodiment also includes step S4: applying ratings to the generated wonderful pet videos, and generating different presentation methods for the wonderful pet videos based on different application rating results; making the presentation method more accurate and the push more reasonable, avoiding the poor user experience caused by the use of the same push method and the cost waste caused by the lack of rational application of resources, thereby further improving the pet owner's viewing experience of the wonderful pet videos. Specifically: The application rating formula is: F = a*b*c*S ’ ; Where F is the application score; S ’ is the basic score; a is the "homogeneity coefficient of the video within today and the recent time range"; b is the density; c is the recommendation coefficient; where:
[0080] For a: when the homogeneity is very low, the value of a is 1.5; when the homogeneity is normal, the value of a is 1; when the homogeneity is high, the value of a is 0.5. Here, it is explained that the homogeneity refers to the degree of similarity between the current video and other wonderful videos in the current day and the nearby time range. Generally, when it occurs less than twice, it is judged that the homogeneity is very low; when it occurs 3-5 times, it is judged that the homogeneity is normal; and when it occurs more than 5 times, it is judged that the homogeneity is high. This can be adjusted according to the actual situation.
[0081] For b: For single-pet and multi-pet families, different coefficients are used to determine the video density within a certain time range. The key point is that there are pets with low frequency of appearance in multi-pet families. When this pet appears, the impact of video density is reduced and the coefficient is increased, so that the video of this pet is better than the videos of other pets. (1) When there is no identity tag, single-pet family: when the number of wonderful videos within 1 hour is less than 1, it is judged that the video density is very low, and the coefficient is 1.5; when the number of wonderful videos within 1 hour is greater than 0 and less than or equal to 2, it is judged that the video density is normal, and the coefficient is 1; when the number of wonderful videos within 1 hour is greater than 2 and less than or equal to 4, it is judged that the video density is slightly high, and the coefficient is 0.7; when the number of wonderful videos within 1 hour is greater than 4, it is judged that the video density is very high. At this time, analyze and evaluate S first. ‘ Score: If S'≥90, the coefficient is 1.5; S ‘ ≥85, the coefficient is 1; otherwise, in other cases, the coefficient is 0.5; Multi-pet family: no matter how many wonderful videos there are in 1 hour, as long as there are multiple confidence levels, that is, multiple pets appear in the same video, the coefficient is 1.5; (2) When there is an identity tag, the judgment rules for single-pet families are the same as those without identity tags; Multi-pet families increase the number of pets that appear less frequently to appear in the screen, so that the video of this pet is better than the videos of other pets. When no matter how many wonderful videos there are in 1 hour and none of them belong to "this pet", the video of this pet that has never appeared has a coefficient of 1.5 (the specific value of the coefficient is only an example and can be adjusted according to the actual situation);
[0082] For c: the recommendation coefficient can be completed using a user recommendation algorithm. The user recommendation algorithm uses computer vision to analyze the operation data reported by the user when browsing the video, and obtains the final recommendation algorithm score. For the convenience of calculation, the recommendation algorithm here is divided into three levels: very recommended; generally recommended; not recommended; when very recommended, the recommendation coefficient is 1.5; when generally recommended, the recommendation coefficient is 1; when not recommended, the recommendation coefficient is 0; (the specific value of the coefficient is only an example and can be adjusted according to the actual situation); and the user recommendation algorithm is relatively mature in the existing technology, and the grading is just a routine operation according to different standards, so it will not be described in detail in this embodiment.
[0083] Furthermore, in this embodiment, the presentation method includes:
[0084] Basic presentation: only displays the generated pet videos. Here, when the app rating is between 60 and 115, the basic presentation is used.
[0085] Key presentation: Not only the generated pet video is displayed, but also the copy that matches the pet video, i.e., the scenario copy. Here, when the app score is 115-170 (excluding 115 and including 170), the key presentation is selected.
[0086] Focused presentation plus immediate push notification: Not only will the generated pet highlights video and the copy that matches the pet highlights video be displayed, but a push notification will also be sent immediately, i.e., the copy will be pushed. Here, when the app rating score is greater than 170, the focus presentation plus immediate push notification will be selected.
[0087] Do not present: Do not display the video, but save it. When the application rating is less than 60, the video will be included in the pet video list in the cloud service-loop recording without presentation, and the automatically edited and synthesized video will replace the fragment video of the same time for display. The "do not present" here only means that it will not be presented on the webpage of the wonderful video.
[0088] In this embodiment, two concepts, scenario-based copy and push copy, are introduced: scenario-based copy is the specific copy of the scenario-based label, and a scenario-based label has multiple scenario-based copy: the selection logic of scenario-based copy: if there is no video with the same scenario-based label on that day, a scenario-based copy is randomly selected and displayed together with the video in the APP to highlight the key points; if there is a video with the same scenario-based label on that day, the scenario-based copy of the adjacent videos avoids the same copy, and one of the other two is selected and displayed together with the video in the APP to highlight the key points; this embodiment takes the following examples: (1) Waiting for dinner: the corresponding scenario-based copy is: My baby is hungry, can we have dinner earlier? I digested a little fast today, and my baby is already hungry~; Why is it not time for dinner yet? Baby is hungry; (2) Eat a midnight snack: The corresponding scenario copy is: I am hungry, let’s have some midnight snack~; Baby is hungry after playing, the baby needs to replenish his energy; Eating gently should not disturb his sleep~; (3) Full: The corresponding scenario copy is: I am full, I am comfortable~; Baby, pat your butt and leave after eating; Eat small meals frequently, and come back later; The generation logic of scenario copy can refer to the following example: Scenario copy: "Mom, I am hungry, can we have dinner earlier~"; The corresponding generation logic is: The pet is recognized to appear and stay within 15 minutes of the automatic food delivery. The generation of copy depends on the pet's behavior and surrounding components.
[0089] Push notifications are sent immediately to the wonderful videos that have been highlighted and have a particularly high score in the application rating system, so that users can watch them at the first time; therefore, the push text is the specific push copy of the scenario label. The push copy is different from the scenario copy. One scenario text has multiple push copies: the selection logic of the push copy: if there is no video with the same scenario label on that day, a push copy is randomly selected for push notification; if there is a video with the same scenario label on that day, the push copy of the adjacent videos avoids the same copy, and one of the other two is selected for push notification; this embodiment is through The following examples are provided for easier understanding: if the scenario label is "waiting for dinner," the push copy could be: "Baby is hungry, hurry up and feed him"; "Baby is already waiting for dinner, hurry up"; "Hey, baby is hungry again, go see what he's been doing today"; if the scenario label is "eating a midnight snack," the push copy could be: "Finally I know what he's doing up late at night"; "It turns out he's the one making the noise at night"; "Baby is hungry after a night of patrolling, time for a midnight snack to replenish his energy"; if the scenario label is "full," the push copy could be: "Report, baby just ate 14g"; "This is the fourth meal of the day"; "Full and ready to go, probably taking a nap after dinner," etc. Of course, the above examples are not intended to be limiting, but are merely for a better understanding of the differences between recommendation copy, scenario-based copy, and scenario labeling.
[0090] The present invention also includes applying a rating to a wonderful pet video after it is generated, so that it can be presented in different ways according to different rating results, making its presentation more accurate and its push more reasonable, avoiding the poor experience brought about by the use of the same push method and the cost waste caused by the unreasonable application of resources, thereby further improving the pet owner's viewing experience of the wonderful pet video.
[0091] Example 2
[0092] A system using the method for generating wonderful pet videos as described in Example 1 above, comprising:
[0093] Data acquisition unit: used to obtain feature data of each frame in the video;
[0094] Wonderfulness evaluation unit: used to evaluate the wonderfulness of feature data; in this unit, the value of each frame of feature data is first calculated, then the average value of all frame feature data values is calculated, and finally the average value is verified to obtain the wonderfulness evaluation result;
[0095] Display unit: used to generate and display wonderful pet videos.
[0096] The system unit of the present invention has a simple composition and stable operation, and can accurately judge the wonderful videos of pets, thereby improving the happiness of pet owners.
[0097] Example 3
[0098] A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements a method for generating a wonderful pet video as described in the above-mentioned embodiment 1. It is worth noting that for those skilled in the art, in addition to implementing the system and various modules provided by this application in a purely computer-readable program code manner, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, editable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system provided by this application and its various modules can be considered as a hardware component, and the modules included therein for implementing various programs can also be regarded as structures within the hardware component, and the modules for implementing various functions can also be regarded as both software programs for implementing the method and structures within the hardware component.
[0099] The examples described in the present invention are merely descriptions of the preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made to the technical solutions of the present invention by engineers and technicians in this field should fall within the scope of protection of the present invention.
Claims
1. A method for generating wonderful pet videos, characterized by: The following steps are involved: S1: Obtain feature data of each frame in the video; S2: Evaluate the wonderfulness of the feature data. The wonderfulness evaluation specifically includes first calculating the value of the feature data of each frame, then calculating the average value of the feature data of all frames, and finally verifying the average value to obtain the wonderfulness evaluation result. Specifically, when the average value is greater than or equal to the set threshold, it is considered to have passed and the subsequent steps are carried out; if it is less than the set threshold, it is considered to have failed and the subsequent steps are not entered. This step also includes establishing a scoring countermeasure process, that is, when the verification is passed, the confidence score of the feature data of each frame is judged. When the confidence score is less than the threshold, it is considered to have failed and the subsequent steps are not entered. When the verification fails, the confidence score of the feature data of each frame is judged. When the confidence score is greater than or equal to the threshold, it is considered to have passed and the subsequent steps are entered. S3: If the requirements for excitement are met, each frame meeting the requirements for excitement is automatically edited and synthesized to generate an exciting pet video.
2. The method for generating a wonderful pet video according to claim 1, wherein: The step S2 specifically includes the following steps: S21: Calculate the value of the feature data of each frame: S = (100% + A + B + C) * Z, where A is the total label coefficient of each frame, B is the human-pet interaction coefficient, C is the image quality coefficient, and Z is the confidence score; S22: Calculate the average value of all frames: S ’ =(S1+S2+S3+……+S N ) / N, where N is the number of all frames; S23: Verify the average value: When the average value is greater than or equal to the set threshold, it is considered passed and proceeds to the subsequent steps; if it is less than the set threshold, it is considered failed and does not proceed to the subsequent steps.
3. The method for generating a wonderful pet video according to claim 1, wherein: The step S1 includes pre-analyzing each frame to obtain feature data in the picture: total label coefficient, human-pet interaction coefficient, picture quality coefficient and confidence; at the same time, the pre-analysis also includes determining whether there is misidentification and / or non-identification phenomenon in the picture.
4. The method for generating a wonderful pet video according to claim 1, wherein: The specific operations of generating the wonderful pet video in step S3 are as follows: S31: Cropping: Cropping is performed at double speed from the original image to the moment the pet appears in each frame, and the image after the pet appears remains unchanged; S32: Splicing: Splicing the previous and next frames. When the interval between the previous and next frames is less than or equal to a threshold, the images are filled in to fill the gap between the previous and next frames. When the interval between the previous and next frames is greater than the threshold, the previous and next frames are directly spliced. S33: Generate cover: Select and display the cover of the video.
5. The method for generating a wonderful pet video according to claim 4, wherein: The step S33 specifically includes the following steps: S331: Selecting a frame with the highest value in the video as the cover material; S332: Pets are the center of the picture; S333: Cutting out the portion of the video without the pet according to the aspect ratio of the original video; S334: The cut-out image is used as the video cover; S335: Determine the boundary value of the video cover.
6. The method for generating a wonderful pet video according to claim 1, wherein: The method further includes step S4: performing application rating on the generated wonderful pet videos, and generating different presentation modes of the wonderful pet videos according to different application rating results.
7. The method for generating a wonderful pet video according to claim 6, wherein: The presenting methods include: basic presenting: only displaying the generated wonderful pet videos; Key presentation: Display not only the generated pet videos but also the copywriting that matches them; Highlight presentation plus immediate push notification: Not only will the generated pet highlights video and the copywriting that matches the pet highlights video be displayed, but push notification will also be sent immediately; Do not present: Do not display the video, but save it.
8. A system using the method for generating wonderful pet videos according to any one of claims 1 to 7, characterized in that: include: Data acquisition unit: used to obtain feature data of each frame in the video; Wonderfulness evaluation unit: used to evaluate the wonderfulness of feature data; Display unit: used to generate and display wonderful pet videos.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a wonderful pet video according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Automatic generation method of video selection set based on content analysis
CN112445935A
Pet video recording and automatic editing method based on AI algorithm
CN113784072A
Video data processing method and device and readable storage medium
CN109587554A
Automatic video editing method and device and portable terminal
CN109819338A