Target video capture method, device, equipment and medium based on image screening
By obtaining the request instruction to look back at the target event, determining the basic video associated with the target event, and synthesizing the target video according to the preset rules, the problem of difficult to find video clips in the prior art is solved, and saving user time and device storage space are achieved.
Patent Information
- Application Number
- CN202310208827.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-06-02
AI Technical Summary
In the prior art, video clips when a specific event occurs are difficult to find, causing inconvenience to users.
By obtaining a request instruction to look back at the target event, at least one continuous basic video associated with the target event is determined, and the video data associated with the target event in the base video is synthesized according to the preset video synthesis rules to generate the target video.
This enables users to find the event-related videos that need to be viewed in a large number of recorded basic videos, saving user time, and target videos are only generated when users need to view, saving device storage space.
Smart Images

Figure CN116208821B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on June 2, 2021, with the invention name “Video playback method, device, electronic device and medium” and application number 202110612998.0. Technical Field
[0002] The present invention relates to the field of video surveillance technology, and in particular to a method, device, equipment and medium for capturing a target video based on image screening. Background Art
[0003] Video surveillance is the physical basis for real-time monitoring of key departments or important places in various industries. Management departments can obtain effective data, images or sound information through it, monitor and remember the process of sudden abnormal events in a timely manner, and provide efficient and timely command and processing. Video surveillance is an important part of the security system. The traditional monitoring system includes front-end cameras, transmission cables, and video monitoring platforms. Cameras can be divided into network digital cameras and analog cameras, which can be used as the collection of front-end video image signals. It is a comprehensive system with strong prevention capabilities. Video surveillance is widely used in many occasions because of its intuitive, accurate, timely and rich information content.
[0004] In recent years, with the rapid development of computers, networks, image processing and transmission technologies, video surveillance technology has also made great progress. Current video surveillance equipment usually records video continuously. Continuous recording makes the video files larger. Large video files store massive amounts of video data. The resulting technical problem is that video clips of specific events are difficult to find, which brings inconvenience to users. Summary of the invention
[0005] In view of this, the embodiments of the present invention provide a method, device, equipment and medium for capturing target videos based on image screening, so as to solve the technical problem in the prior art that video clips when specific events occur are difficult to find, causing inconvenience to users.
[0006] The technical solution adopted by the present invention is:
[0007] The present invention provides a target video capture method based on image screening, the method comprising:
[0008] S1: Obtain a request instruction for replaying a target video corresponding to a target event;
[0009] S2: determining, according to the request instruction, at least one continuous basic video associated with the target event;
[0010] S3: According to a preset video synthesis rule, the video data associated with the target event in the basic video is synthesized to generate the target video, and the target video corresponding to the target event is output.
[0011] Preferably, S2 includes:
[0012] S21: Obtaining the tag information of the target event and the duration corresponding to the target video;
[0013] S22: Determine the first basic video where the target event is located according to the time information of the tag information;
[0014] S23: determining at least one continuous basic video associated with the target event according to the duration corresponding to the target video and the position of the target event in the first basic video;
[0015] The duration of the target video is less than or equal to the duration of the basic video.
[0016] Preferably, the S23 includes:
[0017] S231: Divide the first basic video into a plurality of video segments according to the duration of the target video;
[0018] S232: Determine, according to the time information of the target event, a target video segment of the first basic video to which the target event belongs;
[0019] S233: Determine at least one continuous basic video associated with the target event according to the position information of the target video segment compared to the first basic video and the duration corresponding to the target video.
[0020] Preferably, the S231 includes:
[0021] S2311: Obtain a target duration corresponding to half the duration of the target video;
[0022] S2312: Segment the first basic video according to the target duration to obtain the multiple video segments.
[0023] Preferably, S3 includes:
[0024] S31: Obtaining a preset video synthesis rule of the target video and a duration corresponding to the target video;
[0025] S32: extracting each frame image of the target video data associated with the target event in each of the basic videos according to the preset video synthesis rule, and synthesizing the target video;
[0026] The total duration of each frame image of the target video data is equal to the duration corresponding to the target video.
[0027] Preferably, the step S1 includes:
[0028] S01: Obtain the duration threshold corresponding to generating the target event;
[0029] S02: timing the special events in the target area and generating a reference duration corresponding to the timing duration;
[0030] S03: When the reference duration meets the duration threshold requirement, a key frame image corresponding to the special event within the reference duration is used as an index image of the target event corresponding to the special event.
[0031] Preferably, the step S03 includes:
[0032] S04: Acquire all index images within a specified time period;
[0033] S05: Identify each index image and obtain the confidence of each index image;
[0034] S06: When the confidence of the index image meets the confidence threshold requirement, extract the index image meeting the confidence threshold requirement and create an index image list.
[0035] Preferably, the S06 includes:
[0036] S061: Acquire historical video viewing data of the user, determine event content in the target video that the user is interested in, and determine historical key frame images in the event content;
[0037] S062: assigning a weight value to the similarity between the key frame image of the special event and the historical key frame image;
[0038] S063: Sort each index image in the index image list according to the confidence level and the weight value of each index image.
[0039] The present invention also provides a device, comprising:
[0040] Instruction acquisition module: used to obtain the request instruction for viewing the target event;
[0041] A basic video positioning module: used to determine at least one continuous basic video associated with the target event according to the request instruction;
[0042] Target video synthesis module: used to synthesize the video data associated with the target event in the basic video according to a preset video synthesis rule, and output a target video corresponding to the target event.
[0043] The present invention also provides an electronic device, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, and when the computer program instructions are executed by the processor, any of the above-mentioned methods is implemented.
[0044] The present invention also provides a medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, any of the above-mentioned methods is implemented.
[0045] In summary, the beneficial effects of the present invention are as follows:
[0046] The present invention provides a method, device, equipment and medium for capturing target videos based on image screening, which obtain a request instruction for reviewing a target video corresponding to a target event; according to the request instruction, determine at least one continuous basic video associated with the target event; according to a preset video synthesis rule, synthesize the video data associated with the target event in the basic video to generate the target video, and output the target video corresponding to the target event. The user issues a request instruction for video review, determines the basic video associated with the target event through the request instruction, and then synthesizes the target video to be viewed through the basic video. On the one hand, the user does not need to search for the event-related video that needs to be viewed in a large number of recorded basic videos, saving the user's precious time. On the other hand, the target video is only viewed when the user needs to view it, and according to the preset video synthesis rule, the video data associated with the target event in the basic video is synthesized to generate the target video under the user's request instruction. The target video is not stored in advance, which can save the storage space of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the drawings required for use in the embodiment of the present invention will be briefly introduced below. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work, and these are all within the protection scope of the present invention.
[0048] Figure 1 It is a flowchart of a target video capture method based on image screening in Example 1 of Implementation Mode 1 of the present invention;
[0049] Figure 2 This is a schematic diagram of a process of determining a basic video in Example 1 of Implementation Mode 1 of the present invention;
[0050] Figure 3 This is a schematic diagram of a process of determining a basic video according to duration and position in Example 1 of Implementation Mode 1 of the present invention;
[0051] Figure 4 It is a schematic diagram of a process of segmenting a basic video in Example 1 of the first embodiment of the present invention;
[0052] Figure 5 Schematic diagram of the process of synthesizing a target video in Example 1 of Implementation Mode 1 of the present invention;
[0053] Figure 6 This is a schematic diagram of a process of generating an index image in Example 1 of Implementation Mode 1 of the present invention;
[0054] Figure 7 It is a schematic diagram of the process of establishing an index image list in Example 1 of Implementation Mode 1 of the present invention;
[0055] Figure 8 It is a schematic diagram of a process of sorting index images in an index image list in Example 1 of Embodiment 1 of the present invention;
[0056] Fig. 9 This is a schematic diagram of a process of storing real-time video in Example 1 of Implementation Mode 1 of the present invention;
[0057] Fig.10 Schematic diagram of the process of video extraction and splicing method in Example 2 of Implementation Mode 1 of the present invention;
[0058] Fig.11 It is a structural block diagram of the device in Example 3 of Implementation Mode 2 of the present invention;
[0059] Fig.12 It is a structural block diagram of a video extraction and splicing device in Example 4 of Implementation Mode 2 of the present invention;
[0060] Fig.13 It is a schematic diagram of the structure of an electronic device in the third embodiment of the present invention. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described in conjunction with the drawings in the embodiment of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. In the absence of further restrictions, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements. If there is no conflict, the various features in the embodiments and examples of the present invention can be combined with each other, all within the scope of protection of the present invention.
[0062] Implementation Method 1
[0063] Example 1
[0064] See also Figure 1 , Figure 1 : is a flow chart of a target video capture method based on image screening in Embodiment 1 of the present invention. The target video capture method based on image screening in Embodiment 1 of the present invention comprises:
[0065] S1: Obtain a request instruction for replaying a target video corresponding to a target event;
[0066] Specifically, the user can conveniently issue a request command to review the target video corresponding to the target event by clicking and pressing the APP on the mobile terminal. The target event includes some events generated by the baby's daily activities, such as events related to the baby's crying, clapping, and rolling in bed. The video monitoring equipment records the target event and generates an index tag associated with the target event on the App side, such as the video playback interface on the App side, which displays the image corresponding to the target event.
[0067] S2: determining, according to the request instruction, at least one continuous basic video associated with the target event;
[0068] Specifically, the target event is recorded in chronological order, and at least one continuous basic video associated with the target event is determined based on the time information of the target event or the identifiable action information of the target event. The real-time video recorded by the video surveillance device according to time is stored as the basic video. It should be noted that, in determining at least one continuous basic video associated with the target event, a comprehensive consideration can be given to the start time of the target event, the duration of the target video and the preset video synthesis rules. For example, each basic video is recorded with a duration of one minute. When the start time of the target event is at the tenth second of the fifth minute, if the duration of the target video is thirty seconds, if thirty seconds of video data is taken backward from the tenth second of the fifth minute, that is, from the tenth second of the fifth minute to the fortieth second of the fifth minute, this section of video data is all included in the fifth section of the basic video. Therefore, in this case, only the fifth section of the basic video is extracted; however, if the duration of the target video is also thirty seconds, if fifteen seconds of video data are taken forward and fifteen seconds of video data are taken backward from the tenth second of the fifth minute, the situation is different from the above situation because fifteen seconds of video data are taken forward from the tenth second of the fifth minute, and five seconds of video data are not in the fifth basic video, but in the last five seconds of the fourth basic video. In order to synthesize the target video, two continuous basic videos, the fourth basic video and the fifth basic video, need to be taken. Furthermore, when the target event starts at the fiftieth second of the fifth minute, if thirty seconds of video data are taken from the fiftieth second of the fifth minute backward, two consecutive basic videos, namely the fifth basic video and the sixth basic video, need to be taken.
[0069] S3: According to a preset video synthesis rule, the video data associated with the target event in the basic video is synthesized to generate the target video, and the target video corresponding to the target event is output.
[0070] Specifically, after locating the video data associated with the target event, the video data is cropped and spliced according to preset video synthesis rules to synthesize the target video, which is then output to the user terminal so that the user can view the event video he wants to watch.
[0071] Specifically, the user clicks and presses the index tag of the target event in the video playback interface of the App to trigger a request instruction for playback of the target video corresponding to the target event. The basic video associated with the target event is determined through the request instruction, and the target video to be viewed is synthesized through the basic video. On the one hand, this saves the user from having to search for the event-related videos they need to view among a large number of recorded basic videos, saving the user's valuable time. On the other hand, the target video is only viewed when the user needs to, under the user's request instruction, and according to preset video synthesis rules, the video data associated with the target event in the basic video is synthesized to generate the target video. The target video is not stored in advance, which can save device storage space.
[0072] In one embodiment, see Figure 2 , said S2 comprises:
[0073] S21: Obtaining the tag information of the target event and the duration corresponding to the target video;
[0074] Specifically, when the target event occurs, the corresponding basic video will be marked accordingly. The corresponding marking operation is to generate an event tag that records more information and generate tag information corresponding to the target event. Specifically, the duration of the target video to be generated is less than or equal to the duration of the basic video. The significance of making such a provision is that the user's time and energy are relatively limited, so the duration of the target video should not be too long, and if the duration of the target video is longer than the basic video, the corresponding basic video can be directly extracted, and there is no need to involve video splicing. Therefore, the duration of the target video is preferably less than or equal to the duration of the basic video.
[0075] S22: Determine the first basic video where the target event is located according to the time information of the tag information;
[0076] Specifically, since the tag information of the event tag contains more information, including time information, and since the basic videos are recorded continuously in chronological order, the specific basic video corresponding to the target event can be determined through the time information contained in the event tag information.
[0077] S23: determining at least one of the basic videos associated with the target event according to the duration corresponding to the target video and the position of the target event in the first basic video;
[0078] The duration of the target video is less than or equal to the duration of the basic video.
[0079] Specifically, there may be more than one basic video associated with the target event. For example, if the recording time of each basic video is one minute, and the duration of the target video is set to 30 seconds, and the event starts at the 50th second of a basic video, then its duration of 30 seconds spans the first basic video and the second basic video after the first basic video, specifically from the 50th second of the first basic video to the first 20 seconds of the second basic video. Therefore, when determining the basic video associated with the target event, there is more than one basic video associated with the target event. At this time, the appropriate basic video is selected according to the progress of the target event to be viewed, and the corresponding video data is extracted and spliced to present the target video closely related to the target event to the user.
[0080] See also Figure 3 , the S23 comprises:
[0081] S231: Divide the first basic video into a plurality of video segments according to the duration of the target video;
[0082] Specifically, in order to obtain video data that is more consistent with the target event and make the generated target video closer to the actual needs of the user, in this solution, the first basic video is divided into multiple video segments, and the search is gradually approached within the multiple video segments to obtain more appropriate video data.
[0083] S232: Determine, according to the time information of the target event, a target video segment of the first basic video to which the target event belongs;
[0084] Specifically, in the actual operation process, only one basic video needs to be taken or two adjacent continuous basic videos need to be taken. These two situations are slightly complicated to judge. In order to make the related judgment process faster and simpler, in this embodiment, the start time of the target event is first used to locate a corresponding basic video, and then the basic video is reasonably segmented. The number of basic videos required for quick judgment is achieved according to the position of the start time of the target event in the segmented video. For example, the recording time of each basic video is one minute, and the corresponding duration of the set target video is 30 seconds. When the starting time of the target event occurs at the twenty-fifth second of the fifth minute, the fifth basic video is positioned, and the fifth basic video is segmented every fifteen seconds, which is divided into four segmented videos, namely, the fifth minute to the fourteenth second of the fifth minute; the fifteenth second of the fifth minute to the twenty-ninth second of the fifth minute; the thirtieth second of the fifth minute to the forty-fourth second of the fifth minute; and the forty-fifth second of the fifth minute to the fiftieth second of the fifth minute. If the method of taking fifteen seconds forward and backward from the starting time of the target event is adopted, then when the starting time of the target event is located in the first segmented video or the fourth segmented video, it is necessary to take the two consecutive basic videos before and after, and when the starting time of the target event is located in the second segmented video or the third segmented video, only one basic video at the current position is needed.
[0085] S233: Determine at least one continuous basic video associated with the target event according to the position information of the target video segment compared to the first basic video and the duration corresponding to the target video.
[0086] Specifically, according to the position of the target video segment in the first basic video and the duration of the target video, it can be inferred whether the target video can be completely extracted in the existing first basic video, or whether it needs to be extracted in an adjacent basic video. According to the position information and duration, it can be flexibly determined in the actual operation process, and it is not too limited here. If you need to have a complete understanding of the process before and after the target event, you can extract and splice the two basic videos before and after in chronological order. If the time span of the target event is large, you can also extract and splice more video data to give users a better video playback experience.
[0087] In one embodiment, see Figure 4 , the S231 includes:
[0088] S2311: Obtain a target duration corresponding to half the duration of the target video;
[0089] S2312: Segment the first basic video according to the target duration to obtain the multiple video segments.
[0090] Specifically, by segmenting the first basic video according to the target duration, each video segment will not be too short, and the segmentation of the basic video will not be too scattered while ensuring that the first basic video is segmented reasonably, which helps to improve the efficiency of cutting and splicing and reduce the pressure of data processing. In the process of generating the target video, two video segments associated with the target event are selected, and the target video can be obtained by extracting and splicing these two video segments, making the generation of the target video simpler and faster. In addition, by the position of the starting time of the target event in a specific basic video, that is, the segmented video where the starting time of the target event is located, it is quickly determined whether a basic video needs to be extracted or two adjacent basic videos need to be extracted.
[0091] In one embodiment, see Figure 5 , said S3 comprises:
[0092] S31: Obtaining a preset video synthesis rule of the target video and a duration corresponding to the target video;
[0093] S32: extracting each frame image of the target video data associated with the target event in each of the basic videos according to the preset video synthesis rule, and synthesizing the target video;
[0094] The total duration of each frame image of the target video data is equal to the duration corresponding to the target video.
[0095] Specifically, the target video data is extracted according to continuous video frames, and the target video data is processed and the target video is generated in combination with the set synthesis rule of the target video and the duration of the target video. As for the synthesis rule of the target video, it can be determined in combination with the difficulty of video processing, the computing power of the processor, the actual needs of the user, etc. For example, if the video data processing is more difficult and the computing power of the processor is small, continuous video frames can be intercepted and spliced to generate the target video faster, avoiding more system resource consumption.
[0096] To this end, in one embodiment, see Figure 6 , said S1 before includes:
[0097] S01: Obtain the duration threshold corresponding to generating the target event;
[0098] S02: timing the special events in the target area and generating a reference duration corresponding to the timing duration;
[0099] S03: When the reference duration meets the duration threshold requirement, a key frame image corresponding to the special event within the reference duration is used as an index image of the target event corresponding to the special event.
[0100] Specifically, in order to avoid overly sensitive generation of target videos and avoid troubles caused to users by false alarms, in this solution, by recording the reference duration of the target event, the index image of the target event is generated only when the reference duration meets a certain duration threshold requirement, so as to facilitate subsequent user retrieval. In addition, the video data corresponding to the reference duration can be directly used as the video data of the target video, or the corresponding video data can be obtained from the reference duration forward or backward, and combined with the video data within the reference duration to generate the target video, so as to meet the video playback needs of different users. As mentioned above, the target events include some events generated by the daily activities of infants, such as events related to the crying, clapping, and rolling of infants in bed. More specifically, special events are events that are further screened based on the target events, which are events related to the crying of infants. Since the crying of babies is an event that users are more concerned about, and it involves issues such as the baby's adaptability to the environment and sense of security, in order to better protect the baby and make the baby more adaptable to the environment and feel more secure, users usually pay more attention to special events. Therefore, in this embodiment, a frame image of a special event is extracted as the index image of the target event corresponding to the special event, which can help users better review the moment of the baby's crying, help users analyze the reasons for the baby's crying, and make targeted improvements to achieve scientific parenting.
[0101] For further explanation, please see Figure 7 , the S03 then includes:
[0102] S04: Acquire all index images within a specified time period;
[0103] S05: identifying each of the index images, and obtaining the confidence level of each of the index images;
[0104] S06: When the confidence of the index image meets the confidence threshold requirement, extract the index image meeting the confidence threshold requirement and create an index image list.
[0105] Specifically, all index images include multiple index images of different categories, such as a category of index images of babies crying, or a category of index images of babies crawling and playing; a specified time period is 24 hours a day, and many special events may occur within this time range. Special events are events that are further screened based on target events, and are similar events, such as events related to babies crying, or events related to babies crawling and playing on the bed, etc. By comparing each index image with the images in the database, the different judgment methods for index images will lead to a certain deviation in the recognition of index images. According to the deviation, the confidence of each index image can be calculated accordingly. When the confidence of the index image meets the confidence threshold requirement, the index image is extracted, so as to more accurately locate the relevant basic video according to the index image, so that the correlation between the index image and the target video is higher. For example, all the acquired index images are compared with the images in the database, and the similarity value between each index image and the images in the database is calculated as the confidence. For example, there is an index image, and its confidence is 0.6 after comparison with the images in the remaining database. The confidence threshold is set to 0.8. Since 0.6 is lower than 0.8, the confidence of the index image does not meet the requirements of the confidence threshold, and this index image is discarded; if the confidence of another index image is 0.9, since 0.9 is higher than 0.8, it is determined that the confidence of the index image meets the requirements of the confidence threshold, and the index image is extracted, and all index images that meet the confidence threshold requirements are extracted and an index image list is established.
[0106] For further information, see Figure 8 , the S06 comprises:
[0107] S061: Acquire historical video viewing data of the user, determine event content in the target video that the user is interested in, and determine historical key frame images in the event content;
[0108] Specifically, the event content of the target video that the user is interested in can be determined through the user's historical video viewing data. For example, if the user clicks on a certain type of target video a large number of times and the total playback time is long, it will be determined as the target video that the user is interested in. The event content in the target video that the user is interested in is further determined, and the key frames therein are extracted as historical key frame images to facilitate subsequent identification of whether special events are of interest to the user. Specifically, the key frame image of the special event can be compared with the historical key frame image to determine the similarity between the two.
[0109] S062: assigning a weight value to the similarity between the key frame image of the special event and the historical key frame image;
[0110] S063: Sort each index image in the index image list according to the confidence level and the weight value of each index image.
[0111] Specifically, if the key frame image of a special event is 80% similar to the historical key frame image, a weight value of 0.8 is assigned to it; if the key frame image of a special event is 100% similar to the historical key frame image, a weight value of 1 is assigned to it. According to the weight value and the confidence of the index image, the index images are sorted. Specifically, if an index image has a confidence of 0.6 and a weight of 0.8, the confidence is multiplied by the weight to obtain 0.48; if another index image has a confidence of 0.9 and a weight of 1, the confidence is multiplied by the weight to obtain 0.9, and in the index image list, the latter index image is arranged in front of the former index image. In this embodiment, in order to ensure that the index image that the user can view is more closely associated with the target video, the confidence of the index image is obtained, and each type of index image is sorted according to the confidence. For example, in the index images of babies crying, the index images with higher confidence can be arranged at the front position in the index image list, so that users can quickly click to the most relevant video; or the index images with higher confidence can be arranged at the back position in the index image list. Regardless of whether it is at the back or the front, it is necessary to make the index images with higher confidence easier to find in the index image list, so that users can quickly find the closely related basic videos, and then synthesize more relevant target videos through the basic videos. In addition, with regard to the index images of babies crawling and playing, a similar approach can be adopted, so that the index images with higher confidence can be easily found in the index column, which can also improve the problem of poor user experience caused by inaccurate identification and comparison information of index images. By placing some of the index images with higher confidence in each category in a prominent position in the index image list, users can generate more relevant target videos by clicking on the index images, so that the index images in the prominent position have a higher matching degree with each type of event, thereby providing a better user experience and a better viewing experience. Moreover, combined with the user's historical video viewing data, it not only ensures that the index image is closely related to the target video, but also effectively filters out videos that the user is more interested in.
[0112] Specifically, since the user may not finish watching every target video in the process of watching it, each historical target video has a corresponding viewing progress and number of clicks. The user's interest in a certain historical video is determined based on the user's number of clicks. Then, based on the user's viewing progress of a certain historical video, historical videos with more user clicks but lagging viewing progress are screened out as videos to be watched. The videos to be watched are marked so that they can be easily identified by users, thereby facilitating users to review events. The index information of the videos to be watched can also be pushed to the user end to remind users to check them in time.
[0113] For further explanation, please see Fig. 9 , said S1 before includes:
[0114] S01': Get video storage rules;
[0115] S02': storing the real-time video according to the video storage rule to obtain a plurality of the basic videos;
[0116] The storage location of the basic video includes a video acquisition terminal and / or a mobile terminal.
[0117] Specifically, according to the storage needs of video data, corresponding video storage rules are adopted. The real-time video collected by the video surveillance probe can be reasonably stored using this video storage rule to obtain multiple basic videos, making the video data storage more convenient and reliable, and also convenient for subsequent search and use.
[0118] In summary, the target video capture method based on image screening provided in this embodiment obtains a request instruction for viewing the target video corresponding to the target event; according to the request instruction, at least one basic video associated with the target event is determined; according to a preset video synthesis rule, the video data associated with the target event in the basic video is synthesized to generate the target video, and the target video corresponding to the target event is output. The user issues a request instruction for video playback, and the basic video associated with the target event is determined through the request instruction, and then the target video to be viewed is synthesized through the basic video. On the one hand, the user does not need to find the event-related video that needs to be viewed in a large number of recorded basic videos, saving the user's precious time. On the other hand, the target video is only when the user needs to view it, under the user's request instruction, according to the preset video synthesis rule, the video data associated with the target event in the basic video is synthesized to generate the target video. The target video is not pre-stored, which can save the storage space of the device. In short, there are many ways to generate the target video. Without departing from the substantive content of the present invention, there can be a variety of technical solutions for extraction methods. It should be noted that simple changes to the technical solutions in the application document do not require creative work for those skilled in the art, and therefore fall within the scope of protection of the present invention and will not be repeated here.
[0119] Example 2
[0120] In Example 1, a target video is obtained by extracting a specific video segment and performing corresponding video processing operations. When the duration of the target event exceeds the length of the segmented video, multiple basic videos need to be retrieved and processed to generate the target video. This not only causes the extracted data to be larger and inconvenient, but also makes the extracted video longer, which is not focused enough on the event review, and ultimately affects the user experience. In some embodiments, a method of extracting, splitting and splicing video data is adopted, and the video obtained in this way is more suitable as a target video and is more targeted. Therefore, Example 2 of the present invention proposes a video extraction and splicing method, please refer to Fig.10 , the method comprising:
[0121] S400: Demultiplexing the basic video data into raw stream data;
[0122] Specifically, the recorded basic video is in MP4 format, and the video data in MP4 format is demultiplexed into raw stream data H.264 or H.265 through technical means. H.264 is a digital video compression format jointly proposed by the International Organization for Standardization and the International Telecommunication Union. This is an existing technology and will not be introduced in detail here.
[0123] S410: Mark the raw stream data according to the time sequence and determine the video frame number;
[0124] Specifically, marking the raw stream data according to the time sequence involves the problem of video frame rate, which is the frequency of continuous appearance of images in frame units. In this solution, the frame rate of the recorded basic video is 20 frames per second, and there are 20 frames of images every second, that is, at the first second, its frame number is the 20th frame, and at the second second, its frame number is 20×2=40 frames. It can be seen that the frame number = the number of seconds × frame rate, so the corresponding relationship between the frame number and time of the raw stream data is determined without any doubt.
[0125] S420: splicing the raw stream data associated with the target event according to the video frame number to determine target video data.
[0126] Specifically, by selecting a specific starting time point and a predetermined time, the corresponding frame number can be extracted to obtain the corresponding target data. The target data can be stored locally and wait for the user to extract and view it, or it can be pushed to the user's mobile phone, tablet and other terminal devices through technical means so that the user can view it in time. More specifically, after the target data is collected, the target data can be reassembled into MP4 format data and stored locally, or the MP4 format data can be sent to the user-end device for the user to view. The MP4 format data sent to the user-end device is arranged in chronological order, and the target videos of multiple time nodes will be displayed. The user can choose to watch them. This can be flexibly determined according to the actual situation, and no excessive restrictions are made here.
[0127] By adopting the video extraction and splicing method of the present embodiment, the basic video data is demultiplexed into raw stream data; the raw stream data is marked according to the time sequence to determine the video frame number; the raw stream data associated with the target event is spliced according to the video frame number to determine the target data. The video data can be processed more flexibly. For reviewing the event, the video data at a specific moment can be extracted and flexibly cut and spliced. According to the development of the event, the corresponding video data can be extracted and spliced, so that a video related to the event that satisfies the user can be obtained.
[0128] Implementation Method 2
[0129] Example 3
[0130] The embodiment of the present invention also provides a device, such as Fig.11 As shown, including:
[0131] Instruction acquisition module: used to obtain the request instruction for viewing the target event;
[0132] A basic video positioning module: used to determine at least one continuous basic video associated with the target event according to the request instruction;
[0133] Target video synthesis module: used to synthesize the video data associated with the target event in the basic video according to a preset video synthesis rule, and output a target video corresponding to the target event.
[0134] Using the device of this embodiment, a request instruction for reviewing a target video corresponding to a target event is obtained; according to the request instruction, at least one basic video associated with the target event is determined;
[0135] According to the preset video synthesis rules, the video data associated with the target event in the basic video is synthesized to generate the target video, and the target video corresponding to the target event is output. The user issues a request instruction for video playback, and the basic video associated with the target event is determined through the request instruction, and then the target video to be viewed is synthesized through the basic video. On the one hand, this makes it unnecessary for the user to search for the event-related video that needs to be viewed in a large number of recorded basic videos, saving the user's precious time. On the other hand, the target video is only viewed when the user needs to view it. According to the preset video synthesis rules, the video data associated with the target event in the basic video is synthesized to generate the target video under the user's request instruction. The target video is not stored in advance, which can save the storage space of the device.
[0136] Example 4
[0137] In Example 3, a target video is obtained by extracting specific video segments and performing corresponding video processing operations. When the duration of the target event exceeds the length of the segmented video, multiple basic videos need to be retrieved and processed to generate the target video. This not only causes the extracted data to be larger and inconvenient, but also makes the extracted video longer, which is not focused enough on the event review, and ultimately affects the user experience. In some embodiments, a method of extracting, splitting and splicing video data is adopted, and the video obtained in this way is more suitable as a target video and is more targeted. Therefore, Example 4 of the present invention proposes a video extraction and splicing device, please refer to Fig.12 , the device comprises:
[0138] A demultiplexing module, used for demultiplexing the basic video data into raw stream data;
[0139] A frame number marking module, used to mark the raw stream data according to the time sequence and determine the video frame number;
[0140] The splicing module is used to splice the raw stream data associated with the target event according to the video frame number to determine the target video data.
[0141] By using the video extraction and splicing device of this embodiment, the basic video data is demultiplexed into raw stream data; the raw stream data is marked according to the time sequence to determine the video frame number; the raw stream data from the starting time point to the predetermined time is spliced according to the video frame number to determine the target data. The video data can be processed more flexibly. For reviewing events, the video data at specific time points can be extracted and flexibly cut and spliced. According to the development of the event, the corresponding video data can be extracted and spliced, so that a video related to the event that satisfies the user can be obtained.
[0142] Implementation method three:
[0143] The present invention provides an electronic device and a storage medium, such as Fig.13 As shown, the system comprises at least one processor, at least one memory and computer program instructions stored in the memory.
[0144] Specifically, the processor may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of an embodiment of the present invention, and the electronic device includes at least one of the following: a smart camera, a mobile device with a smart camera, and a wearable device with a smart camera.
[0145] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory may be inside or outside a data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM) or a flash memory or a combination of two or more of these.
[0146] The processor reads and executes computer program instructions stored in the memory to implement any one of the methods for optimizing the intelligent camera detection model using edge computing, the sample confidence threshold selection method, and the model self-training method in the above-mentioned embodiment mode 1.
[0147] In one example, the electronic device may further include a communication interface and a bus, wherein the processor, the memory, and the communication interface are connected via the bus and communicate with each other.
[0148] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.
[0149] Bus includes hardware, software or both, and the parts of electronic equipment are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In suitable cases, bus may include one or more buses. Although the embodiment of the present invention describes and shows a specific bus, the present invention considers any suitable bus or interconnection.
[0150] In summary, the embodiments of the present invention provide a target video capture method based on image screening, a video extraction and splicing method, device, equipment and storage medium.
[0151] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.
[0152] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. Programs or code segments can be stored in machine-readable media, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target video capture method based on image screening, characterized in that: The method comprises: Recording a reference duration of a target event related to an infant's activity, and generating an index image of the target event when the reference duration meets a duration threshold corresponding to generating the target event; Calculating the confidence of each of the index images within a specified time period, extracting all index images that meet the confidence threshold requirements and establishing an index image list; According to each index image in the index image list, locate the relevant basic video, and synthesize the video data associated with the target event in the basic video according to a preset video synthesis rule to generate the target video; The step of calculating the confidence of each index image within a specified time period, extracting all index images that meet the confidence threshold requirement and establishing an index image list includes: The videos with the most clicks and the longest playback time in the historical video viewing data are regarded as the target videos that the users are interested in; According to the target video that the user is interested in, extracting key frames of event content in the target video that the user is interested in as historical key frame images; Assigning a weight value to each index image according to the similarity between the key frame image of the special event in the target event and the historical key frame image; The index images are sorted according to the confidence and weight value of each index image to generate an index image list.
2. The target video capture method based on image screening according to claim 1 is characterized in that: The step of sorting the index images according to the confidence and weight value of each index image to generate an index image list comprises: Sorting the index images according to the products obtained by multiplying the confidence and weight values of the index images, arranging the corresponding index images with larger products before the corresponding index images with smaller products; According to the confidence of each of the index images, some of the index images of each type with higher confidence are placed in a prominent position in the index image list.
3. The target video capture method based on image screening according to claim 1 is characterized in that: The step of locating a related basic video according to each index image in the index image list, and synthesizing video data associated with the target event in the basic video according to a preset video synthesis rule to generate the target video includes: Determine at least one continuous basic video associated with the target event based on a comprehensive consideration of the start time of the target event, the duration of the target video, and a preset video synthesis rule; The basic video is extracted, split and spliced to synthesize a related target video, and the target video is output to a user terminal, so that the user can view the event video he wants to watch.
4. The target video capture method based on image screening according to claim 3 is characterized in that: The step of determining at least one continuous basic video associated with the target event based on a comprehensive consideration of the start time of the target event, the duration of the target video, and a preset video synthesis rule includes: Obtaining time information of the label information of the target event and the duration corresponding to the target video, wherein the duration of the target video is less than or equal to the duration of the base video; Determining a first basic video where the target event is located according to the time information of the tag information; Dividing the first basic video into a plurality of video segments according to the duration corresponding to the target video; The starting time of the target event is used to locate a corresponding basic video, and the basic video is reasonably segmented. According to the position information of the starting time of the target event in the segmented video and the corresponding duration of the target video, at least one continuous basic video associated with the target event is determined.
5. The target video capture method based on image screening according to claim 4 is characterized in that: The dividing the first basic video into a plurality of video segments according to the duration corresponding to the target video comprises: Obtain a target duration corresponding to half the duration of the target video; The first basic video is segmented according to the target duration to obtain the multiple video segments.
6. The target video capture method based on image screening according to claim 3 is characterized in that: The method extracts, splits and splices the basic video to synthesize the relevant target video, and outputs the target video to the user terminal so that the user can view the event video he wants to watch, including: Demultiplexing the basic video into raw stream data; Mark the raw stream data according to the time sequence to determine the video frame number, wherein the frame number = the number of seconds * frame rate; The raw stream data associated with the target event is spliced according to the video frame number to determine the target video.
7. A target video capture device based on image screening, characterized in that: include: An index image acquisition module, used to record a reference duration of a target event related to an infant's activity, and to generate an index image of the target event when the reference duration meets a duration threshold corresponding to the target event; An image list building module is used to calculate the confidence of each index image within a specified time period, extract all index images that meet the confidence threshold requirements and build an index image list; A video synthesis module, for locating a related basic video according to each index image in the index image list, synthesizing the video data associated with the target event in the basic video according to a preset video synthesis rule to generate the target video, and outputting the target video corresponding to the target event; The step of calculating the confidence of each index image within a specified time period, extracting all index images that meet the confidence threshold requirement and establishing an index image list includes: The videos with the most clicks and the longest playback time in the historical video viewing data are regarded as the target videos that the users are interested in; According to the target video that the user is interested in, extracting key frames of event content in the target video that the user is interested in as historical key frame images; Assigning a weight value to each index image according to the similarity between the key frame image of the special event in the target event and the historical key frame image; The index images are sorted according to the confidence and weight value of each index image to generate an index image list.
8. An electronic device, characterized in that: include: At least one processor, at least one memory and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, implement the method according to any one of claims 1 to 6.
9. A medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Video query method, device and equipment and readable storage medium
CN111881320A
Video processing method and device, electronic equipment and storage medium
CN112235613A
Associating classifications with images
US20140341476A1