Method for processing dynamic image, electronic device and terminal device connected thereto
Patent Information
- Application Number
- CN202210827722.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-06
- Filing Date
- 2022-07-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-07-14
AI Technical Summary
[0003]1.现有监看系统,其摄影机自动撷取的影像,并没有考虑到婴幼儿身体的动作变异大小,所以撷取到的影像中,例如相当多是表情相似(例如同样是笑脸),或影片中的声音相似者(例如同样的笑声),但身体并没有明显的动作变化,故即使符合前述撷取影像的条件,但仍会撷取到相当多是身体动作单调且内容重复的影像,仍须在这些影像中通过人工剔除不满意者
[0022]此外,当预设对象为两个以上时可设定人脸的数量不小于身体的数量时,才符合筛选条件,使产生的串接影片有多个预设对象时,能确保截取时点的候选影片中可见每个预设对象的人脸,而能够符合使用者的期望。
Smart Images

Figure CN117241130B_ABST
Abstract
Description
Technical Field
[0001] This invention provides a dynamic image processing technology, and more particularly a dynamic image processing method, electronic device, and connected terminal device. Background Technology
[0002] An existing infant monitoring system uses cameras to automatically capture images based on artificial intelligence, primarily focusing on changes in facial expressions or voice. However, this existing system has the following problems:
[0003] 1. Existing monitoring systems automatically capture images without taking into account the variation in the movements of infants and young children. Therefore, many of the captured images contain similar expressions (e.g., the same smiling face) or similar sounds (e.g., the same laughter), but without significant changes in body movement. Thus, even if the aforementioned conditions for capturing images are met, a considerable number of images with monotonous and repetitive body movements will still be captured, and unsatisfactory images still need to be manually removed from these images.
[0004] 2. Furthermore, even if existing monitoring systems can use changes in facial expressions or voices as conditions for capturing images, they cannot sort and filter based on the intensity of the expressions or voices. For example, when selecting smiling faces, those who laugh loudly are prioritized over those who smile (or vice versa), and when selecting laughter, those with high decibels are prioritized over those with low decibels (or vice versa). Similarly, unsatisfactory images still need to be manually removed from these images.
[0005] 3. In addition, existing monitoring systems usually only capture images of infants and young children. If there are more than two people in the image, such as one person being an infant and another being an adult, the existing monitoring system usually only captures images based on changes in the infant's facial expressions or voice. If the capture condition is met but the camera only shows the adult's body and not their face, the image will still be selected, but it will obviously be considered unsatisfactory.
[0006] Therefore, the main focus of this invention is to solve the aforementioned problems of existing monitoring systems. Summary of the Invention
[0007] The inventor then devoted his mind and effort to research and development, and subsequently developed a method for processing dynamic images, an electronic device and a terminal device connected to it, which can use motion variation as a screening condition in order to achieve that the screened image content has more dynamic motion performance.
[0008] To achieve the above objectives, the present invention provides a method for processing dynamic images, which involves an electronic device communicating with a photographic device, reading and executing executable code, using artificial intelligence to identify a preset object, and processing the preset object as a dynamic image, including the following steps: Preset object identification: identifying the preset object from an initial image captured by the photographic device using artificial intelligence; Image filtering: setting a filtering condition, the filtering condition including detecting a movement variable of the preset object in the initial image, and selecting a capture point in the initial image when the filtering condition meets a threshold; and Video lip stitching: selecting a candidate video lip in the initial image according to the capture point, and combining one or more candidate videos to generate a strung video.
[0009] In one embodiment, the motion variation includes a motion level index (MLI) and a motion proportion index (MPI). The threshold includes a first threshold and a second threshold. The motion level index is calculated from multiple frames containing a first number of frames of the preset object within a preset time period. The difference between the first area occupied by the preset object in the Nth frame and the second area occupied in the (N-1)th frame is calculated, and the difference is compared with the first area. The motion proportion index is obtained by summing the differences in the areas of the multiple frames that are greater than the first threshold to obtain a second number of frames, and the second number is compared with the first number of frames. When the motion proportion index is greater than the second threshold, the selection condition is met.
[0010] In one embodiment, the first area and the second area are rectangular ranges enclosed by four boundary points, and the rectangular ranges are defined by the smallest range that can cover the preset object.
[0011] In one embodiment, the preset object is a specific infant, and the screening criteria also include the initial image containing at least the infant's face.
[0012] In one embodiment, the screening criteria further includes an ambient volume measured by the infant, and the screening criteria further includes the ambient volume within a volume range.
[0013] In one embodiment, the candidate videos are selected by ranking the scores of the infant's facial expressions at each capture point and selecting the highest score; or by ranking the scores of the variation in the area of the action at each capture point and selecting the highest score; or by ranking the scores of the facial area of the preset object at each capture point and selecting the highest score.
[0014] In one embodiment, when the number of preset objects is two or more, and at least one of them is an infant and at least one is an adult, the filtering condition further includes calculating the number of faces and bodies of the infant and the adult in the initial image, and the number of faces and bodies of the adult in the initial image. The filtering condition further includes that the number of faces of the infant and the adult is not less than the number of bodies of the two, comparing the rectangular range occupied by the preset objects in the frame, and calculating the motion area variation based on the larger area.
[0015] In one embodiment, in the image screening step, the capture time corresponding to the candidate video is used as a basis, while the capture time of other similar image content within a predetermined time before and / or after the candidate video is set to be excluded.
[0016] In one embodiment, the video concatenation step involves calculating a segment of time backward from the capture point to set the start point of the candidate video, and / or calculating the segment of time backward from the capture point to set the end point of the candidate video.
[0017] In one embodiment, there are multiple capture points, and the multiple candidate videos selected for each capture point are stored in the electronic device and / or a cloud database, and the multiple candidate videos are concatenated to form the concatenated video.
[0018] The present invention also provides an electronic device for processing moving images, which is communicatively connected to a camera device and a database. The database receives an initial image captured by the camera device and identifies a preset object using artificial intelligence. The electronic device processes the preset object as moving images. The electronic device includes: an intelligent processing unit electrically connected to the camera device or the database to read the initial image, and read and execute executable code to set a filtering condition for selecting a capture point in the initial image when a threshold is met. The filtering condition includes a motion variation. The intelligent processing unit selects a candidate video based on the capture point and combines one or more candidate videos to generate a serialized video.
[0019] The present invention also provides a terminal device that can communicate with the aforementioned electronic device. The terminal device carries an application program, and the terminal device executes the application program to receive push notifications of the cascading video from the electronic device.
[0020] By using filtering criteria including motion variation, it is possible to generate more dynamic video sequences of preset objects to meet the user's expectations.
[0021] Furthermore, users can select high or low levels of motion changes, facial expressions, and / or sound from the filter criteria according to their individual needs, so that the generated video clips can better meet the user's expectations.
[0022] In addition, when there are two or more preset objects, the number of faces can be set to be no less than the number of bodies to meet the filtering conditions. This ensures that when the generated video has multiple preset objects, the face of each preset object can be seen in the candidate video at the capture point, thus meeting the user's expectations. Attached Figure Description
[0023] Figure 1 This is a flowchart of the main steps of the processing method in a specific embodiment of the present invention.
[0024] Figure 2 This is a block diagram illustrating the steps of a specific embodiment of the processing method of the present invention.
[0025] Figure 3 This is a block diagram of an electronic device according to a specific embodiment of the present invention.
[0026] Figure 4 This is a block diagram of an electronic device according to another specific embodiment of the present invention.
[0027] Figure 5 This is a block diagram illustrating the area calculation of a rectangular region according to a specific embodiment of the present invention.
[0028] Figure 6a This is a schematic diagram of the rectangular area before the change in the infant's movement, according to a specific embodiment of the present invention.
[0029] Figure 6b This is a schematic diagram of the rectangular range after the infant's movements change according to a specific embodiment of the present invention.
[0030] Figure 7 This is a block diagram illustrating the selection criteria in a specific embodiment of the present invention.
[0031] Figure 8 This is a block diagram showing whether the variation in the action area of a specific embodiment of the present invention meets the first threshold.
[0032] Figure 9 This is a block diagram showing whether the variation in the action ratio in a specific embodiment of the present invention meets the second threshold.
[0033] Figure 10 This is a schematic diagram of the electronic device displaying relevant data in the background, according to a specific embodiment of the present invention.
[0034] Figure 11a This is a schematic diagram illustrating the interception time point that meets the filtering conditions in a specific embodiment of the present invention.
[0035] Figure 11b This is a schematic diagram illustrating a specific embodiment of the present invention where the interception time point does not meet the filtering criteria.
[0036] Figure 11c This is another schematic diagram illustrating that the interception time point of a specific embodiment of the present invention does not meet the filtering conditions.
[0037] Figure 12 This is a block diagram showing the selected capture time point in a specific embodiment of the present invention.
[0038] Figure 13 This is a schematic diagram of an electronic device displaying a selected time point in the background when an infant is shown in a specific embodiment of the present invention.
[0039] Figure 14 This is a schematic diagram of a selected time point in the background display of an infant and an adult on an electronic device according to a specific embodiment of the present invention.
[0040] Figure 15 This is a block diagram generated from candidate videos by concatenating videos according to a specific embodiment of the present invention.
[0041] Figure 16 This is a schematic diagram of a specific embodiment of the present invention, showing the concatenated video playback to a terminal device.
[0042] Figure Labels
[0043] 100 Handling Method
[0044] 101 Default Object Identification
[0045] 102 Image Screening
[0046] 103 Video Interlocking
[0047] 200 electronic devices
[0048] 300 terminal device
[0049] 400 camera devices
[0050] 500 database
[0051] 501 Body Intelligence Recognition Sub-database
[0052] 502 Facial Intelligent Recognition Sub-database
[0053] 503 Crying Sound Intelligent Recognition Sub-database
[0054] 504 Smile Intelligent Recognition Sub-database
[0055] 10 Intelligent Processing Units
[0056] 20 Wireless Communication Units
[0057] A and B bodies
[0058] P Default Object
[0059] X and Y faces
[0060] V1 Initial Footage
[0061] V2 Candidate Videos
[0062] V3 video concatenation
[0063] A1 First area
[0064] A2 Second Area Detailed Implementation
[0065] To fully understand the purpose, features, and effects of the present invention, the present invention will be described in detail below with reference to the following specific embodiments and accompanying drawings:
[0066] Please refer to Figures 1 to 16 This invention provides a method 100 for processing dynamic images, an electronic device 200, and a terminal device 300 connected thereto. The processing method 100 includes the steps of preset object identification 101, image filtering 102, and video concatenation 103; the electronic device 200 includes an intelligent processing unit 10 and a wireless communication unit 20, wherein:
[0067] The processing method 100 involves an electronic device 200 reading and executing executable code to identify a preset object P using artificial intelligence and performing dynamic image processing on the preset object P, thereby executing... Figure 1 The steps shown are: preset object identification 101, image filtering 102, and video concatenation 103. (And refer to...) Figure 2 The preset object identification step 101 mainly performs the function of identifying whether there is a preset object in the initial image within a preset time period; the image filtering step 102 mainly performs the function of whether the filtering conditions meet the threshold and select the capture time point; and the video concatenation step 103 mainly performs the function of capturing candidate videos at the capture time point and generating a concatenated video, which can be pushed to the terminal device 300.
[0068] The electronic device 200, such as Figure 3 , 4As shown, the system is configured to communicate with a camera device 400 and a database 500. The database 500 receives an initial image V1 captured by the camera device 400, identifies a preset object P using artificial intelligence, and processes the preset object P as a dynamic image using an electronic device 200. The intelligent processing unit 10 is electrically connected to the camera device 300 or the database 500 to read the initial image V1. In the above embodiment, the camera device 400 and the database 500 are external devices independent of the electronic device 200. However, in different embodiments, the camera device 400 and the database 500 may also be integrated into the electronic device 200 and systematized.
[0069] In one embodiment, the camera device 400 is a webcam, and the database 500 is a cloud database (e.g., ...). Figure 3 As shown), after initialization, the photographic device 400 can remotely communicate with the database 500 via the Internet, and log in after completing an authentication process (such as logging in with an account and password) to capture and store images. The database 500 can be a cloud database or a local database belonging to the electronic device 200 (e.g., [database name missing]). Figure 4 (as shown in the figure), or the electronic device 200's local database and cloud database coexist (not shown in the figure).
[0070] When processing method 100 is executed, in the preset object identification 101 step, a preset object P is identified in the initial image V1 captured by the self-photography device 400 using artificial intelligence, and then the image screening 102 step is executed. In one embodiment, the preset object P is specifically an infant, but it is not limited to this; the number of preset objects P can also be two or more, with at least one being an infant and at least one being an adult. After the photography device 400 is started, the preset object identification 101 step cycles through a preset time (e.g., 30 seconds). If the photography device 400 identifies a preset object P in the initial image V1 within the preset time, the image screening 102 step is executed; if no preset object P is identified in the initial image V1 within the preset time, the preset object identification 101 step is repeated in the next preset time. When no preset object P is identified in the initial image V1 within the preset time, it is compared with the preset object P last identified in the previous preset time; however, if no preset object P was identified in the previous preset time, it is defined as no data. The artificial intelligence identification is performed, for example, through an Artificial Neural Network (ANN).
[0071] In step 102 of image screening, a screening condition is set. This screening condition includes detecting a motion variable of a preset object P in the initial image V1. When the screening condition meets a threshold, it is selected as a capture moment in the initial image V1. In one embodiment, the motion variable includes a motion level index (MLI) and a motion proportion index (MPI). In one embodiment, the threshold includes a first threshold and a second threshold.
[0072] The variation in the action area is calculated from multiple frames containing a first number of preset objects P within a preset time period in the initial image V1. This variation is obtained by calculating the difference between the first area A1 occupied by the preset object P in the Nth frame and the second area A2 occupied by the preset object P in the (N-1)th frame, and then comparing this difference with the first area A1 (and referring to...). Figure 7 ).
[0073] In one embodiment, the first area A1 and the second area A2 are rectangular ranges enclosed by four boundary points, and the rectangular ranges are defined by the smallest range that can cover the preset object P. Figure 6a , 6b As shown (and refer to) Figure 5 When the camera device 400 identifies a preset object P in the initial image V1 within a preset time, it obtains the boundary points (x1, y1) and (x2, y2) of the two opposite corners of the rectangular area occupied by the preset object P from the initial image V1, and then calculates the area of the rectangular area (formula: area = |(x2-x1)|*|(y2-y1)|). When the camera device 400 identifies two or more preset objects P in the initial image V1 (not shown in the figure) within a preset time, it obtains the boundary points (x1, y1) and (x2, y2) of the two opposite corners of the rectangular area occupied by each preset object P from the initial image V1, calculates the area of each rectangular area, then compares the rectangular areas occupied in the image, and takes the one with the larger area as the detection target, obtains the boundary points (Tx1, Ty1) and (Tx2, Ty2) of the two opposite corners of the rectangular area occupied by the detection target, and calculates the area of the rectangular area of the detection target (formula: area = (Tx2 - Tx1) * (Ty2 - Ty1)).
[0074] The motion ratio variation is calculated by summing the differences in area across multiple frames that exceed the first threshold to obtain a second count. This second count is obtained by comparing it to the first count. When the motion ratio variation exceeds the second threshold, the selection condition is met (and referenced). Figure 7For example, the first number of frames is, for instance, 90 frames captured within a preset time of 30 seconds, and the first threshold is set to 30% (e.g., Figure 8 As shown), and assuming that the area difference (MLI) of 40 out of these 90 frames is greater than the first threshold of 30%, the motion ratio variation (MPI) is 44%. If the second threshold is set to 30% (as shown), Figure 9 As shown in the image, if the motion ratio variation is 44%, which is greater than the second threshold of 30%, the result of the screening condition meeting the threshold is "yes". If only 20 of the 90 frames have an area difference greater than the first threshold of 10%, then the motion ratio variation is 22%. If the motion ratio variation is less than the second threshold of 30%, the result of the screening condition meeting the threshold is "no". In other words, within each preset 30-second video, the motion change of the preset object P must meet the motion ratio variation of more than 30% to be selected as the cut-off point, ensuring that the preset object P in the video has a high degree of motion change.
[0075] In one embodiment, the filtering criteria further include that the initial image V1 contains at least the infant's face, and that the ambient volume measured by the infant is within a certain volume range. It also includes whether a smile is detected on the infant's face, and whether the infant's crying is detected. When the motion ratio variation is greater than the second threshold, if a smile is further detected on the infant's face (judgment result "yes"), and no crying is detected (judgment result "no"), the judgment result for whether the filtering criteria meet the threshold is "yes"; conversely, even if the motion ratio variation is greater than the second threshold, but no smile is detected on the infant's face (judgment result "no"), or crying is detected (judgment result "yes"), the judgment result for whether the filtering criteria meet the threshold is "no". Figure 10 The image shows a background display screen 201 of the electronic device 200, which indicates the viewing angle, ambient volume, whether the infant is in the scene, motion variability (motion area variability MLI and motion scale variability MPI), preset object type (adult / infant), facial expression (e.g., smile), and event (whether there is crying). Figure 11a The image shown illustrates the images of infants and young children at the selected time points that meet the screening criteria; as shown... Figure 11b The image shown illustrates the point in time when the infant or toddler is not present in the scene and therefore does not meet the screening criteria; as shown... Figure 11c The image shown is a snapshot of an infant crying because they did not meet the screening criteria.
[0076] Furthermore, in the image screening step 102, based on the selected candidate video at the corresponding cut-off time point, the cut-off times of other similar image content within a predetermined time period before and / or after the candidate video are set to be excluded (and refer to...). Figure 12 For example, the preset time can be set within a range of 30 seconds to 2 minutes. Taking 1 minute as an example, within 1 minute before and after the selected time, even if there are those that meet the filtering criteria, they are set to be excluded and not selected.
[0077] The aforementioned initial image V1 contains the detection of the infant's face, according to Figure 13 In the display screen 201 shown (and refer to Tables 1A and 1B below), the data listed at time 03:52:19 shows that the infant's face X was detected, including coordinates x1, y1, x2, y2 as {1446, 29, 1494, 85} and a confidence score of 0.69 (out of a total of 1). In one embodiment, the infant's body A was further detected using coordinates x1, y1, x2, y2 as {1389, 6, 1869, 447} and a confidence score of 0.96. At the same time, the coordinates x1, y1, x2, y2 of body B, body C, face Y, and face Z were all {0, 0, 0, 0}, and the confidence scores were all 0. At this time, the number of detected faces was 1, and the number of detected bodies was also 1.
[0078] Table 1A: (The right side of Table 1A is a continuation of the left side of Table 1B)
[0079]
[0080]
[0081] Table 1B:
[0082]
[0083] In one embodiment, assuming the number of preset objects P is two or more, i.e., at least one of them is an infant and at least one is an adult, the filtering condition further includes calculating the number of faces and bodies of infants and adults in the initial image V1, and the number of faces and bodies of adults in the initial image V1, and further detecting that the number of faces of both infants and adults is not less than the number of bodies of both (and referring to...). Figure 12 ).
[0084] The aforementioned detection of faces and bodies in infants and adults, based on... Figure 14In the display screen 201 shown (and refer to Tables 1A and 1B as above), referring to the data listed at time 03:52:03, including coordinates x1, y1, x2, y2 as {1461, 4, 1896, 450}, and a confidence score of 0.98, it is determined that the infant's body A has been detected, and based on the coordinates x1, y1, x2, y2 as {1416, 29, 1455, 96}, and the confidence score of... Based on data such as 0.65, an infant's face (X) was detected. Furthermore, using coordinates x1, y1, x2, y2 as {1203, 695, 1497, 825} and a confidence score of 0.52, an adult's body (B) was detected. And using coordinates x1, y1, x2, y2 as {1674, 9, 1758, 78} and a confidence score of 0.58, an adult's face (Y) was detected. At the same time, the coordinates x1, y1, x2, y2 of both body (C) and face (Z) were {0, 0, 0, 0}, and their confidence scores were also both 0.
[0085] Continuing from the above, based on the data listed at time 03:52:03, it indicates that the number of infant faces and adult faces detected is 1 each, and the number of infant bodies and adult bodies is also 1 each. At this time, the total number of infant and adult faces (2) equals the total number of adult bodies (2), thus meeting the filtering criteria. Figure 12 The result of determining that the number of faces is not less than the number of bodies is "yes".
[0086] Furthermore, assuming a different initial image V1 (not shown in the figure), if the number of infant bodies and faces detected is 1 each, but only the number of adult faces detected is 1, while the number of adult bodies is 0, then the total number of faces (2) for both infants and adults is greater than the total number of bodies (1), which still meets the screening condition. Figure 12 The result of determining that the number of faces is not less than the number of bodies is still "yes".
[0087] Conversely, assuming a different initial image V1 (not shown in the figure), if the number of infant faces detected is 1, but the number of adult faces is 0, even if the number of bodies for both infants and adults is 1, the total number of faces (1) is less than the total number of bodies (2), thus failing to meet the screening condition. Figure 12 The result of determining that the number of faces is not less than the number of bodies is still "no". Therefore, when the number of preset objects P is two or more, each person's face will be present before being selected, and there will be no image of someone only having a body without a face.
[0088] In step 103 of video concatenation, a candidate video clip V2 is selected from the initial image V1 at the capture time point, and one or more candidate videos V2 are combined to generate a concatenated video V3 (and refer to...). Figure 15 ).
[0089] In one embodiment, in the video concatenation step 103, a segment of time is calculated backward from the capture point to set the start point of the candidate video, and / or the segment of time is calculated backward from the capture point to set the end point of the candidate video V2. In one embodiment, assuming the segment time is set to 5 seconds, 5 seconds can be calculated both backward and backward from the capture point to extract each candidate video with a playback time of 10 seconds from the start point to the end point.
[0090] Furthermore, the selection of candidate videos is based on several methods: ranking the infant's facial expressions at each capture point according to their scores, and selecting the highest-scoring video; or ranking them by the variation in the area of the facial expression at each capture point according to the numerical values, and selecting the highest-scoring video; or ranking them by the facial area of a preset object P at each capture point according to the numerical values, and selecting the highest-scoring video. For example, ranking the infant's facial expressions by scores, such as a smile, might assign a score of 0.3 for a slight smile but a score of 1 for a wide, open laugh. In this case, the video with a score of 1 for a wide laugh would be ranked highest and selected. Similarly, ranking by the variation in the facial expression area at each capture point, and by the facial area of the preset object P at each capture point, is based on the detected variation in the facial expression area and the size of the facial area, with the highest-scoring video being selected. Therefore, the selected time point can be not only those who are smiling, but also those who are laughing heartily; it can also be those who are moving, but those whose movement area varies greatly; it can also be those whose face area is the largest. So it is not enough for people to just have facial expressions to be selected.
[0091] In one embodiment, there are multiple capture points, and the multiple candidate videos V2 selected for each capture point are stored in the local database of the electronic device 200 and / or a cloud database, and the multiple candidate videos V2 are concatenated to form the concatenated video V3.
[0092] In one embodiment, the database 500 further includes a body recognition sub-database 501 for recognizing the infant's body; a face recognition sub-database 502 for recognizing the infant's face; a crying sound recognition sub-database 503 for recognizing the infant's crying sound; and / or a smile recognition sub-database 504 for recognizing the infant's smile.
[0093] The terminal device 300 can be a portable mobile communication device, such as a smartphone, tablet computer, or laptop computer, capable of communicating with the wireless communication unit 20 of the electronic device 200 via the Internet. The terminal device 300 carries an application 301. After executing the application 301 and completing an authentication process (e.g., logging in with an account and password), the user logs in to receive push notifications of the concatenated video V3 from the electronic device 200 (e.g.,...). Figure 16 As shown), users can watch the concatenated video V3 through the terminal device 300.
[0094] It is not difficult to see from the above description that the features of the present invention are:
[0095] 1. The processing method and electronic device for processing dynamic images of the present invention include motion variation as a selection criterion. A preset object P in the initial image must exhibit preset motion variations to meet a threshold and be selected as a capture point. Among the candidate videos selected based on this capture point, the preset object P exhibits more dynamic motion performance, thereby generating a serialized video V3 with rich motion variations of the preset object P to meet the user's expectations. Furthermore, the serialized video V3 can be pushed to a terminal device communicatively connected to the electronic device for playback.
[0096] 2. Furthermore, the processing method and electronic device for processing dynamic images of the present invention can sort and filter according to the degree of the filtering conditions, so as to select those with a high or low degree from the filtering conditions, so that the generated serialized video V3 can better meet the user's expectations.
[0097] 3. In addition, if there are two or more preset objects P in the image, including at least one infant and at least one adult, the number of faces can be no less than the number of bodies to meet the screening criteria, so that at least the face of each preset object P can be seen in the candidate video at the capture time point, so that the generated serialized video V3 can also meet the user's expectations when there are multiple preset objects P.
[0098] The present invention has been disclosed above with reference to preferred embodiments. However, those skilled in the art should understand that these embodiments are for illustrative purposes only and should not be construed as limiting the scope of the invention. It should be noted that all variations and substitutions equivalent to these embodiments should be considered within the scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.
Claims
1. A method for processing dynamic images, comprising an electronic device communicatively connected to a photographic device and reading and executing executable code to identify a preset object using artificial intelligence, and processing the preset object as a dynamic image, comprising the following steps: Preset object recognition: The preset object is identified from an initial image captured by the camera device using artificial intelligence; Image filtering: A filtering condition is set, which includes detecting a motion variation of the preset object in the initial image, and selecting a cropping point in the initial image when the filtering condition meets a threshold; and Video stitching: Select a candidate video from the initial image at the capture time point, and combine one or more of the candidate videos to generate a stitched video; The motion variation includes a motion area variation and a motion proportion variation. The threshold includes a first threshold and a second threshold. The motion area variation is calculated by taking multiple frames containing a first number of the preset object within a preset time period of the initial image, and calculating the difference between the first area occupied by the preset object in the Nth frame and the second area occupied in the (N-1)th frame, and comparing the difference with the first area. The motion proportion variation is obtained by summing the differences in the areas of the multiple frames that are greater than the first threshold to obtain a second number, and comparing the second number with the first number. When the motion proportion variation is greater than the second threshold, the selection condition is met.
2. The method for processing dynamic images according to claim 1, characterized in that, The first area and the second area are rectangular ranges enclosed by four boundary points, and the rectangular ranges are defined by the smallest range that can cover the preset object.
3. The method for processing dynamic images according to claim 2, characterized in that, The preset target is a specific infant or toddler, and the screening criteria also include the initial image containing at least the infant or toddler's face.
4. The method for processing dynamic images according to claim 3, characterized in that, The screening criteria also include an ambient volume measured by the infant, and the screening criteria further include the ambient volume within a volume range.
5. The method for processing dynamic images according to claim 3, characterized in that, The candidate videos are selected by ranking the scores of the infant's facial expressions at each capture point and selecting the highest score; or by ranking the scores of the variation in the area of the action at each capture point and selecting the highest score; or by ranking the scores of the facial area of the preset object at each capture point and selecting the highest score.
6. The method for processing dynamic images according to claim 2, characterized in that, When the number of preset objects is two or more, and at least one of them is an infant and at least one is an adult, the filtering condition also includes calculating the number of faces and bodies of the infant and the adult in the initial image, and the number of faces and bodies of the adult in the initial image. The filtering condition further includes that the number of faces of the infant and the adult is not less than the number of bodies of the two, comparing the rectangular range occupied by these preset objects in the frame, and calculating the motion area variation based on the larger area.
7. The method for processing dynamic images according to claim 1, characterized in that, In the image screening step, the capture time corresponding to the selected candidate video is used as the basis, while the capture time of other similar image content within a predetermined time before and / or after the selected video is set as excluded.
8. The method for processing dynamic images according to claim 1, characterized in that, In the process of stringing together the videos, a segment of time is calculated backward from the capture point to set the starting point of the candidate video, and / or the segment of time is calculated backward from the capture point to set the ending point of the candidate video.
9. The method for processing dynamic images according to claim 8, characterized in that, There are multiple capture points, and the multiple candidate videos selected for each capture point are stored in the electronic device and / or a cloud database, and the multiple candidate videos are strung together to form the strung video.
10. A terminal device communicatively connected to an electronic device performing the method of claim 1, the terminal device carrying an application that executes the application to receive push notifications of the cascaded video from the electronic device.
11. An electronic device for processing moving images, communicatively connected to a photographic device and a database, the database receiving an initial image captured by the photographic device and identifying a preset object using artificial intelligence, the electronic device processing the preset object as moving images, the electronic device comprising: An intelligent processing unit is electrically connected to the camera device or the database to read the initial image, and reads and executes an executable code to set a filtering condition for selecting a capture point in the initial image when a threshold is met. The filtering condition includes a motion variation. The intelligent processing unit selects a candidate video according to the capture point and combines one or more candidate videos to generate a serialized video. The motion variation includes a motion area variation and a motion proportion variation. The threshold includes a first threshold and a second threshold. The motion area variation is calculated by taking multiple frames containing a first number of the preset object within a preset time period of the initial image, and calculating the difference between the first area occupied by the preset object in the Nth frame and the second area occupied in the (N-1)th frame, and comparing the difference with the first area. The motion proportion variation is obtained by summing the differences in the areas of the multiple frames that are greater than the first threshold to obtain a second number, and comparing the second number with the first number. When the motion proportion variation is greater than the second threshold, the selection condition is met.
12. The electronic device for processing moving images according to claim 11, characterized in that, The database is the local database of the electronic device and / or the cloud database.
13. The electronic device for processing moving images according to claim 12, characterized in that, The preset target includes at least one infant, and the database further includes a body recognition sub-database for recognizing the infant's body; a face recognition sub-database for recognizing the infant's face; a crying sound recognition sub-database for recognizing the infant's crying sound; and / or a smile recognition sub-database for recognizing the infant's smile.
14. A terminal device communicatively connected to the electronic device of claim 11, the terminal device carrying an application that executes the application to receive push notifications of the cascading video from the electronic device.
Citation Information
Patent Citations
Systems and methods for selecting media items
CN105247845A
Cataloging video and creating video summaries
US9620168B1