Key picture processing method and device, equipment, medium and program product
By determining the video scene to generate a set of candidate frames, and then performing quality assessment and screening, the problems of insufficient multi-scene adaptation and quality assessment in existing technologies are solved, and the accurate extraction and efficient management of key frames in the video are achieved.
Patent Information
- Application Number
- CN202510780356.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies lack multi-scene adaptability and effective quality assessment mechanisms when extracting key information from videos, resulting in inaccurate and inefficient extraction.
By determining the video scene, generating candidate picture sets, evaluating the picture quality, screening out key picture sets, and managing them, including image enhancement and storage.
It achieves accurate extraction and efficient management of key images from videos, ensuring that the extracted images are of high value and can be used efficiently.
Smart Images

Figure CN120640086A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a key picture processing method, apparatus, device, medium, and program product. Background Art
[0002] With the continuous generation of massive amounts of video, extracting key information from videos is crucial. Effectively extracting this information allows viewers to understand the video content in a timely manner. Deep learning models such as convolutional neural networks are used to identify target objects in videos. However, most of these models are designed for a single scene and lack adaptability to multiple scenes. Furthermore, they lack effective quality assessment and screening mechanisms after recognition, resulting in inaccurate and inefficient extraction of key information from videos.
[0003] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0004] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Some embodiments of the present disclosure provide key frame processing methods, devices, equipment, media, and program products to solve the technical problems mentioned in the above background technology section.
[0006] In a first aspect, some embodiments of the present disclosure provide a key picture processing method, including: determining at least one video scene corresponding to an acquired target video; generating a candidate picture set corresponding to at least one target object in the target video based on the at least one video scene; performing quality evaluation on each candidate picture in the candidate picture set to obtain picture quality information; generating a key picture set based on the picture quality information corresponding to each candidate picture and the candidate picture set; and performing picture management on each key picture in the key picture set.
[0007] Optionally, the above-mentioned generating a candidate picture set corresponding to at least one target object in the above-mentioned target video based on the above-mentioned at least one video scene includes: determining the above-mentioned at least one target object based on the above-mentioned at least one video scene; obtaining a pre-trained multi-object detection model corresponding to the above-mentioned at least one target object; performing video segmentation processing on the above-mentioned target video to obtain a sub-video sequence; for each sub-video in the above-mentioned sub-video sequence, using the above-mentioned multi-object detection model to determine at least one object motion information corresponding to at least one target object in the above-mentioned sub-video; and generating the above-mentioned candidate picture set based on the obtained at least one object motion information sequence.
[0008] Optionally, the above-mentioned quality assessment is performed on each candidate picture in the above-mentioned candidate picture set to obtain picture quality information, including: determining the target object corresponding to the above-mentioned candidate picture as the picture object; determining at least one picture quality assessment index corresponding to the above-mentioned picture object; determining at least one indicator content corresponding to the above-mentioned at least one picture assessment index based on the above-mentioned candidate picture; and generating picture quality information corresponding to the above-mentioned candidate picture based on the above-mentioned at least one indicator content.
[0009] Optionally, the above-mentioned generating a key picture set for the at least one target object based on the picture quality information corresponding to each candidate picture and the above-mentioned candidate picture set includes: preliminarily screening each candidate picture in the above-mentioned candidate picture set according to the picture quality information corresponding to each candidate picture to obtain a first filtered picture set; obtaining at least one scene filtering requirement corresponding to the above-mentioned at least one target object; re-screening each picture in the above-mentioned first filtered picture set according to the above-mentioned at least one scene filtering requirement to obtain a second filtered picture set; deduplicating duplicate pictures or removing similar pictures from each picture in the above-mentioned second filtered picture set to obtain a removed picture set; adjusting the time distribution of each picture in the above-mentioned removed picture set to obtain an adjusted picture set; and performing image enhancement on each picture in the above-mentioned adjusted picture set to obtain a key picture set.
[0010] Optionally, the above-mentioned picture management of each key picture in the above-mentioned key picture set includes: for each key picture in the above-mentioned key picture set, executing the following processing steps: determining the video source and picture timestamp corresponding to the above-mentioned key picture; generating key picture metadata corresponding to the above-mentioned key picture according to the above-mentioned video source and the above-mentioned picture timestamp; executing corresponding storage for the above-mentioned key picture and the above-mentioned key picture metadata; and the above-mentioned method also includes: in response to receiving picture retrieval information for the first target key picture, locating the video source and the corresponding picture position corresponding to the above-mentioned first target key picture according to the key picture metadata corresponding to the above-mentioned first target key picture.
[0011] Optionally, the above-mentioned picture management of each key screen in the above-mentioned key screen set includes: in response to executing atlas creation, creating a customized atlas for each key screen in the above-mentioned key screen set to obtain each atlas adapted to the needs of different scenarios, wherein the above-mentioned atlases support copying, moving and sharing of key screens; in response to executing content sharing, sharing the target atlas and / or the second target key screen determined by the recommendation object to the target terminal; in response to executing batch export of key screens, batch exporting the above-mentioned at least one atlas and / or at least one key screen; in response to executing key screen abnormality marking, determining at least one key screen in the above-mentioned key screen set that represents the existence of a predefined abnormal event as at least one abnormal screen; marking the abnormal label to the above-mentioned at least one abnormal screen to support early warning and backtracking of the abnormal event; in response to executing key screen recommendation, according to the feature preference corresponding to the above-mentioned recommendation object, recommending at least one key screen in the above-mentioned key screen set corresponding to the above-mentioned feature preference to the above-mentioned recommendation object.
[0012] Optionally, the above method also includes: in response to receiving at least one picture search parameter filled in on the user interaction page, according to the above at least one picture search parameter, searching for at least one corresponding key picture from the above key picture set as at least one matching picture, wherein the above at least one picture search parameter includes at least one of the following: picture quality information, number of key pictures, time range, scene requirements, the above user interaction page also displays the association relationship between the above key picture set and the above target video, and the above user interaction page supports the generation standard corresponding to the key pictures; the above at least one matching picture is displayed in the above user interaction page in a target display form.
[0013] In a second aspect, some embodiments of the present disclosure provide a key picture processing device, including: a determination unit, configured to determine at least one video scene corresponding to an acquired target video; a first generation unit, configured to generate a candidate picture set corresponding to at least one target object in the target video based on the at least one video scene; a quality assessment unit, configured to perform quality assessment on each candidate picture in the candidate picture set to obtain picture quality information; a second generation unit, configured to generate a key picture set based on the picture quality information corresponding to each candidate picture and the candidate picture set; and a management unit, configured to perform picture management on each key picture in the key picture set.
[0014] Optionally, the first generation unit can be configured to: determine the at least one target object based on the at least one video scene; obtain a pre-trained multi-object detection model corresponding to the at least one target object; perform video segmentation processing on the target video to obtain a sub-video sequence; for each sub-video in the sub-video sequence, use the multi-object detection model to determine at least one object motion information corresponding to at least one target object in the sub-video; and generate the candidate picture set based on the obtained at least one object motion information sequence.
[0015] Optionally, the quality assessment unit can be configured to: determine the target object corresponding to the above-mentioned candidate picture as the picture object; determine at least one picture quality assessment index corresponding to the above-mentioned picture object; determine at least one indicator content corresponding to the above-mentioned at least one picture assessment index based on the above-mentioned candidate picture; and generate picture quality information corresponding to the above-mentioned candidate picture based on the above-mentioned at least one indicator content.
[0016] Optionally, the second generation unit can be configured to: preliminarily screen each candidate picture in the above-mentioned candidate picture set according to the picture quality information corresponding to each candidate picture to obtain a first filtered picture set; obtain at least one scene filtering requirement corresponding to the above-mentioned at least one target object; re-screen each picture in the above-mentioned first filtered picture set according to the above-mentioned at least one scene filtering requirement to obtain a second filtered picture set; deduplicate or remove similar pictures for each picture in the above-mentioned second filtered picture set to obtain a removed picture set; adjust the time distribution of each picture in the above-mentioned removed picture set to obtain an adjusted picture set; perform image enhancement on each picture in the above-mentioned adjusted picture set to obtain a key picture set.
[0017] Optionally, the management unit can be configured to: for each key picture in the above-mentioned key picture set, perform the following processing steps: determine the video source and picture timestamp corresponding to the above-mentioned key picture; generate key picture metadata corresponding to the above-mentioned key picture based on the above-mentioned video source and the above-mentioned picture timestamp; perform corresponding storage for the above-mentioned key picture and the above-mentioned key picture metadata; and the above-mentioned method also includes: in response to receiving picture retrieval information for the first target key picture, locating the video source and corresponding picture position corresponding to the above-mentioned first target key picture according to the key picture metadata corresponding to the above-mentioned first target key picture.
[0018] Optionally, the management unit can be configured to: in response to executing atlas creation, create a customized atlas for each key screen in the above-mentioned key screen set to obtain each atlas adapted to the requirements of different scenarios, wherein the above-mentioned atlases support copying, moving and sharing of key screens; in response to executing content sharing, share the target atlas and / or the second target key screen determined by the recommendation object to the target terminal; in response to executing batch export of key screens, batch export the above-mentioned at least one atlas and / or at least one key screen; in response to executing key screen abnormality marking, determine at least one key screen in the above-mentioned key screen set that represents the existence of a predefined abnormal event as at least one abnormal screen; mark the abnormal label to the above-mentioned at least one abnormal screen to support early warning and backtracking of abnormal events; in response to executing key screen recommendation, according to the feature preference corresponding to the above-mentioned recommendation object, recommend at least one key screen in the above-mentioned key screen set corresponding to the above-mentioned feature preference to the above-mentioned recommendation object.
[0019] Optionally, the device also includes: in response to receiving at least one picture search parameter filled in on the user interaction page, according to the at least one picture search parameter, searching for at least one corresponding key picture from the key picture set as at least one matching picture, wherein the at least one picture search parameter includes at least one of the following: picture quality information, number of key pictures, time range, scene requirements, the user interaction page also displays the association relationship between the key picture set and the target video, and the user interaction page supports the generation standard corresponding to the key pictures; the at least one matching picture is displayed on the user interaction page in a target display form.
[0020] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0021] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0022] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation manner in the first aspect when executed by a processor.
[0023] The aforementioned embodiments of the present disclosure have the following beneficial effects: Through the key frame processing methods of some embodiments of the present disclosure, key frames for at least one video scene can be accurately extracted from a target video, and efficient key frame management can be achieved. Specifically, the reason for the lack of accuracy and efficiency of the relevant key frames is that deep learning models such as convolutional neural networks are used to identify target objects in videos. However, most of these models are designed for a single scene and lack the ability to adapt to multiple scenes. Furthermore, they lack effective quality assessment and screening mechanisms after identification, resulting in inaccurate and inefficient extraction of key information from the video. Based on this, the key frame processing methods of some embodiments of the present disclosure first determine at least one video scene corresponding to the acquired target video. Here, determining the at least one video scene facilitates the subsequent determination of at least one target object for key frame extraction. Then, based on the at least one video scene, a set of candidate frames corresponding to the at least one target object in the target video can be accurately generated, serving as a preliminary screening set of key frames. Next, quality assessment is performed on each candidate frame in the candidate frame set to obtain image quality information. Through this quality assessment, key frames with more critical content and greater relevance to the at least one target object can be screened from the candidate frame set. Furthermore, based on the image quality information corresponding to each candidate frame and the candidate frame set, a key frame set for the at least one target object can be accurately generated, ensuring that the extracted key frames are of high value. This allows for precise determination of the key content of at least one target object (i.e., one or more target objects) in the target video. Finally, by managing the individual key frames in the key frame set, efficient extraction and utilization of the key frame set can be achieved, enabling a more efficient understanding of the key frame content in the target video. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0025] Figure 1 is a schematic diagram of an application scenario of a key picture processing method according to some embodiments of the present disclosure;
[0026] Figure 2 is a flow chart of some embodiments of the key picture processing method according to the present disclosure;
[0027] Figure 3 is a flowchart of other embodiments of the key picture processing method according to the present disclosure;
[0028] Figure 4 is a schematic structural diagram of some embodiments of the key picture processing device according to the present disclosure;
[0029] Figure 5 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0030] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0031] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0035] Before performing any operations involving the collection, storage, and use of user personal information (such as key screens) in this disclosure, relevant organizations or individuals must fulfill their obligations, including conducting personal information security impact assessments, fulfilling their obligation to inform the personal information subject, and obtaining the prior authorization and consent of the personal information subject.
[0036] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0037] Figure 1 3 is a schematic diagram of an application scenario of a key picture processing method according to some embodiments of the present disclosure.
[0038] exist Figure 1In an application scenario, first, the electronic device 101 can determine at least one video scene 103 corresponding to the acquired target video 102. Then, based on the at least one video scene 103, the electronic device 101 can generate a candidate picture set 105 corresponding to at least one target object 104 in the target video 102. Furthermore, the electronic device 101 can perform a quality assessment on each candidate picture in the candidate picture set 105 to obtain picture quality information. In this application scenario, the candidate picture set 105 corresponds to a picture quality information set 106. Next, the electronic device 101 can generate a key picture set 107 based on the picture quality information corresponding to each candidate picture and the candidate picture set 105. Finally, the electronic device 101 can perform picture management on each key picture in the key picture set 07.
[0039] It should be noted that the electronic device 101 can be hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or it can be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.
[0040] It should be understood that Figure 1 The number of electronic devices in the embodiment is merely illustrative. Any number of electronic devices may be provided according to implementation requirements.
[0041] Continue to refer Figure 2 , shows a process 200 of some embodiments of the key picture processing method according to the present disclosure. The key picture processing method includes the following steps:
[0042] Step 201: Determine at least one video scene corresponding to the acquired target video.
[0043] In some embodiments, the execution subject of the above key picture processing method (for example Figure 1The electronic device 101 shown) can determine at least one video scene corresponding to the acquired target video through a wired connection or a wireless connection. The target video can be a video for which key frame extraction is to be performed. The target video can be acquired through a video input module. The video input module can support the input of various video sources. For example, the target video can be a bird video, a security video, or a face or body video. The video scene can be a scene in which the target object in the target video performs an action. For example, if the target video is a bird video, the corresponding video scene can be a scene of birds flying, or a scene of birds foraging. The target object can be the object of attention of the target video. For example, if the target video is a video of a target user watching birds foraging, the corresponding at least one target object may include: target user, bird. The corresponding at least one video scene may include: a target user viewing scene and a bird foraging scene.
[0044] As an example, the execution entity may determine at least one video scene corresponding to the target video using a target object recognition model and a target object behavior detection model. The target object recognition model may be a neural network model that identifies objects that appear in the target video at a frequency higher than a target frequency. For example, the object recognition model may be a YOLO model. The target object behavior detection model may be a neural network model that generates object behaviors corresponding to the target object. The object behavior may be a series of actions performed by the target object. For example, the object behavior detection model may be a Transformer model.
[0045] Step 202: Generate a candidate frame set corresponding to at least one target object in the target video according to the at least one video scene.
[0046] In some embodiments, the execution subject may generate a set of candidate frames corresponding to at least one target object in the target video based on the at least one video scene. Each video scene may have at least one corresponding target object. That is, in one video scene, the extraction of key frames of at least one target object is supported. Similarly, the video scenes in at least one video scene may also have a one-to-one correspondence with the target objects in at least one target object. The candidate frames may be candidate frames to be further confirmed as key video content. That is, the candidate frames may be frames to be confirmed as key frames of a certain target object. The key frames may be frames including key object content corresponding to the target object. That is, the key frames may be frames in the target video that reflect more semantic content of the target object. For example, if the target object is a bird, the corresponding key frame may be a frame including more morphological content or behavioral content of the bird.
[0047] As an example, first, the target video is subjected to frame extraction processing to obtain a video frame sequence. Then, each video frame in the above video frame sequence is clustered to obtain a video frame cluster set. The number of video frame clusters in the video frame cluster set is the same as the number of video scenes corresponding to the above at least one video scene. Next, the target object corresponding to each video scene in the at least one video scene is determined to obtain at least one target object. Furthermore, the video frame corresponding to each cluster center is input into the object recognition model to obtain the object information corresponding to the cluster center. The object recognition model can be a method for determining the degree of object semantic matching between the video frame and the at least one target object. Then, for each video frame cluster, the video frames whose distance from the cluster center is higher than the target value are deleted from the above video frame cluster to obtain a deleted video frame cluster. Finally, the obtained deleted video frame cluster set is determined as the candidate picture set.
[0048] Step 203: Perform quality evaluation on each candidate picture in the candidate picture set to obtain picture quality information.
[0049] In some embodiments, the execution entity may perform a quality assessment on each candidate picture in the candidate picture set to obtain picture quality information. The picture quality information may represent the degree to which the candidate picture reflects the content corresponding to the target object. In practice, the picture quality information may be in the form of a score or text. For example, the picture quality information may be a numerical value between 0 and 100. A higher numerical value indicates that the candidate picture includes more and clearer semantic content of the object. For information in the form of text, the picture quality information may reflect the performance of the candidate picture under various indicators.
[0050] As an example, the execution entity may first determine the display angle of the target object in the candidate image. Then, determine the clarity of the target object in the candidate image. Next, determine the object content ratio of the target object in the candidate image. Furthermore, determine a first difference between the display angle and the target angle. Determine a second difference between the clarity and the target clarity. Determine a third difference between the object content ratio and the target object content ratio. Finally, perform a weighted sum of the first, second, and third differences to obtain image quality information.
[0051] In some optional implementations of some embodiments, the execution entity may perform quality evaluation on each candidate picture in the candidate picture set to obtain picture quality information, including the following steps:
[0052] The first step is to determine the target object corresponding to the candidate screen as the screen object. The screen object can be the main object displayed when the candidate screen appears. For example, for a screen where the candidate screen is a bird, the corresponding target object can be "bird".
[0053] The second step is to determine at least one image quality assessment metric corresponding to the aforementioned image object. The image quality assessment metric can be an indicator that assesses the quality of the content corresponding to the image. In practice, the image quality assessment metric can assess whether the image contains key content related to the target object. This image quality assessment metric can be used to assess the importance and criticality of the image content. Typically, the at least one image quality assessment metric may include clarity, object integrity, posture suitability, lighting uniformity, and background interference. It should be noted that different image quality assessment metrics may correspond to different image objects. For example, if the image object is a bird, the at least one image quality assessment metric may include bird integrity, posture diversity, image clarity, and lighting suitability. Bird integrity may include the visibility of key bird parts such as wings, head, and tail. Posture diversity may include various postures such as flight (wings outstretched), front-facing feeding, and a complete profile from the side. Image clarity can be comprehensively assessed using parameters such as contrast, sharpness, and noise. Lighting suitability can avoid adverse lighting conditions such as overexposure, underexposure, and strong reflections. For images where the subject is a person, the at least one image quality assessment metric may include: facial clarity, facial angle, body integrity, and lighting uniformity. Facial clarity indicates the degree of discernibility of facial details. Under the facial angle metric, the image quality for the front face is higher than that for the side face. The image quality for the side face is higher than that for the back face. Body integrity can be defined as the degree of visibility of the entire body or half of the body. Lighting uniformity can be defined as the uniformity of light and shadow distribution on the face. For images where the subject is a license plate, the at least one image quality assessment metric may include: vehicle integrity, vehicle clarity, and angle adaptability. Vehicle integrity can be defined as the degree of visibility of the main body components. License plate clarity can be defined as the clarity and recognizability of the license plate characters. Angle adaptability can be defined as the visibility score of the front and side body components. For images where the subject is a package, the at least one image quality assessment metric may include: package integrity and detail clarity. Package integrity can be defined as the degree of visibility of the package boundary. Detail clarity can be defined as the degree of visibility of logos and labels on the package surface. For pets, the at least one image quality assessment metric may include: pet integrity, feature salience, and posture naturalness. Pet integrity may include the visibility of key parts of the pet, such as its head and body. Feature salience may include the prominence of the pet's breed characteristics. Posture naturalness may include the naturalness of the pet's posture.
[0054] The third step is to determine, based on the candidate images, at least one indicator content corresponding to the at least one image evaluation indicator. The image evaluation indicator has corresponding indicator content. The indicator content can be the specific indicator value corresponding to the image evaluation indicator in the candidate images. In practice, different methods for extracting the corresponding indicator content vary depending on the image evaluation indicator.
[0055] As an example, the execution subject may determine, according to the indicator determination method corresponding to each picture evaluation indicator, the actual indicator value corresponding to the picture evaluation indicator reflected in the candidate picture as the indicator content.
[0056] The fourth step is to generate picture quality information corresponding to the candidate picture according to the at least one indicator content.
[0057] As an example, the execution entity may determine to combine at least one indicator content in a key-value pair format to obtain picture quality information.
[0058] As another example, in the first step, for each indicator content, the execution entity may first obtain a table for determining the corresponding indicator scores. Each picture evaluation indicator may have a corresponding, custom-defined table for determining the corresponding indicator scores. The table may represent the correspondence between each indicator content interval and the corresponding indicator score for the picture evaluation indicator. The table is then used to determine the indicator score corresponding to the indicator content. In the second step, the obtained indicator scores are weighted and summed to obtain picture quality information.
[0059] Step 204 : generating a key picture set for the at least one target object according to the picture quality information corresponding to each candidate picture and the candidate picture set.
[0060] In some embodiments, the execution entity may generate a key picture set for the at least one target object based on the picture quality information corresponding to each candidate picture and the candidate picture set. The key picture set includes at least one key picture subset corresponding to the at least one target object. That is, each target object corresponds to at least one key picture subset. The at least one key picture subset is a picture subset in the target video that can reflect the key content corresponding to the target object.
[0061] As an example, the execution subject may select candidate pictures whose corresponding picture quality information is within a target number from the candidate picture set to obtain a key picture set.
[0062] In some optional implementations of some embodiments, the execution entity may generate a key picture set for the at least one target object based on the picture quality information corresponding to each candidate picture and the candidate picture set, including the following steps:
[0063] In the first step, the candidate pictures in the candidate picture set are preliminarily screened based on the picture quality information corresponding to each candidate picture to obtain a first screened picture set. The pictures in the first screened picture set may be those whose corresponding picture quality information meets a target quality requirement. For example, the target quality requirement may be that the picture quality information is higher than the target quality information, or that the picture quality information is within the target number.
[0064] As an example, the execution subject may filter out candidate pictures whose corresponding picture quality information and scores are higher than a preset picture quality score from the candidate picture set to obtain a first filtered picture set.
[0065] As another example, the execution subject may filter out candidate pictures whose corresponding index contents of the key picture evaluation index of the object meet preset index contents from the candidate picture set to obtain a first filtered picture set.
[0066] The second step is to obtain at least one scene screening requirement corresponding to the at least one target object. Each target object has a corresponding scene screening requirement. The scene screening requirement can be a requirement for screening images based on the scene of the image. For example, if the target object is a bird, the corresponding scene screening requirements may include: prioritizing the extraction of images that fully display the bird's features, excluding images that only show a portion (such as the tail), are blurred or obscured, and retaining high-quality images from multiple angles for each bird to ensure diversity. For example, if the target object is a person, the corresponding scene screening requirements may include: prioritizing the extraction of full-body and half-body images with clear front views, excluding images of the back, severely side views, or mostly obscured human bodies, and prioritizing front views with clear facial features for face extraction. For example, if the target object is a vehicle, the corresponding scene screening requirements may include: prioritizing the extraction of fully visible front and side views of vehicles, and prioritizing the extraction of front views with clearly recognizable characters for license plates, excluding distant, blurred, partially obscured vehicle images, and images of stained and blurred license plates. If the target object is a package, the corresponding scene screening requirements may include: prioritizing the extraction of images of handheld or independently placed packages, ensuring clear package boundaries and complete shapes. If the target object is a pet, the corresponding scene screening requirements may include: giving priority to extracting images with clear and complete body and head features, and excluding blurred images due to poor lighting conditions, unclear posture or rapid movement.
[0067] In the third step, according to the at least one scene screening requirement, each picture in the first screening picture set is screened again to obtain a second screening picture set.
[0068] As an example, the execution subject may execute the requirement content corresponding to at least one scene screening requirement to re-screen each screen in the first screening screen set to obtain a second screening screen set.
[0069] In the fourth step, duplicate or similar images are removed from each of the images in the second filtered image set to obtain a removed image set. Duplicate images may be images with identical content. Similar images may be images with a content similarity higher than a target similarity. In practice, the similarity of semantic feature information between the images can be used as content similarity.
[0070] The fifth step is to adjust the time distribution of each picture in the above-mentioned removal picture set to obtain an adjusted picture set. Here, by adjusting the time distribution, the problem of excessive concentration of key pictures in the removal pictures can be avoided.
[0071] As an example, the execution subject may perform frame extraction on a subset of removed pictures in the removed picture set that has picture frame time clustering, so as to adjust the time distribution of each picture and obtain an adjusted picture set. Picture frame time clustering may be that a large number of pictures exist within a certain picture frame time period.
[0072] Step 6: Perform image enhancement on each of the images in the adjusted image set to obtain a key image set. Image enhancement may include, but is not limited to, at least one of the following: enhancing the clarity of the selected images, adjusting parameters such as brightness and contrast to optimize visual effects, and cropping to highlight the target object when necessary.
[0073] In some optional implementations of some embodiments, the execution entity may generate a key picture set for the at least one target object based on the picture quality information corresponding to each candidate picture and the candidate picture set, including the following steps:
[0074] In the first step, candidate pictures whose picture quality information corresponds to scores lower than the target quality information are removed from the candidate picture set to obtain a post-removal picture set and a removed picture set.
[0075] The second step is to perform clustering processing on the first candidate picture set to obtain a first candidate picture cluster set, wherein each candidate picture in the first candidate picture cluster has similar picture feature information.
[0076] The third step is to perform clustering processing on the removed picture set to obtain a second candidate picture cluster set, wherein each second candidate picture in the second candidate picture cluster has similar picture feature information.
[0077] In the fourth step, for each first candidate picture cluster in the first candidate picture cluster set, the following first determination step is performed:
[0078] Sub-step 1: determining a second candidate picture cluster in the second candidate picture cluster set corresponding to the first candidate picture cluster as a target candidate picture cluster.
[0079] Sub-step 2: determining a difference picture set between the first candidate picture cluster and the target candidate picture cluster.
[0080] Sub-step 3: determining at least one removed picture in the first candidate picture cluster.
[0081] Sub-step 4: for each removed frame in the at least one removed frame, perform the second determination step:
[0082] The first sub-step is to determine the picture semantic difference between the removed picture and each difference picture in the difference picture set, wherein the picture semantic difference can be a picture content semantic difference score.
[0083] As an example, first, semantic information of removed image features and semantic information of difference image features corresponding to the removed image and the difference image, respectively, are determined. Then, similarity between the semantic information of removed image features and the semantic information of difference image features is determined as image semantic similarity. Finally, the image semantic similarity is subtracted from 1 to obtain the image semantic difference.
[0084] In the second sub-step, the difference pictures whose semantic differences are less than the target difference value are removed from the difference picture set, and the target difference pictures are taken as the target difference pictures.
[0085] Sub-step 5: removing the target difference picture subset from the difference picture set to obtain a difference picture set after removal.
[0086] Sub-step 6: Fusing the removed difference picture set and the first candidate picture cluster to obtain a fused picture cluster.
[0087] In the fifth step, for each fused picture cluster in the obtained fused picture cluster set, a fused picture whose distance from the cluster center picture is less than a predetermined center distance is determined as a key picture.
[0088] Here, the above-mentioned "first step to fifth step" as one of the inventive points of the present disclosure solves "how to realize the screening of high-quality pictures (i.e., key pictures) from the candidate picture set based on picture quality". Based on this, the present disclosure clusters the first candidate picture set and the second candidate picture set respectively to determine the respective picture clusters. Then, the picture cluster differences are determined by comparing the two picture clusters. Next, the picture cluster differences are verified by the removed picture set to determine whether the picture cluster differences are effectively high-quality pictures. Based on this, the picture clusters generated in the case of higher-quality pictures (i.e., fused picture clusters) can be obtained. Finally, by the constraint of the distance from the cluster center, a high-quality key picture set can be accurately screened out from the fused picture cluster.
[0089] Step 205: Perform picture management on each key picture in the key picture set.
[0090] In some embodiments, the execution entity may manage each key screen in the key screen set, wherein the screen management may include: screen storage, screen enhancement, screen display, and screen deletion.
[0091] In some optional implementations of some embodiments, the execution entity may perform screen management on each key screen in the key screen set, including the following steps:
[0092] For each key screen in the key screen set, perform the following processing steps:
[0093] Sub-step 1: Determine the video source and frame timestamp corresponding to the key frame. The video source can be the video source of the target video corresponding to the key frame. The frame timestamp can be the frame time when the key frame appears in the target video.
[0094] Sub-step 2: Generate key picture metadata corresponding to the key picture based on the video source and the picture timestamp. The key picture metadata may be picture description information corresponding to the key picture. In practice, the key picture metadata may include: video source, timestamp, picture index mapping, scene type, and target object category. The picture index mapping may represent the index mapping relationship between the key picture and the target video.
[0095] Sub-step 3: Execute corresponding storage for the key images and the key image metadata. In practice, the key images and key image metadata can be stored according to object types.
[0096] And the above method also includes:
[0097] In response to receiving the picture retrieval information for the first target key picture, the video source and the picture position corresponding to the first target key picture are located according to the key picture metadata corresponding to the first target key picture.
[0098] In some optional implementations of some embodiments, the execution entity may perform screen management on each key screen in the key screen set, including the following steps:
[0099] The first step is to create a customized atlas for each key screen in the above-mentioned key screen set in response to the execution of atlas creation, so as to obtain various atlases adapted to different scene requirements. Among them, the above-mentioned atlases support the copying, moving and sharing of key screens. Among them, each target object supports the existence of a corresponding atlas. The creation of a customized atlas can be to aggregate key screens according to customized features to generate an atlas under the corresponding scene requirements. For example, the customized features can be to create atlases according to different objects, or to create atlases according to time periods, or to create atlases according to events. Scenario requirements can be object requirements, event requirements, or time period requirements.
[0100] In the second step, in response to executing content sharing, the target atlas and / or the second target key screen determined by the recommendation object are shared to the target terminal. Among them, the second target key screen can be the atlas cover corresponding to the target atlas corresponding to the atlas. In practice, the second target key screen can be a key screen selected by the user from the key image set, or it can be a key screen selected by the priority rule. The priority rule can be set by the user. For example, the priority rule can be "the corresponding screen for people is better than the corresponding screen for vehicles, the corresponding screen for vehicles is better than the corresponding screen for pets, and the superior the picture quality, the better." The recommendation object can be the object of the recommended key screen.
[0101] In a third step, in response to executing the batch export of key pictures, the at least one atlas and / or the at least one key picture are exported in batches.
[0102] In step 4, in response to the key screen anomaly annotation, at least one key screen in the key screen set that represents a predefined abnormal event is identified as at least one abnormal screen. A predefined abnormal event may be an event containing abnormal content. Abnormal events can be customized.
[0103] The fifth step is to mark the abnormal label to at least one of the abnormal screens to support early warning and backtracking of abnormal events.
[0104] In step 6, in response to executing the key screen recommendation, at least one key screen from the key screen set corresponding to the feature preference is recommended to the recommendation subject based on the feature preference of the recommendation subject. Feature preferences may be features preferred by the recommendation subject. For example, if the recommendation subject prefers a target type of vehicle, at least one key screen corresponding to the target type of vehicle will be recommended to the recommendation subject's corresponding terminal.
[0105] In some optional implementations of some embodiments, after step 205, the steps further include:
[0106] In the first step, in response to receiving at least one screen search parameter entered on the user interaction page, at least one corresponding key screen is searched from the key screen set according to the at least one screen search parameter as at least one matching screen. The at least one screen search parameter includes at least one of the following: screen quality information, number of key screens, time range, and scene requirements.
[0107] In the second step, the at least one matching scene is displayed in a target display format on the user interaction page. The user interaction page also displays the association between the key scene set and the target video. The user interaction page supports generation criteria for key scene correspondence. The generation criteria for key scene correspondence can be extraction rules for extracting key scene sets from a video.
[0108] Optionally, the execution entity may further support displaying a set of key images according to a timeline, support zooming and adjusting according to a time span, and provide a function for quickly locating key moments.
[0109] The aforementioned embodiments of the present disclosure have the following beneficial effects: Through the key frame processing methods of some embodiments of the present disclosure, key frames for at least one video scene can be accurately extracted from a target video, and efficient key frame management can be achieved. Specifically, the reason for the lack of accuracy and efficiency of the relevant key frames is that deep learning models such as convolutional neural networks are used to identify target objects in videos. However, most of these models are designed for a single scene and lack the ability to adapt to multiple scenes. Furthermore, they lack effective quality assessment and screening mechanisms after identification, resulting in inaccurate and inefficient extraction of key information from the video. Based on this, the key frame processing methods of some embodiments of the present disclosure first determine at least one video scene corresponding to the acquired target video. Here, determining the at least one video scene facilitates the subsequent determination of at least one target object for key frame extraction. Then, based on the at least one video scene, a set of candidate frames corresponding to the at least one target object in the target video can be accurately generated, serving as a preliminary screening set of key frames. Next, quality assessment is performed on each candidate frame in the candidate frame set to obtain image quality information. Through this quality assessment, key frames with more critical content and greater relevance to the at least one target object can be screened from the candidate frame set. Furthermore, based on the image quality information corresponding to each candidate frame and the candidate frame set, a key frame set for the at least one target object can be accurately generated, ensuring that the extracted key frames are of high value. This allows for precise determination of the key content of at least one target object (i.e., one or more target objects) in the target video. Finally, by managing the individual key frames in the key frame set, efficient extraction and utilization of the key frame set can be achieved, enabling a more efficient understanding of the key frame content in the target video.
[0110] Further references Figure 3 , shows a process 300 of another embodiment of the key picture processing method according to the present disclosure. The key picture processing method includes the following steps:
[0111] Step 301: Determine at least one video scene corresponding to the acquired target video.
[0112] Step 302: Determine the at least one target object according to the at least one video scene.
[0113] In some embodiments, the execution entity (e.g. Figure 1 The electronic device 101) shown can determine the at least one target object based on the at least one video scene.
[0114] Step 303: Obtain a pre-trained multi-object detection model corresponding to the at least one target object.
[0115] In some embodiments, the execution entity may obtain a pre-trained multi-object detection model corresponding to the at least one target object. The multi-object detection model may be an object detection model that supports simultaneous detection of multiple objects. For example, the multi-object detection model may be a neural network model that simultaneously supports detection of multiple object types, such as birds, faces, bodies, vehicles, license plates, packages, and pets.
[0116] Step 304: segment the target video to obtain a sub-video sequence.
[0117] In some embodiments, the execution entity may perform video segmentation processing on the target video to obtain a sub-video sequence.
[0118] As an example, the execution entity may perform video segmentation processing on the target video according to a preset video time window to obtain a sub-video sequence.
[0119] Step 305 : For each sub-video in the sub-video sequence, use the multi-object detection model to determine at least one object motion information corresponding to at least one target object in the sub-video.
[0120] In some embodiments, the execution entity may use the multi-object detection model to determine, for each sub-video in the sub-video sequence, at least one piece of object motion information corresponding to at least one target object in the sub-video. Each target object may have corresponding object motion information. The object motion information may be a motion trajectory corresponding to the target object in the sub-video.
[0121] As an example, first, for each frame in the sub-video, the image is input into the multi-object detection model to obtain the position information corresponding to at least one target object. Then, for each target object, the corresponding object motion information is generated based on the corresponding position information sequence.
[0122] Step 306: Generate the candidate picture set according to the obtained at least one object motion information sequence.
[0123] In some embodiments, the execution entity may generate the candidate picture set according to the obtained at least one object motion information sequence.
[0124] As an example, the execution entity may determine pictures related to at least one object motion information sequence in the target video as candidate pictures to obtain a candidate picture set.
[0125] Step 307: perform quality evaluation on each candidate picture in the candidate picture set to obtain picture quality information.
[0126] Step 308 : Generate a key picture set for the at least one target object based on the picture quality information corresponding to each candidate picture and the candidate picture set.
[0127] Step 309: manage each key screen in the key screen set.
[0128] In some embodiments, the specific implementation of steps 301, 307-300 and the technical effects thereof can be referred to. Figure 2 Steps 201, 203-205 in the corresponding embodiment will not be repeated here.
[0129] from Figure 3 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 3 In some embodiments, the key frame processing method process 300 utilizes a multi-object detection model to simultaneously detect multiple objects within a single video frame, significantly improving target detection efficiency. Furthermore, video segmentation improves processing efficiency, enhances detection accuracy, and balances resources and performance. Furthermore, by extracting object motion information, the candidate frame set is initially and accurately determined, with the frame corresponding to the object motion information being identified as a key frame.
[0130] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a key picture processing device. These device embodiments are similar to Figure 2 Corresponding to the method embodiments shown, the key picture processing device can be specifically applied to various electronic devices.
[0131] like Figure 4 As shown, a key picture processing device 400 includes: a determination unit 401, a first generation unit 402, a quality assessment unit 403, a second generation unit 404, and a management unit 405. The determination unit 401 is configured to determine at least one video scene corresponding to an acquired target video; the first generation unit 402 is configured to generate a candidate picture set corresponding to at least one target object in the target video based on the at least one video scene; the quality assessment unit 403 is configured to perform quality assessment on each candidate picture in the candidate picture set to obtain picture quality information; the second generation unit 404 is configured to generate a key picture set based on the picture quality information corresponding to each candidate picture and the candidate picture set; and the management unit 405 is configured to perform picture management on each key picture in the key picture set.
[0132] In some optional implementations of some embodiments, the first generation unit 402 may be further configured to: determine the at least one target object based on the at least one video scene; obtain a pre-trained multi-object detection model corresponding to the at least one target object; perform video segmentation processing on the target video to obtain a sub-video sequence; for each sub-video in the sub-video sequence, determine at least one object motion information corresponding to at least one target object in the sub-video using the multi-object detection model; and generate the candidate picture set based on the obtained at least one object motion information sequence.
[0133] In some optional implementations of some embodiments, the above-mentioned quality assessment unit 403 can be further configured to: determine the target object corresponding to the above-mentioned candidate picture as the picture object; determine at least one picture quality assessment index corresponding to the above-mentioned picture object; determine at least one indicator content corresponding to the above-mentioned at least one picture assessment index based on the above-mentioned candidate picture; and generate picture quality information corresponding to the above-mentioned candidate picture based on the above-mentioned at least one indicator content.
[0134] In some optional implementations of some embodiments, the second generation unit 404 may be further configured to: preliminarily screen each candidate picture in the candidate picture set according to the picture quality information corresponding to each candidate picture to obtain a first screened picture set; obtain at least one scene screening requirement corresponding to the at least one target object; re-screen each picture in the first screened picture set according to the at least one scene screening requirement to obtain a second screened picture set; deduplicate or remove similar pictures in each picture in the second screened picture set to obtain a removed picture set; adjust the time distribution of each picture in the removed picture set to obtain an adjusted picture set; and perform image enhancement on each picture in the adjusted picture set to obtain a key picture set.
[0135] In some optional implementations of some embodiments, the management unit 405 may be further configured to: for each key picture in the key picture set, perform the following processing steps: determine the video source and picture timestamp corresponding to the key picture; generate key picture metadata corresponding to the key picture based on the video source and the picture timestamp; perform corresponding storage for the key picture and the key picture metadata; and the method further includes: in response to receiving picture retrieval information for a first target key picture, locate the video source and the corresponding picture position corresponding to the first target key picture based on the key picture metadata corresponding to the first target key picture.
[0136] In some optional implementations of some embodiments, the management unit 405 may be further configured to: in response to executing atlas creation, create a customized atlas for each key screen in the key screen set to obtain each atlas adapted to different scenario requirements, wherein each atlas supports copying, moving and sharing of key screens; in response to executing content sharing, share the target atlas and / or the second target key screen determined by the recommendation object to the target terminal; in response to executing batch export of key screens, export the at least one atlas and / or at least one key screen in batches; in response to executing key screen abnormality marking, determine at least one key screen in the key screen set that represents the existence of a predefined abnormal event as at least one abnormal screen; mark the abnormal label to the at least one abnormal screen to support early warning and backtracking of abnormal events; in response to executing key screen recommendation, recommend at least one key screen in the key screen set corresponding to the feature preference to the recommendation object according to the feature preference corresponding to the recommendation object.
[0137] In some optional implementations of some embodiments, the apparatus 400 further includes: a search unit and a display unit (not shown in the figure). The search unit may be configured to: in response to receiving at least one screen search parameter filled in on the user interaction page, search for at least one corresponding key screen from the key screen set according to the at least one screen search parameter as at least one matching screen, wherein the at least one screen search parameter includes at least one of the following: screen quality information, number of key screens, time range, and scene requirements. The display unit may be configured to: display the at least one matching screen on the user interaction page in a target display format, the user interaction page also displays the association relationship between the key screen set and the target video, and the user interaction page supports the generation standard corresponding to the key screen.
[0138] It is understandable that the units described in the key image processing device 400 are similar to those described in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the key picture processing device 400 and the units included therein, and will not be described in detail here.
[0139] Reference below Figure 5 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of the electronic device 101)500. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0140] like Figure 5 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory 502 or a program loaded from a storage device 508 into a random access memory 503. Various programs and data required for the operation of the electronic device 500 are also stored in the random access memory 503. The processing device 501, the read-only memory 502, and the random access memory 503 are connected to each other via a bus 504. An input / output interface 505 is also connected to the bus 504.
[0141] Typically, the following devices may be connected to the input / output interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0142] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 509, or installed from the storage device 508, or installed from the read-only memory 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0143] It should be noted that in some embodiments of the present disclosure, the computer-readable medium mentioned above may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0144] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0145] The computer-readable medium may be included in the electronic device, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: determines at least one video scene corresponding to the acquired target video; generates a set of candidate frames corresponding to at least one target object in the target video based on the at least one video scene; performs a quality assessment on each candidate frame in the candidate frame set to obtain frame quality information; generates a set of key frames for the at least one target object based on the frame quality information corresponding to each candidate frame and the candidate frame set; and performs frame management on each key frame in the key frame set.
[0146] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0148] The units described in some embodiments of the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, they may be described as follows: a processor includes a determination unit, a first generation unit, a quality assessment unit, a second generation unit, and a management unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the determination unit may also be described as a "unit for determining at least one video scene corresponding to the acquired target video."
[0149] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0150] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned key picture processing methods when executed by a processor.
[0151] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A key image processing method, comprising: Determining at least one video scene corresponding to the acquired target video; generating, according to the at least one video scene, a candidate frame set corresponding to at least one target object in the target video; Performing quality evaluation on each candidate picture in the candidate picture set to obtain picture quality information; generating a key picture set for the at least one target object according to the picture quality information corresponding to each candidate picture and the candidate picture set; Screen management is performed on each key screen in the key screen set.
2. The method according to claim 1, wherein The step of generating a candidate picture set corresponding to at least one target object in the target video according to the at least one video scene includes: determining the at least one target object according to the at least one video scene; Obtaining a pre-trained multi-object detection model corresponding to the at least one target object; Performing video segmentation processing on the target video to obtain a sub-video sequence; For each sub-video in the sub-video sequence, determining, using the multi-object detection model, at least one object motion information corresponding to at least one target object in the sub-video; The candidate picture set is generated according to the obtained at least one object motion information sequence.
3. The method according to claim 1, wherein The performing quality evaluation on each candidate picture in the candidate picture set to obtain picture quality information includes: Determine a target object corresponding to the candidate picture as the picture object; Determining at least one picture quality assessment indicator corresponding to the picture object; Determining, based on the candidate screen, at least one indicator content corresponding to the at least one screen evaluation indicator; The picture quality information corresponding to the candidate picture is generated according to the at least one indicator content.
4. The method according to claim 1, wherein Generating a key picture set for the at least one target object according to the picture quality information corresponding to each candidate picture and the candidate picture set includes: performing preliminary screening of each candidate picture in the candidate picture set according to picture quality information corresponding to each candidate picture to obtain a first screened picture set; Obtaining at least one scenario screening requirement corresponding to the at least one target object; According to the at least one scene screening requirement, each screen in the first screening screen set is screened again to obtain a second screening screen set; performing duplicate or similar removal on each picture in the second filtered picture set to obtain a removed picture set; Performing time distribution adjustment on each picture in the removed picture set to obtain an adjusted picture set; Perform image enhancement on each picture in the adjustment picture set to obtain a key picture set.
5. The method according to claim 1, wherein The performing picture management on each key picture in the key picture set includes: For each key picture in the key picture set, the following processing steps are performed: Determine the video source and picture timestamp corresponding to the key picture; generating key picture metadata corresponding to the key picture according to the video source and the picture timestamp; Performing corresponding storage for the key picture and the key picture metadata; And the method further comprises: In response to receiving the picture retrieval information for the first target key picture, the video source and the corresponding picture position of the first target key picture are located according to the key picture metadata corresponding to the first target key picture.
6. The method according to claim 1, wherein The performing picture management on each key picture in the key picture set includes: In response to executing the atlas creation, creating a customized atlas for each key screen in the key screen set to obtain each atlas adapted to different scene requirements, wherein each atlas supports copying, moving, and sharing of key screens; In response to executing content sharing, sharing the target atlas and / or the second target key screen determined by the recommendation object to the target terminal; In response to executing the batch export of key pictures, exporting the at least one atlas and / or the at least one key picture in batches; In response to executing the key screen abnormality marking, determining at least one key screen in the key screen set that represents the presence of a predefined abnormal event as at least one abnormal screen; Marking an abnormality label to the at least one abnormal screen to support early warning and backtracking of abnormal events; In response to executing key picture recommendation, at least one key picture in the key picture set corresponding to the feature preference is recommended to the recommendation object according to the feature preference corresponding to the recommendation object.
7. The method according to claim 1, wherein The method further comprises: In response to receiving at least one picture search parameter entered on the user interaction page, searching for at least one corresponding key picture from the key picture set according to the at least one picture search parameter as at least one matching picture, wherein the at least one picture search parameter includes at least one of the following: picture quality information, number of key pictures, time range, and scene requirements; the user interaction page further displays an association between the key picture set and the target video; and the user interaction page supports generation criteria corresponding to key pictures; The at least one matching picture is displayed on the user interaction page in a target display form.
8. A key picture processing device, comprising: a determining unit configured to determine at least one video scene corresponding to the acquired target video; A first generating unit is configured to generate a candidate frame set corresponding to at least one target object in the target video according to the at least one video scene; a quality assessment unit configured to perform quality assessment on each candidate picture in the candidate picture set to obtain picture quality information; a second generating unit configured to generate a key picture set according to picture quality information corresponding to each candidate picture and the candidate picture set; The management unit is configured to perform picture management on each key picture in the key picture set.
9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.