Method and system for dynamically analyzing, modifying and distributing digital images and videos

By identifying elements and related characteristics in video frames, 3D environment maps are generated, and video scenes are modified based on the mapping, the problem of automatic identification and tracking in video is solved, and fast and effective video content modification and customization is achieved.

CN119996700APending Publication Date: 2025-05-13PANDOODLE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510130525.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-09-04
Filing Date
2019-09-04
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In videos, it is difficult to automatically identify and track scenes, areas, objects, and features, and prior art is difficult to modify and customize video content quickly and effectively.

Method used

By identifying elements in each frame of the video, comparing the relevant characteristics between elements, generating a 3D environment map, and modifying the scenes in the video based on the mapping, the function of quickly replacing or removing video elements is achieved.

Benefits of technology

It realizes rapid and efficient analysis, modification and distribution of digital images or videos, and can perform high-speed and distributed replacement of objects or areas at near real-time speeds, reducing server load and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996700A_ABST
    Figure CN119996700A_ABST
Patent Text Reader

Abstract

A new method for analyzing, modifying and distributing digital images and videos in a fast, efficient, practical and / or cost-effective manner is disclosed. A method of processing a video may employ a different region or object to replace pixels in a frame of a scene having features and characteristics of the identified region or object with another set of pixels. Replacement or other customization of frames and scenes can produce naturally fused videos or images that are indistinguishable by human eyes or other visual systems. In one embodiment, the invention may be used to provide different advertisement elements to different viewers into an image or group of images, or to enable viewers to control elements in a video and add their own preferences or other elements.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In a typical video, the human eye can recognize certain scenes, areas, objects, and features in a variety of ways. Identifying and tracking these same scenes, areas, objects, and features in an automated manner is difficult because multiple different characteristics need to be observed, recognized, and tracked. However, by identifying a scene, area, or object and all the relevant characteristics of that particular scene, area, or object, it is possible to replace the actual pixels in all frames of all scenes that have all the features and characteristics of the identified area or object with another set of pixels that look like they belong to the original frame and scene with a different area or object, so that the human eye or other visual system cannot distinguish them. This can be used, for example, to provide different advertising elements to different viewers in an image or group of images, or to enable viewers to control elements in a video and add their own preferences or other elements. Summary of the invention

[0002] In one embodiment, the video processing method of the present invention is characterized by: (i) identifying one or more elements in each frame of the video; (ii) identifying one or more scenes of the video by comparing elements in each frame with elements in the previous frame and the next frame, wherein frames having a common number of elements greater than a threshold number are considered to be in the same scene; (iii) obtaining one or more relevant characteristics of each element in each frame; (iv) generating a map of the 3D environment in each frame based on relevant features in one or more previous frames and one or more subsequent frames; and (v) modifying one or more scenes in the video based on the mapping.

[0003] In one embodiment, the method disclosed in the present invention is further characterized in that the element in step (i) is an object or a selected area in the scene of the video, and one or more elements in step (i) are identified by comparing with characteristics stored in an object database; wherein the element is automatically detected by a detection algorithm stored in a detection algorithm database, or selected by user input.

[0004] In one embodiment, the method disclosed in the present invention further comprises step (ii), wherein two or more of the scenes are associated via elements in each of the scenes and stored in a scene database.

[0005] In one embodiment, the relevant characteristics in step (iii) include, but are not limited to, position, size, reflection, lighting, shadow, warping, rotation, blur and occlusion.

[0006] In one embodiment, step (v) comprises modifying the one or more scenes by removing the one or more elements and applying the mapping generated in step (iv), averaging the one or more removed elements in each frame within the one or more scenes.

[0007] In one embodiment, step (v) comprises modifying the one or more scenes by warping the desired elements and applying the mapping generated in step (iv) to the desired elements in each frame of the one or more scenes.

[0008] In one embodiment, the method disclosed in the present invention further comprises delivering the modified video in step (v) by streaming or downloading.

[0009] In one embodiment, the present invention provides a method for analyzing, modifying and distributing digital images or videos in a fast, efficient, practical and / or cost-effective manner. In one embodiment, the present invention divides the video into scenes and frames, which can be preprocessed separately and in parallel, and then associated with each other by establishing relationships between the identified objects, regions, frames, scenes and their related metadata. In another embodiment, the system and method are configured to be able to identify scenes in the video and associate them. As another embodiment, objects, regions or parts of objects or regions and some of their features, such as lighting, shadows and / or occlusions are used to calculate a set of algorithms for each pixel, and can be applied to quickly replace or remove them in a customized manner. In another embodiment, element recognition algorithms are used to identify elements within each frame and determine how they are associated with each other. Algorithms for identifying objects, regions or other elements include but are not limited to DRIFT, KAZE, SIFT (scale invariant feature transformation), SURF (accelerated robust features), haar classifiers and FLANN (fast approximate nearest neighbor search library). In a further embodiment, objects and regions in different frames determined to belong to the same scene will be stored in a scene database and can be further used for subsequent recognition. The scene database can specify how scenes are related to each other and store all the information in each scene. In another embodiment, the scene processing server is used to intelligently pass the scene to the scene work node, where the scene can be processed in groups for fast processing. In some embodiments, the characteristics determined in the overall frame can be used to create different types of mappings and generate overall object mappings. In order to detect objects at high speed in the original video and quickly replace the customized video, an identification database can be created to store all information. In some embodiments, the recognition database as described above is further classified into "subsets", which allows single frames containing millions of objects to be quickly processed for multiple nodes, in order to play back customized videos at close to standard buffering and playback speeds. Another aspect of the present invention relates to algorithms for collecting and training image data sets, and can be used for scene preprocessing, frame preprocessing, and replacement or removal stages. The algorithm includes but is not limited to PICO, haar classifiers, and supervised learning. Another aspect of the present invention involves creating a 3D spatial mapping of each frame consisting of all objects, regions, light sources, shadows, occluded objects, and contexts. Another aspect of the present invention allows users to select objects or regions to be replaced or removed. Another aspect of the present invention involves high-speed, distributed object or region replacement in n frames, whereby the alternation process can be close to real-time. Another aspect of the invention allows replacement items to be inserted when pre-downloading the video, which is preferred when no custom insertion is needed but low server load and cost are required. For another embodiment, the video can be split into 1+the number of replaceable elements, whereby the video can retain n non-customized parts and only the customized parts need to be encoded. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In the following detailed description and accompanying drawings, various embodiments of the invention are disclosed.

[0011] Figure 1 is a flow chart illustrating one way in which the present invention may be used to identify scenes, objects, and regions in order to subsequently replace certain objects and regions in all scenes in which they are found.

[0012] Figure 2 It shows how to pre-process scenes and how to associate multiple scenes based on similar characteristics.

[0013] Figure 3 A distributed computing architecture for rapidly replacing elements in a video is shown in order to buffer, stream, encode, or any combination thereof.

[0014] Figure 4 The replacement of elements in the video frame is shown.

[0015] Figure 5 The removal of elements in a video frame is shown.

[0016] Figure 6 A pixel replacement map is shown that is associated with each object or region found in a scene or a group of scenes. DETAILED DESCRIPTION

[0017] The present invention may be implemented in a variety of ways, including as a process; an apparatus; a system; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on or provided by a memory coupled to the processor. In this specification, these embodiments, or any other form that the present invention may take, may be referred to as techniques. In general, the order of steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise stated, a component described as being configured to perform a task, such as a processor or memory, may be implemented as a general-purpose component temporarily configured to perform a task at a given time, or as a specific component manufactured to perform the task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processor cores configured to process data such as computer program instructions.

[0018] The following provides a detailed description of one or more embodiments of the present invention and drawings illustrating the principles of the present invention. The present invention is described in conjunction with these embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses many alternatives, modifications and equivalents. Many specific details are set forth in the following description to provide an in-depth "understanding" of the present invention. These details are provided for exemplary purposes, and the present invention can be practiced according to the claims without some or all of these specific details. For the sake of clarity, technical materials known in the technical field related to the present invention are not described in detail, so as not to unnecessarily obscure the present invention.

[0019] The present invention relates to a system and method in which a video can be broken down into scenes that may be related to each other, as well as objects and regions that exist across single or multiple scenes. By doing so, the present invention allows a user to "understand" the context of a video, or to pre-process the identified video scenes and frames to quickly, effectively, realistically and / or inexpensively customize the video by using the identified objects and regions and other types of metadata associated with the objects or regions or scenes themselves. These elements can then be replaced, altered or removed.

[0020] In some embodiments of the present invention, a system and method for identifying and tracking all or part of a scene, object or region in a video is described. In some embodiments, the method is configured to identify scenes that are related to each other in a video. In some embodiments, in each related scene, an object or a portion of an object or a region or a portion of a region is identified together with related characteristics such as lighting, shadows, and occlusions, and is used to calculate a set of algorithms for each pixel of each object or region, which can be used to quickly replace all pixels and their related characteristics in each object or region. In some embodiments, the method is used to allow a user or machine to replace the identified objects or regions in each scene in the video with a logo, object, or replacement image, so that the final logo, object, or replacement image looks like it has always been there. In some embodiments, the method is used to allow a user or machine to remove an object or region so that it looks like it has never appeared. In some embodiments, the method is used to reconstruct a 3D spatial map for each frame.

[0021] Figure 1 A flow chart is shown illustrating a system for capturing scene, object, region, and other relevant metadata associated with a video and using that information to generate object and region maps and subsequently customize and distribute the video. Figure 1, the system first analyzes the original video, which contains naturally occurring elements captured during the original shooting of the video. In one embodiment of the present invention, scenes and frames can be analyzed in parallel. When both scenes and frames are analyzed and all information has been pre-processed, the present invention can provide metadata around the context of the video and can provide the best suggested replacement areas to the user or other computer program. Once the replacement area is selected for replacement or removal or change, the object or area can be quickly changed based on the pre-processing information and can be sent to different processing mechanisms for buffering or streaming with individually changed frames, or can be queued until all customized frames have been processed and sent to different processing mechanisms such as encoding.

[0022] In some embodiments of the present invention, a system and method for analyzing and associating scenes is described. A scene is one or more frames that are classified as being related in some way in a video. Frames in a video are analyzed in a scene preprocessing phase, where an element recognition algorithm is used to identify elements in each frame of the video to determine which frames are related to each other. The algorithm identifies objects, areas of similar pixels, sequences of continuous actions, lighting, locations, and other elements that can be compared between frames. For example, a car chase sequence can be identified by identifying two cars and the characteristics of each car (color, type, brand), the driver of each car, the surrounding location where the car is traveling, and other identifiable elements in consecutive frames, and assigning a weight to each identifiable object, area, or characteristic so that it can be compared with the previous or next frame. In a different example, bedroom locations can be automatically detected by identifying furniture and related characteristics of each furniture (e.g., color, type, brand, scratches), as well as other elements (e.g., artwork on the wall, carpets, doors, etc.). These objects or regions can be identified by various algorithms, including but not limited to DRIFT, KAZE, SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), haar classifiers, and FLANN (Fast Approximate Nearest Neighbor Search Library). When the number of common elements identified between two consecutive frames or two groups of frames decreases below a threshold number, the scene can be considered to have changed. In a normal scene sequence change, the number of common elements will decrease from a large number in one scene to zero in the next scene. When fading from one scene to another, or when the scene changes gradually, multiple groups of frames can be used to determine the transition point from one frame to another. For example, in the case where one scene begins to gradually fade into another scene, the element recognition algorithm begins to recognize fewer common elements in consecutive frames and recognizes more common elements in a new set of frames in the next scene. The transition point between scenes can be determined in a variety of ways, including the midpoint of the fade transition, which is determined by the number of frames between the last frame of the first scene where the most common elements can be identified and the first frame of the second scene where the most common elements can be identified. For different elements in a scene, the gradual transition points from one scene to another can also be defined differently depending on when the element first fades in or out.

[0023] Once the comparison of objects and regions in the previous or next frame determines that the current frame belongs to a different scene, the previous scene and all its features can be stored in a database, such as Figure 2As shown. This data can be used to identify related scenes with the same characteristics at a later time through scene correlation by comparing to other non-sequential scenes that have data stored. For example, scene 1 may be found to be unrelated to scene 2, but scene 1 may be found to be related to scene 3 based on the same type of comparison that determined that scene 1 and scene 2 are unrelated, and the frames in scene 1 are related. By doing this, all related scenes in a video can be determined even if elements of the scenes are different. For example, if two cars are identified in a series of scenes, but the rest of the elements change, such as in a car chase, then the scenes will be related in different ways by the presence of two fast-moving cars with the same driver. A scene database is developed through this scene correlation to specify how scenes are related to each other, as well as to store all information about each object or area identified in each scene.

[0024] like Figure 3 As shown, the scene processing server is used to intelligently deliver scenes to scene working nodes. These scenes can be processed by continuous grouping, so that each scene can be sent to a node of a specific group for fast processing, because the algorithm requires n previous group and m next group frame data for calculation.

[0025] In some embodiments of the present invention, a system and method for analyzing frames in a video is described. Each frame in the video is analyzed through a frame preprocessing stage to automatically identify all objects and regions by comparing with a database of previously trained objects, regions, places, actions, and other representations, and by looking for continuous spatial regions by examining similar adjacent pixels. Figure 4 As shown, the analysis method improves upon existing methods by identifying / "understanding" and analyzing objects and areas in the video and comparing them to statistically appropriate locations depending on certain determining factors. In this way, the present invention can increase the chances of a good match for a particular item that a user wants to place, replace, or remove.

[0026] like Figure 4 and 5As shown, after identification, relevant characteristics of each object or area in the entire frame can be determined, such as lighting, shadows, warping, rotation, blur, and occlusion. For example, if a bottle is identified in the frame, the surrounding pixels can be detected to determine whether a shadow or reflection is cast, and this information can be used to help determine the light source. In another example, if a bottle is identified in the frame, and there is an object that occludes a portion of the bottle, the size and position of the entire bottle can be calculated, and this information can be used to calculate which part of the frame would be occupied by the bottle if the occluding object did not exist. In this example, the present invention can also detect pixels on the occluding object to determine whether a shadow or reflection is cast, and this information can also be used to help determine things such as light sources and deformations. In a third example, if the present invention determines that the entire frame represents a football game with two players on the field, and one player has a partially or mostly occluded object in his hand, then the player is most likely holding a football, and then the present invention can calculate which part of the ball is revealed and other characteristics such as shadows and deformations, so that this information can be used to change, replace, or remove an object later. Once the relationship between objects or regions and their associated characteristics has been established, each pixel that makes up the final overall region can be used to calculate certain values, such as color, luminosity, and hue. This can be used to create different types of maps that can then be associated with an object or region to produce an overall object map. By storing all of this information in an identification database, all objects and regions in a scene can be retrieved later and "understood" more quickly, and this information can be used to quickly replace parts or entire objects or regions, as the adjustments for each individual pixel in the replacement region have been pre-calculated and only a simple algorithm is needed to create the difference map. The end result is a perfectly blended, altered, replaced, or removed object or region that blends naturally into the scene during playback.

[0027] An identification database is a set of data used to identify specific objects, areas, actions (such as participating in a football game), places (such as cities), or environments (such as beaches). The system uses a variety of methods to collect this data for comparison and subsequent identification of specific objects, areas, actions, places, and environments. The identification database is divided into multiple specific sub-categories of object groupings, which have tags associated with them for identification. After simplifying the database or "dataset" to a specific data set, the method can search n data sets on a specific recognition work node very quickly (less than the time to create or render a frame). This allows the present invention to process a single frame for n nodes, each with their own data set, allowing the present invention to process millions of objects in the time it takes to process a single frame, so that video can be played back at a speed close to standard video buffering and playback.

[0028] Another aspect of the present invention relates to a tool for collecting and training image data sets of specific objects, which can be used for identification purposes in the preprocessing stage and the replacement or removal stage. By using image analysis and training algorithms, such as but not limited to PICO, haar classifiers and supervised learning, the tool can quickly collect and, if necessary, crop image data from locally stored image sets or the Internet by searching for key tags of the required images. After collection and cropping, the image data is converted into a trained metadata file, which can be placed on a server node and later used to identify specific items or groupings of items on a per-thread / node / server basis. In another embodiment, the collection and training process can be completed on a local computer, or can be split to a networked server for faster training. The tool allows testing for multiple data sets to ensure that the trained data set is able to function properly before it is stored in an identification database and deployed on a server.

[0029] Another aspect of the present invention is to employ high-speed detection of items in a video. This process uses methods from an identification database to assist the system by identifying information about the video that helps identify the interests of the viewer. For example, this process can be used to detect faces, logos, specific text, etc. As another example, the identified elements can be further customized by replacement, removal, or other modification. This user interest metadata is used in conjunction with other sets of information that define the user's interests, such as Figure 6 As shown, the present invention can be more specifically targeted at the replacement of objects and select objects that are of greater interest to specific viewers, thereby increasing the relevance to the viewers.

[0030] Another aspect of the present invention involves creating a 3D spatial map of each frame, which includes all objects, areas, light sources, shadows and occluded objects that have been identified, as well as the context of each frame. Since the present invention is able to identify objects, areas, locations, environments and other important data required to fully "understand" a 3D scene, such as shadows, lighting and occlusion, the present invention can reconstruct all or part of a 3D environment by using such data.

[0031] Another aspect of the present invention allows the user to select a replacement area to find specific frames in the video where they think there is a location that needs to be replaced. The algorithm searches through all relevant scenes and previous and subsequent frames to replace the entire area of ​​the video. The user can select an area that they wish to keep as a replacement area, or they can select a single point and allow the system to detect the extent of the replacement area based on the user's input / suggestions.

[0032] Another aspect of the invention relates to high-speed, distributed object or region replacement in n frames. Once the object or region is identified for replacement, modification or removal, and the object map is generated (which may or may not be done before the entire video preprocessing is completed), the system can identify n modification work nodes, which can work on each single frame, the object in the frame has been identified as existing and can be replaced, modified or removed, and each node can process for that specific object or region or a group of overlapping objects or regions. In this way, the modification process of m objects or regions can be close to real time.

[0033] Reference Figure 1 , another aspect of the present invention, which is labeled as "delivery method", allows the option of inserting replacement items when pre-downloading videos, rather than dynamically replacing, building, delivering and rebuilding videos. In the absence of customized insertion, this option is preferred and can reduce server load and cost. In a different embodiment, the video can be divided into 1+replaceable element number parts. In this way, it may not be necessary to re-encode the entire video, and n non-customized parts can be retained, and only the customized parts need to be encoded. In a different embodiment, the original video can be fully retained, and all customized objects or areas can be customized in another set of frames that are completely transparent except for the customized objects or areas. Then, this other set of frames can be synthesized onto the original set of frames, or sent as another set of frames or another video, which can be replayed by a video player that can synchronously play multiple video streams at the same time as the original, non-customized video.

[0034] In one embodiment, the video processing method of the present invention is characterized by: (i) identifying one or more elements in each frame of the video; (ii) identifying one or more scenes of the video by comparing elements in each frame with elements in the previous frame and the next frame, wherein frames having a common number of elements greater than a threshold number are considered to be in the same scene; (iii) obtaining one or more relevant characteristics of each element in each frame; (iv) generating a map of the 3D environment in each frame based on relevant features in one or more previous frames and one or more subsequent frames; and (v) modifying one or more scenes in the video based on the mapping.

[0035] In one embodiment, the method disclosed in the present invention is further characterized in that the element in step (i) can be an object or a selected area in the scene of the video, one or more elements in step (i) are identified by comparing with characteristics stored in an object database, and the elements can be automatically detected by a detection algorithm stored in a detection algorithm database, or selected by user input.

[0036] In one embodiment, the method disclosed in the present invention further comprises step (ii), wherein two or more of the scenes are associated via elements in each of the scenes and stored in a scene database.

[0037] In one embodiment, the relevant characteristics in step (iii) include, but are not limited to, position, size, reflection, lighting, shadow, warping, rotation, blur and occlusion.

[0038] In one embodiment, step (v) comprises modifying the one or more scenes by removing the one or more elements and applying the mapping generated in step (iv), averaging the one or more removed elements in each frame within the one or more scenes.

[0039] In one embodiment, step (v) comprises modifying the one or more scenes by warping the desired elements and applying the mapping generated in step (iv) to the desired elements in each frame of the one or more scenes.

[0040] In one embodiment, the method disclosed in the present invention further comprises delivering the modified video in step (v) by streaming or downloading.

[0041] In one embodiment, the computer-implemented system for processing video of the present invention may be, but need not necessarily be, characterized in that it includes the following steps: (i) identifying one or more elements in each frame of the video; (ii) identifying one or more scenes of the video by comparing elements in each frame with elements in the previous frame and the next frame, wherein frames having a common number of elements greater than a threshold number are considered to be in the same scene; (iii) obtaining one or more relevant characteristics of each element in each frame; (iv) generating a map of the 3D environment in each frame based on relevant features in one or more previous frames and one or more subsequent frames; and (v) modifying one or more scenes in the video based on the mapping.

[0042] In one embodiment, the computer-implemented system disclosed in the present invention is further characterized in that the elements in step (i) may be objects or selected areas in the scene of the video, and one or more elements in step (i) may be identified by comparing with characteristics stored in an object database; wherein the elements may be automatically detected by a detection algorithm stored in a detection algorithm database, or selected by user input.

[0043] In one embodiment, the method disclosed in the present invention further comprises step (ii), wherein two or more of the scenes are associated via elements in each of the scenes and stored in a scene database.

[0044] In one embodiment, the relevant characteristics in step (iii) include, but are not limited to, position, size, reflection, lighting, shadow, warping, rotation, blur and occlusion.

[0045] In one embodiment, step (v) may include modifying the one or more scenes by removing one or more elements and applying the mapping generated in step (iv), averaging the one or more removed elements in each frame within the one or more scenes.

[0046] In one embodiment, step (v) may include modifying the one or more scenes by warping the desired elements and applying the mapping generated in step (iv) to the desired elements in each frame of the one or more scenes.

[0047] In one embodiment, the computer-implemented system disclosed in the present invention may further include delivering the modified video in step (v) by streaming or downloading.

Claims

1. A computer-implemented method for processing content of a video comprising one or more frames, the method comprising: (i) identifying one or more elements in each frame of the video; (ii) identifying one or more scenes of the video by comparing elements in each frame with elements in the previous frame and the next frame, wherein frames having a common number of elements greater than a threshold number are considered to be in the same scene; (iii) obtaining one or more relevant characteristics of each element in each frame; (iv) generating a map of the 3D environment in each frame based on relevant features in one or more previous frames and one or more subsequent frames; and (v) modifying one or more scenes in the video based on the mapping.

2. The method according to claim 1, characterized in that: The one or more elements in step (i) are identified by comparison with characteristics stored in a database of objects.

3. The method according to claim 1, characterized in that: The element is an object or a selected area in the scene of the video.

4. The method according to claim 1, characterized in that: The elements are detected by a detection algorithm stored in a detection algorithm database, or selected by user input.

5. The method according to claim 1, characterized in that: The method further comprises associating two or more of the scenes of step (ii) by elements in each of the scenes, and storing the scenes in a scene database.

6. The method according to claim 1, characterized in that: The relevant properties include position, size, reflection, lighting, shadowing, rotation, warping, blur and occlusion.

7. The method according to claim 1, characterized in that: Step (v) comprises modifying the one or more scenes by removing one or more elements and applying the mapping, averaging the one or more removed elements in each frame within the one or more scenes.

8. The method according to claim 1, characterized in that: Step (v) comprises modifying the one or more scenes by warping the desired elements and applying the mapping to the desired elements in each frame of the one or more scenes.

9. The method according to claim 1, characterized in that: The method further includes delivering the modified video in step (v) by streaming or downloading.