Method and apparatus for stitching image frames including moving objects
By identifying and tracking the moving objects and their trajectories in the image frame, and masking and regenerating them during the stitching process, the problem of moving objects being displayed as static in the panoramic image in the prior art is solved, and the effect of correctly displaying the movement of moving objects and improving image quality is achieved.
Patent Information
- Application Number
- CN202280101344.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-29
- Filing Date
- 2022-12-15
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to correctly process moving objects when generating panoramic images, resulting in the moving objects being displayed as static in the panoramic images, and frame stitching will cause the moving objects to be blurred.
By identifying the moving objects and their motion trajectories in the image frame, the trajectory of the moving object is determined using normalized properties, and the masked part of the moving object is masked and regenerated during the stitching process to generate the correct stitching image.
It realizes the correct display of the motion of moving objects in the panoramic image, avoids blurring of moving objects caused by frame stitching, and improves the quality of generated panoramic images.
Smart Images

Figure CN120112937A_ABST
Abstract
Description
Technical Field
[0001] The present subject matter generally relates to generating stitched images in a panorama, and in particular, to methods and apparatus for generating stitched images by stitching image frames that include one or more moving objects. Background Art
[0002] In existing systems for creating panoramic images, moving objects can be identified so that image stitching can be performed correctly to produce aligned images. This causes problems with image stitching because moving objects overlap in the various frames. Since the moving objects appear across the overlapping areas, the moving objects in the panoramic image cannot produce the effect of motion when the panorama is captured, so the panoramic image is essentially static.
[0003] Panorama is a wide-angle view of photography. Multiple frames covering the wide angle depend on consecutive frames captured for the wide-angle view. The captured frames are stitched to make a single photo. Available panorama solutions are static, i.e. static background or stationary view.
[0004] Currently, there are many problems with the panoramic mode, for example, the generated panoramic image shows all objects as static, while the real view can have some moving objects, and frame stitching will cause blurring of moving objects.
[0005] When the captured frames have moving objects, a correct panoramic image cannot be generated. The reasons for the failure are that the trajectory or overlapping areas of the moving objects are not followed, the trajectory path is not masked and reconstructed, and the frames are stitched together to capture the panoramic image when the phone is in motion without adjusting based on the motion of the moving objects that appear across multiple frames.
[0006] A conventional solution discloses a method of stitching multiple frames to form a wide-angle view static image (static panorama). All objects, whether static or dynamic, are displayed as static in the final generated image.
[0007] Another conventional solution discloses a method for providing information and improving the quality of digital entertainment by using a panoramic video as a counterpart of image stitching. However, this other conventional solution also does not support displaying a moving object as an object in motion.
[0008] Shortcomings of conventional solutions include not following the trajectories or overlapping areas of moving objects, no masking and reconstruction of trajectory paths are performed, they cannot handle to produce liveliness in the image if more than one object is present and moving towards another, and stitching frames to capture a panoramic image while the phone is in motion does not adjust based on the motion of moving objects that appear across multiple frames.
[0009] A solution is needed to overcome the above disadvantages. Summary of the invention
[0010] This Summary is provided to introduce in a simplified form concepts that are further described in the Detailed Description of the present disclosure. This Summary is not intended to identify key inventive concepts or fundamental inventive concepts of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter. In accordance with the purposes of the present disclosure, the present disclosure, as embodied and broadly described herein, describes methods and systems for providing recommendations for maintaining personal hygiene of a user.
[0011] According to an embodiment of the present disclosure, a method for generating a stitched image by an electronic device is disclosed. The method may include: obtaining an input stream including a plurality of frames through an image capture device. The method may include: identifying one or more moving objects and one or more timestamps associated with the movement of the one or more moving objects from a first frame and a second frame selected from the plurality of frames. The method may include: determining a plurality of attributes associated with the one or more moving objects from at least one overlapping region of the first frame and the second frame. The method may include: normalizing the determined plurality of attributes relative to a plurality of device attributes associated with the image capture device. The method may include: determining, based on the normalized plurality of attributes, trajectories associated with the one or more moving objects and the movement region of the one or more moving objects in the first frame and the second frame. The method may include performing one of the following operations: stitching a first frame at a first position where one or more moving objects exist in at least one overlapping region, wherein the trajectory from the first frame is masked to regenerate a masked portion of the second frame; and stitching a first frame at a second position where one or more moving objects do not exist in at least one overlapping region, wherein the trajectory from the first frame is masked to regenerate a masked portion of the second frame. The method may include: generating a stitched image by stitching the first frame with the masked portion of the second frame.
[0012] According to an embodiment of the present disclosure, an electronic device for generating a stitched image is disclosed. The electronic device may include: a memory; and at least one processor coupled to the memory. The at least one processor may be configured to: obtain an input stream including a plurality of frames through an image capture device. The at least one processor may be configured to: identify one or more moving objects and one or more timestamps associated with the movement of the one or more moving objects from a first frame and a second frame selected from the plurality of frames. The at least one processor may be configured to: determine a plurality of attributes associated with the one or more moving objects from at least one overlapping region of the first frame and the second frame. The at least one processor may be configured to: normalize the determined plurality of attributes relative to a plurality of device attributes associated with the image capture device. The at least one processor may be configured to: determine, based on the normalized plurality of attributes, trajectories associated with the one or more moving objects and the movement regions of the one or more moving objects in the first frame and the second frame. At least one processor may be configured to perform one of the following operations: stitching a first frame at a first location of at least one overlapping region where one or more moving objects exist, wherein trajectories from the first frame are masked to regenerate a masked portion of a second frame; and stitching a first frame at a second location of at least one overlapping region where one or more moving objects do not exist, wherein trajectories from the first frame are masked to regenerate a masked portion of the second frame. At least one processor may be configured to generate a stitched image by stitching the first frame with the masked portion of the second frame.
[0013] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing instructions is disclosed. When the instructions are executed by at least one processor of an electronic device, the electronic device performs an operation. The operation may include: obtaining an input stream including multiple frames through an image capture device. The operation may include: identifying one or more moving objects and one or more timestamps associated with the movement of the one or more moving objects from a first frame and a second frame selected from the multiple frames. The operation may include: determining multiple attributes associated with the one or more moving objects from at least one overlapping area of the first frame and the second frame. The operation may include: normalizing the determined multiple attributes relative to multiple device attributes associated with the image capture device. The operation may include: determining trajectories associated with one or more moving objects and the movement area of one or more moving objects in the first frame and the second frame based on the normalized multiple attributes. The operation may include performing one of the following operations: splicing a first frame at a first position where one or more moving objects exist in at least one overlapping area, wherein the trajectory from the first frame is masked to regenerate the mask portion of the second frame; and splicing a first frame at a second position where one or more moving objects do not exist in at least one overlapping area, wherein the trajectory from the first frame is masked to regenerate the mask portion of the second frame. The operation may include generating a stitched image by stitching the first frame with the masked portion of the second frame.
[0014] These aspects and advantages will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A block diagram depicting a method of generating a stitched image by stitching image frames including one or more moving objects is shown according to an embodiment of the present subject matter;
[0016] Figure 2 A schematic block diagram of a system configured to generate a stitched image by stitching image frames including one or more moving objects according to an embodiment of the present subject matter is shown;
[0017] Figure 3 An operational flow chart depicting a process of generating a stitched image by stitching image frames including one or more moving objects in accordance with an embodiment of the present subject matter is shown;
[0018] Figure 4 An architectural diagram depicting a method of generating a stitched image by stitching image frames including one or more moving objects according to an embodiment of the present subject matter is shown;
[0019] FIG5 a shows a diagram depicting a method of selecting a first frame and a second frame from a plurality of frames according to an embodiment of the present subject matter;
[0020] FIG5 b shows an operational flow diagram depicting a process for selecting a first frame and a second frame from a plurality of frames according to an embodiment of the present subject matter; and
[0021] FIG5 c shows a diagram depicting a first stage and a second stage for selecting a first frame and a second frame according to an embodiment of the present subject matter;
[0022] Figure 6 An operational flow chart depicting a process for identifying one or more moving objects in a plurality of frames according to an embodiment of the present subject matter is shown;
[0023] Figure 7 An operational flow chart depicting a process for determining a plurality of properties of one or more moving objects in accordance with an embodiment of the present subject matter is shown;
[0024] Figure 8 An operational flow diagram depicting a process for tracking a path of one or more moving objects in accordance with an embodiment of the present subject matter is shown; and
[0025] Fig. 9 shows a diagram depicting trajectory generation according to an embodiment of the present subject matter;
[0026] Furthermore, those skilled in the art will recognize that the elements in the drawings are shown for the sake of brevity and may not necessarily be drawn to scale. For example, the flow chart illustrates the method in terms of the most prominent steps involved to help improve understanding of the various aspects of the present invention. Furthermore, with respect to the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details relevant to understanding the embodiments of the present invention so as not to obscure the drawings with details that would be apparent to one of ordinary skill in the art having the benefit of the description herein. DETAILED DESCRIPTION
[0027] For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiments illustrated in the drawings and specific language will be used to describe these embodiments. It will be understood, however, that the scope of the invention is not limited thereto and that such changes and further modifications in the illustrated systems and such further applications of the principles of the invention illustrated therein will be readily apparent to those skilled in the art to which the invention pertains.
[0028] Those skilled in the art will understand that both the foregoing general description and the following detailed description are illustrative of the present invention and are not intended to be restrictive thereof.
[0029] References throughout this specification to "one aspect," "another aspect," or similar language indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment," "in another embodiment," and similar language throughout this specification may (but do not necessarily) refer to the same embodiment.
[0030] The terms "comprises", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion such that a process or method that includes a list of steps includes not only those steps but may also include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or subsystems or elements or structures or components preceded by "comprising..." does not exclude the presence of other devices or other subsystems or other elements or other structures or other components or additional devices or additional subsystems or additional elements or additional structures or additional components, without further constraints.
[0031] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs.The systems, methods, and examples provided herein are illustrative only and not limiting.
[0032] For the sake of clarity, the first digit of the reference numeral of each component of the present disclosure indicates the figure number showing the corresponding component. Figure 1 , the reference numerals beginning with the number "1" are shown. Figure 2 The reference numerals beginning with the number "2" are shown in FIG.
[0033] Embodiments of the present subject matter will be described in detail below with reference to the accompanying drawings.
[0034] Figure 1 A block diagram depicting a method 100 for generating a stitched image by stitching image frames including one or more moving objects according to an embodiment of the present subject matter is shown. The method 100 may be implemented in an electronic device. Examples of electronic devices may include, but are not limited to, smartphones, laptops, personal computers (PCs), and tablets. The images and stitched images may be in a panoramic mode.
[0035] At block 102 , method 100 includes capturing, by an image capture device, an input stream of a frame sequence associated with a plurality of frames.
[0036] At block 104 , the method 100 includes identifying one or more moving objects and one or more time stamps associated with movement of the one or more moving objects from a first frame and a second frame selected from the plurality of frames.
[0037] At box 106, method 100 includes: determining multiple attributes associated with one or more moving objects from at least one overlapping region of the first frame and the second frame, wherein the multiple attributes include time that the one or more moving objects spend in the at least one overlapping region, and a frame rate associated with an input stream of the frame sequence.
[0038] At block 108 , the method 100 includes normalizing a plurality of properties captured from the at least one overlapping region relative to a plurality of device properties associated with a capture device that captured the plurality of frames, wherein the normalizing includes associating the plurality of properties with the plurality of device properties.
[0039] At block 110 , the method 100 includes determining tracks associated with one or more moving objects and moving regions of the one or more moving objects in the first frame and the second frame based on the normalized plurality of attributes.
[0040] At box 112, method 100 includes performing one of the following operations: splicing a first frame at a first location where one or more objects exist in at least one overlapping area, wherein the trajectory from the first frame is masked to regenerate the masked portion of the second frame; and splicing the first frame at a second location where one or more objects do not exist in at least one overlapping area, wherein the trajectory from the first frame is masked to regenerate the masked portion of the second frame.
[0041] At block 114 , the method 100 includes generating a stitched image by stitching the first frame with the masked portion of the second frame.
[0042] Figure 2 A schematic block diagram 200 of a system 202 configured to generate a stitched image by stitching image frames including one or more moving objects according to an embodiment of the present subject matter is shown. The method 100 can be implemented in an electronic device. Examples of electronic devices can include, but are not limited to, smartphones, laptops, personal computers (PCs), and tablets. The images and stitched images can be in a panoramic mode.
[0043] In one example embodiment, the system 202 may be a chip incorporated into an electronic device. In another example embodiment, the system 202 may be implemented software, a logic-based program, hardware, configurable hardware, etc. The system 202 includes a processor 204, a memory 206, data 208, a module 210, a resource 212, a capture engine 214, a recognition engine 216, a determination engine 218, a normalization engine 220, a trajectory determination engine 222, a stitching engine 224, and a generation engine 226.
[0044] Processor 204, memory 206, data 208, modules 210, resources 212, capture engine 214, recognition engine 216, determination engine 218, normalization engine 220, trajectory determination engine 222, stitching engine 224, and generation engine 226 may be communicatively coupled to one another.
[0045] In an example, the processor 204 may be a single processing unit or multiple units, all of which may include multiple computing units. The processor 204 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, processor cores, multi-core processors, multiprocessors, state machines, logic circuits, application specific integrated circuits, field programmable gate arrays, and / or any device that manipulates signals based on operational instructions. Among other capabilities, the processor 204 may be configured to retrieve and / or execute computer-readable instructions and / or data stored in the memory 206.
[0046] In an example, the memory 206 may include any non-transitory computer-readable medium known in the art, including, for example, volatile memory (such as static random access memory (SRAM) and / or dynamic random access memory (DRAM)) and / or non-volatile memory (such as read-only memory (ROM), erasable programmable ROM (EPROM), flash memory, hard disk, optical disk and / or magnetic tape). The memory 206 may include data 208. The memory 206 may store instructions. When the instructions are executed by the processor 204, the instructions may cause the electronic device 200 or the processor to perform the operations described herein.
[0047] Data 208 serves as, among other things, a repository for storing data processed, received, and generated by one or more of processor 204 , module 210 , resource 212 , capture engine 214 , recognition engine 216 , determination engine 218 , normalization engine 220 , trajectory determination engine 222 , stitching engine 224 , and generation engine 226 .
[0048] Module 210 may include, among other things, routines, programs, objects, components, data structures, etc. that perform specific tasks or implement data types. Module 210 may also be implemented as a signal processor, state machine, logic circuit, and / or any other device or component that manipulates signals based on operational instructions.
[0049] In addition, the module 210 can be implemented in hardware, instructions executed by at least one processing unit (e.g., processor 204), or a combination thereof. The processing unit can be a general-purpose processor that executes instructions to cause the general-purpose processor to perform operations, or the processing unit can be dedicated to perform the required functions. In another aspect of the present disclosure, the module 210 can be machine-readable instructions (software) that, when executed by a processor / processing unit, can perform any of the described functions.
[0050] In some example embodiments, module 210 may be machine readable instructions (software) that, when executed by a processor / processing unit, perform any of the described functions.
[0051] Resources 212 may be physical and / or virtual components of system 202 that provide intrinsic functionality and / or contribute to the performance of system 202. Examples of resources 212 may include, but are not limited to, memory (e.g., memory 206), power units (e.g., batteries), display units, etc. In addition to processor 204 and memory 206, resources 212 may also include power units / battery units, network units, etc.
[0052] Continuing with the above embodiment, the capture engine 214 can be configured to capture an input stream of a sequence of frames. The sequence of frames can be related to a plurality of frames captured by an image capture device. Examples of image capture devices can include, but are not limited to, cameras, smart phones, video recorders, and CCTV.
[0053] Next, the recognition engine 216 may be configured to: identify one or more moving objects and one or more time stamps associated with the movement of the one or more moving objects. The one or more moving objects and the one or more time stamps may be identified from a first frame and a second frame selected from the plurality of frames. To identify the one or more moving objects, the recognition engine 216 may be configured to: compare a plurality of second frame grids of the second frame with a plurality of first frame grids of the first frame in terms of pixel intensity. The pixel intensity is associated with the plurality of second frame grids and the plurality of first frame grids.
[0054] The recognition engine 216 may be configured to determine that pixel intensities associated with the plurality of second frame grids do not match pixel intensities associated with the plurality of first frame grids. The recognition engine 216 may also be configured to, based on the determination, identify one or more moving objects in the first frame and the second frame. Additionally, the first frame may be a previous frame relative to the current frame, and the second frame may be the current frame. To select the first frame and the second frame, the recognition engine 216 may be configured to perform a timestamp-based comparison of the plurality of frames relative to a quality metric for each frame.
[0055] In addition, each frame can be buffered with a timestamp associated with each frame in the plurality of frames. In addition, the recognition engine 216 can be configured to estimate the quality of each frame in the plurality of frames based on a timestamp-based comparison of the quality metric of each frame. The quality metric can be derived from a power spectral density (PSD) of each frame. In addition, the recognition engine 216 can be configured to select the first frame and the second frame in the plurality of frames based on the estimation.
[0056] To this end, in order to estimate the quality of each frame, the recognition engine 216 can be configured to process the plurality of frames by applying a plurality of machine learning (ML) techniques. The recognition engine 216 can be configured to calculate the PSD associated with each of the processed plurality of frames. The recognition engine 216 can be configured to select at least two frames having a PSD greater than a predetermined threshold from the plurality of frames based on density-based clustering and outlier elimination. The at least two frames can include a first frame and a second frame.
[0057] Continuing with the above embodiment, the determination engine 218 can be configured to determine a plurality of attributes associated with one or more moving objects from at least one overlapping region of the first frame and the second frame. The plurality of attributes can include the time that the one or more moving objects spend in the at least one overlapping region, and a frame rate associated with the input stream of the frame sequence.
[0058] To this end, the normalization engine 220 can be configured to normalize multiple attributes captured from at least one overlapping area. Examples of multiple attributes may include, but are not limited to, one or more of the following: relative motion of one or more moving objects, exchange area ratio of one or more moving objects, color of one or more moving objects, background color, size of one or more moving objects, frame rate, speed of one or more moving objects, and time spent by one or more moving objects in the first frame. Normalization can be performed relative to multiple device attributes associated with a capture device that captures multiple frames. Examples of multiple device attributes may include, but are not limited to, one or more of the following: speed of the image capture device, and direction of movement of the image capture device. Normalization may include associating multiple attributes with multiple device attributes. In addition, associating multiple attributes with multiple device attributes may include: changing the value of one or more of the multiple attributes relative to the value of the multiple device attributes.
[0059] Next, the trajectory determination engine 222 may be configured to determine, based on the normalized multiple attributes, trajectories associated with one or more moving objects and a moving region of the one or more moving objects in the first frame and the second frame. To determine the trajectory, the trajectory determination engine 222 may be configured to determine that the multiple attributes cause the one or more moving objects to move after being normalized. The trajectory determination engine 222 may be configured to detect a direction of motion of the one or more moving objects in the first frame and the second frame. The trajectory determination engine 222 may be configured to generate the trajectory based on downsampling and upsampling of the multiple attributes.
[0060] Thus, the stitching engine 224 can be configured to perform one of a plurality of stitching techniques. The plurality of stitching techniques can include stitching a first frame at a first location where one or more objects are present in at least one overlapping region. Trajectories from the first frame are masked to regenerate a masked portion of a second frame. The plurality of stitching techniques can also include stitching a first frame at a second location where one or more objects are not present in at least one overlapping region. Trajectories from the first frame can be masked to regenerate a masked portion of a second frame.
[0061] Furthermore, the generation engine 226 may be configured to generate a stitched image by stitching the first frame with the masked portion of the second frame.
[0062] The functionality of the engines including capture engine 214 , recognition engine 216 , determination engine 218 , normalization engine 220 , trajectory determination engine 222 , stitching engine 224 , and generation engine may be performed by processor 204 in conjunction with instructions stored in memory 206 for execution by the processor.
[0063] Figure 3 An operational flow chart depicting a process 300 for generating a stitched image by stitching image frames including one or more moving objects according to an embodiment of the present subject matter is shown. The process may be performed by a system 202 incorporated into an electronic device of a user. Generating the stitched image may be based on applying one or more ML techniques. Examples of one or more moving objects may include, but are not limited to, humans, animals, and vehicles.
[0064] At step 302, process 300 may include capturing an input stream of a frame sequence. The frame sequence may be associated with a plurality of frames captured by an image capture device. The input stream may be Figure 2 The image capture engine 214 captures the image mentioned above. In addition, multiple frames can be in a panoramic view.
[0065] At step 304, process 300 may include comparing the plurality of second frame grids of the second frame with the plurality of first frame grids of the first frame in terms of pixel intensity. The pixel intensities are associated with the plurality of second frame grids and the plurality of first frame grids. The comparison may be performed to identify one or more moving objects and one or more timestamps associated with the movement of the one or more moving objects. The identification may be performed by Figure 2 The recognition engine 216 mentioned in the above is executed. One or more moving objects and one or more time stamps can be recognized from a first frame and a second frame selected from a plurality of frames.
[0066] At step 306, process 300 may include determining that pixel intensities associated with the plurality of second frame grids do not match pixel intensities associated with the plurality of first frame grids, and identifying one or more moving objects in the first frame and the second frame based on the determination. Additionally, the first frame may be a previous frame relative to the current frame, and the second frame may be the current frame.
[0067] For the selection of the first frame and the second frame, process 300 may include performing a timestamp-based comparison of the plurality of frames relative to a quality metric of each frame. In addition, each frame may be buffered with a timestamp associated with each of the plurality of frames. Process 300 may also include estimating a quality of each of the plurality of frames based on the timestamp-based comparison of the quality metric of each frame. The quality metric may be derived from the PSD of each frame, and the first frame and the second frame may be selected from the plurality of frames based on the estimation.
[0068] To this end, in order to estimate the quality of each frame, process 300 may include: processing multiple frames by applying multiple machine learning (ML) techniques, and calculating the PSD associated with each of the processed multiple frames. Process 300 may also include: selecting at least two frames having a PSD greater than a predetermined threshold from the multiple frames based on density-based clustering and outlier elimination. The at least two frames may include a first frame and a second frame.
[0069] At step 308, process 300 may include determining a plurality of attributes associated with one or more moving objects from at least one overlapping region of the first frame and the second frame. The determination may be performed by Figure 2 The determination engine 218 mentioned in is executed.
[0070] At step 310, process 300 may include associating a plurality of attributes captured from at least one overlapping region with a plurality of device properties associated with a capture device that captured the plurality of frames. Associating may include changing values of one or more of the plurality of attributes relative to values of the plurality of device properties. Figure 2 The normalization engine mentioned in performs a join to normalize multiple attributes.
[0071] Examples of multiple attributes may include, but are not limited to, one or more of the following: relative motion of one or more moving objects, exchange area ratio of one or more moving objects, color of one or more moving objects, background color, size of one or more moving objects, frame rate, speed of one or more moving objects, and time spent by one or more moving objects in the first frame. Examples of multiple device attributes may include, but are not limited to, one or more of the following: speed of the image capture device, and direction of movement of the image capture device. Normalization may include associating multiple attributes with multiple device attributes. In addition,
[0072] At step 312, process 300 may include determining that the plurality of attributes, after being normalized, cause one or more moving objects to move. Process 300 may also include detecting a direction of motion of the one or more moving objects in the first frame and the second frame. Step 312 may be performed by Figure 2 The trajectory determination engine 222 mentioned in is executed.
[0073] At step 314, process 300 may include generating, by trajectory determination engine 222, trajectories based on downsampling and upsampling of the plurality of attributes. Trajectories of one or more moving objects and moving regions of one or more moving objects in the first and second frames may be determined based on the normalized plurality of attributes.
[0074] At step 316a, process 300 may include: stitching a first frame at a first location where one or more objects exist in at least one overlapping region. Trajectories from the first frame are masked to regenerate masked portions of a second frame. Step 316a may be performed by Figure 2 The splicing engine 224 mentioned in is executed.
[0075] At step 316b, process 300 may include: stitching the first frame at a second location where the one or more objects are not present in the at least one overlapping region. The track from the first frame may be masked to regenerate the masked portion of the second frame. Step 316b may be performed by Figure 2 The splicing engine 224 mentioned in is executed.
[0076] At step 318, process 300 may include: Figure 2 The generation engine 226 mentioned in generates a stitched image by stitching the mask portion of the first frame and the second frame.
[0077] Figure 4 An architectural diagram depicting a method 400 for generating a stitched image by stitching image frames including one or more moving objects is shown in accordance with an embodiment of the present subject matter. The method 400 may be performed by a system 202 incorporated into an electronic device.
[0078] At step 402, method 400 includes performing frame selection. A first frame and a second frame may be selected from an input stream of a sequence of frames associated with a plurality of frames captured by an image capture device. Examples of image capture devices may include, but are not limited to, cameras, video recorders, and CCTV. In an embodiment, a plurality of frames may be buffered to compare a new frame to a previous frame. The first frame may be a previous frame relative to a current frame, and the second frame may be a current frame.
[0079] At step 404, method 400 may include identifying one or more moving objects from the first frame and the second frame and associated various timestamps for the object motion within the first frame and the second frame. The identification may be performed based on comparing a plurality of second frame grids of the second frame with a plurality of first frame grids of the first frame in terms of pixel intensity.
[0080] At step 406, method 400 may include determining a plurality of attributes associated with one or more moving objects from at least one overlapping region of the first frame and the second frame. The determination may be performed by Figure 2 The determination engine 218 mentioned in the above is executed via one or more sensors (such as motion sensors and IMU sensors). In addition, the multiple attributes captured from the at least one overlapping region can be associated with multiple device properties associated with the capture device that captured the multiple frames. The association can include: changing the value of one or more of the multiple attributes relative to the value of the multiple device properties. Figure 2 The normalization engine mentioned in performs association to normalize the multiple attributes. The normalization may include associating the multiple attributes with the multiple device attributes.
[0081] At step 408, method 400 may include generating, by trajectory determination engine 222, trajectories based on downsampling and upsampling of the plurality of attributes. Trajectories of one or more moving objects and moving regions of one or more moving objects in the first and second frames may be determined based on the normalized plurality of attributes.
[0082] At step 410, method 400 may include performing one of a plurality of stitching techniques. The plurality of stitching techniques may include stitching a first frame at a first location of at least one overlapping region where one or more objects are present. Trajectories from the first frame are masked to regenerate a masked portion of a second frame. The plurality of stitching techniques may also include stitching the first frame at a second location of at least one overlapping region where one or more objects are not present. Trajectories from the first frame may be masked to regenerate a masked portion of a second frame.
[0083] At step 412, method 400 may include: Figure 2 The generation engine 226 mentioned in the figure completes the stitching of the image by stitching the masked parts of the first frame and the second frame.
[0084] FIG. 5a shows a diagram depicting a method 500a for selecting a first frame and a second frame from a plurality of frames according to an embodiment of the present subject matter. The method 500a may be Figure 2 The capture engine 214 mentioned in the above is executed. A separate quality metric for each of the plurality of frames can be determined to estimate the quality of each frame. In addition, one or more distorted frames can be removed from the plurality of frames. Next, the remaining frames in the plurality of frames can be buffered with time stamps to perform time-based comparisons on the remaining frames. Equation 1 mentioned below describes the buffering.
[0085] F n = (1-r)F n +rF 0
[0086] F n Is a new frame
[0087] F 0 It's an old frame
[0088] r is the regulator value, which adjusts the rate at which foreground objects are removed from the background
[0089] The "N" frames can be clustered into M clusters, such as σ1, σ2, ..., σM. The salient content of any object or frame can be the visual content of the object or frame, which can be the color, texture or shape of the object or frame. The similarity between two frames is determined by calculating the similarity of the visual content.
[0090] FIG. 5 b shows an operational flow diagram depicting a process 500 b for selecting a first frame and a second frame from a plurality of frames, according to an embodiment of the present subject matter.
[0091] Process 500b includes performing frame buffering as disclosed in Figure 5a, and continuing to perform pooling and convolution on multiple frames. Based on performing frame buffering, pooling and convolution, the quality of each of the multiple frames can be estimated. Pooling can be performed to reduce the number of parameters to be learned and the amount of calculations performed in the network. Convolution can be a per-element matrix multiplication of a kernel (filter) with image pixels. The quality metric can be derived from the power spectral density (PSD) of each frame. In addition, one or more frames can be selected based on determining that the PSD associated with one or more frames is greater than a predetermined threshold. After selecting one or more frames, density-based clustering can be performed to detect and eliminate outliers. Based on this, a first frame and a second frame can be selected.
[0092] FIG. 5c shows a diagram 500c depicting the first and second stages for selecting a first frame and a second frame according to an embodiment of the present subject matter. The first stage may include performing polarization, convolution, determining PSD, and eliminating lower value frames from a plurality of frames. Additionally, the second stage may include performing density-based clustering to detect and eliminate outliers. Based on this, the first frame and the second frame may be selected.
[0093] Figure 6 An operational flow chart depicting a process 600 for identifying one or more moving objects in a plurality of frames according to an embodiment of the present subject matter is shown. One or more moving objects may be identified from a first frame and a second frame in the plurality of frames. The process 600 may be performed by Figure 2 The recognition engine 216 mentioned in is executed.
[0094] The model learned before time t-1 cannot be used directly for detection at time t. To use this model, motion compensation is required. A compensated background model can be constructed that performs motion compensation at time t by merging the statistics of the model at time t-1. A single Gaussian model with age can use a Gaussian distribution to track changes in the motion background.
[0095] If the age of the candidate background model becomes greater than the apparent background model, the models may be swapped and the correct background model used. The candidate background model may remain invalid until the age becomes greater than the apparent background model, at which time the two models may be swapped.
[0096] If in the new frame, the pixel intensity in a particular grid does not match the corresponding grid in the previous frame, it can be concluded that there is a moving object in the grid. The parameters with a tilde can refer to the parameter values of the corresponding grid in the previous frame. Due to motion, the background can change, so the grids can match in different frames. The recognition of one or more frames can be described by the equation 2 mentioned below:
[0097]
[0098] where M and V are the mean and variance of all pixels in grid i, and i is the age of grid i, which refers to the number of consecutive frames that the grid is displayed.
[0099] Since the background can move in different frames, motion compensation can be used to match the grid in consecutive frames.
[0100] For all meshes G (3224) at timestamp t, process 600 may include: first performing a Kanade-Lucas-Tomasi feature tracker (KLT) on the corners of each mesh G(t)I to extract features of the points, and further performing RANSAC [2] to generate a transformation matrix from t to t-1. frame.
[0101] For each grid , process 600 may include: Find matching mesh , and Apply a weighted summation to the grid in frame t-1 covered by i to generate Parameter value.
[0102]
[0103] For each grid, we keep track of two SGMs B and F, and we update only one model at a time. We start by updating B (assuming it is the background model) until
[0104]
[0105] where s is the threshold parameter. We then update F, similarly until
[0106]
[0107] In addition, the models can be exchanged to record the foreground and background models if the following conditions are met: the number of consecutive updates of F is greater than the number of consecutive updates of B, that is,
[0108]
[0109] If the "foreground" stays in the frame longer than the "background", a swap can be performed and there is a high probability that the foreground is really the background. M and V can be the mean and variance of all pixels in grid i, Age of grid i.
[0110] In addition, the model exchange can include an SGM (Single Gaussian Model) that can be configured to track changes in the moving background, if in the new frame, the pixel intensity in the grid is different compared to the corresponding grid in the previous frame, it has a moving object. Two SGMs can be used to record grids related to the background and foreground (moving objects) respectively, so that the pixel intensity in the foreground can not pollute the parameter values in the background Gaussian model.
[0111] Figure 7An operational flow chart depicting a process 700 for determining multiple attributes of one or more moving objects according to an embodiment of the present subject matter is shown. The one or more moving objects may be present in multiple frames. Examples of multiple attributes may include, but are not limited to, one or more of the relative motion of the one or more moving objects, the exchange area ratio of the one or more moving objects, the color of the one or more moving objects, the background color, the size of the one or more moving objects, the frame rate, the speed of the one or more moving objects, and the time spent by the one or more moving objects in the first frame. Multiple attributes may be determined based on background extraction, foreground extraction, edge detection and centroid identification, and speed detection.
[0112] Background extraction, foreground extraction, and speed detection can be performed based on equations 3, 4, and 5 mentioned below. After foreground extraction, the collected image can be converted into a binary image because operations such as edge detection, noise and dilation removal, and object labeling are applicable to binary platforms. If the pixel has coordinates, the position of the vehicle in each frame is used to calculate the speed of the moving object in each frame. i=(a,b) i-1=(e,f), where the center of mass position of the object is shown in the i-th frame and the i-1-th frame, with (a,b) coordinates and (e,f) coordinates.
[0113]
[0114] Where: n is the frame number
[0115] F xy (t n ) is the pixel value at (x,y) in the nth frame;
[0116] k xy (t n ) is the average value of the pixel at (x, y) in the nth frame averaged over the previous j frames, and j is the frame number used to calculate the average of the pixel value.
[0117]
[0118] Where Nxy(tn) is the value of the foreground or background of the image at pixel (x, y) in the nth frame.
[0119] T is the threshold used to distinguish foreground from background.
[0120]
[0121] K is the calibration factor.
[0122] Figure 8An operational flow chart depicting a process 800 for tracking a path of one or more moving objects is shown in accordance with an embodiment of the present subject matter. Process 800 may include determining whether a normalized attribute of one or more moving objects renders the one or more moving objects dynamic. In an embodiment, when it is determined that the one or more moving objects are not dynamic, a static position of the identified object may be generated. In another embodiment, when it is determined that the one or more moving objects are dynamic, process 800 may include performing trajectory generation and frame stitching. Path tracking may be described by equation 6 mentioned below. For each pixel movement (u, v) that is constant within a small neighborhood w
[0123]
[0124]
[0125]
[0126] Fig. 9 A diagram 900 depicting trajectory generation according to an embodiment of the present subject matter is shown. Trajectory generation may be performed by Figure 2 The trajectory determination engine 222 mentioned in the above is executed. In addition, trajectory generation can include downsampling and upsampling. Downsampling can be used to reduce the dimension and obtain the purest features - taking high-dimensional data and projecting it into a low dimension. Upsampling can be used to increase the dimension of low-dimensional data and try to reconstruct the original frame. Trajectory generation can utilize Fig. 9 Equation 6 mentioned in can be used to calculate the velocities in the y and x directions by acquiring two frames. Using frame t+Δt, equation 6, and the pixel trajectory, the next image at t+2*Δ can be generated.
[0127] Although specific language has been used to describe the present disclosure, it is not intended to limit any limitations resulting therefrom. It is obvious to those skilled in the art that various work modifications may be made to the method to implement the inventive concepts taught herein. The accompanying drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that the elements described may be well combined into a single functional element. Alternatively, certain elements may be divided into multiple functional elements. Elements from one embodiment may be added to another embodiment. Obviously, the present disclosure may be embodied in various other ways and practiced within the scope of the appended claims.
Claims
1. A method for generating a spliced image by an electronic device (200), the method include: Obtaining (102) an input stream comprising a plurality of frames via an image capture device; identifying (104) one or more moving objects and one or more time stamps associated with movement of the one or more moving objects from a first frame and a second frame selected from the plurality of frames; determining (106) a plurality of attributes associated with the one or more moving objects from at least one overlapping region of the first frame and the second frame; normalizing the determined plurality of properties relative to a plurality of device properties associated with the image capture device (108); Based on the normalized plurality of attributes, determining ( 110 ) tracks associated with the one or more moving objects and moving regions of the one or more moving objects in the first frame and the second frame; Execute (112) one of the following: splicing the first frame at a first location of the at least one overlapping region where the one or more moving objects exist, wherein the trajectory from the first frame is masked to regenerate a masked portion of the second frame, and splicing the first frame at a second location of the at least one overlapping region where the one or more moving objects are not present, wherein the trajectory from the first frame is masked to regenerate a masked portion of the second frame; and The stitched image is generated (114) by stitching the first frame with the masked portion of the second frame.
2. The method according to claim 1, in, The first frame is a previous frame relative to a current frame, and the second frame is the current frame.
3. The method according to claim 1, in, The multiple properties include one or more of the relative motion of the one or more moving objects, the exchange area ratio of the one or more moving objects, the color of the one or more moving objects, the background color, the size of the one or more moving objects, the frame rate, the speed of the one or more moving objects, and the time spent by the one or more moving objects in the first frame, and the multiple device properties include one or more of the speed of the image capture device and the direction of movement of the image capture device.
4. The method according to claim 1, further comprising: include: performing a timestamp-based comparison of the plurality of frames relative to a quality metric of each frame, wherein each frame of the plurality of frames is buffered with a timestamp associated with the corresponding frame; estimating a quality of each of the plurality of frames based on the timestamp-based comparison of a quality metric for each frame, wherein the quality metric is derived from a power spectral density (PSD) of each of the plurality of frames; and The first frame and the second frame are selected from among the plurality of frames based on the estimation.
5. The method according to claim 4, in, Estimating the quality of each frame includes: Processing the plurality of frames by applying a plurality of machine learning (ML) techniques; calculating a PSD associated with each frame of the processed plurality of frames; and According to density-based clustering and outlier elimination, at least two frames having PSDs greater than a predetermined threshold are selected from the plurality of frames, wherein the at least two frames include the first frame and the second frame.
6. The method according to claim 1, in, Identifying the one or more moving objects includes: comparing a plurality of second frame grids of the second frame to a plurality of first frame grids of the first frame in terms of pixel intensities associated with the plurality of second frame grids and the plurality of first frame grids; determining that pixel intensities associated with the plurality of second frame grids do not match pixel intensities associated with the plurality of first frame grids; and Based on the determination, the one or more moving objects in the first frame and the second frame are identified.
7. The method according to claim 1, in, Determining the trajectory of the one or more moving objects and the moving area of the one or more moving objects includes: determining that the plurality of attributes, after being normalized, causes the one or more moving objects to move; detecting a moving direction of the one or more moving objects in the first frame and the second frame; and The trajectory is generated based on downsampling and upsampling of the plurality of attributes.
8. The method according to claim 1, in, Normalizing the determined plurality of attributes with respect to the plurality of device attributes includes associating the plurality of attributes with the plurality of device attributes.
9. The method according to claim 8, in, Associating the plurality of attributes with the plurality of device properties includes changing values of one or more of the plurality of attributes relative to values of the plurality of device properties.
10. An electronic device (200) for generating a spliced image, the electronic device include: Memory (206); as well as At least one processor (204) coupled to the memory (206), wherein the at least one processor is configured to operate according to the method of one of claims 1 to 9.
11. A non-transitory computer-readable storage medium storing instructions, which, when executed by at least one processor (204) of an electronic device (200), cause the electronic device to perform operations according to the method of any one of claims 1 to 9.