Image processing method, device, and computer-readable recording medium

By automatically identifying scenes in image sequences through image recognition and artificial intelligence algorithms, deleting non-key frames, and generating multimedia files, it solves the time and professional software needs of ordinary users in video or photo post-processing, and achieves fast and efficient image editing.

CN122120597APending Publication Date: 2026-05-29AMTRAN TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AMTRAN TECHNOLOGY CO LTD
Filing Date
2024-11-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, post-processing of videos or photos requires professional software and a lot of time, making it difficult for ordinary users to efficiently complete simple video or photo editing.

Method used

By automatically identifying scenes in image sequences through image recognition and artificial intelligence algorithms, deleting non-primary frames and generating multimedia files, and combining people tracking and voice control to adjust shooting parameters, automatic image sequence condensation is achieved.

Benefits of technology

Multimedia files can be generated quickly without the need for professional software or a lot of time, meeting the simple video or photo editing needs of ordinary users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120597A_ABST
    Figure CN122120597A_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, device and computer readable recording medium. The image processing method comprises judging a scene corresponding to an image sequence, and creating the image sequence based on the scene to generate a multimedia file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an image processing mechanism, and more particularly to an image processing method, apparatus, and computer-readable recording medium. Background Technology

[0002] With the advancement of technology, electronic products equipped with cameras are becoming increasingly common. Therefore, it has become quite convenient for modern people to shoot videos and photos. Sharing videos on various social media platforms is also quite easy. However, before sharing and uploading their work, users need to spend a significant amount of time using photo editing software, image and audio editing software, and other tools to organize and edit the videos. These tools generally require time to learn on one's own.

[0003] However, for most people, time is money, and spending a lot of time and energy editing images is impractical. For example, simple and basic needs like shooting a short video of a home-cooked meal to share, or shooting a short video of assembling a computer or electronic product to guide the workflow of related tasks, are unlikely to be pursued by most people unless the work can bring considerable income.

[0004] Most cameras, camcorders, and camera-equipped smartphones on the market currently emphasize shooting raw videos or photos, without considering post-production issues. However, post-production is what truly determines the practicality of a work. Post-production for videos and photos largely relies on professionals or specialized software and requires a significant amount of time. This time commitment makes the creation and sharing of short videos and standard operating procedures (SOP) files impractical. Summary of the Invention

[0005] This disclosure provides an image processing method, apparatus, and computer-readable recording medium that can automatically condense original image sequences to generate edited multimedia files.

[0006] The image processing method disclosed herein includes: determining the scene corresponding to the image sequence, and processing the image sequence based on the scene to generate a multimedia file.

[0007] In one embodiment disclosed herein, the step of determining the scene corresponding to an image sequence includes: performing an image recognition program on the image sequence to detect multiple objects in the image sequence; and determining the scene corresponding to the image sequence based on the objects. After determining the scene corresponding to the image sequence, the method further includes: based on the scene, identifying multiple non-primary frames in multiple frames of the image sequence, and deleting the non-primary frames from the frames to obtain multiple specified frames. Then, a multimedia file is generated based on the specified frames.

[0008] In one embodiment of this disclosure, the step of determining the scene corresponding to the image sequence based on the plurality of objects includes: classifying each object as a subject object corresponding to the scene or an object that does not belong to the subject object; for each frame of the image sequence, calculating a first number corresponding to the subject object and a second number corresponding to the object object for each frame; and determining whether each frame is a non-primary frame based on the first number and the second number.

[0009] In one embodiment disclosed herein, the step of determining the scene corresponding to the image sequence based on the plurality of objects includes: performing artificial intelligence (AI) operations on each frame of the image sequence to determine the action corresponding to each frame; and determining whether each frame is a non-primary frame based on whether the action is related to the scene.

[0010] In one embodiment disclosed herein, the step of generating a multimedia file based on the plurality of specified frames includes: dividing the plurality of specified frames into a plurality of segments based on the template content corresponding to the scene; and adding a corresponding text label to each segment.

[0011] In one embodiment of this disclosure, the image processing method further includes: activating the image acquisition device to acquire an image sequence; performing a person tracking procedure on the image sequence to identify and track a specified person in the image sequence; and adjusting the image acquisition parameters of the image acquisition device based on the position of the specified person.

[0012] In one embodiment of this disclosure, the image acquisition parameters include a specified angle, and the step of adjusting the image acquisition parameters of the image acquisition device based on the position of the specified person includes: sending an angle adjustment command to the motor module based on the position of the specified person to drive the motor module to rotate the image acquisition device by the specified angle.

[0013] In one embodiment of this disclosure, the image acquisition parameters include a specified focal length, and the step of adjusting the image acquisition parameters of the image acquisition device based on the position of the specified person includes: sending a focal length adjustment command to the image acquisition device based on the position of the specified person, so that the image acquisition device is adjusted to the specified focal length.

[0014] In one embodiment of this disclosure, the image processing method further includes: acquiring a person image of a specified person through the image acquisition device when the image acquisition device is initially started; and performing a face recognition program on the person image to obtain a feature set in the person image, and storing the feature set for use by the person tracking program.

[0015] In one embodiment of the present disclosure, the image processing method further includes: activating an image capturing device to acquire an image sequence; receiving a voice signal from a voice receiving device during the acquisition of the image sequence by the image capturing device; performing a voice recognition program on the voice signal to obtain a device adjustment command; and adjusting the image capturing parameters of the image capturing device based on the device adjustment command.

[0016] In one embodiment of this disclosure, the image processing method further includes: after activating the image acquisition device to acquire an image sequence, transmitting the image sequence to a display for display.

[0017] The image processing apparatus disclosed herein includes: a memory including at least one code segment; an image capturing device; and a processor coupled to the memory and the image capturing device, wherein the processor reads at least one code segment to: determine a scene corresponding to an image sequence, wherein the image sequence is acquired by the image capturing device controlled by the processor, and performs processing on the image sequence based on the scene to generate a multimedia file.

[0018] The non-transient computer-readable recording medium disclosed herein records a program, which, when executed by a processor in an electronic device, performs the following steps: determining the scene corresponding to the image sequence, and creating a multimedia file based on the scene.

[0019] Based on the above, this disclosure can automatically remove unimportant frames from an image sequence and generate a multimedia file corresponding to the current scene based on multiple condensed specified frames. Therefore, users do not need to learn video editing tools or spend a lot of time selecting the frames they wish to retain. Attached Figure Description

[0020] Figure 1 This is a block diagram of an image processing apparatus according to an embodiment of the present disclosure;

[0021] Figure 2 This is a flowchart of an image processing method according to an embodiment of the present disclosure;

[0022] Figure 3 This is a block diagram of an image processing apparatus according to another embodiment of the present disclosure;

[0023] Figure 4 This is a flowchart of an image processing method according to another embodiment of the present disclosure.

[0024] Explanation of reference numerals in the attached figures

[0025] 100: Image processing device

[0026] 110: Processor

[0027] 12: Memory

[0028] 130: Image capturing device

[0029] 320: Motor Module

[0030] 330: Voice receiving device

[0031] 340: Monitor

[0032] 350: Communication Connector

[0033] S205~S210: Steps

[0034] S401~S427: Steps Detailed Implementation

[0035] Figure 1 This is a block diagram of an image processing apparatus according to an embodiment of the present disclosure. Please refer to... Figure 1 The image processing apparatus 100 includes a processor 110, a memory 120, and an image capturing device 130. The processor 110 is coupled to the memory 120 and the image capturing device 130.

[0036] The processor 110 may be implemented using a central processing unit (CPU), a physical processing unit (PPU), a programmable microprocessor, an embedded control chip, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or other similar devices.

[0037] The memory 120 may be implemented using any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk, or other similar devices or combinations thereof. The memory 120 includes one or more code segments, which, after being installed, are executed by the processor 110 to implement the image processing method described later.

[0038] The image capturing device 130 can be a camera employing a charge-coupled device (CCD) lens, a complementary metal-oxide-semiconductor (CMOS) lens, or the like. In one embodiment, the image capturing device 130 may consist of a single camera with a resolution of 12MP (4032×3040), a 120-degree field of view to provide the user with a wider field of view, and a 5x optical zoom lens. During the image capturing process, the image capturing device 130 can zoom in or out to capture more objects; however, this disclosure is not limited to this.

[0039] In one embodiment, the processor 110 and memory 120 may also be integrated into a system on a chip (SOC) with a neural-network processing unit.

[0040] Figure 2 This is a flowchart of an image processing method according to an embodiment of the present disclosure. Please also refer to... Figure 1 and Figure 2 In step S205, the scene corresponding to the image sequence is determined. The image sequence is acquired by an electronic device (e.g., the image processing device 100). Next, in step S210, the image sequence is processed based on the scene to generate a multimedia file.

[0041] In practical applications, the image processing device 100 can be a smart TV, a smart camera, a smartphone, or other image-capturing device. The following embodiments use a smart TV as an example. A smart TV can acquire image sequences through the image-capturing device 130. The processor 110 performs a series of processes on the image sequences to identify the scene corresponding to the image sequences, and then performs post-processing on the image sequences to generate multimedia files for that scene. Therefore, after acquiring an image sequence, the image sequence can be promptly condensed to generate a multimedia file corresponding to the current scene. In the case of a smart TV, the generated multimedia file can also be directly displayed on the TV screen. In the case of a smart camera, the multimedia file can also be presented through the display screen attached to the smart camera.

[0042] Specifically, the processor 110 may first execute an image recognition program on the image sequence to detect multiple objects in the image sequence. In one embodiment, the image recognition program employs an image segmentation neural network module and an object detection module. For each frame in the image sequence, the image segmentation neural network module segments each frame into multiple blocks, and then the object detection module separates each object in each block. The image segmentation neural network module is, for example, MobileNetV3-SSD using the COCO dataset. The object detection module is, for example, YoloV8.

[0043] Next, the processor 110 determines the scene corresponding to the image sequence based on the objects. For example, the location of "recipe preparation" is generally in the kitchen, where there are kitchen utensils, ingredients, and other objects. Therefore, the scene is determined based on the category of the objects. In one embodiment, the processor 110 classifies the detected multiple objects according to a default classification rule. For example, the classification categories include: kitchen utensils, ingredients, cosmetics, computer parts, etc. Then, the processor 110 can further determine the scene corresponding to the image sequence based on the number of objects included in each category according to a predetermined judgment rule. For example, if the number of objects classified as kitchen utensils and ingredients exceeds a certain proportion of the total number of objects, the scene is determined to be "recipe preparation". If the number of objects classified as computer parts, etc., exceeds a certain proportion of the total number of objects, the scene is determined to be "computer assembly". If these objects cannot be classified, it is determined to be a general scene. However, this is only an example and is not intended to be limiting.

[0044] After determining the scene, the processor 110 identifies multiple non-primary frames in the multiple frames of the image sequence based on the scene, and deletes the non-primary frames from the frames to obtain multiple designated frames. In one embodiment, the processor 110 may determine for each frame whether the frame is related to the scene, and set frames that are not related to the scene as non-primary frames.

[0045] For example, processor 110 classifies each object as either a subject object corresponding to the scene or an object that is not a subject object. For each frame of the image sequence, processor 110 calculates a first number corresponding to a subject object and a second number corresponding to an object object for each frame. That is, among the objects included in each frame, the number of subject objects and the number of object objects are calculated. Furthermore, based on the first number and the second number, it is determined whether this frame is classified as a non-primary frame. Subsequently, non-primary frames are removed from the initial multiple frames in the image sequence to obtain multiple primary frames, and the multiple primary frames are set as designated frames.

[0046] In addition, if the number of main frames obtained exceeds the preset number, the processor 110 can further filter out one or more duplicate frames from the multiple main frames based on the similarity between two adjacent main frames in time, delete the duplicate frames from the multiple main frames, and set the main frame that is finally retained as the designated frame.

[0047] In another embodiment, the processor 110 may also perform artificial intelligence (AI) calculations on each frame of the image sequence to determine the action corresponding to each frame, and determine whether each frame is a non-primary frame based on whether the action is related to the scene. For example, in the case of a "recipe making" scene, if the action in a frame is a character adjusting the screen, or the character's current action is not a non-primary event such as handling ingredients or cooking food, then this frame can be set as a non-primary frame.

[0048] Then, the processor 110 generates a multimedia file based on the specified frame. Here, the multimedia file may be, for example, a short video or a slideshow file.

[0049] In one embodiment, the processor 110 classifies multiple objects identified from the image sequence into multiple subject objects corresponding to the scene and multiple object objects not belonging to the subject objects. Furthermore, the processor 110 records the temporal correlation between each subject object and each object object. For example, the processor 110 uses a time-series neural network to perform machine learning operations on the content of the original image sequence to find the temporal correlation between the subject objects and each object object in the image sequence. Based on template content corresponding to the scene, the processor 110 divides the specified frame into multiple segments. Then, the processor 110 adds corresponding text labels to each segment.

[0050] For example, taking the scenario of "recipe creation" as an example, the corresponding template content includes text content used in two stages: the ingredient processing stage and the ingredient cooking stage. The processor 110 can use AI calculations to determine the time segments between the ingredient processing stage and the ingredient cooking stage based on the actions of specified frames. Furthermore, the processor 110 can further use AI calculations to convert main objects such as ingredients and seasonings into text and add corresponding text labels to each segment. For example, it can generate a corresponding title for the entire multimedia file and corresponding text labels for different stages. Additionally, corresponding image labels can also be added.

[0051] For example, during recipe creation, the preparation and cooking order of each ingredient is recorded based on a time sequence. These sequences are interconnected temporally, noting which actions require which objects and the interactions between them. Beyond ingredients, the relative relationships between objects like chairs, tables, and windows can be further analyzed and listed. For instance, in the ingredient preparation stage, the order of using which cooking utensils to process which ingredient first, and then which utensils to process which ingredient subsequently, is recorded. Similarly, in the cooking stage, the order of using which cooking utensils to cook which ingredient first, and then which utensils to cook which ingredient subsequently, is recorded. Based on this, the appearance time of each object and the interactions between different objects are recorded.

[0052] Figure 3 This is a block diagram of an image processing apparatus according to another embodiment of this disclosure. Please refer to... Figure 3 The image processing apparatus 300 includes a processor 110, an image capturing device 130, a motor module 320, a voice receiving device 330, a display 340, and a communication connector 350. The processor 110 is coupled to the image capturing device 130, the motor module 320, the voice receiving device 330, the display 340, and the communication connector 350.

[0053] In this embodiment, the processor 110 is implemented using a System-on-a-Chip (SoC) with a Neural-network Processing Unit (NNPU). During the shooting process, the image-capturing device 130 can capture more objects by zooming in or out of the camera. The motor module 320 drives the image-capturing device 130 to rotate. For example, the image-capturing device 130 has a rotating base, and the motor module 320 drives the rotating base to rotate. The motor module 320 can be a brushless motor, which avoids unnecessary noise during the shooting process of the image-capturing device 130 due to the rotation of the motor module 320 and has a longer service life.

[0054] The voice receiving device 330 is, for example, a microphone array audio inmodule, used to collect ambient sound and generate corresponding voice signals as the sound source for image recording.

[0055] The display 340 is, for example, a light-emitting diode display (LED display), a liquid crystal display (LCD), or an organic light-emitting diode display (OLED display). The processor 110 can drive the display 340 via, for example, an eDP (Embedded DisplayPort) VBO (V-by-One) interface. After the image-capturing device 130 is activated to acquire an image sequence, the image sequence is transmitted to the display 340 for display. The user can view the image in real time on the display 340 to decide whether fine-tuning is needed.

[0056] The communication connector 350 can be a chip or circuit employing Local Area Network (LAN) technology, Wireless LAN (WLAN) technology, or mobile communication technology. An example of a LAN is Ethernet. An example of a WLAN is Wi-Fi. Examples of mobile communication technologies include Global System for Mobile Communications (GSM), third-generation (3G), fourth-generation (4G), and fifth-generation (5G). Connecting to a network via the communication connector 120 enables connection to a cloud server, allowing for the timely uploading of at least one of the prepared multimedia files and the original image sequence to the cloud server for storage. Additionally, multimedia files can be directly published to social media websites or transferred to social media applications.

[0057] Figure 4 This is a flowchart of an image processing method according to another embodiment of this disclosure. Please also refer to... Figure 3 and Figure 4 In step S401, the image processing device 300 is started. Next, in step S403, initialization settings are performed. Here, when the image processing device 300 is turned on for the first time, it prompts the user to perform initialization settings. For example, network parameters are set, including (but not limited to) a service set identifier (SSID), a username and password for connecting to the wireless network.

[0058] In addition, during the initial startup of the image acquisition device 130, the initialization settings also include the following actions. The processor 110 acquires an image of a specified person through the image acquisition device 130. The processor 110 performs a facial recognition procedure on the image to obtain a set of features from the image. For example, the facial recognition procedure is used to separate the facial contours and obtain the corresponding feature set. Then, the acquired feature set corresponding to the specified person is stored for use in subsequent person tracking procedures.

[0059] In step S405, the image acquisition device 130 is activated to acquire an image sequence. In step S407, an image recognition program is executed. Next, in step S409, the scene corresponding to the image sequence is determined. In step S411, the image sequence is condensed. In step S413, a multimedia file is generated. For detailed explanations of steps S407, S409, S411, and S413, please refer to steps S205, S210, S215, and S220 above, respectively.

[0060] After the image capturing device 130 is activated to acquire an image sequence, in step S415, a person tracking program is executed on the image sequence to identify and track a designated person in the image sequence. Accordingly, the processor 110 can adjust the image capturing parameters of the image capturing device 130 based on the position of the designated person. The shooting angle of the image capturing device 130 is adjusted by the person tracking program so that the designated person is positioned as close as possible to a designated position in the frame (e.g., in the center of the frame).

[0061] In one embodiment, the image acquisition parameters include a specified angle at which the image acquisition device 130 is to rotate. For example, in step S417, the processor 110 sends an angle adjustment command to the motor module 320 based on the position of a specified person, so as to drive the motor module 320 to rotate the image acquisition device 130 by the specified angle.

[0062] In one embodiment, the image acquisition parameters include a specified focal length. For example, in step S419, the processor 110 sends a focal length adjustment command to the image acquisition device 130 based on the position of the specified person, causing the image acquisition device 130 to adjust to the specified focal length.

[0063] Alternatively, in step S425, the processor 110 may execute integrated control based on the position of the specified person to send an angle adjustment command to the motor module 320 and a focal length adjustment command to the image capturing device 130, thereby executing steps S417 and S419.

[0064] During the acquisition of an image sequence by the image capturing device 130, in step S421, the processor 110 receives an audio signal from the audio receiving device 330. Then, in step S423, a speech recognition program is executed on the audio signal to obtain a device adjustment command. Accordingly, the processor 110 can adjust the image capturing parameters of the image capturing device 130 based on the device adjustment command. In one embodiment, the image capturing parameters include at least one of contrast, saturation, and color temperature. For example, in step S427, the processor 110 sends a device adjustment command to the image capturing device 130 to perform image quality control.

[0065] In practical applications, if you are not satisfied with the current image quality displayed on the monitor 340, you can input voice signals such as "increase contrast", "increase saturation", and "adjust color temperature" and obtain device adjustment instructions through the voice recognition program, and then adjust the image quality of the monitor 340 according to the device adjustment instructions.

[0066] In addition, if users are not satisfied with the shooting position or focal length of the image capturing device 130, they can also control the motor module 320 or the focal length and / or viewing angle of the image capturing device 130 by inputting voice signals such as "zoom in", "zoom out", "a little to the left", "a little to the right".

[0067] In one embodiment, in step S425, the processor 110 may also perform integrated control based on the results of the person tracking program and the voice recognition program. That is, the processor 110 determines whether to send an angle adjustment command to the motor module 320, whether to send a focus adjustment command to the image capturing device 130, and whether to send a device adjustment command to the image capturing device 130, based on the results of the person tracking program and the voice recognition program. Furthermore, it determines whether to execute any one or a combination of steps S417, S419, and S427.

[0068] In summary, this disclosure can automatically remove unimportant frames from an image sequence and generate multimedia files based on a condensed set of specified frames. Therefore, users do not need to learn video editing tools or spend a lot of time selecting which frames to keep; the above embodiments enable them to quickly and accurately generate the desired short videos or electronic presentation files.

Claims

1. An image processing method, applicable to electronic devices, characterized in that, Includes the following steps: Determine the scene corresponding to the image sequence, wherein the image sequence is acquired by the electronic device; and The image sequence is processed based on the described scene to generate a multimedia file.

2. The image processing method according to claim 1, characterized in that, The steps for determining the scene corresponding to the image sequence include: An image recognition program is executed on the image sequence to detect multiple objects in the image sequence; and The scene corresponding to the image sequence is determined based on the multiple objects. After determining the scene corresponding to the image sequence, the method further includes: Based on the scenario, multiple non-primary frames are identified in multiple frames of the image sequence, and multiple designated frames are obtained by deleting the multiple non-primary frames from the multiple frames. The multimedia file is generated by: The multimedia file is generated based on the specified frames.

3. The image processing method according to claim 2, characterized in that, The steps for determining the scene corresponding to the image sequence based on the multiple objects include: Each of the plurality of objects is classified into a subject object corresponding to the scenario or an object that does not belong to the subject object; For each of the plurality of frames in the image sequence, calculate a first number corresponding to the subject object and a second number corresponding to the object object for each of the plurality of frames; and The determination of whether each of the plurality of frames is one of the plurality of non-primary frames is based on the first quantity and the second quantity.

4. The image processing method according to claim 2, characterized in that, The steps for determining the scene corresponding to the image sequence based on the multiple objects include: Artificial intelligence operations are performed on each of the plurality of frames in the image sequence to determine the action corresponding to each of the plurality of frames; and Based on whether the action is related to the scene, it is determined whether each of the plurality of frames is one of the plurality of non-primary frames.

5. The image processing method according to claim 2, characterized in that, The steps for generating the multimedia file based on the multiple specified frames include: Based on the template content corresponding to the scene, the multiple specified frames are divided into multiple segments; and Add corresponding text labels to each of the multiple sections.

6. The image processing method according to claim 1, characterized in that, Also includes: The imaging device is activated to acquire the image sequence; Perform a person tracking procedure on the image sequence to identify and track a specified person in the image sequence; as well as Based on the location of the specified person, adjust the imaging parameters of the imaging device.

7. The image processing method according to claim 6, characterized in that, The image acquisition parameters include a specified angle, and the step of adjusting the image acquisition parameters of the image acquisition device based on the position of the specified person includes: Based on the position of the specified person, an angle adjustment command is sent to the motor module to drive the motor module to rotate the imaging device by the specified angle.

8. The image processing method according to claim 6, characterized in that, The image acquisition parameters include a specified focal length, and the step of adjusting the image acquisition parameters of the image acquisition device based on the position of the specified person includes: Based on the location of the specified person, a focal length adjustment command is sent to the image capturing device, causing the image capturing device to adjust to the specified focal length.

9. The image processing method according to claim 6, further comprising: When the imaging device is initially activated, the image of the specified person is acquired through the imaging device. as well as A facial recognition program is performed on the image of the person to obtain a set of features from the image of the person, and the set of features is stored for use by the person tracking program.

10. The image processing method according to claim 1, further comprising: The imaging device is activated to acquire the image sequence; During the process of the imaging device acquiring the image sequence, the voice receiving device receives a voice signal. Perform a speech recognition program on the speech signal to obtain device adjustment commands; and The imaging parameters of the imaging device are adjusted based on the device adjustment command.

11. The image processing method according to claim 10, further comprising: After the imaging device is activated to acquire the image sequence, the image sequence is transmitted to the display for display.

12. An image processing apparatus, characterized in that, include: The memory includes at least one code segment; Image capturing device; as well as A processor, coupled to the memory and the imaging device, performs the following steps when it reads the at least one code segment: Determine the scene corresponding to the image sequence, wherein the image sequence is acquired by the imaging device controlled by the processor; and The image sequence is processed based on the described scene to generate a multimedia file.

13. A non-transitory computer-readable recording medium, wherein a program is recorded and executed by a processor in an electronic device to perform the following steps: Determine the scene corresponding to the image sequence, wherein the image sequence is acquired by the electronic device; and The image sequence is processed based on the described scene to generate a multimedia file.