Image processing method and apparatus, readable medium, and electronic device
By using pre-configured scripts on the GPU to perform image pre-processing during image processing, the problem that the image is no longer in its original state after being acquired from the upper link is solved, which improves the efficiency and accuracy of image processing, simplifies the process and reduces time-consuming.
Patent Information
- Application Number
- PCT/CN2024/139079
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-02
- Filing Date
- 2024-12-13
- Publication Date
- 2025-08-07
AI Technical Summary
In the prior art In the image processing process, after the image is acquired from the upper link, the image is no longer in its original state due to changes in size or rotation, which affects the recognition accuracy of the visual algorithm. The existing image preprocessing method is cumbersome and time-consuming.
By using pre-configured scripts on the GPU to perform image pre-processing, defining execution logic, including cropping and mirroring flip operations, avoiding copying image data to the CPU for processing, directly performing image pre-processing on the GPU, and subsequent processing is performed through the preset image processing model.
Improves the efficiency and accuracy of image processing, simplifies image preprocessing process, reduces time-consuming, and supports flexible expansion and updates.
Smart Images

Figure CN2024139079_07082025_PF_FP_ABST
Abstract
Description
Image processing method, device, readable medium and electronic device
[0001] This application claims priority to Chinese Patent Application No. 202410155992.9 filed on February 2, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to an image processing method, an image processing device, a readable medium, and an electronic device. Background Art
[0003] During image processing, special effects can be added to pictures or videos, such as mosaics and stickers. Adding special effects to images requires the use of various visual algorithms, such as face recognition algorithms and portrait segmentation.
[0004] The input to a vision algorithm is image information, which typically comes from a higher-level link. After being processed by the higher-level link, the image often undergoes some changes, requiring restoration to restore it to its original state before entering the algorithm. This means that image data received from the higher-level link must undergo image pre-processing before being fed into the vision algorithm. Summary of the Invention
[0005] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] In a first aspect, the present disclosure provides an image processing method, the method comprising:
[0007] Obtain image data to be processed;
[0008] Performing image pre-processing on the image data to be processed by a pre-configured script to obtain first image data, wherein the script is used to define the execution logic of the image pre-processing;
[0009] After the first image data is input into a preset image processing model, second image data is obtained.
[0010] In a second aspect, the present disclosure provides an image processing device, the device comprising:
[0011] An acquisition module, used for acquiring image data to be processed;
[0012] a first image processing module, configured to perform image pre-processing on the image data to be processed using a pre-configured script to obtain first image data, wherein the script is configured to define an execution logic for the image pre-processing;
[0013] The second image processing module is used to input the first image data into a preset image processing model to obtain second image data.
[0014] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of the present disclosure.
[0015] In a fourth aspect, the present disclosure provides an electronic device, comprising:
[0016] a storage device having a computer program stored thereon;
[0017] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect of the present disclosure.
[0018] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:
[0020] FIG1 is a flow chart showing an image processing method according to an exemplary embodiment;
[0021] FIG2 is a flow chart of an image processing method according to the embodiment shown in FIG1 ;
[0022] FIG3 is a flow chart of an image processing method according to the embodiment shown in FIG1 ;
[0023] FIG4 is a flow chart of an image processing method according to the embodiment shown in FIG3 ;
[0024] FIG5 is a block diagram of an image processing apparatus according to an exemplary embodiment;
[0025] FIG6 is a block diagram of an image processing apparatus according to the embodiment shown in FIG5 ; and
[0026] Fig. 7 is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0034] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0037] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0038] First, the application scenarios of the present disclosure are introduced. The present disclosure is mainly used in the scenario where the image is pre-processed before the image obtained from the upper link is input into the visual algorithm. Since the upper link usually performs processing such as splicing or mirror rotation on the image, but the visual algorithm model is trained based on the original image, therefore, during the model application stage of the visual algorithm, the image input to the algorithm model also needs to be restored (also called image pre-processing) on the basis of splicing or mirror rotation before being input into the visual algorithm.
[0039] For example, the input image obtained from the live broadcast link may be filled with black borders due to the mismatch between the image size and the aspect ratio of the user's terminal screen. However, images filled with black borders are not in line with the expectations of the visual algorithm. Because during the training phase of the visual algorithm model, the training images provided to the algorithm are usually the original images, these non-predictable images will reduce the accuracy of the visual algorithm's recognition. Therefore, the image changes caused by the upper-layer link need to be restored before input into the visual algorithm.
[0040] In the related art, the image data to be processed is copied from the GPU (Graphics Processing Unit) to the CPU (Central Processing Unit), and then the algorithm in the visual library is called on the CPU to perform image pre-processing. This will have the following problems: adaptation is required at the engine level, the algorithm code needs to be modified, and the modified algorithm code needs to be compiled to implement different image pre-processing. The process is cumbersome and not conducive to subsequent expansion and updating. In addition, copying the image data to be processed from the GPU to the CPU, and then calling the algorithm in the visual library on the CPU to perform image pre-processing will take a long time. This mainly includes two parts of time consumption: one part is the time taken to copy from the GPU to the CPU, and the other part is that the speed of processing the image on the CPU is usually much lower than the speed of directly processing the image on the GPU.
[0041] To solve the above problems, the present disclosure provides an image processing method, apparatus, readable medium, and electronic device. Specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0042] FIG1 is a flow chart showing an image processing method according to an exemplary embodiment. As shown in FIG1 , the method includes the following steps:
[0043] In step S101 , image data to be processed is obtained.
[0044] The image data to be processed here refers to image data that needs to be pre-processed. The image pre-processing may include, for example, cropping and / or mirror flipping. The image data to be processed can usually be obtained from the upper link. Among them, the "upper link" usually refers to a series of image processing steps or data transmission processes before the image input algorithm. These steps may include image acquisition equipment (such as cameras, scanners), image transmission processes (such as data transmission, compression, encoding and decoding, etc.) and pre-processing (such as image rotation, translation, scaling and color space conversion, etc.). These processing steps may cause changes to the image, causing the image to no longer be in the original input state.
[0045] For example, if the upper link is a live broadcast link, the host interacts online with the users watching the live broadcast, and the host side can obtain the avatar of the interactive user. Because the image size of the interactive user's avatar is inconsistent with the ratio of the anchor side screen size, the live broadcast link adds a black border to the original image of the interactive user's avatar, that is, the interactive user's avatar obtained by the host side has been spliced (that is, the original image is spliced with the black background). If special effects (such as beautification) are added to the interactive user's avatar, before the image data of the interactive user's avatar is input into the visual algorithm, it needs to be cropped to remove the black border to obtain the original image. Therefore, in this scenario, the image data to be processed is the image data corresponding to the interactive user's avatar obtained by the anchor side and spliced, including the black border.
[0046] In step S102, image pre-processing is performed on the image data to be processed by a pre-configured script to obtain first image data, wherein the script is used to define the execution logic of the image pre-processing.
[0047] A script generally refers to a series of instructions arranged in a specific order to perform a specific task. In this disclosure, the preset script refers to a pre-written script code based on the execution logic of the image pre-processing process, and the script may include, for example, JavaScript, Python script, etc.
[0048] In this step, image pre-processing may be performed on the image data to be processed by using the pre-configured script on the GPU to obtain the first image data.
[0049] Because scripts are typically written in a dynamic language and don't need to be compiled into binary code, but rather executed line by a interpreter, the script-based image pre-processing operations in this disclosure can be executed directly on the GPU, eliminating the need to copy the script to the CPU for pre-processing. This solves the time-consuming issue of performing image pre-processing on the CPU.
[0050] In step S103, the first image data is input into a preset image processing model to obtain second image data.
[0051] Among them, the preset image processing model may include a visual algorithm model, such as a face algorithm model, a portrait segmentation model, etc.
[0052] In the present disclosure, after the preset image processing model is called by the algorithm engine, the first image data can be input into the preset image processing model to obtain the second image data.
[0053] In addition, the present disclosure can enable the algorithm engine to perceive that the input first image data has undergone image pre-processing operations by setting a flag bit, wherein different image pre-processing operations can correspond to different flag bits.
[0054] By adopting the above method, the execution logic of image pre-processing is defined through scripts, thereby realizing flexible completion of image pre-processing through scripts. Subsequent update operations are all converged to the script layer, avoiding modifications at the algorithm code layer, making it easier to implement and more conducive to subsequent expansion and updating. In addition, since scripts are usually programs written in a dynamic language, they do not need to be compiled into binary code, but are executed line by line by an interpreter. Therefore, the present disclosure implements image pre-processing operations based on scripts, which can be executed directly on the GPU without copying to the CPU for image pre-processing operations, thereby solving the time-consuming problem of performing image pre-processing operations on the CPU.
[0055] Figure 2 is a flow chart of an image processing method according to the embodiment shown in Figure 1. In the present disclosure, the image data to be processed includes change data relative to the original captured image. The change data may include image data different from the original captured image and label data for indicating the type of change. For example, in order to adapt to the screen size of the target terminal, the upper link splices a black border on the basis of the original captured image, then the change data may include the pixel position of each pixel point in the black border and a first label indicating that the type of change is splicing. For another example, if the upper link performs a mirror rotation on the original captured image, the change data may include the rotation center and the transformation matrix of the mirror rotation, and may further include a second label indicating that the type of change is mirror rotation. This is merely an example, and the present disclosure does not specifically limit this.
[0056] As shown in FIG2 , step S102 includes the following sub-steps:
[0057] In step S1021, the task type of the image pre-processing to be currently performed is determined through the script according to the change data, and different task types correspond to different image pre-processing operations.
[0058] The task type includes: cropping and / or mirror flipping.
[0059] In method 1 of this step, the type of image pre-processing task to be performed can be determined based on the tag data indicating the type of change carried in the change data. For example, if the tag data carried in the change data is a first tag, it can be determined that the upper link has performed splicing processing on the original captured image. In this case, the type of image pre-processing task to be performed is cropping. If the tag data carried in the change data is a second tag, it can be determined that the upper link has performed mirror rotation processing on the original captured image. In this case, the type of image pre-processing task to be performed is mirror flipping to obtain the original captured image.
[0060] In the second approach, the task type may be determined by the script according to the change data and the model type of the preset image processing model.
[0061] It is understandable that different types of preset image processing models have different requirements for input images. For example, some algorithms may only need to process part of the image (such as for the face algorithm model, usually only the face area on the image is processed). Therefore, before the image obtained by the upper link is input into the preset image processing model, the model area of interest in the image can be cropped to reduce the image size of the input algorithm and improve the efficiency and accuracy of image processing.
[0062] For example, after determining based on the change data that the image obtained by the upper link needs to be cropped and / or mirror-flipped (implemented based on method one), if it is determined based on the model type of the preset image processing model that the model type is a model that only processes a part of the image, the image can be further cropped to obtain the target area of interest of the preset image processing model (for example, the target area corresponding to the face algorithm model is the face image area), and then the image of the target area can be input into the preset image processing model.
[0063] In step S1022, image pre-processing is performed on the image data to be processed by the script according to the task type to obtain first image data.
[0064] In this step, the first affine transformation matrix corresponding to the task type can be generated by the script based on the image data to be processed; the first image data can be obtained by driving the rendering engine through the script to perform image pre-processing on the image data to be processed according to the first affine transformation matrix.
[0065] Different task types correspond to different first affine transformation matrices, and the first affine transformation matrix is a transformation matrix used to implement image pre-processing operations.
[0066] In the present disclosure, by executing a pre-configured script, a first affine transformation matrix corresponding to the current task type can be generated based on the image data to be processed, and then the script drives the rendering engine to implement image pre-processing operations on the image data to be processed based on the vertex shader according to the first affine transformation matrix.
[0067] For example, the script layer can drive the rendering engine to perform a corresponding affine transformation on the input image texture. For example, if the task type is clipping, the first affine transformation matrix used in the vertex shader can be: The examples here are merely illustrative and the present disclosure is not limited thereto.
[0068] FIG3 is a flow chart of an image processing method according to the embodiment shown in FIG1 . As shown in FIG3 , the method further includes the following steps:
[0069] In step S104 , image rendering is performed according to the second image data.
[0070] In one implementation, image rendering may be performed by a rendering engine based on the second image data.
[0071] Since the input image of the preset image processing model is an image after image pre-processing, the coordinate system corresponding to the image changes after the pre-processing operation is performed on the image. Therefore, the coordinate system corresponding to the algorithm result (i.e., the second image data) output by the preset image processing model is the coordinate system of the image after the pre-processing operation. Therefore, in the present disclosure, in order to avoid misalignment when the image is rendered on the screen, before the algorithm result is input into the rendering engine for rendering, the algorithm result can be subjected to an affine inverse transformation to map the algorithm result to the first coordinate system (the first coordinate system refers to the coordinate system of the image before the pre-processing operation (i.e., the image corresponding to the image data to be processed)).
[0072] FIG4 is a flow chart of an image processing method according to the embodiment shown in FIG3 . As shown in FIG4 , step S104 includes the following sub-steps:
[0073] In step S1041, a second affine transformation matrix is determined according to the first affine transformation matrix, where the second affine transformation matrix is used to map the second image data to a first coordinate system corresponding to the image data to be processed.
[0074] After determining the first affine transformation matrix corresponding to the current image pre-processing operation, the script engine can also output the first affine transformation matrix to the algorithm engine so that the algorithm engine can implement an inverse affine transformation of the algorithm result based on the first affine transformation matrix.
[0075] The second image data includes point data and / or mask, that is, the algorithm results that need to be restored include two types, point data and mask data.
[0076] The second affine transformation matrix includes a first transformation matrix and / or a second transformation matrix, the first transformation matrix is used to map the point data to the first coordinate system, and the second transformation matrix is used to map the mask to the first coordinate system.
[0077] In the present disclosure, the first transformation matrix is used to map the point data to the first coordinate system, and the first transformation matrix is the first affine transformation matrix. In this way, the point data can be mapped to the first coordinate system by the following formula:
[0078] Where M represents the first transformation matrix, [x, y, z] T Represents the point data in the coordinate system where the algorithm result is located (i.e., the second coordinate system below), [x , ,y , ,z , ] T represents [x,y,z] T Point data in the first coordinate system.
[0079] In addition, the second transformation matrix in the present disclosure is determined by the following method:
[0080] Acquire first size data of the image data to be processed in the first coordinate system, second size data of the second image data in the second coordinate system, and a transformation matrix of the mask output by the preset image processing model;
[0081] The second coordinate system refers to the coordinate system of the image after image pre-processing operations, and is also the coordinate system corresponding to the algorithm result. The first size data may include the length and width of the image corresponding to the first coordinate system before the image pre-processing operations. The second size data refers to the length and width of the image corresponding to the algorithm result in the second coordinate system.
[0082] The second transformation matrix is determined according to the first affine transformation matrix, the first size data, the second size data, and the transformation matrix of the mask output by the preset image processing model.
[0083] For example, since the algorithm results output by the vision algorithm are generally non-normalized, while the transformation matrix output by the script engine is normalized, it is necessary to perform denormalization on the first affine transformation matrix output by the script engine. This can be achieved in the following way:
[0084] Where dstW and dstH represent the width and length of the image in the second coordinate system, srcW and srcH represent the width and length of the image in the first coordinate system, M represents the first affine transformation matrix, and M′ is the product of the three matrices on the left side of the equal sign.
[0085] Assuming a point v in the first coordinate system (i.e. the original coordinate system), we have the equation: T′M′ -1 v=T″v, therefore, T″=T′M′ -1 .
[0086] Among them, T′ represents the transformation matrix of the mask output by the preset image processing model, M′ -1 represents the inverse matrix of matrix M′, and T″ is the second transformation matrix.
[0087] In step S1042, third image data corresponding to the second image data in the first coordinate system is determined according to the second affine transformation matrix.
[0088] In this step, the second image data can be mapped to the first coordinate system through the second affine transformation matrix based on affine transformation theory to obtain the third image data.
[0089] In step S1043 , image rendering is performed according to the third image data.
[0090] In this step, the third image data may be input into a rendering engine, and image rendering may be implemented by the rendering engine.
[0091] FIG5 is a block diagram of an image processing apparatus according to an exemplary embodiment. As shown in FIG5 , the apparatus includes:
[0092] An acquisition module 501 is used to acquire image data to be processed;
[0093] A first image processing module 502 is configured to perform image pre-processing on the image data to be processed using a pre-configured script to obtain first image data, wherein the script is configured to define execution logic for the image pre-processing;
[0094] The second image processing module 503 is configured to input the first image data into a preset image processing model to obtain second image data.
[0095] Optionally, the image data to be processed includes change data relative to the original acquired image, and the first image processing module 502 is used to determine the task type of image pre-processing to be performed currently through the script based on the change data, and different task types correspond to different image pre-processing operations; according to the task type, the image data to be processed is pre-processed through the script to obtain first image data.
[0096] Optionally, the first image processing module 502 is used to generate a first affine transformation matrix corresponding to the task type through the script based on the image data to be processed; and drive the rendering engine through the script to perform image pre-processing on the image data to be processed according to the first affine transformation matrix to obtain the first image data.
[0097] Optionally, the task type includes: cropping and / or mirror flipping.
[0098] Optionally, the first image processing module 502 is configured to determine the task type through the script according to the change data and the model type of the preset image processing model.
[0099] Optionally, the first image processing module 502 is configured to perform image pre-processing on the image data to be processed by using the pre-configured script on a graphics processing unit (GPU) to obtain the first image data.
[0100] Optionally, FIG6 is a block diagram of an image processing device according to the embodiment shown in FIG5 . As shown in FIG6 , the device further includes:
[0101] The image rendering module 504 is configured to perform image rendering according to the second image data.
[0102] Optionally, the image rendering module 504 is used to determine a second affine transformation matrix based on the first affine transformation matrix, and the second affine transformation matrix is used to map the second image data to a first coordinate system corresponding to the image data to be processed; determine third image data corresponding to the second image data in the first coordinate system based on the second affine transformation matrix; and perform image rendering based on the third image data.
[0103] Optionally, the second image data includes point data and / or a mask, and the second affine transformation matrix includes a first transformation matrix and / or a second transformation matrix, the first transformation matrix is used to map the point data to the first coordinate system, and the second transformation matrix is used to map the mask to the first coordinate system.
[0104] Optionally, the first transformation matrix is the first affine transformation matrix.
[0105] Optionally, the second transformation matrix is determined by:
[0106] Acquire first size data of the image data to be processed in the first coordinate system, second size data of the second image data in the second coordinate system, and a transformation matrix of the mask output by the preset image processing model;
[0107] The second transformation matrix is determined according to the first affine transformation matrix, the first size data, the second size data, and the transformation matrix of the mask output by the preset image processing model.
[0108] Reference is now made to FIG7 , which illustrates a schematic diagram of the structure of an electronic device suitable for implementing an embodiment of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG6 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0109] As shown in Figure 7, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0110] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0111] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0112] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0113] In some embodiments, the client can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can interconnect with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0114] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0115] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain image data to be processed; perform image pre-processing on the image data to be processed through a pre-configured script to obtain first image data, and the script is used to define the execution logic of the image pre-processing; and obtain second image data after inputting the first image data into a preset image processing model.
[0116] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0118] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, an acquisition module may also be described as a "module for acquiring image data."
[0119] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0120] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0121] According to one or more embodiments of the present disclosure, Example 1 provides an image processing method, including:
[0122] Obtain image data to be processed;
[0123] Performing image pre-processing on the image data to be processed by a pre-configured script to obtain first image data, wherein the script is used to define the execution logic of the image pre-processing;
[0124] After the first image data is input into a preset image processing model, the second image data is obtained.
[0125] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the image data to be processed includes change data relative to the original captured image, and performing image pre-processing on the image data to be processed using a preconfigured script to obtain the first image data includes:
[0126] Determining the type of image pre-processing task to be currently performed through the script according to the change data, where different task types correspond to different image pre-processing operations;
[0127] Perform image pre-processing on the image data to be processed through the script according to the task type to obtain first image data.
[0128] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, wherein performing image pre-processing on the image data to be processed by the script according to the task type to obtain the first image data includes:
[0129] Generating a first affine transformation matrix corresponding to the task type through the script according to the image data to be processed;
[0130] The script drives a rendering engine to perform image pre-processing on the image data to be processed according to the first affine transformation matrix to obtain the first image data.
[0131] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 2, wherein the task type includes: cropping and / or mirror flipping.
[0132] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 2, wherein determining the type of image pre-processing task to be currently performed by the script based on the change data includes:
[0133] The task type is determined by the script according to the change data and the model type of the preset image processing model.
[0134] According to one or more embodiments of the present disclosure, Example 6 provides the method of any one of Examples 1-5, wherein performing image pre-processing on the image data to be processed using a preconfigured script to obtain the first image data includes:
[0135] The image data to be processed is subjected to image pre-processing by using the pre-configured script on a graphics processing unit (GPU) to obtain the first image data.
[0136] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 3, further comprising:
[0137] Image rendering is performed according to the second image data.
[0138] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7, wherein performing image rendering according to the second image data includes:
[0139] determining a second affine transformation matrix according to the first affine transformation matrix, where the second affine transformation matrix is used to map the second image data to a first coordinate system corresponding to the image data to be processed;
[0140] determining third image data corresponding to the second image data in the first coordinate system according to the second affine transformation matrix;
[0141] Image rendering is performed according to the third image data.
[0142] According to one or more embodiments of the present disclosure, Example 9 provides the method of Example 8, wherein the second image data includes point data and / or a mask, the second affine transformation matrix includes a first transformation matrix and / or a second transformation matrix, the first transformation matrix is used to map the point data to the first coordinate system, and the second transformation matrix is used to map the mask to the first coordinate system.
[0143] According to one or more embodiments of the present disclosure, Example 10 provides the method of Example 9, wherein the first transformation matrix is the first affine transformation matrix.
[0144] According to one or more embodiments of the present disclosure, Example 11 provides the method of Example 9, wherein the second transformation matrix is determined by:
[0145] Acquire first size data of the image data to be processed in the first coordinate system, second size data of the second image data in the second coordinate system, and a transformation matrix of the mask output by the preset image processing model;
[0146] The second transformation matrix is determined according to the first affine transformation matrix, the first size data, the second size data, and the transformation matrix of the mask output by the preset image processing model.
[0147] According to one or more embodiments of the present disclosure, Example 12 provides an image processing apparatus, the apparatus comprising:
[0148] An acquisition module, used for acquiring image data to be processed;
[0149] a first image processing module, configured to perform image pre-processing on the image data to be processed using a pre-configured script to obtain first image data, wherein the script is configured to define an execution logic for the image pre-processing;
[0150] The second image processing module is used to input the first image data into a preset image processing model to obtain second image data.
[0151] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in any one of Examples 1-11 when executed by a processing device.
[0152] According to one or more embodiments of the present disclosure, Example 14 provides an electronic device, including:
[0153] a storage device having a computer program stored thereon;
[0154] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-11.
[0155] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0156] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0157] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. An image processing method, comprising: Obtain image data to be processed; Performing image pre-processing on the image data to be processed by a pre-configured script to obtain first image data, wherein the script is used to define execution logic of the image pre-processing; After the first image data is input into a preset image processing model, second image data is obtained.
2. The method according to claim 1, wherein The image data to be processed includes change data relative to the original captured image, and the image pre-processing is performed on the image data to be processed using a pre-configured script to obtain the first image data, including: Determining the type of image pre-processing task to be currently performed through the script according to the change data, where different task types correspond to different image pre-processing operations; Perform image pre-processing on the image data to be processed through the script according to the task type to obtain first image data.
3. The method according to claim 2, wherein: The performing image pre-processing on the image data to be processed by the script according to the task type to obtain the first image data includes: Generating a first affine transformation matrix corresponding to the task type through the script according to the image data to be processed; The script drives a rendering engine to perform image pre-processing on the image data to be processed according to the first affine transformation matrix to obtain the first image data.
4. The method according to claim 2 or 3, wherein: The task types include: cropping and / or mirror flipping.
5. The method according to any one of claims 2 to 4, wherein: Determining the type of image pre-processing task to be currently performed by the script according to the change data includes: The task type is determined by the script according to the change data and the model type of the preset image processing model.
6. The method according to any one of claims 1 to 5, wherein: The performing image pre-processing on the image data to be processed by using a pre-configured script to obtain the first image data includes: The image data to be processed is subjected to image pre-processing by using the pre-configured script on a graphics processing unit (GPU) to obtain the first image data.
7. The method according to claim 3, further comprising: Image rendering is performed according to the second image data.
8. The method according to claim 7, wherein: The performing image rendering according to the second image data comprises: determining a second affine transformation matrix according to the first affine transformation matrix, where the second affine transformation matrix is used to map the second image data to a first coordinate system corresponding to the image data to be processed; determining third image data corresponding to the second image data in the first coordinate system according to the second affine transformation matrix; Image rendering is performed according to the third image data.
9. The method according to claim 8, wherein The second image data includes point data and / or a mask, and the second affine transformation matrix includes a first transformation matrix and / or a second transformation matrix. The first transformation matrix is used to map the point data to the first coordinate system, and the second transformation matrix is used to map the mask to the first coordinate system.
10. The method according to claim 9, wherein: The first transformation matrix is the first affine transformation matrix.
11. The method according to claim 9 or 10, wherein: The second transformation matrix is determined by: Acquire first size data of the image data to be processed in the first coordinate system, second size data of the second image data in the second coordinate system, and a transformation matrix of the mask output by the preset image processing model; The second transformation matrix is determined according to the first affine transformation matrix, the first size data, the second size data, and the transformation matrix of the mask output by the preset image processing model.
12. An image processing apparatus, comprising: an acquisition module, configured to acquire image data to be processed; a first image processing module configured to perform image pre-processing on the image data to be processed using a pre-configured script to obtain first image data, wherein the script is used to define an execution logic of the image pre-processing; The second image processing module is configured to input the first image data into a preset image processing model to obtain second image data.
13. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processing device, the image processing method according to any one of claims 1 to 11 is implemented.
14. An electronic device comprising: a storage device storing a computer program; as well as A processing device is configured to execute the computer program in the storage device to implement the image processing method according to any one of claims 1 to 11.
Citation Information
Patent Citations
License plate character recognition method based on data enhancement and data generation
CN111382743A
Image processing method, device and system and computer equipment
CN113808147A
Expression driving method and device, electronic equipment and computer readable storage medium
CN114723736A
Eye ground sugar net image focus segmentation integration method and system
CN114998366A
Image processing method and device and electronic equipment
CN115205307A