Image processing method and device, electronic equipment and storage medium
By performing object recognition and position comparison in the original image and the deformed image, the problem of over-detection of image deformation is solved, and the detection accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202311557612.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art cannot effectively detect whether the image is overdeformed during the dimensional adjustment process, making it difficult to distinguish the image objects, hindering information transmission and reducing aesthetics.
By acquiring the original image and the deformed image, target recognition is performed separately to obtain the object position, and then the deformed image object is mapped into the original image, and position comparison is performed to determine the degree of deformation.
It improves the accuracy of image deformation detection, simplifies calculations, improves detection efficiency, and ensures that image objects can still be accurately identified after deformation.
Smart Images

Figure CN120020915A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer vision technology, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] When content creators edit images, they often resize the images by stretching or compressing them to meet the typesetting size requirements. During the image resizing process, if the transformation is not performed while maintaining the original ratio of the image, it may cause excessive stretching or compression of the object areas in the image. This can make the objects in the image difficult to distinguish, hinder the transmission and expression of image information, and also reduce the aesthetic feeling of the image.
[0003] The existing technologies are mainly used to detect the key points of objects in images and identify the identities of objects, and cannot detect whether the objects are overly deformed. Currently, there is no method to solve the problem of detecting excessive deformation of human faces. Summary of the Invention
[0004] Embodiments of this application provide an image processing method, apparatus, electronic device, computer program product, and computer-readable storage medium, which can improve the accuracy of detecting excessive deformation of images.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] Embodiments of this application provide an image processing method, the method comprising:
[0007] Obtain a first image and a second image, wherein the second image is obtained by adjusting the first image;
[0008] Perform object recognition on the first image and the second image respectively to obtain the first positions of at least one first object in the first image and the second positions of at least one second object in the second image;
[0009] Map the second positions of the at least one second object into the first image to obtain the third positions of the at least one second object;
[0010] Compare the third positions of the at least one second object with the first positions of the at least one first object to determine the first object and the second object with a matching relationship;
[0011] Based on the first object and the second object with the matching relationship, determine the deformation detection result of the second image relative to the first image.
[0012] Embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0013] An image acquisition module, configured to acquire a first image and a second image, where the second image is obtained by adjusting the first image;
[0014] A target recognition module, configured to perform target recognition on the first image and the second image respectively, to obtain the first positions of at least one first object in the first image and the second positions of at least one second object in the second image;
[0015] A position mapping module, configured to map the second positions of the at least one second object into the first image, to obtain the third positions of the at least one second object;
[0016] A position comparison module, configured to compare the third positions of the at least one second object with the first positions of the at least one first object, to determine the first object and the second object with a matching relationship;
[0017] A result determination module, configured to determine a deformation detection result of the second image relative to the first image based on the first object and the second object with the matching relationship.
[0018] An embodiment of the present application provides an electronic device, where the electronic device includes:
[0019] A memory, configured to store computer-executable instructions;
[0020] A processor, configured to implement the image processing method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.
[0021] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, which are configured to implement the image processing method provided by the embodiment of the present application when being executed by a processor.
[0022] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where the computer program or computer-executable instructions implement the image processing method provided by the embodiment of the present application when being executed by a processor.
[0023] The embodiments of the present application have the following beneficial effects:
[0024] By performing object recognition in the first image and the second image, the task of detecting the deformation of the second image relative to the first image is converted into the task of detecting the deformation of the first object in the first image relative to the second object that matches (i.e., there is a matching relationship) in the second image, which simplifies the calculation of image deformation detection and improves the detection efficiency; by mapping the second position of the second object in the second image to determine the first object and the second object with a matching relationship, the comparison of each group of objects in the two images is converted onto the same image, and the change degree of the determined matching objects between different images is more accurately reflected through the position distance, enhancing the accuracy of image deformation detection. Description of the Drawings
[0025] Figure 1 is a schematic structural diagram of the image processing system 100 provided by an embodiment of the present application;
[0026] Figure 2 is a schematic structural diagram of the electronic device provided by an embodiment of the present application;
[0027] Figure 3A is a first flowchart of the image processing method provided by an embodiment of the present application;
[0028] Figure 3B is a second flowchart of the image processing method provided by an embodiment of the present application;
[0029] Figure 3C is a third flowchart of the image processing method provided by an embodiment of the present application;
[0030] Figure 3D is a fourth flowchart of the image processing method provided by an embodiment of the present application;
[0031] Figure 3E is a fifth flowchart of the image processing method provided by an embodiment of the present application;
[0032] Figure 3F is a sixth flowchart of the image processing method provided by an embodiment of the present application;
[0033] Figure 3G is a seventh flowchart of the image processing method provided by an embodiment of the present application;
[0034] Figure 3H is an eighth flowchart of the image processing method provided by an embodiment of the present application;
[0035] Figure 3I is a ninth flowchart of the image processing method provided by an embodiment of the present application;
[0036] Figure 4A is a processing flowchart of the first image provided by an embodiment of the present application;
[0037] Figure 4B is the processing flowchart of the second image provided by the embodiments of the present application;
[0038] Figure 4C is the schematic diagram of determining the object position provided by the embodiments of the present application;
[0039] Figure 4D is the structural schematic diagram of the convolutional neural network model provided by the embodiments of the present application;
[0040] Figure 5 is the schematic diagram of determining the degree of facial deformation provided by the embodiments of the present application;
[0041] Figure 6 is the schematic diagram of the result of face detection provided by the embodiments of the present application;
[0042] Figure 7 is the schematic diagram of the image coordinate system provided by the embodiments of the present application;
[0043] Figure 8 is the schematic diagram of an example of the detection result provided by the embodiments of the present application. Detailed implementation manners
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0045] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0046] If similar descriptions such as "first / second" appear in the application documents, the following description is added. In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0047] In the embodiments of the present application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations during actual application, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.
[0048] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used in the embodiments of this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0049] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0050] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described, and the nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0051] 1) Open Source Computer Vision Library (Opencv), an open-source library widely used in the field of computer vision, and the cascade classifier is one of its core functions.
[0052] 2) Cascade classifier, an object detection algorithm based on machine learning, which can quickly and accurately identify a specified target object in an image or video. The cascade classifier is implemented based on Haar features and the AdaBoost algorithm. Haar features are local features based on the grayscale values of an image, which can effectively describe the shape and texture information of the target object. The AdaBoost algorithm is an ensemble learning algorithm that constructs a strong classifier by combining multiple weak classifiers to improve the accuracy of classification.
[0053] 3) Machine learning model. The concept of machine learning is to train the model by inputting a large amount of training data so that the model can master the potential laws contained in the data, and then accurately classify or predict the newly input data. A machine learning model is essentially a function that accepts data as input and generates output. In the process of using known data to predict unknown data, the model should not only perform well on known samples but also have a similar performance on unknown samples.
[0054] 4) Bounding box, that is, a rectangle circumscribing the contour of the target in the image, which is the smallest rectangle that can accommodate the target.
[0055] 5) Pooling: In image processing, due to the presence of a large amount of redundant information in the image, the statistical information (such as the maximum value or the mean value, etc.) of a certain regional sub-block can be used to characterize the spatial distribution pattern presented by all pixel points in this region, so as to replace the values of all pixel points in the regional sub-block. The pooling operation compresses the feature map of the convolution result. On the one hand, it makes the feature map smaller and simplifies the computational complexity of the network; on the other hand, it performs feature compression, extracts the main features, and retains the main information in the feature map.
[0056] When content creators edit images, they often adjust the image size by stretching or compressing to meet the typesetting size requirements. During the process of adjusting the image size, if the transformation is not carried out while maintaining the original proportion of the image, it may cause the object area in the image to be over-stretched or over-compressed. This will make it difficult to distinguish the objects in the image, hinder the transmission and expression of image information, and at the same time reduce the aesthetic feeling of the image.
[0057] The related technologies only focus on the detection of the positions of objects in the image and the positions of key points in the objects and the recognition of the objects, and do not solve the problem of detecting excessive deformation during the process of adjusting the image size.
[0058] Based on the above analysis, the applicant found that the image processing methods of the related technologies cannot detect whether the image is excessively deformed after deformation. In view of the above problems, the embodiments of the present application provide an image processing method, which can improve the accuracy of detecting excessive image deformation.
[0059] The embodiments of the present application provide an image processing method, device, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of detecting excessive image deformation. The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), smart phones, smart speakers, smart watches, smart TVs, vehicle terminals, and aircraft, or can be implemented as a server. The following will describe the exemplary applications when the electronic device is implemented as a server.
[0060] See Figure 1 , Figure 1 FIG. is a schematic diagram of the architecture of the image processing system 100 provided by the embodiments of the present application. To support an image processing application, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0061] The terminal 400 obtains the first image and the second image, and can display the first image and the second image on the graphical interface 410.
[0062] The server 200 is used to obtain a first image and a second image, detect a deformation detection result of the second image relative to the first image, and send the deformation detection result to the terminal 400.
[0063] In some embodiments, the terminal 400 can independently complete the deformation detection task. For example, the terminal 400 is used to obtain a first image and a second image, determine the deformation detection result of the second image relative to the first image depending on the computing power of the terminal 400 itself, and display the deformation detection result on the graphical interface 410.
[0064] An example of the terminal 400 obtaining the first image and the second image is described below.
[0065] In some embodiments, in a video playback scenario, the terminal 400 is used to play a video. In response to a screenshot operation on a certain frame of the video, the obtained screenshot is used as the first image. After secondary editing, the first image is changed into the second image. At this time, the terminal 400 or the server 200 detects the deformation detection result of the second image relative to the first image.
[0066] In some embodiments, in an instant messaging scenario, the terminal 400 is used to run an instant messaging client, obtain a first image during the communication process, edit the first image through a built-in picture editing function to obtain the second image, and then the terminal 400 automatically analyzes the deformation, or in response to a trigger operation on a detection button, detects whether deformation occurs from the first image to the second image.
[0067] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0068] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0069] The embodiments of the present application can be implemented through artificial intelligence technology. Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0070] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0071] In some embodiments, in an in-vehicle scenario without GPS navigation, when the in-vehicle satellite signal is unavailable, the in-vehicle terminal is used to collect images of the current road, which include multiple road elements, such as road signs, warning signs, etc.; compare the collected images of the current road with multiple frames of high-definition maps near the last positioning position to form a deformation detection result. Use the high-definition map with an undeformed deformation detection result as the starting frame map for the current navigation and perform inertial navigation through a gyroscope.
[0072] See Figure 2 , Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Figure 2 The electronic device 500 shown can be Figure 1 the terminal 400 or the server 200 in Figure 2 . The electronic device 500 includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in
[0073] The processor 510 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0074] The user interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0075] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 550 optionally includes one or more storage devices that are physically located away from the processor 510.
[0076] The memory 550 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0077] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.
[0078] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0079] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.;
[0080] A presentation module 553 for enabling the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with the user interface 530 (e.g., a display screen, a speaker, etc.);
[0081] An input processing module 554 for detecting and translating one or more user inputs or interactions from one of one or more input devices 532.
[0082] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2 Shown is an image processing device 555 stored in the memory 550, which may be software in the form of a program and a plug-in, etc., including the following software modules: an image acquisition module 5551, a target recognition module 5552, a position mapping module 5553, a position comparison module 5554, and a result determination module 5555. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0083] In other embodiments, the device provided by the embodiments of the present application may be implemented in hardware. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the image processing method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0084] In some embodiments, the terminal or the server may implement the image processing method provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions may be commands at the microprogram level, machine instructions, or software instructions. The computer program may be a native program or software module in the operating system; it may be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP and a video playback APP; it may also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer-executable instructions may be instructions in any form, and the above computer programs may be application programs, modules, or plug-ins in any form.
[0085] The exemplary applications and implementations of the server provided in the embodiments of the present application will be combined to illustrate the image processing method provided in the embodiments of the present application.
[0086] See Figure 3A , Figure 3A FIG. is the first flowchart of the image processing method provided in the embodiments of the present application, which can be executed by the above terminal or server, and will be described in combination with Figure 3A the steps shown.
[0087] In step 101, a first image and a second image are obtained, where the second image is obtained by adjusting the first image.
[0088] In some embodiments, the first image may be an unadjusted image, that is, the original image collected by the image sensor, or an adjusted image, such as an image after stretching the original image. The second image may be automatically stretched by the terminal according to the size of the screen for full-screen filling of the first image, or may be formed by editing the first image through an image editing tool.
[0089] Exemplarily, see Figure 4A , Figure 4A is the processing flowchart of the first image provided in the embodiments of the present application. See Figure 4B , Figure 4B is the processing flowchart of the second image provided in the embodiments of the present application. Among them, Figure 4A the screen 601 in is the first image, Figure 4B the screen 701 in is the second image adjusted based on the first image.
[0090] In step 102, target recognition is respectively performed on the first image and the second image to obtain the first position of at least one first object in the first image and the second position of at least one second object in the second image.
[0091] Exemplarily, the object to be recognized can be any object that can be imaged, such as a face, a limb, an item (a water cup, a box), etc.
[0092] In some embodiments, referring to Figure 3B , Figure 3B is the second process schematic diagram of the image processing method provided by the embodiments of the present application. Figure 3A Step 102 of Figure 3B can be implemented by steps 1021 to 1026 of
[0093] In step 1021, a plurality of first candidate regions of the first image and a plurality of second candidate regions of the second image are respectively extracted.
[0094] In some embodiments, a plurality of candidate regions can be extracted from an image through a target localization algorithm such as the Fast R-CNN algorithm, a single-stage algorithm (YOLO, You Only Look Once), a single-step multi-box object detection algorithm (SSD, Single Shot MultiBox Detector), etc.
[0095] Exemplarily, referring to Figure 4C , Figure 4C is the schematic diagram of determining the object position provided by the embodiments of the present application. In Figure 4C , the screen 801 is the first image, the region included in the bounding box 810 is the first candidate region, and the second candidate region of the second image is like the first candidate region of the first image shown in Figure 4C .
[0096] In step 1022, feature extraction is respectively performed on the first image and the second image to obtain a first feature map of the first image and a second feature map of the second image.
[0097] In some embodiments, a feature map of an image can be obtained through a convolutional neural network, wherein the feature map of the first image or the second image can be implemented through single-layer convolution or multi-layer convolution.
[0098] In some embodiments, padding can be performed on the first feature map so that in subsequent steps during mapping, there will be no situation of exceeding the boundary of the first feature map. The size of the first feature map is the same as the size of the first image, and the size of the second feature map is the same as the size of the second image.
[0099] Exemplarily, as shown in Figure 4C , the first image 801 corresponds to the first feature map 802, and the second feature map of the second image is like the first feature map of the first image shown in Figure 4C .
[0100] In step 1023, determine the first feature regions corresponding to the multiple first candidate regions on the first feature map respectively.
[0101] In some embodiments, referring to Figure 3C , Figure 3C is the third process schematic diagram of the image processing method provided by the embodiments of the present application. For Figure 3B each first candidate region in step 1023, it can be implemented through Figure 3C steps 10231 to 10234 of step 1023, which are specifically described below.
[0102] In step 10231, obtain the multiple endpoint coordinates of the first candidate region.
[0103] In some embodiments, since the first candidate region is a rectangle and the shape of the rectangle can be uniquely determined by the two endpoints of the diagonal, the coordinates of only the two endpoints of the diagonal of the first candidate region can be obtained.
[0104] Exemplarily, as Figure 4C shown, taking the lower left corner of the first image 801 as the origin, the lower edge of the first image as the positive X-axis, and the right edge of the first image as the positive Y-axis to establish a coordinate system. Obtain the coordinates of the two endpoints of the diagonal of the first candidate region 810, that is, M(x5, y6) and N(x6, y5).
[0105] In step 10232, determine the first ratio of the size of the first image to the size of the first feature map.
[0106] Exemplarily, as Figure 4C shown, the length of the first image 801 is W and the width is H, and the length of the first feature map 802 is W' and the width is H'. The first ratio of the size of the first image to the size of the first feature map can be divided into the length ratio between the length of the first image and the length of the first feature map, and the width ratio between the width of the first image and the width of the first feature map. Among them, the length ratio is W / W', and the width ratio is H / H'.
[0107] In step 10233, determine the multiple second ratios of the endpoint coordinates of the multiple first candidate regions to the first ratio respectively, and use the multiple second ratios as the multiple endpoint coordinates of the first feature region.
[0108] In some embodiments, perform the following processing for each endpoint coordinate of the first candidate region: divide the endpoint coordinate of the first candidate region by the first ratio, and use the obtained new endpoint coordinate as the endpoint coordinate of the first feature region.
[0109] Exemplarily, as Figure 4CAs shown in the figure, select the endpoint coordinates of the diagonal of the first candidate region 810, that is, M(x5, y6) and N(x6, y5). Use x5 / (W / W′) to represent the abscissa of the endpoint M′ at the upper left corner of the first feature region 820, and use y6 / (H / H′) to represent the ordinate of the endpoint M′ at the upper left corner of the first feature region 820. Let x5 / (W / W') = x5', y6 / (H / H′) = y6', then the coordinates of M′ are (x5′, y6′). Similarly, the coordinates of the lower right corner endpoint N′ of the first feature region 820 can be obtained as (x6′, y5′).
[0110] In step 10234, the region corresponding to the multiple endpoint coordinates of the first feature region on the first feature map is determined as the first feature region corresponding to the first candidate region on the first feature map.
[0111] Exemplarily, as Figure 4C shown, the region 820 enclosed by M′ and N′ as the diagonal endpoints is the first feature region corresponding to the first candidate region 810 on the first feature map 802.
[0112] Continue to refer to Figure 3B In step 1024, determine the second feature regions corresponding to the multiple second candidate regions on the second feature map respectively.
[0113] In some embodiments, for each second candidate region, perform the following operations: obtain the multiple endpoint coordinates of the second candidate region; determine the third ratio of the size of the second image to the size of the second feature map; determine the multiple fourth ratios of the endpoint coordinates of the multiple second candidate regions to the third ratio respectively, and use the multiple fourth ratios as the multiple endpoint coordinates of the second feature region; the region corresponding to the multiple endpoint coordinates of the second feature region on the second feature map is determined as the second feature region corresponding to the second candidate region on the second feature map.
[0114] Exemplarily, for the determination method of the second feature region, refer to Figure 4C the principle of determining the first feature region shown in the figure.
[0115] In step 1025, perform pooling operations on the first feature region and the second feature region respectively to obtain the first feature region vector and the second feature region vector.
[0116] In some embodiments, in image processing, due to the existence of a large amount of redundant information in the image, the statistical information (such as the maximum value or the mean value, etc.) of a certain regional sub-block can be used to characterize the spatial distribution pattern presented by all pixel points in this region, so as to replace the values of all pixel points in the regional sub-block. This is the pooling operation in the convolutional neural network. The pooling operation compresses the feature map of the convolutional result. On the one hand, it makes the feature map smaller and simplifies the network calculation complexity; on the other hand, it performs feature compression, extracts the main features, and retains the main information in the feature map.
[0117] In step 1026, regression processing is respectively performed on the first feature region vector and the second feature region vector to obtain the first position of at least one object in the first image and the second position of at least one second object in the second image.
[0118] In some embodiments, regression processing can also be performed through the CascadeClassifier in the Open Source Computer Vision Library (Opencv) to obtain the positions of the first object and the second object. The training process of the CascadeClassifier is divided into two stages: positive sample training and negative sample training. Positive sample training means that through a series of image samples, Haar features are extracted and weak classifiers are trained, and generally the AdaBoost algorithm is adopted. Negative sample training means that through a series of image samples that are not related to the target object, Haar features are also extracted and weak classifiers are trained. Through the training of these two stages, the CascadeClassifier can learn the features of the target object and has high recognition accuracy.
[0119] When the CascadeClassifier performs target detection, it adopts the method of sliding window. It divides the image into multiple windows of different sizes and classifies and judges each window to determine whether the window contains the target object. In the process of classification and judgment, the CascadeClassifier adopts the idea of cascade, that is, the classification process is divided into multiple levels, and each level consists of multiple weak classifiers. In each level, only when the window passes all the weak classifiers can it enter the classification and judgment of the next level. Through this cascading method, the CascadeClassifier can quickly exclude a large number of negative samples, thereby improving the efficiency of target detection. By judging the window, the first position of the first object and the second position of the second object can be determined according to the position of the window.
[0120] In some embodiments, the offset and scaling factor can be regressed by a bounding box regressor to predict the position of an object. Bounding box regression is mainly because the intersection over union (IoU) between the candidate box region or anchor box and the ground truth bounding box is not high, and the position of the object cannot be detected accurately. At this time, a bounding box close to the ground truth bounding box can be obtained through bounding box regression. A bounding box is generally represented by a 4D vector (x, y, w, h). The first two parameters represent the coordinates of the center of the bounding box, and the last two parameters represent the width and height of the bounding box. The goal of regression is to find a window that makes the original box as close as possible to the ground truth window after mapping. The regression principle is to first perform translation to make the box centers coincide as much as possible, and then perform scale scaling to make the areas close, so as to obtain a bounding box closest to the ground truth bounding box. The position of the first object or the second object is represented by the coordinates of this bounding box.
[0121] Continue to refer to Figure 3A , in step 103, the second positions of at least one second object are mapped into the first image to obtain the third positions of at least one second object.
[0122] In some embodiments, refer to Figure 3D , Figure 3D is the fourth process schematic diagram of the image processing method provided by the embodiments of the present application. For the third position of each second object, it can be determined through Figure 3D steps 1031 to 1033, which will be specifically described below.
[0123] In step 1031, determine the ratio of the size of the first image to the size of the second image.
[0124] In some embodiments, the size includes the width and height of the image. Therefore, the width of the first image, the height of the first image, the width of the second image, and the height of the second image can be determined. Determine the width ratio of the width of the first image and the width of the second image, and the height ratio of the height of the first image and the height of the second image, respectively.
[0125] Exemplarily, as Figure 4A and Figure 4B shown, assume that the width of the first image 601 is W1, the height of the first image is H1, the width of the second image 701 is W2, and the height of the second image 701 is H2. Then the width ratio of the first image to the second image is W1 / W2, and the height ratio of the first image to the second image is H1 / H2.
[0126] In step 1032, determine the product of the endpoint coordinates of the bounding box of the second object at the second position and the ratio, and use the product as the coordinates of the bounding box of the second object at the third position.
[0127] In some embodiments, calculate the product of the ratio of the abscissa of each endpoint of the bounding box of the second object at the second position to the width as the abscissa of each endpoint of the bounding box of the second object at the third position.
[0128] For example, as Figure 4B shown, the screen 702 is the second image, and the endpoint coordinates of the bounding box 460 are taken as the second position of the second object. The screen 703 is a schematic diagram of the third position obtained after mapping the second position of the second object. Among them, the endpoint coordinates of the bounding box 461 are the third position of the second object. Taking the diagonal endpoints C and D of the bounding box 460 as an example, the coordinates of endpoint C are (x1′, y2′), the coordinates of endpoint D are (x2′, y1′), and the product of endpoint C and the ratio of the width is x1′*W1 / W2. Let then the abscissa of the upper left corner endpoint C′ of the bounding box 461 of the second object is x3. Similarly, the abscissa x4 of the lower right corner endpoint D′ of the bounding box 461 of the second object can be obtained.
[0129] In some embodiments, calculate the product of the ratio of the ordinate of each endpoint of the bounding box of the second object at the second position to the height as the ordinate of each endpoint of the bounding box of the second object at the third position.
[0130] For example, as Figure 4B shown, the screen 702 is the second image, and the endpoint coordinates of the bounding box 460 are taken as the second position of the second object. The screen 703 is a schematic diagram of the second image mapped back to the first image. Among them, the bounding box 461 is the third position of the second object. Taking the diagonal endpoints C and D of the bounding box 460 as an example, the coordinates of endpoint C are (x1′, y2′), the coordinates of endpoint D are (x2′, y1′), and the product of endpoint C and the ratio of the height is y2′*H1 / H2. then the ordinate of the upper left corner endpoint C′ of the bounding box 461 of the second object is y4. Similarly, the ordinate y3 of the lower right corner endpoint D′ of the bounding box 461 of the second object can be obtained.
[0131] In step 1033, take the endpoint coordinates of the bounding box of the second object at the third position as the third position of at least one second object.
[0132] For example, as Figure 4B shown, the third position of the second object can be determined according to the endpoint coordinates of the two endpoints of the diagonal of the bounding box. That is, the coordinates of the upper left corner endpoint C′ of the bounding box 461 of the second object are (x3, y4), and the coordinates of the lower right corner endpoint D′ of the bounding box 461 of the second object are (x4, y3). Take the coordinates of C′ and D′ as the third position of the second object.
[0133] Continue to refer to Figure 3A, in step 104, the third positions of at least one second object are compared with the first positions of at least one first object to determine the first and second objects with a matching relationship.
[0134] In some embodiments, referring to Figure 3E , Figure 3E is the fifth flowchart of the image processing method provided by the embodiments of the present application. Figure 3A Step 104 of Figure 3E can be implemented by steps 1041 to 1045 of
[0135] In step 1041, a first set is determined, where the first set includes the first positions of each first object.
[0136] In some embodiments, there are multiple first objects in the first image, and a first set is constructed according to the first positions of the multiple first objects.
[0137] For example, if the first positions of at least one first object in the first image can be expressed as F = (x1, y1, x2, y2), where x1 and y1 respectively represent the abscissa and ordinate of the upper left corner endpoint of the bounding box of the first object, and x2 and y2 respectively represent the abscissa and ordinate of the lower right corner endpoint of the bounding box of the first object, then the first set is expressed as S f = {F1, F2, F3,...}, where F1, F2, F3, etc. represent the first positions of multiple first objects in the first image.
[0138] In step 1042, a second set is determined, where the second set includes the third positions of each second object.
[0139] In some embodiments, there are multiple second objects in the second image. When mapping the multiple second objects back to the first image, there are multiple corresponding third positions, and a second set is constructed according to the third positions of the multiple second objects.
[0140] For example, the second set is expressed as S_f′ = {F1′, F2′, F3′,...}, where F1′, F2′, F3′, etc. represent the third positions of multiple second objects.
[0141] In some embodiments, for each first position in the first set, it is implemented through steps 1043 to 1045.
[0142] In step 1043, any one first position is selected from the first set as the position to be measured.
[0143] For example, the first set can be traversed in the order from left to right to process each first position in the first set. For example, for the first set Sf = {F1, F2, F3, ……}. First, take F1 as the object to be measured. After F1 is processed, then select F2 as the object to be measured, and so on, until all the first objects are processed as the positions to be measured.
[0144] In step 1044, determine the distance between each third position in the second set and the position to be measured.
[0145] In some embodiments, calculate the differences in the abscissa and ordinate between the position to be measured and the third position, calculate the sum of the squares of the abscissa difference and the square of the ordinate difference, and then take the square root after adding the two squares to obtain the distance between each third position and the position to be measured.
[0146] In step 1045, establish a matching relationship based on the second object corresponding to the third position with the smallest distance from the position to be measured and the first object corresponding to the position to be measured.
[0147] In some embodiments, select the position to be measured with the smallest distance from the third position. Then, there is a matching relationship between the first object corresponding to the position to be measured and the second object corresponding to the third position. Among them, the fact that there is a matching relationship between the first object and the second object actually means that the first object and the second object are actually the same object.
[0148] In the embodiment of the present application, map the second position of the second object back to the first image, and find the first object and the second object with a matching relationship by calculating the distance between the third position of the second object obtained by mapping and the first position of the first object. Convert the comparison of each group of objects in the two images to the same image, and more clearly and accurately determine the degree of change of the image through the position distance, improving the accuracy of image deformation detection.
[0149] In some embodiments, refer to Figure 3F , Figure 3F is the sixth process schematic diagram of the image processing method provided by the embodiment of the present application. Before step 104, execute Figure 3F steps 201 to 204 as follows for specific description.
[0150] In step 201, determine the size of the bounding box of the second object according to the endpoint coordinates of the bounding box of the second object.
[0151] Exemplarily, as Figure 4B shown, take the bounding box 460 as the bounding box of the second object. Among them, the endpoint coordinates of the two endpoints C and D of the diagonal of the bounding box 460 are endpoint C(x1′, y2′) and endpoint D(x2′, y1′) respectively. Then, the width of the bounding box 460 of the second object is x2′ - x1′, and the height is y2′ - y1′.
[0152] In step 202, obtain the size of the second image.
[0153] Exemplarily, as Figure 4B shown, the size includes width and height. The width of the bounding box 460 of the second object is x2′ - x1′, and the height is y2′ - y1′.
[0154] In step 203, determine the size ratio between the size of the bounding box of the second object and the size of the second image.
[0155] In some embodiments, respectively determine the width ratio between the width of the bounding box of the second object and the width of the second image, and determine the height ratio between the height of the bounding box of the second object and the height of the second image.
[0156] Exemplarily, as Figure 4B shown, the screen 701 is the second image. If the width of the second image 701 is W2, the height of the second image 701 is H2, the width of the bounding box 460 of the second object is x2′ - x1′, and the height is y2′ - y1′, then the width ratio between the width of the bounding box of the second object and the width of the second image is (x2′ - x1′) / W2, and the height ratio is (y2′ - y1′) / H2.
[0157] In step 204, filter out the bounding boxes with a size ratio less than the size ratio threshold.
[0158] In some embodiments, the size ratio threshold includes a width ratio threshold and a height ratio threshold. Filtering out the bounding box is actually to filter out smaller second objects. The size ratio threshold is the ratio between the size of the smallest recognizable bounding box (i.e., the smallest size of the object in the bounding box) and the size of the image including the smallest bounding box.
[0159] Exemplarily, as Figure 4B shown, if the width ratio threshold in the size ratio threshold is 0.5, the height ratio threshold is 0.6, the width ratio of the bounding box 480 is 0.3, and the height ratio threshold is 0.2, then filter out the bounding box 480.
[0160] In some embodiments, the size ratio threshold is determined by a machine learning model. The machine learning model can be a convolutional neural network model. Refer to Figure 4D , Figure 4D which is the structural schematic diagram of the convolutional neural network model provided by the embodiments of the present application. In Figure 4DAmong them, a convolutional neural network is a hierarchical feature extractor used to extract features at increasingly high levels. As the receptive field of the features becomes larger and larger, the features change from local to global. A convolutional neural network usually includes the following layers: an input layer, a convolutional layer, a pooling layer, and an output layer (fully connected layer + activation function layer (softmax layer)).
[0161] Exemplarily, other machine learning regression models can also be used to determine the size ratio threshold, such as: linear regression model, decision tree regression model, random forest regression model, neural network regression model, gradient boosting based on decision tree (LightGBM, Light Gradient Boosting Machine) regression model.
[0162] In some embodiments, refer to Figure 3G , Figure 3G is the seventh process schematic diagram of the image processing method provided by the embodiments of the present application. The training process of the machine learning model is implemented through Figure 3G Steps 301 to 304 are specifically described below.
[0163] In step 301, a plurality of original image samples and a plurality of adjusted image samples corresponding to the original image samples are obtained.
[0164] In some embodiments, the adjusted image sample is obtained by editing the original image sample, such as: image stretching, image cropping, image bending, etc.
[0165] In step 302, the true size ratio threshold calibrated for the original image sample is obtained.
[0166] In some embodiments, the true size ratio threshold can be set manually.
[0167] In step 303, a machine learning model is called based on the original image sample and the corresponding adjusted image sample to obtain a predicted size ratio threshold.
[0168] In some embodiments, the machine learning model is for performing a size ratio threshold prediction operation to obtain a predicted size ratio threshold.
[0169] In step 304, the loss between the true size ratio threshold and the predicted size ratio threshold is determined, and the parameters of the machine learning model are updated according to the loss to obtain an updated machine learning model.
[0170] In some embodiments, the loss between the true size ratio threshold and the predicted size ratio threshold can be determined by a loss function, such as: quadratic loss function, logarithmic loss function, absolute loss function, Hinge loss function, etc.
[0171] In some embodiments, the size ratio threshold is preset and obtained through a ratio threshold setting interface; in step 204, the preset size ratio threshold is obtained, where the size ratio threshold is obtained through a threshold setting operation, and the threshold setting operation is received through the threshold setting interface.
[0172] By setting the size ratio threshold in the embodiments of the present application, smaller objects in the image that do not have recognition value are filtered out, improving the accuracy of image detection.
[0173] In some embodiments, the first position of the first object includes the endpoint coordinates of the bounding box of the first object, such as the endpoint coordinates of any diagonal of the bounding box; before step 104, according to the endpoint coordinates of the bounding box of the first object, determine the size of the bounding box of the first object; determine the size of the first image; determine the size ratio between the size of the bounding box of the first object and the size of the first image; filter out the bounding boxes with a size ratio less than the preset size ratio threshold.
[0174] In some embodiments, refer to Figure 3H , Figure 3H is the eighth process schematic diagram of the image processing method provided by the embodiments of the present application. Before step 104, execute Figure 3H steps 401 to 402 of
[0175] In step 401, based on the first position, obtain the endpoint coordinates of each endpoint of the bounding box of the first object.
[0176] Exemplarily, as Figure 4A shown, in the screen 602, obtain the endpoint coordinates of each endpoint of the bounding boxes 420, 430, 440, 450, or obtain the coordinates of the two endpoints of the diagonal of each bounding box. For example, obtain the coordinates (x1, y1) of endpoint A and the coordinates (x2, y2) of endpoint B of the diagonal of the bounding box 420.
[0177] In step 402, for each first object, in response to the endpoint coordinates of any endpoint of the bounding box of the first object exceeding the coordinate range of the first image, filter out the bounding box of the first object.
[0178] Exemplarily, as Figure 4A shown, the horizontal coordinates of the upper left corner endpoint and the lower left corner endpoint of the bounding box 450 of the first object exceed the range of the coordinate axes of the first image 602, so the bounding box 450 is filtered out.
[0179] In some embodiments, the second position is the endpoint coordinates of the bounding box of the second object; before step 104, the endpoint coordinates of each endpoint of the bounding box of the second object are obtained based on the second position; for each second object, in response to the endpoint coordinates of any endpoint of the bounding box of the second object exceeding the coordinate range of the second image, the bounding box of the second object is filtered out, that is, the second object is filtered out.
[0180] Exemplarily, as Figure 4B shown, the horizontal coordinates of the upper left corner endpoint and the lower left corner endpoint of the bounding box 490 of the second object exceed the range of the coordinate axes of the second image 702, so the bounding box 490 is filtered out.
[0181] By filtering out the bounding box coordinates of the objects that exceed the image coordinate range in the embodiments of the present application, in fact, the incomplete objects are filtered out, reducing the calculation cost and improving the detection efficiency.
[0182] Continuing to refer to Figure 3A , in step 105, based on the first object and the second object with a matching relationship, the deformation detection result of the second image relative to the first image is determined.
[0183] In some embodiments, referring to Figure 3I , Figure 3I is the ninth process schematic diagram of the image processing method provided by the embodiments of the present application. Figure 3A Step 105 of Figure 3I can be implemented by steps 1051 to 1053 of
[0184] In step 1051, the first object and the second object with a matching relationship are respectively used as the first object to be measured and the second object to be measured.
[0185] Exemplarily, as Figure 4B shown, the second object included in the bounding box 460 in the screen 702 has a matching relationship with the first object included in the bounding box 420 in the screen 703, then the first object included in the bounding box 420 is used as the first object to be measured, and the second object included in the bounding box 460 is used as the second object to be measured.
[0186] In step 1052, in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to the difference threshold, stretching excessive is taken as the deformation detection result of the second image relative to the first image.
[0187] Exemplarily, as Figure 4B shown, in the screen 703, when the third position of the second object to be measured is the position of the bounding box 471, the first position of the first object to be measured is the position of the bounding box 430. Let the endpoint coordinates of the diagonal of the bounding box 471 be endpoint E(m1, n1) and F(m2, n2) respectively, and the endpoint coordinates of the diagonal of the bounding box 430 be endpoint G′(p1, q1) and H′(p2, q2) respectively. The difference threshold is set to a horizontal coordinate difference threshold of 1 and a vertical coordinate difference threshold of 2. If m2 - m1 > 1 or m2 - m1 = 1, then the deformation detection result of the second image relative to the first image is stretching excessive. Similarly, in any of the cases where p2 - p1 > 1 or p2 - p1 = 1 or n2 - n1 > 2 or n2 - n1 = 2 or q2 - q1 > 2 or q2 - q1 = 2, the deformation detection result of the second image relative to the first image is considered to be stretching excessive.
[0188] In some embodiments, obtain the quantity ratio between the quantity of the first object to be measured and the quantity of the first object; in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to the difference threshold, and the quantity ratio being less than the quantity ratio threshold, take stretching excessive as the deformation detection result of the second image relative to the first image.
[0189] Exemplarily, in the case where the position difference between the third position of the second object to be measured and the first position of the first object to be measured is greater than or equal to the difference threshold, if the quantity of the first object to be measured is 4, the quantity of the first object is 2, and the quantity ratio threshold is 3, at this time the quantity ratio is 2, and the quantity ratio is less than the quantity ratio threshold, then the deformation detection result of the second image relative to the first image is stretching excessive.
[0190] In step 1053, in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, take not stretching excessive as the deformation detection result of the second image relative to the first image.
[0191] Exemplarily, as Figure 4BAs shown, in the screen 703, when the third position of the second object to be measured is the position of the bounding box 461, the first position of the first object to be measured is the position of the bounding box 420. Let the endpoint coordinates of the diagonal of the bounding box 461 be the endpoints C'(x3, y4) and D'(x4, y3) respectively, and the endpoint coordinates of the diagonal of the bounding box 420 be the endpoints A(x1, y2) and B(x2, y1) respectively. The difference threshold is set to a horizontal coordinate difference threshold of 1 and a vertical coordinate difference threshold of 2. If x4 - x3 < 1, the deformation detection result of the second image relative to the first image is that it is not over-stretched. Similarly, in any of the cases where x2 - x1 < 1 or y4 - y3 < 2 or y2 - y1 < 2, it is considered that the deformation detection result of the second image relative to the first image is that it is not over-stretched.
[0192] In some embodiments, obtain the quantitative ratio between the quantity of the second object to be measured and the quantity of the second object. In response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, and the quantitative ratio being less than the quantitative ratio threshold, take not over-stretched as the deformation detection result of the second image relative to the first image.
[0193] Exemplarily, in the case where the position difference between the third position of the second object to be measured and the first position of the first object to be measured is less than the difference threshold, if the quantity of the first object to be measured is 4, the quantity of the first object is 2, and the quantitative ratio threshold is 3, and at this time the quantitative ratio is 2, and the quantitative ratio is less than the quantitative ratio threshold, then the deformation detection result of the second image relative to the first image is not over-stretched.
[0194] In some embodiments, the position difference between the third position of the second object to be measured and the first position of the first object to be measured can be used as the deformation detection result.
[0195] In some embodiments, determine the stretching level corresponding to the position difference and the quantitative ratio as the deformation detection result, where different stretching levels correspond to different intervals of the position difference and the quantitative ratio, and the higher the stretching level, the larger the value of the interval.
[0196] In the embodiments of the present application, by setting the difference threshold and the quantitative ratio threshold, the third position of the second object after mapping is further screened, filtering out the second objects that do not correspond to the first object after mapping, reducing the influence of the second objects that do not meet the preset conditions on the detection result, and improving the accuracy of detecting image over-deformation.
[0197] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0198] In scenarios of instant messaging or video playback, when content creators edit images, they often adjust the image size by stretching or compressing to meet the typesetting size requirements, or the captured pictures are distorted due to the stretching of the video frame during screenshotting. During the process of adjusting the image size, if the transformation is not carried out while maintaining the original ratio of the image, it may cause the objects contained in the image to be over-stretched or over-compressed, that is, the image is overly distorted. This makes it difficult to distinguish the objects contained in the image, hinders the transmission and expression of image information, and also reduces the aesthetic feeling of the image. Through the image processing method provided by the embodiments of the present application, the positions of the objects contained in the original image (i.e., the first image) and the distorted image (i.e., the second image) are obtained, the object contained in the distorted image (i.e., the first object) is mapped back to the corresponding position in the original image, and compared one by one with the position of the object contained in the original image (i.e., the first position of the first object), so as to determine whether the image is overly distorted.
[0199] Taking the scenario of determining the degree of face distortion as an example, refer to Figure 5 , Figure 5 which is the schematic diagram of determining the degree of face distortion provided by the embodiments of the present application.
[0200] In Figure 5 , it is mainly divided into two modules. The first module is to perform face detection and filtering. First, face area detection is carried out. Refer to Figure 6 , Figure 6 which is the schematic diagram of the face detection result provided by the embodiments of the present application. As shown in Figure 6 , based on the deep learning face area detector, face area detection is carried out. The detector returns the upper left corner and lower right corner coordinates of the circumscribed rectangle of the face area. Among them, the picture 610 is the image, and the two face areas detected by the face area detector are within the bounding box 620 and the bounding box 630.
[0201] Among them, the positions of the upper left corner and lower right corner of the rectangle are represented by the coordinates within the image. Refer to Figure 7 , Figure 7 which is the schematic diagram of the image coordinate system provided by the embodiments of the present application. As shown in Figure 7 , the coordinate system in the image is: taking the upper left corner of the image as the origin, the horizontal direction as the X axis, and the vertical direction as the Y axis. That is, the X coordinate represents the column where the pixel is located in the image, and the Y coordinate represents the row where the pixel is located.
[0202] For example, use P=(x, y) to represent the coordinates of a point. x and y respectively represent the horizontal and vertical positions of this point in the image, that is, the serial numbers of the column and row where it is located.
[0203] Among them, a face detection result (i.e., the first position of at least one first object in the first image) can be expressed as F = (x1, y1, x2, y2), where x1 and y1 respectively represent the horizontal and vertical positions of the upper left corner of the rectangle in the image, that is, the serial numbers of the column and row where it is located, and x2 and y2 represent the horizontal and vertical positions of the lower right corner of the rectangle in the image.
[0204] Express the output result of face detection in the image as S_f = {F1, F2, F3,...}; express the output result of face detection in the stretched image as: S_f' = {F1', F2', F3',...}.
[0205] Secondly, filter out smaller face regions (i.e., bounding boxes). In the image, some face regions in the background do not provide main information, and their deformation situations can be ignored. In addition, some incomplete face regions and face regions with a large deviation from the front angle in the image do not need to consider their deformation situations during image adjustment. Therefore, the face detection results in the image need to be screened.
[0206] Smaller face regions are usually in the background, and the cropping of these faces usually does not affect the beauty of the image and the conveyance of effective information. Therefore, such face regions should be filtered out. Among them, the width of the face region can be calculated by subtracting the horizontal coordinate of the upper left corner from the horizontal coordinate of the lower right corner in the face detection result. The calculation method is: width = x2 - x1, where width represents the width of the face region, x1 represents the horizontal coordinate of the upper left corner of the face region, and x2 represents the horizontal coordinate of the lower right corner of the face region.
[0207] The height of the face region can be calculated by subtracting the vertical coordinate of the upper left corner from the vertical coordinate of the lower right corner of the key points. The calculation method is: height = y2 - y1, where height represents the width of the face region, y1 represents the vertical coordinate of the upper left corner of the face region, and y2 represents the vertical coordinate of the lower right corner of the face region.
[0208] To filter out smaller face regions, a minimum height ratio coefficient MIN_HEIGHT_Ratio and a minimum width ratio coefficient MIN_WIDTH_Ratio (i.e., size ratio thresholds) are preset in advance. These two coefficients respectively represent the proportions of the height and width of the smallest face region that needs to be determined in the height and width of the image. For example, if the image height is 1000 pixels and MIN_HEIGHT_Ratio is set to 0.05, then the minimum height of the face region that needs to be determined is 50 pixels. Denote the image width as W and the image height as H. Then the filtering condition is: when width < W * MIN_WIDTH_Ratio or height < H * MIN_HEIGHT_Ratio, filter out this face detection result.
[0209] Finally, filter the detection results of the face range exceeding the image boundary. The face positions returned by some face detection methods may exceed the image range, which means that the face may be incomplete in the original image. Therefore, these text boxes are not detected. Among them, the filtering condition is: when x2 > W or y2 > H or x1 < 0 or y1 < 0, then filter out this face detection result.
[0210] After face detection and filtering of the detection results, remove the face detection results that do not meet the conditions in the two sets S_f and S_f', and obtain a series of face regions that meet the conditions. If S_f is not empty after filtering, then next judge whether the image stretching makes the face region unrecognizable
[0211] Continue to refer to Figure 5 , the second module is the determination of the degree of face deformation. First, perform coordinate mapping of the face region. After the image is stretched, the position of the face region should be mapped back to the original image coordinate system, and then the face deformation situation is determined. Denote the width of the original image as W, the height as H, the width of the stretched image as W', and the height as H'. Then the coordinates of the face detection results in the stretched figure are mapped back to the coordinates of the original image in the following way: the abscissa mapping is: x' / W' * W → x', and the ordinate mapping is: y' / H' * H → y', where x' represents the abscissa after mapping, and y' represents the ordinate after mapping.
[0212] Then, judge the face situation. After mapping, match the set of face detection results S_f (i.e., the first set) in the original image and the set of face detection results S_f' (i.e., the second set) on the stretched image one by one according to the position. The matching method is: compare the coordinate positions of the detection results in S_f with the detection results in the mapped S_f' one by one. If the coordinate position deviation is within the Margin range (i.e., the difference threshold), then it is considered that the two mark the same face region. After matching, judge the stretching situation according to the matching situation.
[0213] If any face detection result in S_f cannot be matched to the corresponding face detection result in S_f', it means that the image stretching has caused the image to be deformed severely, resulting in the face detection model being unable to recognize the face region in the image. At this time, it is determined that the image stretching has caused excessive face deformation.
[0214] If all face detection results in S_f can be matched in S_f', it means that the face region can still be accurately recognized after the image is stretched. Then it is determined that the image stretching has no impact on the face detection results.
[0215] The above method can be applied to the detection of improper face cropping in images. Refer to Figure 8 , Figure 8It is an exemplary schematic diagram of the detection result provided by an embodiment of the present application. Figure 8 It shows an example of excessive deformation of a human face detected due to image stretching.
[0216] By means of object recognition in the first image and the second image, the task of detecting the deformation of the second image relative to the first image in an embodiment of the present application is converted into the task of detecting the deformation of the first object in the first image relative to the second object that matches (i.e., there is a matching relationship) in the second image, which simplifies the calculation of image deformation detection and improves the detection efficiency; mapping the object in the deformed image to the original image to facilitate judging whether the object in the deformed image is complete after mapping, more accurately reflecting the degree of change of the determined matching object between different images through the position distance, and removing the incomplete object in the image or the object that exceeds the image edge after mapping through a preset threshold, only focusing on the relatively complete object area, avoiding misjudgment of image deformation, and improving the accuracy of detecting excessive image deformation.
[0217] Next, continue to describe the exemplary structure of the implementation of the image processing device 555 provided by the embodiment of the present application as a software module. In some embodiments, as Figure 2 shown, the software module stored in the image processing device 555 in the memory 550 may include:
[0218] An image acquisition module 5551, configured to acquire a first image and a second image, where the first image is an unadjusted image, and the second image is obtained by adjusting the first image.
[0219] An object recognition module 5552, configured to perform object recognition on the first image and the second image respectively to obtain the first position of at least one first object in the first image and the second position of at least one second object in the second image.
[0220] A position mapping module 5553, configured to map the second position of at least one second object to the first image to obtain the third position of at least one second object.
[0221] A position comparison module 5554, configured to compare the third position of at least one second object with the first position of at least one first object to determine the first object and the second object with a matching relationship.
[0222] A result determination module 5555, configured to determine the deformation detection result of the second image relative to the first image based on the first object and the second object with a matching relationship.
[0223] In some embodiments, the target recognition module 5552 is further configured to respectively extract a plurality of first candidate regions of the first image and a plurality of second candidate regions of the second image; respectively perform feature extraction on the first image and the second image to obtain a first feature map of the first image and a second feature map of the second image; determine first feature regions corresponding to the plurality of first candidate regions on the first feature map respectively; determine second feature regions corresponding to the plurality of second candidate regions on the second feature map respectively; respectively perform pooling operations on the first feature regions and the second feature regions to obtain a first feature region vector and a second feature region vector; respectively perform regression processing on the first feature region vector and the second feature region vector to obtain a first position of at least one object in the first image and a second position of at least one second object in the second image.
[0224] In some embodiments, for each first candidate region, the target recognition module 5552 is further configured to obtain a plurality of endpoint coordinates of the first candidate region; determine a first ratio between the size of the first image and the size of the first feature map; determine a plurality of second ratios between the endpoint coordinates of the plurality of first candidate regions and the first ratio respectively, and use the plurality of second ratios as the plurality of endpoint coordinates of the first feature region; determine the region corresponding to the plurality of endpoint coordinates of the first feature region on the first feature map as the first feature region corresponding to the first candidate region on the first feature map.
[0225] In some embodiments, the position mapping module 5553 is further configured to determine a ratio between the size of the first image and the size of the second image; determine the product of the endpoint coordinates of the bounding box of the second object at the second position and the ratio, and use the product as the coordinates of the bounding box of the second object at the third position; use the endpoint coordinates of the bounding box of the second object at the third position as the third position of at least one second object.
[0226] In some embodiments, the position comparison module 5554 is further configured to determine a first set, where the first set includes the first position of each first object; determine a second set, where the second set includes the third position of each second object; for each first position in the first set, perform the following operations: select any one first position from the first set as the position to be measured; determine the distance between each third position in the second set and the position to be measured; establish a matching relationship based on the second object corresponding to the third position with the smallest distance from the position to be measured and the first object corresponding to the position to be measured.
[0227] In some embodiments, the position comparison module 5554 is further configured to determine the size of the bounding box of the second object according to the endpoint coordinates of the bounding box of the second object; obtain the size of the second image; determine the size ratio between the size of the bounding box of the second object and the size of the second image; filter out the bounding boxes with a size ratio less than the size ratio threshold.
[0228] In some embodiments, the position comparison module 5554 is further configured to obtain a plurality of original image samples and a plurality of adjusted image samples corresponding to the original image samples; obtain a true size ratio threshold calibrated for the original image samples; call a machine learning model based on the original image samples and the corresponding adjusted image samples to obtain a predicted size ratio threshold; determine the loss between the true size ratio threshold and the predicted size ratio threshold, and update the parameters of the machine learning model according to the loss to obtain an updated machine learning model.
[0229] In some embodiments, the size ratio threshold is pre-set and is obtained through a ratio threshold setting interface. The position comparison module 5554 is further configured to obtain the pre-set size ratio threshold, where the size ratio threshold is obtained through a threshold setting operation, and the threshold setting operation is received through the threshold setting interface.
[0230] In some embodiments, the position comparison module 5554 is further configured to obtain the endpoint coordinates of each endpoint of the bounding box of the first object based on the first position; for each first object, in response to the endpoint coordinates of any endpoint of the bounding box of the first object exceeding the coordinate range of the first image, filter out the bounding box of the first object.
[0231] In some embodiments, the result determination module 5555 is further configured to use the first object and the second object with a matching relationship as the first object to be measured and the second object to be measured respectively; in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to the difference threshold, take over-stretching as the deformation detection result of the second image relative to the first image; in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, take non-over-stretching as the deformation detection result of the second image relative to the first image.
[0232] In some embodiments, the result determination module 5555 is further configured to obtain the quantity ratio between the quantity of the first object to be measured and the quantity of the first object; in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to the difference threshold, and the quantity ratio being less than the quantity ratio threshold, take over-stretching as the deformation detection result of the second image relative to the first image; in response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, and the quantity ratio being less than the quantity ratio threshold, take non-over-stretching as the deformation detection result of the second image relative to the first image.
[0233] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, causing the electronic device to execute the image processing method described above in the embodiment of the present application.
[0234] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the image processing method provided by the embodiment of the present application. For example, Figure 3A the image processing method shown.
[0235] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0236] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0237] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (such as files storing one or more modules, subroutines, or code portions).
[0238] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0239] In summary, through the embodiments of the present application, by performing object recognition in the first image and the second image, the task of detecting the deformation of the second image relative to the first image is converted into the task of detecting the deformation of the first object in the first image relative to the second object that matches (i.e., there is a matching relationship) in the second image, which simplifies the calculation of image deformation detection and improves the detection efficiency; the object in the deformed image is mapped to the original image to facilitate determining whether the object in the deformed image is complete after mapping. The degree of change of the determined matching objects between different images is more accurately reflected by the position distance, and the incomplete objects in the image or the objects that exceed the image edge after mapping are removed through a preset threshold, and only the relatively complete object regions are concerned, avoiding misjudgment of image deformation and improving the accuracy of detecting excessive image deformation.
[0240] As described above, the above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a first image and a second image, wherein the second image is obtained by adjusting the first image; Performing target recognition on the first image and the second image respectively to obtain a first position of at least one first object in the first image and a second position of at least one second object in the second image; Mapping the second position of the at least one second object into the first image to obtain a third position of the at least one second object; comparing the third position of the at least one second object with the first position of the at least one first object to determine the first object and the second object that have a matching relationship; Based on the first object and the second object having the matching relationship, a deformation detection result of the second image relative to the first image is determined.
2. The method according to claim 1, characterized in that The second position is the endpoint coordinates of the bounding box of the second object; Before comparing the third position of the at least one second object with the first position of the at least one first object, the method further includes: Determining a size of the bounding box of the second object according to the endpoint coordinates of the bounding box of the second object; Acquire the size of the second image; determining a size ratio between a size of a bounding box of the second object and a size of the second image; The bounding boxes whose size ratio is smaller than a size ratio threshold are filtered out.
3. The method according to claim 2, characterized in that Before filtering out the bounding boxes whose size ratio is less than a size ratio threshold, the method further includes: The size ratio threshold is obtained, wherein the size ratio threshold is set by a threshold setting operation, and the threshold setting operation is received through a threshold setting interface.
4. The method according to claim 2, characterized in that: The size ratio threshold is determined by a machine learning model; The training method of the machine learning model includes: Acquire a plurality of original image samples and a plurality of adjusted image samples corresponding to the original image samples; Obtaining a real size ratio threshold calibrated for the original image sample; Calling the machine learning model based on the original image sample and the corresponding adjusted image sample to obtain a predicted size ratio threshold; Determine the loss between the true size ratio threshold and the predicted size ratio threshold, and update the parameters of the machine learning model according to the loss to obtain the updated machine learning model.
5. The method according to any one of claims 1 to 4, characterized in that: The first position is the coordinates of the endpoints of the bounding box of the first object; Before comparing the third position of the at least one second object with the first position of the at least one first object, the method further includes: Acquire endpoint coordinates of each endpoint of the bounding box of the first object based on the first position; For each of the first objects, in response to the endpoint coordinates of any endpoint of the bounding box of the first object exceeding the coordinate range of the first image, the bounding box of the first object is filtered out.
6. The method according to any one of claims 1 to 4, characterized in that: Mapping the second position of the at least one second object into the first image to obtain a third position of the at least one second object includes: The following processing is performed for each of the second objects: determining a ratio of a size of the first image to a size of the second image; Determine the product of the endpoint coordinates of the enclosing box of the second object at the second position and the ratio, and use the product as the endpoint coordinates of the enclosing box of the second object at the third position; The endpoint coordinates of the enclosing box of the second object at the third position are used as the third position of the second object.
7. The method according to any one of claims 1 to 4, characterized in that: The determining, based on the first object and the second object having the matching relationship, a deformation detection result of the second image relative to the first image includes: The first object and the second object having the matching relationship are used as a first object to be tested and a second object to be tested respectively; In response to a position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to a difference threshold, taking overstretching as a deformation detection result of the second image relative to the first image; In response to a position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, non-excessive stretching is used as a deformation detection result of the second image relative to the first image.
8. The method according to claim 7, characterized in that When the first object and the second object having the matching relationship are used as the first object to be tested and the second object to be tested respectively, the method further includes: Obtaining a quantity ratio between the quantity of the first objects to be tested and the quantity of the first objects; In response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to a difference threshold, taking overstretching as a deformation detection result of the second image relative to the first image, comprising: In response to a position difference between the third position of the second object to be measured and the first position of the first object to be measured being greater than or equal to the difference threshold, and the quantity ratio being less than the quantity ratio threshold, taking overstretching as a deformation detection result of the second image relative to the first image; In response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, taking no over-stretching as a deformation detection result of the second image relative to the first image, comprising: In response to the position difference between the third position of the second object to be measured and the first position of the first object to be measured being less than the difference threshold, and the quantity ratio being less than the quantity ratio threshold, non-over-stretching is taken as the deformation detection result of the second image relative to the first image.
9. The method according to any one of claims 1 to 4, characterized in that: The comparing the third position of the at least one second object with the first position of the at least one first object to determine the first object and the second object having a matching relationship includes: determining a first set, wherein the first set includes the first position of each of the first objects; determining a second set, wherein the second set includes the third position of each of the second objects; For each of the first positions in the first set, perform the following operations: Selecting any one of the first positions from the first set as a position to be measured; Determine the distance between each of the third positions in the second set and the position to be measured; A matching relationship is established based on the second object corresponding to the third position having the smallest distance from the position to be measured and the first object corresponding to the position to be measured.
10. The method according to any one of claims 1 to 4, characterized in that: The performing target recognition on the first image and the second image respectively to obtain a first position of at least one object in the first image and a second position of at least one second object in the second image includes: extracting a plurality of first candidate regions of the first image and a plurality of second candidate regions of the second image respectively; Performing feature extraction on the first image and the second image respectively to obtain a first feature map of the first image and a second feature map of the second image; Determine first feature regions corresponding to the plurality of first candidate regions on the first feature map respectively; Determine second feature regions corresponding to the plurality of second candidate regions on the second feature map respectively; Performing pooling operations on the first feature region and the second feature region respectively to obtain a first feature region vector and a second feature region vector; Regression processing is performed on the first feature region vector and the second feature region vector respectively to obtain a first position of at least one object in the first image and a second position of at least one second object in the second image.
11. The method according to claim 10, characterized in that The determining first feature regions corresponding to the plurality of first candidate regions respectively on the first feature map includes: For each of the first candidate regions, perform the following operations: Obtaining coordinates of multiple endpoints of the first candidate area; determining a first ratio of a size of the first image to a size of the first feature map; Determine a plurality of second ratios of the endpoint coordinates of the plurality of first candidate regions to the first ratio, and use the plurality of second ratios as a plurality of endpoint coordinates of the first feature region; An area corresponding to the coordinates of multiple endpoints of the first feature area on the first feature map is determined as the first feature area corresponding to the first candidate area on the first feature map.
12. An image processing device, characterized in that: The device comprises: An image acquisition module, used to acquire a first image and a second image, wherein the second image is obtained by adjusting the first image; a target recognition module, configured to perform target recognition on the first image and the second image respectively, and obtain a first position of at least one first object in the first image and a second position of at least one second object in the second image; a position mapping module, configured to map the second position of the at least one second object into the first image to obtain a third position of the at least one second object; a position comparison module, configured to compare the third position of the at least one second object with the first position of the at least one first object to determine the first object and the second object having a matching relationship; The result determination module is used to determine a deformation detection result of the second image relative to the first image based on the first object and the second object having the matching relationship.
13. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions; A processor, configured to implement the image processing method according to any one of claims 1 to 11 when executing the computer executable instructions stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.