Image processing method and device, equipment, medium and program product
Through the edge detection model of unsupervised learning and shape constraint processing, the problem of manually labeling data in the prior art image edge detection is solved, and more efficient and accurate edge detection results are achieved.
Patent Information
- Application Number
- CN202410015170.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
The existing deep network models require a large amount of manual annotation of training data in image edge detection, which is time-consuming and labor-intensive and the recognition results are not accurate enough.
The edge detection model is trained using unsupervised learning, and the edge detection results are optimized in combination with shape constraint processing. By acquiring the edge image of the source image and multi-scale resolution images for edge detection, and using a lightweight neural network model for feature extraction and edge detection.
The workload of manually labeling training data is reduced, the accuracy and efficiency of image edge detection is improved, and the object areas in the image can be more accurately identified.
Smart Images

Figure CN120259351A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to an image processing method, an image processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In the field of image processing, it often involves edge (or contour) detection processing of images such as game images, traffic images, live images, etc., so as to identify target objects in the images (such as game controls, game characters, vehicles, etc.), so as to facilitate subsequent processing of the identified target objects.
[0003] Currently, generally, deep network models such as the Faster RCNN (Faster Regions With Convolutional Neural Networks) model, the SSD (Single Shot Multibox Detector) model, and the RefineDet (a refined detection network) model are used to identify and detect images, so as to obtain corresponding edge detection results. The above models usually need to be trained with a large amount of manually labeled training data, which is time-consuming and laborious; in addition, there may be certain errors in the model recognition results and they are not accurate enough. Summary of the Invention
[0004] Embodiments of the present application propose an image processing method, apparatus, device, medium, and program product, which can perform edge detection on an image through an edge detection model after unsupervised training, and further perform shape constraint processing on the edge detection result obtained by the model detection, which can improve the accuracy of edge detection.
[0005] On the one hand, embodiments of the present application provide an image processing method, and the method includes:
[0006] Obtain a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales, one image scale corresponds to one resolution image, and M is a positive integer;
[0007] Based on the M resolution images and the first edge image, call an edge detection model to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and determine an edge detection result of the source image based on the edge detection result of the first edge image; the edge detection model is obtained after model training by using an unsupervised learning method; the edge detection result includes at least one detected edge point;
[0008] Analyze the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; each connected region is composed of multiple edge points;
[0009] Perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0010] On the one hand, an embodiment of the present application provides an image processing apparatus, which includes:
[0011] An acquisition unit, configured to acquire a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales, where one image scale corresponds to one resolution image, and M is a positive integer;
[0012] A processing unit, configured to perform edge detection processing on the first edge image by invoking an edge detection model based on the M resolution images and the first edge image to obtain an edge detection result of the first edge image, and determine an edge detection result of the source image based on the edge detection result of the first edge image; the edge detection model is obtained after model training by using an unsupervised learning method; the edge detection result includes at least one detected edge point;
[0013] The processing unit is further configured to analyze the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; each connected region is composed of multiple edge points;
[0014] The processing unit is further configured to perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0015] In a possible implementation manner, the processing unit is further configured to perform the following operations:
[0016] Perform denoising processing on the source image to obtain a denoised source image;
[0017] Use an edge detection operator to calculate the gradient values of P pixel points in the denoised source image;
[0018] Based on the gradient values of the P pixel points, perform filtering processing on the P pixel points in a non-maximum suppression manner to obtain Q edge pixels, where P and Q are both positive integers and Q≤P;
[0019] Perform double-threshold processing on the Q edge pixels to divide the Q edge pixels into q1 strong edge pixels and q2 weak edge pixels, where q1 and q2 are both positive integers;
[0020] Perform edge suppression processing on the q2 weak edge pixels to obtain a first edge image of the source image.
[0021] In a possible implementation, when M = 3, the M resolution images include a first resolution image, a second resolution image, and a third resolution image; the image scale of the source image includes a source width W * a source height H; the processing unit is further configured to perform the following operations:
[0022] Perform multi-scale scaling processing on the source width and source height of the source image respectively according to a preset ratio to obtain a first resolution image, a second resolution image, and a third resolution image;
[0023] The image scale of the first resolution image is a first width X * a first height X / 2;
[0024] The image scale of the second resolution image is a second width X / 2 * a second height X / 4;
[0025] The image scale of the third resolution image is a third width X / 4 * a third height X / 8;
[0026] Wherein, the source width W of the source image is the same as or different from the scaled widths X, X / 2, X / 4; and, the source height H of the source image is the same as or different from the scaled heights X / 2, X / 4, X / 8.
[0027] In a possible implementation, the image input channels of the edge detection model include a color channel and an edge channel; the first edge image is used as the input image of the edge channel, and any one of the resolution images is used as the input image of the color channel; wherein, the edge detection model includes: a convolutional layer, an upsampling layer, and a pooling layer;
[0028] The processing unit, based on the M resolution images and the first edge image, calls the edge detection model to perform edge detection processing on the first edge image to obtain the edge detection result of the first edge image, which is used to perform the following operations:
[0029] Use the convolutional layer to perform feature extraction processing on the M resolution images and the first edge image to obtain the convolutional image features of the source image;
[0030] Use the upsampling layer to perform size amplification processing on the convolutional image features to obtain the processed convolutional image features;
[0031] Use the pooling layer to perform pooling processing on the feature dimensions of the processed convolutional image features to obtain the second edge image of the source image;
[0032] Based on the first edge image and the second edge image, obtain the edge detection result of the first edge image.
[0033] In a possible implementation, the convolutional layer includes: a first-scale convolutional layer, a second-scale convolutional layer, and a third-scale convolutional layer. The edge detection model further includes a concatenation layer. The processing unit uses the convolutional layer to perform feature extraction processing on M resolution images and the first edge image to obtain the convolutional image features of the source image for performing the following operations:
[0034] Use the first-scale convolutional layer to perform feature extraction processing on the first resolution image and the first edge image to obtain the first convolutional features of the first resolution image;
[0035] Use the second-scale convolutional layer to perform feature extraction processing on the second resolution image and the first edge image to obtain the second convolutional features of the second resolution image;
[0036] Use the third-scale convolutional layer to perform feature extraction processing on the third resolution image and the first edge image to obtain the third convolutional features of the third resolution image;
[0037] Call the concatenation layer to perform concatenation processing on the first convolutional features, the second convolutional features, and the third convolutional features to obtain the convolutional image features of the source image.
[0038] In a possible implementation, the first edge image includes I edge points, and the second edge image includes J edge points, where both I and J are positive integers. The processing unit obtains the edge detection result of the first edge image based on the first edge image and the second edge image for performing the following operations:
[0039] Obtain the first pixel value of the i-th edge point in the first edge image, where i is a positive integer and 1 ≤ i ≤ I;
[0040] Obtain the second pixel value of the j-th edge point in the second edge image, where j is a positive integer and 1 ≤ j ≤ J;
[0041] Perform an intersection operation on the first pixel value of the i-th edge point in the first edge image and the second pixel value of the j-th edge point in the second edge image to obtain an operation result;
[0042] Obtain the edge detection result of the first edge image based on the operation result;
[0043] Wherein, the i-th edge point in the first edge image and the j-th pixel point in the second edge image refer to the pixel points at the same position in the source image.
[0044] In a possible implementation, the edge detection result includes an edge detection image composed of multiple edge points. The processing unit analyzes the edge detection result of the source image to obtain N connected regions in the source image for performing the following operations:
[0045] Perform dilation processing on the edge detection image to obtain the dilated edge detection image;
[0046] Traverse each pixel point in the dilated edge detection image to obtain the traversal result;
[0047] Determine N connected regions in the source image according to the traversal result.
[0048] In a possible implementation, the edge detection image includes at least one edge point; any pixel point in the dilated edge detection image is represented as pixel point a; the processing unit traverses each pixel point in the dilated edge detection image to obtain the traversal result for performing the following operations:
[0049] In the dilated edge detection image, obtain K associated points associated with pixel point a, where K is a positive integer;
[0050] If any one of the K associated points does not belong to the edge points in the edge detection image, mark pixel point a as a background point;
[0051] Perform pixel point expansion with pixel point a as the center point;
[0052] After stopping the pixel point expansion, obtain one or more pixel points that have been visited, and obtain the traversal result according to each visited pixel point.
[0053] In a possible implementation, the processing unit performs shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions for performing the following operations:
[0054] Perform shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions;
[0055] Calculate the ratio between the area of each candidate region and the area of the source image;
[0056] If the ratio reaches the preset ratio threshold, use the candidate region as the object region of the source image.
[0057] In a possible implementation, the control shape constraint rule is used to indicate performing shape constraint processing on the first connected region among the N connected regions according to a rectangular region; the processing unit performs shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions for performing the following operations:
[0058] Determine the first center point and at least one first edge point of the first connected region, where the first edge point refers to the edge point included in the first connected region;
[0059] Calculate a first horizontal distance and a second vertical distance between a first edge point and a first center point;
[0060] Determine a rectangular area according to the first center point, the first horizontal distance, and the second vertical distance;
[0061] Calculate a first intersection over union (IoU) between the first connected area and the rectangular area;
[0062] If the first IoU is greater than or equal to a first preset threshold, determine the first connected area as a candidate area.
[0063] In a possible implementation, a control shape constraint rule is used to indicate performing shape constraint processing on a second connected area among N connected areas according to a circular area; the processing unit performs shape constraint processing on the N connected areas according to the control shape constraint rule to determine one or more candidate areas from the N connected areas for performing the following operations:
[0064] Determine a second center point and at least one second edge point of the second connected area, where the second edge point refers to an edge point included in the second connected area;
[0065] Calculate a reference distance between the second edge point and the second center point;
[0066] Determine a circular area according to the second center point and the reference distance;
[0067] Calculate a second IoU between the second connected area and the circular area;
[0068] If the second IoU is greater than or equal to a second preset threshold, determine the second connected area as a candidate area.
[0069] In a possible implementation, the processing unit is further configured to perform the following operations:
[0070] Determine an associated area associated with a reference connected area from the N connected areas; where the reference connected area includes the first connected area or the second connected area;
[0071] Combine the reference connected area and the associated area to obtain an updated area;
[0072] Perform shape constraint processing on the updated area according to the rectangular area or the circular area to obtain a target edge detection result of the source image.
[0073] In a possible implementation, the processing unit is further configured to perform the following operations:
[0074] Obtain a sample data set, where the sample data set includes multiple sample images and the edge annotation results of each sample image, and the edge annotation result of any sample image is obtained by annotating at least one object;
[0075] Construct an initial neural network model and use the sample data set to train the initial neural network model;
[0076] When the initial neural network model reaches the model convergence condition, stop training the initial neural network model and use the initial neural network model after stopping training as the edge detection model.
[0077] In a possible implementation, the processing unit uses the sample data set to train the initial neural network model for performing the following operations:
[0078] Call the initial neural network model to perform edge detection processing on any sample image in the sample data set to obtain the edge detection result of any sample image;
[0079] Perform integration operations on the edge annotation results of any sample image in the sample data set to obtain the integrated sample annotation result;
[0080] Calculate the model loss of any sample image according to the sample annotation result of any sample image and the edge detection result of any sample image;
[0081] Iteratively adjust the model parameters of the initial neural network model according to the model loss of any sample image.
[0082] In a possible implementation, the source image is a game image in a target game, and the object area includes: a control area, a character area, and an item area; after the processing unit performs shape constraint processing on N connected areas to obtain the target edge detection result of the source image, it is further used to perform the following operations:
[0083] If the object area is the control area, perform response time test processing on the game control indicated by the control area during the operation of the target game;
[0084] If the object area is the character area, perform movement control processing on the game character indicated by the character area during the operation of the target game;
[0085] If the object area is the item area, perform pickup control processing on the game item indicated by the item area during the operation of the target game.
[0086] On the one hand, an embodiment of the present application provides a computer device, which includes a processor, an input device, an output device, and a memory; a computer program is stored in the memory; when the computer program is executed by the processor, the above-mentioned image processing method is executed.
[0087] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned image processing method is executed.
[0088] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned image processing method is executed.
[0089] In an embodiment of the present application, a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales can be obtained, where one image scale corresponds to one resolution image, and M is a positive integer; based on the M resolution images and the first edge image, an edge detection model is called to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and an edge detection result of the source image is determined based on the edge detection result of the first edge image. Here, the edge detection model is obtained after model training using an unsupervised learning method, and the edge detection result includes at least one detected edge point; by analyzing the edge detection result of the source image, N connected regions in the source image can be obtained, where N is a positive integer; among them, each connected region is composed of multiple edge points; shape constraint processing is performed on the N connected regions to determine at least one object region in the source image from the N connected regions. Thus, on the one hand, the present application can perform edge detection processing on the first edge image of the source image using an edge detection model trained by an unsupervised learning method, without the need for manual annotation of training data, which can reduce the workload of users and thus improve the efficiency of image processing; on the other hand, the present application uses the first edge image of the source image as the input of the model, which can provide more edge information for the model and make the edge processing of the model more accurate; on the further hand, the present application further performs shape constraint processing on the edge detection result output by the model, so that the detection result of the model can be optimized, thereby improving the accuracy of the edge detection processing of the source image. Description of the Drawings
[0090] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0091] Figure 1 It is a schematic diagram of the principle of an image processing solution provided by an embodiment of the present application;
[0092] Figure 2 It is a schematic diagram of the architecture of an image processing system provided by an embodiment of the present application;
[0093] Figure 3 It is a schematic diagram of the process of an image processing method provided by an embodiment of the present application;
[0094] Figure 4 It is a schematic diagram of the structure of an edge detection model provided by an embodiment of the present application;
[0095] Figure 5 It is a schematic diagram of the scene of the object area of a source image provided by an embodiment of the present application;
[0096] Figure 6 It is a schematic diagram of the process of another image processing method provided by an embodiment of the present application;
[0097] Figure 7 It is a schematic diagram of a game scene of image processing provided by an embodiment of the present application;
[0098] Figure 8 It is a schematic diagram of the structure of an image processing device provided by an embodiment of the present application;
[0099] Figure 9 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0100] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0101] The present application provides an image processing solution, which is applicable to any object detection game scenes such as game control detection, game character detection, game item detection, etc.; the present application can train an edge detection model by using an unsupervised model training method, which can reduce the workload and improve the image processing efficiency; and can further optimize the edge detection result output by the model by using shape constraint processing, thereby improving the accuracy of edge detection. Please refer to Figure 1 , Figure 1 which is a schematic diagram of the principle of an image processing solution provided by an embodiment of the present application. As Figure 1As shown in the figure, the principle of the image processing solution provided by this application is roughly as follows:
[0102] ① Obtain an edge data set. Specifically, when implemented, an edge data set (i.e., a sample data set) can be obtained from an open-source database (such as the edge database BSDS500). This edge data set can be composed of 500 sample images and the corresponding edge annotation results for each sample image. Among them, the edge annotation results of each sample image can be annotated by multiple (such as five) objects. Since the generality of the edge contour is relatively strong, it can be directly used for image edge extraction in the game field in the future.
[0103] ② Construct a lightweight deep model. Specifically, a lightweight neural network model can be used as the initial model for training. Here, the lightweight neural network model refers to a network model with fewer model parameters or fewer network layers. In addition, the initial neural network model has the ability of edge detection, and this model can be a neural network model with any network structure. This application does not make specific limitations on the model structure.
[0104] ③ Train the deep model. Based on the edge data set, the lightweight neural network model is trained in an unsupervised learning manner, and the trained neural network model is used as the edge detection model.
[0105] ④ Extract the edges of the game image. Specifically, when implemented, the source image to be processed, the first edge image of the source image, and M resolution images corresponding to the source image at M image scales can be obtained. One image scale corresponds to one resolution image, and M is a positive integer. Based on the M resolution images and the first edge image, the edge detection model is called to perform edge detection processing on the source image, and the edge detection result of the source image is obtained. The edge detection result includes at least one edge point detected from the source image.
[0106] ⑤ Screen regions based on shape priors. Specifically, when implemented, by analyzing the edge detection result of the source image output by the model, N connected regions in the source image can be obtained, where N is a positive integer. Each connected region is composed of multiple edge points. Shape constraint processing is performed on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0107] As can be seen above, on the one hand, the present application can obtain sample data for model training from an existing edge database and train an edge detection model using unsupervised learning. Thus, the edge detection model is used to perform edge detection processing on the source image without the need for manual annotation of training data, which can reduce the workload of users and improve the efficiency of image processing. On the other hand, the present application uses the first edge image of the source image as the input of the model, which can provide more edge information for the model and make the edge processing of the model more accurate. On the further hand, the present application further performs shape constraint processing on the edge detection result output by the model, so that the detection result of the model can be optimized, thereby improving the accuracy of the edge detection processing of the source image.
[0108] The following will introduce the key technical terms involved in the present application in detail.
[0109] I. Source image, first edge image, and resolution image.
[0110] The source image refers to the image on which edge detection processing needs to be performed. The source image can be an image in any scenario. For example, the source image can include: game images in game screens, video images in videos, traffic images on traffic roads, and landscape images, etc. The present application does not specifically limit the type of the source image. In addition, the source image can be obtained from a database or can be collected in real time.
[0111] The first edge image refers to the image obtained by performing edge extraction processing on the source image using an edge extraction algorithm. The edge extraction algorithm here can include but is not limited to: Canny algorithm, Roberts algorithm, Laplacian algorithm, Sobel algorithm, etc. It should be noted that the present application uses the first edge image as one of the input images of the edge detection model, which is beneficial to providing more image edge information for the model, thereby making the edge detection processing of the model on the source image more accurate.
[0112] The resolution image refers to the image obtained by performing scaling processing on the image scale of the source image. The image scale here can include, for example: width and height. And after performing one-time scale scaling processing on the source image, a resolution image corresponding to the source image can be obtained, that is, one image scale corresponds to one resolution image. For example, if the image scale after scaling the source image is 512*256, the first resolution image can be obtained; another example, if the image scale after scaling the source image is 256*128, the second resolution image can be obtained; still another example, if the image scale after scaling the source image is 128*64, the third resolution image can be obtained. That is to say, different image scales correspond to images with different resolutions.
[0113] II. Connected region and object region.
[0114] A connected region, as the name implies, refers to a region that is connected. A so-called connected region is a closed region without breakpoints. For example, polygonal regions such as circles, rectangles, and triangles can all be considered connected regions. In this application, a connected region is composed of multiple edge points detected by an edge detection model. An edge point refers to a pixel point in the source image that is detected as an object edge by the edge detection model.
[0115] An object region refers to a region determined from a connected region that contains a specified object. The specified object here can be set according to the business requirements of the business scenario. For example, the types of specified objects can include: items, people, etc. Among them, the specified objects corresponding to different business scenarios can be the same or different. For example, the specified objects in a game scenario can include: game controls, game characters, game items (such as tools, ammunition), etc.; and the specified objects in a traffic scenario can include: vehicles, passers-by, etc.
[0116] III. Shape Constraint Processing.
[0117] Shape constraint processing refers to a processing operation performed on a connected region according to a preset shape. For example, the preset shape includes: any polygon such as a circle, a rectangle, a triangle, etc. For different types of objects, different preset shapes can be used to perform shape constraint processing on the connected region, so as to screen out the corresponding type of object region from the connected region. For example, if the object to be detected is a game control, a circular shape can be used to perform shape constraint processing on the connected region; and if the object to be detected is a vehicle, a rectangular shape can be used to perform shape constraint processing on the connected region, and so on. After performing shape constraint processing on the connected region, the object region that matches the preset shape can be further screened out from the connected region, so that the object region to be detected can be more accurately screened out from the connected region.
[0118] IV. Artificial Intelligence.
[0119] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence; artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields and includes both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0120] The image processing solution proposed in this application may involve machine learning technology and computer vision technology in the field of artificial intelligence. Specifically, on the one hand, this application can use machine learning technology to train an initial neural network model to obtain an edge detection model, so as to be able to call the edge detection model to perform edge detection processing on the source image and obtain the edge detection result of the source image; on the other hand, during the process of calling the edge detection model to perform edge detection processing on the source image, computer vision technology can be used to extract and analyze the features of the source image, so as to be able to more accurately detect multiple edge points from the source image.
[0121] V. Cloud technology.
[0122] In the image processing solution proposed in this application, there are a large number of data calculation services and data storage services involved, so a large amount of computer operation costs are required. Then, cloud technology can be used to provide data calculation services and data storage services for this solution, so as to better perform image processing. Specifically, an edge detection model can be called based on the data calculation service to perform edge detection processing on the source image, so as to obtain the edge detection result of the source image, or an edge detection model can be obtained after model training using the unsupervised learning method based on the data calculation service; in addition, the edge detection result of the source image can be stored based on the data storage service, so as to facilitate subsequent data analysis of the edge detection result. Among them, cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model, which can form a resource pool, be used as needed, and is flexible and convenient. Among them, cloud technology can include cloud storage technology. The so-called cloud storage is a new concept extended and developed on the basis of the cloud computing concept. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world.
[0123] VI. Blockchain.
[0124] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The following explains related concepts such as blockchain systems, blockchain nodes, and block structures.
[0125] In this application, there are many types of voice wake-up data involved in the image processing process. Optionally, this application can send the voice wake-up data to the blockchain for storage. Based on the characteristics of the blockchain such as immutability and traceability, data tampering or leakage can be avoided, thereby improving the data security and reliability in the image processing process.
[0126] It should be specifically noted that there are many data involved in the image processing process of this application, such as: source image, first edge image, M resolution images, and edge detection results of the source image, etc. When the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the processes of collecting, using, and processing relevant data need to comply with relevant laws, regulations, and standards of the country and region, conform to the principles of legality, legitimacy, and necessity, and do not involve obtaining data types prohibited or restricted by laws and regulations. In some optional embodiments, the relevant data involved in the embodiments of this application is obtained after individual authorization from the object. Additionally, when obtaining individual authorization from the object, the purpose of the relevant data involved needs to be stated to the object.
[0127] The following specifically introduces the architecture diagram of the image processing system provided by this application.
[0128] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the architecture of an image processing system provided by an embodiment of this application. As Figure 2 shown, the architecture diagram of this image processing system may at least include: a cluster of terminal devices and a server 204. Among them, the cluster of terminal devices may include at least one terminal device, such as: terminal device 201, terminal device 202, terminal device 203, etc. The embodiments of this application do not specifically limit the number of terminal devices in the cluster of terminal devices, and the number of devices can be flexibly changed according to different requirements of business scenarios. For example, the number of game devices in a game scenario is 2, and the number of vehicle devices in a traffic scenario is 3, etc. Specifically, any terminal device can be directly or indirectly connected to the server 204 through wired or wireless communication methods.
[0129] Any computer device (terminal device or server) in the image processing system provided by this application can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a mobile internet device (MID), a vehicle, an in-vehicle device, a roadside device, a smart robot, an aircraft, a wearable device, such as smart devices like smart watches, smart bracelets, pedometers, etc., and a virtual reality device, etc.
[0130] Any computer device (terminal device or server) in the image processing system provided by this application can also be a server. Specifically, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0131] It can be understood that the types of the various computer devices in the image processing system of this application can be the same or different. For example, the terminal device 201 can be a mobile phone, the terminal device 202 can be a game device, and the terminal device 203 can be a vehicle device; again, for example, the terminal device 201, the terminal device 202, and the terminal device 203 can all be mobile phones, and the server 204 can be a server. This application does not limit the quantity and type of each computer device in the image processing system. Taking the terminal device 201 and the server 204 as examples, the specific process of image processing involved in the embodiments of this application will be briefly described below.
[0132] ① In the game scene of the target game running on the terminal device 201, the terminal device 201 can obtain the game image in the game screen, use the game image as the source image, obtain the first edge image of the source image, and obtain M resolution images corresponding to the source image at M image scales respectively, where one image scale corresponds to one resolution image.
[0133] ② The terminal device 201 sends the obtained source image, the first edge image of the source image, and the M resolution images to the server 204.
[0134] ③ The server 204 can, based on the M resolution images and the first edge image, call the edge detection model to perform edge detection processing on the first edge image to obtain the edge detection result of the first edge image, and determine the edge detection result of the source image based on the edge detection result of the first edge image; among them, the edge detection model is obtained after model training using the unsupervised learning method; the edge detection result includes at least one detected edge point.
[0135] ④ The server 204 returns the edge detection result of the source image to the terminal device 201.
[0136] ⑤ The terminal device 201 analyzes the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; among them, each connected region is composed of multiple edge points.
[0137] ⑥ The terminal device 201 performs shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions, such as regions of objects like game controls, game characters, game items, etc.
[0138] It should be noted that the above process is only an example and does not specifically limit the steps executed by the terminal device 201 and the server 204. Optionally, analyzing the edge detection result of the source image to obtain the N connected regions in the source image; and performing shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions, the above steps can also be executed by the server 204; in addition, the above process can also be executed independently by the terminal device 201 or the server 204.
[0139] In a possible implementation manner, the image processing system provided in this application can be deployed in a blockchain system, that is, the terminal device 201, the terminal device 202, the terminal device 203, and the server 204 can all be used as node devices in the blockchain system, and the relevant data involved in the above image processing process (such as the source image, the first edge image, the M resolution images, and the edge detection result of the source image, etc.) are all stored on the blockchain, so that the specific processing process of the source image in this application can be executed on the blockchain, which can not only ensure the fairness and impartiality of the image processing process, but also make the image processing process traceable, improving the security and reliability of the image processing process.
[0140] The image processing system provided by this application enables a computer device (any terminal device or server) to obtain a source image to be processed, a first edge image of the source image, and M resolution images corresponding to the source image at M image scales respectively, where one image scale corresponds to one resolution image and M is a positive integer. Based on the M resolution images and the first edge image, an edge detection model is called to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and an edge detection result of the source image is determined based on the edge detection result of the first edge image. Here, the edge detection model is obtained after model training using an unsupervised learning method, and the edge detection result includes at least one detected edge point. By analyzing the edge detection result of the source image, N connected regions in the source image can be obtained, where N is a positive integer. Each connected region is composed of multiple edge points. Shape constraint processing is performed on the N connected regions to determine at least one object region in the source image from the N connected regions. Thus, on the one hand, this application can perform edge detection processing on the first edge image using an edge detection model trained through an unsupervised learning method without manual annotation of training data, which can reduce the workload of users and improve the efficiency of image processing. On the other hand, this application uses the first edge image of the source image as the input of the model, which can provide more edge information for the model and make the edge processing of the model more accurate. On the further hand, this application further performs shape constraint processing on the edge detection result output by the model, so it can optimize the model detection result and improve the accuracy of the edge detection processing of the source image.
[0141] It can be understood that the image processing system described in the embodiments of this application is to more clearly illustrate the technical solutions of the embodiments of this application and does not constitute a limitation on the technical solutions provided by the embodiments of this application. As is known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0142] The following describes specific embodiments related to the image processing solution with reference to the accompanying drawings.
[0143] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an image processing method provided by an embodiment of this application. This method can be executed by a computer device (such as a terminal device or a server) in the image processing system shown in Figure 2 As shown in Figure 3 the image processing method mainly includes but is not limited to the following steps S301 - S304:
[0144] S301: Obtain the source image to be processed, the first edge image of the source image, and M resolution images corresponding to the source image at M image scales respectively, where one image scale corresponds to one resolution image, and M is a positive integer.
[0145] (1) The process of obtaining the source image.
[0146] In a possible implementation, the shooting device can be called to collect the environmental image in the service scenario in real time, and the environmental image collected in real time is used as the source image. For example, if the service scenario is a game scenario and the computer device is a mobile phone, the game image in the game screen can be intercepted in real time and used as the source image; another example is that if the service scenario is a traffic scenario and the computer device is a Road Side Unit (RSU), the traffic image in the traffic road can be collected by calling the RSU in real time and used as the source image. In this implementation, the source image is obtained in real time, and the real-time obtained source image is beneficial to timely conduct service analysis on the corresponding service scenario and has strong timeliness.
[0147] In another possible implementation, the source image can be obtained from the image database. The image database includes at least one collected image, and the required image can be selected from the image database according to the service requirements. For example, the game image can be selected from the image database as the source image, or the landscape image can be selected from the image database as the source image, etc. In this implementation, the source image can be obtained from the existing database without real-time collection, which can improve the efficiency of image acquisition.
[0148] (2) The process of obtaining the first edge image.
[0149] In a possible implementation, the edge extraction algorithm can be used to perform edge extraction processing on the source image to generate the first edge image. The edge extraction algorithm here can include but is not limited to: Canny extraction algorithm, Roberts algorithm, Laplacian algorithm, Sobel algorithm, etc. This application does not limit the edge extraction algorithm. Since the Canny algorithm can generate a refined, complete and continuous edge map, the following takes the Canny algorithm as an example to illustrate the edge extraction process of the source image:
[0150] ① Denoise the source image to obtain the denoised source image. Specifically, the Gaussian filter can be used to denoise the source image, which can smooth the image and filter out the noise.
[0151] ②Adopt an edge detection operator to calculate the gradient values of P pixel points in the denoised source image. Specifically, the edges in the source image can point in various directions. Therefore, the Canny algorithm uses four operators to detect the horizontal, vertical, and diagonal edges in the source image. The edge detection operators (such as Roberts, Prewitt, Sobel, etc.) return the first derivative values in the horizontal Gx direction and the vertical Gy direction, from which the gradient value G and the direction theta of the pixel point can be determined.
[0152] ③Based on the gradient values of the P pixel points, filter the P pixel points by non-maximum suppression to obtain Q edge pixel points, where both P and Q are positive integers and Q ≤ P. Specifically, non-maximum suppression is an edge thinning technique, and its role is to "thin" the edges. After calculating the gradient values of the source image, the edges extracted only based on the gradient values are still very blurred. Therefore, non-maximum suppression can be used to suppress all gradient values other than the local maximum to 0. The algorithm for non-maximum suppression of each pixel point in the gradient image includes: First, compare the gradient intensity of the current pixel point with the two pixel points along the positive and negative gradient directions; if the gradient intensity of the current pixel point is the largest compared to the other two pixel points, then this pixel point is retained as an edge point, otherwise this pixel point will be suppressed.
[0153] ④Perform double-threshold processing on the Q edge pixel points to divide the Q edge pixel points into q1 strong edge pixel points and q2 weak edge pixel points, where both q1 and q2 are positive integers. Specifically, the double threshold includes a high threshold and a low threshold, and these two thresholds play a key role in the double-threshold processing step of the Canny algorithm. Specifically, they are used to divide pixel points into three categories: strong edge pixel points, weak edge pixel points, and non-edge pixel points. I. If the gradient value of a certain pixel point is greater than the high threshold, it is regarded as a strong edge pixel point. Strong edge pixel points are basically part of the real edge, so they will be retained in the final result; II. If the gradient value of a certain pixel point is less than the low threshold, it is regarded as a non-edge pixel point, and non-edge pixel points will be discarded in the final result; III. If the gradient value of a certain pixel point is between the high threshold and the low threshold, it is regarded as a weak edge pixel point. Weak edge pixel points may be part of the real edge or caused by noise. Only when a weak edge pixel point is adjacent to a strong edge pixel point will it be retained in the final result. In the embodiments of the present application, the low threshold can be set to 30 and the high threshold can be set to 120. Practice shows that in the case of these two thresholds, the results of Canny detection can include most of the contours (i.e., edges). Through the Canny edge image, the results output by the model can be better constrained. The task of the edge detection model is to remove some details within the object in the Canny edge image and strengthen the contour of the object.
[0154] ⑤Perform edge suppression processing on the q2 weak edge pixel points to obtain the first edge image of the source image. Finally, based on the above q1 strong edge pixel points and the weak edge pixel points obtained through edge suppression processing, the first edge image of the source image is obtained.
[0155] (3) The process of obtaining M resolution images.
[0156] In a possible implementation, perform multi-scale scaling processing on the source image to obtain M resolution images corresponding to the source image at M image scales respectively. Specifically, assuming M = 3, the M resolution images may include: the first resolution image, the second resolution image, and the third resolution image, and the image scale of the source image includes the source width W and the source height H; then, the source width W and the source height H of the source image can be respectively subjected to multi-scale scaling processing according to a preset ratio (for example, 2:1) to obtain the first resolution image, the second resolution image, and the third resolution image. It should be understood that the number M of resolution images can be set according to service requirements, and this application does not make specific limitations on this. Among them, the source width W of the source image may or may not be the same as the scaled widths X, X / 2, X / 4, and the source height H of the source image may or may not be the same as the scaled heights X / 2, X / 4, X / 8. For example, W = 1024, H = 720, X = 512, that is, the scale of the source image can be 1024*720, then the image scale of the first resolution image after scale scaling can be X*X / 2 (for example, 512*216), the image scale of the second resolution image can be X / 2*X / 4 (for example, 216*128), and the image scale of the third resolution image is X / 4*X / 8 (for example, 128*64); another example is W = X = 512, H = 216, that is, the scale of the source image can be 512*216, the image scale of the first resolution image after scale scaling can be X*X / 2 (for example, 512*216); the image scale of the second resolution image can be X / 2*X / 4 (for example, 216*128); the image scale of the third resolution image can be X / 4*X / 8 (for example, 128*64). Generally speaking, this application does not make specific limitations on the values of W, H, and X.
[0157] S302: Based on the M resolution images and the first edge image, call the edge detection model to perform edge detection processing on the first edge image to obtain the edge detection result of the first edge image, and determine the edge detection result of the source image based on the edge detection result of the first edge image; the edge detection model is obtained after model training using an unsupervised learning method; the edge detection result includes at least one detected edge point.
[0158] In a possible implementation, after obtaining the source image, the first edge image, and M resolution images, the M resolution images and the first edge image can be directly used as the input images of the edge detection model; optionally, image enhancement processing can also be performed on the source image, the first edge image, and the M resolution images. The image enhancement processing here can include any one or more of image smoothing, image sharpening, scale scaling, etc. Then, the enhanced images are used as the input images of the edge detection model for model recognition. By using this method, the input images can be made clearer, which is beneficial for the model to more accurately recognize the input images and can improve the accuracy of the model for edge detection processing of the source image.
[0159] Specifically, the image input channel of the edge detection model is a four-channel, and the four channels specifically include: a color channel and an edge channel. Among them, the color channel refers to the RGB three channels, specifically including: the Red channel, the Green channel, and the Blue channel. The first edge image is used as the input image of the edge channel, and any one of the resolution images is used as the input image of the color channel (i.e., the RGB three channels). The edge detection model involved in this application mainly includes: a convolutional layer, an upsampling layer, and a pooling layer. Among them, ① the main function of the convolutional layer is: to extract the local features of the input image, realize feature mapping through convolutional operations. Each neuron in the convolutional layer is connected to a local area of the input image and learns features on this local area. Through the stacking of multiple convolutional layers, the network can learn higher-level and more abstract feature representations; ② the upsampling layer is mainly used to enlarge the size of the input image or feature map and restore the resolution of the image; ③ the pooling layer is used to reduce the dimension of the feature map, reduce the amount of calculation and the number of parameters, thereby reducing the risk of model overfitting.
[0160] In a possible implementation, the computer device, based on the M resolution images and the first edge image, calls the edge detection model to perform edge detection processing on the first edge image to obtain the edge detection result of the first edge image, including the following process: using the convolutional layer to perform feature extraction processing on the M resolution images and the first edge image to obtain the convolutional image features of the source image; using the upsampling layer to perform size enlargement processing on the convolutional image features to obtain the processed convolutional image features; using the pooling layer to perform pooling processing on the feature dimension of the processed convolutional image features to obtain the second edge image of the source image; based on the first edge image and the second edge image, obtaining the edge detection result of the first edge image.
[0161] Further, since the first edge image is an edge image obtained by performing edge extraction processing on the source image, then the edge detection result obtained after performing edge detection processing on the first edge image by invoking the edge detection model can be considered as an edge detection result of the source image. That is to say, the essence of performing edge detection by the edge detection model is to perform edge detection on the source image, so as to obtain the edge detection result of the source image. Therefore, in this application, determining the edge detection result of the source image based on the edge detection result of the first edge image can specifically be taking the edge detection result of the first edge image as the edge detection result of the source image. Since the first edge image is an edge image obtained by performing preliminary edge extraction on the source image using the canny extraction algorithm, the essence of the edge detection model performing edge detection processing on the source image can be considered as performing edge detection processing on the first edge image of the source image, so that the edge detection result of the first edge image can be taken as the edge detection result of the source image. In this implementation manner, since the first edge image can provide more edge information of the source image, the method of invoking the edge extraction model to perform edge detection processing on the first edge image can improve the efficiency of edge detection processing and the accuracy of the edge detection result compared with directly performing edge detection processing on the source image.
[0162] The processing process of the edge detection model will be introduced in detail below.
[0163] (1) Processing process of the convolutional layer.
[0164] Optionally, the above convolutional layer may include: a first-scale convolutional layer, a second-scale convolutional layer, and a third-scale convolutional layer. The edge detection model further includes a concatenation layer. Then, the computer device uses the convolutional layer to perform feature extraction processing on the M resolution images and the first edge image to obtain the convolutional image features of the source image, which specifically includes the following process: using the first-scale convolutional layer to perform feature extraction processing on the first resolution image and the first edge image to obtain the first convolutional features of the first resolution image; using the second-scale convolutional layer to perform feature extraction processing on the second resolution image and the first edge image to obtain the second convolutional features of the second resolution image; using the third-scale convolutional layer to perform feature extraction processing on the third resolution image and the first edge image to obtain the third convolutional features of the third resolution image; invoking the concatenation layer to perform concatenation processing on the first convolutional features, the second convolutional features, and the third convolutional features to obtain the convolutional image features of the source image. Here, the concatenation processing may specifically include adding the first convolutional features, the second convolutional features, and the third convolutional features, and taking the operation result as the convolutional image features of the source image.
[0165] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an edge detection model provided by an embodiment of the present application. AsFigure 4 As shown in the figure, the edge detection model includes: a first-scale convolutional layer, a second-scale convolutional layer, and a third-scale convolutional layer. Among them, the first-scale convolutional layer is used to perform feature extraction processing on the first-resolution image and the first edge image (such as the Canny edge image); the second-scale convolutional layer is used to perform feature extraction processing on the second-resolution image and the first edge image (such as the Canny edge image); the third-scale convolutional layer is used to perform feature extraction processing on the third-resolution image and the first edge image (such as the Canny edge image). As can be seen from the above, the input image of the edge detection model is a four-channel image with three resolutions. The first three channels are RGB channels, and the fourth channel is the detection result of the Canny edge (i.e., the Canny edge image). Here, using the Canny edge image as the input image of the model can provide more edge information for the model, which is beneficial for the model to more accurately extract the image edge. For example Figure 4 As shown in the figure, the edge detection model involved in this application includes input images with three resolutions, which are respectively used to detect images (such as game controls) with three scales: large, medium, and small. For the 512*256 scale (i.e., the first-scale convolutional layer), the first convolutional features of the input image are extracted through 4 convolutional layers. For the 256*128 scale (i.e., the second-scale convolutional layer), the second convolutional features of the input image are extracted through 3 convolutional layers. For the 128*64 scale (i.e., the third-scale convolutional layer), the third convolutional features of the input image are extracted through 2 convolutional layers; subsequently, the convolutional features extracted from the three scales are cascaded through a cascade layer, and the number of output channels is 720 (cascading the output features of 3 convolutional layers, and the number of channels of the output features of each convolutional layer is 240, so the number of channels after cascading is 720).
[0166] (2) The processing process of the upsampling layer.
[0167] For example Figure 4 As shown in the figure, after the convolutional image features with 720 output channels are output by the cascade layer, 4 upsampling layers can be used to process the convolutional image features, so as to output the second edge image. Specifically, the upsampling layer here is implemented through transposed convolution. A certain number of zeros (i.e., zero-padding processing) are inserted around each element of the input feature map (i.e., the convolutional image features) to expand the size of the convolutional image features. The number of zero-padding depends on the required stride and the size of the convolutional kernel. Perform a transposed convolution operation on the expanded convolutional image features. Here, the transposed convolution operation needs to use the same convolutional kernel as the forward convolution, and the convolutional kernel here can be pre-learned or learned through training. Add the result of the transposed convolution operation to the output feature map (i.e., the second edge image). In this way, the size of the output feature map will be larger than the size of the input feature map, that is, the second edge image can be obtained.
[0168] (3) Processing process of the pooling layer.
[0169] Optionally, the output feature map of the upsampling layer (i.e., the second edge image) can be subjected to pooling processing, such as: any one of average pooling, weighted pooling, and Region Of Interest pooling (ROI pooling); the purpose of pooling processing is to reduce the dimension of the feature map, reduce the amount of calculation and the number of parameters, thereby reducing the risk of model overfitting.
[0170] (4) Process of finding the intersection of edge maps.
[0171] Specifically, the first edge image (i.e., the Canny edge image) and the second edge image (such as Figure 4 the result output by the upsampling layer or the pooling layer above) are subjected to an intersection operation, so as to obtain the edge detection result (i.e., the edge detection image) output by the edge detection model. Among them, the intersection operation here can include: any one or more of bitwise AND operation, average operation, and weighted operation.
[0172] In a possible implementation, the first edge image includes I edge points, and the second edge image includes J edge points, where both I and J are positive integers; the computer device obtains the edge detection result of the source image based on the first edge image and the second edge image, which specifically includes the following process: obtaining the first pixel value of the i-th edge point in the first edge image, where i is a positive integer and 1 ≤ i ≤ I; obtaining the second pixel value of the j-th edge point in the second edge image, where j is a positive integer and 1 ≤ j ≤ J; performing an intersection operation on the first pixel value of the i-th edge point in the first edge image and the second pixel value of the j-th edge point in the second edge image to obtain an operation result; obtaining the edge detection result of the source image based on the operation result; where the i-th edge point in the first edge image and the j-th pixel point in the second edge image refer to the pixel points at the same position in the source image. For example, taking the intersection operation method as a bitwise AND operation, the so-called bitwise AND operation means performing a bitwise AND operation on the corresponding pixel values in two binary images (i.e., the first edge image and the second edge image). Among them, the pixel points with a pixel value of 1 after the bitwise AND operation indicate that the pixel values of the pixel points at this position in the two binary images are both 1; otherwise, the pixel points with a pixel value of 0 indicate that the pixel values of the pixel points at this position in the two binary images are not all 1.
[0173] In summary, from (1) to (4) above, since the edge points detected by Canny can be used as the object edges (or contours) finally detected, the Canny edge image can better constrain the results output by the model. The task of the edge detection model is to remove some details inside the object in the Canny edge image and strengthen the contour of the object. Then, after the model detects the second edge image, performing an intersection operation on the first edge image and the second edge image can further improve the accuracy of the edge detection image output by the model.
[0174] S303: Analyze the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; each of the connected regions is composed of multiple of the edge points.
[0175] Specifically, since the edge detection result includes at least one edge point detected from the source image, the essence of analyzing the edge detection result of the source image is to analyze and process each of the detected edge points, so as to obtain N connected regions in the source image, where one connected region is composed of multiple detected edge points.
[0176] In a possible implementation, the computer device analyzes the edge detection result of the source image to obtain N connected regions in the source image, which specifically includes the following steps: First, perform dilation processing on the edge detection image to obtain the dilated edge detection image; then traverse each pixel point in the dilated edge detection image to obtain the traversal result; finally, determine the N connected regions in the source image according to the traversal result. Specifically, when implemented, the edge detection image includes at least one edge point; any pixel point in the dilated edge detection image is represented as pixel point a; the computer device traverses each pixel point in the dilated edge detection image to obtain the traversal result, which specifically includes the following process: In the dilated edge detection image, obtain K associated points associated with pixel point a, where K is a positive integer; if any of the K associated points does not belong to the edge points in the edge detection image, mark pixel point a as a background point; perform pixel point expansion with pixel point a as the center point; after stopping the pixel point expansion, obtain one or more pixel points that have been visited, and obtain the traversal result according to each of the visited pixel points.
[0177] The following details the specific process of obtaining the connected regions of the source image.
[0178] (1) Image dilation processing.
[0179] In specific implementation, a dilation operation can be performed on the edge detection image using a 2*2 dilation kernel. The so-called dilation operation is a process of merging all background points in contact with an object into the object, so as to expand the boundary outwards, which can be used to fill holes or breakpoints in the object. Generally speaking, the dilation operation can use a 3x3 kernel (dilation kernel). Use this dilation kernel to scan each pixel point in the edge detection image, and then perform an "AND" operation between the kernel and the edge detection image it covers; if all are 0, then the pixel value of this pixel point in the edge detection image is 0, otherwise, the pixel value of this pixel point is 1; the edge detection image after dilation will visually expand by one circle.
[0180] (2) Traverse the dilated edge detection image to determine the connected regions.
[0181] I. Traverse each pixel point in the dilated edge detection image. For any pixel point a, it can be determined whether K pixel points (such as 8 pixel points) around this pixel point a are all edge points. If not, mark this pixel point a as a background point and skip it from the connected region.
[0182] II. If a pixel point a is marked as a background point, then continue with pixel expansion with this pixel a as the center point. The specific method is to check the 8 pixel points adjacent to this pixel point a. If one or more of these pixel points (such as pixel point b) have not been visited, then start from this pixel point b and continue with pixel expansion. This process will continue until there are no new expansion points. At this time, all pixel points expanded from the central pixel belong to the same connected region.
[0183] III. After completing one pixel expansion, it is necessary to skip all pixel points within the current connected region to avoid repeated access. Then continue to traverse other pixel points in the image and repeat the above processes I and II until all pixel points have been visited, and then N connected regions in the source image can be determined.
[0184] As can be seen from the above steps (1)-(2), the present application can further process the edge detection image output by the model, that is, traverse each pixel point in the image and perform pixel expansion processing, so as to effectively and accurately determine valid connected regions from the source image during the iteration process of pixel expansion.
[0185] S304: Perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0186] In a possible implementation, the computer device performs shape constraint processing on N connected regions to determine at least one object region in the source image from the N connected regions. The specific process is as follows: The control shape constraint rule can be used to perform shape constraint processing on the N connected regions to determine one or more candidate regions; calculate the ratio between the area of each candidate region and the area of the source image; if the ratio reaches a preset ratio threshold, the candidate region is used as the object region of the source image. Among them, the control shape rule is used to define the preset shape of the object to be detected in the source image. In different business scenarios, the preset shapes of the objects to be detected can be the same or different. For example, in a game scenario, the object to be detected is a game control, and the control shape rule is used to define the preset shape corresponding to the game control, including a circle or a rectangle; in a traffic scenario, the object to be detected is a vehicle, and the control shape rule is used to define the preset shape corresponding to the vehicle as a rectangle, and so on.
[0187] The following details the specific process of determining the object region from the N connected regions.
[0188] (1) Shape constraint processing to determine candidate regions.
[0189] ① Perform shape constraint processing according to a rectangle. In a possible implementation, the control shape constraint rule is used to indicate that shape constraint processing is performed on the first connected region among the N connected regions according to a rectangular region; perform shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions. The specific process is as follows: Determine the first center point and at least one first edge point of the first connected region. The first edge point refers to the edge point included in the first connected region; calculate the first horizontal distance and the second vertical distance between the first edge point and the first center point; determine a rectangular region according to the first center point, the first horizontal distance, and the second vertical distance; calculate the first intersection-over-union ratio between the first connected region and the rectangular region; if the first intersection-over-union ratio is greater than or equal to the first preset threshold, the first connected region is determined as a candidate region. Among them, the first connected region can be any one of the N connected regions.
[0190] Specifically, taking the object to be detected as a game control as an example, since most game controls are rectangular or circular, when detecting a game control, whether its shape is rectangular or circular will be considered. For the first connected region, the center point of it (i.e., the first center point) and the boundary points of the region (i.e., the first edge points) can be calculated, and the farthest horizontal distance (i.e., the first horizontal distance) and the farthest vertical distance (i.e., the second vertical distance) from these boundary points to the center point can be calculated. Here, the first horizontal distance and the second vertical distance can be calculated according to the position coordinates of each pixel point. Then, according to the above information (the first center point, the first horizontal distance, the second vertical distance), a rectangular region can be constructed. This rectangular region includes length and width. Here, the length is twice the first horizontal distance, and the width is twice the second vertical distance. Finally, after determining the rectangular region, the first intersection over union (IOU) between the first connected region and the rectangular region can be calculated. The so-called intersection over union (IOU) is an index used to measure the overlapping degree of two regions, which means the ratio of the intersection to the union. Specifically, if the value of the first intersection over union is greater than or equal to the first preset threshold (such as 0.9), then it can be considered that the overlapping degree of the first connected region and the rectangular region is relatively high, and the first connected region is regarded as a candidate region of the game control.
[0191] ②Perform shape constraint processing according to a circle. In a possible implementation manner, the control shape constraint rule is used to indicate that shape constraint processing is performed on the second connected region among the N connected regions according to a circular region; perform shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions. The specific process is as follows: Determine the second center point and at least one second edge point of the second connected region. The second edge point refers to the edge point included in the second connected region; calculate the reference distance between the second edge point and the second center point; determine a circular region according to the second center point and the reference distance; calculate the second intersection over union between the second connected region and the circular region; if the second intersection over union is greater than or equal to the second preset threshold, then determine the second connected region as a candidate region.
[0192] Similarly, for the second connected region, its center point (the second center point) and the maximum distance (i.e., the reference distance) between the region boundary points (i.e., the second edge points) and the center point can be calculated. Here, the reference distance can also be calculated according to the position coordinates of each pixel point. Then, according to the above information (the second center point, the reference distance), a circular region can be constructed. The circular region includes the center and the radius. Here, the center is the second center point, and the radius is the reference distance. Finally, after determining the circular region, the second intersection-over-union ratio between the second connected region and the circular region can be calculated. If the value of the second intersection-over-union ratio is greater than or equal to the second preset threshold (e.g., 0.9), where the second preset threshold can be the same as or different from the first preset threshold, then it can be considered that the overlapping degree between the second connected region and the circular region is relatively high, and the second connected region is regarded as a candidate region of the game control.
[0193] ③ Perform shape constraint according to the associated region.
[0194] Optionally, an associated region associated with the reference connected region can also be determined from the N connected regions; where the reference connected region includes the first connected region or the second connected region; the reference connected region and the associated region are combined to obtain an updated region; shape constraint processing is performed on the updated region according to a rectangular region or a circular region to obtain the target edge detection result of the source image. Specifically, since the connected region may be a part of the game control, and there may be a situation where multiple connected regions form a game control, then this application can also combine the connected region and the adjacent region to generate a new region (i.e., the updated region), and then use the control shape constraint method shown in the above ① and ② for region filtering, so as to further determine the candidate region of the game control from the N connected regions.
[0195] (2) Determine the object region from the candidate regions according to the area ratio.
[0196] In specific implementation, after determining each candidate region, the ratio between the area of each candidate region and the area of the source image can be calculated; here, the preset ratio threshold of the game control can be set in advance according to experience between 1 / 500 and 1 / 25. That is to say, when the ratio between the area of a certain candidate region and the area of the source image is within the interval of the preset ratio threshold [1 / 500, 1 / 25], then the candidate region can be used as the final object region; otherwise, the candidate region can be deleted. In the above manner, at least one object region in the source image can be determined from multiple candidate regions. Please refer to Figure 5 , Figure 5 which is a schematic diagram of the scenario of the object region of the source image provided by the embodiment of this application. As Figure 5As shown, the source image can be a game image in a game screen. Based on the image processing method described in the embodiments of the present application, after performing edge detection processing on the game image, at least one game control can be determined from the game image. For example, the object area corresponding to the game control can include Figure 5 the object areas corresponding to S501, S502, S503, and S504 in
[0197] As shown in steps (1)-(2) above, the present application can, after determining N connected regions, further perform shape constraint processing on each connected region using shape prior information, so that the objects to be detected can better meet the actual business requirements. Further, after determining each candidate region, the present application makes a final screening according to the area ratio between the candidate region and the source image, so as to more accurately and effectively determine each object area in the source image.
[0198] In the embodiments of the present application, a source image to be processed, a first edge image of the source image, and M resolution images corresponding to the source image at M image scales can be obtained. One image scale corresponds to one resolution image, and M is a positive integer. Based on the M resolution images and the first edge image, an edge detection model is called to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and an edge detection result of the source image is determined based on the edge detection result of the first edge image. Here, the edge detection model is obtained after model training using an unsupervised learning method, and the edge detection result includes at least one detected edge point. By analyzing the edge detection result of the source image, N connected regions in the source image can be obtained, where N is a positive integer. Each connected region is composed of multiple edge points. Shape constraint processing is performed on the N connected regions to determine at least one object area in the source image from the N connected regions. Thus, on the one hand, the present application can perform edge detection processing on the source image using an edge detection model trained by an unsupervised learning method without manual annotation of training data, which can reduce the workload of users and thus improve the efficiency of image processing. On the other hand, the present application uses the first edge image of the source image as the input of the model, which can provide more edge information for the model and make the edge processing of the model more accurate. On the other hand, the present application further performs shape constraint processing on the edge detection result output by the model, so that the model detection result can be optimized, thereby improving the accuracy of the edge detection processing of the source image.
[0199] Please refer to Figure 6 , Figure 6 which is a schematic diagram of another image processing method provided by the embodiments of the present application. This method can be performed byFigure 2 is executed by a computer device (such as a terminal device or a server) in the image processing system shown. As Figure 6 shown, the image processing method mainly includes but is not limited to the following steps S601 - S607:
[0200] S601: Obtain a sample data set, where the sample data set includes multiple sample images and the edge annotation results of each sample image, and the edge annotation result of any sample image is obtained by at least one object annotation.
[0201] In a possible implementation manner, an edge data set can be obtained from an open - source database (such as the edge database BSDS500). The edge data set can be composed of 500 sample images and the corresponding edge annotation results of each sample image. Among them, each sample image can be annotated by multiple (such as five) objects. Since the generality of the edge contour is relatively strong, it can be directly used for the edge extraction processing of images subsequently. In this implementation manner, the sample data set is obtained from an existing edge database, and there is no need for additional manual image annotation, so the workload can be reduced and the processing efficiency can be improved.
[0202] S602: Construct an initial neural network model and use the sample data set to train the initial neural network model.
[0203] Specifically, a lightweight deep - network model can be constructed as the initial network model. A lightweight deep - network model refers to a network model with fewer model parameters or fewer network layers. For example, the initial neural network model can include but is not limited to: CNN (Convolutional neural networks) model, FasterRCNN (Faster Regions With Convolutional Neural Networks) model, SSD (Single Shot Multibox Detector) model, RefineDet (a refined detection network) model, and RPN (Region Proposal Network) model, etc. The embodiments of the present application do not specifically limit the model structure of the initial neural network model, as long as the initial neural network model has the ability to detect the edges of images.
[0204] The model training process of the initial neural network model will be described in detail below.
[0205] In a possible implementation, the computer device uses a sample data set to train an initial neural network model, mainly including the following steps (1)-(4):
[0206] (1) Invoke the initial neural network model to perform edge detection processing on any sample image in the sample data set, and obtain the edge detection result of any sample image. Optionally, the initial neural network model may include: a convolutional layer, an upsampling layer, and a pooling layer. ① The main function of the convolutional layer is: to extract local features of the input image, implement feature mapping through convolution operations. Each neuron in the convolutional layer is connected to a local area of the input image and learns features on this local area. Through the stacking of multiple convolutional layers, the network can learn higher-level and more abstract feature representations; ② The upsampling layer is mainly used to enlarge the size of the input image or feature map and restore the resolution of the image; ③ The pooling layer is used to reduce the dimension of the feature map, reduce the amount of computation and the number of parameters, thereby reducing the risk of model overfitting. In addition, the detailed process of the initial neural network model for processing any sample image can refer to Figure 3 the processing process of the edge detection model for the source image in step S302 of the embodiment. This application embodiment will not be elaborated here.
[0207] (2) Integrate and calculate the edge annotation results of any sample image in the sample data set to obtain the integrated sample annotation result. Specifically, the edge annotation result of any sample image may include edge annotation images obtained by annotating three objects. For example, it includes: edge annotation image 1, edge annotation image 2, edge annotation image 3 (where one object corresponds to one edge annotation image). Then, the edge annotation images of these three objects can be integrated and calculated. The integration calculation here may include: any one of average calculation and weighted calculation. For example, perform an average calculation on edge annotation image 1, edge annotation image 2, and edge annotation image 3 (specifically, add multiple edge annotation images annotated for the same sample image and then divide by the total number of edge annotation images included in the sample image), and the average edge image of the sample image can be obtained, and this average edge image is used as the sample annotation result of the current sample image. In this way, a sample image has multiple object annotations, and the annotation results of each object can be integrated, thereby improving the accuracy of the sample annotation result.
[0208] (3) Calculate the model loss of any sample image according to the sample annotation result of any sample image and the edge detection result of any sample image. Specifically, a loss function can be used to calculate the model loss of any sample image. The loss function here can include, but is not limited to: Loss function, Mean Squared Error (MSE) function, SmoothL1 Loss function, etc. For example, in this application, the Loss function can be used to calculate the model loss, and the formula of this Loss function is as follows:
[0209]
[0210] In the above formula, n is the number of pixel points included in the sample image, and y p is the pixel value corresponding to the p-th pixel point in the average edge image of the current sample image, and y' p is the pixel value corresponding to the p-th pixel point in the edge image output by the model. The goal of this loss function is to reduce the pixel difference between the edge image estimated by the network and the true average edge image.
[0211] (4) Iteratively adjust the model parameters of the initial neural network model according to the model loss of any sample image. Specifically, for each sample image in the sample data set, calculate the model loss according to the above steps (1)-(3), and then optimize the model parameters of the initial network model by minimizing the above model loss. By analogy, perform iterative training on the initial neural network model.
[0212] S603: When the initial neural network model reaches the model convergence condition, stop training the initial neural network model, and use the initial neural network model after stopping training as the edge detection model.
[0213] Specifically, the model convergence condition here can include any one of the following: when the number of training times of the initial neural network model reaches the preset training threshold, for example, 100 times, then the initial neural network model meets the model convergence condition; when the error between the edge detection result of any sample image and the edge annotation result of the sample image is less than the error threshold, then the initial neural network model meets the model convergence condition; when the change between the edge detection results obtained by two adjacent trainings of the initial neural network model is less than the change threshold, then the initial neural network model meets the model convergence condition.
[0214] S604: Obtain the source image to be processed, the first edge image of the source image, and M resolution images corresponding to the source image at M image scales respectively.
[0215] S605: Based on the M resolution images and the first edge image, call an edge detection model to perform edge detection processing on the source image to obtain the edge detection result of the source image.
[0216] S606: Analyze the edge detection result of the source image to obtain N connected regions in the source image.
[0217] S607: Perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0218] It should be noted that for the detailed steps executed by the computer device in steps S604 - S607 of the embodiments of the present application, reference may specifically be made to Figure 3 the relevant processes in steps S301 - S304 in the embodiments. The embodiments of the present application will not elaborate herein.
[0219] In a possible implementation, the source image is a game image in a target game, and the object region includes: a control region, a character region, and an item region. After the computer device performs shape constraint processing on the N connected regions to obtain the target edge detection result of the source image, the following operations may also be performed: ① If the object region is a control region, during the running of the target game, perform response time test processing on the game control indicated by the control region. For example, if the game control is a game control (such as Figure 5 shown in S501, S502, S503, S504), then after detecting each game control in the game using the above method of the present application, an automated test of the response time of the game control can be performed to facilitate subsequent testing and optimization of the relevant functions of the target game; ② If the object region is a character region, during the running of the target game, perform movement control processing on the game character indicated by the character region. For example, after identifying an enemy character in the target game, the enemy character can be automatically attacked, thereby enhancing the user's gaming experience; ③ If the object region is an item region, during the running of the target game, perform pick-up control processing on the game item indicated by the item region. For example, after identifying game items such as tools and ammunition in the target game, the game item can be automatically picked up, which can also enhance the user's gaming experience.
[0220] The following uses the accompanying drawings to give an example of the edge detection processing process in a game scenario.
[0221] Please refer to Figure 7 , Figure 7 which is a schematic diagram of a game scenario for image processing provided by the embodiments of the present application. As Figure 7As shown in the figure, the game scenario mainly involves: a game terminal and a game server. Among them, the target game runs on the game terminal, and a trained edge detection model is deployed on the game server. In the test scenario of the target game, ① the game terminal can run the target game, obtain the game screen (i.e., game image) presented at the current moment of the target game, and transmit the game image to the game server in real time as the source image; optionally, the game terminal can also perform edge extraction processing on the current game image to obtain the first edge image of the game image (such as a Canny edge image); in addition, the game terminal can also perform multi-scale scaling processing on the current game image to obtain multiple resolution images of the game image at multiple resolutions; ② after receiving the game image, the first edge image, and the multiple resolution images, the game server can call the edge detection model to perform edge detection processing on the game image to obtain the edge detection result of the current game image; ③ the game server returns the edge detection result to the game terminal, and the game terminal can analyze the edge detection result to obtain multiple connected regions in the game image, and perform shape constraint processing on these connected regions to determine at least one game control in the game image (such as the game control shown in S701) from the multiple connected regions; ④ the game terminal can trigger an automatic click on the current game control to display the detailed purchase page (such as the page shown in S702), so as to test whether the relevant functions in the current game are effective and reliable. It can be seen that this application can automatically detect game controls in the game screen in the game scenario, which is beneficial to realizing the automated test of game functions.
[0222] In the embodiments of the present application, an existing edge database can be used to construct a multi-scale lightweight initial neural network model, and the model is trained using an unsupervised learning method to obtain an edge detection model. It is possible to train a general edge detection model without adding a new game database. Since this application does not require additional labeled data, it can reduce the workload of users and improve the efficiency of model training; in addition, the embodiments of the present application propose a general detection scheme for an edge detection model and shape-constrained unsupervised game controls, which can optimize the edge detection result according to the shape, so the detection effect and accuracy of game controls can be further improved.
[0223] The following elaborates on the image processing device provided in the embodiments of the present application.
[0224] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an image processing device provided in the embodiments of the present application. As Figure 8As shown in the figure, the image processing device 800 can be applied to the computer device (such as a terminal device or a server) mentioned in the foregoing embodiments. Specifically, the image processing device 800 can be a computer program (including program code) running on the computer device. For example, the image processing device 800 is an application software. The image processing device 800 can be used to execute the corresponding steps in the image processing method provided in the embodiments of the present application. Specifically, when implemented, the image processing device 800 can specifically include:
[0225] An acquisition unit 801, configured to acquire a source image to be processed, a first edge image of the source image, and M resolution images corresponding to the source image at M image scales respectively. One image scale corresponds to one resolution image, and M is a positive integer;
[0226] A processing unit 802, configured to perform edge detection processing on the first edge image by invoking an edge detection model based on the M resolution images and the first edge image, obtain an edge detection result of the first edge image, and determine an edge detection result of the source image based on the edge detection result of the first edge image. The edge detection model is obtained by training the model in an unsupervised learning manner. The edge detection result includes at least one detected edge point;
[0227] The processing unit 802 is further configured to analyze the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer. Each connected region is composed of multiple edge points;
[0228] The processing unit 802 is further configured to perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0229] In a possible implementation manner, the processing unit 802 is further configured to perform the following operations:
[0230] Perform denoising processing on the source image to obtain a denoised source image;
[0231] Use an edge detection operator to calculate the gradient values of P pixel points in the denoised source image;
[0232] Based on the gradient values of the P pixel points, perform filtering processing on the P pixel points in a non-maximum suppression manner to obtain Q edge pixel points, where both P and Q are positive integers and Q ≤ P;
[0233] Perform double-threshold processing on the Q edge pixel points to divide the Q edge pixel points into q1 strong edge pixel points and q2 weak edge pixel points, where both q1 and q2 are positive integers;
[0234] Perform edge suppression processing on the q2 weak edge pixel points to obtain the first edge image of the source image.
[0235] In a possible implementation, when M = 3, the M resolution images include a first-resolution image, a second-resolution image, and a third-resolution image; the image scale of the source image includes a source width W * a source height H; the processing unit 802 is further configured to perform the following operations:
[0236] Perform multi-scale scaling processing on the source width and source height of the source image respectively according to a preset ratio to obtain a first-resolution image, a second-resolution image, and a third-resolution image;
[0237] wherein, the image scale of the first-resolution image is a first width X * a first height X / 2;
[0238] the image scale of the second-resolution image is a second width X / 2 * a second height X / 4;
[0239] the image scale of the third-resolution image is a third width X / 4 * a third height X / 8;
[0240] wherein, the source width W of the source image is the same as or different from the scaled widths X, X / 2, X / 4; and, the source height H of the source image is the same as or different from the scaled heights X / 2, X / 4, X / 8.
[0241] In a possible implementation, the image input channels of the edge detection model include a color channel and an edge channel; the first edge image serves as the input image of the edge channel, and any one of the resolution images serves as the input image of the color channel; wherein, the edge detection model includes: a convolutional layer, an upsampling layer, and a pooling layer;
[0242] Based on the M resolution images and the first edge image, the processing unit 802 calls the edge detection model to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, which is used to perform the following operations:
[0243] Perform feature extraction processing on the M resolution images and the first edge image by using the convolutional layer to obtain convolutional image features of the source image;
[0244] Perform size amplification processing on the convolutional image features by using the upsampling layer to obtain processed convolutional image features;
[0245] Perform pooling processing on the feature dimensions of the processed convolutional image features by using the pooling layer to obtain a second edge image of the source image;
[0246] Based on the first edge image and the second edge image, obtain the edge detection result of the first edge image.
[0247] In a possible implementation, the convolutional layer includes: a first-scale convolutional layer, a second-scale convolutional layer, and a third-scale convolutional layer. The edge detection model further includes a cascading layer. The processing unit 802 uses the convolutional layer to perform feature extraction processing on M resolution images and the first edge image to obtain the convolutional image features of the source image, for performing the following operations:
[0248] Use the first-scale convolutional layer to perform feature extraction processing on the first resolution image and the first edge image to obtain the first convolutional features of the first resolution image;
[0249] Use the second-scale convolutional layer to perform feature extraction processing on the second resolution image and the first edge image to obtain the second convolutional features of the second resolution image;
[0250] Use the third-scale convolutional layer to perform feature extraction processing on the third resolution image and the first edge image to obtain the third convolutional features of the third resolution image;
[0251] Call the cascading layer to perform cascading processing on the first convolutional features, the second convolutional features, and the third convolutional features to obtain the convolutional image features of the source image.
[0252] In a possible implementation, the first edge image includes I edge points, and the second edge image includes J edge points, where both I and J are positive integers. The processing unit 802 obtains the edge detection result of the first edge image based on the first edge image and the second edge image, for performing the following operations:
[0253] Obtain the first pixel value of the i-th edge point in the first edge image, where i is a positive integer and 1 ≤ i ≤ I;
[0254] Obtain the second pixel value of the j-th edge point in the second edge image, where j is a positive integer and 1 ≤ j ≤ J;
[0255] Perform an intersection operation on the first pixel value of the i-th edge point in the first edge image and the second pixel value of the j-th edge point in the second edge image to obtain an operation result;
[0256] Based on the operation result, obtain the edge detection result of the first edge image;
[0257] Wherein, the i-th edge point in the first edge image and the j-th pixel point in the second edge image refer to the pixel points at the same position in the source image.
[0258] In a possible implementation, the edge detection result includes an edge detection image composed of multiple edge points. The processing unit 802 analyzes the edge detection result of the source image to obtain N connected regions in the source image, for performing the following operations:
[0259] Perform dilation processing on the edge detection image to obtain the dilated edge detection image;
[0260] Traverse each pixel point in the dilated edge detection image to obtain a traversal result;
[0261] Determine N connected regions in the source image according to the traversal result.
[0262] In a possible implementation, the edge detection image includes at least one edge point; any pixel point in the dilated edge detection image is represented as pixel point a; the processing unit 802 traverses each pixel point in the dilated edge detection image to obtain a traversal result for performing the following operations:
[0263] In the dilated edge detection image, obtain K associated points associated with pixel point a, where K is a positive integer;
[0264] If any one of the K associated points does not belong to the edge points in the edge detection image, mark pixel point a as a background point;
[0265] Perform pixel point expansion with pixel point a as the center point;
[0266] After stopping the pixel point expansion, obtain one or more pixel points that have been visited, and obtain a traversal result according to each visited pixel point.
[0267] In a possible implementation, the processing unit 802 performs shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions for performing the following operations:
[0268] Perform shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions;
[0269] Calculate the ratio between the area of each candidate region and the area of the source image;
[0270] If the ratio reaches a preset ratio threshold, use the candidate region as the object region of the source image.
[0271] In a possible implementation, the control shape constraint rule is used to indicate performing shape constraint processing on the first connected region among the N connected regions according to a rectangular region; the processing unit 802 performs shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions for performing the following operations:
[0272] Determine the first center point and at least one first edge point of the first connected region, where the first edge point refers to the edge point included in the first connected region;
[0273] Calculate a first horizontal distance and a second vertical distance between a first edge point and a first center point;
[0274] Determine a rectangular area according to the first center point, the first horizontal distance, and the second vertical distance;
[0275] Calculate a first intersection over union (IoU) between the first connected area and the rectangular area;
[0276] If the first IoU is greater than or equal to a first preset threshold, determine the first connected area as a candidate area.
[0277] In a possible implementation, the control shape constraint rule is used to indicate that shape constraint processing is performed on a second connected area among the N connected areas according to a circular area; the processing unit 802 performs shape constraint processing on the N connected areas according to the control shape constraint rule to determine one or more candidate areas from the N connected areas for performing the following operations:
[0278] Determine a second center point and at least one second edge point of the second connected area, where the second edge point refers to an edge point included in the second connected area;
[0279] Calculate a reference distance between the second edge point and the second center point;
[0280] Determine a circular area according to the second center point and the reference distance;
[0281] Calculate a second IoU between the second connected area and the circular area;
[0282] If the second IoU is greater than or equal to a second preset threshold, determine the second connected area as a candidate area.
[0283] In a possible implementation, the processing unit 802 is further configured to perform the following operations:
[0284] Determine an associated area associated with a reference connected area from the N connected areas; where the reference connected area includes the first connected area or the second connected area;
[0285] Combine the reference connected area and the associated area to obtain an updated area;
[0286] Perform shape constraint processing on the updated area according to the rectangular area or the circular area to obtain a target edge detection result of the source image.
[0287] In a possible implementation, the processing unit 802 is further configured to perform the following operations:
[0288] Obtain a sample data set, where the sample data set includes multiple sample images and the edge annotation results of each sample image, and the edge annotation result of any sample image is obtained by annotating at least one object;
[0289] Construct an initial neural network model and use the sample data set to train the initial neural network model;
[0290] After the initial neural network model reaches the model convergence condition, stop training the initial neural network model and use the initial neural network model after stopping training as the edge detection model.
[0291] In a possible implementation, the processing unit 802 uses the sample data set to train the initial neural network model and is used to perform the following operations:
[0292] Call the initial neural network model to perform edge detection processing on any sample image in the sample data set to obtain the edge detection result of any sample image;
[0293] Perform integration operations on the edge annotation results of any sample image in the sample data set to obtain the integrated sample annotation result;
[0294] Calculate the model loss of any sample image according to the sample annotation result of any sample image and the edge detection result of any sample image;
[0295] Iteratively adjust the model parameters of the initial neural network model according to the model loss of any sample image.
[0296] In a possible implementation, the source image is a game image in the target game, and the object area includes: a control area, a character area, and an item area; after the processing unit 802 performs shape constraint processing on the N connected regions to obtain the target edge detection result of the source image, it is also used to perform the following operations:
[0297] If the object area is the control area, perform response time test processing on the game control indicated by the control area during the operation of the target game;
[0298] If the object area is the character area, perform movement control processing on the game character indicated by the character area during the operation of the target game;
[0299] If the object area is the item area, perform pick-up control processing on the game item indicated by the item area during the operation of the target game.
[0300] In the embodiments of the present application, a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales can be obtained, where one image scale corresponds to one resolution image, and M is a positive integer. Based on the M resolution images and the first edge image, an edge detection model is called to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and an edge detection result of the source image is determined based on the edge detection result of the first edge image. Here, the edge detection model is obtained after model training using an unsupervised learning method, and the edge detection result includes at least one detected edge point. By analyzing the edge detection result of the source image, N connected regions in the source image can be obtained, where N is a positive integer. Each connected region is composed of multiple edge points. Shape constraint processing is performed on the N connected regions to determine at least one object region in the source image from the N connected regions. It can be seen that, on the one hand, the present application can perform edge detection processing on the source image using an edge detection model trained by an unsupervised learning method without manual annotation of training data, which can reduce the workload of users and thus improve the efficiency of image processing. On the other hand, the present application uses the first edge image of the source image as the input of the model, which can provide more edge information for the model and make the edge processing of the model more accurate. On the further hand, the present application further performs shape constraint processing on the edge detection result output by the model, so that the detection result of the model can be optimized, thereby improving the accuracy of the edge detection processing of the source image.
[0301] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 900 is used to execute the steps performed by the terminal device or the server in the foregoing method embodiment. The computer device 900 includes: one or more processors 901; one or more input devices 902, one or more output devices 903, and a memory 904. The foregoing processors 901, input devices 902, output devices 903, and memory 904 are connected through a bus 905. Among them, the memory 904 is used to store a computer program, and the computer program includes program instructions. Specifically, the processor 901 is used to call the program instructions stored in the memory 904 to perform the following operations:
[0302] Obtain a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales, where one image scale corresponds to one resolution image, and M is a positive integer;
[0303] Based on M resolution images and the first edge image, call an edge detection model to perform edge detection processing on the first edge image to obtain the edge detection result of the first edge image, and determine the edge detection result of the source image based on the edge detection result of the first edge image; the edge detection model is obtained after model training using an unsupervised learning method; the edge detection result includes at least one detected edge point;
[0304] Analyze the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; among them, each connected region is composed of multiple edge points;
[0305] Perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
[0306] In a possible implementation manner, the processor 901 is further configured to perform the following operations:
[0307] Perform denoising processing on the source image to obtain the denoised source image;
[0308] Adopt an edge detection operator to calculate the gradient values of P pixel points in the denoised source image;
[0309] Based on the gradient values of the P pixel points, perform filtering processing on the P pixel points in a non-maximum suppression manner to obtain Q edge pixel points, where both P and Q are positive integers and Q ≤ P;
[0310] Perform double-threshold processing on the Q edge pixel points to divide the Q edge pixel points into q1 strong edge pixel points and q2 weak edge pixel points, where both q1 and q2 are positive integers;
[0311] Perform edge suppression processing on the q2 weak edge pixel points to obtain the first edge image of the source image.
[0312] In a possible implementation manner, when M = 3, the M resolution images include a first resolution image, a second resolution image, and a third resolution image; the image scale of the source image includes the source width W * the source height H; the processor 901 is further configured to perform the following operations:
[0313] Perform multi-scale scaling processing on the source width and source height of the source image respectively according to a preset ratio to obtain the first resolution image, the second resolution image, and the third resolution image;
[0314] The image scale of the first resolution image is the first width X * the first height X / 2;
[0315] The image scale of the second resolution image is the second width X / 2 * the second height X / 4;
[0316] The image scale of the third-resolution image is the third width X / 4 * the third height X / 8;
[0317] Wherein, the source width W of the source image is the same as or different from the scaled widths X, X / 2, X / 4; and the source height H of the source image is the same as or different from the scaled heights X / 2, X / 4, X / 8.
[0318] In a possible implementation, the image input channels of the edge detection model include a color channel and an edge channel; the first edge image is used as the input image of the edge channel, and any resolution image is used as the input image of the color channel; wherein, the edge detection model includes: a convolutional layer, an upsampling layer, and a pooling layer;
[0319] The processor 901 invokes the edge detection model to perform edge detection processing on the first edge image based on the M resolution images and the first edge image, and obtains the edge detection result of the first edge image for performing the following operations:
[0320] Performing feature extraction processing on the M resolution images and the first edge image by using the convolutional layer to obtain the convolutional image features of the source image;
[0321] Performing size amplification processing on the convolutional image features by using the upsampling layer to obtain the processed convolutional image features;
[0322] Performing pooling processing on the feature dimensions of the processed convolutional image features by using the pooling layer to obtain the second edge image of the source image;
[0323] Based on the first edge image and the second edge image, obtaining the edge detection result of the first edge image.
[0324] In a possible implementation, the convolutional layer includes: a first-scale convolutional layer, a second-scale convolutional layer, and a third-scale convolutional layer, and the edge detection model further includes a cascade layer; the processor 901 performs feature extraction processing on the M resolution images and the first edge image by using the convolutional layer to obtain the convolutional image features of the source image for performing the following operations:
[0325] Performing feature extraction processing on the first resolution image and the first edge image by using the first-scale convolutional layer to obtain the first convolutional features of the first resolution image;
[0326] Performing feature extraction processing on the second resolution image and the first edge image by using the second-scale convolutional layer to obtain the second convolutional features of the second resolution image;
[0327] Performing feature extraction processing on the third resolution image and the first edge image by using the third-scale convolutional layer to obtain the third convolutional features of the third resolution image;
[0328] Call the cascading layer to cascade the first convolutional feature, the second convolutional feature, and the third convolutional feature to obtain the convolutional image feature of the source image.
[0329] In a possible implementation, the first edge image includes I edge points, the second edge image includes J edge points, and both I and J are positive integers; the processor 901 obtains the edge detection result of the first edge image based on the first edge image and the second edge image for performing the following operations:
[0330] Obtain the first pixel value of the i-th edge point in the first edge image, where i is a positive integer and 1 ≤ i ≤ I;
[0331] Obtain the second pixel value of the j-th edge point in the second edge image, where j is a positive integer and 1 ≤ j ≤ J;
[0332] Perform an intersection operation on the first pixel value of the i-th edge point in the first edge image and the second pixel value of the j-th edge point in the second edge image to obtain an operation result;
[0333] Based on the operation result, obtain the edge detection result of the first edge image;
[0334] Wherein, the i-th edge point in the first edge image and the j-th pixel point in the second edge image refer to the pixel points at the same position in the source image.
[0335] In a possible implementation, the edge detection result includes an edge detection image composed of multiple edge points; the processor 901 analyzes the edge detection result of the source image to obtain N connected regions in the source image for performing the following operations:
[0336] Perform a dilation process on the edge detection image to obtain a dilated edge detection image;
[0337] Traverse each pixel point in the dilated edge detection image to obtain a traversal result;
[0338] Determine N connected regions in the source image according to the traversal result.
[0339] In a possible implementation, the edge detection image includes at least one edge point; any pixel point in the dilated edge detection image is represented as pixel point a; the processor 901 traverses each pixel point in the dilated edge detection image to obtain a traversal result for performing the following operations:
[0340] In the dilated edge detection image, obtain K associated points associated with pixel point a, where K is a positive integer;
[0341] If none of the K associated points belong to the edge points in the edge detection image, the pixel point a is marked as a background point;
[0342] Perform pixel point expansion with the pixel point a as the center point;
[0343] After stopping the pixel point expansion, obtain one or more pixel points that have been visited, and obtain the traversal result according to each visited pixel point.
[0344] In a possible implementation, the processor 901 performs shape constraint processing on N connected regions to determine at least one object region in the source image from the N connected regions for performing the following operations:
[0345] Perform shape constraint processing on the N connected regions according to the control shape constraint rules to determine one or more candidate regions from the N connected regions;
[0346] Calculate the ratio between the area of each candidate region and the area of the source image;
[0347] If the ratio reaches the preset ratio threshold, the candidate region is used as the object region of the source image.
[0348] In a possible implementation, the control shape constraint rules are used to indicate performing shape constraint processing on the first connected region among the N connected regions according to a rectangular region; the processor 901 performs shape constraint processing on the N connected regions according to the control shape constraint rules to determine one or more candidate regions from the N connected regions for performing the following operations:
[0349] Determine the first center point and at least one first edge point of the first connected region, where the first edge point refers to the edge point included in the first connected region;
[0350] Calculate the first horizontal distance and the second vertical distance between the first edge point and the first center point;
[0351] Determine the rectangular region according to the first center point, the first horizontal distance, and the second vertical distance;
[0352] Calculate the first intersection over union between the first connected region and the rectangular region;
[0353] If the first intersection over union is greater than or equal to the first preset threshold, the first connected region is determined as the candidate region.
[0354] In a possible implementation, the control shape constraint rule is used to indicate that shape constraint processing is performed on the second connected region among the N connected regions according to a circular region; the processor 901 performs shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions for performing the following operations:
[0355] Determine the second center point and at least one second edge point of the second connected region, where the second edge point refers to the edge point included in the second connected region;
[0356] Calculate the reference distance between the second edge point and the second center point;
[0357] Determine a circular region according to the second center point and the reference distance;
[0358] Calculate the second intersection-over-union ratio between the second connected region and the circular region;
[0359] If the second intersection-over-union ratio is greater than or equal to the second preset threshold, determine the second connected region as a candidate region.
[0360] In a possible implementation, the processor 901 is further configured to perform the following operations:
[0361] Determine an associated region associated with the reference connected region from the N connected regions; where the reference connected region includes the first connected region or the second connected region;
[0362] Combine the reference connected region and the associated region to obtain an updated region;
[0363] Perform shape constraint processing on the updated region according to a rectangular region or a circular region to obtain the target edge detection result of the source image.
[0364] In a possible implementation, the processor 901 is further configured to perform the following operations:
[0365] Obtain a sample data set, where the sample data set includes multiple sample images and the edge annotation result of each sample image, and the edge annotation result of any sample image is obtained by annotating at least one object;
[0366] Construct an initial neural network model and train the initial neural network model using the sample data set;
[0367] When the initial neural network model reaches the model convergence condition, stop training the initial neural network model and use the initial neural network model after stopping training as the edge detection model.
[0368] In a possible implementation, the processor 901 uses a sample data set to train an initial neural network model for performing the following operations:
[0369] Call the initial neural network model to perform edge detection processing on any sample image in the sample data set to obtain the edge detection result of any sample image;
[0370] Perform integration operations on the edge annotation results of any sample image in the sample data set to obtain the integrated sample annotation results;
[0371] Calculate the model loss of any sample image according to the sample annotation result of any sample image and the edge detection result of any sample image;
[0372] Iteratively adjust the model parameters of the initial neural network model according to the model loss of any sample image.
[0373] In a possible implementation, the source image is a game image in a target game, and the object area includes: a control area, a character area, and an item area; after the processor 901 performs shape constraint processing on N connected areas to obtain the target edge detection result of the source image, it is further used to perform the following operations:
[0374] If the object area is a control area, perform response time test processing on the game control indicated by the control area during the running of the target game;
[0375] If the object area is a character area, perform movement control processing on the game character indicated by the character area during the running of the target game;
[0376] If the object area is an item area, perform pickup control processing on the game item indicated by the item area during the running of the target game.
[0377] In the embodiments of the present application, a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales can be obtained, where one image scale corresponds to one resolution image, and M is a positive integer. Based on the M resolution images and the first edge image, an edge detection model is called to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and an edge detection result of the source image is determined based on the edge detection result of the first edge image. Here, the edge detection model is obtained after model training using an unsupervised learning method, and the edge detection result includes at least one detected edge point. By analyzing the edge detection result of the source image, N connected regions in the source image can be obtained, where N is a positive integer. Each connected region is composed of multiple edge points. Shape constraint processing is performed on the N connected regions to determine at least one object region in the source image from the N connected regions. Thus, on the one hand, the present application can perform edge detection processing on the source image using an edge detection model trained by an unsupervised learning method, without the need for manual annotation of training data, which can reduce the workload of users and thus improve the efficiency of image processing. On the other hand, the present application uses the first edge image of the source image as the input of the model, which can provide more edge information for the model and make the edge processing of the model more accurate. On the further hand, the present application further performs shape constraint processing on the edge detection result output by the model, so that the model detection result can be optimized, thereby improving the accuracy of the edge detection processing of the source image.
[0378] In the above embodiments, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0379] In addition, it should be noted here that: The embodiments of the present application also provide a computer storage medium, and the computer storage medium stores a computer program, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the methods in the corresponding embodiments described above. Therefore, the description will not be repeated here. For the technical details not disclosed in the embodiments of the computer storage medium involved in the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located at one place, or executed on multiple computer devices distributed at multiple places and interconnected through a communication network.
[0380] According to one aspect of the present application, embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can execute the methods in the corresponding embodiments described above. Therefore, details will not be repeated here.
[0381] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data processing device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0382] The foregoing disclosure is only for the preferred embodiments of the present application, and of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. An image processing method, characterized in that, Including: Obtaining a source image to be processed, a first edge image of the source image, and M resolution images corresponding to the source image at M image scales respectively, where one image scale corresponds to one resolution image, and M is a positive integer; Based on the M resolution images and the first edge image, calling an edge detection model to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, and determining an edge detection result of the source image based on the edge detection result of the first edge image; The edge detection model is obtained after model training using an unsupervised learning method; The edge detection result includes at least one detected edge point; Analyzing the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; wherein each of the connected regions is composed of multiple of the edge points; Performing shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
2. The method according to claim 1, characterized in that, The method further includes: Performing denoising processing on the source image to obtain a denoised source image; Using an edge detection operator to calculate gradient values of P pixel points in the denoised source image; Based on the gradient values of the P pixel points, filtering the P pixel points using a non-maximum suppression method to obtain Q edge pixel points, where both P and Q are positive integers and Q ≤ P; Performing double-threshold processing on the Q edge pixel points to divide the Q edge pixel points into q1 strong edge pixel points and q2 weak edge pixel points, where both q1 and q2 are positive integers; Performing edge suppression processing on the q2 weak edge pixel points to obtain the first edge image of the source image.
3. The method according to claim 1, characterized in that, When M = 3, the M resolution images include a first resolution image, a second resolution image, and a third resolution image; the image scales of the source image include a source width W and a source height H; the method further includes: Performing multi-scale scaling processing on the source width and source height of the source image respectively according to a preset ratio to obtain the first resolution image, the second resolution image, and the third resolution image; The image scale of the first resolution image is a first width X * a first height X / 2; The image scale of the second resolution image is a second width X / 2 * a second height X / 4; The image scale of the third resolution image is a third width X / 4 * a third height X / 8; Wherein, the source width W of the source image is the same as or different from the scaled widths X, X / 2, X / 4; and the source height H of the source image is the same as or different from the scaled heights X / 2, X / 4, X / 8.
4. The method according to claim 3, wherein The image input channels of the edge detection model include a color channel and an edge channel; the first edge image is used as the input image of the edge channel, and any one of the resolution images is used as the input image of the color channel; wherein, the edge detection model includes: a convolutional layer, an upsampling layer, and a pooling layer; Based on the M resolution images and the first edge image, calling an edge detection model to perform edge detection processing on the first edge image to obtain an edge detection result of the first edge image, including: Using the convolutional layer to perform feature extraction processing on the M resolution images and the first edge image to obtain convolutional image features of the source image; Using the upsampling layer to perform size amplification processing on the convolutional image features to obtain the processed convolutional image features; Using the pooling layer to perform pooling processing on the feature dimensions of the processed convolutional image features to obtain a second edge image of the source image; Based on the first edge image and the second edge image, obtaining an edge detection result of the first edge image.
5. The method according to claim 4, wherein The convolutional layer includes: a first-scale convolutional layer, a second-scale convolutional layer, and a third-scale convolutional layer, and the edge detection model further includes a cascading layer; using the convolutional layer to perform feature extraction processing on the M resolution images and the first edge image to obtain convolutional image features of the source image, including: Using the first-scale convolutional layer to perform feature extraction processing on the first resolution image and the first edge image to obtain first convolutional features of the first resolution image; Using the second-scale convolutional layer to perform feature extraction processing on the second resolution image and the first edge image to obtain second convolutional features of the second resolution image; Using the third-scale convolutional layer to perform feature extraction processing on the third resolution image and the first edge image to obtain third convolutional features of the third resolution image; Calling the cascading layer to perform cascading processing on the first convolutional features, second convolutional features, and third convolutional features to obtain convolutional image features of the source image.
6. The method according to claim 5, characterized in that, The first edge image includes I edge points, and the second edge image includes J edge points, where both I and J are positive integers; based on the first edge image and the second edge image, obtaining an edge detection result of the first edge image, including: Obtaining a first pixel value of the i-th edge point in the first edge image, where i is a positive integer and 1 ≤ i ≤ I; Obtaining a second pixel value of the j-th edge point in the second edge image, where j is a positive integer and 1 ≤ j ≤ J; Performing an intersection operation on the first pixel value of the i-th edge point in the first edge image and the second pixel value of the j-th edge point in the second edge image to obtain an operation result; Based on the operation result, obtaining an edge detection result of the first edge image; Wherein, the i-th edge point in the first edge image and the j-th pixel point in the second edge image refer to pixel points at the same position in the source image.
7. The method according to claim 1, characterized in that, The edge detection result includes an edge detection image composed of multiple edge points; analyzing the edge detection result of the source image to obtain N connected regions in the source image, including: Performing dilation processing on the edge detection image to obtain a dilated edge detection image; Traversing each pixel point in the dilated edge detection image to obtain a traversal result; Determine N connected regions in the source image according to the traversal result.
8. The method according to claim 7, characterized in that, The edge detection image includes at least one edge point; any pixel point in the dilated edge detection image is represented as pixel point a; traversing each pixel point in the dilated edge detection image to obtain a traversal result, including: In the dilated edge detection image, obtain K associated points associated with pixel point a, where K is a positive integer; If any one of the K associated points does not belong to the edge points in the edge detection image, mark pixel point a as a background point; Perform pixel point expansion with pixel point a as the center point; After stopping the pixel point expansion, obtain one or more pixel points that have been visited, and obtain the traversal result according to each of the visited pixel points.
9. The method according to claim 1, characterized in that, Performing shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions, including: Performing shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions; Calculate the ratio between the area of each candidate region and the area of the source image; If the ratio reaches a preset ratio threshold, use the candidate region as the object region of the source image.
10. The method according to claim 9, characterized in that, The control shape constraint rule is used to indicate performing shape constraint processing on the first connected region among the N connected regions according to a rectangular region; performing shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions, including: Determine the first center point and at least one first edge point of the first connected region, where the first edge point refers to the edge point included in the first connected region; Calculate the first horizontal distance and the second vertical distance between the first edge point and the first center point; Determine a rectangular region according to the first center point, the first horizontal distance, and the second vertical distance; Calculate the first intersection-over-union ratio between the first connected region and the rectangular region; If the first intersection-over-union ratio is greater than or equal to a first preset threshold, determine the first connected region as a candidate region.
11. The method according to claim 10, characterized in that The control shape constraint rule is used to indicate performing shape constraint processing on the second connected region among the N connected regions according to a circular region; performing shape constraint processing on the N connected regions according to the control shape constraint rule to determine one or more candidate regions from the N connected regions, including: Determine the second center point and at least one second edge point of the second connected region, where the second edge point refers to the edge point included in the second connected region; Calculate the reference distance between the second edge point and the second center point; Determine a circular region according to the second center point and the reference distance; Calculate the second intersection-over-union ratio between the second connected region and the circular region; If the second intersection-over-union ratio is greater than or equal to a second preset threshold, determine the second connected region as a candidate region.
12. The method according to any one of claims 9-11, characterized in that, The method further includes: Determine an associated region associated with a reference connected region from the N connected regions; wherein, the reference connected region includes the first connected region or the second connected region; Combine the reference connected region and the associated region to obtain an updated region; Perform shape constraint processing on the updated region according to a rectangular region or a circular region to obtain the target edge detection result of the source image.
13. The method according to claim 1, wherein The method further includes: Obtain a sample data set, the sample data set includes a plurality of sample images and the edge annotation result of each sample image, and the edge annotation result of any sample image is obtained by annotating at least one object; Construct an initial neural network model and use the sample data set to train the initial neural network model; When the initial neural network model reaches the model convergence condition, stop training the initial neural network model, and use the initial neural network model after stopping training as the edge detection model.
14. The method according to claim 13, wherein The using the sample data set to train the initial neural network model includes: Call the initial neural network model to perform edge detection processing on any sample image in the sample data set to obtain the edge detection result of any sample image; Perform integration operation on the edge annotation result of any sample image in the sample data set to obtain the integrated sample annotation result; Calculate the model loss of any sample image according to the sample annotation result of any sample image and the edge detection result of any sample image; Iteratively adjust the model parameters of the initial neural network model according to the model loss of any sample image.
15. The method according to claim 1, characterized in that, The source image is a game image in a target game, and the object region includes: a control region, a character region, and an item region; after performing shape constraint processing on the N connected regions to obtain the target edge detection result of the source image, it further includes: If the object region is a control region, perform response time test processing on the game control indicated by the control region during the running of the target game; If the object region is a character region, perform movement control processing on the game character indicated by the character region during the running of the target game; If the object region is an item region, perform pickup control processing on the game item indicated by the item region during the running of the target game.
16. An image processing apparatus, characterized in that, Includes: An acquisition unit, configured to acquire a source image to be processed, a first edge image of the source image, and M resolution images respectively corresponding to the source image at M image scales, one image scale corresponding to one resolution image, and M being a positive integer; A processing unit, configured to, based on the M resolution images and the first edge image, call an edge detection model to perform edge detection processing on the first edge image, obtain an edge detection result of the first edge image, and determine an edge detection result of the source image based on the edge detection result of the first edge image; the edge detection model is obtained by performing model training in an unsupervised learning manner; the edge detection result includes at least one detected edge point; The processing unit is further configured to analyze the edge detection result of the source image to obtain N connected regions in the source image, where N is a positive integer; each of the connected regions is composed of a plurality of the edge points; The processing unit is further configured to perform shape constraint processing on the N connected regions to determine at least one object region in the source image from the N connected regions.
17. A computer device, characterized in that, Comprising: A storage device and a processor; A memory, in which one or more computer programs are stored; A processor, configured to load the one or more computer programs to implement the image processing method according to any one of claims 1-15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the image processing method according to any one of claims 1-15.
19. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the image processing method according to any one of claims 1-15.