Picture main body detection method and related device
By extracting and reducing the dimensions of the detected images, the target pixels at the candidate edges are selected, and the problems of difficulty and low efficiency of picture subject recognition in the prior art are solved, and efficient and accurate picture subject detection is achieved.
Patent Information
- Application Number
- CN202410107734.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the recognition of the picture subject usually uses contour detection and edge detection, which has problems such as high detection difficulty, low efficiency and prone to identification errors, especially when there are strong patterns on the edges.
By extracting the image to be detected, a feature map is obtained, and dimensionality reduction is performed in the width and height directions, the target pixels at the associated candidate edge are filtered out, and the main area of the picture is redefined as key point detection, which is simplified into horizontal, vertical and horizontal key point detection in the image.
The efficiency and accuracy of screen subject detection are improved, ensuring that the screen subject area can be determined through up to four target pixels, reducing detection complexity and improving detection accuracy.
Smart Images

Figure CN120374655A_ABST
Abstract
Description
Background Art
[0002] Since the aspect ratios of content such as videos and images are inconsistent with the aspect ratios specified by the content display platform, when uploading content, it is necessary to use the content as the main body of the picture and add a border outside the main body of the picture to conform to the aspect ratio of the content display platform. To implement functions such as content recommendation and content detection, the content display platform needs to perform semantic understanding on the content. However, the additional border often affects the semantic understanding result, so it is necessary to identify the main body of the picture.
[0003] In related technologies, the identification of the main body of the picture usually adopts two methods: contour detection and edge detection. In contour detection, the square main body in the content is identified as the main body of the picture, while in edge detection, the straight line of the border is identified to determine the main body of the picture.
[0004] However, whether it is edge detection or contour detection, it is to detect the significant features in the content (such as the straight line of the border, the square contour of the main body of the picture), which requires a large amount of logical design, the detection difficulty is large, and the detection efficiency is affected. Especially when there are strongly contrasting patterns at the edge of the main body of the picture, it is difficult to accurately detect the straight line. At the same time, due to the strong diversity of the border styles, the identification of the main body of the picture is relatively difficult, and there may be situations of misidentification and inaccurate identification. Summary of the Invention
[0005] Embodiments of the present application provide a method and related device for detecting the main body of a picture to improve the detection accuracy and detection efficiency of the main body of the picture.
[0006] In a first aspect, embodiments of the present application provide a method for detecting the main body of a picture, including:
[0007] Performing feature extraction on the image to be detected to obtain a feature map, where the main body area of the image to be detected is rectangular;
[0008] Reducing the dimension of the feature map in the width direction to obtain a column of candidate pixels having the same height as the feature map, and each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected that passes through the associated candidate pixel;
[0009] Reducing the dimension of the feature map in the height direction to obtain a row of candidate pixels having the same width as the feature map, and each candidate pixel in the row of candidate pixels is associated with a candidate high edge, and each candidate high edge is: a column of pixels in the image to be detected that passes through the associated candidate pixel;
[0010] Based on the respective values of the column of candidate pixels and the row of candidate pixels, at least one target pixel is selected from the column of candidate pixels and the row of candidate pixels;
[0011] Based on the candidate edges respectively associated with the at least one target pixel, obtain the main body region of the image to be detected, where each candidate edge is a candidate wide edge or a candidate high edge.
[0012] In a second aspect, an embodiment of the present application provides a main body detection device for an image, including:
[0013] A feature extraction unit, configured to perform feature extraction on the image to be detected to obtain a feature map, where the main body region of the image to be detected is a rectangle;
[0014] A column vector extraction unit, configured to perform dimensionality reduction on the feature map in the width direction to obtain a column of candidate pixels having the same height as the feature map, and each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected passing through the associated candidate pixel;
[0015] A row vector extraction unit, configured to perform dimensionality reduction on the feature map in the height direction to obtain a row of candidate pixels having the same width as the feature map, and each candidate pixel in the row of candidate pixels is associated with a candidate high edge, and each candidate high edge is: a column of pixels in the image to be detected passing through the associated candidate pixel;
[0016] A pixel screening unit, configured to screen out at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels;
[0017] A region determination unit, configured to obtain the main body region of the image to be detected based on the candidate edges respectively associated with the at least one target pixel, where each candidate edge is a candidate wide edge or a candidate high edge.
[0018] As a possible implementation manner, the pixel value is assigned by a first assignment method. In the first assignment method, a first pixel value is used to represent that the candidate edge associated with the candidate pixel belongs to the edge of the main body region of the image, and a second pixel value is used to represent that the candidate edge associated with the candidate pixel does not belong to the edge of the main body region of the image;
[0019] Then, when screening out at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels, the pixel screening unit is specifically configured to:
[0020] Screen out at least one candidate pixel with a value of the first pixel value from the column of candidate pixels and the row of candidate pixels as the target pixel.
[0021] As a possible implementation, the pixel values are assigned using a second assignment method. In the second assignment method, a set of pixel values between a first pixel value and a second pixel value is used to represent the probability that the candidate wide edge associated with the candidate pixel belongs to the edge of the main area of the picture.
[0022] When screening at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels, the pixel screening unit is specifically configured to:
[0023] Based on the respective values of the column of candidate pixels and the row of candidate pixels, screen out at least one reference pixel that meets the set local value condition from the column of candidate pixels and the row of candidate pixels;
[0024] Select at least one target pixel whose value reaches the set first pixel value threshold from the at least one reference pixel.
[0025] As a possible implementation, when screening at least one reference pixel that meets the set local value condition from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels, the pixel screening unit is specifically configured to:
[0026] For a set of candidate pixels, the set of candidate pixels being the column of candidate pixels or the row of candidate pixels, perform the following operations:
[0027] Based on the value change situation of the set of candidate pixels, divide the set of candidate pixels to obtain at least one candidate pixel set;
[0028] According to the pixel value magnitude, screen out candidate pixels whose values reach the set second pixel value threshold from at least one candidate pixel set respectively as at least one reference pixel that meets the set local value condition, where the second pixel value threshold is less than the first pixel value threshold.
[0029] As a possible implementation, when performing feature extraction on the image to be detected to obtain a feature map, the feature extraction unit is specifically configured to:
[0030] Use a feature extraction model to perform feature extraction on the image to be detected to obtain a feature map; the resolution of the feature map conforms to the set resolution range.
[0031] As a possible implementation, when performing feature extraction on the image to be detected to obtain a feature map, the feature extraction unit is specifically configured to:
[0032] Using a feature extraction model, perform feature extraction on the detected image to obtain a feature map; the feature extraction model is iteratively trained using a sample image set, and each sample image in the sample image set is associated with a target label, and each target label is used to indicate the main body area of the picture in the associated sample image.
[0033] As a possible implementation manner, the feature extraction unit is further configured to obtain a target label associated with a sample image in the following manner:
[0034] Based on two true width edges of the main body area of the picture in the one sample image, obtain a corresponding column of sample pixels; among the column of sample pixels, the sample pixels with a value of the first pixel value are associated with the true width edge.
[0035] Based on two true height edges of the main body area of the picture in the one sample image, obtain a corresponding row of sample pixels; among the row of sample pixels, the sample pixels with a value of the first pixel value are associated with the true height edge.
[0036] Based on the column of sample pixels and the row of sample pixels, obtain the target label associated with the one sample image.
[0037] As a possible implementation manner, when obtaining a corresponding column of sample pixels based on two true width edges of the main body area of the picture in the one sample image, the feature extraction unit specifically is configured to:
[0038] Based on the height of the feature map output by the feature extraction network and in combination with a set initial pixel value, construct a column of initial sample pixels.
[0039] When there is at least one initial sample pixel located at the two true width edges in the column of initial sample pixels, use the first pixel value to assign values to the at least one initial sample pixel.
[0040] Based on the column of initial sample pixels after assignment, obtain a corresponding column of sample pixels.
[0041] As a possible implementation manner, after using the first pixel value to assign values to the at least one initial sample pixel and before obtaining a corresponding column of sample pixels based on the column of initial sample pixels after assignment, the feature extraction unit is further configured to:
[0042] Use a second pixel value to assign values to the other initial sample pixels in the column of initial sample pixels except the at least one initial sample pixel; or,
[0043] Use the pixel values between the first pixel value and the second pixel value to assign values to the other initial sample pixels in the column of initial sample pixels except for the at least one initial sample pixel.
[0044] As a possible implementation, when obtaining the main area of the image to be detected based on the candidate edges respectively associated with the at least one target pixel, the area determination unit is specifically configured to:
[0045] If the number of target pixels is four, obtain the main area of the image to be detected based on the candidate edges respectively associated with the four target pixels;
[0046] If the number of target pixels is less than four, obtain the main area of the image to be detected by combining the candidate edges respectively associated with the at least one target pixel and the edges of the image to be detected.
[0047] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above method.
[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method in any of the above aspects.
[0049] In a fifth aspect, an embodiment of the present application provides a computer program product. The program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the method in any of the above aspects.
[0050] In the embodiment of the present application, feature extraction is performed on the image to be detected to obtain a feature map. The main area of the image to be detected is a rectangle; dimensionality reduction is respectively performed on the feature map in the width direction and the height direction to obtain a column of candidate pixels having the same height as the feature map, and a row of candidate pixels having the same width as the feature map. Then, based on the pixel values, at least one target pixel is screened out from the column of candidate pixels and the row of candidate pixels, and then the main area of the image to be detected is obtained based on the candidate edges respectively associated with the at least one target pixel, where each candidate edge is: a row of pixels or a column of pixels in the image to be detected that passes through the associated candidate pixel.
[0051] In the embodiments of the present application, the recognition of the main area of the screen is redefined as the detection of key points in the width and height directions of the image. Compared with directly detecting straight line segments or rectangular areas, the detection process is simpler, thus improving the detection efficiency. In addition, due to the characteristics of the image being horizontal and vertical, at most four target pixels can be used to determine the main area of the image to be detected. The method for detecting the main area of the screen based on pixel points improves the detection efficiency of the main area of the screen while ensuring the accuracy of the detection of the main area of the screen.
[0052] Other features and advantages of the present application will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0053] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0054] Figure 1 It is a schematic diagram of an application scenario provided in the embodiments of the present application;
[0055] Figure 2 It is a schematic flowchart of a method for detecting the main area of the screen provided in the embodiments of the present application;
[0056] Figure 3 It is a schematic diagram of the image 1 to be detected provided in the embodiments of the present application;
[0057] Figure 4 It is a schematic diagram of the image 2 to be detected provided in the embodiments of the present application;
[0058] Figure 5 It is a schematic diagram of the image 3 to be detected provided in the embodiments of the present application;
[0059] Figure 6 It is a schematic diagram of a feature extraction process provided in the embodiments of the present application;
[0060] Figure 7 It is a schematic diagram of a column vector provided in the embodiments of the present application;
[0061] Figure 8 It is a schematic diagram of a row vector provided in the embodiments of the present application;
[0062] Figure 9 It is a schematic diagram of the first assignment method provided in the embodiments of the present application;
[0063] Figure 10 Schematic diagram of the second assignment method provided in the embodiment of the present application;
[0064] Figure 11 Schematic diagram of the first canvas style provided in the embodiment of the present application;
[0065] Figure 12 Schematic diagram of the second canvas style provided in the embodiment of the present application;
[0066] Figure 13A Schematic diagram of the third canvas style provided in the embodiment of the present application;
[0067] Figure 13B Schematic diagram of the fourth canvas style provided in the embodiment of the present application;
[0068] Figure 14 Schematic diagram of the fifth canvas style provided in the embodiment of the present application;
[0069] Figure 15 Schematic diagram of the sixth canvas style provided in the embodiment of the present application;
[0070] Figure 16 Schematic diagram of the seventh canvas style provided in the embodiment of the present application;
[0071] Figure 17 Schematic diagram of the HRNet architecture provided in the embodiment of the present application;
[0072] Figure 18 Schematic flow chart of a model training method provided in the embodiment of the present application;
[0073] Figure 19 Schematic diagram of a smoothing vector provided in the embodiment of the present application;
[0074] Figure 20 Schematic logical diagram of a target pixel determination process provided in the embodiment of the present application;
[0075] Figure 21 Schematic logical diagram of a main subject detection process in the picture provided in the embodiment of the present application;
[0076] Figure 22 Schematic structural diagram of a main subject detection device provided in the embodiment of the present application;
[0077] Figure 23 Schematic structural diagram of an electronic device provided in the embodiment of the present application. Detailed implementation manners
[0078] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of this application. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the technical solutions of this application.
[0079] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0080] It can be understood that in the specific implementation of this application, when it comes to data such as the image to be detected, when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0081] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0082] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0083] Computer vision technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition, monitoring, and measurement, and further performs graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pretrained models in the field of vision such as Swin Transformer, Vision Transformer (ViT), Vision MoE (V-MoE) based on the Mixture of Experts (MoE), and Masked Autoencoders (MAE) can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, Optical Character Recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0084] The key technologies of speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods. Large model technology has brought changes to the development of speech technology. Pretrained models that follow the Transformer architecture such as WavLM and UniSpeech have strong generalization and versatility and can excellently complete speech processing tasks in various directions.
[0085] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves disciplines such as computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, is developed from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0086] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.
[0087] Autopilot technology refers to the vehicle's ability to drive itself without driver operation. It usually includes technologies such as high-precision maps, environmental perception, computer vision, behavior decision-making, path planning, and motion control. Autopilot includes multiple development paths such as single-vehicle intelligence, vehicle-road collaboration, and networked cloud control. Autopilot technology has broad application prospects. Currently, in addition to the fields of logistics, public transportation, taxis, and intelligent transportation, it will be further developed in the future.
[0088] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autopilot, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0089] The solution provided by the embodiments of this application involves the application of machine learning technology, specifically: using a feature extraction model (or feature extraction network) to extract the features of the image to be detected, obtain the corresponding feature map, and then locate the main area of the image to be detected based on the dimensionality reduction result of the feature map.
[0090] In the related art, the recognition of the main area of the image usually adopts two methods: contour detection and edge detection. In contour detection, the square main body in the content is recognized as the main area of the image, while in edge detection, the straight lines of the border are recognized to determine the main area of the image.
[0091] However, whether it is edge detection or contour detection, they both detect the significant features in the content (such as the straight lines of the border, the square contour of the main area of the image), which requires a large amount of logical design, has a high detection difficulty, and affects the detection efficiency. Especially when there are strongly contrasting patterns at the edge of the main area of the image, it is difficult to accurately detect the straight lines. At the same time, due to the strong diversity of the border styles, it is difficult to identify the main area of the image, and there may be cases of misidentification or inaccurate identification. Moreover, the main area of the image is often a large area, and it is difficult to accurately locate the detection and detect the border in the fine details.
[0092] In the embodiments of this application, the features of the image to be detected are extracted to obtain a feature map. The main area of the image to be detected is rectangular; the feature map is dimensionally reduced in the width direction and the height direction respectively to obtain a column of candidate pixels with the same height as the feature map, and a row of candidate pixels with the same width as the feature map. Then, based on the pixel values, at least one target pixel is selected from the column of candidate pixels and the row of candidate pixels. Then, based on the candidate edges associated with each of the at least one target pixel, the main area of the image to be detected is obtained, where each candidate edge is: a row of pixels or a column of pixels in the image to be detected that passes through the associated candidate pixel.
[0093] Redefining the recognition of the main area of the image as the detection of key points in the width direction and the height direction of the image, compared with directly detecting straight line segments or rectangular areas, the detection process is simpler, thus improving the detection efficiency. In addition, due to the characteristics of the image being horizontal and vertical, therefore, the main area of the image to be detected can be determined by at most four target pixels. The method of detecting the main area of the image based on pixel points improves the detection efficiency of the main area of the image while ensuring the accuracy and precision of the detection of the main area of the image.
[0094] Refer to Figure 1As shown in the figure, it is a schematic diagram of an application scenario provided in an embodiment of the present application. This application scenario includes a terminal device 110 and a server 120. The number of terminal devices 110 can be one or more. The number of servers 120 can also be one or more. The present application does not make specific limitations on the number of terminal devices 110 and servers 120.
[0095] In an embodiment of the present application, the terminal device 110 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. The terminal device 110 supports a client for content recommendation, and the server 120 is the background server corresponding to the client.
[0096] The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0097] The terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.
[0098] In some embodiments, the server 120 extracts features from the image to be detected to obtain a feature map. The main body area of the image to be detected is rectangular; the feature map is dimensionally reduced in the width direction to obtain a column of candidate pixels having the same height as the feature map. Each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected that passes through the associated candidate pixel; the feature map is dimensionally reduced in the height direction to obtain a row of candidate pixels having the same width as the feature map. Each candidate pixel in the row of candidate pixels is associated with a candidate high edge, and each candidate high edge is: a column of pixels in the image to be detected that passes through the associated candidate pixel; based on the respective values of the column of candidate pixels and the row of candidate pixels, at least one target pixel is selected from the column of candidate pixels and the row of candidate pixels; based on the candidate edges associated with the at least one target pixel, the main body area of the image to be detected is obtained, where each candidate edge is a candidate wide edge or a candidate high edge.
[0099] It should be noted that the main body detection method of the image mentioned in the embodiment of the present application can be executed by the server or the terminal device, or jointly executed by the server and the terminal device, and this is not limited. This article only takes the server as an example for illustration.
[0100] Refer to Figure 2 As shown, it is a schematic flowchart of a method for detecting the main subject of a picture provided in an embodiment of the present application. This method is applied to a terminal device or a server, and the specific process is as follows:
[0101] S201. Extract features from the image to be detected to obtain a feature map. The main subject area of the image to be detected is a rectangle.
[0102] In an embodiment of the present application, the image to be detected can be a certain frame in a video, the cover of a video, a picture, etc., but is not limited thereto.
[0103] The other parts of the image to be detected except the main subject area can be called the border area. The method for detecting the main subject of the image to be detected can also be understood as the method for detecting the border of the image to be detected. The border area can be a black border or an area containing information of types such as video, text, and image.
[0104] As a possible implementation, the main subject area can be located in the border area or on one side of the border area (such as the left side, the right side, the upper side, the lower side, etc.). In this article, the relative position relationship between the main subject area and the border area is not limited, but is not limited thereto.
[0105] Taking the image 1 to be detected as an example, refer to Figure 3 As shown, the main subject area of the image 1 to be detected is shown by the dashed box. There are border areas on both the upper side and the lower side of the main subject area. Among them, the main subject area shows a picture of a skirt rotating. The upper border shows the upper half of the picture, and the lower border shows the lower half of the picture.
[0106] Taking the image 2 to be detected as an example, refer to Figure 4 As shown, the main subject area of the image 2 to be detected is shown by the dashed box. There are border areas on both the upper side and the lower side of the main subject area. Among them, the main subject area shows a picture of making popcorn. The upper border shows the text description "The popcorn can't be popped out. Use a hammer to help. The result is too shocking", and the lower border shows other pictures (such as a house in the distance).
[0107] Taking the image 3 to be detected as an example, refer to Figure 5 As shown, the main subject area of the image 3 to be detected is shown by the dashed box. There are border areas on both the upper side and the lower side of the main subject area. Among them, the main subject area shows a picture of an exam. The upper border and the main subject area show the text description "Elect two elective courses and get a high-paying job", and the lower border is a black border.
[0108] Since the relative positional relationship between the main subject area and the border area in the image to be detected 1, the image to be detected 2, and the image to be detected 3 is the same, only the image to be detected 1 will be used as an example for illustration below.
[0109] Refer to Figure 6 As shown, feature extraction is performed on the image to be detected 1 to obtain the feature map of the image to be detected 1. The feature extraction is implemented using a feature extraction network, and the feature extraction network is of the HRNet structure. Cuboids of different sizes in the feature extraction network represent feature maps at different resolutions. In the HRNet structure, the feature map always maintains a high resolution and combines with low-resolution feature maps to generate the final feature map. The finally generated feature map maintains a high resolution. For the specific HRNet structure, please refer to Figure 17 . Subsequently, a row vector and a column vector can be extracted from the final feature map. The row vector contains a row of candidate pixels, and the column vector contains a column of candidate pixels. Exemplarily, the resolution of the feature map obtained by performing feature extraction on the image to be detected 1 is 512×512×3, where 3 represents the number of channels. In the feature extraction network, downsampling the 512×512×3 feature map can obtain a low-resolution feature map of 128×128×16.
[0110] S202. Dimension reduction is performed on the feature map in the width direction to obtain a column of candidate pixels having the same height as the feature map. Each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected that passes through the associated candidate pixel.
[0111] In the embodiments of the present application, a column of candidate pixels having the same height as the feature map can also be referred to as a column feature vector or a column vector. Each element in the column vector corresponds to a candidate pixel.
[0112] Suppose the width of the feature map is fw and the height of the feature map is fh. When performing dimension reduction on the feature map in the width direction, as a possible case, if the feature map is a single-channel feature map, that is, the feature dimension of the feature map is fw×fh, then the feature map can be directly dimension-reduced in the width direction to obtain a column vector with a feature dimension of 1×fh; as another possible case, if the feature map is a multi-channel feature map, that is, the feature dimension of the feature map is fw×fh×c, where c represents the number of channels, then on each channel, the feature map is dimension-reduced in the width direction to obtain a corresponding sub-vector of 1×fh, and then, the 1×fh sub-vectors corresponding to multiple channels are vector-fused to obtain a column vector with a feature dimension of 1×fh, where the vector fusion can be implemented by vector addition, but is not limited thereto.
[0113] For example, refer to Figure 7As shown in the figure, assume that the height of the feature map is 12, and the feature map is dimensionally reduced in the width direction to obtain a column vector of 1×12. The column vector contains 12 candidate pixels, including pixel L1, pixel L2, …, pixel L12. Among them, pixel L1 is associated with candidate wide edge 1, and candidate wide edge 1 is a row of pixels (i.e., the first row of pixels) passing through pixel L1 in the image to be detected. Pixel L2 is associated with candidate wide edge 2, and candidate wide edge 2 is a row of pixels (i.e., the second row of pixels) passing through pixel L2 in the image to be detected. Similarly, pixel L12 is associated with candidate wide edge 12, and candidate wide edge 12 is a row of pixels (i.e., the twelfth row of pixels) passing through pixel L12 in the image to be detected.
[0114] S203. Dimensionally reduce the feature map in the height direction to obtain a row of candidate pixels with the same width as the feature map. Each candidate pixel in the row of candidate pixels is associated with a candidate high edge. Each candidate high edge is: a column of pixels passing through the associated candidate pixel in the image to be detected. It should be noted that in the embodiments of the present application, the execution order between S203 and S202 is not limited. S203 can be executed first, or S202 can be executed first.
[0115] In the embodiments of the present application, a row of candidate pixels with the same width as the feature map can also be referred to as a row feature vector or a row vector. Each element in the row vector corresponds to a candidate pixel.
[0116] Assume that the width of the feature map is fw and the height of the feature map is fh. When dimensionally reducing the feature map in the height direction, as a possible case, if the feature map is a single-channel feature map, that is, the feature dimension of the feature map is fw×fh, then the feature map can be directly dimensionally reduced in the height direction to obtain a row vector with a feature dimension of fw×1. As another possible case, if the feature map is a multi-channel feature map, that is, the feature dimension of the feature map is fw×fh×c, where c represents the number of channels, then on each channel, the feature map is dimensionally reduced in the height direction to obtain a corresponding sub-vector with a feature dimension of fw×1. Then, vector fusion is performed on the sub-vectors with a feature dimension of fw×1 corresponding to multiple channels to obtain a row vector with a feature dimension of fw×1. Among them, vector fusion can be achieved by vector addition, but is not limited thereto.
[0117] For example, refer to Figure 8As shown in the figure, assume that the width of the feature map is 11. The feature map is dimensionally reduced in the height direction to obtain a 10×1 row vector. The row vector contains 11 candidate pixels, which include pixel H1, pixel H2, …, pixel H11. Among them, pixel H1 is associated with candidate high edge 1, and candidate high edge 1 is a column of pixels (i.e., the first column of pixels) passing through pixel H1 in the image to be detected. Pixel H2 is associated with candidate high edge 2, and candidate high edge 2 is a column of pixels (i.e., the second column of pixels) passing through pixel H2 in the image to be detected. Similarly, pixel H11 is associated with candidate high edge 11, and candidate high edge 11 is a column of pixels (i.e., the 11th column of pixels) passing through pixel H11 in the image to be detected.
[0118] S204. Based on the respective values of a column of candidate pixels and a row of candidate pixels, at least one target pixel is selected from the column of candidate pixels and the row of candidate pixels.
[0119] In the embodiments of the present application, different target pixel selection methods can be adopted according to different assignment methods of pixel values. Among them, the pixel value can adopt the first assignment method or the second assignment method.
[0120] The first assignment method uses the first pixel value to represent that the candidate edge associated with the candidate pixel belongs to the edge of the main body area of the picture, and uses the second pixel value to represent that the candidate edge associated with the candidate pixel does not belong to the edge of the main body area of the picture. Among them, the candidate edge is the candidate wide edge or the candidate high edge. Specifically, when the candidate pixel is a candidate pixel in a row of candidate pixels, then the candidate edge is the candidate high edge. When the candidate pixel is a candidate pixel in a column of candidate pixels, then the candidate edge is the candidate wide edge.
[0121] That is to say, the first assignment method is binary assignment. Exemplarily, the first pixel value is 1 and the second pixel value is 0. When the pixel value is 1, it represents that the candidate edge (candidate wide edge or candidate high edge) associated with the pixel belongs to the edge of the main body area of the picture. When the pixel value is 0, it represents that the candidate edge (candidate wide edge or candidate high edge) associated with the pixel belongs to the edge of the main body area of the picture.
[0122] Specifically, in the case where the pixel value adopts the first assignment method, based on the respective values of a column of candidate pixels and a row of candidate pixels, at least one candidate pixel with a value of the first pixel value is selected from the column of candidate pixels and the row of candidate pixels as the target pixel.
[0123] For example, refer to Figure 9As shown, the values of a row of candidate pixels are all 0. Among a column of candidate pixels, the pixel values of pixel L2 and pixel L11 are 1, and the values of the remaining pixels are all 0. Based on the values of a column of candidate pixels and a row of candidate pixels respectively, pixel L2 and pixel L11 with a value of 1 are selected from the column of candidate pixels and the row of candidate pixels as target pixels.
[0124] The second assignment method uses a set of pixel values to represent the probability that the candidate wide edge associated with the candidate pixel belongs to the edge of the main area of the picture. Among them, the pixel range of a set of pixel values is from the first pixel value to the second pixel value. The pixel range can be an open interval, a closed interval, or a semi-open and semi-closed interval, and no limitation is made in this regard, so it will not be elaborated here. The continuous pixel values can be arithmetic or non-arithmetic, and no limitation is made in this regard.
[0125] Exemplarily, refer to Figure 10 As shown, the pixel range of a set of continuous pixel values is between 0 and 1. The values of a row of candidate pixels are all 0. Among a column of candidate pixels, the values of pixel L1, pixel L2,..., pixel L12 are 0.8, 1, 0.8, 0.4, 0, 0, 0, 0, 0.4, 0.8, 1, 0.8 in sequence.
[0126] When the pixel value adopts the second assignment method, the local optimal solution can be first screened out from each candidate pixel (including a column of candidate pixels and a row of candidate pixels), and then the target pixel can be determined from the local optimal solution. Specifically, when executing S204, the following steps can be adopted but are not limited to:
[0127] Step 1: Based on the values of a column of candidate pixels and a row of candidate pixels respectively, at least one reference pixel that meets the set local value condition is screened out from the column of candidate pixels and the row of candidate pixels.
[0128] Among them, the local value condition can be used to screen out the local optimal solution, but is not limited to this.
[0129] As a possible implementation manner, the screening of the reference pixel is implemented by the following methods but is not limited to:
[0130] For a set of candidate pixels, where the set of candidate pixels is a column of candidate pixels or a row of candidate pixels, the following operations are performed:
[0131] Based on the value change situation of a set of candidate pixels, the set of candidate pixels is divided to obtain at least one candidate pixel set;
[0132] According to the size of pixel values, candidate pixels whose values reach a set second pixel value threshold are screened out from at least one candidate pixel set as at least one reference pixel that meets the set local value condition, wherein the second pixel value threshold is less than the first pixel value threshold.
[0133] Take only one column of candidate pixels as an example, see Figure 20 As shown, based on the value changes of a column of candidate pixels, a group of candidate pixels is divided to obtain three candidate pixel sets, which are L1 to L6, L7 to L9, and L9 to L12. Assuming that the second pixel value threshold is 0.5, according to the pixel value size, the candidate pixels with a value of 0.5 are screened out from the three candidate pixel sets respectively as reference pixels that meet the set local value conditions. The reference pixels can also be called local optimal solutions.
[0134] Step 2: Select at least one target pixel whose value reaches a set first pixel value threshold from at least one reference pixel.
[0135] For example, see Figure 20 As shown, it is assumed that the first pixel value threshold is 0.9, the reference pixels include pixel L3, pixel L8 and pixel L10, and the target pixels with a value of 0.9 are selected from pixel L3, pixel L8 and pixel L10, and the target pixels include pixel L3 and pixel L10.
[0136] S205 . Obtain a main area of the image to be detected based on candidate edges respectively associated with at least one target pixel, wherein each candidate edge is a candidate wide edge or a candidate high edge.
[0137] In actual application, the image to be detected and the main area of the picture are both rectangular, and the edge of the image to be detected is parallel to the edge of the main area of the picture. Therefore, the main area of the picture can be composed of two edges in the width direction and two edges in the length direction.
[0138] See also Figure 11 , when the main area of the picture is surrounded by a border, there are two target pixels in a column of candidate pixels and two target pixels in a row of candidate pixels. Obviously, there are four target pixels for a column of candidate pixels and a row of candidate pixels.
[0139] See also Figure 12 , Figure 13A , Figure 13B , Figure 14 and Figure 15, in the case of one side of the main area of the picture, there is one or two target pixels in a column of candidate pixels, or there is one or two target pixels in a row of candidate pixels. Obviously, for a column of candidate pixels and a row of candidate pixels, there may be one target pixel, two target pixels or three target pixels.
[0140] See Figure 16 , in the case that there is no border in the image to be detected, there are no target pixels in either a column of candidate pixels or a row of candidate pixels.
[0141] Obviously, in the embodiments of the present application, the main area of the image to be detected can be determined by at most four target pixels. The method for detecting the main area of the picture based on pixel points is simpler than the method for detecting the main area of the picture based on edges or shapes. While ensuring the accuracy of the main area detection of the picture, the detection efficiency of the main area of the picture is improved.
[0142] Next, the process of determining the main area of the picture will be described for different distributions of the main area of the picture.
[0143] Specifically, when executing S205, there are but not limited to the following situations:
[0144] Situation A: If the number of target pixels is four, the main area of the image to be detected is obtained based on the candidate edges respectively associated with the four target pixels. Exemplarily, the rectangular area enclosed by the four candidate edges can be used as the main area of the image to be detected.
[0145] Refer to Figure 11 As shown, the target pixels include four pixels A, B, C, and D. Among them, A is associated with the candidate high edge a, B is associated with the candidate high edge b, C is associated with the candidate wide edge c, and D is associated with the candidate wide edge d. The rectangular area enclosed by the candidate high edge a, the candidate high edge b, the candidate wide edge c, and the candidate wide edge d is used as the main area of the image to be detected.
[0146] Situation B: If the number of target pixels is less than four, the main area of the image to be detected is obtained based on the candidate edges respectively associated with at least one target pixel, in combination with the edges of the image to be detected.
[0147] Situation B: The number of target pixels is three, that is, there are candidate edges on three sides of the main area of the picture and no candidate edge on one side. At this time, the main area of the image to be detected is obtained based on the candidate edges respectively associated with the three target pixels, in combination with the edges of the image to be detected. Exemplarily, the rectangular area enclosed by the three candidate edges and the edges of the image to be detected is used as the main area of the image to be detected.
[0148] Refer to Figure 12As shown, the target pixels include three pixels A, B, and C. Among them, A is associated with the candidate high edge a, B is associated with the candidate high edge b, and C is associated with the candidate wide edge c. The rectangular area enclosed by the candidate high edge a, the candidate high edge b, the candidate wide edge c, and the edge of the image to be detected is taken as the main area of the image to be detected.
[0149] Case C: The number of target pixels is two, that is, there are candidate edges on both sides of the main area of the image, and there are no candidate edges on both sides. At this time, based on the candidate edges associated with the two target pixels respectively, combined with the edge of the image to be detected, the main area of the image to be detected is obtained. Exemplarily, the rectangular area enclosed by the three candidate edges and the edge of the image to be detected is taken as the main area of the image to be detected.
[0150] Case C1: The two candidate edges are candidate edges in the same direction. For example, both of the two candidate edges are candidate wide edges or candidate high edges.
[0151] Specifically, the rectangular area enclosed by the candidate wide edges associated with the two target pixels respectively and the two edges of the image to be detected in the height direction is taken as the main area of the image to be detected; or, the rectangular area enclosed by the candidate high edges associated with the two target pixels respectively and the two edges of the image to be detected in the width direction is taken as the main area of the image to be detected.
[0152] Refer to Figure 13A As shown, the target pixels include two pixels C and D. Among them, C is associated with the candidate wide edge c, and D is associated with the candidate wide edge d. The rectangular area enclosed by the candidate wide edge c, the candidate wide edge d, and the two edges of the image to be detected in the width direction is taken as the main area of the image to be detected.
[0153] Refer to Figure 13B As shown, the target pixels include two pixels A and B. Among them, A is associated with the candidate high edge a, and B is associated with the candidate high edge b. The rectangular area enclosed by the candidate high edge a, the candidate high edge b, and the two edges of the image to be detected in the width direction is taken as the main area of the image to be detected.
[0154] Case C2: The two candidate edges are candidate edges in different directions, that is, the two candidate edges include a candidate wide edge and a candidate high edge.
[0155] As a possible implementation, considering that the area of the main area of the image in the image to be detected accounts for a relatively large proportion, therefore, among the rectangular areas enclosed by the candidate wide edges associated with the two target pixels respectively and the edge of the image to be detected, the rectangular area with the largest area is taken as the main area of the image to be detected.
[0156] Refer to Figure 14As shown, the target pixel includes two pixels, A and C. Among them, A is associated with a candidate high edge a, and C is associated with a candidate wide edge c. The candidate high edge a, the candidate wide edge c, and the edges of the image to be detected can form four rectangular regions. The rectangular region with the largest area among the four rectangular regions is used as the main body region of the image to be detected.
[0157] Case D: The number of target pixels is one, that is, there is a candidate edge on one side of the main body region of the image, and there are no candidate edges on the other three sides.
[0158] As a possible implementation, considering that the main body region of the image occupies a relatively large area in the image to be detected, therefore, among the rectangular regions formed by the candidate wide edges associated with each target pixel and the edges of the image to be detected, the rectangular region with the largest area is used as the main body region of the image to be detected.
[0159] Refer to Figure 15 As shown, the target pixel includes A. Among them, A is associated with a candidate high edge a. The candidate high edge a and the edges of the image to be detected can form two rectangular regions. The rectangular region with the largest area among the two rectangular regions is used as the main body region of the image to be detected.
[0160] Case E: Refer to Figure 16 , the number of target pixels is zero, that is, there are no candidate edges on all four sides of the main body region of the image. At this time, the entire image to be detected is used as the main body region.
[0161] In some embodiments, when performing feature extraction on the image to be detected, a feature extraction model can be used to perform feature extraction on the image to be detected to obtain a feature map; the resolution of the feature map conforms to a set resolution range. That is to say, the feature extraction model needs to output a feature map that conforms to the set resolution range, so as to accurately locate the target pixel, and then ensure the detection accuracy of the main body region of the image and improve the detection accuracy. Among them, the set resolution range can be set according to the actual business, as long as it meets the high-resolution sampling requirements of the business.
[0162] The backbone network of the feature extraction model can adopt a model structure for image segmentation. As an example, the feature extraction model can be implemented using a U-Net structure. Generally speaking, models for image segmentation often adopt a U-Net structure, that is, sampling the feature map from high resolution to low resolution, and then restoring the high-resolution representation from the low-resolution representation.
[0163] As another example, the feature extraction model can also be implemented using the HRNet structure. Generally, in the U-Net structure, the input image is first downsampled to reduce the image resolution, and then upsampled to increase the resolution to obtain the final feature extraction result. The method of retaining the stronger feature maps during downsampling and then restoring the image resolution during upsampling will lose a certain degree of spatial information, and there will be a certain quantization error when the feature extraction result is used for subsequent detection. HRNet adopts a design method of parallel network structures, always keeping one branch for high-resolution feature extraction processing, thereby protecting rich spatial information and reducing the generation of quantization errors.
[0164] Refer to Figure 17 As shown, it is a partial schematic diagram of an HRNet structure provided in an embodiment of the present application. Figure 17 In it, the horizontal direction represents the depth of the network, and the vertical direction represents the scale of the network. The value of the depth represents the network structure of the feature extraction part of HRnet. The network feature extraction part can be divided into multiple stages (for example, depth 1 to 2 is one stage, and depth 3 to 7 is one stage).
[0165] In actual implementation, before the start of each stage, a feature map with a smaller resolution is added, and feature maps of different scales are obtained through interpolation upsampling and convolutional downsampling respectively, and the feature maps of the same scale are fused to ensure that the initial feature map combines the features of the feature maps of different scales in the previous stage. During each stage, a residual neural network (ResNet) can be used to perform deep learning on the feature maps of each scale respectively. In the last stage, the final output feature map can be obtained based on the feature maps of 3 different scales.
[0166] Such as Figure 17 As shown, in the HRNet structure, a high-resolution feature map is always maintained, and feature maps of different scales are combined, and finally a high-resolution feature map is output for downstream tasks.
[0167] In some embodiments, feature extraction is performed on the image to be detected to obtain a feature map, including:
[0168] Using the feature extraction model, perform feature extraction on the image to be detected to obtain a feature map; the feature extraction model is obtained by iterative training using a sample image set, and each sample image in the sample image set is associated with a target label, and each target label is used to indicate the main area of the picture in the associated sample image.
[0169] Refer to Figure 18As shown in the figure, it is a schematic flowchart of a model training method provided in an embodiment of the present application, and this method is applied to a terminal device or a server. During iterative training, all sample images are divided into specified batches, and training is performed based on the sample images in each batch. Since the steps executed during training for each batch in each iteration process are similar, here, the training for one batch is taken as an example for illustration.
[0170] S1801. Use a feature extraction model to extract features from each sample image respectively to obtain corresponding sample feature maps. For details, refer to S201.
[0171] S1802. In the width direction, reduce the dimension of each obtained sample feature map respectively to obtain corresponding column vectors, and a column vector contains a column of candidate pixels. For details, refer to S202.
[0172] S1803. In the height direction, reduce the dimension of each obtained feature map respectively to obtain corresponding row vectors, and a row vector contains a row of candidate pixels. For details, refer to S203.
[0173] S1804. Based on the column vectors and row vectors corresponding to each sample feature map obtained, and in combination with the target labels associated with each sample feature map, obtain a model loss. Among them, the model loss can adopt the MSE loss, but is not limited thereto.
[0174] S1805. Determine whether the model convergence condition is satisfied. If so, execute S1807; otherwise, execute S1806 and enter the training of the next batch.
[0175] In the embodiment of the present application, the model convergence condition may include at least one of the following conditions:
[0176] (1) The model loss is not greater than a preset loss value threshold.
[0177] (2) The number of iterations reaches a preset upper limit value.
[0178] S1806. Adjust the model parameters based on the model loss.
[0179] S1807. Output the feature extraction model.
[0180] In some implementation manners, the target label associated with a sample image can be obtained through the following method:
[0181] Based on two true wide sides of the main body area of the picture in a sample image, obtain a corresponding column of sample pixels; among a column of sample pixels, the sample pixels with a value of the first pixel value are associated with the true wide side;
[0182] Based on two true high edges in the main area of the picture in a sample image, obtain a corresponding row of sample pixels; among the row of sample pixels, the sample pixels with the first pixel value are associated with the true high edges.
[0183] Based on a column of sample pixels and a row of sample pixels, obtain a target label associated with a sample image.
[0184] That is to say, in the embodiments of the present application, the target label is composed of a row vector (including a row of sample pixels) and a column vector (including a column of sample pixels). Among them, the row vector is used to indicate the true high edge of the sample image, and the column vector is used to indicate the true wide edge of the sample image. In this way, according to the target label, the model is trained so that the model can output a more accurate prediction result under the guidance of the target label, improving the accuracy of the model prediction result.
[0185] Among them, as a possible implementation manner, a column of sample pixels can be obtained in the following but not limited ways:
[0186] Based on the height of the feature map output by the feature extraction network and in combination with the set initial pixel value, construct a column of initial sample pixels;
[0187] When there is at least one initial sample pixel located at two true wide edges among a column of initial sample pixels, use the first pixel value to assign values to the at least one initial sample pixel;
[0188] Based on the column of initial sample pixels after value assignment, obtain the corresponding column of sample pixels.
[0189] That is to say, for a row of sample pixels, use the first pixel value to assign values to the initial sample pixels associated with the true wide edge among a column of initial sample pixels.
[0190] According to the different value assignment methods of pixel values, for the other initial sample pixels except the initial sample pixels associated with the true wide edge among a row of sample pixels, the following methods can be used for value assignment:
[0191] If the first value assignment method is used for pixel value assignment, in the first value assignment method, use the first pixel value to represent that the candidate edge associated with the candidate pixel belongs to the edge of the main area of the picture, and use the second pixel value to represent that the candidate edge associated with the candidate pixel does not belong to the edge of the main area of the picture. Therefore, use the second pixel value to assign values to the other initial sample pixels except the at least one initial sample pixel among a column of initial sample pixels.
[0192] For example, assume that the pixels L2 and L11 in the sample pixels are associated with a true wide edge. Then, in a column of initial sample pixels, the pixel values of pixels L2 and L11 are 1, and the pixel values of the remaining pixels are all 0. In the embodiments of the present application, the vector obtained by the first assignment method can also be called a hard vector. There are only two values in the hard vector, and its processing is relatively simple, which can improve the data processing efficiency to a certain extent.
[0193] If the second assignment method is used for pixel assignment, a set of continuous pixel values is used to represent the probability that the candidate wide edge associated with the candidate pixel belongs to the edge of the main body area of the picture. Therefore, according to the distance between other initial sample pixels and at least one initial sample pixel among the initial sample pixels in a column, and in combination with the set pixel range, other initial sample pixels are assigned values in sequence.
[0194] For example, refer to Figure 19 As shown, assume that the pixels L2 and L11 in the sample pixels are associated with a true wide edge. The pixel range of a set of continuous pixel values is between 0 and 1. In a column of initial sample pixels, the pixel values of pixels L2 and L11 are 1. According to the distance from L2 and L11, the respective pixel values of pixels L1, L2, …, L12 are 0.8, 1, 0.8, 0.4, 0, 0, 0, 0, 0.4, 0.8, 1, 0.8. In the embodiments of the present application, the hard vector can be converted into a soft vector (or called a soft vector) through, but not limited to, Gaussian smoothing, so as to solve the problems of unstable hard target convergence and large error.
[0195] According to the different assignment methods of pixel values, for the assignment of other initial sample pixels except the initial sample pixels associated with the true wide edge in a row of sample pixels, the following methods can be used:
[0196] If the first assignment method is used for pixel assignment, in the first assignment method, the first pixel value is used to represent that the candidate edge associated with the candidate pixel belongs to the edge of the main body area of the picture, and the second pixel value is used to represent that the candidate edge associated with the candidate pixel does not belong to the edge of the main body area of the picture. Therefore, the second pixel value is used to assign values to other initial sample pixels except at least one initial sample pixel in a row of initial sample pixels.
[0197] For example, assume that the pixels H5 and H10 in the sample pixels are associated with a true high edge. Then, in a row of initial sample pixels, the pixel values of pixels H5 and H10 are 1, and the pixel values of the remaining pixels are all 0.
[0198] If the second assignment method is used for pixel assignment, a set of pixel values is used to represent the probability that the candidate wide edge associated with the candidate pixel belongs to the edge of the main area of the picture. Therefore, according to the distances between other initial sample pixels and at least one initial sample pixel among a row of initial sample pixels, combined with the set pixel range, the other initial sample pixels are assigned values in turn.
[0199] Since the assignment process for a row of sample pixels is similar to that for a column of sample pixels, it will not be elaborated here.
[0200] It should be noted that the above description is only based on the case where the resolution of the output feature map is the same as that of the image to be detected. In actual application, the resolution of the output feature map may also be different from that of the image to be detected.
[0201] Refer to Figure 21 As shown, it is a logical schematic diagram of a main area detection process provided in an embodiment of the present application.
[0202] The image to be detected includes the image to be detected a, the image to be detected b, and the image to be detected c. Among them, the main area of the image to be detected a is a certain frame of a game event video, and other videos are displayed in the border area; the main area of the image to be detected b is a certain frame of a certain life sharing video, and the image to be detected b has an extremely narrow border; the main area of the image to be detected c is a certain frame of a certain news video, and there is text in the border of the image to be detected c.
[0203] For the image to be detected a, the image to be detected b, and the image to be detected c, feature extraction is performed respectively to obtain corresponding feature maps, and then dimensionality reduction is performed on the feature maps to obtain row vectors and column vectors. Then, target pixels are determined according to the row vectors and column vectors, and then the main area of the picture is determined according to the target pixels. For the specific main area detection, refer to Figure 2 , which will not be elaborated here. Figure 21 The detected main area of the picture is indicated by a thick dashed line in
[0204] In addition, comparative experiments were conducted on the YOLOv5 model (directly outputting the main body area of the picture using the YOLOv5 model), HRNet + edge regression (i.e., in the embodiments of the present application, using the HRNet structure for feature extraction, and extracting row vectors and column vectors from the generated feature map to further determine the main body area of the picture), and HRNet + edge regression + target smoothing (on the basis of HRNet + edge regression, using the second assignment method for pixel assignment). The accuracies (precision) of the three were 81.4%, 92.3%, and 95.9% respectively, and the recall rates were 84.3%, 93.5%, and 94.2% respectively. Obviously, the effect of directly detecting the main body area of the picture through the detection model (YOLOv5 model) is not very good. The feature extraction model based on edge regression has made great progress compared to the detection model, and target smoothing has a relatively obvious improvement on the effect of the model.
[0205] Under the premise of the error between two pixels, the accuracy of HRNet + edge regression + target smoothing is 95.89%, and the recall rate is 94.22%.
[0206] In actual business, about 30% of the small video covers contain canvases and borders. Apply the feature extraction model based on edge regression and target smoothing to the preprocessing of video classification, and use the main body part of the preprocessed picture as the input of the video classification model. The accuracy and recall rate of the video classification model are improved from 92% and 87% to 93% and 90%. For Figure 21 example, Figure 21 in, in scenarios such as extremely narrow borders and complex backgrounds, the main body area of the picture can still be accurately located.
[0207] Based on the same inventive concept, the embodiments of the present application provide a main body detection device for pictures. As Figure 22 shown, it is a schematic structural diagram of the main body detection device 2200, which may include:
[0208] A feature extraction unit 2201, configured to perform feature extraction on the image to be detected to obtain a feature map, and the main body area of the image to be detected is rectangular;
[0209] A column vector extraction unit 2202, configured to perform dimensionality reduction on the feature map in the width direction to obtain a column of candidate pixels having the same height as the feature map, and each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected passing through the associated candidate pixel;
[0210] A row vector extraction unit 2203 is configured to reduce the dimension of the feature map in the height direction to obtain a row of candidate pixels having the same width as the feature map. Each candidate pixel in the row of candidate pixels is associated with a candidate high edge, and each candidate high edge is: a column of pixels in the image to be detected that passes through the associated candidate pixel.
[0211] A pixel screening unit 2204 is configured to screen out at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels.
[0212] A region determination unit 2205 is configured to obtain the main body region of the image to be detected based on the candidate edges respectively associated with the at least one target pixel, where each candidate edge is a candidate wide edge or a candidate high edge.
[0213] As a possible implementation, the pixel values are assigned by a first assignment method. In the first assignment method, a first pixel value is used to represent that the candidate edge associated with the candidate pixel belongs to the edge of the main body region of the image, and a second pixel value is used to represent that the candidate edge associated with the candidate pixel does not belong to the edge of the main body region of the image.
[0214] When screening out at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels, the pixel screening unit 2204 is specifically configured to:
[0215] Based on the respective values of the column of candidate pixels and the row of candidate pixels, screen out at least one candidate pixel with the value of the first pixel value from the column of candidate pixels and the row of candidate pixels as the target pixel.
[0216] As a possible implementation, the pixel values are assigned by a second assignment method. In the second assignment method, a set of pixel values between the first pixel value and the second pixel value is used to represent the probability that the candidate wide edge associated with the candidate pixel belongs to the edge of the main body region of the image.
[0217] When screening out at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels, the pixel screening unit 2204 is specifically configured to:
[0218] Based on the respective values of the column of candidate pixels and the row of candidate pixels, screen out at least one reference pixel that meets the set local value condition from the column of candidate pixels and the row of candidate pixels.
[0219] From the at least one reference pixel, at least one target pixel whose value reaches a set first pixel value threshold is selected.
[0220] As a possible implementation manner, when screening at least one reference pixel that meets the set local value condition from the column of candidate pixels and the row of candidate pixels based on the values of the column of candidate pixels and the row of candidate pixels respectively, the pixel screening unit 2204 is specifically configured to:
[0221] For a group of candidate pixels, where the group of candidate pixels is the column of candidate pixels or the row of candidate pixels, perform the following operations:
[0222] Based on the value change situation of the group of candidate pixels, divide the group of candidate pixels to obtain at least one candidate pixel set;
[0223] According to the pixel value size, respectively select candidate pixels whose values reach the set second pixel value threshold from at least one candidate pixel set as at least one reference pixel that meets the set local value condition, where the second pixel value threshold is less than the first pixel value threshold.
[0224] As a possible implementation manner, when performing feature extraction on the image to be detected to obtain a feature map, the feature extraction unit 2201 is specifically configured to:
[0225] Use a feature extraction model to perform feature extraction on the image to be detected to obtain a feature map; the resolution of the feature map conforms to the set resolution range.
[0226] As a possible implementation manner, when performing feature extraction on the image to be detected to obtain a feature map, the feature extraction unit 2201 is specifically configured to:
[0227] Use a feature extraction model to perform feature extraction on the image to be detected to obtain a feature map; the feature extraction model is obtained by iterative training using a sample image set, and each sample image in the sample image set is associated with a target label, and each target label is used to indicate the main subject area in the associated sample image.
[0228] As a possible implementation manner, the feature extraction unit 2201 is further configured to obtain the target label associated with a sample image through the following method:
[0229] Based on two true wide sides of the main subject area in the one sample image, obtain a corresponding column of sample pixels; among the column of sample pixels, the sample pixels with the value of the first pixel value are associated with the true wide sides.
[0230] Obtain a corresponding row of sample pixels based on two true high edges of the main subject area in the one sample image; among the row of sample pixels, the sample pixels with the value of the first pixel value are associated with the true high edges.
[0231] Based on the column of sample pixels and the row of sample pixels, obtain the target label associated with the one sample image.
[0232] As a possible implementation manner, when obtaining a corresponding column of sample pixels based on two true wide edges of the main subject area in the one sample image, the feature extraction unit 2201 is specifically configured to:
[0233] Based on the height of the feature map output by the feature extraction network and in combination with the set initial pixel value, construct a column of initial sample pixels;
[0234] When there is at least one initial sample pixel located at the two true wide edges in the column of initial sample pixels, use the first pixel value to assign values to the at least one initial sample pixel;
[0235] Based on the column of initial sample pixels after assignment, obtain a corresponding column of sample pixels.
[0236] As a possible implementation manner, after using the first pixel value to assign values to the at least one initial sample pixel and before obtaining a corresponding column of sample pixels based on the column of initial sample pixels after assignment, the feature extraction unit 2201 is further configured to:
[0237] Use the second pixel value to assign values to the other initial sample pixels in the column of initial sample pixels except the at least one initial sample pixel; or,
[0238] Use the pixel value between the first pixel value and the second pixel value to assign values to the other initial sample pixels in the column of initial sample pixels except the at least one initial sample pixel.
[0239] As a possible implementation manner, when obtaining the main subject area of the to-be-detected image based on the candidate edges respectively associated with the at least one target pixel, the area determination unit 2205 is specifically configured to:
[0240] If the number of target pixels is four, obtain the main subject area of the to-be-detected image based on the candidate edges respectively associated with the four target pixels;
[0241] If the number of target pixels is less than four, obtain the main subject area of the to-be-detected image based on the candidate edges respectively associated with the at least one target pixel and in combination with the edges of the to-be-detected image.
[0242] For the convenience of description, the above parts are divided into respective modules (or units) according to their functions and described separately. Of course, when implementing the present application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.
[0243] Regarding the device in the above embodiments, the specific manner in which each unit executes the request has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0244] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0245] Based on the same inventive concept, an embodiment of the present application further provides an electronic device. In one embodiment, the electronic device can be a server or a terminal device. Refer to Figure 23 As shown, it is a schematic structural diagram of a possible electronic device provided in an embodiment of the present application. Figure 23 In it, the electronic device 2300 includes: a processor 2310 and a memory 2320.
[0246] Among them, the memory 2320 stores a computer program executable by the processor 2310. By executing the instructions stored in the memory 2320, the processor 2310 can execute the steps of the above-mentioned method for detecting the main body of the picture.
[0247] The memory 2320 can be a volatile memory, such as a random-access memory (RAM); the memory 2320 can also be a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 2320 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2320 can also be a combination of the above memories.
[0248] The processor 2310 may include one or more central processing units (CPUs) or be a digital processing unit, etc. When the processor 2310 executes the computer program stored in the memory 2320, the above-mentioned method for detecting the main body of the picture is implemented.
[0249] In some embodiments, the processor 2310 and the memory 2320 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.
[0250] In the embodiments of the present application, the specific connection medium between the above-mentioned processor 2310 and the memory 2320 is not limited. In the embodiments of the present application, taking the connection between the processor 2310 and the memory 2320 through a bus as an example, the bus is Figure 23 described by a thick line in []. The connection manners between other components are only for illustrative purposes and are not limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 23 only a thick line is used to describe it in []. However, it does not describe that there is only one bus or one type of bus.
[0251] Based on the same inventive concept, the embodiments of the present application provide a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above-mentioned method for detecting the main body of the picture. In some possible implementation manners, each aspect of the method for detecting the main body of the picture provided by the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the above-mentioned method for detecting the main body of the picture. For example, the electronic device can execute as Figure 2 the steps shown in [].
[0252] The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0253] The program product of the embodiments of the present application may adopt a CD-ROM and include a computer program, and can run on an electronic device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a computer program, and the computer program can be used by or in combination with a command execution system, device, or apparatus.
[0254] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a computer program for use by or in combination with a command execution system, device, or apparatus.
[0255] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0256] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for detecting the main subject of a picture, characterized in that Method: Extract features from the image to be detected to obtain a feature map, where the main area of the image to be detected is rectangular; Reduce the dimension of the feature map in the width direction to obtain a column of candidate pixels with the same height as the feature map. Each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected that passes through the associated candidate pixel; Reduce the dimension of the feature map in the height direction to obtain a row of candidate pixels with the same width as the feature map. Each candidate pixel in the row of candidate pixels is associated with a candidate high edge, and each candidate high edge is: a column of pixels in the image to be detected that passes through the associated candidate pixel; Based on the values of the column of candidate pixels and the row of candidate pixels respectively, select at least one target pixel from the column of candidate pixels and the row of candidate pixels; Based on the candidate edges associated with the at least one target pixel respectively, obtain the main area of the image to be detected, where each candidate edge is a candidate wide edge or a candidate high edge.
2. The method according to claim 1, wherein The pixel values are assigned using a first assignment method. In the first assignment method, a first pixel value is used to represent that the candidate edge associated with the candidate pixel belongs to the edge of the main area of the image, and a second pixel value is used to represent that the candidate edge associated with the candidate pixel does not belong to the edge of the main area of the image; Then the step of selecting at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the values of the column of candidate pixels and the row of candidate pixels respectively includes: Based on the values of the column of candidate pixels and the row of candidate pixels respectively, select at least one candidate pixel with a value of the first pixel value from the column of candidate pixels and the row of candidate pixels as the target pixel.
3. The method according to claim 1, characterized in that, The pixel values are assigned using a second assignment method. In the second assignment method, a set of pixel values between the first pixel value and the second pixel value is used to represent the probability that the candidate wide edge associated with the candidate pixel belongs to the edge of the main area of the image; Then the step of selecting at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the values of the column of candidate pixels and the row of candidate pixels respectively includes: Based on the values of the column of candidate pixels and the row of candidate pixels respectively, select at least one reference pixel that meets the set local value condition from the column of candidate pixels and the row of candidate pixels; Select at least one target pixel with a value reaching the set first pixel value threshold from the at least one reference pixel.
4. The method according to claim 3, characterized in that, The step of selecting at least one reference pixel that meets the set local value condition from the column of candidate pixels and the row of candidate pixels based on the values of the column of candidate pixels and the row of candidate pixels respectively includes: For a group of candidate pixels, the group of candidate pixels is the column of candidate pixels or the row of candidate pixels, and perform the following operations: Based on the value change situation of the group of candidate pixels, divide the group of candidate pixels to obtain at least one candidate pixel set; According to the pixel value, at least one candidate pixel is selected from at least one candidate pixel set, and the candidate pixel whose value reaches the set second pixel value threshold is used as at least one reference pixel that meets the set local value condition, where the second pixel value threshold is less than the first pixel value threshold.
5. The method according to any one of claims 1-4, characterized in that, The feature extraction of the image to be detected to obtain a feature map includes: Using a feature extraction model to perform feature extraction on the image to be detected to obtain a feature map; the resolution of the feature map meets the set resolution range.
6. The method according to any one of claims 1-4, characterized in that, The feature extraction of the image to be detected to obtain a feature map includes: Using a feature extraction model to perform feature extraction on the image to be detected to obtain a feature map; the feature extraction model is obtained by iterative training using a sample image set, and each sample image in the sample image set is associated with a target label, and each target label is used to indicate the main area of the picture in the associated sample image.
7. The method according to claim 6, wherein The target label associated with a sample image is obtained by the following method: Based on the two true wide edges of the main area of the picture in the one sample image, a column of sample pixels is obtained; among the column of sample pixels, the sample pixels with the first pixel value are associated with the true wide edge. Based on the two true high edges of the main area of the picture in the one sample image, a row of sample pixels is obtained; among the row of sample pixels, the sample pixels with the first pixel value are associated with the true high edge. Based on the column of sample pixels and the row of sample pixels, the target label associated with the one sample image is obtained.
8. The method according to claim 7, wherein The obtaining of a column of sample pixels based on the two true wide edges of the main area of the picture in the one sample image includes: Based on the height of the feature map output by the feature extraction network and in combination with the set initial pixel value, a column of initial sample pixels is constructed. When there is at least one initial sample pixel located on the two true wide edges in the column of initial sample pixels, the first pixel value is used to assign values to at least one initial sample pixel. Based on the column of initial sample pixels after assignment, a corresponding column of sample pixels is obtained.
9. The method according to claim 8, wherein After the assignment of at least one initial sample pixel with the first pixel value and before the obtaining of a corresponding column of sample pixels based on the column of initial sample pixels after assignment, it further includes: Using the second pixel value to assign values to the other initial sample pixels in the column of initial sample pixels except the at least one initial sample pixel; or, Using a pixel value between the first pixel value and the second pixel value to assign values to the other initial sample pixels in the column of initial sample pixels except the at least one initial sample pixel.
10. The method according to any one of claims 1-4, characterized in that, The obtaining of the main area of the picture of the image to be detected based on the candidate edges associated with each of the at least one target pixel includes: If the number of target pixels is four, the main area of the picture of the image to be detected is obtained based on the candidate edges associated with each of the four target pixels. If the number of target pixels is less than four, based on the candidate edges respectively associated with the at least one target pixel and in combination with the edges of the image to be detected, the main area of the image to be detected is obtained.
11. A device for detecting a main subject in a picture, characterized in that It includes: A feature extraction unit for extracting features from the image to be detected to obtain a feature map, wherein the main area of the image to be detected is rectangular; A column vector extraction unit for reducing the dimension of the feature map in the width direction to obtain a column of candidate pixels having the same height as the feature map, and each candidate pixel in the column of candidate pixels is associated with a candidate wide edge, and each candidate wide edge is: a row of pixels in the image to be detected passing through the associated candidate pixel; A row vector extraction unit for reducing the dimension of the feature map in the height direction to obtain a row of candidate pixels having the same width as the feature map, and each candidate pixel in the row of candidate pixels is associated with a candidate high edge, and each candidate high edge is: a column of pixels in the image to be detected passing through the associated candidate pixel; A pixel screening unit for screening out at least one target pixel from the column of candidate pixels and the row of candidate pixels based on the respective values of the column of candidate pixels and the row of candidate pixels; A region determination unit for obtaining the main area of the image to be detected based on the candidate edges respectively associated with the at least one target pixel, wherein each candidate edge is a candidate wide edge or a candidate high edge.
12. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, It includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of claims 1 to 10.
14. A computer program product, characterized in that, It includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of any one of claims 1 to 10.