Image Segmentation Method, Device, Electronic Device and Storage Medium
By searching the target image and segmented block images for whole pixels to obtain motion estimation information and inputting it into segmented neural network, the problem of high complexity of Libaom AV1 encoder is solved, the encoding efficiency is improved, and it is suitable for real-time communication applications.
Patent Information
- Application Number
- CN202011606637.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-12-28
AI Technical Summary
The Libaom AV1 encoder has high motion estimation and coding complexity, making it difficult to adapt to real-time communication application scenarios.
By searching the target image and the target segmented block image separately, the motion estimation information is obtained, and inputting it into the target segmented neural network, reducing the computational complexity of the motion estimation and neural network.
Without reducing image segmentation accuracy and encoding quality, encoding efficiency is improved and adapted to real-time communication application scenarios such as video conferencing and video capture scenarios where cameras capture videos.
Smart Images

Figure CN114693714B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia processing technologies, and in particular, to an image segmentation method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of multimedia technologies nowadays, encoders have attracted much attention. For example, H.264 / 265 / 266 encoders. However, H.264 / 265 / 266 encoders require high patent fees in practical applications. Based on this situation, the AOM Alliance (Alliance for Open Media) has proposed the first generation of royalty-free high-performance codec AV1 (aomedia video 1). AV1 has a performance improvement of more than 30% compared to the H.265 encoder, and there is only a 5-6% performance difference compared to the latest H.266 encoder. Moreover, the complexity at the best performance of AV1 does not exceed half of the best performance complexity of H.266. Therefore, AV1 has received more attention in the industry and has more possibilities for applications to be implemented.
[0003] In the encoding process, it is necessary to determine whether the current image or the segmented block image of the current image needs to be further segmented. Based on this, Libaom AV1 (aomedia video encoder for av1, a reference codec for AV1) provides different gears to correspond to different encoding speeds and encoding qualities. Starting from gear 1 to gear 5, each gear has a corresponding neural network NN to determine whether the current image or the segmented block image of the current image needs to be further segmented, improving the encoding speed. However, the complexity of motion estimation and encoding in the gears provided by Libaom AV1 is relatively high, resulting in a relatively high processing complexity of the segmentation neural network NN (neural networks), and it cannot be well applied to real-time communication application scenarios. Summary of the Invention
[0004] In view of the above-mentioned existing technical problems, this application proposes an image segmentation method, apparatus, electronic device, and storage medium.
[0005] According to one aspect of this application, an image segmentation method is provided, and the method includes:
[0006] Obtain a target image and a set of segmentation styles;
[0007] Extract a target segmentation style from the set of segmentation styles;
[0008] Based on the target segmentation style and the target image, obtain a target segmented block image;
[0009] In the reference image, perform a whole-pixel search on the target image and the target segmented block image respectively to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmented block image and the reference image;
[0010] Input the first motion estimation information and the second motion estimation information into the target segmentation neural network corresponding to the target segmentation style to obtain a target segmentation result;
[0011] Perform segmentation processing on the target image according to the target segmentation result.
[0012] According to another aspect of the present application, there is provided an image segmentation device, including:
[0013] A target image and segmentation style set acquisition module, configured to acquire a target image and a segmentation style set;
[0014] A target segmentation style extraction module, configured to extract a target segmentation style from the segmentation style set;
[0015] A target segmented block image acquisition module, configured to obtain a target segmented block image based on the target segmentation style and the target image;
[0016] A motion estimation information acquisition module, configured to perform a whole-pixel search on the target image and the target segmented block image respectively in the reference image to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmented block image and the reference image;
[0017] A target segmentation result acquisition module, configured to input the first motion estimation information and the second motion estimation information into the target segmentation neural network corresponding to the target segmentation style to obtain a target segmentation result;
[0018] An image segmentation processing module, configured to perform segmentation processing on the target image according to the target segmentation result.
[0019] According to another aspect of the present application, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above method.
[0020] According to another aspect of the present application, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, when the computer program instructions are executed by a processor, the above method is implemented.
[0021] By performing whole-pixel search on the target image and the target segmentation block image respectively, the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmentation block image and the reference image can be obtained, which can reduce the computational complexity of motion estimation; and using the first motion estimation information and the second motion estimation information as the input of the target segmentation neural network reduces the input information of the target segmentation neural network, thereby reducing the processing complexity of the target segmentation neural network. Without reducing the image segmentation accuracy and coding quality, the coding efficiency is improved. This image segmentation scheme can effectively adapt to real-time communication application scenarios, for example, it can be effectively applied to coding scenarios based on LP (low delay P-frame, low delay unidirectional reference frame) (such as coding scenarios in video conferencing, coding scenarios for videos captured by cameras).
[0022] Other features and aspects of the present application will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings included in and constituting a part of this specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present application and are used to explain the principles of the present application.
[0024] Figure 1 FIG. shows a schematic diagram of an application system provided according to an embodiment of the present application.
[0025] Figure 2 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application.
[0026] Figure 3 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application.
[0027] Figure 4 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application.
[0028] Figure 5 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application.
[0029] Figure 6 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application.
[0030] Figure 7 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application.
[0031] Figure 8The flowchart shows a method for performing a whole-pixel search on a target image and a target segmentation block image in a reference image according to an embodiment of the present application, to obtain first motion estimation information between the target image and the reference image and second motion estimation information between the target segmentation block image and the reference image.
[0032] Figure 9 The block diagram shows an image segmentation device according to an embodiment of the present application.
[0033] Figure 10 The block diagram shows an electronic device for image segmentation according to an embodiment of the present application. Detailed implementation manners
[0034] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0035] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior or better than other embodiments.
[0036] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0037] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0038] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely applied in multiple fields. The solution provided by the embodiments of the present application involves technologies such as machine learning / deep learning, and is specifically described through the following embodiments:
[0039] Please refer to Figure 1 , Figure 1 The schematic diagram shows an application system provided according to an embodiment of the present application. The application system can be used for the image segmentation method of the present application. As Figure 1As shown in the figure, the application system may at least include a server 01 and a terminal 02.
[0040] In the embodiment of the present application, the server 01 may be a server cluster or a distributed system composed of multiple physical servers, or may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0041] In the embodiment of the present application, the terminal 02 may include physical devices of types such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. The physical device may also include software running on the physical device, such as application programs, etc. In the embodiment of the present application, the operating system running on the terminal 02 may include, but is not limited to, Android system, IOS system, Linux, Windows, etc.
[0042] In the embodiment of this specification, the above terminal 02 and server 01 may be directly or indirectly connected by wired or wireless communication methods, and the present application does not make any limitations in this regard.
[0043] The terminal 02 may be used to provide image segmentation processing for users. Users may upload a target image or video to be encoded on the terminal 02. The ways for the terminal 02 to provide image segmentation processing for users may include, but are not limited to, the application program way, the web page way, etc.
[0044] In a specific embodiment, when the server 02 is a distributed system, the distributed system may be a blockchain system. When the distributed system is a blockchain system, it may be formed by multiple nodes (any form of computing device connected to the network, such as a server, a user terminal). In the distributed system, any machine such as a server or a terminal can join and become a node. The node includes a hardware layer, a middle layer, an operating system layer, and an application layer. Specifically, the functions of each node in the blockchain system may involve the following functions:
[0045] 1) Routing, which is a basic function of the node and is used to support communication between nodes.
[0046] In addition to the routing function, the node may also have the following functions:
[0047] 2) An application, which is used to be deployed in a blockchain, implements a specific business according to actual business requirements, records data related to the implemented functions to form recorded data, carries a digital signature in the recorded data to indicate the source of the task data, and sends the recorded data to other nodes in the blockchain system. When other nodes verify the source and integrity of the recorded data successfully, the recorded data is added to a temporary block.
[0048] It should be noted that the following shows a possible order of steps. In fact, it is not necessarily limited to strictly follow this order. Some steps can be executed in parallel without depending on each other.
[0049] Specifically, Figure 2 The flowchart showing the image segmentation method according to an embodiment of the present application is shown. As Figure 2 shown, the method may include:
[0050] S201, obtain a target image and a set of segmentation styles.
[0051] In the embodiments of the present specification, the target image may refer to an image to be segmented. The image to be segmented may refer to a frame image or a segmented block image in a frame image. The frame image may be a frame image in a video to be encoded. The present application does not limit this. The set of segmentation styles may be preset based on the set of segmentation styles in AV1. For example, the set of segmentation styles may include all or part of the segmentation styles in the set of segmentation styles in AV1. The set of segmentation styles in AV1 may include 11 segmentation styles such as square average quartering, horizontal bisection, vertical bisection, horizontal quartering, vertical quartering, horizontal T-shaped, and vertical T-shaped. The present application does not limit this either.
[0052] In the embodiments of the present specification, an image to be segmented may be obtained as the target image, and a preset set of segmentation styles may be obtained.
[0053] S203, extract a target segmentation style from the set of segmentation styles.
[0054] In the embodiments of the present specification, a segmentation style may be extracted from the set of segmentation styles as the target segmentation style. The present application does not limit the extraction method.
[0055] S205, obtain a target segmented block image based on the target segmentation style and the target image.
[0056] In the embodiments of this specification, the target image can be segmented based on the target segmentation style to obtain the target segmented block image. For example, when the target segmentation style is the segmentation style of evenly dividing a square into four parts, the target image can be segmented based on this segmentation style of evenly dividing a square into four parts to obtain four target segmented block images. These four target segmented block images can form the target image.
[0057] S207. In the reference image, perform whole-pixel search on the target image and the target segmented block image respectively to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmented block image and the reference image.
[0058] In the embodiments of this specification, based on a preset search algorithm, in the reference image, perform whole-pixel search on the target image and the target segmented block image respectively to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmented block image and the reference image. The preset search algorithm can include the large diamond search algorithm, and this application does not limit this. Among them, the reference image can refer to at least one frame image in the encoded images.
[0059] In the embodiments of this specification, the first motion estimation information can include the pixel difference information between the target image and the reference image. For example, it can include the pixel variance information VAR (variance), the sum of squared pixel difference information SSD (Sum of Squared Difference), or it can also include the pixel motion vector information. This application does not limit this.
[0060] Correspondingly, the second motion estimation information can include the pixel difference information between the target segmented block image and the reference image. For example, it can include the pixel variance information VAR, the sum of squared pixel difference information SSD, or it can also include the pixel motion vector information. This application does not limit this.
[0061] S209. Input the first motion estimation information and the second motion estimation information into the target segmentation neural network corresponding to the target segmentation style to obtain the target segmentation result.
[0062] In the embodiments of this specification, among the 1-5 gears of AV1, the higher the gear, the lower the accuracy and the faster the speed. Each gear corresponds to multiple segmentation neural networks, and these multiple segmentation neural networks can correspond to multiple segmentation styles. Based on this, according to the encoding accuracy requirement, the target segmentation neural network corresponding to the target segmentation style can be obtained from the corresponding gear of AV1. For example, if the encoding accuracy requirement is relatively low, the target segmentation neural network corresponding to the target segmentation style can be obtained from the multiple segmentation neural networks corresponding to gear 5. Among them, the segmentation neural network can be a neural network including a convolutional layer and an activation layer.
[0063] In the embodiments of this specification, the first motion estimation information and the second motion estimation information may be input into a target segmentation neural network corresponding to a target segmentation style to obtain a target segmentation result. The target segmentation result may be a segmentation probability, that is, the target segmentation result may be a value. The segmentation probability may correspond to the target segmentation style and may represent the probability that the target image is segmented according to the target segmentation style. The higher the segmentation probability, the higher the possibility that the target image is segmented according to the target segmentation style.
[0064] Optionally, in addition to the above-mentioned first motion estimation information and second motion estimation information, the input information of the target segmentation neural network may further include the resolution information of the target image and the resolution information of the target segmentation block image. This application does not limit this.
[0065] S2011. Perform segmentation processing on the target image according to the target segmentation result.
[0066] In the embodiments of this specification, the segmentation style for segmentation processing may be determined according to the target segmentation result, so that the target image can be segmented based on the segmentation style for segmentation processing. In one example, as Figure 3 shown, the target segmentation style may be any segmentation style in the segmentation style set; that is, when traversing each segmentation style in the segmentation style set, the target segmentation result in S2011 may include the target segmentation result corresponding to any segmentation style in the segmentation style set. This step S2011 may include:
[0067] S301. Obtain the segmentation probability corresponding to each segmentation style in the segmentation style set;
[0068] S303. Obtain the segmentation style corresponding to the maximum segmentation probability;
[0069] S305. Perform segmentation processing on the target image according to the segmentation style corresponding to the maximum segmentation probability.
[0070] In the embodiments of this specification, the segmentation probability corresponding to each segmentation style in the segmentation style set may be obtained, so that the segmentation style corresponding to the maximum segmentation probability can be obtained, and the target image can be segmented according to the segmentation style corresponding to the maximum segmentation probability.
[0071] By performing whole-pixel search on the target image and the target segmentation block image respectively, the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmentation block image and the reference image can be obtained, which can reduce the computational complexity of motion estimation; and using the first motion estimation information and the second motion estimation information as the input of the target segmentation neural network reduces the input information of the target segmentation neural network, thereby reducing the processing complexity of the target segmentation neural network. On the basis of not reducing the image segmentation accuracy and coding quality, the coding efficiency is improved. This image segmentation scheme can effectively adapt to real-time communication application scenarios, for example, it can be effectively applied to LP-based coding scenarios (such as coding scenarios in video conferencing and coding scenarios for videos captured by cameras).
[0072] Figure 4 FIG. shows a flowchart of an image segmentation method according to an embodiment of the present application. As Figure 4 shown, in a possible implementation, step S203 may include:
[0073] S401, obtaining the priority corresponding to the segmentation style in the segmentation style set.
[0074] In practical applications, different segmentation styles have different segmentation complexities and corresponding coding qualities. In the embodiments of this specification, the corresponding priorities can be set for different segmentation styles by combining the segmentation complexity of the segmentation style and the influence degree on the coding quality. In an optional embodiment, the segmentation complexity and the influence degree on the coding quality of different segmentation styles can be quantified according to a preset quantization rule, and the quantization value corresponding to the segmentation complexity of different segmentation styles and the quantization value corresponding to the influence degree on the coding quality can be obtained. Further, the weighted value of the quantization value corresponding to the segmentation complexity of different segmentation styles and the quantization value corresponding to the influence degree on the coding quality is obtained, and this weighted value can be used as the priority corresponding to the segmentation style. Among them, the lower the segmentation complexity and the higher the influence degree on the coding quality, the higher the weighted value and the higher the corresponding priority. The present application does not limit this. There may be a corresponding relationship between the segmentation style and the priority, so that the priority corresponding to the segmentation style in the segmentation style set can be obtained based on this corresponding relationship between the segmentation style and the priority.
[0075] As an example, for the segmentation style set in AV1, it includes square mean quartering, horizontal bisection, vertical bisection, horizontal quartering, vertical quartering, horizontal T-shaped, vertical T-shaped, etc. The order of the set priorities from high to low can be: square mean quartering, horizontal bisection, vertical bisection, horizontal quartering, vertical quartering, horizontal T-shaped, vertical T-shaped. The present application does not limit this, and the priority can be set according to actual needs.
[0076] S403. Extract a target segmentation style from the segmentation style set according to the priority corresponding to the segmentation style.
[0077] In the embodiments of this specification, a segmentation style can be extracted from the segmentation style set as the target segmentation style according to the order from high to low priority.
[0078] Correspondingly, in a possible implementation, as Figure 5 shown, step S2011 may include:
[0079] S501. Obtain the segmentation threshold corresponding to the target segmentation neural network;
[0080] S503. When the target segmentation result is greater than the segmentation threshold, perform segmentation processing on the target image according to the target segmentation style.
[0081] In the embodiments of this specification, a mapping relationship between each segmentation neural network and the segmentation threshold can be preset. The segmentation threshold can be set according to experience or can be obtained through pre-training. The segmentation threshold corresponding to the target segmentation neural network can be obtained according to this mapping relationship. And it can be determined whether the target segmentation result is greater than the segmentation threshold. When the target segmentation result is greater than the segmentation threshold, segmentation processing can be performed on the target image according to the target segmentation style.
[0082] By extracting the target segmentation style from the segmentation style set according to the priority corresponding to the segmentation style, when the target segmentation result is greater than the segmentation threshold, the target image can be directly segmented according to the target segmentation style. It is not necessary to traverse all the segmentation styles in the segmentation style set, improving the efficiency of image segmentation and being able to effectively meet the real-time requirements.
[0083] Figure 6 The flowchart of an image segmentation method according to an embodiment of the present application is shown. As Figure 6 shown, in a possible implementation, before step S209, the method may further include:
[0084] S601. Obtain the set of segmentation neural networks corresponding to the preset profile in the video coding format AV1.
[0085] In the embodiments of this specification, the preset gear may refer to gear 5 in AV1 (cpuused = 5). The segmentation neural networks in the segmentation neural network set corresponding to the existing gear 5 have lower accuracy than those in the segmentation neural network sets corresponding to other gears. However, the segmentation neural networks corresponding to gears 1-5 all perform integer-pixel search and sub-pixel search (1 / 2 and 1 / 4 pixel search) to obtain motion estimation information as input information. That is, the accuracy and complexity of the input information of gear 5 are the same as those of other gears. Based on the relatively low accuracy of the segmentation neural network of gear 5, it is possible to simplify the input information of gear 5 without affecting the output accuracy of the segmentation neural networks in the segmentation neural network set corresponding to gear 5. Based on this, setting the preset gear to gear 5 in AV1 can effectively meet the coding requirements in the LP low-latency application scenario. In such a low-resolution real-time application scenario, it is possible to improve the coding efficiency while not reducing the coding quality.
[0086] In the embodiments of this specification, a segmentation neural network set corresponding to a preset gear in the video coding format AV1 can be obtained.
[0087] S603. Obtain the correspondence between the segmentation neural networks in the segmentation neural network set and the segmentation styles in the segmentation style set.
[0088] In the embodiments of this specification, the segmentation neural network may be one that has been set in AV1; or it may be obtained by training based on the segmentation neural networks that have been set in AV1. This application does not make any limitations in this regard. Thus, according to the correspondence between the segmentation neural networks in AV1 and the segmentation styles, the correspondence between the segmentation neural networks in the segmentation neural network set and the segmentation styles in the segmentation style set can be set.
[0089] S605. According to the correspondence, obtain the target segmentation neural network corresponding to the target segmentation style from the segmentation neural network set.
[0090] In the embodiments of this specification, the target segmentation neural network corresponding to the target segmentation style can be obtained from the segmentation neural network set according to the correspondence between the segmentation neural networks in the segmentation neural network set and the segmentation styles in the segmentation style set.
[0091] By obtaining the target segmentation neural network from the segmentation neural network set corresponding to the preset gear of AV1, the corresponding target segmentation neural network can be selected according to actual needs, so that image segmentation can be applied to different application scenarios.
[0092] Figure 7 The flowchart of an image segmentation method according to an embodiment of the present application is shown. As Figure 7 shown, in a possible implementation, when the target image is a block image, the method may further include:
[0093] S701. Obtain the segmentation style information of the adjacent block image of the block image.
[0094] In the embodiments of this specification, the block image may be one of at least two segmented block images obtained by segmenting a frame of image; the adjacent block image may be another one of the at least two segmented block images. In this frame of image, this one segmented block image is adjacent to this other one segmented block image. It should be noted that the adjacent block image may be a segmented block image that has been segmented. Thus, the segmentation style information of the adjacent block image of the block image can be obtained. The segmentation style information may be one of the segmentation style sets.
[0095] Correspondingly, step S209 may include:
[0096] S703. Input the segmentation style information, the first motion estimation information, and the second motion estimation information into the target segmentation neural network to obtain the target segmentation result.
[0097] In the embodiments of this specification, the segmentation style information, the first motion estimation information, and the second motion estimation information may be input into the target segmentation neural network to obtain the target segmentation result.
[0098] It should be noted that when the input of the target segmentation neural network includes the segmentation style information, the corresponding segmentation threshold needs to be reset, and this setting can be set according to experience or can be set according to training. This application does not make a limitation in this regard. Correspondingly, the target image can be segmented according to whether the target segmentation result is greater than the reset segmentation threshold.
[0099] By adding the segmentation style information of the adjacent block image, the input of the target segmentation neural network has segmentation reference information, improving the accuracy of the input information, and thus the accuracy of image segmentation can be improved. And as long as the segmentation threshold for determining whether to segment is reset, there is no need to retrain the segmentation neural network, reducing the technical complexity and improving the image segmentation efficiency.
[0100] Figure 8 The flowchart shows a method for performing whole-pixel search on a target image and a target segmented block image respectively in a reference image according to an embodiment of the present application to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmented block image and the reference image. As Figure 8 shown, in a possible implementation, this step S207 may include:
[0101] S801. In the reference image, perform a whole-pixel search on the target image and the target segmentation block image respectively to determine the first target matching block in the reference image that matches the target image and the second target matching block in the reference image that matches the target segmentation block image.
[0102] In the embodiments of this specification, based on a preset search algorithm, such as the large diamond search algorithm, in the reference image, perform a whole-pixel search on the target image to determine the first target matching block in the reference image that matches the target image; and based on the preset search algorithm, such as the large diamond search algorithm, perform a whole-pixel search on the target segmentation block image to determine the second target matching block in the reference image that matches the target segmentation block image.
[0103] S803. Obtain the first pixel variance information and the first sum of squared pixel differences information between the target image and the first target matching block.
[0104] In the embodiments of this specification, calculate the difference between the pixel values of the target image and the pixel values of the first target matching block to obtain the first pixel variance information VAR and the first sum of squared pixel differences information SSD between the target image and the first target matching block.
[0105] S805. Obtain the second pixel variance information and the second sum of squared pixel differences information between the target segmentation block image and the second target matching block; this step can refer to S803 and will not be elaborated here.
[0106] S807. Use the first pixel variance information and the first sum of squared pixel differences information as the first motion estimation information;
[0107] S809. Use the second pixel variance information and the second sum of squared pixel differences information as the second motion estimation information.
[0108] In the embodiments of this specification, the first pixel variance information and the first sum of squared pixel differences information can be used as the first motion estimation information; and the second pixel variance information and the second sum of squared pixel differences information can be used as the second motion estimation information.
[0109] The first motion estimation information and the second motion estimation information obtained through whole-pixel search reduce the computational complexity of motion estimation, reduce the information input to the target segmentation neural network, and improve the image segmentation efficiency.
[0110] Figure 9 The block diagram of an image segmentation device according to an embodiment of the present application is shown. As Figure 9 shown, the device may include:
[0111] A target image and segmentation style set acquisition module 901, configured to acquire a target image and a segmentation style set;
[0112] A target segmentation style extraction module 903, configured to extract a target segmentation style from the segmentation style set;
[0113] A target segmentation block image acquisition module 905, configured to obtain a target segmentation block image based on the target segmentation style and the target image;
[0114] A motion estimation information acquisition module 907, configured to perform a whole-pixel search on the target image and the target segmentation block image respectively in a reference image, to obtain first motion estimation information between the target image and the reference image and second motion estimation information between the target segmentation block image and the reference image;
[0115] A target segmentation result acquisition module 909, configured to input the first motion estimation information and the second motion estimation information into a target segmentation neural network corresponding to the target segmentation style, to obtain a target segmentation result;
[0116] An image segmentation processing module 9011, configured to perform segmentation processing on the target image according to the target segmentation result.
[0117] By performing a whole-pixel search on the target image and the target segmentation block image respectively, first motion estimation information between the target image and the reference image and second motion estimation information between the target segmentation block image and the reference image are obtained, which can reduce the computational complexity of motion estimation; and the first motion estimation information and the second motion estimation information are used as inputs to the target segmentation neural network, reducing the input information of the target segmentation neural network, thereby reducing the processing complexity of the target segmentation neural network, and improving the coding efficiency without reducing the image segmentation accuracy and coding quality. This image segmentation scheme can effectively adapt to real-time communication application scenarios, for example, it can be effectively applied to LP-based coding scenarios (such as coding scenarios in video conferencing, coding scenarios for videos captured by cameras).
[0118] In a possible implementation manner, the target segmentation style extraction module 903 may include:
[0119] A priority acquisition unit, configured to acquire priorities corresponding to the segmentation styles in the segmentation style set;
[0120] A target segmentation style extraction unit, configured to extract the target segmentation style from the segmentation style set according to the priorities corresponding to the segmentation styles.
[0121] In a possible implementation manner, the image segmentation processing module 9011 may include:
[0122] A segmentation threshold acquisition unit, configured to acquire a segmentation threshold corresponding to the target segmentation neural network;
[0123] A first image segmentation processing unit, configured to perform segmentation processing on the target image according to the target segmentation style when the target segmentation result is greater than the segmentation threshold.
[0124] In a possible implementation manner, when the target segmentation result is a segmentation probability; the target segmentation style is any segmentation style in the segmentation style set; the image segmentation processing module 9011 may include:
[0125] A segmentation probability acquisition unit, configured to acquire a segmentation probability corresponding to each segmentation style in the segmentation style set;
[0126] A target segmentation style acquisition unit, configured to acquire the segmentation style corresponding to the maximum segmentation probability;
[0127] A second image segmentation processing unit, configured to perform segmentation processing on the target image according to the segmentation style corresponding to the maximum segmentation probability.
[0128] In a possible implementation manner, the apparatus may further include:
[0129] A segmentation neural network set acquisition module, configured to acquire a set of segmentation neural networks corresponding to a preset gear in the video coding format AV1;
[0130] A correspondence relationship acquisition module between the segmentation neural network set and the segmentation style, configured to acquire the correspondence relationship between the segmentation neural networks in the segmentation neural network set and the segmentation styles in the segmentation style set;
[0131] A target segmentation neural network acquisition module, configured to acquire the target segmentation neural network corresponding to the target segmentation style from the segmentation neural network set according to the correspondence relationship.
[0132] In a possible implementation manner, when the target image is a block image, the apparatus further includes:
[0133] A segmentation style information acquisition module, configured to acquire segmentation style information of adjacent block images of the block image;
[0134] Correspondingly, the target segmentation result acquisition module 909 may include:
[0135] A target segmentation result acquisition unit, configured to input the segmentation style information, the first motion estimation information, and the second motion estimation information into the target segmentation neural network to obtain a target segmentation result.
[0136] In a possible implementation, the motion estimation information acquisition module 907 may include:
[0137] A matching block acquisition unit, configured to perform full-pixel search on the target image and the target segmentation block image respectively in the reference image, and determine a first target matching block in the reference image that matches the target image and a second target matching block in the reference image that matches the target segmentation block image;
[0138] A first pixel information acquisition unit, configured to acquire first pixel variance information and first sum of squared pixel difference information between the target image and the first target matching block;
[0139] A second pixel information acquisition unit, configured to acquire second pixel variance information and second sum of squared pixel difference information between the target segmentation block image and the second target matching block;
[0140] A first motion estimation information acquisition unit, configured to use the first pixel variance information and the first sum of squared pixel difference information as the first motion estimation information;
[0141] A second motion estimation information acquisition unit, configured to use the second pixel variance information and the second sum of squared pixel difference information as the second motion estimation information.
[0142] Regarding the device in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0143] On the other hand, the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image segmentation method provided in the above various optional implementations.
[0144] Figure 10 The block diagram of an electronic device for image segmentation according to an embodiment of the present application is shown. The electronic device may be a server, and its internal structure diagram may be as Figure 10As shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for image segmentation.
[0145] Those skilled in the art can understand that Figure 10 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0146] In an exemplary embodiment, there is also provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the image segmentation method as in the embodiment of this application.
[0147] In an exemplary embodiment, there is also provided a storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the image segmentation method in the embodiment of this application.
[0148] In an exemplary embodiment, there is also provided a computer program product containing instructions, when it runs on a computer, the computer executes the image segmentation method in the embodiment of this application.
[0149] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0150] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0151] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An image segmentation method, characterized in that, The method includes: Obtaining a target image and a set of segmentation styles; Extracting a target segmentation style from the set of segmentation styles; Based on the target segmentation style and the target image, obtaining a target segmented block image; In a reference image, performing whole-pixel search on the target image and the target segmented block image respectively, to obtain first motion estimation information between the target image and the reference image and second motion estimation information between the target segmented block image and the reference image; Inputting the first motion estimation information and the second motion estimation information into a target segmentation neural network corresponding to the target segmentation style, to obtain a target segmentation result; the target segmentation result represents the probability that the target image is segmented according to the target segmentation style; Performing segmentation processing on the target image according to the target segmentation result; Wherein, the performing segmentation processing on the target image according to the target segmentation result includes: determining a segmentation style for segmentation processing according to the target segmentation result, and performing segmentation processing on the target image based on the segmentation style for segmentation processing.
2. The method according to claim 1, characterized in that, It further includes: The extracting a target segmentation style from the set of segmentation styles includes: Obtaining the priorities corresponding to the segmentation styles in the set of segmentation styles; Extracting the target segmentation style from the set of segmentation styles according to the priorities corresponding to the segmentation styles.
3. The method according to claim 2, characterized in that The performing segmentation processing on the target image according to the target segmentation result includes: Obtaining a segmentation threshold corresponding to the target segmentation neural network; When the target segmentation result is greater than the segmentation threshold, performing segmentation processing on the target image according to the target segmentation style.
4. The method according to claim 1, characterized in that, When the target segmentation result is a segmentation probability; the target segmentation style is any segmentation style in the set of segmentation styles; the performing segmentation processing on the target image according to the target segmentation result includes: Obtaining the segmentation probabilities corresponding to each segmentation style in the set of segmentation styles; Obtaining the segmentation style corresponding to the maximum segmentation probability; Performing segmentation processing on the target image according to the segmentation style corresponding to the maximum segmentation probability.
5. The method according to claim 1, characterized in that, Before the inputting the first motion estimation information and the second motion estimation information into a target segmentation neural network corresponding to the target segmentation style to obtain a target segmentation result, the method further includes: Obtaining a set of segmentation neural networks corresponding to a preset gear in the video coding format AV1; Obtaining the correspondence between the segmentation neural networks in the set of segmentation neural networks and the segmentation styles in the set of segmentation styles; According to the correspondence, obtaining the target segmentation neural network corresponding to the target segmentation style from the set of segmentation neural networks.
6. The method according to claim 1, characterized in that, When the target image is a block image, the method further includes: Obtaining the segmentation style information of adjacent block images of the block image; The inputting the first motion estimation information and the second motion estimation information into the target segmentation neural network to obtain a target segmentation result includes: Input the segmentation style information, the first motion estimation information, and the second motion estimation information into the target segmentation neural network to obtain a target segmentation result.
7. The method according to claim 1, characterized in that The step of performing a whole-pixel search on the target image and the target segmentation block image respectively in the reference image to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmentation block image and the reference image includes: In the reference image, perform a whole-pixel search on the target image and the target segmentation block image respectively to determine the first target matching block in the reference image that matches the target image and the second target matching block in the reference image that matches the target segmentation block image; Obtain the first pixel variance information and the first sum of squared pixel difference information between the target image and the first target matching block; Obtain the second pixel variance information and the second sum of squared pixel difference information between the target segmentation block image and the second target matching block; Use the first pixel variance information and the first sum of squared pixel difference information as the first motion estimation information; Use the second pixel variance information and the second sum of squared pixel difference information as the second motion estimation information.
8. An image segmentation device, characterized in that, Includes: A target image and segmentation style set acquisition module for acquiring a target image and a segmentation style set; A target segmentation style extraction module for extracting a target segmentation style from the segmentation style set; A target segmentation block image acquisition module for obtaining a target segmentation block image based on the target segmentation style and the target image; A motion estimation information acquisition module for performing a whole-pixel search on the target image and the target segmentation block image respectively in the reference image to obtain the first motion estimation information between the target image and the reference image and the second motion estimation information between the target segmentation block image and the reference image; A target segmentation result acquisition module for inputting the first motion estimation information and the second motion estimation information into the target segmentation neural network corresponding to the target segmentation style to obtain a target segmentation result; The target segmentation result represents the probability that the target image is segmented according to the target segmentation style; An image segmentation processing module for performing segmentation processing on the target image according to the target segmentation result; Wherein, the image segmentation processing module is further configured to determine a segmentation style for segmentation processing according to the target segmentation result, and perform segmentation processing on the target image based on the segmentation style for segmentation processing.
9. The device according to claim 8, wherein The target segmentation style extraction module includes: A priority acquisition unit for acquiring the priority corresponding to the segmentation style in the segmentation style set; A target segmentation style extraction unit for extracting the target segmentation style from the segmentation style set according to the priority corresponding to the segmentation style.
10. The device according to claim 9, characterized in that, The image segmentation processing module includes: A segmentation threshold acquisition unit for acquiring the segmentation threshold corresponding to the target segmentation neural network; A first image segmentation processing unit, configured to perform segmentation processing on the target image according to the target segmentation style when the target segmentation result is greater than the segmentation threshold value.
11. The device according to claim 8, characterized in that, When the target segmentation result is a segmentation probability; the target segmentation style is any segmentation style in the segmentation style set; the image segmentation processing module includes: A segmentation probability acquisition unit, configured to acquire the segmentation probability corresponding to each segmentation style in the segmentation style set; A target segmentation style acquisition unit, configured to acquire the segmentation style corresponding to the maximum segmentation probability; A second image segmentation processing unit, configured to perform segmentation processing on the target image according to the segmentation style corresponding to the maximum segmentation probability.
12. The device according to claim 8, wherein The apparatus further includes: A segmentation neural network set acquisition module, configured to acquire a segmentation neural network set corresponding to a preset gear in the video coding format AV1; A correspondence relationship acquisition module between the segmentation neural network set and the segmentation style, configured to acquire the correspondence relationship between the segmentation neural networks in the segmentation neural network set and the segmentation styles in the segmentation style set; A target segmentation neural network acquisition module, configured to acquire the target segmentation neural network corresponding to the target segmentation style from the segmentation neural network set according to the correspondence relationship.
13. The device according to claim 8, wherein, When the target image is a block image, the apparatus further includes: A segmentation style information acquisition module, configured to acquire the segmentation style information of adjacent block images of the block image; Correspondingly, the target segmentation result acquisition module includes: A target segmentation result acquisition unit, configured to input the segmentation style information, the first motion estimation information, and the second motion estimation information into the target segmentation neural network to obtain a target segmentation result.
14. The device according to claim 8, characterized in that, The motion estimation information acquisition module includes: A matching block acquisition unit, configured to perform full-pixel search on the target image and the target segmentation block image respectively in the reference image, and determine a first target matching block in the reference image that matches the target image and a second target matching block in the reference image that matches the target segmentation block image; A first pixel information acquisition unit, configured to acquire first pixel variance information and first pixel difference sum of squares information between the target image and the first target matching block; A second pixel information acquisition unit, configured to acquire second pixel variance information and second pixel difference sum of squares information between the target segmentation block image and the second target matching block; A first motion estimation information acquisition unit, configured to use the first pixel variance information and the first pixel difference sum of squares information as the first motion estimation information; A second motion estimation information acquisition unit, configured to use the second pixel variance information and the second pixel difference sum of squares information as the second motion estimation information.
15. An electronic device, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 7.
16. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
17. A computer program product, characterized in that, including computer instructions which, when executed by a processor, cause a computer to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target detection method, device, apparatus, and storage medium for continuous images
CN109272509A
Image recognition method and device based on neural network and computer equipment
CN111709422A