Image segmentation method and device, electronic equipment and storage medium
By standardizing banking business images using the Unet3+ neural network and MSCAM module, and employing scale and channel feature extraction modules for region division and segmentation, the system addresses the issues of low efficiency and error-proneness in bill recognition and document verification in banking operations, achieving higher segmentation accuracy and faster business processing speed.
Patent Information
- Application Number
- CN202511050915.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
In banking operations, especially in key areas such as bill recognition and document verification, there are problems of low efficiency and high error rates, with a high risk of errors due to human intervention.
An image segmentation method is adopted. After standardizing and normalizing the image using the Unet3+ neural network and MSCAM module, the scale feature extraction module and channel feature extraction module are used for region division and segmentation. This enhances the weight of key information regions, suppresses the interference of irrelevant regions, captures the inter-channel dependency, and improves the segmentation accuracy.
It improves the accuracy and efficiency of image segmentation, reduces the need for manual intervention, and enhances the accuracy and speed of banking business processing.
Smart Images

Figure CN120954009A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image segmentation method, apparatus, electronic device, and storage medium. Background Technology
[0002] In the current paradigm of basic banking operations, the application of technology has significantly improved the efficiency of some business processes. However, the efficiency bottlenecks, excessive manual intervention, and resulting error rates inherent in traditional processing models remain prominent. Especially in critical areas such as document recognition, document verification, and signature verification, manual intervention is not only time-consuming but also exacerbates the risk of errors under high workloads. Summary of the Invention
[0003] This invention provides an image segmentation method, apparatus, electronic device, and storage medium to solve the problems of low efficiency and error-proneness when verifying documents, etc.
[0004] According to one aspect of the present invention, an image segmentation method is provided, comprising:
[0005] A first image is determined; the first image is the image obtained after standardizing and normalizing the second image; the second image is the image that records the change information when the resource information is modified;
[0006] The first image is divided into a first region using a first model to obtain a first position. The first model includes a scale feature extraction module and a channel feature extraction module. The scale feature extraction module is used to fuse first information and second information. The first information is the edge contour and texture information of the content within the first image. The second information represents the correlation between the content within the first image. The channel feature extraction module is used to obtain the dependency relationship between each channel of the feature map output by the scale feature extraction module. The first region contains information generated when resource information is modified. The first position is the location of the first region in the first image.
[0007] The first image is segmented based on the first position to obtain a third image.
[0008] According to another aspect of the present invention, an image segmentation apparatus is provided, comprising:
[0009] The first image determination module is used to determine a first image; the first image is an image obtained by standardizing and normalizing a second image; the second image is an image that records change information when resource information is modified;
[0010] A first location determination module is used to divide the first image into a first region using a first model to obtain a first location. The first model is configured with a scale feature extraction module and a channel feature extraction module. The scale feature extraction module is used to fuse first information and second information. The first information is the edge contour and texture information of the content in the first image. The second information represents the correlation between the content in the first image. The channel feature extraction module is used to obtain the dependency relationship between each channel of the feature map output by the scale feature extraction module. The first region contains information generated when resource information is modified. The first location is the position of the first region in the first image.
[0011] The third image determination module is used to segment the first image based on the first position to obtain a third image.
[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image segmentation method according to any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the image segmentation method according to any embodiment of the present invention.
[0017] The technical solution of this invention involves determining a first image, which eliminates differences between images, thereby improving segmentation accuracy. A first model is used to divide the first image into a first region to obtain a first position. The introduction of the first model can handle multi-scale, complex structural, and detailed information contained in the first image, providing a more accurate and detailed first position. The first image is then segmented based on the first position to obtain a third image, improving the segmentation accuracy of the third image. This method extracts key information from the first image through a scale feature extraction module and a channel feature extraction module configured in the first model. It can enhance the weight of key information regions in the first image, suppress irrelevant regions, and capture inter-channel dependencies to recalibrate the weights of each channel. This allows the obtained first position to accurately focus on the region of interest, thereby improving recognition accuracy.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart of an image segmentation method provided in an embodiment of the present invention;
[0021] Figure 2 A structural diagram of a first model provided in an embodiment of the present invention;
[0022] Figure 3 A processing flowchart of an MSCAM module provided in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of an image segmentation device provided in an embodiment of the present invention;
[0024] Figure 5 A schematic diagram of the structure of an electronic device for implementing the image segmentation method of this invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects of operation and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] Figure 1 This is a flowchart illustrating an image segmentation method provided in an embodiment of the present invention. This embodiment is applicable to extracting transaction information from transaction orders generated by banking transactions. The method can be executed by an image segmentation device, which can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities. Figure 1 As shown, the method includes:
[0028] S110. Determine the first image; the first image is the image obtained after standardizing and normalizing the second image; the second image is the image that records the change information when the resource information is modified.
[0029] The second image is an image generated when the object changes the resource information in its resource information account, which can prove its identity and the change record information generated after the resource information is modified.
[0030] For example, the second image can be at least one of the following: a ticket image, or an image of the identity information of the object being operated on.
[0031] Specifically, when an object modifies resource information within its resource information account, a ticket is generated based on the resource modification information, and an image of the object's identity information is obtained. The obtained ticket image and identity information image are used as the second image. After standardization and normalization processing, the second image is converted to a size that the first model can recognize, thus obtaining the first image.
[0032] The above steps, which process the second image, are to eliminate the differences between the images, thereby improving the generalization ability of the first model.
[0033] S120. The first image is divided into a first region using the first model to obtain a first position. The first model is configured with a scale feature extraction module and a channel feature extraction module. The scale feature extraction module is used to fuse first information and second information. The first information is the edge contour and texture information of the content in the first image. The second information represents the relationship between the content in the first image. The channel feature extraction module is used to obtain the dependency relationship between each channel of the feature map output by the scale feature extraction module. The first region contains information generated when the resource information is modified. The first position is the position of the first region in the first image.
[0034] The first model consists of a Unet3+ neural network, a scale feature extraction module, and a channel feature extraction module.
[0035] The Unet3+ neural network mainly consists of an encoder and a decoder. The encoder extracts features from the first image through multiple convolutions and max pooling. The decoder concatenates the detailed features extracted by the encoder at different scales.
[0036] For example, the first model is as follows Figure 2 As shown, an MSCAM module is added between the encoder and decoder in the Unet3+ neural network. The MSCAM module consists of a scale feature extraction module and a channel feature extraction module.
[0037] The scale feature extraction module extracts the first and second information from the first image and then fuses the obtained information.
[0038] The first information includes the edge contour of the text information within the first image, the texture information of the first image, and the color information.
[0039] Furthermore, the first information includes at least one of the following: the edge outline of text information, the texture information of different regions in the first image, and the color information of different content in the first image.
[0040] For example, taking a ticket as an example, the texture of the font on the ticket is different from the texture of the background, and the color and texture of different characters are also different.
[0041] The second information is used to characterize the correlation between the first information and the logical relationship and attribute information between the different contents contained in the first image.
[0042] For example, taking a bill as an example, the text outline representing the amount on the bill is dependent on the string of numbers on the same line. Different strings of numbers represent different attributes, such as line number, amount, or account information for accessing resources.
[0043] The channel feature extraction module is used to generate a channel feature map from the output feature map of the scale feature extraction module, thereby obtaining the importance of features in different channels, and determining the location of the first region based on the obtained importance of features.
[0044] The first area includes at least: information about the first component that modifies resource information and information about the operation object.
[0045] The first component is a platform capable of modifying resource information. For example, the first component can be at least one of the following: a self-service counter and a queuing service counter.
[0046] Specifically, the first model performs two convolutions on the first image using a first preset convolution kernel and passes them to MSCAM modules at different scales for feature refinement. Simultaneously, it performs downsampling through max pooling layers and convolutional layers, and then passes the samples to MSCAM modules at different scales for further downsampling, and so on, until a feature map of detailed features at a preset scale is obtained after downsampling. The obtained feature map is then upsampled and passed to MSCAM modules at different scales for feature refinement. Finally, all the obtained feature maps are combined and output to obtain a comprehensive feature map. Based on the obtained comprehensive feature map, the first region in the first image is labeled to obtain the first position.
[0047] The steps described above, through precise spatial location guidance and the capture of inter-channel dependencies, enhance the feature representation of the first region in the first image and suppress interference from irrelevant regions, thereby improving segmentation accuracy. This helps reduce errors and omissions, and improves the accuracy of business processing.
[0048] S130. The first image is segmented according to the first position to obtain the third image.
[0049] Specifically, a cropping frame is drawn based on the first position, and the first image is cut according to the cropping frame to obtain the third image.
[0050] The above steps, by acquiring the third image, can reduce the need for manual intervention when checking the first image, avoid the bottleneck of manual processing, and greatly speed up the processing speed.
[0051] Optionally, determine the first image, including steps A1-A3:
[0052] Step A1: Determine the second image.
[0053] Specifically, when the object being operated modifies the resource information in its resource information account, corresponding record information is generated based on the modification operation, and the obtained record information is used to generate an image to obtain the second image.
[0054] Furthermore, when the object being operated modifies the resource information, the identity information of the object being operated is simultaneously verified and captured, resulting in an image of the object's identity information, which is then used as a second image.
[0055] Step A2: Standardize and normalize the second image to obtain the fourth image.
[0056] Specifically, the second image is standardized by converting its pixel values to a normal distribution with a mean of 0 and a standard deviation of 1. After standardization, the standardized second image is normalized by scaling its pixel values to between [0,1] and [-1,1] to obtain the fourth image.
[0057] Furthermore, standardization can be achieved through the following formula:
[0058]
[0059] Where σ is the standard deviation of the second image pixels; μ is the mean of the second image pixels; x′ is the standardized pixel value; and x is the pixel value.
[0060] Furthermore, normalization can be achieved using the following formula:
[0061]
[0062] Where x′ is the standardized pixel value; min(x′) is the minimum standardized pixel value; max(x′) is the maximum standardized pixel value; and x″ is the normalized pixel value.
[0063] The above steps, including normalization and standardization, are to eliminate the dimensional differences between different pixel values, accelerate the training convergence of the first model, and improve the stability and generalization ability of the first model.
[0064] Step A3: Adjust the size of the fourth image according to the input size of the first model to obtain the first image.
[0065] Specifically, the input dimensions of the first model are obtained, and the image size of the fourth image is adjusted according to the obtained input dimensions to obtain the first image.
[0066] Optionally, the first image is divided into a first region using the first model to obtain the first position, including steps B1-B4:
[0067] Step B1: Downsample the first image according to the first preset convolution kernel to obtain the first feature.
[0068] Specifically, the first image is convolved twice according to the first preset convolution kernel to obtain the first feature.
[0069] The first preset convolutional kernel size is 3*3.
[0070] Step B2: Extract scale features from the first feature using the scale feature extraction module to obtain the second feature.
[0071] Specifically, the scale feature extraction module extracts encoded features from the first feature and convolves these encoded features using a second preset convolution kernel, using the convolved feature as the sixth detail feature. The scale feature extraction module extracts decoded features from the first feature and convolves them using the second preset convolution kernel. After convolution, the features are activated using a first activation function to obtain the seventh detail feature. The sixth and seventh detail features are concatenated, and then convolved using the second preset convolution kernel. After convolution, the features are activated using a second activation function to obtain spatial location weights. These spatial location weights are multiplied by the encoded features and then concatenated with the decoded features to obtain the second feature.
[0072] Step B3: Extract features from the second feature using the channel feature extraction module to obtain the third feature.
[0073] Specifically, the channel feature extraction module performs two-dimensional average pooling on the second feature, adds the processed detail features element by element, activates them through the second activation function, and then multiplies them element by element with the second feature to obtain the third feature.
[0074] The steps described above, including the use of scale feature extraction and channel feature extraction modules, allow for the dynamic selection and weighting of features at different scales during the decoding stage. This is more flexible than static feature fusion methods, better adapting to objects and details at different scales. It also allows for greater focus on key regions and details in the image, especially in areas with complex edges and textures. Furthermore, it better handles scale variations and occlusion issues, improving segmentation performance in complex scenes. Simultaneously, it facilitates more effective feature fusion between feature maps of different scales, which is crucial for capturing a balance between global and local features.
[0075] Step B3: Mark the first region from the first image based on the third feature and determine the first position.
[0076] Specifically, the content within the first image is filtered and labeled based on the third feature to obtain the first region. The first position is determined based on the position of the first region in the first image.
[0077] Optionally, the first image is downsampled according to the first preset convolution kernel to obtain the first feature, including steps C1-C4:
[0078] Step C1: Perform shallow feature extraction on the first image according to the first preset convolution kernel to obtain the first detail features; the shallow features are the contour, texture and color features in the first image.
[0079] Specifically, the first image is convolved twice according to the first preset convolution kernel to extract the contour, texture and color features in the first image to obtain the first detail features.
[0080] Step C2: The first detail feature is downsampled by a pooling layer and convolved by a first preset convolution kernel to obtain the second detail feature.
[0081] Specifically, the first detail feature is downsampled by a max pooling layer, and then convolved twice by a first preset convolution kernel to obtain the second detail feature.
[0082] Step C3: Downsample the second detail feature and convolve it using the first preset convolution kernel, then upsample it using the first preset convolution kernel to obtain the third detail feature.
[0083] Specifically, the second detail feature is first downsampled using a max pooling layer and a first preset convolutional kernel. After downsampling, it is upsampled using the first preset convolutional kernel to restore the scale and obtain the third detail feature.
[0084] Step C4: Generate the first feature based on the first detail feature, the second detail feature, and the third detail feature.
[0085] Specifically, the third detail feature and the second detail feature are combined to obtain the first intermediate feature. The first intermediate feature and the first detail feature are then processed by the scale feature extraction module and the channel feature extraction module for scale feature extraction. A downsampling convolution is then performed using a first preset convolution kernel. After the downsampling convolution, an upsampling convolution is performed using the first preset convolution kernel to obtain the upsampled detail feature. The first detail feature, the upsampled detail feature, and the third detail feature are then combined to obtain the second intermediate feature. The first intermediate feature and the second intermediate feature are used as the first feature.
[0086] Furthermore, when at least two downsampling operations are required, the downsampling can be superimposed using the steps described above.
[0087] For example, such as Figure 2As shown, downsampling was performed through four rounds of pooling and convolution. Assume that after the first convolution, feature 1 is obtained; after one more pooling and convolution, feature 1 becomes feature 2; after another pooling and convolution, feature 2 becomes feature 3; after another pooling and convolution, feature 3 becomes feature 4; and after another pooling and convolution, feature 4 is upsampled to obtain feature 5. As shown in the figure, features 5, 4, 3, 2, and 1 are combined to obtain the first intermediate feature. The first intermediate feature is then extracted using the MSCAM module and convolved with a first preset convolution kernel to obtain feature 6. Features 6, 5, 3, 2, and 1 are then combined to obtain the second intermediate feature. The second intermediate feature is then extracted using the MSCAM module and convolved with a first preset convolution kernel to obtain feature 7. Features 7, 6, 5, 2, and 1 are then combined to obtain the third intermediate feature. The third intermediate feature is then extracted using the MSCAM module and convolved with a first preset convolution kernel to obtain feature 8. Features 8, 7, 6, 5, and 1 are then combined to obtain the fourth intermediate feature. Finally, the first, second, third, and fourth intermediate features are combined to obtain the first feature.
[0088] Optionally, the first feature is generated based on the first detail feature, the second detail feature, and the third detail feature, including steps D1-D3:
[0089] Step D1: Combine the second and third detail features to obtain the fourth detail feature.
[0090] Specifically, the second and third detail features are scale-balanced, that is, their scales are converted to the same scale, and then they are spliced together to obtain the fourth detail feature.
[0091] Step D2: Extract features from the fourth detail feature using the scale feature extraction module and the channel feature extraction module, and then perform convolution to obtain the fifth detail feature.
[0092] Specifically, the fourth detail feature is input into the scale feature extraction module and the channel feature extraction module for detail feature extraction. The obtained detail feature is downsampled by the first preset convolution kernel and then upsampled once to obtain the fifth detail feature.
[0093] Step D3: Determine the first feature based on the first detailed feature, the third detailed feature, and the fifth detailed feature.
[0094] Specifically, the first, third, and fifth detail features are scale-balanced, and after scale transformation, they are spliced together to obtain the first feature.
[0095] Optionally, scale feature extraction is performed on the first feature using the scale feature extraction module to obtain the second feature, including steps E1-E5:
[0096] Step E1: Obtain the encoded features from the first feature and perform convolution on the encoded features to obtain the sixth detail feature.
[0097] The encoded features are the features obtained after downsampling the first image through the first preset convolution kernel each time.
[0098] Specifically, the scale feature extraction module extracts the encoded features from the first feature, and convolves the encoded features according to the second preset convolution kernel, and uses the convolved features as the sixth detail features.
[0099] The second preset convolution kernel is a 1*1 convolution kernel.
[0100] Step E2: Obtain the decoding features from the first features, and perform convolution and activation of the decoding features using the first activation function to obtain the seventh detail features.
[0101] The decoding features are the features of the first image upsampled each time through the first preset convolution kernel and / or the features output by the next-level scale feature extraction module and the convolution calculation.
[0102] Specifically, the scale feature extraction module extracts decoding features from the first feature and performs convolution through the second preset convolution kernel. After convolution, it is activated by the first activation function to obtain the seventh detail feature.
[0103] The first activation function is the ReLU activation function (Rectified Linear Unit), defined as: f(x) = max(θ, x). When the input value is greater than 0, the output is the input value; when the input value is less than or equal to 0, the output is 0.
[0104] Step E3: Concatenate the sixth and seventh detail features, and then perform convolution and activation with the second activation function to obtain the eighth detail feature.
[0105] Specifically, the obtained sixth and seventh detail features are concatenated, and then convolved sequentially through the second preset convolution kernel. After convolution, the features are activated by the second activation function to obtain spatial location weights, which are then used as the eighth detail feature.
[0106] The second activation function is the sigmoid activation function, which can map the convolutional features to the range (0,1).
[0107] Among them, spatial location weights are used to characterize the interdependencies of detailed features within spatial units after convolution.
[0108] Step E4: Multiply the eighth detail feature with the encoded feature to obtain the ninth detail feature.
[0109] Specifically, the obtained eighth detail feature is multiplied with the coding feature to apply the eighth detail feature to the coding feature, thus obtaining the ninth detail feature.
[0110] Step E5: Concatenate the ninth detail feature with the decoded feature to obtain the second feature.
[0111] Specifically, following the order of the ninth detail feature first and the decoding feature second, the ninth detail feature and the decoding feature are concatenated to obtain the second feature.
[0112] For example, such as Figure 3 As shown, the scale feature extraction module extracts encoded features. The sixth detail feature is obtained by convolution using a 1x1 kernel; the decoded features are then processed. The seventh detail feature is obtained by performing convolution using a 1x1 kernel and then activating it using a ReLU activation function. The sixth and seventh detail features are then concatenated and sequentially convolved using a 1x1 kernel, followed by activation using a sigmoid activation function to obtain the eighth detail feature, g. en The eighth detail feature g en With coding features Multiplying the products yields the ninth detail feature x′. en The ninth detail feature x′ en With decoding features By splicing the components together, the second feature is obtained.
[0113] Furthermore, the calculation process of the scale feature extraction module can be represented as follows:
[0114]
[0115] Where ω1, ω2, and ω3 are 1*1 convolutions; C represents the concatenation operation. σ is the ReLU activation function; σ is the sigmoid activation function. For element-wise multiplication; C3 = C1 + C2 is the dimension of the second feature.
[0116] The above steps enable the scale feature extraction module to enhance the weight of the first region and suppress the influence of irrelevant regions in the first image, thereby improving the accuracy of recognition.
[0117] Optionally, the second feature is extracted using the channel feature extraction module to obtain the third feature, including steps F1-F3:
[0118] Step F1: Perform two-dimensional average pooling on the second feature to obtain the tenth detail feature.
[0119] Specifically, the channel feature extraction module performs two-dimensional average pooling on the second feature to obtain the tenth detail feature in C3×1×1 dimensions.
[0120] Among them, the tenth detailed feature includes features representing global context information. Features that effectively remove invalid information
[0121] The two-dimensional average pooling process includes global average pooling and global max pooling.
[0122] Step F2: Add the tenth detail features and activate them using the second activation function to obtain the eleventh detail feature.
[0123] Specifically, the tenth detail feature is added element by element and activated by the second activation function to obtain the eleventh detail feature.
[0124] Step F3: Multiply the second feature and the eleventh detail feature to obtain the third feature.
[0125] Specifically, the second feature and the eleventh detail feature are multiplied element by element to obtain the third feature.
[0126] For example, such as Figure 3 As shown, the channel feature extraction module extracts the second feature. Two-dimensional average pooling is performed to obtain the tenth detail feature (C3×1×1 dimensional). The spatial information in the tenth detail feature is then summed element-wise and activated using a second activation function to obtain the eleventh detail feature. The second and eleventh detail features are then multiplied element-wise to obtain the third feature F. out .
[0127] Furthermore, the calculation process of the channel feature extraction module can be represented as follows:
[0128]
[0129] Where ω1, ω2, and ω0 are 1*1 convolutions; σ is the sigmoid activation function; For element-wise multiplication; This is an element-wise addition.
[0130] In the above steps, the channel feature extraction module adaptively models multi-scale features along the channel dimension and recalibrates the weights of each channel, thereby improving the expressive power of the features.
[0131] The technical solution of this embodiment involves determining a first image, which eliminates differences between images, thereby improving the accuracy of segmentation. A first model is used to divide the first image into a first region to obtain a first position. The introduction of the first model can handle multi-scale, complex structural, and detailed information contained in the first image, providing a more accurate and detailed first position. The first image is then segmented based on the first position to obtain a third image, improving the accuracy of the third image segmentation. This method extracts key information from the first image through the scale feature extraction module and channel feature extraction module configured in the first model, enhancing the weight of key information regions in the first image, suppressing irrelevant regions, and capturing inter-channel dependencies to recalibrate the weights of each channel. This allows the obtained first position to accurately focus on the region of interest, thereby improving the accuracy of recognition.
[0132] Figure 4 This is a schematic diagram of an image segmentation device provided in an embodiment of the present invention. This embodiment is applicable to extracting transaction information from transaction slips generated by banking transactions. The image segmentation device can be implemented in hardware and / or software, and can be configured in any electronic device with network communication capabilities. Figure 4 As shown, the device includes: a first image determination module 210, a first position determination module 220, and a third image determination module 230, wherein:
[0133] First image determination module 210: used to determine a first image; the first image is an image obtained after standardizing and normalizing the second image; the second image is an image that records change information when resource information is modified;
[0134] First position determination module 220: used to divide the first image into a first region using a first model to obtain a first position; the first model is configured with a scale feature extraction module and a channel feature extraction module; the scale feature extraction module is used to fuse first information and second information; the first information is the edge contour and texture information of the content in the first image; the second information represents the correlation between the content in the first image; the channel feature extraction module is used to obtain the dependency relationship between each channel of the feature map output by the scale feature extraction module; the first region contains information generated when the resource information is modified; the first position is the position of the first region in the first image;
[0135] Third image determination module 230: used to segment the first image according to the first position to obtain the third image.
[0136] Optionally, the first image determination module 210 includes:
[0137] Second image determination unit: used to determine the second image;
[0138] Fourth image determination unit: used to standardize and normalize the second image to obtain the fourth image;
[0139] First image determination unit: used to adjust the size of the fourth image according to the input size of the first model to obtain the first image.
[0140] Optionally, the first position determination module 220 includes:
[0141] First feature determination unit: used to downsample the first image according to the first preset convolution kernel to obtain the first feature;
[0142] The second feature determination unit is used to extract scale features from the first feature based on the scale feature extraction module to obtain the second feature.
[0143] The third feature determination unit is used to extract features from the second feature based on the channel feature extraction module to obtain the third feature;
[0144] First position determination unit: used to mark a first region from the first image based on a third feature and determine the first position.
[0145] Optionally, the first feature determining unit includes:
[0146] First detail feature determination subunit: used to perform shallow feature extraction on the first image according to the first preset convolution kernel to obtain the first detail feature; the shallow feature is the contour, texture and color feature in the first image;
[0147] The second detail feature determination subunit is used to downsample the first detail feature through a pooling layer and convolve it through a first preset convolution kernel to obtain the second detail feature.
[0148] The third detail feature determination subunit is used to downsample the second detail feature and convolve it through the first preset convolution kernel, and then upsample it through the first preset convolution kernel to obtain the third detail feature;
[0149] First feature determination subunit: used to generate the first feature based on the first detailed feature, the second detailed feature and the third detailed feature.
[0150] Optionally, the first feature determines the sub-unit, specifically used for:
[0151] The second and third detailed features are combined to obtain the fourth detailed feature;
[0152] The fourth detail feature is extracted using the scale feature extraction module and the channel feature extraction module, and then convolution is performed to obtain the fifth detail feature.
[0153] The first feature is determined based on the first detailed feature, the third detailed feature, and the fifth detailed feature.
[0154] Optionally, the second feature determining unit includes:
[0155] The sixth detail feature determination subunit is used to obtain the encoded features from the first feature and to convolve the encoded features to obtain the sixth detail feature.
[0156] The seventh detail feature determination subunit is used to obtain the decoding features from the first features, and to perform convolution and activation of the decoding features by the first activation function to obtain the seventh detail features;
[0157] The eighth detail feature determination subunit is used to concatenate the sixth and seventh detail features, and then perform convolution and activation with the second activation function to obtain the eighth detail feature.
[0158] The ninth detail feature determination subunit is used to multiply the eighth detail feature with the encoded feature to obtain the ninth detail feature;
[0159] The second feature determination subunit is used to concatenate the ninth detail feature with the decoding feature to obtain the second feature.
[0160] Optionally, the third feature determining unit includes:
[0161] The tenth detail feature determination subunit is used to perform two-dimensional average pooling on the second feature to obtain the tenth detail feature.
[0162] The eleventh detail feature determination subunit is used to add the tenth detail feature and activate it through the second activation function to obtain the eleventh detail feature;
[0163] The third feature determines the sub-unit: it is used to multiply the second feature and the eleventh detail feature to obtain the third feature.
[0164] The image segmentation apparatus provided in the embodiments of the present invention can execute the image segmentation method provided in any of the embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the image segmentation method. For details, please refer to the relevant operations of the image segmentation method in the foregoing embodiments.
[0165] Figure 5This is a schematic diagram of the structure of an electronic device for implementing the image segmentation method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0166] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0167] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0168] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image segmentation methods.
[0169] In some embodiments, the image segmentation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image segmentation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image segmentation method by any other suitable means (e.g., by means of firmware).
[0170] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0171] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0172] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0174] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0175] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0176] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0177] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An image segmentation method, characterized in that, include: A first image is determined; the first image is the image obtained by standardizing and normalizing the second image. The second image is an image that records the changes when resource information is modified; The first image is divided into a first region using the first model to obtain the first position; The first model is equipped with a scale feature extraction module and a channel feature extraction module; The scale feature extraction module is used to fuse the first information and the second information; the first information is the edge contour and texture information of the content in the first image; the second information represents the correlation between the content in the first image; the channel feature extraction module is used to obtain the dependency relationship between each channel of the feature map output by the scale feature extraction module; the first region contains information generated when the resource information is modified; the first position is the position of the first region in the first image. The first image is segmented based on the first position to obtain a third image.
2. The method according to claim 1, characterized in that, Determining the first image includes: Determine the second image; The second image is standardized and normalized to obtain the fourth image; The fourth image is resized according to the input dimensions of the first model to obtain the first image.
3. The method according to claim 1, characterized in that, The step of dividing the first image into a first region using a first model to obtain a first position includes: The first image is downsampled according to the first preset convolution kernel to obtain the first feature; The scale feature extraction module extracts scale features from the first feature to obtain the second feature. The second feature is extracted using the channel feature extraction module to obtain the third feature; The first region is marked from the first image based on the third feature, and the first position is determined.
4. The method according to claim 3, characterized in that, The step of downsampling the first image according to the first preset convolution kernel to obtain the first feature includes: The first image is subjected to shallow feature extraction based on the first preset convolution kernel to obtain the first detail feature; the shallow feature is the contour, texture and color feature in the first image; The first detailed feature is downsampled by a pooling layer and then convolved by a first preset convolution kernel to obtain the second detailed feature. The second detailed feature is downsampled and convolved through a first preset convolution kernel, and then upsampled through the first preset convolution kernel to obtain the third detailed feature; A first feature is generated based on the first detailed feature, the second detailed feature, and the third detailed feature.
5. The method according to claim 4, characterized in that, The step of generating the first feature based on the first detail feature, the second detail feature, and the third detail feature includes: The second and third detailed features are combined to obtain the fourth detailed feature; The fourth detail feature is extracted using the scale feature extraction module and the channel feature extraction module, and then convolved to obtain the fifth detail feature. The first feature is determined based on the first detailed feature, the third detailed feature, and the fifth detailed feature.
6. The method according to claim 3, characterized in that, The step of extracting scale features from the first feature using the scale feature extraction module to obtain the second feature includes: The encoded features are obtained from the first feature, and the encoded features are convolved to obtain the sixth detail feature; Decoding features are obtained from the first features, and the decoding features are convolved and activated by the first activation function to obtain the seventh detail features; The sixth and seventh detail features are concatenated and then convolved and activated by the second activation function to obtain the eighth detail feature; The eighth detail feature is multiplied by the encoded feature to obtain the ninth detail feature; The ninth detail feature is concatenated with the decoded feature to obtain the second feature.
7. The method according to claim 3, characterized in that, The step of extracting features from the second feature using the channel feature extraction module to obtain the third feature includes: The second feature is subjected to two-dimensional average pooling to obtain the tenth detail feature; The tenth detail feature is added together and activated by the second activation function to obtain the eleventh detail feature; The second feature and the eleventh detailed feature are multiplied together to obtain the third feature.
8. An image segmentation apparatus, characterized in that, include: The first image determination module is used to determine the first image; The first image is the image obtained after standardizing and normalizing the second image; The second image is an image that records the changes when resource information is modified; The first location determination module is used to divide the first image into a first region using a first model to obtain a first location; The first model is equipped with a scale feature extraction module and a channel feature extraction module; The scale feature extraction module is used to fuse the first information and the second information; the first information is the edge contour and texture information of the content in the first image; the second information represents the correlation between the content in the first image; the channel feature extraction module is used to obtain the dependency relationship between each channel of the feature map output by the scale feature extraction module; the first region contains information generated when the resource information is modified; the first position is the position of the first region in the first image. The third image determination module is used to segment the first image based on the first position to obtain a third image.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image segmentation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image segmentation method according to any one of claims 1-7.