Scanning Method, Device, Electronic Device, Storage Medium, and Program Product

By performing content detection and pixel classification on the image, the third area is obtained and scanned, which solves the problem of low scanning efficiency in the prior art and realizes fast and efficient scanning image acquisition.

CN115116065BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210622323.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-05-30
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

The prior art consumes a lot of computing power when obtaining scanned copies of paper documents through photos, resulting in low scanning efficiency.

Method used

By acquiring images, performing content detection and pixel classification, an area composed of pixels of key types is determined, and region expansion processing is performed based on these pixels to obtain a third area and scan it.

Benefits of technology

It realizes rapid acquisition of scanned images, improves scanning efficiency, reduces complex computing processes, and improves the robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116065B_ABST
    Figure CN115116065B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a scanning method, apparatus, electronic device, storage medium, and program product. Embodiments of the present application can obtain an image, perform content detection on the image to obtain a first region where the content to be measured is located in the image, perform pixel classification on the image to obtain the types of pixels, and determine a region composed of all pixels of key types as a second region. Based on the pixels of key types in the second region, perform region expansion processing on the first region to obtain a third region, and scan the third region to obtain a scanned image of the third region. In the embodiments of the present application, in order to make the scanned image have complete file content, the third region can include the entire file and the file content in the file, and when scanning the third region, the irrelevant background between the image and the third region can be removed, and this scanning process does not involve complex calculation processes. Thus, this solution can improve the scanning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and particularly to a scanning method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] In recent years, with the increasing number of application scenarios for scanned documents, using a printer to scan paper documents has increased the difficulty of obtaining scanned documents. Therefore, currently, users often use their mobile phones to take pictures of paper documents and use scanning software to remove the background in the photo other than the paper document, enabling users to conveniently obtain the scanned document corresponding to the paper document.

[0003] However, currently, obtaining the scanned document corresponding to a paper document from a photo requires a large amount of computing power, resulting in low scanning efficiency. Summary of the Invention

[0004] Embodiments of this application provide a scanning method, apparatus, electronic device, storage medium, and program product, which can improve scanning efficiency.

[0005] Embodiments of this application provide a scanning method, including:

[0006] Obtain an image, where the image includes the content to be measured;

[0007] Perform content detection on the image to obtain a first region where the content to be measured is located in the image;

[0008] Perform pixel classification on the image to obtain the types of pixels, and determine the region composed of all pixels of key types as a second region, where the first region is included in the second region;

[0009] Based on the pixels of key types in the second region, perform region expansion processing on the first region to obtain a third region;

[0010] Scan the third region to obtain a scanned image of the third region.

[0011] Embodiments of this application also provide a scanning apparatus, including:

[0012] An obtaining unit, configured to obtain an image, where the image includes the content to be measured;

[0013] A detection unit, configured to perform content detection on the image to obtain a first region where the content to be measured is located in the image;

[0014] A classification unit, configured to perform pixel classification on the image to obtain the types of pixels, and determine the region composed of all pixels of key types as a second region, where the first region is included in the second region;

[0015] An expansion unit for performing region expansion processing on a first region based on pixels of a key type in a second region to obtain a third region;

[0016] A scanning unit for scanning the third region to obtain a scanned image of the third region.

[0017] In some embodiments, classifying pixels of an image to obtain the types of the pixels, and determining the region composed of all pixels of the key types as the second region includes:

[0018] Performing binarization processing on the image to obtain the pixel values of each pixel in the image;

[0019] Determining the type of the pixel according to the pixel value of the pixel;

[0020] Taking the types of the pixels corresponding to the pixels in the first region as the key types, and determining the region composed of all pixels of the key types as the second region.

[0021] In some embodiments, performing region expansion processing on the first region based on pixels of a key type in the second region to obtain a third region includes:

[0022] Obtaining a preset unit length and the side lengths of each boundary in the first region;

[0023] Dividing the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions;

[0024] Performing region expansion on the first region according to the pixels of the key type in the sub-regions to obtain a third region.

[0025] In some embodiments, the boundary includes a first type of boundary and a second type of boundary, and the first type of boundary and the second type of boundary intersect. Dividing the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions includes:

[0026] Taking the quotient of the side length of the first type of boundary and the preset unit length as the first quantity;

[0027] Taking the quotient of the side length of the second type of boundary and the preset unit length as the second quantity;

[0028] Performing first quantity equal division on the first type of boundary and second quantity equal division on the second type of boundary, thereby dividing the first region into the first quantity multiplied by the second quantity of sub-regions.

[0029] In some embodiments, performing region expansion on the first region according to the pixels of the key type in the sub-regions to obtain a third region includes:

[0030] Determining the sub-regions located on the outermost sides of the first quantity multiplied by the second quantity of sub-regions;

[0031] Based on the pixels of the key type in the outermost sub-region, perform region expansion on the first region to obtain the third region.

[0032] In some embodiments, performing region expansion on the first region based on the pixels of the key type in the outermost sub-region to obtain the third region includes:

[0033] Determine the number of outermost key sub-regions, where the outermost key sub-region is the outermost sub-region composed of pixels of the key type;

[0034] When the number meets the first preset condition, perform region expansion on the first region to obtain the third region.

[0035] In some embodiments, performing content detection on an image to obtain the first region where the content to be detected is located in the image includes:

[0036] Perform convolution processing on the image to obtain the target probability corresponding to the pixel, where the target probability is the probability that the pixel belongs to the content to be detected;

[0037] According to the target probability, determine the loss value of the pixel belonging to the content to be detected;

[0038] Determine the target pixel according to the loss value, where the target pixel is the pixel belonging to the content to be detected;

[0039] According to all the target pixels, determine the first region of the content to be detected in the image.

[0040] In some embodiments, performing convolution processing on the image to obtain the target probability corresponding to the pixel includes:

[0041] Perform convolution processing on the image to obtain multiple feature maps of different sizes;

[0042] Perform feature fusion on the multiple feature maps of different sizes to obtain a fused feature map;

[0043] Perform convolution processing on the fused feature map to obtain the target probability corresponding to the pixel.

[0044] In some embodiments, according to the target probability, determining the loss value of the pixel belonging to the content to be detected includes:

[0045] Perform convolution processing on the image to obtain the threshold corresponding to the pixel;

[0046] According to the difference between the target probability and the threshold, determine the loss value of the pixel belonging to the content to be detected.

[0047] In some embodiments, before classifying the pixels of the image to obtain the type of the pixel, it further includes:

[0048] Determine the proportion of the first region occupied in the image;

[0049] When the proportion meets the second preset condition, perform pixel classification on the image to obtain the types of pixels.

[0050] In some embodiments, after scanning the third region to obtain a scanned image of the third region, it further includes:

[0051] Measure the angle of the content to be measured in the scanned image to obtain an offset angle;

[0052] According to the offset angle, correct the angle of the content to be measured in the scanned image to obtain a new scanned image.

[0053] An embodiment of the present application further provides an electronic device, including a memory storing multiple instructions; a processor loads instructions from the memory to execute the steps in any one of the scanning methods provided by the embodiments of the present application.

[0054] An embodiment of the present application further provides a computer-readable storage medium storing multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the scanning methods provided by the embodiments of the present application.

[0055] An embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps in any one of the scanning methods provided by the embodiments of the present application are implemented.

[0056] An embodiment of the present application can obtain an image including the content to be measured; perform content detection on the image to obtain the first region where the content to be measured is located in the image; perform pixel classification on the image to obtain the types of pixels, and determine the region composed of all pixels of key types as the second region, and the first region is included in the second region; based on the pixels of key types in the second region, perform region expansion processing on the first region to obtain the third region; scan the third region to obtain a scanned image of the third region.

[0057] In this application, content detection can be performed on an image to identify the first region where the content to be measured is located in the image, and pixel classification can be performed on the image to obtain the types of pixels. Among them, the types of pixels include key types. The region composed of pixels of the key type is used as the second region, and the first region is included in the second region, that is, the types of pixels in the first region and the types of pixels in the second region are the same. Thus, it can be known that the content to be measured is in the second region. The region expansion of the first region can be controlled by pixels of the key type, and the expansion stops when it reaches pixels that do not belong to the key type. Thus, the third region can be quickly obtained, and the third region can be scanned to obtain a scanned image of the third region. When scanning the file content in the image by this scanning method, the scanned image can have a complete file and file content, and there is no irrelevant background between the image and the third region in the scanned image. This scanning process does not involve complex calculation processes. Therefore, this application can quickly obtain a scanned image, thereby improving the scanning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0059] Figure 1a is a schematic diagram of the scenario of the scanning method provided by the embodiment of the present application;

[0060] Figure 1b is a schematic diagram of the region in the scanning method provided by the embodiment of the present application;

[0061] Figure 1c is a schematic flowchart of the scanning method provided by the embodiment of the present application;

[0062] Figure 1d is a schematic flowchart of the application of the scanning method provided by the embodiment of the present application in a terminal and a server;

[0063] Figure 2a is a schematic diagram of the application of the scanning method provided by the embodiment of the present application in a translation text scenario;

[0064] Figure 2b is a schematic diagram of the application of the network structure of DBNet provided by the embodiment of the present application in a text detection scenario;

[0065] Figure 3 is a schematic diagram of the structure of the scanning device provided by the embodiment of the present application;

[0066] Figure 4It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0068] An embodiment of the present application provides a scanning method, device, electronic device, storage medium, and program product.

[0069] Among them, the scanning device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC); the server can be a single server or a server cluster composed of multiple servers.

[0070] In some embodiments, the scanning device can also be integrated in multiple electronic devices. For example, the scanning device can be integrated in multiple servers, and the scanning method of the present application is implemented by multiple servers.

[0071] In some embodiments, the server can also be implemented in the form of a terminal.

[0072] For example, referring to Figure 1a , the server can obtain an image from the terminal through the network, and the image includes the content to be measured; perform content detection on the image to obtain the first area where the content to be measured is located in the image; perform pixel classification on the image to obtain the type of the pixel, and determine the area composed of all pixels of the key types as the second area, and the first area is included in the second area; based on the pixels of the key types in the second area, perform area expansion processing on the first area to obtain the third area; scan the third area to obtain a scanned image of the third area, and the server sends the scanned image to the terminal through the network.

[0073] In some embodiments, referring to Figure 1b , after the user takes a photo of a file containing text content, a photo 01 as shown in Figure 1b can be obtained. Among them, the photo 01 includes the file 03, and the text content 02 is displayed on the file 03. To ensure the integrity of the text content 02 when scanning the file, the set scanning area can be set as area 04, so that area 04 completely includes the entire file 03, the text content 02 in the file 03, and the irrelevant background between area 04 and photo 01 can be removed.

[0074] Among them, with reference to Figure 1a and Figure 1b , after obtaining the user's permission or consent, the server can obtain an image (i.e., photo 01) from the terminal through the network. After obtaining the image, the server performs content detection on the image to identify the first area where the content to be measured is located in the image (i.e., the area where the text content 02 is located), and classifies the pixels of the image to obtain the types of the pixels. Among them, the types of the pixels include key types. The area composed of the pixels of the key types is used as the second area (i.e., file 03), and the first area is included in the second area, that is, the types of the pixels in the first area are the same as those in the second area. Thus, it can be known that the content to be measured is in the second area. The area expansion of the first area can be controlled through the pixels of the key types, so that the expansion stops when it reaches the pixels that do not belong to the key types. Thus, the third area (i.e., scanning area 04) can be quickly obtained, and the third area is scanned to obtain a scanned image of the third area. The server sends the scanned image to the terminal through the network.

[0075] When scanning the file content in the image by this scanning method, the scanned image can have a complete file and file content, and there is no irrelevant background between the image and the third area in the scanned image. This scanning process does not involve complex calculation processes. Therefore, this application can quickly obtain a scanned image, thereby improving the scanning efficiency.

[0076] In addition, since this application does not rely on mathematical algorithms and the logical rules used by mathematical algorithms, this application is less likely to experience situations such as crashing due to complex algorithms and has good robustness.

[0077] The following will be described in detail respectively. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.

[0078] Artificial Intelligence (AI) is a technology that uses digital computers to simulate human beings' perception of the environment, acquisition of knowledge, and use of knowledge. This technology can enable machines to have functions similar to human perception, reasoning, and decision-making. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0079] Among them, computer vision (CV) is a technology that uses a computer to replace the human eye to perform operations such as identifying and measuring target images and further processing them. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. For example, image processing technologies such as image coloring and image stroke extraction.

[0080] In this embodiment, a scanning method based on image semantic understanding involving artificial intelligence is provided, as Figure 1c shown, the specific process of this scanning method can be as follows:

[0081] 110. Obtain an image, which includes the content to be measured.

[0082] Among them, the content to be measured is the content waiting to be detected in the image. For example, the content to be measured can be the text waiting to be translated, or the content waiting to be searched, or the content waiting to be recognized, and so on.

[0083] Among them, the image is an image containing the content to be measured. For example, the image can be an image containing the content waiting to be translated, or an image containing the content waiting to be searched, or an image containing the content waiting to be recognized, and so on.

[0084] Among them, the method for obtaining the image can be obtained through the camera on the terminal, or through terminal screenshots, or obtained from the terminal's photo gallery, or an image sent from other terminals, and so on.

[0085] For example, when the user needs to translate the content to be measured, the terminal takes a photo of the file containing the content to be measured, thereby obtaining an image including the content to be measured waiting to be translated.

[0086] 120. Perform content detection on the image to obtain the first region where the content to be measured is located in the image.

[0087] Among them, content detection is used to detect the content to be measured in the image. For example, content detection can be to detect text, animals, items, etc. in the image.

[0088] Among them, the first region is the region where the content to be measured is located in the image. For example, the first region can be the region where the text to be measured is located in the image, or the region where the target to be measured is located in the image, and the target can be an animal, an item, etc.

[0089] For example, the image includes text to be measured. By performing content detection on the image, the text to be measured in the image can be recognized through a text box, and the area enclosed by the text box is the first area.

[0090] In some embodiments, in order to recognize the content to be measured in the image, content detection is performed on the image to obtain the first area where the content to be measured is located in the image, including:

[0091] 121. Perform convolution processing on the image to obtain the target probability corresponding to the pixel, where the target probability is the probability that the pixel belongs to the content to be measured.

[0092] Among them, the pixel is the basic unit that constitutes the image. For example, the image can be composed of multiple pixels, and some pixels in the image form the content to be measured.

[0093] Among them, the target probability is the probability that the pixel belongs to the content to be measured. For example, the target probability can be the probability that the pixel belongs to the text to be measured, or the probability that the pixel belongs to the target to be measured, and so on.

[0094] For example, after performing convolution processing on the image, the target probability corresponding to each pixel in the image is obtained, and the target probability can be 0.1, 0.7, 0.9, and so on.

[0095] In some embodiments, in order to recognize the probability that each pixel in the image belongs to the content to be measured, convolution processing is performed on the image to obtain the target probability corresponding to the pixel, including:

[0096] Perform convolution processing on the image to obtain multiple feature maps of different sizes;

[0097] Perform feature fusion on the multiple feature maps of different sizes to obtain a fused feature map;

[0098] Perform convolution processing on the fused feature map to obtain the target probability corresponding to the pixel.

[0099] Among them, the feature map is a set of features extracted from the image, and the feature points in the feature map correspond to the pixels in the image. The feature points include spatial features, color features, and so on of the corresponding pixels.

[0100] Among them, the fused feature map is the feature map corresponding to the fusion of multiple feature maps of different sizes.

[0101] Among them, during feature fusion, according to the positions of the feature points in the feature map, the feature points in the multiple feature maps of different sizes are corresponded and then fused.

[0102] For example, after performing convolution processing on the fused feature map, the target probability corresponding to each pixel in the image can be obtained.

[0103] 122. Determine the loss value of a pixel belonging to the content to be measured according to the target probability.

[0104] Among them, the loss value is the loss of the pixel belonging to the content to be measured. For example, the loss value can be the loss of the pixel belonging to the text to be measured, or the loss of the pixel belonging to the target to be measured, and so on.

[0105] In some embodiments, in order to calculate the loss of a pixel belonging to the content to be measured, determining the loss value of the pixel belonging to the content to be measured according to the target probability includes:

[0106] Perform convolution processing on the image to obtain the threshold corresponding to the pixel;

[0107] Determine the loss value of the pixel belonging to the content to be measured according to the difference between the target probability and the threshold.

[0108] Among them, the threshold is the adaptive binary threshold corresponding to the pixel, and this binary threshold is used to divide the pixels of the image into two categories. For example, the threshold corresponding to the pixel can be 0.7, or 0.3. According to the thresholds corresponding to all pixels in the image, the pixels in the image can be divided into pixels with a threshold of 0.7 and pixels with a threshold of 0.3.

[0109] For example, the threshold corresponding to the pixel belonging to the content to be measured can be 0.3, and the threshold corresponding to the pixel not belonging to the content to be measured can be 0.7.

[0110] The method for determining the loss value of the pixel belonging to the content to be measured according to the difference between the target probability and the threshold is:

[0111] The loss function corresponding to the loss value can be: GEloss = -ylog(f(x)) - (1 - y)log(1 - (f(x));

[0112] Among them, -ylog(f(x)) is the loss value of the pixel belonging to the content to be measured, and -(1 - y)log(1 - (f(x)) is the loss value of the pixel not belonging to the content to be measured.

[0113] Among them, f(x) = B i,j , x = P i,j -T i,j ;

[0114] Among them, B i,j is the value corresponding to participating in the binarization of the pixels in the image, P i,j is the probability that the pixel at the (i, j) position in the image belongs to the content to be measured, and T i,j is the threshold corresponding to the pixel at the (i, j) position in the image.

[0115] If x > 0, then

[0116] If x < 0, then

[0117] where k is a coefficient associated with the loss value, and its value can be determined according to the actual situation. For example, in order to have a larger loss value when classifying pixels incorrectly, k = 50 can be set.

[0118] 123. Determine the target pixels according to the loss value, and the target pixels are the pixels belonging to the content to be measured.

[0119] Among them, the target pixels are the pixels corresponding to the loss value less than the preset loss value, and the preset loss value is used to measure the loss value of the pixels belonging to the content to be measured. For example, the target pixels can be the pixels belonging to the text to be measured, or the pixels belonging to the target to be measured, and so on.

[0120] For example, the preset loss value is 0.3. When the loss value of the pixel belonging to the content to be measured is not greater than 0.3, then the pixel is a target pixel; when the loss value of the pixel belonging to the content to be measured is greater than 0.3, then the pixel does not belong to the target pixels.

[0121] The method for determining the target pixels according to the loss value is:

[0122] Through the above calculation formulas (1) and (2), it can be obtained for binarizing the pixels in the image, so that the obtained value can be used to classify the pixels to determine the target pixels in the image.

[0123] Among them, is the value corresponding to the binarization of the pixel, and through it can indicate that the pixel at the position (i, j) in the image belongs to the content to be measured.

[0124] 124. Determine the first region of the content to be measured in the image according to all the target pixels.

[0125] For example, if the target pixels are the pixels belonging to the text to be measured, the first region of the text to be measured in the image can be obtained through all the target pixels in the image.

[0126] 130. Classify the pixels in the image to obtain the types of the pixels, and determine the region composed of all the pixels of the key types as the second region, and the first region is included in the second region.

[0127] Among them, the type of pixel is used to distinguish the pixels in the image that belong to the carrier of the content to be measured, and the carrier of the content to be measured is used to record the content to be measured. For example, if the carrier of the content to be measured is paper, the type of pixel is used to distinguish the pixels in the image that belong to the paper. If the carrier of the content to be measured is a certificate, the type of pixel is used to distinguish the pixels in the image that belong to the certificate. If the carrier of the content to be measured is an electronic page, the type of pixel is used to distinguish the pixels in the image that belong to the electronic page, and so on.

[0128] Among them, the pixels of the key type are the pixels in the image that belong to the carrier of the content to be measured. For example, if the carrier for recording the content to be measured is paper, the pixels of the key type are the pixels that belong to the paper. If the carrier for recording the content to be measured is a certificate, the pixels of the key type are the pixels that belong to the certificate. If the carrier for recording the content to be measured is an electronic page, the pixels of the key type are the pixels that belong to the electronic page, and so on.

[0129] Among them, the second region is composed of the pixels in the image that belong to the carrier of the content to be measured. For example, the second region can be composed of the pixels of the paper belonging to the text to be measured, or can be composed of the pixels of the certificate belonging to the text to be measured, or can be composed of the pixels of the electronic page belonging to the text to be measured, and so on.

[0130] For example, the image includes a piece of paper on which the content to be measured is recorded. After pixel classification of the image, the pixels belonging to the carrier of the content to be measured can be distinguished from the image, so that the region composed of all the pixels of the key type is determined as the second region. Since the content to be measured is recorded on the paper, the first region where the content to be measured is located in the image is within the second region.

[0131] In some embodiments, in order to distinguish the carrier belonging to the content to be measured from the image, pixel classification is performed on the image to obtain the type of pixel, and the region composed of all the pixels of the key type is determined as the second region, including:

[0132] Perform binarization processing on the image to obtain the pixel value of each pixel in the image;

[0133] Determine the type of pixel according to the pixel value of the pixel;

[0134] Use the type of pixel corresponding to the pixels in the first region as the key type, and determine the region composed of all the pixels of the key type as the second region.

[0135] Among them, the binarization process is to set the gray value of the pixel points on the image to 0 or 255. The gray value is related to the brightness, that is, the gray value is related to the depth of color in the image. The carrier of the content to be measured and the background have different brightness in the image, that is, there is a critical gray value between the gray value of the carrier of the content to be measured in the image and the gray value of the background in the image. The gray value greater than the critical gray value in the image is set as the maximum gray value, and the gray value less than the critical gray value is set as the minimum gray value, so as to realize the binarization process of the image.

[0136] Among them, the pixel value of the pixel is the pixel value after binarization processing. For example, the pixel value of the pixel includes 0 or 255 in the gray value.

[0137] Among them, the gray value of 0 corresponds to one type of pixel, and the gray value of 255 corresponds to another type of pixel.

[0138] Among them, the key type is the type of pixel corresponding to the gray value of the pixel that makes up the content to be measured. For example, if the gray value of the pixel that makes up the content to be measured is 0, then the type of pixel corresponding to the gray value of 0 is the key type; if the gray value of the pixel that makes up the content to be measured is 255, then the type of pixel corresponding to the gray value of 255 is the key type.

[0139] Among them, the second area is composed of pixels of the key type. For example, if the key type of pixel is the pixel with a gray value of 0, then the second area is composed of all pixels with a gray value of 0; if the key type of pixel is the pixel with a gray value of 255, then the second area is composed of all pixels with a gray value of 255.

[0140] In some embodiments, considering that when there are multiple images, the multiple images include images that do not contain the content to be measured and images that contain the content to be measured. In order to avoid excessive processing of images that do not contain the content to be measured, before classifying the pixels of the image to obtain the type of pixel, it further includes:

[0141] Determine the proportion of the first area occupied in the image;

[0142] When the proportion meets the second preset condition, classify the pixels of the image to obtain the type of pixel.

[0143] Among them, the proportion is the proportion of the number of pixels in the first area among all the pixels in the image. For example, when the image does not contain the content to be measured, the proportion of the first area occupied in the image is 0; when the image contains the content to be measured, the proportion of the first area occupied in the image is greater than 0.

[0144] Among them, the second preset condition is used to distinguish the presence of the content to be measured in the image. For example, when the content to be measured is text, if the image does not contain text, the proportion of the first region in the image is 0, that is, pixel classification of the image is not performed. If the image contains text, the proportion of the first region in the image is 0.2, that is, 0.2 is greater than 0, so as to determine that the image contains the content to be measured, and then pixel classification of the image is performed to obtain the type of the pixel.

[0145] For example, as Figure 1d shown, it can also be that the terminal quickly detects the image, that is, through the quick detection, the proportion of the first region in the image can be determined, so as to identify whether there is the content to be measured in the image. After the content to be measured exists, the image is uploaded to the server through the network, and the server scans the image.

[0146] 140. Based on the pixels of the key type in the second region, perform region expansion processing on the first region to obtain the third region.

[0147] Among them, the expansion processing is used to divide the boundary between the background in the image and the carrier of the content to be measured.

[0148] Among them, the boundary of the third region is used to divide the background in the image and the carrier of the content to be measured, and the third region includes the content to be measured.

[0149] For example, in order to obtain the scanned image containing the content to be measured in the image, since the second region includes the first region, based on the pixels of the key type in the second region, perform region expansion processing on the first region, and the boundary between the background and the carrier of the content to be measured can be divided in the image, so as to obtain the third region containing the content to be measured.

[0150] In some embodiments, considering that the pixels of the key type can divide the background in the image and the carrier of the content to be measured, in order to realize the orderly expansion of the first region, based on the pixels of the key type in the second region, perform region expansion processing on the first region to obtain the third region, including:

[0151] 141. Obtain the preset unit length and the side length of each boundary in the first region.

[0152] Among them, the preset unit length is used to measure the side length of the boundary. For example, the preset unit length can be in pixels, or can be in centimeters, etc.

[0153] Among them, the side length of the boundary is the length of the boundary of the first region. For example, the length of the boundary of the first region can be equal to the number of pixels forming the boundary of the first region, and the length of the boundary of the first region can also be in centimeters, etc.

[0154] 142. Divide the first region according to the quotient of the side length of the boundary and the preset unit length to obtain multiple sub-regions.

[0155] Among them, the quotient of the side length of the boundary and the preset unit length is the value obtained by dividing the side length of the boundary by the preset unit length. For example, if the length of the boundary is equal to 100 pixels and the preset unit length is equal to 10 pixels, the quotient is equal to 10. If the length of the boundary is equal to 3 centimeters and the preset unit length is equal to 1 centimeter, the quotient is equal to 3, and so on.

[0156] Among them, the multiple sub-regions form the first region.

[0157] For example, if the first region is equal to 200 pixels * 100 pixels and the quotient is equal to 20, then the first region is divided into 10 * 5 sub-regions.

[0158] In some embodiments, in order to be able to divide the first region according to the quotient, the boundary includes a first type of boundary and a second type of boundary, the first type of boundary and the second type of boundary intersect, and divide the first region according to the quotient of the side length of the boundary and the preset unit length to obtain multiple sub-regions, including:

[0159] Take the quotient of the side length of the first type of boundary and the preset unit length as the first quantity;

[0160] Take the quotient of the side length of the second type of boundary and the preset unit length as the second quantity;

[0161] Equally divide the first type of boundary into the first quantity of parts, and equally divide the second type of boundary into the second quantity of parts, so as to divide the first region into the product of the first quantity and the second quantity of sub-regions.

[0162] Among them, the first type of boundary is any one of the boundaries that make up the first region. For example, if the shape of the first region is a rectangle, the first type of boundary can be the long side or the wide side of the first region. If the shape of the first region is a parallelogram, the first type of boundary can be any one of the boundaries that make up the parallelogram, and so on.

[0163] Among them, the side length of the first type of boundary is equal to the first quantity of preset unit lengths. For example, if the side length of the first type of boundary is 100 pixels and the preset unit length is equal to 10 pixels, the first quantity is equal to 10.

[0164] Among them, the second type of boundary is the boundary that intersects the first boundary. For example, if the shape of the first region is a rectangle and the first type of boundary is the long side of the first region, the second type of boundary is the wide side of the first region. If the shape of the first region is a parallelogram and the first type of boundary is one side of the first region, the second type of boundary is the side of the first region that intersects the first type of boundary, and so on.

[0165] Among them, the side length of the second type of boundary is equal to the second quantity of preset unit lengths. For example, if the side length of the second type of boundary is 50 pixels and the preset unit length is equal to 10 pixels, then the second quantity is equal to 5.

[0166] For example, if the first quantity is equal to 10, the first type of boundary is equally divided into 10 parts. If the second quantity is equal to 5, the second type of boundary is equally divided into 5 parts, so as to divide the first region into 10 * 5 sub-regions.

[0167] 143. Expand the first region according to the pixels of the key type in the sub-region to obtain a third region.

[0168] For example, if the pixels of the key type can be pixels with a gray value of 0, the preset unit length is equal to 10 pixels, and the sub-region has 10 * 10 pixels, then expand the first region according to the pixels of the key type among the 10 * 10 pixels in the sub-region to obtain a third region.

[0169] In some embodiments, considering reducing the amount of calculation required for region expansion, expanding the first region according to the pixels of the key type in the sub-region to obtain a third region includes:

[0170] Determine the sub-regions located on the outermost side of the first quantity multiplied by the second quantity of sub-regions;

[0171] Expand the first region based on the pixels of the key type in the outermost sub-region to obtain a third region.

[0172] Among them, the outermost sub-region is the outermost sub-region among multiple sub-regions. For example, if the shape of the first region is a rectangle and the first region is divided into 10 * 5 sub-regions, that is, multiple sub-regions are arranged in 10 columns and 5 rows in the image, the outermost sub-regions include the sub-regions in the first column, the sub-regions in the last column, the sub-regions in the first row, and the sub-regions in the last row.

[0173] For example, if the shape of the first region is a rectangle, the first region can be expanded in the direction where the sub-regions in the first column are located based on the pixels of the key type in the sub-regions in the first column, and the sub-regions in the last column, the sub-regions in the first row, and the sub-regions in the last row are also the same as the sub-regions in the first column above, so that the first region is expanded in the direction where they are located, thereby obtaining a third region.

[0174] In some embodiments, considering that the carrier recording the content to be measured may be inclined to a certain extent in the image, in order to uniformly expand the boundary of the first region when the first region expands to the boundary of the second region, expanding the first region based on the pixels of the key type in the outermost sub-region to obtain a third region includes:

[0175] Determine the number of outermost key sub-regions. The outermost key sub-region is the outermost sub-region composed of pixels of the key type;

[0176] When the number meets the first preset condition, perform region expansion on the first region to obtain the third region.

[0177] Among them, the outermost key sub-region is the outermost sub-region composed of pixels of the key type. For example, if the outermost sub-region consists of 10*10 pixels of the key type, then this outermost sub-region is the outermost key sub-region.

[0178] Among them, the number is the number of outermost key sub-regions on one side of the first region. For example, if the shape of the first region is a rectangle, the number can be the number of outermost key sub-regions in the sub-regions of the first column, or the number of outermost key sub-regions in the sub-regions of the last column, or the number of outermost key sub-regions in the sub-regions of the first row, or the number of outermost key sub-regions in the sub-regions of the last row, and so on.

[0179] Among them, the first preset condition is used to measure the number of outermost key sub-regions on one side of the first region.

[0180] For example, the first preset condition is used to measure the number of outermost key sub-regions in the sub-regions of the first column. The value for measuring the number is equal to N×α, where N is the number of outermost key sub-regions, and α can be any value between (0, 1]. The outermost sub-region is the sub-region of the first column, and the number of sub-regions in the first column is 5, and the number of outermost key sub-regions is 5. Then when the number of outermost key sub-regions is 5 which is greater than 5*0.8, expand the first region in the direction where the sub-regions of the first column are located. The sub-regions of the last column, the first row, and the last row of the first region are also like the sub-regions of the first column, and expand the first region in its corresponding direction to obtain the third region.

[0181] For example, set a preset expansion step length d. When there is a sub-region composed of pixels of the key type in the sub-regions of the first column, then this sub-region is the outermost key sub-region and the count is incremented by one. If there are 5 outermost key sub-regions in the sub-regions of the first column, then the count is 5. If the count of 5 is greater than the number N of the outermost sub-regions multiplied by α, then expand the first region by a length of d in the direction where the sub-regions of the first column are located. Repeat the above expansion steps until the first region cannot be expanded in the direction where the sub-regions of the first column are located, where the value of α is set according to the actual situation.

[0182] 150. Scan the third region to obtain the scanned image of the third region.

[0183] Among them, the scanned image is the image corresponding to the content scanned in the third area. For example, the third area includes the content to be measured and the texture, color, etc. of the carrier recording the content to be measured. At the same time, it can also include part of the background in the image. After scanning the third area, the scanned image can only present the content to be measured in the image, the texture, color of the carrier recording the content to be measured, and part of the background in the image.

[0184] For example, if the value of the above-mentioned α is not 100%, when expanding in the first area, there may be part of the background in the image in the third area, so that in addition to the content to be measured, the texture, color, etc. of the carrier recording the content to be measured in the scanned image, it also includes part of the background in the image.

[0185] In some embodiments, as Figure 1d shown, in order to display the content to be measured without angular deviation, after scanning the third area to obtain the scanned image of the third area, it further includes:

[0186] Measuring the angle of the content to be measured in the scanned image to obtain the offset angle;

[0187] According to the offset angle, correcting the angle of the content to be measured in the scanned image to obtain a new scanned image.

[0188] Among them, the offset angle is the inclination angle of the content to be measured in the scanned image. For example, the offset angle can be that the content to be measured is offset M degrees to the right, or the content to be measured is offset M degrees to the left. M can be any angle.

[0189] There are various ways of angle measurement. For example, the angle measurement function in the graphics algorithm library can be used for angle measurement.

[0190] For example, in some embodiments, the angle measurement can be to obtain the angle of the content to be measured in the image by using Opencv (an open-source graphics algorithm library). The specific method is as follows:

[0191] Preprocess the scanned image. The specific preprocessing includes removing the edges, noise, filtering, binarization, etc. in the scanned image, so that the preprocessed scanned image only shows the content to be measured, and then use the minimum bounding rectangle method (minAreaRect()) of Opencv to operate. This method will return the rotation angle of the minimum bounding rectangle of the content to be measured.

[0192] Among them, the new scanned image is the scanned image where the content to be measured is located after angle correction.

[0193] For example, if the content to be measured in the scanned image is offset 10° to the left, then the content to be measured in the scanned image is offset 10° to the right, so that the content to be measured can be displayed upright in the scanned image and a new scanned image is obtained.

[0194] In some embodiments, such as Figure 1d shown, in order to enhance the clarity of the display of the content to be measured, after scanning the third region to obtain a scanned image of the third region, the following steps are further included:

[0195] Enhance the brightness of the scanned image.

[0196] As can be seen from the above, the embodiments of the present application can obtain an image, which includes the content to be measured; perform content detection on the image to obtain the first region where the content to be measured is located in the image; perform pixel classification on the image to obtain the type of the pixel, and determine the region composed of all pixels of the key types as the second region, and the first region is included in the second region; based on the pixels of the key types, perform region expansion processing on the first region to obtain the third region; scan the third region to obtain a scanned image of the third region.

[0197] In the embodiments of the present invention, content detection can be performed on the image to identify the first region where the content to be measured is located in the image, and pixel classification can be performed on the image to obtain the type of the pixel. Among them, the type of the pixel includes the key type, and the region composed of the pixels of the key type is used as the second region, and the first region is included in the second region, that is, the type of the pixels in the first region is the same as the type of the pixels in the second region. Thus, it can be known that the content to be measured is in the second region. The region expansion of the first region can be controlled by the pixels of the key type, and the expansion stops when it reaches the pixels that do not belong to the key type, so that the third region can be obtained, and the third region can be scanned to obtain a scanned image of the third region. When scanning the file content in the image by this scanning method, the scanned image can have a complete file and file content, and there is no irrelevant background between the image and the third region in the scanned image. This scanning process does not involve complex calculation processes. Therefore, the present application can quickly obtain the scanned image, thereby improving the scanning efficiency.

[0198] According to the method described in the above embodiments, further detailed description will be made below.

[0199] In this embodiment, as Figure 2a shown in the schematic diagram of applying the scanning method in the translation text scenario, taking the content to be measured as the translation text as an example, the method of the embodiments of the present application will be described in detail.

[0200] The specific process of a scanning method is as follows:

[0201] 210. Obtain an image containing the content to be measured.

[0202] 220. Based on the text detection algorithm, obtain the first region where the content to be measured is located in the image.

[0203] For example, when using a text detection algorithm to detect the content to be measured in an image, a text semantic segmentation map is generated. In the text semantic segmentation map, the area inside the minimum bounding rectangle enclosing the text is shown in white, and the rest of the image is black. Based on the text semantic segmentation map, the area of the minimum bounding rectangle in the image can be used as the first area.

[0204] Among them, the text detection algorithm can be (Real-time Scene Text Detection with Differentiable Binarization, DBNet) in the scene text recognition model, or it can be other segmentation-based text detection algorithms.

[0205] For example, as Figure 2b shown in the schematic diagram of the network structure of DBNet applied in the text detection scenario, so that the first area where the content to be measured is located in the image can be obtained.

[0206] 221. DBNet first obtains multiple feature maps of different sizes through a series of Convolutional Neural Network (CNN) operations (Backbone).

[0207] 222. The feature maps are fused through upsampling to obtain a fused feature map. Then, the fused feature map is passed through a convolution operation to obtain the target probability corresponding to the pixel and the threshold corresponding to the pixel.

[0208] 223. According to the target probability corresponding to the pixel and the threshold corresponding to the pixel, determine the loss value of the pixel belonging to the content to be measured.

[0209] 224. Determine the target pixels according to the loss value. The target pixels are the pixels belonging to the content to be measured;

[0210] 225. According to all the target pixels, determine the first area of the content to be measured in the image.

[0211] 230. Perform binarization processing on the image to obtain the pixel value of each pixel in the image.

[0212] 240. Determine the type of the pixel according to the pixel value of the pixel.

[0213] 250. Take the type of the pixel corresponding to the pixel in the first area as the key type, and determine the area composed of all the pixels of the key type as the second area.

[0214] 260. Take the quotient of the side length of the first type of boundary and the preset unit length as the first quantity, and take the quotient of the side length of the second type of boundary and the preset unit length as the second quantity, so as to divide the first area into the first quantity multiplied by the second quantity of sub-areas.

[0215] 270. If all the pixels within the product of the preset unit length and the preset unit length of pixels in the outermost sub-region are pixels of the key type, increment the count by one. If the final count of the outermost sub-region is greater than the number of the outermost sub-regions multiplied by a, expand the first region in the direction where the outermost sub-region is located.

[0216] 280. Repeat step 270 until the first region cannot be expanded in the direction where the outermost sub-region is located, and obtain the third region.

[0217] 290. Scan the third region to obtain the scanned image of the third region.

[0218] As can be seen from the above, this paper proposes a scanning method based on a text detection algorithm for locating the text part in a picture. Specifically, this method first uses the DBNet text detection model to detect the text in the picture. On the basis of obtaining the first region, it uses an image augmentation algorithm based on binarization to optimize the first region. Finally, it obtains the third region and scans to obtain the scanned image of the third region. Through manual evaluation, the accuracy rate of this method is as high as 90%, achieving the expected available effect.

[0219] To better implement the above method, the embodiment of this application also provides a scanning device, which can be specifically integrated in an electronic device. The electronic device can be a device such as a terminal or a server. Among them, the terminal can be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer, or a personal computer; the server can be a single server or a server cluster composed of multiple servers.

[0220] For example, in this embodiment, taking the scanning device being specifically integrated in an electronic device as an example, the method of the embodiment of this application will be described in detail.

[0221] For example, as Figure 3 shown, the scanning device may include an acquisition unit 310, a detection unit 320, a classification unit 330, an expansion unit 340, and a scanning unit 350, as follows:

[0222] (1). Acquisition unit 310.

[0223] The acquisition unit 310 is used to acquire an image, and the image includes the content to be measured.

[0224] (2). Detection unit 320.

[0225] The detection unit 320 is used to perform content detection on the image to obtain the first region where the content to be measured is located in the image.

[0226] In some embodiments, content detection is performed on an image to obtain a first region where the content to be detected is located in the image, including:

[0227] Performing convolution processing on the image to obtain a target probability corresponding to a pixel, where the target probability is the probability that the pixel belongs to the content to be detected;

[0228] Determining a loss value of the pixel belonging to the content to be detected according to the target probability;

[0229] Determining target pixels according to the loss value, where the target pixels are pixels belonging to the content to be detected;

[0230] Determining the first region of the content to be detected in the image according to all the target pixels.

[0231] In some embodiments, performing convolution processing on the image to obtain a target probability corresponding to a pixel, including:

[0232] Performing convolution processing on the image to obtain multiple feature maps of different sizes;

[0233] Performing feature fusion on the multiple feature maps of different sizes to obtain a fused feature map;

[0234] Performing convolution processing on the fused feature map to obtain a target probability corresponding to a pixel.

[0235] In some embodiments, determining a loss value of the pixel belonging to the content to be detected according to the target probability, including:

[0236] Performing convolution processing on the image to obtain a threshold corresponding to a pixel;

[0237] Determining a loss value of the pixel belonging to the content to be detected according to the difference between the target probability and the threshold.

[0238] (III). Classification unit 330.

[0239] The classification unit 330 is configured to perform pixel classification on the image to obtain the type of the pixel, and determine the region composed of all pixels of the key types as the second region, where the first region is included in the second region.

[0240] In some embodiments, performing pixel classification on the image to obtain the type of the pixel, and determining the region composed of all pixels of the key types as the second region, including:

[0241] Performing binarization processing on the image to obtain the pixel value of each pixel in the image;

[0242] Determining the type of the pixel according to the pixel value of the pixel;

[0243] Take the type of the pixel corresponding to the pixel in the first region as the key type, and determine the region composed of all pixels of the key type as the second region.

[0244] In some embodiments, before classifying the pixels of the image to obtain the type of the pixels, it further includes:

[0245] Determine the proportion occupied by the first region in the image;

[0246] When the proportion meets the second preset condition, classify the pixels of the image to obtain the type of the pixels.

[0247] (4). Expansion unit 340.

[0248] The expansion unit 340 is used to perform region expansion processing on the first region based on the pixels of the key type in the second region to obtain a third region.

[0249] In some embodiments, performing region expansion processing on the first region based on the pixels of the key type in the second region to obtain a third region includes:

[0250] Obtain a preset unit length and the side length of each boundary in the first region;

[0251] Divide the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions;

[0252] Perform region expansion on the first region according to the pixels of the key type in the sub-regions to obtain a third region.

[0253] In some embodiments, the boundary includes a first type of boundary and a second type of boundary, and the first type of boundary and the second type of boundary intersect. Dividing the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions includes:

[0254] Take the quotient of the side length of the first type of boundary and the preset unit length as the first quantity;

[0255] Take the quotient of the side length of the second type of boundary and the preset unit length as the second quantity;

[0256] Perform the first quantity equal division on the first type of boundary and the second quantity equal division on the second type of boundary, so as to divide the first region into the first quantity multiplied by the second quantity of sub-regions.

[0257] In some embodiments, performing region expansion on the first region according to the pixels of the key type in the sub-regions to obtain a third region includes:

[0258] Determine the sub-regions located on the outermost side of the first quantity multiplied by the second quantity of sub-regions;

[0259] Based on the pixels of the key type in the outermost sub-region, perform region expansion on the first region to obtain a third region.

[0260] In some embodiments, performing region expansion on the first region based on the pixels of the key type in the outermost sub-region to obtain a third region includes:

[0261] Determine the number of outermost key sub-regions, where the outermost key sub-region is the outermost sub-region composed of pixels of the key type;

[0262] When the number meets the first preset condition, perform region expansion on the first region to obtain a third region.

[0263] (V). Scanning unit 350.

[0264] The scanning unit 350 is configured to scan the third region to obtain a scanned image of the third region.

[0265] In some embodiments, after scanning the third region to obtain a scanned image of the third region, it further includes:

[0266] Measure the angle of the content to be measured in the scanned image to obtain an offset angle;

[0267] According to the offset angle, correct the angle of the content to be measured in the scanned image to obtain a new scanned image.

[0268] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the method embodiments described above, which will not be elaborated here.

[0269] As can be seen from the above, the scanning device in this embodiment obtains an image through the acquisition unit, and the image includes the content to be measured; the detection unit performs content detection on the image to obtain the first region where the content to be measured is located in the image; the classification unit classifies the pixels of the image to obtain the type of the pixels, and determines the region composed of all pixels of the key type as the second region, and the first region is included in the second region; the expansion unit performs region expansion processing on the first region based on the pixels of the key type in the second region to obtain a third region; the scanning unit scans the third region to obtain a scanned image of the third region.

[0270] Thus, the embodiment of the present application can quickly obtain a scanned image and improve the scanning efficiency.

[0271] The embodiments of the present application further provide an electronic device, which may be a device such as a terminal, a server, etc. Among them, the terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, and so on; the server may be a single server or a server cluster composed of multiple servers, and so on.

[0272] In some embodiments, the scanning device may also be integrated in multiple electronic devices. For example, the scanning device may be integrated in multiple servers, and the scanning method of the present application may be implemented by multiple servers.

[0273] In this embodiment, the electronic device of this embodiment will be described in detail by taking the example of a mobile terminal. For example, as Figure 4 shown, it shows a schematic structural diagram of the mobile terminal involved in the embodiments of the present application. Specifically:

[0274] The mobile terminal may include components such as a processor 410 with one or more processing cores, a memory 420 with one or more computer-readable storage media, a power supply 430, an input module 440, and a communication module 450. Those skilled in the art can understand that Figure 4 the mobile terminal structure shown in does not constitute a limitation on the mobile terminal, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0275] The processor 410 is the control center of the mobile terminal, connecting various parts of the entire mobile terminal through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 420, and calling data stored in the memory 420, it executes various functions of the mobile terminal and processes data. In some embodiments, the processor 410 may include one or more processing cores; in some embodiments, the processor 410 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 410 either.

[0276] The memory 420 can be used to store software programs and modules. The processor 410 executes various functional applications and data processing by running the software programs and modules stored in the memory 420. The memory 420 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile terminal. In addition, the memory 420 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 420 can also include a memory controller to provide the processor 410 with access to the memory 420.

[0277] The mobile terminal further includes a power supply 430 for powering each component. In some embodiments, the power supply 430 can be logically connected to the processor 410 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 430 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0278] The mobile terminal may further include an input module 440, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0279] The mobile terminal may further include a communication module 450. In some embodiments, the communication module 450 can include a wireless module. The mobile terminal can perform short-distance wireless transmission through the wireless module of the communication module 450, thereby providing users with wireless broadband Internet access. For example, the communication module 450 can be used to help users send and receive emails, browse web pages, and access streaming media, etc.

[0280] Although not shown, the mobile terminal may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 410 in the mobile terminal will load the executable files corresponding to the processes of one or more application programs into the memory 420 according to the following instructions, and the processor 410 will run the application programs stored in the memory 420 to implement various functions as follows:

[0281] Obtain an image, where the image includes the content to be measured;

[0282] Perform content detection on the image to obtain the first area where the content to be measured is located in the image;

[0283] Classify the pixels of the image to obtain the types of the pixels, and determine the region composed of all the pixels of the key types as the second region, where the second region includes the first region;

[0284] Based on the pixels of the key types in the second region, perform region expansion processing on the first region to obtain the third region;

[0285] Scan the third region to obtain a scanned image of the third region.

[0286] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0287] As can be seen from the above, content detection can be performed on the image to identify the first region where the content to be measured is located in the image, and pixel classification can be performed on the image to obtain the types of the pixels. Among them, the types of the pixels include key types. The region composed of the pixels of the key types is used as the second region, and the second region includes the first region, that is, the types of the pixels in the first region and the types of the pixels in the second region are the same. Thus, it can be known that the content to be measured is in the second region. The region expansion of the first region can be controlled by the pixels of the key types, and the expansion stops when it reaches the pixels that do not belong to the key types. Thus, the third region can be quickly obtained, and the third region can be scanned to obtain a scanned image of the third region. When scanning the file content in the image by this scanning method, the scanned image can have a complete file and file content, and there is no irrelevant background between the image and the third region in the scanned image. This scanning process does not involve complex calculation processes. Therefore, this application can quickly obtain a scanned image, thereby improving the scanning efficiency.

[0288] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0289] Therefore, an embodiment of this application provides a computer-readable storage medium, in which multiple instructions are stored. The instructions can be loaded by a processor to execute the steps in any one of the scanning methods provided by the embodiments of this application. For example, the instructions can execute the following steps:

[0290] Obtain an image, where the image includes the content to be measured;

[0291] Perform content detection on the image to obtain the first region where the content to be measured is located in the image;

[0292] Classify the pixels of the image to obtain the types of the pixels, and determine the region composed of all the pixels of the key types as the second region, where the second region includes the first region;

[0293] Based on the pixels of the key type in the second region, perform region expansion processing on the first region to obtain a third region;

[0294] Scan the third region to obtain a scanned image of the third region.

[0295] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0296] According to one aspect of the present application, there is provided a computer program product, including a computer program / instructions, which are stored in a computer-readable storage medium. The processor of the electronic device reads the computer program / instructions from the computer-readable storage medium, and the processor executes the computer program / instructions, so that the electronic device executes the methods provided in various alternative implementations of the content search aspect in the above embodiments.

[0297] Since the instructions stored in the storage medium can execute the steps in any of the scanning methods provided in the embodiments of the present application, therefore, the beneficial effects that can be achieved by any of the scanning methods provided in the embodiments of the present application can be realized. For details, see the previous embodiments and will not be repeated here.

[0298] The above has introduced in detail a scanning method, device, electronic device, storage medium and program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A scanning method, characterized in that, comprising: acquiring an image, where the image includes the content to be measured; performing content detection on the image to obtain a first region where the content to be measured is located in the image; performing pixel classification on the image to obtain the types of pixels, and determining a region composed of all pixels of key types as a second region, where the first region is included in the second region; performing region expansion processing on the first region based on the pixels of key types in the second region to obtain a third region; The performing region expansion processing on the first region based on the pixels of key types in the second region to obtain a third region includes: acquiring a preset unit length and the side lengths of each boundary in the first region; dividing the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions; performing region expansion on the first region according to the pixels of key types in the sub-regions to obtain a third region; scanning the third region to obtain a scanned image of the third region.

2. The scanning method according to claim 1, characterized in that, The performing pixel classification on the image to obtain the types of pixels, and determining a region composed of all pixels of key types as a second region includes: performing binarization processing on the image to obtain the pixel values of each pixel in the image; determining the type of the pixel according to the pixel value of the pixel; taking the types of the pixels corresponding to the pixels in the first region as key types, and determining a region composed of all pixels of the key types as a second region.

3. The scanning method according to claim 1, characterized in that, the boundary includes a first type of boundary and a second type of boundary, the first type of boundary and the second type of boundary intersect, and the dividing the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions includes: taking the quotient of the side length of the first type of boundary and the preset unit length as a first quantity; taking the quotient of the side length of the second type of boundary and the preset unit length as a second quantity; equally dividing the first type of boundary into the first quantity of parts, and equally dividing the second type of boundary into the second quantity of parts, so as to divide the first region into the first quantity multiplied by the second quantity of sub-regions.

4. The scanning method according to claim 3, characterized in that, The performing region expansion on the first region according to the pixels of key types in the sub-regions to obtain a third region includes: determining the outermost sub-regions among the first quantity multiplied by the second quantity of sub-regions; performing region expansion on the first region based on the pixels of key types in the outermost sub-regions to obtain a third region.

5. The scanning method according to claim 4, characterized in that, The performing region expansion on the first region based on the pixels of key types in the outermost sub-regions to obtain a third region includes: Determine the number of outermost key sub-regions, where the outermost key sub-region is the outermost sub-region composed of pixels of the key type; When the number meets the first preset condition, perform region expansion on the first region to obtain a third region.

6. The scanning method according to claim 1, characterized in that the performing content detection on the image to obtain a first region where the content to be measured is located in the image includes: performing convolution processing on the image to obtain a target probability corresponding to a pixel, where the target probability is the probability that the pixel belongs to the content to be measured; determining a loss value of the pixel belonging to the content to be measured according to the target probability; determining target pixels according to the loss value, where the target pixels are pixels belonging to the content to be measured; determining the first region of the content to be measured in the image according to all the target pixels.

7. The scanning method according to claim 6, characterized in that the performing convolution processing on the image to obtain a target probability corresponding to a pixel includes: performing convolution processing on the image to obtain multiple feature maps of different sizes; performing feature fusion on the multiple feature maps of different sizes to obtain a fused feature map; performing convolution processing on the fused feature map to obtain a target probability corresponding to a pixel.

8. The scanning method according to claim 6, characterized in that the determining a loss value of the pixel belonging to the content to be measured according to the target probability includes: performing convolution processing on the image to obtain a threshold corresponding to a pixel; determining a loss value of the pixel belonging to the content to be measured according to the difference between the target probability and the threshold.

9. The scanning method according to claim 1, characterized in that before the performing pixel classification on the image to obtain the type of the pixel, further includes: determining the proportion of the first region occupied in the image; when the proportion meets the second preset condition, performing pixel classification on the image to obtain the type of the pixel.

10. The scanning method according to claim 1, characterized in that after the performing scanning on the third region to obtain a scanned image of the third region, further includes: performing angle measurement on the content to be measured in the scanned image to obtain an offset angle; correcting the angle of the content to be measured in the scanned image according to the offset angle to obtain a new scanned image.

11. A scanning device, characterized in that includes: an acquisition unit, configured to acquire an image, where the image includes the content to be measured; a detection unit, configured to perform content detection on the image to obtain a first region where the content to be measured is located in the image; a classification unit, configured to perform pixel classification on the image to obtain the type of the pixel, and determine a second region composed of all pixels of the key type, where the second region includes the first region; an expansion unit, configured to perform region expansion processing on the first region based on the pixels of the key type in the second region to obtain a third region; Performing region expansion processing on the first region based on the pixels of the key type in the second region to obtain a third region, including: obtaining a preset unit length and the side lengths of each boundary in the first region; dividing the first region according to the quotient of the side length of the boundary and the preset unit length to obtain a plurality of sub-regions; performing region expansion on the first region according to the pixels of the key type in the sub-regions to obtain a third region; A scanning unit, configured to scan the third region to obtain a scanned image of the third region.

12. An electronic device, characterized in that it includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in the scanning method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the scanning method according to any one of claims 1 to 10.

14. A computer program product, characterized in that it includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps in the scanning method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Target detection method and related device

    CN112989872A