Paper Localization Method and Device Based on Unsupervised Learning

Through the unsupervised learning self-organizing neural network method, the outline of paper in the picture is extracted, which solves the problem of inaccurate paper positioning in the prior art, and improves the robustness and accuracy of text recognition.

CN114723930BActive Publication Date: 2025-06-24INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210344286.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-02
Publication Date
2025-06-24
Estimated Expiration
2042-04-02

AI Technical Summary

Technical Problem

In the prior art, in text recognition scenarios, it is difficult to accurately locate paper with inclined, folded or messy backgrounds, resulting in poor robustness of text recognition models.

Method used

Using an autonomous neural network method based on unsupervised learning, the two-dimensional pixel value array of the picture is input to the input layer of the autonomous neural network, and the outline of the paper image in the picture is determined through the two-dimensional topology of the output layer.

Benefits of technology

It realizes accurate positioning of paper in the picture, improves the accuracy of subsequent text recognition, and does not require labeling data and high hardware requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114723930B_ABST
    Figure CN114723930B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method and device for paper positioning based on unsupervised learning, which can be used in the financial field or other technical fields. The method includes: inputting a two-dimensional pixel value array of a picture containing a paper image into the input layer of a self-organizing neural network; when the state function of the neurons in the output layer of the self-organizing neural network reaches a convergence condition, obtaining the two-dimensional topological structure of the output layer; and determining the contour of the paper image in the picture according to the two-dimensional topological structure. The present invention achieves the beneficial effect of accurately positioning the paper in the picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a method and device for paper positioning based on unsupervised learning. Background Art

[0002] In many text recognition scenarios, the paper containing text in the recognized image usually has the situations of being inclined, folded or having a cluttered background, resulting in generally poor robustness of the text recognition model in practical applications. Therefore, it is necessary to locate the paper in the image before the model recognizes the text, so as to perform affine transformation. However, factors such as background and shadow in the actual scenario will greatly interfere with the positioning of the paper, so the correct affine transformation cannot be performed, and thus the preprocessing process cannot be of any help to the subsequent recognition.

[0003] It can be seen that the prior art lacks an effective solution for positioning the paper in the image. Summary of the Invention

[0004] In order to solve at least one of the above technical problems in the background art, the present invention proposes a method and device for paper positioning based on unsupervised learning.

[0005] To achieve the above object, according to one aspect of the present invention, there is provided a method for paper positioning based on unsupervised learning, the method comprising:

[0006] Inputting a two-dimensional pixel value array of a picture containing a paper image into an input layer of a self-organizing neural network;

[0007] When the state function of the neurons in the output layer of the self-organizing neural network reaches a convergence condition, obtaining a two-dimensional topological structure of the output layer;

[0008] Determining the contour of the paper image in the picture according to the two-dimensional topological structure.

[0009] Optionally, before inputting the two-dimensional pixel value array of the picture containing the paper image into the input layer of the self-organizing neural network, it further comprises:

[0010] Determining the number of neurons in the output layer according to the number of pixel points of the picture;

[0011] Determining the number of layers of the output layer and the number of neurons in each layer according to the number of neurons and the aspect ratio of the width and height of the paper corresponding to the paper image, wherein the output layer adopts a two-dimensional topological structure, and the number of neurons in each layer of the output layer is the same;

[0012] Constructing the output layer according to the number of layers of the output layer and the number of neurons in each layer.

[0013] Optionally, determining the contour of the paper image in the picture according to the two-dimensional topological structure specifically includes:

[0014] Connecting the neurons in the two-dimensional topological structure to obtain the contour of the paper image in the picture.

[0015] Optionally, inputting the two-dimensional pixel value array of the picture containing the paper image into the input layer of the self-organizing neural network specifically includes:

[0016] Inputting the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel of the picture into the input layer respectively;

[0017] Obtaining the two-dimensional topological structure of the output layer specifically includes:

[0018] Respectively obtaining the two-dimensional topological structures of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel.

[0019] Optionally, determining the contour of the paper image in the picture according to the two-dimensional topological structure specifically includes:

[0020] Fusing and superimposing the two-dimensional topological structures of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel to obtain a fusion and superposition result;

[0021] Determining the contour of the paper image in the picture according to the fusion and superposition result.

[0022] Optionally, determining the contour of the paper image in the picture according to the fusion and superposition result specifically includes:

[0023] Connecting the neurons in the fusion and superposition result to obtain the contour of the paper image in the picture.

[0024] Optionally, the convergence conditions include: the activation radius function output is 0, the time variable reaches the range boundary, and the states of all neurons no longer update.

[0025] To achieve the above object, according to another aspect of the present invention, a paper positioning device based on unsupervised learning is provided, and the device includes:

[0026] An input unit for inputting the two-dimensional pixel value array of the picture containing the paper image into the input layer of the self-organizing neural network;

[0027] A two-dimensional topology acquisition unit, configured to acquire the two-dimensional topology of the output layer when the state function of the neurons in the output layer of the self-organizing neural network reaches a convergence condition;

[0028] A contour determination unit, configured to determine the contour of the paper image in the picture according to the two-dimensional topology.

[0029] To achieve the above object, according to another aspect of the present invention, there is also provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned paper positioning method based on unsupervised learning are implemented.

[0030] To achieve the above object, according to another aspect of the present invention, there is also provided a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the above-mentioned paper positioning method based on unsupervised learning are implemented.

[0031] To achieve the above object, according to another aspect of the present invention, there is also provided a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned paper positioning method based on unsupervised learning are implemented.

[0032] The beneficial effects of the present invention are as follows:

[0033] In the embodiment of the present invention, by inputting the two-dimensional pixel value array of the picture containing the paper image into the input layer of the self-organizing neural network, and acquiring the two-dimensional topology of the output layer when the state function of the neurons in the output layer of the self-organizing neural network reaches a convergence condition, and finally determining the contour of the paper image in the picture according to the two-dimensional topology. The present invention uses the self-organizing neural network to extract the contour of the paper, achieving the beneficial effect of accurately positioning the paper in the picture. Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0035] Figure 1 It is the first flowchart of the paper positioning method based on unsupervised learning in the embodiment of the present invention;

[0036] Figure 2It is the second flowchart of the paper positioning method based on unsupervised learning in the embodiments of the present invention;

[0037] Figure 3 It is the third flowchart of the paper positioning method based on unsupervised learning in the embodiments of the present invention;

[0038] Figure 4 It is the first structural block diagram of the paper positioning device based on unsupervised learning in the embodiments of the present invention;

[0039] Figure 5 It is the second structural block diagram of the paper positioning device based on unsupervised learning in the embodiments of the present invention;

[0040] Figure 6 It is a schematic diagram of a computer device in the embodiments of the present invention. Detailed implementation manners

[0041] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0042] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0043] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0044] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0045] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.

[0046] It should be noted that the paper positioning method and device based on unsupervised learning of the present invention can be applied to the financial field and can also be applied to other technical fields.

[0047] The present invention proposes a paper positioning method under complex backgrounds based on unsupervised learning. Through an image enhancement algorithm and a self-organizing neural network algorithm, the present invention matches a two-dimensional network topology with a paper reference benchmark on a picture, thereby extracting the contour information of the paper. Subsequently, subsequent steps such as corner extraction and affine transformation can be combined to improve the accuracy rate of subsequent character recognition.

[0048] The present invention uses the method of unsupervised learning, which can effectively avoid the problem that general algorithms cannot separate the target object and complex background noise. The self-organizing neural network (SOM) can screen out low-frequency paper edge features and remove high-frequency complex background features, thereby improving the accuracy rate of paper pose recognition; at the same time, compared with deep learning, it does not require annotation, has no high hardware requirements and training costs.

[0049] Figure 1 is the first flowchart of the paper positioning method based on unsupervised learning in an embodiment of the present invention. As Figure 1 shown, in an embodiment of the present invention, the paper positioning method based on unsupervised learning of the present invention includes steps S101 to S103.

[0050] Step S101: Input the two-dimensional pixel value array of the picture containing the paper image into the input layer of the self-organizing neural network.

[0051] In an embodiment of the present invention, the number of pixels of the picture is: I×J, where I is the row pixel value of the picture, J is the column pixel value of the picture, and both I and J are positive integers. In an embodiment of the present invention, the two-dimensional pixel value array of the picture can be represented by X(i,j), where i∈[1,I] and j∈[1,J].

[0052] In the present invention, the self-organizing neural network consists of an input layer and an output layer (also known as a competitive layer). In an embodiment of the present invention, the input layer has two neurons.

[0053] In an embodiment of the present invention, before the above-mentioned step S101, the present invention also preprocesses and enhances the picture, thereby improving the effect of determining the contour of the paper image in the picture. In an embodiment of the present invention, preprocessing the picture includes: unifying the format and size of the picture, for example, unifying the picture depth to 3 and the size to a unified specification, etc. In an embodiment of the present invention, enhancing the picture includes: simply screening the features of the image.

[0054] Step S102, when the state function of the neurons in the output layer of the self-organizing neural network reaches the convergence condition, obtain the two-dimensional topological structure of the output layer.

[0055] In the present invention, the output layer performs competitive learning according to the data of the output layer until the state function of the neurons in the output layer reaches the convergence condition, and the two-dimensional topological structure of the output layer at this time is obtained.

[0056] In an embodiment of the present invention, the number of neurons in the output layer is determined according to the number of pixel points of the picture, thereby improving the effect of determining the contour of the paper image.

[0057] In an embodiment of the present invention, the output layer adopts a two-dimensional topological structure, and the number of neurons in each layer of the output layer is the same. In an embodiment of the present invention, the number of layers of the output layer and the number of neurons in each layer are determined according to the number of pixel points of the picture and the aspect ratio of the width and height of the paper corresponding to the paper image, so that the neurons in the output layer can better aggregate at the edge of the paper image, improving the effect of determining the contour of the paper image.

[0058] In an embodiment of the present invention, the structure of the output layer can be simply expressed as ω(m,n), where m is the number of neurons in each layer and n is the number of layers of the output layer, where 0 ≤ m ≤ I and 0 ≤ n ≤ J. Considering that the optimal solution of the output layer should form a two-dimensional topological structure and be consistent with the object to be self-trained in structure, so as to more accurately generalize the characteristics of the object, so 0 < m ≤ I and 0 < n ≤ J.

[0059] In an embodiment of the present invention, the convergence condition includes: the activation radius function outputs 0, the time variable reaches the range boundary, and the states of all neurons no longer update.

[0060] Step S103, determine the contour of the paper image in the picture according to the two-dimensional topological structure.

[0061] In an embodiment of the present invention, since the edges of the paper image in the picture have significant differences from the features of other background parts, the output layer can filter out the low-frequency paper edge features and remove the high-frequency complex background features. When the state function of the neurons in the output layer reaches the convergence condition, the neurons in the output layer aggregate at the edges of the paper image, and the two-dimensional topological structure at this time can already reflect the contour of the paper image.

[0062] In another embodiment of the present invention, the present invention can also connect the neurons in the two-dimensional topological structure, and its contour is the contour feature of the paper image in the picture. Because after unsupervised learning, the output layer filters out the noise features and retains the contour features of the paper image. At the same time, the two-dimensional topological structure of the output layer expresses this contour feature in two dimensions.

[0063] Therefore, by simply connecting the node positions output by the output layer, a complete network structure similar to the paper contour can be formed. By performing the Hough transform on this network structure, the contour features without noise interference can be extracted. This contour can be used for subsequent affine transformation, so that the affine transformation has no noise interference.

[0064] The present invention can filter out the low-frequency paper edge features and remove the high-frequency complex background features through a self-organizing neural network (SOM), thereby improving the accuracy of determining the contour of the paper image in the picture. At the same time, compared with deep learning, the present invention does not require annotation, has no high hardware requirements and training costs.

[0065] In the present invention, the output layer performs competitive learning based on the data of the output layer. In competitive learning, the neurons in the output layer compete with each other to be activated, and the activated neurons are called winning neurons in the competition. In an embodiment of the present invention, the present invention can find a certain neuron point that best matches a certain point in the image by calculating the Euclidean distance or cosine similarity between the pixel points in the image and the neurons, that is, the winning neuron. If the Euclidean distance is used, the similarity matching formula is:

[0066]

[0067] In formula (1), ω m,n is the neuron point of the output layer, X i,j is the pixel point in the image, and d(X i,j ) is the Euclidean distance between the neuron point ω m,n and the pixel point X i,j .

[0068] In the present invention, an activation function is also provided. The activation function can adjust the weights between the winning neuron and ordinary neurons. The adjustment relationship is from far to near, and the weights decrease from large to small. Among them, the neuron with the largest weight is the winning neuron, and the neuron with a weight of 0 is the neuron beyond a certain distance. As time progresses, at each time point, the winning neuron needs to be updated according to the activation function to ensure that the winning neuron is in the best activation state.

[0069] In an embodiment of the present invention, the activation function can specifically be:

[0070]

[0071] In formula (2), h m,n is the activation state of a certain neuron, and d m,n is the distance of this neuron from the winning neuron node. When d = 0, it is the winning neuron itself, and it is h max ; the nodes near the winning neuron have weights slightly smaller than h max ; when d → ∞, h ≈ 0, then this node has no weight, will not be activated, and will not participate in the competition for the winning neuron.

[0072] Corresponding to the image, this activation function will activate the neurons near the winning neuron ω m,n and update the competition state of the neurons by assigning certain weights and calculating the surrounding pixel values, so as to obtain the pixel positions in the area with the most obvious features for output.

[0073] In the present invention, an activation radius function (which can also be called a learning rate function) is also provided. The activation radius function describes the relationship between the radius with the winning neuron as the origin and its length decreasing with time. The range enclosed by this radius is generally called the winning neighborhood. Only the neurons within this range can participate in the competition for the winning neuron, and neurons at different distances from the initial winning neuron will obtain different weights through the activation function (according to the properties of the activation function, the magnitude of the weight decreases with the increase of the distance).

[0074] In an embodiment of the present invention, the activation radius function can specifically be:

[0075]

[0076] In formula (3), σ0 is the initial radius, and this radius can be set to a relatively large value; n is the time variable; τ is the time constant, which is related to the decay rate of the radius.

[0077] In the present invention, the initial two-dimensional topological structure, activation function, and activation radius function of the constructed output layer are brought into the calculation of the states of each neuron. The state of each neuron at time n + 1 is:

[0078] ω m,n (n + 1)= ω m,n (n)-σ(n)h i,j (n)(x i,j -ω m,n ) (4)

[0079] In formula (4), ω m,n (n + 1) is the state of neuron point ω m,n at time n + 1, ω m,n (n) is the state of neuron point ω m,n at time n, σ(n) is the activation radius function, h i,j (n) is the activation function.

[0080] Generally, the convergence limit of this function has the following situations: the output of the activation radius function is 0; the time variable reaches the boundary of the range; all neuron states are no longer updated. The specific convergence conditions can be judged according to the actual situation.

[0081] Figure 2 is the second flowchart of the paper positioning method based on unsupervised learning in the embodiments of the present invention. As Figure 2 shown, in an embodiment of the present invention, before the above step S101, the paper positioning method based on unsupervised learning of the present invention includes steps S201 to S203.

[0082] Step S201, determine the number of neurons in the output layer according to the number of pixel points of the picture.

[0083] In an embodiment of the present invention, the number of pixels of the picture is: I×J. In an embodiment of the present invention, the number of neurons in the output layer is a preset percentage of the number of pixels of the picture. In an embodiment of the present invention, the number of neurons in the output layer does not exceed N, where N is a positive integer. In a specific embodiment of the present invention, N is 2000.

[0084] Step S202, determine the number of layers in the output layer and the number of neurons in each layer according to the number of neurons and the aspect ratio of the width and height of the paper corresponding to the paper image, where the output layer adopts a two-dimensional topological structure and the number of neurons in each layer of the output layer is the same.

[0085] In an embodiment of the present invention, the ratio of the number of layers in the output layer to the number of neurons in each layer is the same as the aspect ratio of the width and height of the paper corresponding to the paper image, thereby improving the accuracy of determining the contour of the paper image.

[0086] Step S203, construct the output layer according to the number of layers in the output layer and the number of neurons in each layer.

[0087] In one embodiment of the present invention, inputting the two-dimensional pixel value array of the picture containing the paper image into the input layer of the self-organizing neural network in step S101 specifically includes:

[0088] Inputting the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel of the picture into the input layer respectively.

[0089] In one embodiment of the present invention, since the picture can be stored in the format of three two-dimensional arrays (R, G, and B channels), the arrays of the three channels can be calculated separately and then fused and superimposed. The arrays of the three channels are respectively represented as: X R (i,j), X G (i,j), X B (i,j), where X R (i,j) is the two-dimensional array of the R channel, X G (i,j) is the two-dimensional array of the G channel, X B (i,j) is the two-dimensional array of the B channel.

[0090] In one embodiment of the present invention, obtaining the two-dimensional topological structure of the output layer in step S102 specifically includes:

[0091] Respectively obtaining the two-dimensional topological structures of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel.

[0092] As Figure 3 shown, in one embodiment of the present invention, determining the contour of the paper image in the picture according to the two-dimensional topological structure in step S103 specifically includes step S301 and step S302.

[0093] Step S301, fusing and superimposing the two-dimensional topological structures of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel respectively to obtain a fusion and superposition result.

[0094] Step S302, determining the contour of the paper image in the picture according to the fusion and superposition result.

[0095] The present invention first separately determines the two-dimensional topological structures of the R, G, and B channels of the picture, and then fuses and superimposes the two-dimensional topological structures, thereby achieving the beneficial effect of more accurately determining the contour of the paper image in the picture.

[0096] In one embodiment of the present invention, the present invention can connect each neuron in the fusion and superposition result, and its contour is the contour feature of the paper image in the picture.

[0097] As can be seen from the above embodiments, the paper positioning method based on unsupervised learning of the present invention has at least achieved the following beneficial effects:

[0098] 1. By using the method of unsupervised learning, the present invention enhances the anti-noise ability and improves the accuracy of subsequent affine transformation;

[0099] 2. Compared with the deep learning method, the present invention adopts an unsupervised algorithm, which does not need to be trained with a large amount of data and only needs to perform operations on a single sample. Therefore, it can be applied to different scenarios, while the deep learning method needs to train different models according to different scenarios;

[0100] 3. The present invention requires less computing power. Unsupervised learning only involves simple matrix operations and does not require hardware such as GPUs to participate. Compared with the deep learning method, the implementation of the present invention is more convenient.

[0101] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0102] Based on the same inventive concept, the embodiment of the present invention also provides a paper positioning device based on unsupervised learning, which can be used to implement the paper positioning method based on unsupervised learning described in the above embodiments, as described in the following embodiments. Since the principle of the paper positioning device based on unsupervised learning to solve problems is similar to that of the paper positioning method based on unsupervised learning, the embodiments of the paper positioning device based on unsupervised learning can refer to the embodiments of the paper positioning method based on unsupervised learning, and the repeated parts will not be described again. As used hereinafter, the term "unit" or "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0103] Figure 4 is the first structural block diagram of the paper positioning device based on unsupervised learning in the embodiment of the present invention, as Figure 4 shown, in one embodiment of the present invention, the paper positioning device based on unsupervised learning of the present invention includes:

[0104] An input unit 1, configured to input a two-dimensional pixel value array of a picture containing a paper image into the input layer of the self-organizing neural network;

[0105] A two-dimensional topology acquisition unit 2, configured to acquire the two-dimensional topology of the output layer when the state function of the neurons in the output layer of the self-organizing neural network reaches a convergence condition;

[0106] A contour determination unit 3, configured to determine the contour of the paper image in the picture according to the two-dimensional topology;

[0107] Figure 5 It is the second structural block diagram of the paper positioning device based on unsupervised learning in the embodiments of the present invention. As Figure 5 shown, in an embodiment of the present invention, the paper positioning device based on unsupervised learning of the present invention further includes:

[0108] A neuron number determination unit 4, configured to determine the number of neurons in the output layer according to the number of pixels of the picture;

[0109] An output layer structure determination unit 5, configured to determine the number of layers of the output layer and the number of neurons in each layer according to the number of neurons and the aspect ratio of the width and height of the paper corresponding to the paper image, wherein the output layer adopts a two-dimensional topology, and the number of neurons in each layer of the output layer is the same;

[0110] An output layer construction unit 6, configured to construct the output layer according to the number of layers of the output layer and the number of neurons in each layer.

[0111] In an embodiment of the present invention, the input unit 1 is specifically configured to input the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel of the picture into the input layer respectively.

[0112] In an embodiment of the present invention, the two-dimensional topology acquisition unit 2 is specifically configured to acquire the two-dimensional topology of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel respectively.

[0113] In an embodiment of the present invention, the contour determination unit 3 specifically includes:

[0114] A fusion superposition result determination module, configured to fuse and superpose the two-dimensional topologies of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel respectively to obtain a fusion superposition result;

[0115] A contour determination module, configured to determine the contour of the paper image in the picture according to the fusion superposition result.

[0116] To achieve the above object, according to another aspect of the present application, a computer device is also provided. As Figure 6 shown, the computer device includes a memory, a processor, a communication interface, and a communication bus. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps in the method of the above embodiment are implemented.

[0117] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.

[0118] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and units, such as the corresponding program units in the method embodiment of the present invention above. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor executes various functional applications and work data processing of the processor, that is, the method in the above method embodiment is implemented.

[0119] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0120] The one or more units are stored in the memory and, when executed by the processor, execute the method in the above embodiment.

[0121] The specific details of the above computer device can be understood by referring to the corresponding relevant descriptions and effects in the above embodiment, and will not be elaborated here.

[0122] To achieve the above object, according to another aspect of the present application, there is also provided a computer-readable storage medium storing a computer program, and when the computer program is executed in a computer processor, it implements the steps in the above-mentioned paper positioning method based on unsupervised learning. Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiment methods, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memories.

[0123] To achieve the above object, according to another aspect of the present application, there is also provided a computer program product including a computer program / instructions, and when the computer program / instructions are executed by a processor, they implement the steps of the above-mentioned paper positioning method based on unsupervised learning.

[0124] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0125] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A paper positioning method based on unsupervised learning, characterized in that including: inputting a two-dimensional pixel value array of a picture including a paper image into an input layer of a self-organizing neural network; when a state function of neurons in an output layer of the self-organizing neural network reaches a convergence condition, the neurons in the output layer aggregate at the edge of the paper image, and acquiring a two-dimensional topological structure of the output layer; determining a contour of the paper image in the picture according to the two-dimensional topological structure; wherein, the determining the contour of the paper image in the picture according to the two-dimensional topological structure specifically includes: connecting the neurons in the two-dimensional topological structure to obtain the contour of the paper image in the picture.

2. The method for paper positioning based on unsupervised learning according to claim 1, wherein Before inputting the two-dimensional pixel value array of the picture including the paper image into the input layer of the self-organizing neural network, it further includes: determining the number of neurons in the output layer according to the number of pixel points of the picture; determining the number of layers of the output layer and the number of neurons in each layer according to the number of neurons and the aspect ratio of the width and height of the paper corresponding to the paper image, wherein the output layer adopts a two-dimensional topological structure, and the number of neurons in each layer of the output layer is the same; constructing the output layer according to the number of layers of the output layer and the number of neurons in each layer.

3. The method for paper positioning based on unsupervised learning according to claim 1, wherein The inputting the two-dimensional pixel value array of the picture including the paper image into the input layer of the self-organizing neural network specifically includes: respectively inputting the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel of the picture into the input layer; The acquiring the two-dimensional topological structure of the output layer specifically includes: respectively acquiring the two-dimensional topological structure of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel.

4. The method for paper positioning based on unsupervised learning according to claim 3, wherein The determining the contour of the paper image in the picture according to the two-dimensional topological structure specifically includes: fusing and superposing the two-dimensional topological structures of the output layer corresponding to the two-dimensional pixel value array of the R channel, the two-dimensional pixel value array of the G channel, and the two-dimensional pixel value array of the B channel to obtain a fusion and superposition result; determining the contour of the paper image in the picture according to the fusion and superposition result.

5. The method for paper positioning based on unsupervised learning according to claim 4, wherein The determining the contour of the paper image in the picture according to the fusion and superposition result specifically includes: connecting the neurons in the fusion and superposition result to obtain the contour of the paper image in the picture.

6. The method for paper positioning based on unsupervised learning according to claim 1, characterized in that The convergence condition includes: the activation radius function outputs 0, the time variable reaches the range boundary, and the states of all neurons no longer update.

7. A paper positioning device based on unsupervised learning, characterized in that, including: an input unit for inputting a two-dimensional pixel value array of a picture including a paper image into an input layer of a self-organizing neural network; a two-dimensional topological structure acquisition unit for, when a state function of neurons in an output layer of the self-organizing neural network reaches a convergence condition, the neurons in the output layer aggregate at the edge of the paper image, and acquiring a two-dimensional topological structure of the output layer; a contour determination unit for determining a contour of the paper image in the picture according to the two-dimensional topological structure; Specifically, the contour determination unit is configured to connect the neurons in the two-dimensional topological structure to obtain the contour of the paper image in the picture.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.