Image-based risk text processing methods, apparatus, devices, and storage media
By using text recognition models and knowledge graph technology, the problem of incomplete risk information in images with mixed handwriting and printed text in traditional optical character recognition technology has been solved, enabling accurate identification and desensitization of risk text in financial business images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2026-03-13
Smart Images

Figure CN116740735B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus, computer device, storage medium, and computer program product for processing image-based risky text. Background Technology
[0002] With the development of artificial intelligence, a method for automatically identifying risk information in financial business images has emerged. This technology uses traditional optical character recognition (OCR) technology to identify character text in images, thereby identifying risk information in the images and processing that risk information.
[0003] In the above technical solution, much of the information in the image is handwritten by the customer, resulting in inconsistent font formats and a mix of handwritten and printed text. In this scenario, traditional optical character recognition technology cannot accurately identify the text content in the image, making it impossible to fully and accurately process the risk information in the image. Summary of the Invention
[0004] Therefore, it is necessary to provide an image risk text processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can completely and accurately process risk text in financial business text images, addressing the aforementioned technical problems.
[0005] Firstly, this application provides a method for processing image-based risky text. The method includes:
[0006] A financial business text image is acquired and input into a pre-built text recognition model. The text feature acquisition layer in the text recognition model is used to obtain the text content features corresponding to the character units contained in the financial business text image.
[0007] The text content features are input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character unit;
[0008] The text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image;
[0009] Based on the text characters, risk text regions are identified from the financial business text image, and the risk text regions are desensitized to obtain the desensitized financial business text image.
[0010] In one embodiment, the step of inputting the text content features into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character unit includes: inputting the text content features into the region feature acquisition layer, obtaining the character unit position region corresponding to the character unit through the region feature acquisition layer; and superimposing the character unit position region with the text content features to obtain the text region features corresponding to the character unit.
[0011] In one embodiment, the number of character units is multiple; the step of inputting the text region features into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image includes: inputting the text region features corresponding to each character unit into the text recognition layer, obtaining the positional correlation between each character unit through the text recognition layer; and obtaining the text characters contained in the financial business text image based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer.
[0012] In one embodiment, obtaining the positional correlation between each character unit through the text recognition layer includes: obtaining the target text region features corresponding to each character unit from the text region features corresponding to each character unit through the text recognition layer; obtaining the distance information between each character unit based on the target text region features corresponding to each character unit; and obtaining the positional correlation between each character unit based on the distance information between each character unit.
[0013] In one embodiment, obtaining the text content features corresponding to the character units contained in the financial business text image through the text feature acquisition layer in the text recognition model includes: extracting text content features of the financial business text image at multiple scales through the text feature acquisition layer to obtain multiple initial text content features corresponding to the financial business text image; and superimposing the multiple initial text features according to a preset superposition probability to obtain the text content features corresponding to the character units.
[0014] In one embodiment, identifying risk text regions from the financial business text image based on the text characters includes: retrieving the text characters using a pre-set financial business knowledge graph to obtain retrieval results corresponding to the text characters; identifying risk text characters in the financial business text image based on the retrieval results, and taking the area where the risk text characters are located as the risk text region in the financial business text image.
[0015] In one embodiment, the step of desensitizing the risk text region to obtain a desensitized financial business text image includes: blurring the risk text region and selecting a preset number of target pixels from the pixels of the blurred risk text region; changing the color of the target pixels to a preset color to obtain the desensitized financial business text image.
[0016] Secondly, this application also provides an image-based risk text processing device. The device includes:
[0017] The text content feature acquisition module is used to acquire financial business text images and input the financial business text images into a pre-built text recognition model. Through the text feature acquisition layer in the text recognition model, the text content features corresponding to the character units contained in the financial business text images are obtained.
[0018] The text region feature acquisition module is used to input the text content features into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character unit.
[0019] The text character acquisition module is used to input the text region features into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image;
[0020] The text image desensitization module is used to identify risk text regions from the financial business text image based on the text characters, and to desensitize the risk text regions to obtain a desensitized financial business text image.
[0021] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0022] A financial business text image is acquired and input into a pre-built text recognition model. The text feature acquisition layer in the text recognition model is used to obtain the text content features corresponding to the character units contained in the financial business text image.
[0023] The text content features are input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character unit;
[0024] The text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image;
[0025] Based on the text characters, risk text regions are identified from the financial business text image, and the risk text regions are desensitized to obtain the desensitized financial business text image.
[0026] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0027] A financial business text image is acquired and input into a pre-built text recognition model. The text feature acquisition layer in the text recognition model is used to obtain the text content features corresponding to the character units contained in the financial business text image.
[0028] The text content features are input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character unit;
[0029] The text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image;
[0030] Based on the text characters, risk text regions are identified from the financial business text image, and the risk text regions are desensitized to obtain the desensitized financial business text image.
[0031] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0032] A financial business text image is acquired and input into a pre-built text recognition model. The text feature acquisition layer in the text recognition model is used to obtain the text content features corresponding to the character units contained in the financial business text image.
[0033] The text content features are input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character unit;
[0034] The text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image;
[0035] Based on the text characters, risk text regions are identified from the financial business text image, and the risk text regions are desensitized to obtain the desensitized financial business text image.
[0036] The aforementioned image risk text processing method, apparatus, computer equipment, storage medium, and computer program product acquire a financial business text image and input it into a pre-constructed text recognition model. Through the text feature acquisition layer in the text recognition model, the text content features corresponding to the character units contained in the financial business text image are obtained. The text content features are then input into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character units. The text region features are then input into the text recognition layer in the text recognition model to obtain the text characters contained in the financial business text image. Based on the text characters, risk text regions are identified from the financial business text image, and the risk text regions are desensitized to obtain a desensitized financial business text image. This application obtains the text content features corresponding to the character units contained in the financial business text image through the text feature acquisition layer of the text recognition model, then obtains the text region features corresponding to the above character units through the region feature acquisition layer of the text recognition model, and then obtains the text characters contained in the financial business text image through the text recognition layer of the text recognition model. Finally, based on the text characters, the risk text in the financial business text image is identified, and the risk text is desensitized, which can completely and accurately process the risk text in the financial business text image. Attached Figure Description
[0037] Figure 1 This is a flowchart illustrating an image risk text processing method in one embodiment;
[0038] Figure 2 This is a flowchart illustrating the process of obtaining text region features in one embodiment;
[0039] Figure 3 This is a schematic diagram of the process for obtaining text characters in one embodiment;
[0040] Figure 4 This is a schematic diagram of the process for obtaining location relevance in one embodiment;
[0041] Figure 5 This is a framework diagram of an image risk text processing model in one embodiment;
[0042] Figure 6 This is a framework diagram of a feature pyramid neural network in one embodiment;
[0043] Figure 7 This is a framework diagram of a correlation graph convolutional neural network in one embodiment;
[0044] Figure 8 This is a structural block diagram of an image risk text processing device in one embodiment;
[0045] Figure 9Internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0046] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0047] It should be noted that the terms "first / second" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when permitted. It should be understood that the objects distinguished by "first / second" can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.
[0048] In one embodiment, as Figure 1 shown, a method for processing image risk texts is provided. In this embodiment, the method is illustrated by taking its application to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0049] Step S101: Obtain a financial business text image, input the financial business text image into a pre-constructed text recognition model, and obtain the text content features corresponding to the character units included in the financial business text image through the text feature acquisition layer in the text recognition model.
[0050] Among them, the financial business text image is a picture of a document generated in financial business. For example, the financial business text image can be a picture of a user's payment voucher. At the same time, in addition to standard printed text, there may also be handwritten text in the financial business text image. The pre-constructed text recognition model is a text recognition model for financial business text images. The text feature acquisition layer is a jumping feature pyramid neural network structure, and this text feature acquisition layer can be used to extract text content features from the financial business text image. Then, the character unit is the character unit that constitutes the text characters in the financial business text image. This character unit can form a text character alone, or multiple text character units can form a text character. For example, the character unit "金" can form the text character "金" alone, and the character units "日" and "月" can form the text character "明". It should be noted that the spacing between handwritten text characters is not standardized, and it is easy to combine character units into incorrect text characters. Finally, the text content feature refers to the text content feature corresponding to the above character unit.
[0051] Specifically, the original financial business text image is obtained from the financial business database. The original financial business text image is preprocessed by correction, noise reduction and binarization to obtain the financial business text image. The financial business text image is input into the pre-built text recognition model. The jump feature pyramid neural network of the text feature acquisition layer in the text recognition model is used to extract features from the character units contained in the financial business text image to obtain the text content features corresponding to the character units.
[0052] Step S102: Input the text content features into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character units.
[0053] Among them, the region feature acquisition layer is a region extraction network structure. This region feature acquisition layer can be used to extract multiple different positional region anchor boxes corresponding to text region features (i.e. character units). The text region features are the features of the region of the above character units superimposed with the text content.
[0054] Specifically, the text content features are input into the region feature acquisition layer in the text recognition model. The region extraction network in the region feature acquisition layer obtains multiple different positional region anchor boxes corresponding to the character unit. The multiple positional region anchor boxes are superimposed with the text content features to obtain the text region features corresponding to the character unit.
[0055] Step S103: Input the text region features into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image.
[0056] The text recognition layer is a correlation graph convolutional neural network structure, and the text characters are the characters in the financial business text image.
[0057] Specifically, the text region features are input into the text recognition layer of the text recognition model. Based on the text region features, the graph nodes and edge relationships of the correlation graph convolutional neural network of the text recognition layer are obtained. Then, the text characters contained in the financial business text image are obtained through the correlation graph convolutional neural network of the text recognition layer.
[0058] Step S104: Based on text characters, identify the risk text region from the financial business text image, and perform desensitization processing on the risk text region to obtain the desensitized financial business text image.
[0059] The risk text region is the area in the financial business text image where the risk text character is located, and the desensitization process involves deleting or blurring the risk text region.
[0060] Specifically, by searching and comparing text characters, risk text characters are obtained from the text characters, thereby obtaining the risk text region where the risk text characters are located in the financial business text image, and deleting or blurring and adding noise to the risk text region to obtain the desensitized financial business text image.
[0061] In the aforementioned image risk text processing method, a financial business text image is acquired and input into a pre-constructed text recognition model. The text feature acquisition layer of the text recognition model obtains the text content features corresponding to the character units contained in the financial business text image. These text content features are then input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character units. Finally, the text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image. Based on these text characters, risk text regions are identified from the financial business text image, and these risk text regions are then desensitized to obtain a desensitized financial business text image. This application uses the text feature acquisition layer of the text recognition model to obtain the text content features corresponding to the character units contained in the financial business text image. Then, it uses the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the aforementioned character units. Next, it uses the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image. Finally, based on these text characters, it identifies the risk text in the financial business text image and desensitizes the risk text, thus enabling complete and accurate processing of risk text in financial business text images.
[0062] In one embodiment, such as Figure 2 As shown, the text content features are input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character units, including the following steps:
[0063] Step S201: Input the text content features into the region feature acquisition layer, and obtain the character unit position region corresponding to the character unit through the region feature acquisition layer.
[0064] The character unit location area is as described above, which is the location area of the character unit in the financial business text image. It should be noted that there can be multiple character unit location areas, specifically multiple rectangular anchor box areas.
[0065] Specifically, the text content features are input into the region feature acquisition layer, and the region feature acquisition layer obtains multiple rectangular anchor box regions corresponding to the character units.
[0066] Step S202: Overlay the character unit position region with the text content features to obtain the text region features corresponding to the character unit.
[0067] Specifically, multiple rectangular anchor regions corresponding to a character unit are superimposed with the text content features of that character unit to obtain the text region features corresponding to the character unit.
[0068] In this embodiment, by obtaining the character unit position region corresponding to the character unit and superimposing the character unit position region with the text content features, the text region features corresponding to the character unit can be accurately obtained.
[0069] In one embodiment, such as Figure 3 As shown, there are multiple character units; the text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image, including the following steps:
[0070] Step S301: Input the text region features corresponding to each character unit into the text recognition layer, and obtain the positional correlation between each character unit through the text recognition layer.
[0071] Among them, positional relevance is the degree of correlation between each character unit based on position. Specifically, positional relevance is used to characterize whether there is a relationship between two character units, and positional relevance can be represented by an adjacency matrix.
[0072] Specifically, the text region features corresponding to each character unit are input into the text recognition layer. The positional correlation between each character unit is obtained through the correlation graph convolutional neural network in the text recognition layer based on the distance relationship between each character unit.
[0073] Step S302: Based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer, the text characters contained in the financial business text image are obtained.
[0074] The model parameters are the weight matrices learned by backpropagation of the correlation graph convolutional neural network.
[0075] Specifically, as follows:
[0076] Z = S·X·W;
[0077] Where Z represents the text feature map output by the text recognition layer, i.e., the text characters contained in the financial business text image; S is the adjacency matrix of the graph nodes in the correlation graph convolutional neural network, i.e., the positional correlation between each character unit; X is the input information of the graph nodes in the correlation graph convolutional neural network, i.e., the positional correlation between each character unit; and W is the weight matrix learned for backpropagation of the correlation graph convolutional neural network, i.e., the model parameters. Using the above formula, based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer, the text characters contained in the financial business text image are calculated.
[0078] In this embodiment, by obtaining the positional correlation between each character unit, and based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer, the text characters contained in the financial business text image can be accurately obtained.
[0079] In one embodiment, such as Figure 4 As shown, the positional correlation between each character unit is obtained through the text recognition layer, including the following steps:
[0080] Step S401: Through the text recognition layer, obtain the target text region features corresponding to each character unit from the text region features corresponding to each character unit.
[0081] Among them, the target text region feature is the text region feature that best represents the location region of each character among multiple text region features.
[0082] Specifically, through the text recognition layer, based on preset calculation parameters, the target text region features corresponding to each character unit are obtained from the text region features corresponding to each character unit.
[0083] Step S402: Based on the target text region features corresponding to each character unit, obtain the distance information between each character unit.
[0084] The distance information is the ratio of the distance between each character unit to a preset value.
[0085] Specifically, based on the target anchor frame region contained in the target text region features corresponding to each character unit, the center position of each target anchor frame region is obtained, and then the ratio of the distance between each center position to a preset value is obtained as the distance information between each character unit.
[0086] Step S403: Based on the distance information between each character unit, obtain the positional correlation between each character unit.
[0087] Specifically, the distance information between each character unit is assigned a corresponding weight. When the weight meets the preset condition, the two character units corresponding to the weight are related. When the weight does not meet the preset condition, the two character units corresponding to the weight are not related. The correlation between each character unit is used as the positional correlation between each character unit.
[0088] In this embodiment, the target text region features are obtained from the text region features corresponding to each character unit, and then the distance information between each character unit is obtained based on the target text region features. Based on this distance information, the positional correlation between each character unit can be accurately obtained.
[0089] In one embodiment, the text content features corresponding to the character units contained in the financial business text image are obtained through the text feature acquisition layer in the text recognition model, including the following steps:
[0090] Through the text feature acquisition layer, multi-scale text content features are extracted from the financial business text image to obtain multiple initial text content features corresponding to the financial business text image; according to the pre-set superposition probability, the multiple initial text features are superimposed to obtain the text content features corresponding to the character unit.
[0091] Among them, multiple initial text content features are text content features of different scales in the financial business text image, while the superposition probability is the probability of superimposing multiple text content features of different scales.
[0092] Specifically, the text feature acquisition layer uses a skip feature pyramid network to extract text content features from low to high scales in the financial business text image until the text content features at the character unit scale are extracted. Finally, multiple initial text content features are obtained. Since some information (such as positional information) will be lost during the feature extraction process from low to high scale, the multiple initial text features are superimposed according to a pre-set superposition probability to obtain the text content features corresponding to the character unit.
[0093] In this embodiment, the skip feature pyramid network of the text feature acquisition layer is used to obtain initial text content features at multiple different scales corresponding to the financial business text image. The multiple initial text features are superimposed to accurately obtain the text content features corresponding to the character unit. At the same time, the text content features contain other information of the character unit in addition to the text content of the character unit, such as position information, so that the position region (anchor box region) corresponding to the character unit can be obtained based on the position information.
[0094] In one embodiment, identifying risk text regions from a financial business text image based on text characters includes the following steps:
[0095] By searching the text characters using a pre-set financial business knowledge graph, the search results corresponding to the text characters are obtained. Based on the search results, risk text characters in the financial business text image are identified, and the area where the risk text characters are located is taken as the risk text area in the financial business text image.
[0096] The financial business knowledge graph is OwnThink's Chinese open-source knowledge graph. Search results are text characters or text strings specific to this Chinese open-source knowledge graph; for example, a search result could be a text string containing a phone number. Risk text characters are text characters containing risk-related content or text characters within a text string.
[0097] Specifically, by inputting text characters from a financial business text image into a Chinese open-source knowledge graph for retrieval, the search results corresponding to a certain text character or a certain text string can be retrieved. For example, the search result may be a text string that is a phone number. According to the preset risk content library, the string corresponding to the search result (phone number) is identified as risk content. Then, all the characters contained in the string are risk text characters, and the area where the risk text characters are located is taken as the risk text area in the financial business text image.
[0098] In this embodiment, the text characters are retrieved using a Chinese open-source knowledge graph to obtain the retrieval results corresponding to the text characters. Then, based on the retrieval results, the risk text characters in the financial business text image are identified, and the area where the risk text characters are located is taken as the risk text area in the financial business text image, thus accurately obtaining the risk text area.
[0099] In one embodiment, the risk text area is de-identified to obtain a de-identified financial business text image, including the following steps:
[0100] The risk text area is blurred, and a preset number of target pixels are selected from the pixels of the blurred risk text area; the color of the target pixels is changed to a preset color to obtain the desensitized financial business text image.
[0101] The blurring process is Gaussian blurring, and the target pixels are randomly selected pixels in the risk text area. The preset color can be black or white.
[0102] Specifically, the center of the risky text region is used as the reference point, and other points are calculated according to their respective weights based on a Gaussian normal distribution to obtain a weighted average. The closer the risky region is to the center, the larger the value; the farther away from the center, the smaller the value, thus ultimately achieving the Gaussian blur effect. The calculation process is shown below:
[0103]
[0104] Where G represents the Gaussian function, x and y represent the horizontal and vertical distances from the central reference point, respectively, e is the natural base, and σ is the standard deviation.
[0105] Additionally, salt-and-pepper noise can be added to the risk text region that has undergone Gaussian blurring, causing random black and white dots to appear, further visually interfering with the risk information. The calculation process is as follows:
[0106] NP = SP × (1 - SNR);
[0107] Wherein, NP represents the number of pixels to be noise-added (preset number), SP is the total number of pixels in the risk area, and SNR is the signal-to-noise ratio, which ranges from [0,1].
[0108] In this embodiment, by applying Gaussian blurring and salt-and-pepper noise processing to the risk text region, a complete and accurate desensitized financial business text image can be obtained.
[0109] In one application embodiment, an image risk text processing method for financial business images is provided. This method processes image risk text using a pre-obtained image risk text processing model. The framework structure of the model is as follows: Figure 5 As shown, the image risk text processing model includes: a text feature acquisition layer using a jump feature pyramid neural network; a region feature acquisition layer using a region extraction network; a text recognition layer using a correlation graph convolutional neural network; a risk text recognition layer using OwnThink's Chinese open-source knowledge graph; and a risk text desensitization layer using Gaussian blur and salt-and-pepper noise techniques. First, the financial business image is input into the above image risk text processing model. The jump feature pyramid neural network in the text feature acquisition layer acquires the text content features corresponding to the character units contained in the financial business image. Then, the region extraction network in the region feature acquisition layer acquires the text region features corresponding to the aforementioned character units. Next, the correlation graph convolutional neural network in the text recognition layer identifies the text characters contained in the financial business image. Finally, the risk text recognition layer identifies the risk text in the financial business image, and the risk text desensitization layer performs Gaussian blur and salt-and-pepper noise processing on the risk text, ultimately outputting the risk text desensitized financial business image.
[0110] The detailed steps are as follows:
[0111] 1. Image preprocessing for financial transactions
[0112] Given the complex features of the image scene, image preprocessing is required first. Tilt detection and rotation correction methods from the OpenCV library are used to correct the angle of the text in the image. Simultaneously, an adaptive Wiener filter is used to smooth areas with high local variance in the image to reduce background noise. Finally, the entire image is binarized to enhance the readability of the text in the financial business image.
[0113] 2. Obtain text content features through the text feature acquisition layer.
[0114] Leveraging the multi-scale structure of feature pyramids, text content features at different scales can be effectively extracted. Considering the varying complexity of features across different levels, and that while semantic information gradually concentrates as the feature pyramid deepens, some rich low-level features are lost during convolution operations, skip connections can be used to superimpose feature information from different levels. Figure 6 As shown, spatial location information from low-level features is superimposed to help locate semantic information in high-level features, thereby obtaining the text content features of each character unit in the financial business image. Specifically, the hierarchical connections in the backbone path use a 3×3 convolution operation with a stride of 1, while cross-level connections use a 3×3 convolution operation with a stride of m, where m represents the number of levels crossed. This cross-level skip feature pyramid achieves information fusion between different levels, improving the effective extraction of features. To further generalize the model's feature extraction capability, skip connections between different levels are performed with a certain probability, where the probability of a skip connection is p = 0.5.
[0115] 3. Obtain character text through the text recognition layer
[0116] After obtaining the feature map using the jump feature pyramid, a Region Proposal Network (RPN) is used to extract the features of each candidate region as graph nodes. After obtaining different character feature nodes, different weights are assigned based on different distance ratios, and these different weights are used as the edge relationships between the character feature nodes to construct a correlation graph convolutional neural network, such as... Figure 7 As shown. Since graph convolutional networks can extract spatial features from topological graphs and optimize and update the state information of each node in the graph structure, they can simultaneously perform convolutional computations on non-Euclidean data for both node and edge relationships. Therefore, the computation formula for a correlation graph convolutional neural network is as follows:
[0117]
[0118] Among them, g θLet represent the diagonal matrix composed of Fourier transforms; ⊙ represents the Hadamard product, which is a multiplication operation between matrices of the same order; x represents the input of the graph convolutional network; θ is the parameter in the Chebyshev polynomial convolution kernel; I N Let D represent the degree matrix of each node, and A represent the adjacency matrix of each node. Considering that repeatedly using this formula in deep networks may lead to vanishing or exploding gradients, normalizing the above formula yields:
[0119]
[0120] in, It is the adjacency matrix of an undirected graph with self-loop connections, and Θ represents the parameter weight matrix of the graph convolution kernel, and Z represents the output matrix of the graph convolution after normalization, that is, the output features of each convolutional layer.
[0121] Since the nodes of a graph convolutional network model represent features extracted from each image by a jump feature pyramid, and the edges of the graph structure represent the positional relationships between features, then for the graph structure vector G = (V, E, W) of the graph convolutional network, where V = {x1, x2, ... x...} N Let} represent N feature nodes of the text character, E represent the positional correlation between feature nodes, and W represent the learnable weight coefficients. Then, the positional correlation between two feature nodes i and j is:
[0122] S(x1, x2) = h(x i ) T h′(x j );
[0123] Where h(xi) = wxi, h'(xj) = w'xj, w and w' are d×d dimensional learnable weight matrices. The correlation between each character can be calculated from the above formula. The value of each edge of the graph node is standardized so that the sum of the edge values of each node is 1. The standardized values can form the adjacency matrix of the graph. Substituting the adjacency matrix formed by the above formula into the graph convolutional network. From this, we can obtain:
[0124] Z = S·X·W;
[0125] Where Z is the feature map output by the convolutional layer, i.e., the character text, with a size of N×d; S is the adjacency matrix of the graph nodes, with a size of N×N; X is a vector of size N×d composed of N region features extracted from each video frame, which is the node input information of the graph convolutional network; W is a d×d dimensional weight matrix that can be learned through backpropagation of the graph convolutional network. After graph convolutional learning and calculation, the text information in the image is obtained and stored in a vector.
[0126] 4. Desensitization of split text
[0127] Knowledge graphs contain a wide range of knowledge sources, with relatively clear relationships between data, and can be updated along with the development of existing knowledge. Currently, knowledge graph technology is widely used in text and image fields such as intelligent search and personalized recommendation. Based on the practical needs of text content recognition in this invention, the character text output by the relevance graph convolutional neural network is used as an entity. OwnThink's Chinese open-source knowledge graph is used to identify risks in the text content. If risky text such as long strings of numbers (bank accounts, phone numbers, etc.) or illegal or harmful content is found, the risk area in the financial business image is determined by locating the risk text using spatial location information from low-level visual features. The center of the risk area is used as the reference point, and other points are calculated according to the corresponding weights of a Gaussian normal distribution to obtain a weighted average. The closer the risk area is to the center, the larger the value; the farther away from the center, the smaller the value, thus ultimately achieving a Gaussian blur effect. The calculation process is as follows:
[0128]
[0129] Where G represents the Gaussian function, x and y represent the horizontal and vertical distances from the central reference point, respectively, e is the natural base, and σ is the standard deviation.
[0130] Furthermore, a salt-and-pepper noise template can be added to the risk areas that have undergone Gaussian blurring, causing random black and white dots to appear, further visually interfering with the risk information. The calculation process is as follows:
[0131] NP = SP × (1 - SNR);
[0132] Where NP represents the number of pixels to be noise-added, SP is the total number of pixels in the risk area, and SNR is the signal-to-noise ratio, which ranges from [0,1].
[0133] The risk text is processed by Gaussian blurring and salt-and-pepper noise reduction through a risk text desensitization layer, and finally the financial business image after risk text desensitization is output.
[0134] The aforementioned image risk text processing method for financial business images involves inputting the financial business image into the image risk text processing model. A jump feature pyramid neural network in the text feature acquisition layer extracts the text content features corresponding to the character units within the financial business image. Then, a region extraction network in the region feature acquisition layer extracts the text region features corresponding to the aforementioned character units. Next, a correlation graph convolutional neural network in the text recognition layer identifies the text characters contained in the financial business image. Finally, a risk text recognition layer identifies the risk text in the financial business image, and a risk text desensitization layer applies Gaussian blurring and salt-and-pepper noise reduction to the risk text. The final output is the desensitized financial business image. This method can completely and accurately desensitize risk text in financial business text images.
[0135] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0136] Based on the same inventive concept, this application also provides an image risk text processing apparatus for implementing the image risk text processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image risk text processing apparatus embodiments provided below can be found in the limitations of the image risk text processing method described above, and will not be repeated here.
[0137] In one embodiment, such as Figure 8 As shown, an image risk text processing device is provided, including: a text content feature acquisition module 801, a text region feature acquisition module 802, a text character acquisition module 803, and a text image desensitization module 804, wherein:
[0138] The text content feature acquisition module 801 is used to acquire financial business text images and input the financial business text images into a pre-built text recognition model. Through the text feature acquisition layer in the text recognition model, the text content features corresponding to the character units contained in the financial business text images are obtained.
[0139] The text region feature acquisition module 802 is used to input text content features into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character units.
[0140] The text character acquisition module 803 is used to input text region features into the text recognition layer in the text recognition model to obtain the text characters contained in the financial business text image;
[0141] The text image desensitization module 804 is used to identify risk text regions from financial business text images based on text characters, and to desensitize the risk text regions to obtain desensitized financial business text images.
[0142] In one embodiment, the text region feature acquisition module 802 is further configured to input text content features to the region feature acquisition layer, acquire the character unit position region corresponding to the character unit through the region feature acquisition layer, and superimpose the character unit position region with the text content features to obtain the text region feature corresponding to the character unit.
[0143] In one embodiment, the text character acquisition module 803 is further configured to input the text region features corresponding to each character unit to the text recognition layer, and obtain the positional correlation between each character unit through the text recognition layer; based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer, the text characters contained in the financial business text image are obtained.
[0144] In one embodiment, the text character acquisition module 803 is further configured to obtain the target text region features corresponding to each character unit from the text region features corresponding to each character unit through the text recognition layer; obtain the distance information between each character unit based on the target text region features corresponding to each character unit; and obtain the positional correlation between each character unit based on the distance information between each character unit.
[0145] In one embodiment, the text content feature acquisition module 801 is further configured to extract text content features from the financial business text image at multiple scales through the text feature acquisition layer to obtain multiple initial text content features corresponding to the financial business text image; and to superimpose the multiple initial text features according to a preset superposition probability to obtain the text content features corresponding to the character unit.
[0146] In one embodiment, the text image desensitization module 804 is further configured to retrieve text characters through a pre-set financial business knowledge graph to obtain retrieval results corresponding to the text characters; based on the retrieval results, identify risk text characters in the financial business text image, and use the area where the risk text characters are located as the risk text area in the financial business text image.
[0147] In one embodiment, the text image desensitization module 804 is further configured to blur the risk text region and select a preset number of target pixels from the pixels of the blurred risk text region; change the color of the target pixels to a preset color to obtain the desensitized financial business text image.
[0148] Each module in the aforementioned image risk text processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0149] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements an image-based text processing method. The display screen can be an LCD screen or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0150] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0152] The text image of financial business is acquired and input into a pre-built text recognition model. The text feature acquisition layer in the text recognition model is used to obtain the text content features corresponding to the character units contained in the text image of financial business.
[0153] The text content features are input into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character units;
[0154] By inputting the text region features into the text recognition layer of the text recognition model, the text characters contained in the financial business text image are obtained;
[0155] Based on text characters, risk text regions are identified from financial business text images, and the risk text regions are desensitized to obtain desensitized financial business text images.
[0156] In one embodiment, when the processor executes the computer program, it further performs the following steps: inputting text content features to a region feature acquisition layer, obtaining the character unit position region corresponding to the character unit through the region feature acquisition layer; and superimposing the character unit position region with the text content features to obtain the text region features corresponding to the character unit.
[0157] In one embodiment, when the processor executes the computer program, it further implements the following steps: inputting the text region features corresponding to each character unit into the text recognition layer, obtaining the positional correlation between each character unit through the text recognition layer; and obtaining the text characters contained in the financial business text image based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer.
[0158] In one embodiment, when the processor executes the computer program, it further performs the following steps: obtaining the target text region features corresponding to each character unit from the text region features corresponding to each character unit through the text recognition layer; obtaining the distance information between each character unit based on the target text region features corresponding to each character unit; and obtaining the positional correlation between each character unit based on the distance information between each character unit.
[0159] In one embodiment, when the processor executes the computer program, it further performs the following steps: extracting text content features from the financial business text image at multiple scales through the text feature acquisition layer to obtain multiple initial text content features corresponding to the financial business text image; and superimposing the multiple initial text features according to a pre-set superposition probability to obtain the text content features corresponding to the character unit.
[0160] In one embodiment, when the processor executes the computer program, it further performs the following steps: retrieving text characters through a pre-set financial business knowledge graph to obtain the retrieval results corresponding to the text characters; based on the retrieval results, identifying risk text characters in the financial business text image, and taking the area where the risk text characters are located as the risk text area in the financial business text image.
[0161] In one embodiment, when the processor executes the computer program, it further performs the following steps: blurring the risk text region and selecting a preset number of target pixels from the pixels of the blurred risk text region; changing the color of the target pixels to a preset color to obtain a desensitized financial business text image.
[0162] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0163] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0164] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0165] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0166] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0167] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for processing image-based risk text, characterized in that, The method includes: A financial business text image is acquired and input into a pre-built text recognition model. The text feature acquisition layer in the text recognition model is used to obtain the text content features corresponding to the character units contained in the financial business text image. The text content features are input into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character unit; The text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image; Based on the text characters, risk text regions are identified from the financial business text image, and the risk text regions are desensitized to obtain a desensitized financial business text image. The number of character units is multiple; the text region features are input into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image, including: The text region features corresponding to each character unit are input into the text recognition layer, and the positional correlation between each character unit is obtained through the text recognition layer. Based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer, the text characters contained in the financial business text image are obtained. The positional relevance is represented by an adjacency matrix, and the positional relevance between each character unit is calculated through the following steps: The text region features corresponding to each character unit are multiplied by the learnable weight matrix to obtain the transformed feature representation. The initial positional correlation between each pair of character units is obtained by transposing the transformed feature representations of each pair of character units and then multiplying them. The initial positional correlation is standardized to obtain the positional correlation between each character unit.
2. The method according to claim 1, characterized in that, The step of inputting the text content features into the region feature acquisition layer of the text recognition model to obtain the text region features corresponding to the character unit includes: The text content features are input into the region feature acquisition layer, and the region feature acquisition layer obtains the character unit position region corresponding to the character unit. The text region feature corresponding to the character unit is obtained by superimposing the character unit position region with the text content feature.
3. The method according to claim 1, characterized in that, The method of obtaining the positional correlation between each character unit through the text recognition layer also includes: The text recognition layer obtains the target text region features corresponding to each character unit from the text region features corresponding to each character unit. Based on the target text region features corresponding to each character unit, the distance information between each character unit is obtained; Based on the distance information between each character unit, the positional correlation between each character unit is obtained.
4. The method according to claim 1, characterized in that, The step of obtaining the text content features corresponding to the character units contained in the financial business text image through the text feature acquisition layer in the text recognition model includes: Through the text feature acquisition layer, multi-scale text content features are extracted from the financial business text image to obtain multiple initial text content features corresponding to the financial business text image. The multiple initial text features are superimposed according to a pre-set superposition probability to obtain the text content features corresponding to the character unit.
5. The method according to claim 1, characterized in that, The step of identifying risk text regions from the financial business text image based on the text characters includes: The text characters are retrieved using a pre-set financial business knowledge graph to obtain the search results corresponding to the text characters; Based on the search results, risk text characters in the financial business text image are identified, and the area where the risk text characters are located is taken as the risk text area in the financial business text image.
6. The method according to claim 5, characterized in that, The process of desensitizing the risk text region to obtain the desensitized financial business text image includes: The risk text region is blurred, and a preset number of target pixels are selected from the pixels of the blurred risk text region. The color of the target pixel is changed to a preset color to obtain the desensitized financial business text image.
7. An image risk text processing device, characterized in that, The device includes: The text content feature acquisition module is used to acquire financial business text images and input the financial business text images into a pre-built text recognition model. Through the text feature acquisition layer in the text recognition model, the text content features corresponding to the character units contained in the financial business text images are obtained. The text region feature acquisition module is used to input the text content features into the region feature acquisition layer in the text recognition model to obtain the text region features corresponding to the character unit. The text character acquisition module is used to input the text region features into the text recognition layer of the text recognition model to obtain the text characters contained in the financial business text image; The text image desensitization module is used to identify risk text regions from the financial business text image based on the text characters, and to desensitize the risk text regions to obtain the desensitized financial business text image. The number of character units is multiple; the text character acquisition module is also used to input the text region features corresponding to each character unit to the text recognition layer, and obtain the positional correlation between each character unit through the text recognition layer; based on the text region features corresponding to each character unit, the positional correlation between each character unit, and the model parameters of the text recognition layer, the text characters contained in the financial business text image are obtained. The positional relevance is represented by an adjacency matrix, and the positional relevance between each character unit is calculated through the following steps: The text region features corresponding to each character unit are multiplied by the learnable weight matrix to obtain the transformed feature representation. The initial positional correlation between each pair of character units is obtained by transposing the transformed feature representations of each pair of character units and then multiplying them. The initial positional correlation is standardized to obtain the positional correlation between each character unit.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Character recognition method and related product
CN109583449A
Method and apparatus for inputing handwriting
CN1106548A
Method and device for desensitizing document image, electronic equipment and medium
CN112380566A