Risk prediction method and device, electronic equipment, medium and program product
By detecting and segmenting the text area of the image and fusing the image vector for risk prediction, the problem of traditional OCR technology ignoring the low confidence text and visual characteristics is improved, and the accuracy and efficiency of risk prediction are improved.
Patent Information
- Application Number
- CN202510142760.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional OCR technology is prone to ignore text with low confidence when processing text, resulting in the loss of important information, affecting the accuracy of risk prediction, and at the same time ignores the visual characteristics of text.
By detecting text area and segmenting the image to be processed, the image vector of each character is obtained, and the overall image vector is fused with the character image vector, and the risk prediction results are output using the trained model.
It avoids misleading information caused by text recognition errors, improves the accuracy of risk prediction, and improves the efficiency and stability of the model by processing image information.
Smart Images

Figure CN120107948A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a risk prediction method, device, electronic device, computer-readable medium and computer program product. Background Art
[0002] Traditional OCR (Optical Character Recognition) technology is a technology that converts text in an image into electronic text. OCR technology can be used to identify and extract text content in images or videos, and to determine whether the text content is appropriate through risk prediction. However, traditional OCR text processing methods have some limitations. First, the OCR recognition model usually only outputs the text with the highest confidence, which means that if an image contains multiple words, but only some of the words have high recognition confidence, then other words with lower confidence may be ignored. Therefore, this method may lead to the loss of important information, thereby affecting the accuracy of risk prediction. Secondly, the method of risk prediction based on OCR technology mainly focuses on the text content of the input image, ignoring the visual features of the text, such as font, color, size, etc. Summary of the invention
[0003] Multiple aspects of the present application provide a risk prediction method, apparatus, electronic device, computer-readable medium, and computer program product.
[0004] In one aspect of the present application, a risk prediction method is provided, wherein the method comprises:
[0005] Perform text region detection on the image to be processed to obtain one or more text regions containing text in the image to be processed;
[0006] Based on the character images corresponding to the characters in the text area, a corresponding character string vector is obtained, wherein the character string vector is used to represent the total vector of the image vectors of the detected characters;
[0007] Based on the entire image to be processed, a corresponding image overall vector is obtained, wherein the image overall vector is used to represent the entire image to be processed;
[0008] Fusing the character string vector with the overall image vector to obtain an image fusion vector;
[0009] The trained target model is used to output the corresponding risk prediction result based on the image fusion vector of the image to be processed.
[0010] In one aspect of the present application, a risk prediction device is provided, wherein the device comprises:
[0011] A device for performing text region detection on an image to be processed to obtain one or more text regions containing text in the image to be processed;
[0012] A device for obtaining a corresponding character string vector based on the character image corresponding to each character in the text area, wherein the character string vector is used to represent the total vector of the image vector of each detected character;
[0013] A device for obtaining a corresponding image overall vector based on the entire image to be processed, wherein the image overall vector is used to represent the entire image to be processed;
[0014] A device for fusing the character string vector and the overall image vector to obtain an image fusion vector;
[0015] A device for outputting corresponding risk prediction results based on the image fusion vector of the image to be processed through a trained target model.
[0016] Another aspect of the present application is an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of an embodiment of the present application.
[0017] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.
[0018] Another aspect of the present application provides a computer program product, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.
[0019] In the solution provided in the embodiment of the present application, text area detection and image segmentation are performed on the image to be processed to obtain images corresponding to each character contained in the image to be processed, and the overall image vector of the image to be processed and the image vector of each character are fused, and the trained model is used to output the risk prediction result based on the fused image vector. This risk prediction method does not require the recognition of the text content contained in the input image, thereby avoiding misleading information caused by text recognition errors. Since the input to the model encoder for processing is all image information, all images can be converted using the image encoder, which facilitates training and reasoning, improves efficiency, and reduces the noise caused by the fusion of multimodal information in the text recognition method, thereby improving the accuracy of the model risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0022] Figure 1 A schematic diagram of a risk prediction method provided in an embodiment of the present application is shown;
[0023] Figure 2 A schematic diagram of an exemplary model structure according to an embodiment of the present application is shown;
[0024] Figure 3 A schematic diagram of the structure of a device for risk prediction according to an embodiment of the present application is shown;
[0025] Figure 4 A schematic diagram of the structure of a device suitable for implementing the solution in the embodiment of the present application is shown.
[0026] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0028] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPU), input / output interface, network interface and memory.
[0029] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0030] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer program instructions, data structures, modules of programs or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0031] Figure 1 The flowchart of a risk prediction method provided by an embodiment of the present application is shown. The method at least includes step S101, step S102, step S103, step S104 and step S105.
[0032] In actual scenarios, the execution subject of the method can be a network device, or an application running on a network device. The network device includes but is not limited to a network host, a single network server, a set of multiple network servers, or a set of computers based on cloud computing, which can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, which is a virtual computer composed of a group of loosely coupled computers.
[0033] In the scenario of risk prediction for video manuscripts, the traditional method is to train a risk prediction model based on text recognition. Since the model only outputs the text with the highest confidence, other text image information is often lost. The embodiment of the present application performs text area detection and image segmentation on the video frame to be processed to obtain the image corresponding to each character contained in the video frame, and fuses the overall image vector of the video frame with the image vector of each character, and uses the trained model to output the risk prediction result based on the fused image vector. This method does not have the step of recognizing text content in the traditional method, but outputs text visual information as an intermediate result to avoid the impact caused by text content recognition.
[0034] Reference Figure 1 In step S101, text region detection is performed on the image to be processed to obtain one or more text regions containing text in the image to be processed.
[0035] The images to be processed may include various types of images containing text content, such as video frames and the like.
[0036] The method determines the text area of the image to be processed by using OCR (Optical Character Recognition) technology or other methods for detecting text in an image.
[0037] Optionally, before performing text detection, the image to be processed is preprocessed. The preprocessing includes but is not limited to at least any one of the following:
[0038] 1) Denoising: used to eliminate noise in the image and improve image quality;
[0039] 2) Binarization: used to convert the image into black and white, so that the contrast between the text and the background is higher, which is convenient for subsequent processing;
[0040] 3) Rotation correction processing: If the text in the image is tilted, rotation correction is required to make the text horizontal or vertical.
[0041] In step S102, based on the character images corresponding to the characters in the text area, corresponding character string vectors are obtained, where the character string vectors are used to represent the total vector of the image vectors of the detected characters.
[0042] Specifically, the step S102 includes steps S1021 to S1023.
[0043] In step S1021, for each text region, image segmentation is performed on the text region to obtain character images corresponding to each character in the text region.
[0044] According to one embodiment, step S1021 further includes step S10211 and step S10212.
[0045] In step S10211, the text area is segmented into multiple parts by performing image segmentation on the text area, each part corresponding to a separate character, wherein a text detection algorithm is used to identify text lines and single characters in the image.
[0046] Optionally, the image segmentation result is corrected. For example, after image segmentation, multiple characters written together may be obtained, and further segmentation is required to obtain separate characters.
[0047] In step S10212, based on the image segmentation result, the character image corresponding to each character is obtained.
[0048] Optionally, for each character, the method uses the picture corresponding to the text box area of the character as the character picture. For example, for the detected character "cat", the corresponding character picture is a cat.
[0049] Optionally, the method performs cropping processing based on the pictures corresponding to the text box areas of the respective characters according to predetermined sizes and shapes, thereby obtaining a plurality of character images with the same shapes and sizes.
[0050] In step S1022, each character image is converted into a corresponding character image vector.
[0051] Optionally, the method obtains a character image vector corresponding to each character image by pixel value conversion or the like.
[0052] Optionally, the method extracts image feature vectors of each character image through a feature extraction network, thereby using the obtained image feature vectors as character image vectors. For example, the image features can be mapped to the semantic space of the text through a multi-layer perceptron (MLP) structure to obtain the corresponding image feature vectors.
[0053] In step S1023, the character image vectors corresponding to each character image are merged to obtain a corresponding character string vector.
[0054] The method may combine the character image vectors corresponding to each character through a concatenation operation. Alternatively, the method may combine the character image vectors through a more complex structure, such as a recurrent neural network (RNN), to capture the sequential relationship between characters.
[0055] According to the first example of the present application, Ii represents a set of character images obtained by text detection. Then, for each character image, it is converted into a character image vector vi by a vectorization function f, vi=f(Ii). Among them, the function f can be a simple pixel value conversion or a complex feature extraction network. Then, all the character image vectors {v1, v2,…, vn} are merged into a total vector V of a sentence, V=g({v1, v2,…, vn}), where g is a function for merging all character vectors into a total vector representation V of a sentence.
[0056] Continue to refer to Figure 1 To illustrate, in step S103, based on the entire image to be processed, a corresponding overall image vector is obtained, and the overall image vector is used to represent the overall image to be processed.
[0057] Similar to the above operation of obtaining the character image vector, the image to be processed is converted into a vector as a whole. The whole image vector represents the whole content of the image to be processed.
[0058] Optionally, a convolutional neural network (CNN) may be used to extract image features corresponding to the image to be processed as the overall image vector.
[0059] Continuing to explain the first example, for the image to be processed I, by V image = h(I) to get the overall image vector V image =h(I). The function is used to convert the entire image I into an image vector.
[0060] In step S104, the character string vector and the overall image vector are fused to obtain an image fusion vector.
[0061] Optionally, the fusion process includes concatenation, wherein the method concatenates the string vector and the whole picture vector through a concatenation operation to obtain an image fusion vector. The image fusion vector integrates text and visual information to obtain a more comprehensive image content representation.
[0062] Optionally, the fusion process may also include other feature fusion methods, such as weighted summation.
[0063] Continuing with the first example, for the obtained string vector V and the image overall vector V image , through V final =φ(V,V image ) to get the fusion vector. Among them, φ is a fusion function that combines the string vector V and the image vector V image Fusion is performed to obtain the fused vector representation V final .
[0064] In step S105, the trained target model is used to output a corresponding risk prediction result based on the image fusion vector of the image to be processed.
[0065] According to one embodiment, the target model is a risk sub-model, which is used to output corresponding risk prediction results based on the image fusion features of the input image.
[0066] The risk prediction results include but are not limited to the risk level or the determination result of whether there is a risk.
[0067] Optionally, the target model predicts the risk level of the input image through a classifier structure, such as a support vector machine (SVM) or a neural network.
[0068] Continuing to explain the first example above, based on the obtained image fusion vector V final , through P = C (V final ) to get the risk prediction result. Among them, P represents the risk assessment value of the image, C represents the classification model, which is based on the image fusion vector V final To predict the risk level of the input image.
[0069] According to one embodiment, the method trains the target model through steps S106 and S107:
[0070] In step S106, image fusion vectors corresponding to multiple target sample images are obtained.
[0071] The target sample images are used for training the model, and the target sample images may include various types of images containing inappropriate text content, such as video frames, etc.
[0072] Similar to the process of steps S101 to S104 above, the image fusion features of the target sample image are obtained through the following steps: text area detection is performed on the target sample image to obtain one or more text areas containing text in the target sample image; based on the character images corresponding to the characters in the text area, a corresponding string vector is obtained, and the string vector is used to represent the total vector of the image vectors of the detected characters; based on the entire target sample image, a corresponding overall image vector is obtained, and the overall image vector is used to represent the overall image to be processed; the string vector and the overall image vector are fused to obtain an image fusion vector.
[0073] In step S107, a target model is trained based on the image fusion vectors of the plurality of target sample images, so that the target model learns to output corresponding risk prediction results based on the image fusion vectors.
[0074] The risk prediction results include but are not limited to the risk level or the determination result of whether there is a risk.
[0075] According to the method of the embodiment of the present application, text area detection and image segmentation are performed on the image to be processed to obtain images corresponding to each character contained in the image to be processed, and the overall image vector of the image to be processed and the image vector of each character are fused, and the trained model is used to output the risk prediction result based on the fused image vector. This risk prediction method does not require the recognition of the text content contained in the input image, thereby avoiding misleading information caused by text recognition errors. Since the input to the model encoder for processing is all image information, it is possible to use the image encoder to convert all images, which facilitates training and reasoning, improves efficiency, and reduces the noise caused by the fusion of multimodal information in the text recognition method, thereby improving the accuracy of the model risk prediction.
[0076] Refer to the following Figure 2 The exemplary model structure shown is used to illustrate the embodiments of the present application.
[0077] Reference Figure 2 , the model in this example is used to perform risk prediction tasks to predict the risk of input images.
[0078] For an input image, a character image string and image information corresponding to the input image are obtained, wherein the character image string includes images of each character obtained by performing OCR detection and image segmentation on the input image.
[0079] Next, the character image string is input to the vocabulary encoder for processing, and the vocabulary encoder obtains a string vector, which is used to represent the total vector of the image vectors of each detected character. The image information is input to the image encoder for processing, and the image encoder obtains an overall image vector, which is used to represent the overall input image.
[0080] The obtained string vector and the overall image vector are fused and spliced through concatenation.
[0081] Next, the fused and concatenated vectors are input into a multi-layer perceptron (MLP), which is a feed-forward neural network used to further process the fused and concatenated vectors and generate the final risk prediction results.
[0082] The model in this example outputs the risk prediction result based on the fusion vector of the input image and the image vector of each character. This method does not need to recognize the text content contained in the input image, avoiding the impact caused by text recognition errors.
[0083] Figure 3 A schematic structural diagram of a device for risk prediction according to an embodiment of the present application is shown.
[0084] The device includes: a device for performing text area detection on an image to be processed to obtain one or more text areas containing text in the image to be processed (hereinafter referred to as "text area detection device 101"), a device for obtaining a corresponding character string vector based on the character images corresponding to each character in the text area, and the character string vector is used to represent the total vector of the image vectors of each detected character (hereinafter referred to as "character vector acquisition device 102"), a device for obtaining a corresponding image overall vector based on the entire image to be processed, and the image overall vector is used to represent the entire image to be processed (hereinafter referred to as "overall vector acquisition device 103"), a device for fusing the character string vector and the image overall vector to obtain an image fusion vector (hereinafter referred to as "vector fusion device 104"), and a device for outputting a corresponding risk prediction result based on the image fusion vector of the image to be processed through a trained target model (hereinafter referred to as "risk prediction device 105").
[0085] Reference Figure 3 The text region detection device 101 performs text region detection on the image to be processed to obtain one or more text regions containing text in the image to be processed.
[0086] The images to be processed may include various types of images containing text content.
[0087] The method determines the text area of the image to be processed by using OCR (Optical Character Recognition) technology or other methods for detecting text in an image.
[0088] Optionally, before performing text detection, the image to be processed is preprocessed. The preprocessing includes but is not limited to at least any one of the following:
[0089] 1) Denoising: used to eliminate noise in the image and improve image quality;
[0090] 2) Binarization: used to convert the image into black and white, so that the contrast between the text and the background is higher, which is convenient for subsequent processing;
[0091] 3) Rotation correction processing: If the text in the image is tilted, rotation correction is required to make the text horizontal or vertical.
[0092] The character vector acquisition device 102 obtains a corresponding character string vector based on the character image corresponding to each character in the text area, and the character string vector is used to represent the total vector of the image vectors of each detected character.
[0093] Specifically, the character vector acquisition device 102 includes a character image acquisition device, a character vector conversion device and a character vector merging device.
[0094] The character image acquisition device obtains character images corresponding to each character in the text area by performing image segmentation on the text area.
[0095] According to one embodiment, the character image acquisition device divides the text area into multiple parts by performing image segmentation on the text area, each part corresponding to a separate character, wherein a text detection algorithm is used to identify text lines and single characters in the image.
[0096] Optionally, the image segmentation result is corrected. For example, after image segmentation, multiple characters written together may be obtained, and further segmentation is required to obtain separate characters.
[0097] Next, the character image acquisition device obtains the character image corresponding to each character based on the image segmentation result.
[0098] Optionally, for each character, the character image acquisition device uses the picture corresponding to the text box area of the character as the character picture. For example, for the detected character "cat", the corresponding character picture is.
[0099] Optionally, the character image acquisition device performs cropping processing based on the picture corresponding to the text box area of each character according to a predetermined size and shape, thereby obtaining a plurality of character images with the same shape and size.
[0100] The character vector conversion device converts each character image into a corresponding character image vector.
[0101] Optionally, the character vector conversion device obtains the character image vector corresponding to each character image by pixel value conversion or the like.
[0102] Optionally, the character vector conversion device extracts the image feature vector of each character image through a feature extraction network, thereby using the obtained image feature vector as the character image vector. For example, the image feature can be mapped to the semantic space of the text through a multi-layer perceptron (MLP) structure to obtain the corresponding image feature vector.
[0103] The character vector merging device merges the character image vectors corresponding to each character image to obtain a corresponding character string vector.
[0104] The character vector merging device may merge the character image vectors corresponding to each character through a concatenation operation. Alternatively, the method may merge the character image vectors through a more complex structure, such as a recurrent neural network (RNN), to capture the sequential relationship between characters.
[0105] According to the first example of the present application, Ii represents a set of character images obtained by text detection. Then, for each character image, it is converted into a character image vector vi by a vectorization function f, vi=f(Ii). Among them, the function f can be a simple pixel value conversion or a complex feature extraction network. Then, all the character image vectors {v1, v2,…, vn} are merged into a total vector V of a sentence, V=g({v1, v2,…, vn}), where g is a function for merging all character vectors into a total vector representation V of a sentence.
[0106] Continue to refer to Figure 3 To explain, the overall vector acquisition device 103 acquires the corresponding image overall vector based on the entire image to be processed, and the image overall vector is used to represent the overall image to be processed.
[0107] Similar to the above operation of obtaining the character image vector, the image to be processed is converted into a vector as a whole. The whole image vector represents the whole content of the image to be processed.
[0108] Optionally, the overall vector acquisition device 103 may extract image features corresponding to the image to be processed by a convolutional neural network (CNN) as the overall image vector.
[0109] Continuing to explain the first example, for the image to be processed I, by V image = h(I) to get the overall image vector V image =h(I). The function is used to convert the entire image I into an image vector.
[0110] The vector fusion device 104 fuses the character string vector and the overall image vector to obtain an image fusion vector.
[0111] Optionally, the fusion process includes concatenation, and the vector fusion device 104 concatenates the string vector and the overall picture vector through a concatenation operation to obtain an image fusion vector. The image fusion vector integrates text and visual information to obtain a more comprehensive image content representation.
[0112] Optionally, the fusion process may also include other feature fusion methods, such as weighted summation.
[0113] Continuing with the first example, for the obtained string vector V and the image overall vector V image , through V final =φ(V,V image ) to get the fusion vector. Among them, φ is a fusion function that combines the string vector V and the image vector Vimage Fusion is performed to obtain the fused vector representation V final .
[0114] The risk prediction device 105 outputs a corresponding risk prediction result based on the image fusion vector of the image to be processed through the trained target model.
[0115] According to one embodiment, the target model is a risk sub-model, which is used to output corresponding risk prediction results based on the image fusion features of the input image.
[0116] The risk prediction results include but are not limited to the risk level or the determination result of whether there is a risk.
[0117] Optionally, the target model predicts the risk level of the input image through a classifier structure, such as a support vector machine (SVM) or a neural network.
[0118] Continuing to explain the first example above, based on the obtained image fusion vector V final , through P = C (V final ) to get the risk prediction result. Among them, P represents the risk assessment value of the image, C represents the classification model, which is based on the image fusion vector V final To predict the risk level of the input image.
[0119] According to one embodiment, the device further comprises a model training device.
[0120] The model training device obtains image fusion vectors corresponding to multiple target input images.
[0121] The target input image is used to train the model, and the target input image may include various types of images containing text content, such as video frames, etc.
[0122] Similar to the process of steps S101 to S104 above, the image fusion features of the target sample image are obtained by the following operations: text area detection is performed on the target sample image to obtain one or more text areas containing text in the target sample image; based on the character images corresponding to the characters in the text area, a corresponding string vector is obtained, and the string vector is used to represent the total vector of the image vectors of the detected characters; based on the entire target sample image, a corresponding overall image vector is obtained, and the overall image vector is used to represent the overall image to be processed; the string vector and the overall image vector are fused to obtain an image fusion vector.
[0123] Next, the model training device trains the target model based on the image fusion vectors of multiple target input images, so that the target model learns to output corresponding risk prediction results based on the image fusion vectors.
[0124] The risk prediction results include but are not limited to the risk level or the determination result of whether there is a risk.
[0125] According to the device of the embodiment of the present application, text area detection and image segmentation are performed on the image to be processed to obtain images corresponding to each character contained in the image to be processed, and the overall image vector of the image to be processed and the image vector of each character are fused, and the trained model is used to output the risk prediction result based on the fused image vector. This risk prediction method does not require recognition of the text content contained in the input image, thereby avoiding misleading information caused by text recognition errors. Since the input to the model encoder for processing is all image information, the image encoder can be used to convert all images, which facilitates training and reasoning, improves efficiency, and reduces the noise caused by the fusion of multimodal information in the text recognition method, thereby improving the accuracy of the model risk prediction.
[0126] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application, and the method corresponding to the electronic device may be the method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.
[0127] The electronic device may be a user device, or a device formed by integrating a user device and a network device through a network, or may be an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablet computers, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, which can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.
[0128] Figure 4The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown, and the device 1200 includes a central processing unit (CPU, Central Processing Unit) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM, Random Access Memory) 1203. In RAM1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O, Input / Output) interface 1205 is also connected to bus 1204.
[0129] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker, etc.; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet.
[0130] In particular, the methods and / or embodiments in the embodiments of the present application may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above functions defined in the method of the present application are executed.
[0131] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application described above.
[0132] Specifically, the present embodiment may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device, or device.
[0133] Computer readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0134] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0135] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0136] The flow chart or block diagram in the accompanying drawings shows the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0138] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0139] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0141] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0143] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
Claims
1. A risk prediction method, wherein: The method comprises: Perform text region detection on the image to be processed to obtain one or more text regions containing text in the image to be processed; Based on the character images corresponding to the characters in the text area, a corresponding character string vector is obtained, wherein the character string vector is used to represent the total vector of the image vectors of the detected characters; Based on the entire image to be processed, a corresponding image overall vector is obtained, wherein the image overall vector is used to represent the entire image to be processed; Fusing the character string vector with the overall image vector to obtain an image fusion vector; The trained target model is used to output the corresponding risk prediction result based on the image fusion vector of the image to be processed.
2. The method according to claim 1, wherein: The obtaining of the corresponding character string vector based on the character image corresponding to each character in the text area includes: For each text region, the character image corresponding to each character in the text region is obtained by performing image segmentation on the text region; Convert each character image into a corresponding character image vector respectively; The corresponding character string vector is obtained by merging the character image vector and the character picture vector corresponding to each character image.
3. The method according to claim 2, wherein: For each text region, the character images corresponding to the characters in the text region are obtained by performing image segmentation on the text region, including: By performing image segmentation on the text area, the text area is divided into a plurality of parts, each part corresponding to a separate character; Based on the image segmentation result, the character image corresponding to each character is obtained.
4. The method according to claim 3, wherein: The obtaining of character images corresponding to each character based on the image segmentation result includes: Based on the pictures corresponding to the text box areas of the respective characters, the pictures are cut out according to a predetermined size and shape, thereby obtaining a plurality of character images having the same shape and size.
5. The method according to any one of claims 1 to 4, wherein: The method further comprises: Obtain image fusion vectors corresponding to multiple target sample images; The target model is trained based on the image fusion vectors of multiple target sample images, so that the target model learns to output corresponding risk prediction results based on the image fusion vectors.
6. The method according to claim 5, wherein: The step of obtaining image fusion vectors corresponding to a plurality of target sample images comprises: Performing text region detection on the target sample image to obtain one or more text regions containing text in the target sample image; Based on the character images corresponding to the characters in the text area, a corresponding character string vector is obtained, wherein the character string vector is used to represent the total vector of the image vectors of the detected characters; Based on the entire target sample image, a corresponding image overall vector is obtained, wherein the image overall vector is used to represent the overall image to be processed; The character string vector and the overall image vector are fused to obtain an image fusion vector corresponding to the target sample image.
7. A device for risk prediction, wherein: The device comprises: A device for performing text region detection on an image to be processed to obtain one or more text regions containing text in the image to be processed; A device for obtaining a corresponding character string vector based on the character image corresponding to each character in the text area, wherein the character string vector is used to represent the total vector of the image vector of each detected character; A device for obtaining a corresponding image overall vector based on the entire image to be processed, wherein the image overall vector is used to represent the entire image to be processed; A device for fusing the character string vector and the overall image vector to obtain an image fusion vector; A device for outputting corresponding risk prediction results based on the image fusion vector of the image to be processed through a trained target model.
8. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.