Method for processing image data and related device
By identifying and fusion of position and visual feature information on the text content in the video, the problem of insufficient accuracy in risk prediction in traditional OCR technology is solved, and more efficient risk prediction is achieved.
Patent Information
- Application Number
- CN202510142769.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
AI Technical Summary
When traditional OCR text analysis recognizes text content in video, it fails to fully consider the impact of the position of the text in the image and the visual characteristics on risk prediction, resulting in insufficient accuracy of risk prediction.
By identifying the text content in the image, text detection data, including the location information and visual feature information of text instances, and fuse these information to output risk prediction results based on the trained target model.
The accuracy of risk prediction is improved, and the ability to recognize inappropriate content is enhanced by combining text content and visual characteristics such as its position, text size or clarity in the image.
Smart Images

Figure CN120107949A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, computer-readable medium and computer program product for processing image data. Background Art
[0002] When managing video content, OCR technology is used to identify text content in videos and compare it with the keyword library in the content management system to reduce the appearance of inappropriate content. This method helps reduce the workload of manual processing and enhances the consistency and efficiency of content management. However, traditional OCR text analysis usually only predicts risks based on the identified text data, without considering the position of the identified text in the image and the influence of characteristics such as the size and color of the text. Summary of the invention
[0003] Multiple aspects of the present application provide a method, an apparatus, an electronic device, a computer-readable medium, and a computer program product for processing image data.
[0004] In one aspect of the present application, a method for processing image data is provided, wherein the method comprises:
[0005] Get the image to be processed;
[0006] By performing text recognition processing on the image to be processed, corresponding text detection data is obtained, wherein the text detection data includes position information and visual feature information of one or more recognized text instances;
[0007] Use the trained target model to output the corresponding risk prediction results based on the text information and text detection data corresponding to the image to be processed.
[0008] In one aspect of the present application, a device for processing image data is provided, wherein the device comprises:
[0009] A device for acquiring an image to be processed;
[0010] Device for obtaining corresponding text detection data by performing text recognition processing on the image to be processed, wherein the text detection data includes position information and visual feature information of one or more recognized text instances;
[0011] A device for using a trained target model to output corresponding risk prediction results based on text information and text detection data corresponding to the image to be processed.
[0012] Another aspect of the present application is an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of an embodiment of the present application.
[0013] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.
[0014] Another aspect of the present application provides a computer program product, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.
[0015] In the solution provided in the embodiment of the present application, the text content detected from the image and the position information and visual feature information obtained by text detection are fused, and the target model is trained based on the fused information, so that the target model outputs the corresponding risk prediction results based on the text content, as well as the visual features such as the position of the text in the image, text size or clarity, thereby improving the accuracy of risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0017] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0018] Figure 1 A flow chart of a method for processing image data according to an embodiment of the present application is shown;
[0019] Figure 2 A schematic diagram of an exemplary model structure according to an embodiment of the present application is shown;
[0020] Figure 3 A schematic diagram showing the structure of an apparatus for processing image data according to an embodiment of the present application is shown;
[0021] Figure 4 A schematic diagram of the structure of a device suitable for implementing the solution in the embodiment of the present application is shown.
[0022] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0024] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPU), input / output interface, network interface and memory.
[0025] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0026] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer program instructions, data structures, modules of programs or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0027] Figure 1 A flow chart of a method for processing image data according to an embodiment of the present application is shown. The method at least includes step S101, step S102 and step S103.
[0028] In actual scenarios, the execution subject of the method can be a network device, or an application running on a network device, wherein the network device includes but is not limited to a network host, a single network server, a plurality of network server sets, or a collection of computers based on cloud computing, and can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, wherein cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.
[0029] Different from the traditional way of predicting risks based on detected text data, the method according to the embodiment of the present application fuses the text content with the position information and visual feature information obtained by text detection, and trains the target model based on the fused information, so that the target model outputs the corresponding risk prediction results based on the text content, and visual features such as the position of the text in the image, text size or clarity. For example, according to the following prior knowledge: inappropriate content often appears in the lower left and upper right corners of the video; the font is small and blurred; these inappropriate contents are mostly chat messages, artistic characters, etc.; these inappropriate contents often appear in continuous conversations and have contextual relevance. The target model obtained by the training method of the embodiment of the present application can realize the learning of the above knowledge and apply it to risk prediction tasks, thereby improving the accuracy of risk prediction.
[0030] Refer to the following Figure 1 To explain, in step S101, an image to be processed is obtained.
[0031] The images to be processed may include various types of images containing text content, such as video frames and the like.
[0032] In step S102, text recognition processing is performed on the image to be processed to obtain corresponding text detection data.
[0033] The method uses OCR (Optical Character Recognition) technology or other methods for detecting text in an image to identify text instances contained in the image to be processed.
[0034] The text detection data includes location information and visual feature information of one or more identified text instances.
[0035] The location information includes but is not limited to the location coordinates of one or more detected text instances.
[0036] According to one embodiment, the method obtains the location information by the following steps:
[0037] The OCR technology is used to perform text recognition on the image to be processed to obtain boundary points of one or more recognized text instances; then, the obtained boundary points are converted into position information of one or more text instances.
[0038] According to one embodiment, the method extracts visual feature information corresponding to each text instance for the one or more recognized text instances.
[0039] The visual feature information is, but is not limited to, information indicating at least one of the following:
[0040] 1) Font type;
[0041] 2) Text direction;
[0042] 3) Text color;
[0043] 3) Text clarity; includes various information that can be used to indicate the degree of blurriness of text.
[0044] If the visual features include text sharpness, the blurriness can be assessed by analyzing the sharpness of the text edges or using specialized image processing techniques.
[0045] According to one embodiment, the text detection data also includes confidence information. The method obtains one or more text instances by performing text recognition on the image to be processed using OCR technology, and calculates the confidence information of the recognized text corresponding to each text instance.
[0046] The confidence information reflects the credibility of the recognition result and can be used as an indicator of the possibility of the recognition text being correct, thereby serving as a basis for judging the risk level of the text.
[0047] According to one embodiment, the text detection data further includes association indication information, and the method further includes step S104 and step S105.
[0048] The correlation information is used to indicate the correlation between multiple text instances detected in the image.
[0049] The correlation includes but is not limited to:
[0050] 1) Based on the relevance of the positions in the image; for example, the positions in the image are close;
[0051] 2) Relevance based on the visual features of the text; for example, having the same or similar font color, size, or clarity.
[0052] In step S104, a predetermined clustering algorithm is used to group multiple text instances obtained by text detection according to the position information and visual feature information.
[0053] The clustering algorithms include but are not limited to K-means, DBSCAN, etc.
[0054] In step S105, a combination label is assigned to each group as association indication information, wherein the combination label is used to indicate that multiple text instances belonging to a combination are associated.
[0055] Continue to refer to Figure 1 To illustrate, in step S103, the trained target model is used to output the corresponding risk prediction result based on the text information and text detection data corresponding to the image to be processed.
[0056] The risk prediction results include but are not limited to the risk level or the determination result of whether there is a risk.
[0057] According to one embodiment, the target model is a risk sub-model, which is used to output corresponding risk prediction results based on text information and text detection data of the input image.
[0058] According to one embodiment, step S103 includes step S1031 to step S1034.
[0059] In step S1031, text feature representations corresponding to the one or more text instances are obtained.
[0060] The text feature representation includes text feature representation corresponding to the text content of one or more text instances. Specifically, the text feature representation corresponding to the text content is obtained by performing word embedding.
[0061] Optionally, the text feature representation may also include text feature representations corresponding to other text information. For example, for a video image, the text feature representation may also include text feature representations corresponding to text information such as the video title, tags, and uploader information.
[0062] In step S1032, a feature representation corresponding to the text detection data is obtained.
[0063] Specifically, for the position information, visual feature information and other information contained in the text detection data, the information is converted into a vector of a fixed dimension as a feature representation corresponding to the text detection data.
[0064] In step S1033, the text feature representation and the feature representation corresponding to the text detection data are input into the target model for fusion processing through a multi-head attention mechanism to obtain fused feature information that combines the text content and the text detection data.
[0065] In step S1034, the risk prediction result output by the target model based on the fused feature information is obtained.
[0066] According to one embodiment, wherein the image to be processed includes continuous video frames, the method further includes step S106.
[0067] In step S106, text detection data of a plurality of consecutive video frames is obtained.
[0068] In step S103, the trained target model is used to output corresponding risk prediction results based on the text information and text detection data corresponding to the video frame to be processed and multiple video frames before and after it.
[0069] Specifically, for each video frame in the continuous video frames, the method uses the trained target model to obtain the risk prediction results of each video based on the text information and text detection data of the video frame. Then, by comparing and analyzing the risk prediction results of each video frame, the final risk prediction result of the video frame to be processed is obtained.
[0070] For example, the trained target model is used to predict the risk of each video frame corresponding to a video and obtain the risk score corresponding to each video frame, and the average of these risk scores is calculated. If the difference between the risk score of one of the video frames to be processed and the calculated average value is greater than a predetermined threshold, the risk score of the video frame to be processed is corrected. The correction can be based on statistical methods, such as considering the standard deviation of the score, or using more complex machine learning techniques, such as anomaly detection algorithms, to re-evaluate and adjust the risk score.
[0071] According to one embodiment, the method trains the target model through steps S107 to S109.
[0072] In step S107, corresponding text detection data is acquired by performing text recognition on a plurality of target sample images, where the text detection data includes position information and visual feature information of one or more recognized text instances.
[0073] The target sample images are used for training the model, and the target sample images may include various types of images containing inappropriate text content, such as video frames, etc.
[0074] In step S108, the text content of the one or more text instances and the text detection data are fused to obtain corresponding fused feature information.
[0075] Similar to the process of steps S1031 to S1033 above, the fused feature information of the target sample image is obtained through the following steps: obtaining the text feature representation corresponding to the one or more text instances; obtaining the feature representation corresponding to the text detection data; inputting the text feature representation and the feature representation corresponding to the text detection data into the target model to perform fusion processing through a multi-head attention mechanism to obtain fused feature information that integrates the text content and the text detection data.
[0076] In step S109, a target model is trained based on the fused feature information of multiple target sample images, so that the target model outputs a risk prediction result based on the text content in the target sample image and the corresponding text detection data through training learning.
[0077] According to one embodiment, the method further includes step S110 and step S111.
[0078] In step S110, prior knowledge information corresponding to the risk prediction task is obtained.
[0079] Specifically, the pre-stored prior knowledge information is acquired, or the prior knowledge information corresponding to the risk prediction task is collected in real time.
[0080] For example, for scenarios with risky video content, relevant prior knowledge may include: locations where high-risk text content often appears, such as the lower left or upper right corner of the video; specific visual features, such as small and blurred fonts; text content types, such as chat messages and artistic text; combination patterns, for example, appearing in combinations of multiple sentences in conversation.
[0081] In step S111, the prior knowledge information is encoded to obtain a feature representation of the prior knowledge information.
[0082] The corresponding feature representation obtained by encoding the prior knowledge information enables the relevant data of the prior knowledge information to be understood and processed by the target model. The encoding results obtained by the encoding process include but are not limited to weights and biases, rules, feature vectors, etc. For example, if a feature has a strong correlation with high-risk content, a higher weight can be assigned to the feature.
[0083] In step S109, the target model is trained by combining the fused feature information of multiple target sample images and the feature representation of the prior knowledge information, so that the target model adjusts parameters by comparing the fused feature information and the prior knowledge.
[0084] The target model in this embodiment not only learns the features extracted from the input data, but also combines prior knowledge to perform risk prediction, thereby improving the accuracy of risk prediction.
[0085] According to the method of the embodiment of the present application, the text content detected from the image and the position information and visual feature information obtained by text detection are fused, and the target model is trained based on the fused information, so that the target model outputs the corresponding risk prediction results based on the text content, as well as the visual features such as the position of the text in the image, text size or clarity, thereby improving the accuracy of risk prediction.
[0086] Refer to the following Figure 2 The exemplary model structure shown is used to illustrate the method of the embodiment of the present application.
[0087] Reference Figure 2 The model in this example is based on CLIP (Contrastive Language-Image Pre-training), which outputs corresponding risk prediction results based on the text information of the input image, as well as visual features such as the position of the text in the image, the size of the text, or the clarity of the text. The data to be processed in this example is a video image containing text on a video website.
[0088] Below Figure 2 The structure shown is introduced:
[0089] OCR: Optical Character Recognition, used to recognize and extract text information from input images.
[0090] Poster Image&Extracted Text: The poster image input to the model and the text extracted from the poster image.
[0091] CLIP: CLIP is a model that can learn by pairing images and text. The CLIP model in this example includes a visual encoder and a textual encoder, which are used to process image and text data to extract features, respectively. The visual encoder is implemented by the ViT-B / 32 model, and the textual encoder is implemented by the BERT base model (expressed as BERTbase).
[0092] Visual Feature&Textual Feature: refer to visual features and textual features respectively.
[0093] LayerNorm: Layer normalization, used to stabilize the training process and reduce internal covariate shift.
[0094] Projection: The projection layer is used to map the features of different modalities into a common space for fusion.
[0095] Multi-Head Cross Attention: Multi-head attention mechanism, which is used in cross attention, can capture different correlations in different subspaces by computing multiple different attention distributions in parallel. The multi-head attention mechanism enables the model to focus on different features of the input sequence at different levels, improving the modeling ability of complex data.
[0096] Add&Norm: Residual connections (Add) and layer normalization (Norm). Residual connections help alleviate the gradient vanishing problem, while layer normalization helps speed up the training process and improve the stability of the model.
[0097] FFN (Feed Forward Network): The feedforward network is another important component in Transformer, which performs independent and identical operations on the representation of each position of the input, usually consisting of two linear transformations.
[0098] Multi-Head Self Attention: Multi-Head Self Attention Mechanism, the multi-head self-attention mechanism splits the input sequence into multiple "heads", each head learns a different part of the input representation, and then merges these representations to capture different aspects of the information.
[0099] When the model of this example is applied to the actual risk prediction task of video content, the OCR model will be used to perform OCR recognition on the video image to be processed to obtain text detection data corresponding to the video image, where the text detection data may include the position, font type, font direction, font color, font clarity, etc. of each detected text instance.
[0100] The text information corresponding to the video image is input into the text encoder of the model for processing to obtain corresponding text features. The text information includes the text contained in the video image, the video title, the video tag, etc. The video image is input into the visual encoder of the model to obtain corresponding visual features, which may include the position, font type, font direction, font color and font clarity of each detected text instance.
[0101] For the visual features output by the visual encoder and the text features output by the text encoder, the model processes them through a multi-head attention mechanism, residual connection (Add) and layer normalization (Norm), and performs risk prediction based on the trained multi-head attention weights, and outputs the corresponding risk score.
[0102] The model in this example fuses the text content of the input image with the position information and visual feature information obtained by text detection, and outputs corresponding risk prediction results based on the text content, as well as visual features such as the position of the text in the image, text size or clarity, thereby improving the accuracy of risk prediction.
[0103] Figure 3 A schematic diagram of the structure of an apparatus for processing image data provided in an embodiment of the present application is shown.
[0104] The device includes: a device for acquiring an image to be processed (hereinafter referred to as "image acquisition device 101"), a device for acquiring corresponding text detection data by performing text recognition processing on the image to be processed (hereinafter referred to as "text detection device 102"), and a device for outputting corresponding risk prediction results based on text information and text detection data corresponding to the image to be processed using a trained target model (hereinafter referred to as "risk prediction device 103").
[0105] Reference Figure 3 , the image acquisition device 101 acquires the image to be processed.
[0106] The images to be processed may include various types of images containing text content, such as video frames and the like.
[0107] The text detection device 102 performs text recognition processing on the image to be processed to obtain corresponding text detection data.
[0108] The method uses OCR (Optical Character Recognition) technology or other methods for detecting text in an image to identify text instances contained in the image to be processed.
[0109] The text detection data includes location information and visual feature information of one or more identified text instances.
[0110] The location information includes but is not limited to the location coordinates of one or more detected text instances.
[0111] According to one embodiment, the data acquisition device 101 obtains the location information by performing the following operations:
[0112] The OCR technology is used to perform text recognition on the image to be processed to obtain boundary points of one or more recognized text instances; then, the obtained boundary points are converted into position information of one or more text instances.
[0113] According to one embodiment, the data acquisition device 101 extracts visual feature information corresponding to each of the identified one or more text instances.
[0114] The visual feature information is, but is not limited to, information indicating at least one of the following:
[0115] 1) Font type;
[0116] 2) Text direction;
[0117] 3) Text color;
[0118] 3) Text clarity; includes various information that can be used to indicate the degree of blurriness of text.
[0119] If the visual features include text sharpness, the blurriness can be assessed by analyzing the sharpness of the text edges or using specialized image processing techniques.
[0120] According to an embodiment, the text detection data further includes confidence information. The data acquisition device 101 obtains one or more text instances by performing text recognition on the image to be processed using OCR technology, and calculates the confidence information of the recognized text corresponding to each text instance.
[0121] The confidence information reflects the credibility of the recognition result and can be used as an indicator of the possibility of correct text recognition, thereby serving as a basis for assessing the risk level of the text.
[0122] According to one embodiment, the text detection data further includes association indication information, and the device further includes text grouping means and label assigning means.
[0123] The correlation information is used to indicate the correlation between multiple text instances detected in the image.
[0124] The correlation includes but is not limited to:
[0125] 1) Based on the relevance of the positions in the image; for example, the positions in the image are close;
[0126] 2) Relevance based on the visual features of the text; for example, having the same or similar font color, size, or clarity.
[0127] The text grouping device uses a predetermined clustering algorithm to group multiple text instances obtained by text detection according to the position information and visual feature information.
[0128] The clustering algorithms include but are not limited to K-means, DBSCAN, etc.
[0129] The label assigning device assigns a combination label to each group as association indication information, wherein the combination label is used to indicate that multiple text instances belonging to a combination are associated.
[0130] Continue to refer to Figure 3 To illustrate, the risk prediction device 103 uses the trained target model to output corresponding risk prediction results based on the text information and text detection data corresponding to the image to be processed.
[0131] The risk prediction results include but are not limited to the risk level or the determination result of whether there is a risk.
[0132] According to one embodiment, the target model is a risk sub-model, which is used to output corresponding risk prediction results based on text information and text detection data of the input image.
[0133] According to one embodiment, the step S103 further includes a text feature acquisition device, a visual feature acquisition device, a feature fusion processing device and a prediction result acquisition device.
[0134] The text feature acquisition device acquires text feature representations corresponding to the one or more text instances.
[0135] The text feature representation includes text feature representation corresponding to the text content of one or more text instances. Specifically, the text feature representation corresponding to the text content is obtained by performing word embedding.
[0136] Optionally, the text feature representation may also include text feature representations corresponding to other text information. For example, for a video image, the text feature representation may also include text feature representations corresponding to text information such as the video title, tags, and uploader information.
[0137] The visual feature acquisition device acquires the feature representation corresponding to the text detection data.
[0138] Specifically, for the position information, visual feature information and other information contained in the text detection data, the information is converted into a vector of a fixed dimension as a feature representation corresponding to the text detection data.
[0139] The feature fusion processing device inputs the text feature representation and the feature representation corresponding to the text detection data into the target model to perform fusion processing through a multi-head attention mechanism to obtain fused feature information that combines the text content and the text detection data.
[0140] The prediction result acquisition device acquires the risk prediction result output by the target model based on the fused feature information.
[0141] According to one embodiment, the image to be processed includes continuous video frames, and the device also includes a context acquisition device.
[0142] The context acquisition device acquires text detection data of multiple consecutive video frames.
[0143] The risk prediction device 103 uses the trained target model to output corresponding risk prediction results based on the text information and text detection data corresponding to the video frame to be processed and multiple video frames before and after it.
[0144] Specifically, for each video frame in the continuous video frames, the risk prediction device 103 uses the trained target model to obtain the risk prediction results of each video based on the text information and text detection data of the video frame. Then, by comparing and analyzing the risk prediction results of each video frame, the final risk prediction result of the video frame to be processed is obtained.
[0145] For example, the trained target model is used to predict the risk of each video frame corresponding to a video and obtain the risk score corresponding to each video frame, and the average of these risk scores is calculated. If the difference between the risk score of one of the video frames to be processed and the calculated average value is greater than a predetermined threshold, the risk score of the video frame to be processed is corrected. The correction can be based on statistical methods, such as considering the standard deviation of the score, or using more complex machine learning techniques, such as anomaly detection algorithms, to re-evaluate and adjust the risk score.
[0146] According to one embodiment, the apparatus comprises a model training apparatus.
[0147] The model training device performs text recognition on multiple target sample images to obtain corresponding text detection data, where the text detection data includes location information and visual feature information of one or more recognized text instances.
[0148] The target sample images are used for training the model, and the target sample images may include various types of images containing inappropriate text content, such as video frames, etc.
[0149] Next, the model training device fuses the text content and text detection data of the one or more text instances to obtain corresponding fusion feature information.
[0150] Similar to the process of steps S1031 to S1033 above, the fused feature information of the target sample image is obtained through the following steps: obtaining the text feature representation corresponding to the one or more text instances; obtaining the feature representation corresponding to the text detection data; inputting the text feature representation and the feature representation corresponding to the text detection data into the target model to perform fusion processing through a multi-head attention mechanism to obtain fused feature information that integrates the text content and the text detection data.
[0151] Next, the model training device trains the target model based on the fused feature information of the multiple target sample images, so that the target model outputs the risk prediction result based on the text content in the target sample image and the corresponding text detection data through training learning.
[0152] According to one embodiment, the device further comprises a priori acquisition means and a priori encoding means.
[0153] The prior acquisition device acquires prior knowledge information corresponding to the risk prediction task.
[0154] Specifically, the pre-stored prior knowledge information is acquired, or the prior knowledge information corresponding to the risk prediction task is collected in real time.
[0155] For example, for scenarios with risky video content, relevant prior knowledge may include: locations where high-risk text content often appears, such as the lower left or upper right corner of the video; specific visual features, such as small and blurred fonts; text content types, such as chat messages and artistic text; combination patterns, for example, appearing in combinations of multiple sentences in conversation.
[0156] The priori encoding device obtains the feature representation of the priori knowledge information by encoding the priori knowledge information.
[0157] The corresponding feature representation obtained by encoding the prior knowledge information enables the relevant data of the prior knowledge information to be understood and processed by the target model. The encoding results obtained by the encoding process include but are not limited to weights and biases, rules, feature vectors, etc. For example, if a feature has a strong correlation with high-risk content, a higher weight can be assigned to the feature.
[0158] The model training device combines the fused feature information of multiple target sample images and the feature representation of the prior knowledge information to train the target model, so that the target model adjusts parameters by comparing the fused feature information and the prior knowledge.
[0159] The target model in this embodiment not only learns the features extracted from the input data, but also combines prior knowledge to perform risk prediction, thereby improving the accuracy of risk prediction.
[0160] According to the device of the embodiment of the present application, by fusing the text content detected from the image with the position information and visual feature information obtained by text detection, and training the target model based on the fused information, the target model is enabled to output corresponding risk prediction results based on the text content, as well as visual features such as the position of the text in the image, text size or clarity, thereby improving the accuracy of risk prediction.
[0161] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application, and the method corresponding to the electronic device may be the method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.
[0162] The electronic device may be a user device, or a device formed by integrating a user device and a network device through a network, or may be an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablet computers, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, which can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.
[0163] Figure 4 The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown, and the device 1200 includes a central processing unit (CPU, Central Processing Unit) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM, Random Access Memory) 1203. In RAM1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O, Input / Output) interface 1205 is also connected to bus 1204.
[0164] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker, etc.; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet.
[0165] In particular, the methods and / or embodiments in the embodiments of the present application may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above functions defined in the method of the present application are executed.
[0166] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application described above.
[0167] Specifically, the present embodiment may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device, or device.
[0168] Computer readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0169] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0170] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0171] The flow chart or block diagram in the accompanying drawings shows the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0172] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0173] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0174] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0176] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0178] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
Claims
1. A method for processing image data, wherein: The method comprises: Get the image to be processed; By performing text recognition processing on the image to be processed, corresponding text detection data is obtained, wherein the text detection data includes position information and visual feature information of one or more recognized text instances; Use the trained target model to output the corresponding risk prediction results based on the text information and text detection data corresponding to the image to be processed.
2. The method according to claim 1, wherein: The method of using the trained target prediction model to output the corresponding risk prediction results based on the text information and text detection data corresponding to the image to be processed includes: Obtaining text feature representations corresponding to the one or more text instances; Obtaining feature representation corresponding to the text detection data; Inputting the text feature representation and the feature representation corresponding to the text detection data into the target model to perform fusion processing through a multi-head attention mechanism to obtain fused feature information that combines the text content and the text detection data; Obtain the risk prediction result output by the target model based on the fused feature information.
3. The method according to claim 1, wherein: The image to be processed is a continuous video frame, and the method further includes: Obtain text detection data for multiple consecutive video frames; The method of using the trained target model to output corresponding risk prediction results based on the text information and text detection data corresponding to the image to be processed includes: The trained target prediction model is used to output the corresponding risk prediction results based on the text information and text detection data corresponding to the video frame to be processed and multiple video frames before and after it.
4. The method according to claim 1, wherein: The obtaining of corresponding text detection data by performing text recognition processing on the image to be processed includes: Use OCR technology to perform text recognition on the image to be processed, and obtain the boundary points of one or more recognized text instances; The obtained boundary points are converted into the location information of one or more text instances.
5. The method according to any one of claims 1 to 4, wherein: The text detection data also includes associated indication information, and the method further includes: Using a predetermined clustering algorithm, grouping multiple text instances obtained by text detection according to the position information and visual feature information; A combination label is allocated to each group as association indication information, wherein the combination label is used to indicate that multiple text instances belonging to a combination are associated.
6. The method according to any one of claims 1 to 4, wherein: The method further comprises: By performing text recognition on a plurality of target sample images, corresponding text detection data is obtained, wherein the text detection data includes position information and visual feature information of one or more recognized text instances; Fusing the text content of the one or more text instances with the text detection data to obtain corresponding fusion feature information; The target model is trained based on the fused feature information of multiple target sample images, so that the target model outputs risk prediction results based on the text content in the target sample images and the corresponding text detection data through training learning.
7. The method according to claim 6, wherein: The method further comprises: Obtaining prior knowledge information corresponding to the risk prediction task; By encoding the prior knowledge information, a feature representation of the prior knowledge information is obtained; The target model training based on the fusion feature information of multiple target sample images includes: The target model is trained by combining the fused feature information of multiple target sample images and the feature representation of the prior knowledge information, so that the target model adjusts parameters by comparing the fused feature information and the prior knowledge.
8. A device for processing image data, wherein: The device comprises: A device for acquiring an image to be processed; A device for obtaining corresponding text detection data by performing text recognition processing on the image to be processed, wherein the text detection data includes position information and visual feature information of one or more recognized text instances; A device for using a trained target model to output corresponding risk prediction results based on text information and text detection data corresponding to the image to be processed.
9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.