OCR (Optical Character Recognition) method and system based on deep learning
By using deep learning technology to identify and analyze negative tags, the problem of detecting and analyzing the impact of negative tags in user-written content is solved, thereby improving content quality and user experience.
Patent Information
- Application Number
- CN202511012847.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies make it difficult to detect and analyze in real time the multi-dimensional impact of negative tags in user-written content on subsequent content, resulting in reduced content quality and poor user experience.
Using deep learning technology, a pre-trained deep learning model is used to identify local areas of negative markers, and combined with semantic analysis, the most suitable output timing is planned to provide users with multi-dimensional impact reminders.
It realizes automatic detection of negative tags and multi-dimensional impact analysis, improves content quality and user experience, and enhances the accuracy and intelligence level of OCR.
Smart Images

Figure CN120808362A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of deep learning, in particular to an OCR recognition method and system based on deep learning. BACKGROUND
[0002] Currently, when writing technical documents, legal files or academic papers, etc., users often face the challenge of maintaining logical consistency and accurately expressing semantics. Negative markers in the text, as possible key points of logical transitions, have an important influence on the semantic direction and structural layout of subsequent content. However, users often have difficulty in real-time perception of the multidimensional influence of these markers during the writing process, especially when the content is complex or lengthy, this problem is more pronounced. In the prior art, although some writing assistance tools based on natural language processing (NLP) can provide grammar checking or style optimization suggestions, these tools are mostly limited to surface-level improvements, lack the ability to analyze deep semantics and logical relationships, and cannot provide real-time feedback on the potential influence of negative markers.
[0003] In addition, the combination of existing OCR technology and writing assistance functions is still insufficient, and the recognition capabilities of OCR are not fully utilized to improve the intelligent level of content creation. Users often need to manually check and correct logical errors, which is not only inefficient, but also may cause content quality to decline due to negligence.
[0004] Therefore, there is an urgent need for a technical solution that integrates OCR recognition and semantic analysis functions through deep learning technology, real-time detects negative markers in user-written content, analyzes their multidimensional influence on subsequent content, and provides feedback to users in an intelligent manner, thereby improving content quality and user experience. SUMMARY
[0005] One of the purposes of the present application is to provide an OCR recognition method and system based on deep learning to solve the problems pointed out in the background art.
[0006] In the first aspect, the present application provides an OCR recognition method based on deep learning, comprising: When the user writes content, a pre-trained deep learning model is used to recognize the local area with negative markers in the written content; analyze the multidimensional influence of the local area on the user's subsequent content writing; plan the output timing point that is most conducive to the user's understanding of the multidimensional influence; At the output timing point, output multidimensional influence reminder information to the user.
[0007] Optionally, the pre-training step of the deep learning model comprises: Collecting an original sample set of text region images containing negation marks; Based on semantic adversarial generation technology, gradient negation mark enhancement is performed on the original sample set to synthesize a negation mark sample set with visual gradient features; The main task branch of the deep learning network is trained using the negation mark sample set to output the position region coordinates and type probability distribution of the negation mark; The auxiliary task branch of the deep learning network is trained using the convolutional features extracted by the main task branch to predict the conflict intensity value between the position region and the context semantics; Based on the position region coordinates, the visual blur index of the corresponding region is calculated; Based on the preset blur threshold, the negation mark sample set is divided into a clear mark group and a fuzzy mark group according to the visual blur index; The main task branch is iteratively retrained in stages in the order from the clear mark group to the fuzzy mark group; The deep learning network that has completed the iterative retraining is used as the final deep learning model.
[0008] Optionally, the negation mark at least includes: text strikethrough, negation symbol, content highlight annotation of negation color system, and user preset negation keyword.
[0009] Optionally, the analysis of the multidimensional impact of the local region on the user's subsequent content writing includes: Analyzing a plurality of sets of abnormal features before and after the negation of the local region in the full text of the content; For each set of abnormal features, a set of abnormal features is matched with its preset representative local region impact possibility on the user's subsequent content writing and its possible verification rule, and based on the possible verification rule, the impact possibility is verified according to the user's subsequent content, and when the verification is passed, the impact possibility is used as a single-dimensional impact contributed by the set of abnormal features; The single-dimensional impacts contributed by different sets of abnormal features are aggregated to obtain the multidimensional impact of the local region on the user's subsequent content writing.
[0010] Optionally, the output timing point that is most conducive to the user's understanding of the multidimensional impact includes: Based on the subsequent content written by the user and the operation behavior in the subsequent content writing process, the user's future content writing thought direction is continuously and real-time predicted, and a thought direction chain is formed; wherein different content writing thought directions are connected in time sequence to form a thought direction chain; When the steady-state part and the mutation part in the thought direction chain meet the conditions for determining the output timing, the time when the new steady-state part is about to be formed and the new mutation part that has been formed in the thought direction chain arrive first will be used as the output timing point that is most conducive to the user's understanding of the multi-dimensional impact; among them, the steady-state part is a continuous number of content writing thought directions in the thought direction chain whose two-way consistency exceeds the first consistency threshold; the mutation part is a number of content writing thought directions in the thought direction chain whose two-way consistency does not exceed the second consistency threshold and is less than a preset first number.
[0011] Optionally, the output timing determination condition includes: All mutation parts combined represent that the degree to which the user is close to understanding the multi-dimensional impact exceeds the preset degree threshold; and, The correlation between each of more than a preset second number of steady-state portions and the multi-dimensional influence exceeds a preset correlation threshold.
[0012] Optionally, the outputting multi-dimensional impact reminder information to the user includes: Generate templates for preset reminder messages that impact multi-dimensional matching; Generate multi-dimensional impact reminder information based on the reminder information generation template and multi-dimensional impact; Display multi-dimensional impact reminder information to users.
[0013] Optional deep learning-based OCR recognition methods also include: Improve the deep learning model training at preset time intervals.
[0014] Optional deep learning-based OCR recognition methods also include: Allows users to define negative tags.
[0015] In a second aspect, an embodiment of the present invention provides an OCR recognition system based on deep learning, comprising: The deep learning OCR module is used to use a pre-trained deep learning model to perform OCR to identify local areas with negative markers in the content when the user writes the content; Impact analysis module, used to analyze the multi-dimensional impact of local areas on users' subsequent content writing; Timing planning module, used to plan the output timing that is most conducive to users understanding the multi-dimensional impact; The reminder output module is used to output multi-dimensional impact reminder information to the user at the output timing point.
[0016] The present invention has achieved the following beneficial effects: Through deep learning technology, especially in the field of OCR recognition, the automatic detection and multi-dimensional influence analysis of negative markers in user-written content are realized. The pre-trained deep learning model is used to identify the markers in real time, analyze the influence of the negative markers in the local area, and provide feedback to the user through an intelligent reminder mechanism. Overall, the combination of visual recognition and semantic understanding not only improves the accuracy of OCR, but also significantly improves user experience and content quality.
[0017] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and achieved by the structure particularly pointed out in the written description and the accompanying drawings.
[0018] The technical solutions of the present application will be further described in detail below with the help of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation of the present application. In the drawings: Figure 1 A schematic diagram of an OCR recognition method based on deep learning in an embodiment of the present application; Figure 2 A schematic diagram of an OCR recognition system based on deep learning in an embodiment of the present application. DETAILED DESCRIPTION
[0020] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not limit the present application.
[0021] The research and development idea of the present application is to develop an OCR method capable of efficiently recognizing local areas with negative markers in user-written content by combining the advantages of deep learning technology, computer vision and natural language processing. These negative markers include but are not limited to text strikethrough, negative symbols, negative color highlighting annotations and user-predefined negative keywords, etc. The core goal of the research and development is to use a pre-trained deep learning model to automatically detect the visual and semantic features of these markers, and analyze their influence on subsequent content writing in combination with context analysis, and finally improve the writing quality and efficiency of users through an intelligent reminder mechanism.
[0022] Traditional OCR techniques mainly focus on recognizing text content, and are difficult to effectively handle complex scenarios with negative markers. For example, in handwritten or electronic documents, users may indicate discarded content through strikeout or mark problems through highlight annotations, but the diversity and ambiguity of these markers pose challenges to recognition.
[0023] Therefore, the present method selects deep learning as the technical basis because of its strong feature extraction capability and modeling ability for non-linear patterns. By constructing a diversified training sample set and designing a multi-task network structure, the method can simultaneously capture the visual features and semantic meanings of negative markers, and then provide real-time multi-dimensional impact reminders for users. This method is particularly suitable for scenarios that require high accuracy and user experience, such as document editing and review in government systems.
[0024] Figure 1 A flowchart of an OCR recognition method based on deep learning is provided for the embodiments of the present application, as shown in Figure 1 The method comprises: 101. When the user is writing content, use a pre-trained deep learning model to OCR recognize the local area with negative markers in the written content.
[0025] This step is the basis of the entire method, which uses a pre-trained deep learning model to perform real-time OCR recognition on the text image written by the user. The core function of the model is to detect and locate the local area with negative markers, which can appear in various forms, including at least: text strikeout (such as a line of text indicating deletion or discard), negative symbol (such as "X" or "-", indicating negation or cancellation), content highlight annotation in negative color system (such as red or gray highlight, marking problems or content that needs to be modified), and user-defined negative keywords (such as "error" and "not applicable", which are defined by the user to adapt to specific needs).
[0026] Provide a specific implementation example: when the user inputs content, the system captures the text image in real time, and the deep learning model starts to intervene. The model first divides the image into multiple regions, generates candidate boxes using the region proposal network (RPN), then classifies and locates each candidate box, and finally outputs the region with negative markers and its type. Specifically, in the context of document editing in a government system, the user is drafting a policy statement document. In the second paragraph, the user has crossed out the sentence "The existing system will continue to run" with a strikeout, and has written "Discarded" in red highlight annotation next to it. The system detects the text area covered by the strikeout and the highlight annotation area through OCR recognition, and labels their types as "strikeout" and "negative color system annotation", respectively.
[0027] The pre-training step of the deep learning model includes: 201. Collect an original sample set of text region images containing negative markers.
[0028] In this step, the training of the model relies on a high-quality data set. This step aims to collect diverse text image samples, ensuring that the samples cover various types of negative markers and application scenarios. Sample sources include handwritten notes, printed documents, electronic documents, etc., and marker forms include strikeout, symbol, highlight annotation, and keywords.
[0029] Provide a specific implementation example: technicians collect data from actual applications, such as scanning users' handwritten manuscripts, extracting screenshots of electronic documents, etc. Each image needs to be manually labeled with the location and type of negative marker to form an original sample set. Specifically, in the data sorting of a government system, 1000 document images are collected, of which 500 contain strikeout, 300 contain negative symbols, and 200 contain highlight annotations. The samples include different handwriting, fonts, and background noise to improve the robustness of the model.
[0030] 202. Based on semantic adversarial generation technology, gradient negative marker enhancement is performed on the original sample set to synthesize a negative marker sample set with visual gradient features.
[0031] In this step, the number of original samples is limited and it is difficult to cover all possible marker styles. To enhance the generalization ability of the model, semantic adversarial generation technology (Semantic Adversarial Generation) is used to generate synthetic samples with visual gradient features by adding small perturbations to the original images. This method can simulate situations such as marker blur, tilt, or color change.
[0032] Provide a specific implementation example: use a generative adversarial network (GAN) to superimpose gradient perturbations on the original image. For example, for an image containing a strikeout, the generator will add noise or change the thickness around the strikeout while maintaining semantic consistency. The discriminator then judges whether the synthetic sample is realistic, and finally generates a diverse set of enhanced samples. For example, for an image containing a "wrong" red highlight, the enhancement process generates three variants: the highlight color changes to dark red, the highlight area edge is blurred, and the highlight area is tilted by 5 degrees. These synthetic samples are added to the training set to form a negative marker sample set.
[0033] 203. Use the negative marker sample set to train the main task branch of the deep learning network, output the location region coordinates and type probability distribution of the negative marker.
[0034] In this step, the main task branch is responsible for the detection and classification of negative markers, based on existing target detection frameworks such as Faster R-CNN. Its network includes convolutional layers to extract image features, RPN to generate candidate regions, and a classifier to predict marker types. During training, cross-entropy loss function is used to optimize classification accuracy, and IoU loss is used to optimize positioning accuracy. The input is the enhanced sample image, and the output is the position coordinates and type probability distribution of each marker (e.g. deletion line probability 0.9, highlight annotation probability 0.1).
[0035] A specific implementation example is provided: in a government document image, the deletion line region is detected, and the output coordinates are [100, 200, 300, 220], and the type probability distribution is [deletion line: 0.92, negative symbol: 0.04, highlight annotation: 0.03, keyword: 0.01].
[0036] 204、Synchronize the convolutional features extracted by the main task branch to train the auxiliary task branch of the deep learning network to predict the conflict intensity value of the position region and the context semantics.
[0037] In this step, simply detecting the position and type of markers is not enough to understand their impact. The auxiliary task branch uses the convolutional features extracted by the main task to analyze the semantic conflict intensity between the marker region and the context. For example, whether the sentence crossed out by the deletion line will cause the subsequent logic to be inconsistent. The auxiliary branch is based on fully connected layers and RNN, with the feature map and surrounding text sequence of the main task as input, and the conflict intensity value (a floating point number between 0 and 1) as output. During training, the mean square error loss is used, and the conflict degree annotated by humans is used as a reference.
[0038] A specific implementation example is provided: in a government document, the deletion line crosses out "the existing system will continue to run", and the auxiliary branch analyzes and outputs a conflict intensity of 0.8, indicating that the content referring to "the existing system" may be invalid due to the deletion.
[0039] 205、Based on the position region coordinates, calculate the visual blur index of the corresponding region.
[0040] In this step, the clarity of negative markers affects the difficulty of recognition. This step calculates the visual blur index of the marker region through image processing techniques, which is used for subsequent sample classification. Blur is affected by factors such as edge sharpness and contrast. Specifically, the existing Sobel operator is used to calculate the edge gradient, and the region contrast is combined to calculate the weighted average to generate a blur score (0 to 1, 0 for clear, 1 for blurred). For example: the blur calculation result of a deletion line region is 0.6, because the handwriting is light and has noise.
[0041] 206、Based on the preset blur threshold, divide the negative marker sample set into clear marker group and blurred marker group according to the visual blur index.
[0042] In this step, fuzzy mark recognition is more difficult and requires staged training. Set a fuzziness threshold (e.g., 0.5) and divide the samples into a clear group (fuzziness < 0.5) and a fuzzy group (fuzziness ≥ 0.5).
[0043] 207. The main task branch is iteratively retrained in stages from the clear label group to the fuzzy label group.
[0044] In this step, training is performed in stages to improve the model's adaptability to fuzzy samples. First, the clear group is used to train basic capabilities, and then the fuzzy group is used to optimize robustness. For example, in the first stage, the clear group is trained for 10 epochs. In the second stage, the fuzzy group is added for another 5 epochs, gradually adjusting the model parameters.
[0045] 208. The deep learning network that has completed iterative retraining is used as the final deep learning model.
[0046] In this step, after multiple stages of training, the model achieves high accuracy and robustness, ready for practical applications. The final model parameters are saved and deployed to the OCR system.
[0047] 102. Analyze the multi-dimensional impact of local areas on users’ subsequent content writing.
[0048] In this step, after identifying the local area of negative markers, it is necessary to evaluate its impact on subsequent content, including logical coherence, grammatical completeness, semantic consistency, etc. For example, deleting a paragraph may cause subsequent citations to become invalid.
[0049] 103. Plan the output timing that is most conducive to users understanding the multi-dimensional impact.
[0050] In this step, if the multi-dimensional impact reminder information is directly output to the user, the user's current thinking state may not be able to clearly understand the multi-dimensional impact for a while, and it may also interrupt the user's thinking. Therefore, it is necessary to plan the output timing that is most conducive to the user's understanding of the multi-dimensional impact.
[0051] 104. At the output timing point, output multi-dimensional impact reminder information to the user.
[0052] In this step, multi-dimensional impact reminders alert users to the multi-dimensional impacts that may occur in a local area. For example, in a government document, the system might prompt the user: "You have deleted the description of the existing system in paragraph 2. This may affect the logic in paragraph 3. We recommend checking."
[0053] A specific implementation example of such an overall method is provided: in a government system migration scenario, a user drafts a device update plan. The first paragraph of the document reads: "Old devices will be decommissioned next year," but the user crosses out "decommissioned" with a strikeout and adds a red annotation "continue to use." The system identifies the strikeout (coordinates: [50, 80, 200, 100]) and the annotation (coordinates: [210, 80, 300, 100]) through step 101. Step 102 analysis finds that paragraph 2 mentions "new devices replace old devices," which conflicts with "continue to use" after the strikeout, with a conflict intensity of 0.75. Step 103 triggers a reminder after the user completes paragraph 1, and step 104 prompts: "You have deleted 'decommissioned,' which may conflict with the replacement plan in paragraph 2. Please adjust." The user modifies paragraph 2 to "new devices supplement old devices" based on the prompt, ensuring document consistency.
[0054] In summary, steps 101 to 104 achieve automatic detection and multi-dimensional impact analysis of negation markers in user-written content through deep learning technology, especially in the field of OCR recognition. Step 101 uses a pre-trained deep learning model to identify markers in real time, step 102 analyzes the impact of negation markers in local areas, and steps 103 and 104 provide feedback to users through intelligent reminder mechanisms. Overall, the combination of visual recognition and semantic understanding not only improves the accuracy of OCR, but also significantly improves user experience and content quality.
[0055] Steps 201 to 208 build high-performance deep learning models through systematic data collection, enhancement and training strategies. Steps 201 and 202 ensure data diversity, steps 203 and 204 achieve dual optimization of vision and semantics, steps 205 to 207 improve robustness through phased training, and step 208 outputs the final model. Overall, effectively deal with the complexity of negation markers, ensure the stability and accuracy of the model in different scenarios.
[0056] In some embodiments, the 102, analyzing the multi-dimensional impact of the local area on the user's subsequent content writing, includes: 301, analyze a plurality of sets of abnormal feature sets before and after the negation in the local area in the full text of the content.
[0057] In this step, the system conducts a thorough analysis of the full text of the content, focusing on the state of the local region before negation and the state of the local region after negation. By comparing the features of the content in these two states, the system can identify a number of sets of abnormal features. These sets of abnormal features reflect the specific changes or effects that the negation of the local region has on the full text content. Abnormal features refer to significant changes in the content that occur as a result of the negation of the local region. These changes can be in the form of abnormal patterns in sentiment, theme, tone, vocabulary usage, etc. Based on existing semantic analysis techniques, the text features (such as word frequency, sentence structure, sentiment, etc.) before and after the negation of the local region can be compared to extract abnormal feature combinations.
[0058] A specific implementation example is provided as follows: Suppose the user is writing an article about the use of a mobile phone, and the content before the negation of the local region is "The performance of this mobile phone is very good, especially the processor speed is very fast." After the negation of the local region, the content becomes "The performance of this mobile phone is good." (The word "very" and the phrase "especially the processor speed is fast" are deleted), and the analysis results in a set of abnormal features: the sentiment shifts from positive ("very good" and "especially the processor speed is fast") to relatively normal ("good" performance only).
[0059] 302、For each set of abnormal features, match the set of abnormal features with its pre-set possible influence of the local region on the user's subsequent content writing and its possible verification rule, based on the possible verification rule, verify the possible influence according to the subsequent content written by the user, and when the verification is passed, take the possible influence as the single-dimensional influence contributed by the set of abnormal features.
[0060] In this step, the system matches each set of abnormal features with pre-set possible influences, which represent the potential effects of the local region on the user's subsequent content writing. Each possible influence also corresponds to a set of verification rules for determining whether the influence is valid. The system uses the verification rules to test each possible influence based on the subsequent content written by the user. If the subsequent content matches the expected patterns or features of the verification rules, the verification is passed. When the verification is passed, the possible influence is considered as the single-dimensional influence contributed by the set of abnormal features.
[0061] Continuing with the example of the mobile phone article, a specific implementation example is provided as follows: Suppose the subsequent content is "I am very satisfied with the processor of this mobile phone, for example, when processing files, the mobile phone runs very fast", and the possible influence is that the subsequent content cannot echo the excellent performance of the mobile phone, especially the processor, as described in the previous text. The verification rule is to verify whether the user specifically describes the excellent performance of the mobile phone, especially the processor, and if so, the verification is passed. When the subsequent content is verified using this verification rule, it is found that the verification is passed, and the possible influence is generated as a single-dimensional influence.
[0062] 303. Aggregate the single-dimension impacts contributed by different sets of abnormal features, obtaining the multi-dimension impact of the local region on the user's subsequent content writing.
[0063] In this step, the single-dimension impacts contributed by different sets of abnormal features are integrated. By synthesizing these single-dimension impacts, the system obtains the multi-dimension impact of the local region on the user's subsequent content writing. This multi-dimension impact provides a more comprehensive perspective, helping to understand how the negation of the local region affects the user's writing behavior and content development direction on multiple levels.
[0064] In summary, through steps 301 to 303, the system can systematically analyze and quantify the impact of the negation of the local region on the user's subsequent content writing. Step 301 identifies sets of abnormal features, step 302 confirms single-dimension impacts through verification rules, and step 303 aggregates single-dimension impacts to obtain multi-dimension impacts. Overall, the specific impact of the negation of the local region is revealed, and reliable data support is provided for subsequent output of multi-dimension impact reminder information to the user.
[0065] In some embodiments, the 103 plans the output timing point that is most conducive to the user's understanding of the multi-dimension impact, including: 401. Continuously and real-time predict the user's future content writing thought direction based on the user's subsequent content writing and operation behavior during the subsequent content writing process, and record to form a thought direction chain. Among them, different content writing thought directions are connected in time sequence to form a thought direction chain.
[0066] In this step, the system continuously and real-time predicts the user's future content writing thought direction based on the user's subsequent content writing and operation behavior (such as input speed, modification frequency, pause time, etc.) during the subsequent content writing process. These predicted directions are connected in chronological order to form a thought direction chain. The thought direction chain is essentially a dynamic record of the user's writing thought path, reflecting the user's focus and intention changes during the writing process. The content writing thought direction refers to the user's writing theme, intention or logical tendency at a certain moment or in a certain paragraph, for example: the user may focus on describing the problem in a paragraph, and turn to proposing a solution in another paragraph. The content writing thought direction represented by different written subsequent content and operation behavior can be analyzed in advance, and the mapping relationship between different written subsequent content and operation behavior and content writing thought direction is established. When predicting, directly query the mapping relationship to determine the content writing thought direction.
[0067] 402、When the steady-state part and the mutation part in the train of thought direction chain meet the output opportunity determination condition, the first-to-time of the new steady-state part to be formed and the new mutation part to be formed in the train of thought direction chain is taken as the output opportunity point most conducive to the user's understanding of the multidimensional influence. The steady-state part is a continuous plurality of content writing train of thought directions in the train of thought direction chain, the consistency degree of which between any two directions exceeds the first consistency degree threshold. The mutation part is less than a preset first number of content writing train of thought directions in the train of thought direction chain, the consistency degree of which between any two directions does not exceed the second consistency degree threshold.
[0068] In this step, the consistency degree refers to the consistency degree between two content writing train of thought directions. The first consistency degree threshold is a threshold representing a relatively large consistency degree. The second consistency degree threshold is a threshold representing a relatively small consistency degree. The preset first number can be 3. The steady-state part as a whole reflects the continuous and less changeable writing train of thought of the user. On the contrary, the mutation part as a whole reflects the continuous and suddenly changeable writing train of thought of the user. When the steady-state part and the mutation part in the train of thought direction chain meet the output opportunity determination condition, it means that the output opportunity can be determined at this time. The first-to-time of the new steady-state part to be formed (a new continuous and less changeable writing train of thought situation will be generated) and the new mutation part to be formed (a continuous and suddenly changeable writing train of thought situation has been generated) in the train of thought direction chain is taken as the output opportunity point most conducive to the user's understanding of the multidimensional influence. This avoids the generation of a new continuous and less changeable writing train of thought situation, which leads to a lower degree of acceptance of thinking, or the output immediately after the generation of a continuous and suddenly changeable writing train of thought situation, which leads to a higher degree of acceptance of thinking.
[0069] The output opportunity determination condition includes: Condition one, all mutation parts jointly represent that the degree of the user's approaching understanding of the multidimensional influence exceeds a preset degree threshold.
[0070] And, Condition two, the correlation degree between more than a preset second number of steady-state parts and the multidimensional influence respectively exceeds a preset correlation degree threshold.
[0071] In the first condition, the degree of the user approaching to understand the multidimensional influence represented by the mutation part combination refers to the degree of the user approaching to understand the multidimensional influence when the user generates the content writing thought direction in the mutation part. Taking the example of the mobile phone article above, for example, the multidimensional influence is that the performance of the mobile phone, especially the processor, is excellent, which cannot be echoed with the previous description, and the content writing thought direction in the mutation part is to re-describe the performance of the mobile phone, especially the processor, which is excellent. Then it represents that the user may realize that the negation of the deleted part in the previous description should be revoked. Different degrees of the user approaching to understand the multidimensional influence represented by the mutation part combination can be pre-set. The degree threshold refers to a threshold representing a greater degree of the user approaching to understand the multidimensional influence.
[0072] In the second condition, the second number can be 2. The correlation between each of the steady-state parts and the multidimensional influence refers to the correlation between the content writing thought direction in the steady-state part and the multidimensional influence. Different correlations between each of the steady-state parts and the multidimensional influence can be pre-set. The correlation threshold refers to a threshold representing a greater correlation.
[0073] At this time, when the steady-state part and the mutation part in the thought direction chain meet the output opportunity determination condition, it means that the user's thinking state is sufficient to enter the multidimensional influence related to the multidimensional influence and is close enough to understand the multidimensional influence. At this time, it is most suitable to determine the output opportunity point. Avoid outputting in the steady-state part, at this time, the user may focus on the current thought, and the willingness to accept new information is lower. Or use the end point of the mutation part, at this time, the user's thinking is active and adaptable, and it is easier to accept the reminding information of the multidimensional influence.
[0074] In summary, steps 401 to 402 can accurately determine the most suitable output opportunity point in the user's writing process by predicting the user's content writing thought direction in real time and forming a thought direction chain, thereby significantly improving the user's understanding efficiency of complex multidimensional influence and avoiding interference with the user's thought. Specifically, the thought path is dynamically recorded based on the user's subsequent content and operation behavior, and the steady-state part and the mutation part are identified in the thought direction chain. When both meet the output opportunity condition, for example, the user is close to understanding the multidimensional influence and the thinking activity is high, the output point is selected as the first time to form a new steady-state and the mutation part has been formed.
[0075] In some embodiments, the 104, outputting multidimensional influence reminding information to the user, comprises: 501, match a pre-set reminding information generation template for the multidimensional influence.
[0076] 502, based on the reminding information generation template, generate multidimensional influence reminding information according to the multidimensional influence.
[0077] 503, display the multidimensional influence reminding information to the user.
[0078] The reminding information generation template matched with different multi-dimension influences can be preset in advance, and the multi-dimension influence reminding information is generated based on the template, and finally displayed to the user, so as to realize outputting the multi-dimension influence reminding information to the user.
[0079] In some embodiments, the deep learning-based OCR recognition method further comprises: The deep learning model is improved and trained at every preset time interval.
[0080] The deep learning model can also be improved and trained at a regular time, so as to continuously improve its working capacity.
[0081] In some embodiments, the deep learning-based OCR recognition method further comprises: The user can customize the negative mark.
[0082] The user can customize the negative mark according to his / her preference and individuality, so as to improve the user experience.
[0083] Figure 2 A schematic diagram of an OCR recognition system based on deep learning is provided for the embodiments of the present application, as shown in the figure, the system comprises: Figure 2 The deep learning OCR module 100 is used for, when the user writes content, using a pre-trained deep learning model to OCR recognize a local area with a negative mark in the content; The influence analysis module 200 is used for analyzing the multi-dimension influence of the local area on the user's subsequent content writing; The timing point planning module 300 is used for planning an output timing point most beneficial for the user to understand the multi-dimension influence; The reminding output module 400 is used for outputting the multi-dimension influence reminding information to the user at the output timing point.
[0084] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and the equivalent technologies thereof, the present application also intends to include these modifications and variations.
Claims
1. An OCR recognition method based on deep learning, characterized in that: include: When users write content, OCR uses a pre-trained deep learning model to identify local areas with negative markers in the written content; Analyze the multi-dimensional impact of local areas on users' subsequent content writing; Plan the output timing that is most conducive to users understanding the multi-dimensional impact; At the output timing point, multi-dimensional impact reminder information is output to the user.
2. The deep learning-based OCR recognition method according to claim 1, wherein: The pre-training steps of the deep learning model include: Collecting an original sample set of text region images containing negative labels; Based on semantic adversarial generation technology, the original sample set is enhanced with gradient negative labels to synthesize a negative label sample set with visual gradient features; The main task branch of the deep learning network is trained using the negatively labeled sample set to output the location area coordinates and type probability distribution of the negatively labeled samples. The convolutional features extracted by the main task branch are used to train the auxiliary task branch of the deep learning network to predict the conflict intensity value between the location area and the context semantics. Based on the coordinates of the location area, calculate the visual fuzziness index of the corresponding area; Based on the preset fuzziness threshold and the visual fuzziness index, the negatively labeled sample set is divided into a clear label group and a fuzzy label group; Iteratively retrain the main task branch in stages, starting from the clear label group to the fuzzy label group. The deep learning network that has completed iterative retraining is used as the final deep learning model.
3. The deep learning-based OCR recognition method according to claim 1, wherein: The negative mark includes at least: text strikethrough, negative symbol, content highlight annotation in negative color and negative keywords preset by the user.
4. The deep learning-based OCR recognition method according to claim 1, wherein: The analysis of the multi-dimensional impact of a local area on the user's subsequent content writing includes: Analyze multiple sets of abnormal features of the full text before and after the local area is denied; For each set of unusual feature sets, the influence of the preset representative local area on the user's subsequent content writing and its possible verification rules are matched for the unusual feature set. Based on the possible verification rules, the influence possibility is verified according to the subsequent content written by the user. When the verification passes, the influence possibility is used as the single-dimensional influence contributed by the unusual feature set; Aggregate the single-dimensional influences contributed by different groups of unusual feature sets to obtain the multi-dimensional influence of local areas on users' subsequent content writing.
5. The deep learning-based OCR recognition method according to claim 1, wherein: The above plan is most helpful for users to understand the output timing of multi-dimensional impacts, including: Based on the subsequent content written by the user and the operational behavior during the subsequent content writing process, the user's future content writing ideas are continuously predicted in real time and recorded to form a chain of ideas. Among them, different content writing ideas are connected in time sequence to form a chain of ideas. When the steady-state part and the mutation part in the thought direction chain meet the conditions for determining the output timing, the time when the new steady-state part is about to be formed and the new mutation part that has been formed in the thought direction chain arrive first will be used as the output timing point that is most conducive to the user's understanding of the multi-dimensional impact; among them, the steady-state part is a continuous number of content writing thought directions in the thought direction chain whose two-way consistency exceeds the first consistency threshold; the mutation part is a number of content writing thought directions in the thought direction chain whose two-way consistency does not exceed the second consistency threshold and is less than a preset first number.
6. The deep learning-based OCR recognition method according to claim 5, wherein: The output timing determination conditions include: All mutation parts combined represent that the user is close to understanding the multi-dimensional impact beyond the preset degree threshold; and, The correlation between each of more than a preset second number of steady-state portions and the multi-dimensional influence exceeds a preset correlation threshold.
7. The deep learning-based OCR recognition method according to claim 1, wherein: Outputting multi-dimensional impact reminder information to the user includes: Generate templates for preset reminder messages that impact multi-dimensional matching; Generate multi-dimensional impact reminder information based on the reminder information generation template and multi-dimensional impact; Display multi-dimensional impact reminder information to users.
8. The deep learning-based OCR recognition method according to claim 1, wherein: Also includes: Improve the deep learning model training at preset time intervals.
9. The deep learning-based OCR recognition method according to claim 1, wherein: Also includes: Allows users to define negative tags.
10. An OCR recognition system based on deep learning, characterized in that: include: The deep learning OCR module is used to use a pre-trained deep learning model to perform OCR to identify local areas with negative markers in the content when the user writes the content; Impact analysis module, used to analyze the multi-dimensional impact of local areas on users' subsequent content writing; Timing planning module, used to plan the output timing that is most conducive to users understanding the multi-dimensional impact; The reminder output module is used to output multi-dimensional impact reminder information to the user at the output timing point.