Floating window real-time OCR recognition effect evaluation and dynamic optimization method, device and equipment

CN122780932APending Publication Date: 2026-09-18CHENGDU BOOK SOUND TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610838924.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明提供了一种悬浮窗实时0CR识别效果评测与动态优化方法、装置及设备,用以解决现有悬浮窗OCR识别技术无法根据识别效果评测结果对OCR识别流程进行动态优化,导致识别准确率与识别速度难以符合用户需求的问题

Benefits of technology

本发明提供的悬浮窗实时0CR识别效果评测与动态优化方法、装置及设备,首先获取悬浮窗显示区域对应的实时图像数据,并利用初始OCR识别模型对实时图像数据进行文字检测与字符识别,得到实时OCR识别结果;随后,通过识别准确率评测和识别速度评测对当前识别结果进行多维度效果分析,从而判断当前OCR识别结果是否满足预设OCR识别要求。当识别结果不满足要求时,不再沿用固定OCR识别流程,而是依据评测结果在预设优化策略集合中动态获取对应的目标优化策略,并针对当前场景对实时图像数据和/或初始OCR识别模型进行优化处理,例如对文字检测参数、字符识别参数以及模型结构参数进行调整,从而生成优化后的目标图像数据或目标OCR识别模型;之后再将优化后的目标图像数据或目标OCR识别模型重新投入OCR识别流程中进行再次识别与评测,直至识别结果满足对应的识别要求。通过上述方式,使OCR识别能够根据不同悬浮窗场景下的识别状态动态调整识别策略,在保证识别准确率的同时兼顾识别速度,从而能够根据悬浮窗OCR识别效果动态调整对应优化策略,从而实现识别准确率与识别速度的协同优化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780932A_ABST
    Figure CN122780932A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, solves the problem that the existing floating window OCR recognition technology cannot dynamically optimize the OCR recognition process according to the recognition effect evaluation result, leading to the difficulty of the recognition accuracy and speed to meet the user's demand, and provides a floating window real-time OCR recognition effect evaluation and dynamic optimization method, device and equipment. The method comprises: obtaining real-time image data of a floating window display area, inputting an initial OCR recognition model, evaluating the output real-time OCR recognition result, and according to the evaluation result, when it is judged that the real-time OCR recognition result does not meet the OCR recognition requirement, obtaining the target image data and / or the target OCR recognition model after optimization processing according to the target optimization strategy, until it is judged that the real-time OCR recognition result meets the OCR recognition requirement, and the real-time OCR recognition result is output. The present application can dynamically adjust the corresponding optimization strategy according to the floating window OCR recognition effect, thereby realizing the collaborative optimization of the recognition accuracy and speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, and equipment for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window. Background Technology

[0002] With the continuous development of mobile terminals, smart wearable devices, and desktop interaction systems, floating window technology has been widely applied in scenarios such as game assistance, instant messaging, video subtitle extraction, online translation, remote work, and multi-task collaborative processing. In these application scenarios, users typically need to extract and recognize text information in the floating window display area in real time to achieve functions such as chat content parsing, subtitle translation, task prompt acquisition, and rapid information input. Therefore, real-time OCR recognition technology for floating windows has gradually become an important technical means to improve human-computer interaction efficiency and user experience. Especially under dynamic interfaces, multi-window overlays, and complex background conditions, how to achieve fast, stable, and accurate recognition of text content in the floating window area has become an important research direction in the field of smart terminal text interaction.

[0003] Most existing floating window OCR recognition technologies employ a fixed recognition process, which involves uniformly recognizing and processing the floating window image using a pre-set OCR recognition model and directly outputting the recognition result. While some existing technologies can achieve basic text detection and character recognition, they typically lack mechanisms for evaluating recognition performance and adaptive optimization for dynamic floating window scenarios. When the floating window scenario involves interface occlusion, font size changes, background interference, dynamic refresh, or image compression distortion, issues such as missed text recognition, incorrect recognition, and decreased recognition speed can easily occur. Furthermore, existing technologies generally cannot dynamically adjust the OCR recognition model or image processing parameters based on the recognition accuracy and speed evaluation results under different scenarios. This makes it difficult to maintain a balance between recognition accuracy and real-time performance, thus affecting the overall stability and practicality of floating window OCR recognition.

[0004] Therefore, how to implement a real-time performance evaluation method for OCR recognition results in floating window scenarios, and dynamically obtain corresponding optimization strategies based on the evaluation results to balance recognition accuracy and recognition speed, has become an urgent technical problem to be solved. Summary of the Invention

[0005] In view of this, the present invention provides a method, apparatus and equipment for real-time OCR recognition effect evaluation and dynamic optimization of floating window, in order to solve the problem that the existing floating window OCR recognition technology cannot dynamically optimize the OCR recognition process based on the recognition effect evaluation results, resulting in recognition accuracy and recognition speed that are difficult to meet user needs.

[0006] The technical solution adopted in this invention is: In a first aspect, the present invention provides a method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window, the method comprising: Acquire real-time image data corresponding to the display area of ​​the floating window in the floating window operation scenario; The real-time image data is input into a preset initial OCR recognition model to obtain the real-time OCR recognition result; The real-time OCR recognition result is evaluated according to the preset multi-dimensional recognition effect evaluation strategy, and the real-time OCR recognition result is judged to meet the preset OCR recognition requirements based on the evaluation result. The multi-dimensional recognition effect evaluation strategy includes recognition accuracy evaluation and recognition speed evaluation. When it is determined that the real-time OCR recognition result meets the OCR recognition requirements, the real-time OCR recognition result is output. When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the corresponding target optimization strategy is obtained from the preset optimization strategy set according to the evaluation result; According to the target optimization strategy, the real-time image data and / or the initial OCR recognition model are optimized to obtain the optimized target image data and / or target OCR recognition model. The target image data is used as the new real-time image data and / or the target OCR recognition model is used as the new initial OCR recognition model. The process of inputting the real-time image data into the preset initial OCR recognition model to obtain the real-time OCR recognition result is repeated until it is determined that the real-time OCR recognition result meets the OCR recognition requirements, and then the real-time OCR recognition result is output.

[0007] Preferably, the step of inputting the real-time image data into a preset initial OCR recognition model to obtain the real-time OCR recognition result includes: The real-time image data is input into an initial text detection model constructed based on a differentiable binarization algorithm. The initial text detection model is used to calculate the probability information and threshold information of the text regions in the real-time image data to generate the corresponding binary image of the text regions. The connected regions in the binary image of the text region are extracted, and corresponding text region detection boxes are generated based on the boundary contours of each connected region. Based on each of the text region detection boxes, the corresponding text regions in the real-time image data are cropped to obtain multiple text region images; The pixel projection distribution results of each text region image in the horizontal direction are obtained respectively, and the corresponding text line region is determined according to the pixel projection distribution results in the horizontal direction; The pixel projection distribution results of each text line region in the vertical direction are statistically analyzed, and the candidate regions of characters in each text line region are separated according to the pixel interval width between adjacent characters to obtain multiple character images; The size of each character image is normalized, and the normalized character image is input into an initial character recognition model constructed based on a convolutional recurrent neural network. The initial character recognition model is used to identify the character category corresponding to each character image to obtain the corresponding character recognition result and character recognition confidence. According to the arrangement order of each character recognition result in the corresponding text line area, the character recognition results are concatenated to generate the corresponding text recognition result; Based on the confidence scores of each character recognition, the confidence scores of the corresponding text recognition results are statistically analyzed, and the text recognition results whose statistically obtained text recognition confidence scores meet the preset confidence conditions are determined as the real-time OCR recognition results.

[0008] Preferably, the initial text detection model is trained through the following steps: Multiple sample images containing floating window display interfaces are obtained, and the text regions in each sample image are labeled with rectangular boxes to obtain a text detection training dataset. Each sample image in the text detection training dataset is scaled, and the scaled sample images are uniformly resized according to a preset size to obtain standard training images. Based on the text region annotation results in each of the standard training images, generate corresponding text region probability label maps and text region threshold label maps. Each of the standard training images is input into a text detection network constructed based on a differentiable binarization algorithm, and the text detection network outputs the corresponding text region prediction probability map and text region prediction threshold map respectively. Based on the difference between the predicted probability map of the text region and the corresponding probability label map of the text region, the corresponding probability map loss value is calculated, and based on the difference between the predicted threshold map of the text region and the corresponding threshold label map of the text region, the corresponding threshold map loss value is calculated. Based on the probability graph loss value and the threshold graph loss value, the network parameters in the text detection network are iteratively updated until the preset training termination condition is met, and the initial text detection model is obtained. The initial character recognition model is trained through the following steps: Multiple character sample images are acquired, and the character categories corresponding to each character sample image are labeled to obtain a character recognition training dataset; The character sample images in the character recognition training dataset are normalized to obtain standard character images; Each of the standard character images is input into a character recognition network constructed based on a convolutional recurrent neural network. The local features of the characters in each of the standard character images are extracted through a convolutional feature extraction layer to obtain the corresponding character feature map. Each of the character feature maps is input into the cyclic sequence recognition layer, and the character sequence features in the character feature maps are extracted to obtain the corresponding character sequence features; Based on the character sequence features, the character category corresponding to each of the standard character images is predicted to obtain the corresponding character category prediction result; Calculate the corresponding character recognition loss value based on the difference between the character category prediction result and the corresponding character category annotation result; Based on the character recognition loss value, the network parameters in the character recognition network are iteratively updated until the preset training termination condition is met, thus obtaining the initial character recognition model.

[0009] Preferably, the step of evaluating the real-time OCR recognition result according to a preset multi-dimensional recognition effect evaluation strategy, and determining whether the real-time OCR recognition result meets the preset OCR recognition requirements based on the evaluation result, includes: The real-time OCR recognition result is matched with the preset standard recognition result corresponding to the real-time image data. Based on the matching result, the number of correctly matched characters, the number of incorrectly recognized characters, and the number of missing characters are counted. Based on the number of correctly matched characters, the number of incorrectly identified characters, and the number of missed characters, combined with the total number of characters identified, the corresponding character recognition accuracy is calculated, and the character recognition accuracy is determined as the corresponding recognition accuracy evaluation result. The number of valid image frames that have completed OCR recognition processing within a preset statistical time period is obtained, and the corresponding OCR recognition speed is calculated based on the number of valid image frames and the preset statistical time period. The OCR recognition speed is then determined as the corresponding recognition speed evaluation result. The recognition accuracy evaluation result is compared with a preset accuracy threshold, and the recognition speed evaluation result is compared with a preset speed threshold. If the recognition accuracy evaluation result is greater than or equal to the accuracy threshold, and the recognition speed evaluation result is greater than or equal to the speed threshold, then the real-time OCR recognition result meets the OCR recognition requirements. If the recognition accuracy evaluation result is less than the accuracy threshold and / or the recognition speed evaluation result is less than the speed threshold, then the real-time OCR recognition result does not meet the OCR recognition requirements.

[0010] Preferably, when it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, obtaining the corresponding target optimization strategy from the preset optimization strategy set based on the evaluation result includes: When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the Mask R-CNN instance segmentation algorithm is used to detect occlusion areas in the real-time image data to determine whether there are occlusion areas in the real-time image data and obtain the judgment result. Based on the judgment result, a first scene type for the floating window operation scene is determined, wherein the first scene type includes an occluded scene or an unoccluded scene; Based on the mapping relationship between the first scene type and the preset scene type and optimization strategy, a subset of candidate optimization strategies corresponding to the first scene type is determined in the set of optimization strategies, wherein the subset of candidate optimization strategies includes a subset of optimization strategies for occluded scenes or a subset of optimization strategies for non-occluded scenes. Based on the evaluation results, the corresponding target optimization strategy is selected from the subset of candidate optimization strategies.

[0011] Preferably, selecting the corresponding target optimization strategy from the subset of candidate optimization strategies based on the evaluation results includes: If the evaluation result is that the recognition accuracy does not meet the requirements but the recognition speed does meet the requirements, then at least one of the following optimization strategies is selected from the candidate optimization strategy subset: adjusting the text region binarization threshold to a first preset threshold range, adjusting the text region detection box intersection-union ratio threshold to a first preset intersection-union ratio range, adjusting the character image to a first preset input size, and adjusting the number of hidden layer nodes of the convolutional recurrent neural network to a first preset number of nodes range, as the target optimization strategy. If the evaluation result is that the recognition accuracy meets the requirements but the recognition speed does not meet the requirements, then at least one of the following optimization strategies is selected from the candidate optimization strategy subset: adjusting the text region binarization threshold to a second preset threshold range, adjusting the text region detection box intersection-union ratio threshold to a second preset intersection-union ratio range, adjusting the character image to a second preset input size, and adjusting the number of hidden layer nodes of the convolutional recurrent neural network to a second preset node number range as the target optimization strategy. Wherein, the value corresponding to the first preset threshold interval is less than the value corresponding to the second preset threshold interval, the image resolution corresponding to the first preset input size is greater than the image resolution corresponding to the second preset input size, and the number of nodes corresponding to the first preset node number interval is greater than the number of nodes corresponding to the second preset node number interval. If the evaluation result is that the recognition accuracy and recognition speed do not meet the requirements, then a combined optimization strategy including the text region binarization threshold adjustment strategy, the text region detection box intersection-union threshold adjustment strategy, the character image input size adjustment strategy, and the convolutional recurrent neural network hidden layer node number adjustment strategy is selected from the candidate optimization strategy subset as the target optimization strategy.

[0012] Preferably, if the evaluation result is that the recognition accuracy and recognition speed do not meet the requirements, then a combined optimization strategy including a text region binarization threshold adjustment strategy, a text region detection box intersection-union threshold adjustment strategy, a character image input size adjustment strategy, and a convolutional recurrent neural network hidden layer node number adjustment strategy is selected from the candidate optimization strategy subset as the target optimization strategy, including: Obtain the strategy parameter combination corresponding to multiple candidate combination optimization strategies in the candidate optimization strategy subset, wherein each strategy parameter combination includes a text region binarization threshold parameter, a text region detection box intersection-union ratio threshold parameter, a character image input size parameter, and a convolutional recurrent neural network hidden layer node number parameter; Based on the combination of strategy parameters, the real-time image data is subjected to OCR recognition processing to obtain multiple candidate OCR recognition results. The recognition accuracy and recognition speed of each candidate OCR recognition result are evaluated to obtain the corresponding candidate recognition accuracy evaluation results and candidate recognition speed evaluation results. Based on the evaluation results of the accuracy of each candidate recognition and the evaluation results of the speed of each candidate recognition, a target candidate combination optimization strategy that simultaneously satisfies the preset accuracy condition and the preset speed condition is selected. If there are multiple target candidate combination optimization strategies, then based on the degree of deviation between the recognition accuracy evaluation result and the recognition speed evaluation result corresponding to each target candidate combination optimization strategy, the target candidate combination optimization strategy with the smallest deviation degree is selected as the target optimization strategy; If there is no target candidate combination optimization strategy that satisfies the preset accuracy condition and the preset speed condition, then the target candidate combination optimization strategy with the smallest combined deviation of the recognition accuracy evaluation result and the recognition speed evaluation result from the preset threshold is selected as the target optimization strategy.

[0013] Preferably, the OCR recognition requirement is obtained through the following steps: In response to the OCR recognition command input by the user, the OCR recognition command is parsed to obtain the second scene type of the floating window running scene, wherein the second scene type includes game scene, video playback scene, instant messaging scene, web browsing scene and office document scene; The OCR recognition requirements are determined based on the mapping relationship between the second scene type and the preset scene type and recognition requirements.

[0014] Secondly, the present invention provides a device for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window, the device comprising: The image acquisition module is used to acquire real-time image data corresponding to the display area of ​​the floating window in the floating window operation scenario; The OCR recognition module is used to input the real-time image data into a preset initial OCR recognition model to obtain the real-time OCR recognition result; The evaluation module is used to evaluate the real-time OCR recognition result according to the preset multi-dimensional recognition effect evaluation strategy, and to determine whether the real-time OCR recognition result meets the preset OCR recognition requirements based on the evaluation result. The multi-dimensional recognition effect evaluation strategy includes recognition accuracy evaluation, recognition speed evaluation, and occlusion robustness evaluation. The first output module for recognition results is used to output the real-time OCR recognition result when it is determined that the real-time OCR recognition result meets the OCR recognition requirements; The optimization strategy acquisition module is used to acquire the corresponding target optimization strategy from the preset optimization strategy set according to the evaluation result when it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements. An optimization processing module is used to optimize the real-time image data and / or the initial OCR recognition model according to the target optimization strategy, so as to obtain the optimized target image data and / or target OCR recognition model. The second output module for recognition results is used to take the target image data as the new real-time image data and / or take the target OCR recognition model as the new initial OCR recognition model, return to execute the step of inputting the real-time image data into the preset initial OCR recognition model to obtain the real-time OCR recognition result, until it is determined that the real-time OCR recognition result meets the OCR recognition requirements, and output the real-time OCR recognition result.

[0015] Thirdly, embodiments of the present invention also provide an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect described above.

[0016] In summary, the beneficial effects of the present invention are as follows: The present invention provides a method, apparatus, and device for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window. First, real-time image data corresponding to the display area of ​​the floating window is acquired. Then, an initial OCR recognition model is used to perform text detection and character recognition on the real-time image data to obtain a real-time OCR recognition result. Subsequently, the recognition result is analyzed from multiple dimensions through recognition accuracy and recognition speed evaluations to determine whether the current OCR recognition result meets the preset OCR recognition requirements. When the recognition result does not meet the requirements, the fixed OCR recognition process is no longer used. Instead, based on the evaluation results, a corresponding target optimization strategy is dynamically acquired from a preset optimization strategy set. The real-time image data and / or the initial OCR recognition model are then optimized for the current scenario, for example, by adjusting text detection parameters, character recognition parameters, and model structure parameters, thereby generating optimized target image data or a target OCR recognition model. The optimized target image data or target OCR recognition model is then re-entered into the OCR recognition process for re-recognition and evaluation until the recognition result meets the corresponding recognition requirements. By employing the above methods, OCR recognition can dynamically adjust its recognition strategy based on the recognition status under different floating window scenarios, ensuring both recognition accuracy and speed. This allows for dynamic adjustment of corresponding optimization strategies based on the floating window OCR recognition effect, thereby achieving synergistic optimization of recognition accuracy and speed. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.

[0018] Figure 1 This is a schematic diagram of the overall workflow of the floating window real-time OCR recognition effect evaluation and dynamic optimization method in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating how, in Embodiment 1 of the present invention, a target optimization strategy is obtained from a preset set of optimization strategies based on the evaluation results. Figure 3 This is a structural block diagram of the floating window real-time OCR recognition effect evaluation and dynamic optimization device in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the electronic device in Embodiment 3 of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the terms "center," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, the element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Where there is no conflict, embodiments of the present invention and the various features thereof can be combined with each other, all of which are within the scope of protection of the present invention. Example 1

[0020] Please see Figure 1 Embodiment 1 of the present invention discloses a method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window, the method comprising: Acquire real-time image data corresponding to the display area of ​​the floating window in the floating window operation scenario; Specifically, the floating window scenario typically refers to an independent display window that floats above the interface of a mobile terminal, tablet device, or desktop system. These windows are usually small in size, transparently overlaid, dynamically refreshed, and displayed in parallel with the underlying application interface. Therefore, their image content is easily affected by factors such as complex background textures, dynamic image changes, partial occlusion, and resolution compression. Real-time image data refers to the window image frame data acquired continuously at preset time intervals at the current moment. It can be represented as RGB image data, grayscale image data, or screen frame buffer data, reflecting the actual visual content of the floating window display area in its current operating state. The system window management interface obtains the window hierarchy information, boundary coordinate information, and display area size parameters of the current floating window, and establishes a corresponding area acquisition range based on the floating window's position on the screen. Subsequently, the display content in the corresponding area is periodically captured through a screen capture module, a GPU frame buffer reading module, or a system graphics interface, forming a continuous sequence of video frame images as the real-time image data.

[0021] The real-time image data is input into a preset initial OCR recognition model to obtain the real-time OCR recognition result; Specifically, after acquiring real-time image data, the data is input into the text detection network of the initial OCR recognition model. The text detection network is used to locate regions in the entire image that may contain text content. It extracts features from the texture structure, edge distribution, and character arrangement patterns in the image using a convolutional neural network and outputs the corresponding text region probability distribution results. In another implementation, the text detection network can generate a text region probability map using pixel-level semantic segmentation. By thresholding and binarizing the probability map, candidate text regions are formed, and then connected component analysis is used to extract the contours of continuous text regions. For each detected text region, a bounding box, a rotated bounding box, or a polygonal bounding box can be further calculated to accurately identify the location range of the corresponding text region. After completing the text region localization, based on the position coordinates of each text region, the corresponding local text region image is cropped from the original real-time image, and character arrangement analysis is performed on each local region. Since the text in the floating window may be arranged horizontally, vertically, or in an irregular layout, pixel projection analysis, contour spacing analysis, or connected component clustering are used to separate the character boundaries in the text region. Pixel projection analysis typically identifies the whitespace between characters by statistically analyzing the pixel distribution density in the horizontal or vertical direction of an image, thereby determining the character segmentation position. Connected region clustering, on the other hand, merges character regions belonging to the same text line by analyzing the distance relationships between character contours. After character region separation, each character image is input into a character recognition network for character category determination. The character recognition network employs a structure combining convolutional neural networks and recurrent neural networks. The convolutional layers are used to extract local texture features of characters, such as stroke structure, edge direction, and character contour features; the recurrent structure is used to learn the contextual relationships between character sequences, thereby improving the ability to distinguish between similar characters. For each character image, the recognition network outputs a corresponding character category probability distribution and determines the final character category based on the maximum probability value. It also outputs a corresponding character recognition confidence score to reflect the reliability of the current character recognition result. After obtaining the character recognition results, multiple characters are sequentially concatenated according to their arrangement order in the original text region to form a complete text recognition result. For positions with low-confidence characters, error correction can be performed by combining contextual semantic rules, word frequency statistics rules, or language models. For example, when character combinations that do not conform to normal language structure appear in the recognition results, the corresponding character category can be corrected using character context probability, thereby further improving the accuracy of text recognition.

[0022] The real-time OCR recognition result is evaluated according to the preset multi-dimensional recognition effect evaluation strategy, and the real-time OCR recognition result is judged to meet the preset OCR recognition requirements based on the evaluation result. The multi-dimensional recognition effect evaluation strategy includes recognition accuracy evaluation and recognition speed evaluation. Specifically, after OCR recognition is completed, the recognition quality and operational efficiency of the current recognition result are jointly analyzed to determine whether the current OCR recognition status meets the requirements of the floating window real-time application. The recognition accuracy evaluation measures the consistency between the OCR recognition result and the actual text content, while the recognition speed evaluation measures the system's efficiency in completing OCR recognition processing per unit time, thus avoiding the problem of focusing solely on recognition accuracy while neglecting real-time response capabilities. In practice, the real-time OCR recognition result is first matched character-level with a preset standard recognition result, and the number of correctly recognized characters, incorrectly recognized characters, and missed characters are counted. The corresponding character recognition accuracy is then calculated based on the total number of characters. For cases where text arrangement is offset, an edit distance algorithm can be used to dynamically align the character sequence to improve the reliability of accuracy statistics. Simultaneously, the start and end times of OCR recognition processing are recorded, and the number of valid image frames successfully processed within a preset time period is counted. The corresponding OCR recognition speed is calculated based on the number of valid frames and the time length to reflect the real-time processing capability of the current model. After obtaining the recognition accuracy evaluation results and recognition speed evaluation results, they are compared with preset accuracy thresholds and speed thresholds, respectively. When both recognition accuracy and recognition speed meet the corresponding threshold requirements, it indicates that the current OCR recognition model can guarantee text recognition accuracy and meet the real-time processing needs of dynamic floating window scenarios. If either indicator fails to meet the requirements, it indicates that the current recognition process has insufficient recognition accuracy or insufficient model operating efficiency, requiring further dynamic optimization. By jointly evaluating recognition accuracy and recognition speed, both OCR recognition quality and real-time operating performance can be considered simultaneously, improving the stability and adaptability of OCR recognition results in complex floating window operating scenarios, and providing a clear evaluation basis for the selection of subsequent optimization strategies.

[0023] When it is determined that the real-time OCR recognition result meets the OCR recognition requirements, the real-time OCR recognition result is output. When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the corresponding target optimization strategy is obtained from the preset optimization strategy set according to the evaluation result; Specifically, once it is determined that the current OCR recognition result meets the preset recognition requirements, it indicates that the current recognition process has reached the expected operating standards in terms of text recognition accuracy and real-time processing efficiency. At this point, the corresponding text content can be output as the final valid recognition result. The OCR recognition requirements are comprehensive performance conditions pre-set for different floating window application scenarios. These include not only the accuracy of text recognition but also the recognition response capability per unit time. For example, in instant messaging scenarios, real-time recognition is emphasized, while in office document scenarios, character recognition accuracy is emphasized. The verified text recognition result is output to the display interface, input method module, translation module, or upper-level business program. It can also include the corresponding character confidence score, recognition timestamp, and text region location information for further use by subsequent modules. Because the output result has undergone dual evaluation of recognition accuracy and recognition speed, the problem of erroneous characters directly participating in business processing can be reduced, improving the stability of the floating window's real-time interaction process.

[0024] When the current OCR recognition result fails to meet the preset OCR recognition requirements, it indicates that the current recognition process has insufficient recognition performance. The optimization strategy set here refers to a pre-established combination of various OCR parameter adjustment schemes and model adjustment schemes, used to perform corresponding optimization processing for different recognition anomalies. The so-called target optimization strategy is the optimization scheme with the highest adaptability selected by the system from multiple candidate optimization schemes based on the current evaluation results. In the specific implementation process, the source of abnormal indicators in the current evaluation results is first analyzed. For example, a decrease in recognition accuracy may be caused by blurred text edges, background occlusion, character adhesion, or insufficient image resolution, while a decrease in recognition speed may be related to an excessively large input image size, too many text detection areas, or excessive computational load of neural network inference. Subsequently, according to different anomaly types, corresponding optimization directions are matched in the optimization strategy set. For example, when the recognition accuracy is low, the binarization sensitivity of the text region can be increased, the character input resolution can be increased, or the number of character feature extraction layer nodes can be increased to enhance the ability to recognize details; when the recognition speed is low, the input image size can be reduced, the candidate text region range can be narrowed, or the number of hidden layer nodes in the network can be reduced to reduce the computational load of the model. For situations where both recognition accuracy and recognition speed are abnormal, multiple optimization strategies can be combined and invoked to synchronously adjust image processing parameters and OCR model parameters, thereby making the subsequent OCR recognition process more adaptable to the current floating window operating scenario and improving overall recognition stability and dynamic adaptive capability.

[0025] According to the target optimization strategy, the real-time image data and / or the initial OCR recognition model are optimized to obtain the optimized target image data and / or target OCR recognition model. Specifically, after determining the target optimization strategy, the input image data or OCR recognition model parameters in the current OCR recognition process are adjusted accordingly to improve the current recognition effect. The target optimization strategy is an automatically selected adaptive optimization scheme based on the aforementioned evaluation results; essentially, it is a parameter adjustment rule or model optimization rule corresponding to the current recognition anomaly type. The target image data is the optimized image input data, whose text region edge clarity, contrast, or noise level is more suitable for OCR recognition processing compared to the original real-time image data. The target OCR recognition model is the recognition model with adjusted parameters, whose network structure parameters, detection thresholds, or inference configuration are more suitable for the current floating window operation scenario. When it is determined that the recognition problem mainly stems from image quality, image optimization processing can be prioritized for the real-time image data. For example, adaptive brightness enhancement can be performed on images with insufficient brightness in the text region, background suppression and contrast enhancement can be performed on images with strong background interference, sharpening processing can be performed on images with blurred text edges, and filtering and noise reduction processing can be performed on random noise in dynamic images. For floating window scenarios with partial occlusion or transparent overlay interference, effective text regions can be extracted through region segmentation, retaining only the main text part for OCR recognition, thereby improving the distinguishability of text regions.

[0026] When the recognition problem is determined to stem primarily from mismatched OCR model parameters, model optimization is performed on the initial OCR recognition model. For example, the binarization threshold in the text detection stage can be adjusted to make text region extraction more sensitive; the cross-union threshold of detection boxes can be adjusted to reduce false screening of overlapping detection boxes; the character input size can be adjusted to improve the detail retention of small characters; or the number of hidden layer nodes in the convolutional recurrent neural network can be adjusted to change the balance between the model's feature extraction capability and inference computation. In some implementations, OCR recognition models of different complexities can be dynamically switched, for example, a high-precision model can be used on high-performance devices, while a lightweight model can be used on resource-constrained devices to ensure real-time operating efficiency.

[0027] The target image data is used as the new real-time image data and / or the target OCR recognition model is used as the new initial OCR recognition model. The process of inputting the real-time image data into the preset initial OCR recognition model to obtain the real-time OCR recognition result is repeated until it is determined that the real-time OCR recognition result meets the OCR recognition requirements, and then the real-time OCR recognition result is output.

[0028] Specifically, the target image data is image data that has undergone image enhancement, noise reduction, contrast adjustment, or text region optimization. Compared to the original real-time image data, its text features are clearer. The target OCR recognition model is an OCR recognition model with adjusted parameters, whose detection threshold, network structure parameters, or character recognition configuration are more suitable for the current floating window scenario. The optimized target image data is used as input data for the next round of OCR recognition, or the optimized target OCR recognition model is used as a new recognition model in subsequent recognition processing, thus forming a cyclical processing mechanism of "recognition—evaluation—optimization—re-recognition". After completing one optimization process, the system calls back the OCR recognition process, performs text detection, character segmentation, and character recognition processing on the updated image data or the updated OCR model, and recalculates the corresponding recognition accuracy and recognition speed evaluation results. If the current evaluation results still do not meet the preset OCR recognition requirements, the corresponding optimization strategy is selected again based on the new evaluation results, and the image data or model parameters are further adjusted. For example, if the recognition accuracy is still low after optimization, the text region enhancement intensity or character feature extraction capability will be further increased; if the recognition speed decreases significantly, the input resolution or the number of network computing nodes will be reduced to decrease the model inference time. During each iteration, the system retains the current evaluation results and corresponding optimization parameters to analyze the impact of different optimization schemes on the recognition effect, thereby avoiding ineffective repeated adjustments. When it is detected that the current OCR recognition result simultaneously meets the preset accuracy threshold and speed threshold, the cyclic optimization process stops, and the current recognition result is output as the final valid OCR recognition result. This closed-loop dynamic optimization method avoids the problem of traditional OCR systems' fixed parameters being unable to adapt to complex floating window scenarios, enabling the system to continuously and adaptively adjust the recognition strategy according to the current operating state, thereby improving OCR recognition stability, recognition accuracy, and real-time processing capabilities.

[0029] Preferably, the step of inputting the real-time image data into a preset initial OCR recognition model to obtain the real-time OCR recognition result includes: The real-time image data is input into an initial text detection model constructed based on a differentiable binarization algorithm. The initial text detection model is used to calculate the probability information and threshold information of the text regions in the real-time image data to generate the corresponding binary image of the text regions. The connected regions in the binary image of the text region are extracted, and corresponding text region detection boxes are generated based on the boundary contours of each connected region. Based on each of the text region detection boxes, the corresponding text regions in the real-time image data are cropped to obtain multiple text region images; The pixel projection distribution results of each text region image in the horizontal direction are obtained respectively, and the corresponding text line region is determined according to the pixel projection distribution results in the horizontal direction; The pixel projection distribution results of each text line region in the vertical direction are statistically analyzed, and the candidate regions of characters in each text line region are separated according to the pixel interval width between adjacent characters to obtain multiple character images; The size of each character image is normalized, and the normalized character image is input into an initial character recognition model constructed based on a convolutional recurrent neural network. The initial character recognition model is used to identify the character category corresponding to each character image to obtain the corresponding character recognition result and character recognition confidence. According to the arrangement order of each character recognition result in the corresponding text line area, the character recognition results are concatenated to generate the corresponding text recognition result; Based on the confidence scores of each character recognition, the confidence scores of the corresponding text recognition results are statistically analyzed, and the text recognition results whose statistically obtained text recognition confidence scores meet the preset confidence conditions are determined as the real-time OCR recognition results.

[0030] Specifically, a text detection model based on a differentiable binarization algorithm is first used to automatically locate text regions in real-time images. Differentiable binarization is a text detection method that dynamically learns binarization thresholds during neural network training. Compared to traditional fixed-threshold binarization, it automatically adjusts the segmentation boundary between text and background based on the text features of different regions, thus improving text detection capabilities in complex backgrounds. In practice, the system inputs real-time image data into the initial text detection model, which outputs text region probability information and text region threshold information. The text region probability information represents the likelihood of each pixel belonging to a text region, while the text region threshold information represents the dynamic segmentation threshold for the corresponding region. The system then combines these two types of information to generate a binary image of the text regions, clearly distinguishing them from the background. Next, a connected component extraction algorithm is used to aggregate continuously distributed text pixel regions in the binary image, and corresponding text region detection boxes are generated based on the boundary contours of each connected component, thus completing the text region localization in the entire image. Compared to traditional sliding window detection methods, this approach reduces interference from complex backgrounds and improves the detection stability of small fonts and low-contrast text.

[0031] After text region detection, the corresponding text region images are cropped from the original real-time image based on the detection bounding boxes for each text region. Since a single text region typically contains multiple lines of text, the system further utilizes pixel projection distribution to segment the text lines. Horizontal pixel projection involves counting the number of effective pixels in each line of the image and analyzing pixel density changes at different locations to determine concentrated text content areas and interline blank areas. When the pixel density of a region continuously increases, it can be determined that a text line exists in that region; when the pixel density significantly decreases, it can be determined as a text line gap area, thus achieving text line region localization. After segmenting the text lines, the system continues to perform vertical pixel projection statistics on each text line region. By analyzing the pixel distribution in different columns, it identifies the blank spaces between characters. Combining the pixel spacing width between adjacent characters, continuous character regions can be separated into multiple independent character images. This processing method effectively solves problems such as character adhesion, uneven character spacing, and small font overlap in floating window scenarios, improving the accuracy of subsequent character recognition.

[0032] In the character recognition stage, the size of each character image is first normalized to convert characters of different sizes into a fixed input size, ensuring a consistent input structure for the neural network. Then, the normalized character images are input into a character recognition model built on a convolutional recurrent neural network (RNN). The convolutional neural network primarily extracts local texture features of characters, such as character edges, stroke directions, and structural contours; the recurrent neural network learns the contextual relationships between character sequences, thereby improving the ability to distinguish between similar characters. The model outputs the character category and character recognition confidence score for each character. The character recognition confidence score indicates the reliability of the current character prediction result; a higher score indicates greater certainty in the model's recognition of that character. Finally, the system sequentially concatenates the character recognition results according to the order of characters in the text line to form a complete text recognition result.

[0033] To further improve the reliability of OCR recognition results, statistical analysis is performed on the confidence scores of each character. For example, the confidence scores of multiple characters in the same text can be averaged, weighted, or filtered for the lowest value to obtain the overall text recognition confidence score. Only when the overall text recognition confidence score reaches a preset confidence level will the system determine the corresponding text result as a valid real-time OCR recognition result; if the confidence score is insufficient, it indicates a significant risk of misrecognition, and the system can proceed to a subsequent dynamic optimization process. Through these methods, the system can not only perform text detection and character recognition in complex floating window scenarios but also effectively filter recognition results based on recognition confidence, thereby improving the overall stability and reliability of OCR recognition results.

[0034] Preferably, the initial text detection model is trained through the following steps: Multiple sample images containing floating window display interfaces are obtained, and the text regions in each sample image are labeled with rectangular boxes to obtain a text detection training dataset. Each sample image in the text detection training dataset is scaled, and the scaled sample images are uniformly resized according to a preset size to obtain standard training images. Based on the text region annotation results in each of the standard training images, generate corresponding text region probability label maps and text region threshold label maps. Each of the standard training images is input into a text detection network constructed based on a differentiable binarization algorithm, and the text detection network outputs the corresponding text region prediction probability map and text region prediction threshold map respectively. Based on the difference between the predicted probability map of the text region and the corresponding probability label map of the text region, the corresponding probability map loss value is calculated, and based on the difference between the predicted threshold map of the text region and the corresponding threshold label map of the text region, the corresponding threshold map loss value is calculated. Based on the probability graph loss value and the threshold graph loss value, the network parameters in the text detection network are iteratively updated until the preset training termination condition is met, and the initial text detection model is obtained. Specifically, a text detection model is essentially a deep learning network used to locate text regions in an image. The core of its training process lies in enabling the model to gradually learn the feature differences between text regions and background regions. In practice, multiple sample images containing floating window displays are first acquired. These sample images can come from different scenarios such as game interfaces, video playback interfaces, chat windows, or office interfaces to ensure the training data has scene diversity. Then, rectangular bounding boxes are used to label the text regions in each sample image, thus forming the corresponding text detection training dataset. Since text in floating windows often has issues such as transparent backgrounds, dynamic interference, and font size changes, training with samples from multiple scenarios can improve the model's ability to learn complex text features.

[0035] After data annotation, the sample images in the training dataset undergo unified preprocessing. Specifically, the original sample images are first scaled to ensure that images of different resolutions are converted to a uniform ratio, preventing network training instability caused by excessive image size differences. Then, the scaled images are uniformly resized according to a preset input size to obtain standard training images. This process ensures consistent tensor dimensions for subsequent neural network inputs, thereby improving batch training efficiency. Following this, the system generates corresponding text region probability label maps and text region threshold label maps based on the text region annotation results in each standard training image. The text region probability label map represents the true probability distribution of each pixel belonging to a text region, while the text region threshold label map represents the optimal segmentation threshold for different regions when performing text-to-background separation. Compared to traditional fixed threshold methods, this approach allows the model to learn more adaptable dynamic binarization boundaries for different regions, improving text extraction capabilities in complex backgrounds.

[0036] Next, standard training images are input into a text detection network built based on a differentiable binarization algorithm for training. The core of the differentiable binarization algorithm lies in transforming the traditional non-differentiable binarization process into a continuously differentiable process that can participate in the backpropagation calculation of the neural network, enabling the model to automatically learn the optimal text segmentation method during training. The text detection network outputs a text region prediction probability map and a text region prediction threshold map. The prediction probability map indicates which regions in the image the model believes may belong to text regions, while the prediction threshold map represents the dynamic segmentation thresholds automatically learned by the model for different regions. Subsequently, the system calculates the degree of difference between the prediction results and the true labels. For example, by comparing the pixel differences between the text region prediction probability map and the text region probability label map, the probability map loss value is obtained; by comparing the deviation between the prediction threshold map and the threshold label map, the threshold map loss value is obtained. The loss value essentially reflects the magnitude of the error between the model's current prediction result and the true result; the smaller the loss value, the closer the model's prediction effect is to the true labeled result.

[0037] Finally, based on the probabilistic graphical loss and the threshold graphical loss, the network parameters in the text detection network are iteratively updated. Specifically, the influence of each network parameter on the loss value can be calculated using the gradient backpropagation algorithm, and optimization algorithms can be used to adjust the convolutional layer weight parameters, bias parameters, and threshold learning parameters to make the prediction results in the next round closer to the true labels. As the number of training epochs increases, the model's ability to distinguish between text region features and background features gradually improves. When the loss value drops to a preset range, or when the training epochs reach a preset termination condition, parameter updates are stopped, resulting in the initial text detection model that has been trained. Through the above training method, the model can possess stronger text detection capabilities in complex scenes, improving the accuracy and stability of text region extraction in dynamic floating window environments.

[0038] The initial character recognition model is trained through the following steps: Multiple character sample images are acquired, and the character categories corresponding to each character sample image are labeled to obtain a character recognition training dataset; The character sample images in the character recognition training dataset are normalized to obtain standard character images; Each of the standard character images is input into a character recognition network constructed based on a convolutional recurrent neural network. The local features of the characters in each of the standard character images are extracted through a convolutional feature extraction layer to obtain the corresponding character feature map. Each of the character feature maps is input into the cyclic sequence recognition layer, and the character sequence features in the character feature maps are extracted to obtain the corresponding character sequence features; Based on the character sequence features, the character category corresponding to each of the standard character images is predicted to obtain the corresponding character category prediction result; Calculate the corresponding character recognition loss value based on the difference between the character category prediction result and the corresponding character category annotation result; Based on the character recognition loss value, the network parameters in the character recognition network are iteratively updated until the preset training termination condition is met, thus obtaining the initial character recognition model.

[0039] Specifically, a character recognition model is essentially a deep learning model used to determine character categories. Its core objective is to learn the structural feature differences between different characters. In practice, multiple character sample images are first acquired, and the character category corresponding to each sample is labeled, thus forming a character recognition training dataset. Character category labeling refers to assigning a corresponding real character content to each character image, such as numbers, letters, Chinese characters, or special symbols. To improve the model's adaptability to characters in complex scenes, the training data can include character samples with different fonts, sizes, brightness levels, and background conditions, enabling the model to learn more character variation features.

[0040] After constructing the dataset, the system performs size normalization on each character sample image. Size normalization refers to adjusting character images of different sizes to a fixed input size, thereby ensuring consistency in the subsequent neural network input structure. Since the character sizes within the floating window may vary significantly, uniform size processing reduces the impact of character scaling differences on model training stability. Subsequently, the standard character images are input into a character recognition network built on a convolutional recurrent neural network. The convolutional recurrent neural network is a deep learning structure that combines convolutional feature extraction capabilities with sequence feature learning capabilities. The convolutional feature extraction layer in the network is mainly used to extract local texture features of characters, such as stroke edges, transition structures, and character contour information, and generate corresponding character feature maps. Character feature maps are essentially high-dimensional representations of key feature information in the original character images, which can be used to enhance the distinguishability between different characters.

[0041] After obtaining the character feature maps, the system further inputs these maps into a recurrent sequence recognition layer to learn the sequential association features between characters. The recurrent sequence recognition layer typically employs a recurrent neural network structure, which leverages the contextual relationships between preceding and following characters to improve the recognition accuracy of similar characters. For example, for characters with similar structures, the arrangement patterns of preceding and following characters can be used to assist in determining the current character category. After recurrent sequence feature extraction, the system obtains the corresponding character sequence features and predicts the character category corresponding to each standard character image based on these features, outputting the corresponding character category prediction result.

[0042] Next, the character category prediction results are compared with the pre-labeled true character categories, and the corresponding character recognition loss value is calculated. The character recognition loss value reflects the degree of error between the current model's prediction result and the true result; the smaller the loss value, the closer the model's recognition result is to the true character category. In specific implementation, the difference between the predicted probability distribution and the true label can be calculated using the cross-entropy loss function. Subsequently, the system iteratively updates the network parameters in the character recognition network based on the character recognition loss value, continuously adjusting the convolutional layer parameters, recurrent layer parameters, and classification layer parameters through the gradient backpropagation algorithm, enabling the model to gradually learn a more accurate character feature representation ability. When the loss value drops to a preset range or the training epochs reach a preset termination condition, parameter updates are stopped, resulting in the initial character recognition model that has been trained.

[0043] Preferably, the step of evaluating the real-time OCR recognition result according to a preset multi-dimensional recognition effect evaluation strategy, and determining whether the real-time OCR recognition result meets the preset OCR recognition requirements based on the evaluation result, includes: The real-time OCR recognition result is matched with the preset standard recognition result corresponding to the real-time image data. Based on the matching result, the number of correctly matched characters, the number of incorrectly recognized characters, and the number of missing characters are counted. Based on the number of correctly matched characters, the number of incorrectly identified characters, and the number of missed characters, combined with the total number of characters identified, the corresponding character recognition accuracy is calculated, and the character recognition accuracy is determined as the corresponding recognition accuracy evaluation result. The number of valid image frames that have completed OCR recognition processing within a preset statistical time period is obtained, and the corresponding OCR recognition speed is calculated based on the number of valid image frames and the preset statistical time period. The OCR recognition speed is then determined as the corresponding recognition speed evaluation result. The recognition accuracy evaluation result is compared with a preset accuracy threshold, and the recognition speed evaluation result is compared with a preset speed threshold. If the recognition accuracy evaluation result is greater than or equal to the accuracy threshold, and the recognition speed evaluation result is greater than or equal to the speed threshold, then the real-time OCR recognition result meets the OCR recognition requirements. If the recognition accuracy evaluation result is less than the accuracy threshold and / or the recognition speed evaluation result is less than the speed threshold, then the real-time OCR recognition result does not meet the OCR recognition requirements.

[0044] Specifically, the standard recognition result refers to the predetermined correct text content, which can be derived from manual annotation results, historical verification results, or preset standard text, and serves as a benchmark for comparing OCR recognition results. First, the real-time OCR recognition result is matched character-by-character with the corresponding standard recognition result. The accuracy of the current recognition result is analyzed by comparing each character individually, and the number of correctly matched characters, the number of incorrectly recognized characters, and the number of missed characters are counted. The number of incorrectly recognized characters represents the number of characters in the OCR output that do not match the actual characters, and the number of missed characters represents the number of characters present in the actual text that the OCR failed to recognize. To improve matching accuracy, an edit distance algorithm can be used to dynamically align the character sequence, avoiding statistical errors caused by character position offsets.

[0045] After completing the character matching statistics, the corresponding character recognition accuracy is calculated based on the total number of characters recognized. Character recognition accuracy reflects the correctness of the current OCR model in recognizing text content; the higher the value, the closer the OCR recognition result is to the actual text content. In practice, the number of correctly matched characters can be taken as the number of effectively recognized characters, and the overall recognition ratio can be calculated by combining the number of incorrectly recognized characters with the number of missed characters. For example, the corresponding recognition accuracy can be obtained by dividing the number of correctly matched characters by the total number of characters recognized. This method can intuitively reflect the character recognition capability in the OCR process and provide a quantitative basis for subsequent optimization strategy selection.

[0046] In addition to recognition accuracy, OCR recognition speed is also evaluated. Specifically, the number of valid image frames that successfully complete OCR recognition processing within a preset time period is counted. A valid image frame refers to an image frame that completes the entire OCR recognition process and successfully outputs the recognition result. Then, the OCR recognition speed is calculated based on the number of valid image frames and the corresponding statistical time length; for example, the number of image frames processed per unit time is calculated to reflect real-time processing capabilities. Since floating window scenarios typically involve frequent dynamic refreshes and rapid interface changes, OCR recognition speed directly affects the actual smoothness of the system's interaction.

[0047] After obtaining the accuracy and speed evaluation results, they are compared with preset accuracy and speed thresholds, respectively. The accuracy threshold limits the minimum reliability of the OCR recognition result, while the speed threshold limits the system's minimum real-time processing capability. When both accuracy and speed meet their respective threshold requirements, the current OCR process ensures both accuracy and real-time performance, and the result is deemed to meet the OCR requirements. If the accuracy is below the threshold, there is a high risk of misidentification; if the speed is below the threshold, the OCR processing efficiency is insufficient. If either indicator fails to meet the requirements, the result is deemed not to meet the OCR requirements, and the process enters a subsequent dynamic optimization phase.

[0048] Preferably, please refer to Figure 2 When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the step of obtaining the corresponding target optimization strategy from the preset optimization strategy set based on the evaluation result includes: When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the Mask R-CNN instance segmentation algorithm is used to detect occlusion areas in the real-time image data to determine whether there are occlusion areas in the real-time image data and obtain the judgment result. Based on the judgment result, a first scene type for the floating window operation scene is determined, wherein the first scene type includes an occluded scene or an unoccluded scene; Based on the mapping relationship between the first scene type and the preset scene type and optimization strategy, a subset of candidate optimization strategies corresponding to the first scene type is determined in the set of optimization strategies, wherein the subset of candidate optimization strategies includes a subset of optimization strategies for occluded scenes or a subset of optimization strategies for non-occluded scenes. Based on the evaluation results, the corresponding target optimization strategy is selected from the subset of candidate optimization strategies.

[0049] Specifically, the optimization strategy set is a pre-established collection of multiple OCR optimization schemes, with different optimization strategies corresponding to different scenario problems, such as text occlusion, background interference, image blurring, or excessive model computation. The target optimization strategy is the most suitable optimization scheme selected from multiple candidate optimization schemes based on the current evaluation results and scenario status.

[0050] In practice, when the current OCR recognition accuracy or speed fails to meet preset requirements, the Mask R-CNN instance segmentation algorithm is first used to detect occlusion regions in the real-time image data. Mask R-CNN is a deep learning-based instance segmentation algorithm that can not only detect target locations but also output pixel-level contour information of the target region, thus more accurately distinguishing between text regions and occluded regions. After the real-time image is input into the Mask R-CNN network, the network extracts features from the target region in the image and outputs the corresponding occlusion region mask result. Subsequently, the system determines whether there are occlusion regions affecting OCR recognition in the current real-time image based on the area, position, and degree of overlap between the occluded region and the text region. For example, when the occluded region covers the text region to a preset proportion, it can be determined that there is a text occlusion problem in the current scene.

[0051] After obtaining the occlusion detection results, the system further determines the first scene type corresponding to the current floating window's operating scenario. An occluded scene refers to a scenario where the text area is covered by pop-ups, icons, dynamic layers, or interface elements; a non-occluded scene indicates that the text area is completely intact, with only ordinary background interference or image quality issues. Subsequently, based on the preset mapping relationship between scene types and optimization strategies, the system selects a corresponding subset of candidate optimization strategies from the optimization strategy set. Specifically, when the current scene is an occluded scene, the system prioritizes calling the occluded scene optimization strategy subset, such as performing occluded area compensation, local text enhancement, occluded area reconstruction, or character context correction optimizations; when the current scene is a non-occluded scene, it prioritizes calling the non-occluded scene optimization strategy subset, such as adjusting the binarization threshold, optimizing character input size, or reducing model inference complexity.

[0052] After filtering the subset of candidate optimization strategies, the final target optimization strategy is further determined based on the current evaluation results. For example, when the recognition accuracy is low, optimization strategies that enhance text edge details or increase character resolution can be prioritized; when the recognition speed is low, optimization strategies that reduce the input image size or reduce the number of model nodes can be prioritized. By classifying the scene first and then filtering the strategy, the problem of using a fixed optimization method for all scenes can be avoided. This allows the system to dynamically select a more suitable OCR optimization scheme based on the current floating window operating state, thereby improving the stability and adaptability of OCR recognition in complex scenes.

[0053] Preferably, selecting the corresponding target optimization strategy from the subset of candidate optimization strategies based on the evaluation results includes: If the evaluation result is that the recognition accuracy does not meet the requirements but the recognition speed does meet the requirements, then at least one of the following optimization strategies is selected from the candidate optimization strategy subset: adjusting the text region binarization threshold to a first preset threshold range, adjusting the text region detection box intersection-union ratio threshold to a first preset intersection-union ratio range, adjusting the character image to a first preset input size, and adjusting the number of hidden layer nodes of the convolutional recurrent neural network to a first preset number of nodes range, as the target optimization strategy. If the evaluation result is that the recognition accuracy meets the requirements but the recognition speed does not meet the requirements, then at least one of the following optimization strategies is selected from the candidate optimization strategy subset: adjusting the text region binarization threshold to a second preset threshold range, adjusting the text region detection box intersection-union ratio threshold to a second preset intersection-union ratio range, adjusting the character image to a second preset input size, and adjusting the number of hidden layer nodes of the convolutional recurrent neural network to a second preset node number range as the target optimization strategy. Wherein, the value corresponding to the first preset threshold interval is less than the value corresponding to the second preset threshold interval, the image resolution corresponding to the first preset input size is greater than the image resolution corresponding to the second preset input size, and the number of nodes corresponding to the first preset node number interval is greater than the number of nodes corresponding to the second preset node number interval. If the evaluation result is that the recognition accuracy and recognition speed do not meet the requirements, then a combined optimization strategy including the text region binarization threshold adjustment strategy, the text region detection box intersection-union threshold adjustment strategy, the character image input size adjustment strategy, and the convolutional recurrent neural network hidden layer node number adjustment strategy is selected from the candidate optimization strategy subset as the target optimization strategy.

[0054] Specifically, based on the types of problems encountered during the current OCR recognition process, a more suitable parameter adjustment scheme is automatically selected, enabling targeted optimization of subsequent OCR recognition processes. Among these, the text region binarization threshold is a crucial parameter in the text region extraction stage, distinguishing text regions from background regions in an image. For example, when the grayscale value of a pixel in an image is higher than a preset threshold, the pixel is determined to belong to a text region; if it is lower than the threshold, it is determined to be a background region. If the current recognition accuracy is low but the recognition speed is normal, it indicates that the system's computing power is sufficient, but some text details are not being extracted correctly. For example, in a game floating window scene, when small white text is superimposed on a dynamic background, a high binarization threshold may misidentify some light-colored text edges as background regions, resulting in missing text. Therefore, the binarization threshold is adjusted to the first preset threshold range, i.e., a lower threshold range, so that more text edge regions are preserved, thereby improving the completeness of text detection.

[0055] Simultaneously, the system adjusts the intersection-union (IU) threshold for text region detection boxes. IU refers to the ratio of the overlapping area of ​​two detection boxes to the total coverage area, primarily used to filter out duplicate detection boxes. For example, when multiple detection boxes simultaneously cover the same text region, the system typically retains the highest-scoring detection box and deletes other detection boxes with high overlap. If the recognition accuracy is low, it indicates that some text regions may be mistakenly deleted; therefore, the system lowers the IU threshold to retain more candidate text regions. For instance, in a chat window where adjacent text is closely spaced, a higher IU threshold might incorrectly merge adjacent text, while lowering the threshold retains more independent text regions, improving text separation.

[0056] In addition, the input size of the character image and the number of hidden layer nodes in the convolutional recurrent neural network will be increased. The character input size refers to the image resolution before the character image enters the recognition model. For example, when the original input size is 32×32 pixels, some complex Chinese character strokes may appear blurred; after adjusting to 64×64 pixels, the character edge details will be clearer, thus improving character recognition accuracy. The number of hidden layer nodes in the convolutional recurrent neural network controls the model's feature learning ability; the more nodes, the more complex the character features the model can learn. For example, in scenarios with large font deformations in bullet screen text, increasing the number of hidden layer nodes allows the model to learn more character contour variation patterns, thereby improving the ability to recognize complex characters. Since the current recognition speed already meets the requirements, the system can allow for some increased computational load in exchange for higher recognition accuracy.

[0057] When the recognition accuracy meets the requirements but the recognition speed is insufficient, it indicates that the current model's processing computation is too high. In this case, the system will adjust the aforementioned parameters in reverse. For example, increasing the binarization threshold ensures that only more obvious text regions are retained, reducing the number of regions to be processed subsequently; increasing the cross-union threshold of detection boxes allows duplicate detection boxes to be filtered out more quickly, reducing the computational load on invalid regions. Simultaneously, the system will reduce the character input size. For example, reducing the character input size from 64×64 to 32×32 significantly reduces the number of pixels in the convolution calculation process, thus reducing character recognition time. Furthermore, the number of hidden layer nodes will be reduced to shrink the network parameter scale. For example, when the number of hidden layer nodes was originally 512, the model's computational load was high; reducing it to 256 reduces the number of matrix operations in the feature calculation process, thereby improving the overall OCR recognition speed.

[0058] When both recognition accuracy and speed fail to meet requirements, it indicates poor overall adaptability of the current OCR recognition process. For example, in a video playback floating window scenario, there is both dynamic background interference and insufficient system resources; adjusting a single parameter alone often fails to solve the problem. Therefore, a combined optimization strategy is adopted, adjusting multiple parameters simultaneously. For instance, the binarization threshold can be appropriately lowered to enhance text edge extraction, while the character input size and the number of hidden layer nodes can be appropriately reduced to control the overall model computation, bringing both recognition accuracy and speed closer to the target requirements. Through this multi-parameter joint optimization approach, the OCR processing flow can be dynamically adjusted according to different floating window operating states, improving overall recognition stability and real-time adaptability in complex scenarios.

[0059] Preferably, if the evaluation result is that the recognition accuracy and recognition speed do not meet the requirements, then a combined optimization strategy including a text region binarization threshold adjustment strategy, a text region detection box intersection-union threshold adjustment strategy, a character image input size adjustment strategy, and a convolutional recurrent neural network hidden layer node number adjustment strategy is selected from the candidate optimization strategy subset as the target optimization strategy, including: Obtain the strategy parameter combination corresponding to multiple candidate combination optimization strategies in the candidate optimization strategy subset, wherein each strategy parameter combination includes a text region binarization threshold parameter, a text region detection box intersection-union ratio threshold parameter, a character image input size parameter, and a convolutional recurrent neural network hidden layer node number parameter; Based on the combination of strategy parameters, the real-time image data is subjected to OCR recognition processing to obtain multiple candidate OCR recognition results. The recognition accuracy and recognition speed of each candidate OCR recognition result are evaluated to obtain the corresponding candidate recognition accuracy evaluation results and candidate recognition speed evaluation results. Based on the evaluation results of the accuracy of each candidate recognition and the evaluation results of the speed of each candidate recognition, a target candidate combination optimization strategy that simultaneously satisfies the preset accuracy condition and the preset speed condition is selected. If there are multiple target candidate combination optimization strategies, then based on the degree of deviation between the recognition accuracy evaluation result and the recognition speed evaluation result corresponding to each target candidate combination optimization strategy, the target candidate combination optimization strategy with the smallest deviation degree is selected as the target optimization strategy; If there is no target candidate combination optimization strategy that satisfies the preset accuracy condition and the preset speed condition, then the target candidate combination optimization strategy with the smallest combined deviation of the recognition accuracy evaluation result and the recognition speed evaluation result from the preset threshold is selected as the target optimization strategy.

[0060] Specifically, when both OCR recognition accuracy and speed fail to meet requirements, the system automatically selects the optimal strategy from multiple parameter combinations to achieve the best overall performance. This avoids the problem of adjusting a single parameter failing to balance recognition accuracy and operational efficiency. Here, the strategy parameter combination refers to a configuration scheme combining multiple OCR optimization parameters, with each combination corresponding to a different OCR recognition operation mode. For example, one parameter combination might use a lower text region binarization threshold, a larger character input size, and fewer hidden layer nodes, while another might use a medium binarization threshold, a smaller input size, and more hidden layer nodes. Different parameter combinations result in different recognition effects and processing speeds.

[0061] In practice, the system first obtains multiple candidate optimization strategy combinations and their corresponding strategy parameter combinations from a subset of candidate optimization strategies. Among these, the text region binarization threshold parameter controls the segmentation sensitivity between text and background regions; the text region detection box intersection-union (IUU) threshold parameter controls the stringency of candidate detection box selection; the character image input size parameter controls the character image resolution; and the number of hidden layer nodes in the convolutional recurrent neural network controls the model's feature learning ability and computational complexity. For example, in a certain parameter combination, the system might set the binarization threshold to a lower value to enhance the extraction of light-colored text, while adjusting the character input size to 64×64 pixels to improve character detail resolution. However, to avoid excessive computation, the number of hidden layer nodes is reduced to a smaller range to balance the overall recognition speed.

[0062] Subsequently, the system performs OCR recognition processing on the real-time image data based on various strategy parameter combinations. That is, the system repeatedly executes text detection, character segmentation, and character recognition processes using different parameter configurations, obtaining corresponding candidate OCR recognition results. For example, the first set of parameters may improve recognition accuracy but slow down processing speed, while the second set of parameters may process faster but reduce the ability to recognize some complex characters. Afterward, the system performs recognition accuracy and recognition speed evaluations on each candidate OCR recognition result, thereby obtaining the candidate recognition accuracy and candidate recognition speed results corresponding to each parameter combination.

[0063] After obtaining all evaluation results, the system first selects target candidate optimization strategies that simultaneously meet preset accuracy and speed requirements. For example, the preset recognition accuracy must be greater than 95%, and the recognition speed must reach 20 frames per second. When a parameter combination meets both conditions, it indicates that the parameter combination can balance recognition accuracy and real-time processing capability. If multiple target candidate optimization strategies meet the conditions, the system further compares the degree of deviation between different strategies. Here, the degree of deviation refers to the comprehensive difference between the current recognition accuracy and recognition speed relative to the preset target values. For example, one strategy may achieve an accuracy of 98%, but its speed is only slightly above the threshold; another strategy may have both accuracy and speed closer to the ideal target value. In this case, the system prioritizes the scheme with the smallest comprehensive deviation as the final target optimization strategy, thereby avoiding excessive bias towards a single performance indicator.

[0064] If none of the candidate optimization strategies can simultaneously meet the accuracy and speed requirements, the system further calculates the overall deviation of each parameter combination from a preset threshold. For example, one set of parameters may have slightly lower accuracy than the target value, but significantly better speed than other solutions; another set of parameters may have higher accuracy but severely insufficient speed. In this case, the system will comprehensively compare the overall difference between the two indicators and the target threshold, and select the parameter combination with the smallest overall deviation as the target optimization strategy. In this way, even if the ideal requirements cannot be fully met in complex scenarios, the system can still automatically select the OCR optimization scheme with the best overall performance under the current conditions.

[0065] Preferably, the OCR recognition requirement is obtained through the following steps: In response to the OCR recognition command input by the user, the OCR recognition command is parsed to obtain the second scene type of the floating window running scene, wherein the second scene type includes game scene, video playback scene, instant messaging scene, web browsing scene and office document scene; The OCR recognition requirements are determined based on the mapping relationship between the second scene type and the preset scene type and recognition requirements.

[0066] Specifically, the system dynamically determines the corresponding OCR recognition requirements based on the actual usage needs of the current floating window operation scenario, thereby avoiding the problem of using a fixed recognition standard for all scenarios. The OCR recognition command is the control information input by the user when triggering the OCR function. It can be expressed as clicking a recognition button, a voice recognition command, a shortcut key trigger command, or an automatic recognition request from the floating window. After receiving the OCR recognition command, the system first parses the command content to identify the current floating window's operating scenario type. The second scenario type indicates the business usage scenario corresponding to the current OCR function, such as a game scenario, video playback scenario, instant messaging scenario, web browsing scenario, and office document scenario. Different scenarios have significantly different requirements for OCR recognition accuracy and real-time performance.

[0067] In practice, the current floating window's operating environment is identified by combining the current foreground application type, window title information, interface element characteristics, and user operation behavior. For example, when game frame rendering features, dynamic health bar icons, or real-time bullet screen elements are detected on the current interface, it can be identified as a game scene; when a video player progress bar, playback control area, or continuous dynamic image frames are detected, it can be identified as a video playback scene; when chat bubbles, contact lists, and message input boxes are present on the interface, it can be identified as an instant messaging scene. For office document scenes, table areas, document layout structures, or continuous text areas can be further identified to improve the accuracy of scene classification.

[0068] After completing scene type identification, the system automatically determines the corresponding OCR recognition requirements based on the pre-established mapping relationship between scene types and recognition requirements. This mapping relationship is essentially a configuration rule between different scenes and corresponding OCR performance metrics. For example, in game and video playback scenarios, due to the rapid changes in the interface, the system prioritizes increasing OCR recognition speed requirements and appropriately reduces the accuracy requirements for some complex characters to ensure real-time responsiveness. In office document scenarios, because the text content is usually relatively stable, the system increases recognition accuracy requirements, such as requiring higher character recognition accuracy and a lower missed recognition rate. For instant messaging scenarios, the system needs to balance recognition speed and accuracy to ensure that chat messages can be recognized quickly and accurately.

[0069] For example, when a user activates the OCR translation function in a game's floating window, the system can automatically increase the OCR recognition speed threshold to a higher level to ensure that subtitles or task text can be translated in real time. Conversely, when a user activates the OCR extraction function in an office document's floating window, the system will increase the character recognition accuracy threshold to reduce typos and omissions in the document content. Through this scenario-adaptive approach, the system can dynamically adjust the OCR recognition standards according to different floating window operating environments, improving the relevance, flexibility, and adaptability of the OCR recognition process to practical applications. Example 2

[0070] Please see Figure 3 Embodiment 2 of the present invention also provides a device for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window, the device comprising: The image acquisition module is used to acquire real-time image data corresponding to the display area of ​​the floating window in the floating window operation scenario; The OCR recognition module is used to input the real-time image data into a preset initial OCR recognition model to obtain the real-time OCR recognition result; The evaluation module is used to evaluate the real-time OCR recognition result according to the preset multi-dimensional recognition effect evaluation strategy, and to determine whether the real-time OCR recognition result meets the preset OCR recognition requirements based on the evaluation result. The multi-dimensional recognition effect evaluation strategy includes recognition accuracy evaluation, recognition speed evaluation, and occlusion robustness evaluation. The first output module for recognition results is used to output the real-time OCR recognition result when it is determined that the real-time OCR recognition result meets the OCR recognition requirements; The optimization strategy acquisition module is used to acquire the corresponding target optimization strategy from the preset optimization strategy set according to the evaluation result when it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements. An optimization processing module is used to optimize the real-time image data and / or the initial OCR recognition model according to the target optimization strategy, so as to obtain the optimized target image data and / or target OCR recognition model. The second output module for recognition results is used to take the target image data as the new real-time image data and / or take the target OCR recognition model as the new initial OCR recognition model, return to execute the step of inputting the real-time image data into the preset initial OCR recognition model to obtain the real-time OCR recognition result, until it is determined that the real-time OCR recognition result meets the OCR recognition requirements, and output the real-time OCR recognition result.

[0071] Specifically, the floating window real-time OCR recognition effect evaluation and dynamic optimization device provided in this embodiment of the invention first acquires real-time image data corresponding to the floating window display area, and uses an initial OCR recognition model to perform text detection and character recognition on the real-time image data to obtain real-time OCR recognition results. Subsequently, the current recognition results are analyzed in multiple dimensions through recognition accuracy evaluation and recognition speed evaluation to determine whether the current OCR recognition results meet the preset OCR recognition requirements. When the recognition results do not meet the requirements, the fixed OCR recognition process is no longer used. Instead, the corresponding target optimization strategy is dynamically acquired from the preset optimization strategy set based on the evaluation results, and the real-time image data and / or the initial OCR recognition model are optimized for the current scenario. For example, the text detection parameters, character recognition parameters, and model structure parameters are adjusted to generate optimized target image data or target OCR recognition models. Then, the optimized target image data or target OCR recognition models are re-entered into the OCR recognition process for recognition and evaluation again until the recognition results meet the corresponding recognition requirements. By employing the above methods, OCR recognition can dynamically adjust its recognition strategy based on the recognition status under different floating window scenarios, ensuring both recognition accuracy and speed. This allows for dynamic adjustment of corresponding optimization strategies based on the floating window OCR recognition effect, thereby achieving synergistic optimization of recognition accuracy and speed. Example

[0072] In addition, combined Figure 1 The floating window real-time OCR recognition effect evaluation and dynamic optimization method of Embodiment 1 of the present invention described herein can be implemented by an electronic device. Figure 4 A schematic diagram of the hardware structure of the electronic device provided in Embodiment 3 of the present invention is shown.

[0073] Electronic devices may include processors and memory storing computer program instructions.

[0074] Specifically, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.

[0075] The memory may include a large-capacity storage device for data or instructions. For example, and not limitingly, the memory may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to a data processing device. In a particular embodiment, the memory is a non-volatile solid-state memory. In a particular embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0076] The processor reads and executes computer program instructions stored in the memory to implement any of the floating window real-time OCR recognition effect evaluation and dynamic optimization methods in the above embodiments.

[0077] In one example, the electronic device may also include a communication interface and a bus. For example, Figure 4 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.

[0078] The communication interface is mainly used to enable communication between various modules, devices, units and / or equipment in the embodiments of the present invention.

[0079] A bus, including hardware, software, or both, couples components of the device together. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, a bus may include one or more buses. While specific buses are described and illustrated in embodiments of the invention, the invention contemplates any suitable bus or interconnect.

[0080] In summary, the embodiments of the present invention provide a method, apparatus, and device for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window.

[0081] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0082] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0083] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant locality, and corresponding operation entry points shall be provided for the user to choose to authorize or refuse.

[0084] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0085] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A method for evaluating and dynamically optimizing the real-time OCR recognition performance of a floating window, characterized in that, The method includes: Acquire real-time image data corresponding to the display area of ​​the floating window in the floating window operation scenario; The real-time image data is input into a preset initial OCR recognition model to obtain the real-time OCR recognition result; The real-time OCR recognition result is evaluated according to the preset multi-dimensional recognition effect evaluation strategy, and the real-time OCR recognition result is judged to meet the preset OCR recognition requirements based on the evaluation result. The multi-dimensional recognition effect evaluation strategy includes recognition accuracy evaluation and recognition speed evaluation. When it is determined that the real-time OCR recognition result meets the OCR recognition requirements, the real-time OCR recognition result is output. When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the corresponding target optimization strategy is obtained from the preset optimization strategy set according to the evaluation result; According to the target optimization strategy, the real-time image data and / or the initial OCR recognition model are optimized to obtain the optimized target image data and / or target OCR recognition model. The target image data is used as the new real-time image data and / or the target OCR recognition model is used as the new initial OCR recognition model. The process of inputting the real-time image data into the preset initial OCR recognition model to obtain the real-time OCR recognition result is repeated until it is determined that the real-time OCR recognition result meets the OCR recognition requirements, and then the real-time OCR recognition result is output.

2. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window according to claim 1, characterized in that, The step of inputting the real-time image data into a preset initial OCR recognition model to obtain the real-time OCR recognition result includes: The real-time image data is input into an initial text detection model constructed based on a differentiable binarization algorithm. The initial text detection model is used to calculate the probability information and threshold information of the text regions in the real-time image data to generate the corresponding binary image of the text regions. The connected regions in the binary image of the text region are extracted, and corresponding text region detection boxes are generated based on the boundary contours of each connected region. Based on each of the text region detection boxes, the corresponding text regions in the real-time image data are cropped to obtain multiple text region images; The pixel projection distribution results of each text region image in the horizontal direction are obtained respectively, and the corresponding text line region is determined according to the pixel projection distribution results in the horizontal direction; The pixel projection distribution results of each text line region in the vertical direction are statistically analyzed, and the candidate regions of characters in each text line region are separated according to the pixel interval width between adjacent characters to obtain multiple character images; The size of each character image is normalized, and the normalized character image is input into an initial character recognition model constructed based on a convolutional recurrent neural network. The initial character recognition model is used to identify the character category corresponding to each character image to obtain the corresponding character recognition result and character recognition confidence. According to the arrangement order of each character recognition result in the corresponding text line area, the character recognition results are concatenated to generate the corresponding text recognition result; Based on the confidence scores of each character recognition, the confidence scores of the corresponding text recognition results are statistically analyzed, and the text recognition results whose statistically obtained text recognition confidence scores meet the preset confidence conditions are determined as the real-time OCR recognition results.

3. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of the floating window according to claim 2, characterized in that, The initial text detection model is trained through the following steps: Multiple sample images containing floating window display interfaces are obtained, and the text regions in each sample image are labeled with rectangular boxes to obtain a text detection training dataset. Each sample image in the text detection training dataset is scaled, and the scaled sample images are uniformly resized according to a preset size to obtain standard training images. Based on the text region annotation results in each of the standard training images, generate corresponding text region probability label maps and text region threshold label maps. Each of the standard training images is input into a text detection network constructed based on a differentiable binarization algorithm, and the text detection network outputs the corresponding text region prediction probability map and text region prediction threshold map respectively. Based on the difference between the predicted probability map of the text region and the corresponding probability label map of the text region, the corresponding probability map loss value is calculated, and based on the difference between the predicted threshold map of the text region and the corresponding threshold label map of the text region, the corresponding threshold map loss value is calculated. Based on the probability graph loss value and the threshold graph loss value, the network parameters in the text detection network are iteratively updated until the preset training termination condition is met, and the initial text detection model is obtained. The initial character recognition model is trained through the following steps: Multiple character sample images are acquired, and the character categories corresponding to each character sample image are labeled to obtain a character recognition training dataset; The character sample images in the character recognition training dataset are normalized to obtain standard character images; Each of the standard character images is input into a character recognition network constructed based on a convolutional recurrent neural network. The local features of the characters in each of the standard character images are extracted through a convolutional feature extraction layer to obtain the corresponding character feature map. Each of the character feature maps is input into the cyclic sequence recognition layer, and the character sequence features in the character feature maps are extracted to obtain the corresponding character sequence features; Based on the character sequence features, the character category corresponding to each of the standard character images is predicted to obtain the corresponding character category prediction result; Calculate the corresponding character recognition loss value based on the difference between the character category prediction result and the corresponding character category annotation result; Based on the character recognition loss value, the network parameters in the character recognition network are iteratively updated until the preset training termination condition is met, thus obtaining the initial character recognition model.

4. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window according to claim 1, characterized in that, The step of evaluating the real-time OCR recognition result according to a preset multi-dimensional recognition effect evaluation strategy, and determining whether the real-time OCR recognition result meets the preset OCR recognition requirements based on the evaluation result, includes: The real-time OCR recognition result is matched with the preset standard recognition result corresponding to the real-time image data. Based on the matching result, the number of correctly matched characters, the number of incorrectly recognized characters, and the number of missing characters are counted. Based on the number of correctly matched characters, the number of incorrectly identified characters, and the number of missed characters, combined with the total number of characters identified, the corresponding character recognition accuracy is calculated, and the character recognition accuracy is determined as the corresponding recognition accuracy evaluation result. The number of valid image frames that have completed OCR recognition processing within a preset statistical time period is obtained, and the corresponding OCR recognition speed is calculated based on the number of valid image frames and the preset statistical time period. The OCR recognition speed is then determined as the corresponding recognition speed evaluation result. The recognition accuracy evaluation result is compared with a preset accuracy threshold, and the recognition speed evaluation result is compared with a preset speed threshold. If the recognition accuracy evaluation result is greater than or equal to the accuracy threshold, and the recognition speed evaluation result is greater than or equal to the speed threshold, then the real-time OCR recognition result meets the OCR recognition requirements. If the recognition accuracy evaluation result is less than the accuracy threshold and / or the recognition speed evaluation result is less than the speed threshold, then the real-time OCR recognition result does not meet the OCR recognition requirements.

5. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window according to claim 1, characterized in that, When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the step of obtaining the corresponding target optimization strategy from the preset optimization strategy set based on the evaluation result includes: When it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements, the Mask R-CNN instance segmentation algorithm is used to detect occlusion areas in the real-time image data to determine whether there are occlusion areas in the real-time image data and obtain the judgment result. Based on the judgment result, a first scene type for the floating window operation scene is determined, wherein the first scene type includes an occluded scene or an unoccluded scene; Based on the mapping relationship between the first scene type and the preset scene type and optimization strategy, a subset of candidate optimization strategies corresponding to the first scene type is determined in the set of optimization strategies, wherein the subset of candidate optimization strategies includes a subset of optimization strategies for occluded scenes or a subset of optimization strategies for non-occluded scenes. Based on the evaluation results, the corresponding target optimization strategy is selected from the subset of candidate optimization strategies.

6. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window according to claim 5, characterized in that, The step of selecting the corresponding target optimization strategy from the subset of candidate optimization strategies based on the evaluation results includes: If the evaluation result is that the recognition accuracy does not meet the requirements but the recognition speed does meet the requirements, then at least one of the following optimization strategies is selected from the candidate optimization strategy subset: adjusting the text region binarization threshold to a first preset threshold range, adjusting the text region detection box intersection-union ratio threshold to a first preset intersection-union ratio range, adjusting the character image to a first preset input size, and adjusting the number of hidden layer nodes of the convolutional recurrent neural network to a first preset number of nodes range, as the target optimization strategy. If the evaluation result is that the recognition accuracy meets the requirements but the recognition speed does not meet the requirements, then at least one of the following optimization strategies is selected from the candidate optimization strategy subset: adjusting the text region binarization threshold to a second preset threshold range, adjusting the text region detection box intersection-union ratio threshold to a second preset intersection-union ratio range, adjusting the character image to a second preset input size, and adjusting the number of hidden layer nodes of the convolutional recurrent neural network to a second preset node number range as the target optimization strategy. Wherein, the value corresponding to the first preset threshold interval is less than the value corresponding to the second preset threshold interval, the image resolution corresponding to the first preset input size is greater than the image resolution corresponding to the second preset input size, and the number of nodes corresponding to the first preset node number interval is greater than the number of nodes corresponding to the second preset node number interval. If the evaluation result is that the recognition accuracy and recognition speed do not meet the requirements, then a combined optimization strategy including the text region binarization threshold adjustment strategy, the text region detection box intersection-union threshold adjustment strategy, the character image input size adjustment strategy, and the convolutional recurrent neural network hidden layer node number adjustment strategy is selected from the candidate optimization strategy subset as the target optimization strategy.

7. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window according to claim 6, characterized in that, If the evaluation result indicates that the recognition accuracy and recognition speed do not meet the requirements, then a combined optimization strategy, including a text region binarization threshold adjustment strategy, a text region detection box intersection-union threshold adjustment strategy, a character image input size adjustment strategy, and a convolutional recurrent neural network hidden layer node number adjustment strategy, is selected from the candidate optimization strategy subset as the target optimization strategy. Obtain the strategy parameter combination corresponding to multiple candidate combination optimization strategies in the candidate optimization strategy subset, wherein each strategy parameter combination includes a text region binarization threshold parameter, a text region detection box intersection-union ratio threshold parameter, a character image input size parameter, and a convolutional recurrent neural network hidden layer node number parameter; Based on the combination of strategy parameters, the real-time image data is subjected to OCR recognition processing to obtain multiple candidate OCR recognition results. The recognition accuracy and recognition speed of each candidate OCR recognition result are evaluated to obtain the corresponding candidate recognition accuracy evaluation results and candidate recognition speed evaluation results. Based on the evaluation results of the accuracy of each candidate recognition and the evaluation results of the speed of each candidate recognition, a target candidate combination optimization strategy that simultaneously satisfies the preset accuracy condition and the preset speed condition is selected. If there are multiple target candidate combination optimization strategies, then based on the degree of deviation between the recognition accuracy evaluation result and the recognition speed evaluation result corresponding to each target candidate combination optimization strategy, the target candidate combination optimization strategy with the smallest deviation degree is selected as the target optimization strategy; If there is no target candidate combination optimization strategy that satisfies the preset accuracy condition and the preset speed condition, then the target candidate combination optimization strategy with the smallest combined deviation of the recognition accuracy evaluation result and the recognition speed evaluation result from the preset threshold is selected as the target optimization strategy.

8. The method for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window according to any one of claims 1-7, characterized in that, The OCR recognition requirements are obtained through the following steps: In response to the OCR recognition command input by the user, the OCR recognition command is parsed to obtain the second scene type of the floating window running scene, wherein the second scene type includes game scene, video playback scene, instant messaging scene, web browsing scene and office document scene; The OCR recognition requirements are determined based on the mapping relationship between the second scene type and the preset scene type and recognition requirements.

9. A device for evaluating and dynamically optimizing the real-time OCR recognition effect of a floating window, characterized in that, The device includes: The image acquisition module is used to acquire real-time image data corresponding to the display area of ​​the floating window in the floating window operation scenario; The OCR recognition module is used to input the real-time image data into a preset initial OCR recognition model to obtain the real-time OCR recognition result; The evaluation module is used to evaluate the real-time OCR recognition result according to the preset multi-dimensional recognition effect evaluation strategy, and to determine whether the real-time OCR recognition result meets the preset OCR recognition requirements based on the evaluation result. The multi-dimensional recognition effect evaluation strategy includes recognition accuracy evaluation, recognition speed evaluation, and occlusion robustness evaluation. The first output module for recognition results is used to output the real-time OCR recognition result when it is determined that the real-time OCR recognition result meets the OCR recognition requirements; The optimization strategy acquisition module is used to acquire the corresponding target optimization strategy from the preset optimization strategy set according to the evaluation result when it is determined that the real-time OCR recognition result does not meet the OCR recognition requirements. An optimization processing module is used to optimize the real-time image data and / or the initial OCR recognition model according to the target optimization strategy, so as to obtain the optimized target image data and / or target OCR recognition model. The second output module for recognition results is used to take the target image data as the new real-time image data and / or take the target OCR recognition model as the new initial OCR recognition model, return to execute the step of inputting the real-time image data into the preset initial OCR recognition model to obtain the real-time OCR recognition result, until it is determined that the real-time OCR recognition result meets the OCR recognition requirements, and output the real-time OCR recognition result.

10. An electronic device, characterized in that, include: At least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method as described in any one of claims 1-8.