Data processing method and device, equipment, storage medium and program product
By combining user-side fingertip swipe gestures with the recognition gateway's fingertip recognition and text correction technology, the problem of typos and omissions in live stream subtitles was solved, enabling rapid subtitle correction and improved user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-17
AI Technical Summary
The live streaming subtitle system has problems with typos and omissions, making it difficult for viewers to provide quick and effective feedback, which leads to a decline in the live streaming effect and a poor user experience.
By using the user's fingertip swipe operation, the system utilizes a recognition gateway for fingertip recognition and text recognition, combined with accidental touch detection, model self-learning, and secondary error correction technology to achieve rapid correction of subtitle errors.
It enables timely error correction of live broadcast subtitles, improves subtitle accuracy and user experience, reduces user workload, and enhances system performance.
Smart Images

Figure CN121879657A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of data identification and data interaction, and in particular to a data processing method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the rapid development of live streaming technology, live streaming has become an important form of entertainment and information acquisition in people's daily lives. During live streams, subtitles, as a crucial means of information transmission, have a significant impact on the overall effect. However, current live streaming subtitle systems often suffer from typos and omissions, making it difficult for viewers to provide quick and effective feedback, leading to a decline in the live stream's quality and a poor viewer experience. Summary of the Invention
[0003] To address the aforementioned technical problems, embodiments of the present invention provide a data processing method, apparatus, device, storage medium, and program product.
[0004] The data processing method provided in this application embodiment is applied to an identification gateway and includes: The receiving terminal sends fingertip information; the fingertip information is generated by the user performing a swipe operation on a target area of the terminal, and the swipe operation is a valid operation. Finger tip recognition is performed on the fingertip information to obtain a sliding area; text recognition is performed on the sliding area to obtain the first text in the sliding area; The terminal sends a confirmation message, which is used to confirm that the first text in the sliding area has been corrected. The first text in the sliding area is subjected to rule-based error correction and / or text error correction to obtain the second text; The second text is sent to the terminal, which is used to replace the first text in the sliding area with the second text.
[0005] The data processing method provided in this application embodiment is applied to a terminal and includes: In response to a user's swipe gesture on a target area, determine whether the swipe gesture is a valid gesture; If the sliding operation is determined to be a valid operation, the fingertip information corresponding to the sliding operation is determined and the fingertip information is sent to the recognition gateway. The recognition gateway is used to perform fingertip recognition on the fingertip information and to perform text recognition on the obtained sliding area. In response to the user's release operation on the target area, the sliding area is highlighted; in response to the user's confirmation operation on the sliding area, confirmation information is sent to the recognition gateway, the confirmation information being used to confirm the correction of the first text in the sliding area; The system receives the second text sent by the identification gateway and replaces the first text in the sliding area with the second text; the second text is obtained by the identification gateway after performing rule correction and / or text correction on the first text in the sliding area.
[0006] The processing device provided in this application includes a processor and a memory, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs stored in the memory to execute any of the above-described data processing methods.
[0007] The computer-readable storage medium provided in this application embodiment is used to store a computer program that causes a computer to execute any of the above-described data processing methods.
[0008] The computer program product provided in this application includes computer program instructions that cause a computer to execute any of the above-described data processing methods.
[0009] In the technical solution of this application embodiment, the identification gateway receives fingertip information sent by the terminal; the fingertip information is information generated by the user performing a swipe operation on a target area of the terminal, and the swipe operation is a valid operation; fingertip recognition is performed on the fingertip information to obtain a swipe area; text recognition is performed on the swipe area to obtain a first text in the swipe area; confirmation information is received from the terminal, which is used to confirm the correction of the first text in the swipe area; rule-based error correction and / or text-based error correction are performed on the first text in the swipe area to obtain a second text; the second text is sent to the terminal, which uses the second text to replace the first text in the swipe area. In this way, when a user swipes on a target area of the terminal using their fingertip, the recognition gateway can immediately identify the user's error correction intention based on the received fingertip information. It then performs fingertip recognition and text recognition to obtain the swipe area and the erroneous text within it. If the user determines that the erroneous text in the swipe area needs correction, the recognition gateway will perform secondary correction and send the correct text to the terminal so that the terminal can directly replace the erroneous text with the correct text. This solution not only directly corrects text subtitles based on the user's fingertip operations but also replaces subtitles while improving the accuracy and reliability of the correction. This achieves accurate recognition and efficient processing of user operations, reduces the burden of repetitive operations for users, provides a more convenient and intuitive interaction method, and improves user experience and system performance. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application; Figure 2This is a flowchart illustrating another data processing method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the rapid feedback method for live subtitle errors provided in this application embodiment; Figure 4 This is a schematic diagram of the fingertip recognition process provided in an embodiment of this application; Figure 5 This is a schematic diagram of the process of model self-optimization and updating after enabling the model fine-tuning function according to the embodiments of this application; Figure 6A This is a schematic diagram of the interface where the selected area is highlighted after selecting an incorrect subtitle with a fingertip, as provided in an embodiment of this application. Figure 6B This is a schematic diagram of the interface where a confirmation button pops up above the selected area after the fingertip is released, as provided in an embodiment of this application. Figure 6C This is a schematic diagram of the interface where the error message is displayed as the correct message after the submission confirmation button is pressed, as provided in this application embodiment. Figure 7 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application; Figure 8 This is a schematic diagram of another data processing device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the processing device provided in the embodiments of this application. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0012] In the description of the embodiments of this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, in the embodiments of this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0013] In the description of the embodiments of this application, the term "correspondence" may indicate that there is a direct or indirect correspondence between two things, or that there is an association between two things, or that there is a relationship of instruction and being instructed, configuration and being configured, etc.
[0014] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.
[0015] With the rapid development of live streaming technology, live streaming has become an important form of entertainment and information acquisition in people's daily lives. During live streaming, subtitles, as a crucial means of information transmission, have a significant impact on the overall effect. However, relevant live streaming subtitle systems often suffer from typos and omissions, and viewers find it difficult to provide quick and effective feedback, leading to a decline in the live streaming experience and a poor viewing experience. Specifically, these problems include the following: (1) Cumbersome operation and low accuracy. Currently, when viewers find typos in the live broadcast subtitles, they usually need to manually find the feedback entry, fill out a detailed feedback form, and may even need to take screenshots and attach text descriptions. When manually filling out the feedback form, viewers may provide inaccurate feedback due to unclear expression or omission of key information. This operation process is relatively cumbersome, which not only interrupts the viewing experience but may also cause viewers to give up on providing feedback due to the complexity of the operation. At the same time, the staff cannot obtain accurate information and cannot efficiently locate the problem or solve it quickly.
[0016] (2) Low feedback efficiency leads to low audience participation. Due to the complexity of the feedback process, viewers often have to wait a long time to receive a response after submitting feedback. This delay not only affects viewers' participation and satisfaction with the live stream, but may also prevent typos from being corrected in a timely manner, further affecting the quality of the live stream. At the same time, due to the complexity and inefficiency of the feedback process, many viewers may choose not to participate in the feedback, which reduces viewers' participation and interactivity with the live stream, and also affects the user stickiness of the live stream platform.
[0017] (3) Lack of real-time capability. The feedback methods for related issues often fail to provide real-time feedback and correction. Even if viewers submit feedback, the subtitle issues may not be resolved in a timely manner during the live broadcast due to lengthy processing procedures or insufficient personnel.
[0018] To address the aforementioned technical issues, this application proposes a rapid error feedback method. This method achieves timely detection and correction of errors in live-stream subtitles through real-time feedback from the user and rapid processing in the backend. By simplifying the feedback process, improving feedback efficiency, ensuring feedback accuracy, achieving real-time feedback and correction, and enhancing user engagement, it effectively solves the problems existing in related technologies, improving the accuracy of live-stream subtitles and the user's viewing experience. Specifically, it achieves rapid resolution of live-stream subtitle errors by organically combining three core stages: user fingertip recognition of subtitle errors, accidental touch detection, model self-learning, and secondary error correction technology. This results in more accurate subtitles being displayed to the user, and also enables accurate recognition and efficient processing of user operations, improving user experience and system performance.
[0019] To facilitate understanding of the technical solutions of the embodiments of this application, the technical solutions of this application are described in detail below through specific embodiments. The above-mentioned related technologies are optional solutions and can be arbitrarily combined with the technical solutions of the embodiments of this application, all of which fall within the protection scope of the embodiments of this application. The embodiments of this application include at least some of the following contents.
[0020] This application proposes a data processing method applied to an identification gateway. Specifically, the identification gateway is a gateway device capable of calling other functional engines such as a fingertip recognition engine, an optical character recognition (OCR) engine, a rule-based error correction engine, a text error correction engine, and a corpus annotation engine. Figure 1 This is a flowchart illustrating the data processing method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps: Step 101: Receive the fingertip information sent by the terminal.
[0021] In this embodiment, the user first performs a fingertip swipe operation on the target area of the terminal. After recognizing the user's fingertip swipe operation, the terminal immediately determines whether the operation is a mis-touch. If the operation is determined to be valid, the terminal sends the fingertip information generated during the user's swipe operation to the backend, which then forwards the fingertip information to the identification gateway. The identification gateway then receives the fingertip information sent by the terminal from the backend. The fingertip information refers to the information generated by the user's swipe operation on the target area of the terminal, including but not limited to the starting point coordinates, ending point coordinates, swipe trajectory, swipe speed, and time. The swipe operation is a valid operation and is not a mis-touch.
[0022] Here, the terminal can also send the screen image and / or audio data corresponding to the user's swipe operation to the recognition gateway via the backend, so as to improve the recognition gateway's accuracy in recognizing erroneous text.
[0023] It's important to note that the data generated by a user's fingertip swipe operation varies depending on the target area of the terminal. For example, if the target area is a live stream with subtitles, a fingertip swipe on the subtitles will generate a corresponding screen image, audio information, and fingertip data. Conversely, if the target area is a reading scenario with subtitles, a fingertip swipe on the subtitles will generate a corresponding screen image and fingertip data.
[0024] Step 102: Perform fingertip recognition on the fingertip information to obtain the sliding area.
[0025] In this embodiment, after receiving the user's fingertip information, the recognition gateway invokes the fingertip recognition engine to identify the fingertip swiping trajectory and fingertip position information included in the fingertip information, in order to extract the swiping area corresponding to the user's fingertip swiping operation. The fingertip information includes the fingertip swiping trajectory, and the fingertip swiping trajectory includes the fingertip position information. Specifically, the boundary of the swiping area is determined by the starting and ending points of the fingertip, and the swiping area is typically represented as a rectangle. For example, if the user swipes from the upper left corner of the screen to the lower right corner, a rectangular area will be defined based on the swiping path; this rectangular area represents the location of the subtitle that the user wishes to correct.
[0026] Here, the fingertip recognition engine is a biometric recognition method based on computer vision and machine learning technology. It determines the range of motion and location information of the fingertip by analyzing the unique features of the fingertip (such as fingerprint patterns, fingertip shape, etc.).
[0027] In some implementations, step 102 specifically includes: The fingertip recognition engine is invoked to identify the fingertip position information, and the coordinates of the fingertip start point and the fingertip end point are obtained; The sliding area is determined based on the coordinates of the fingertip's starting point and ending point.
[0028] Here, the fingertip swipe trajectory includes image information of the fingertip swiping on the target area. This image information records the position information (fingert position information) during the fingertip swipe process. The recognition gateway calls the fingertip recognition engine to dynamically identify the fingertip position information included in the fingertip swipe trajectory. During this recognition process, the fingertip recognition engine identifies the fingertip in the image information in real time through edge detection and other methods, and extracts its edge contour. Then, it calculates the convexity and concavity on the edge contour points, including the start point, end point, and valley point of the concavity, and calculates the curvature of each point on the edge contour. Contour points with curvature greater than the curvature threshold are identified as fingertip points, while points corresponding to noisy convex concavities with valley depth less than the depth threshold are removed. Then, taking the upper left corner of the image information as the origin and the length and width of the image information as the reference, it records the coordinate information corresponding to different positions of the fingertip point during the swipe process to determine the start point coordinate information and end point coordinate information during the fingertip swipe process. Thus, the swipe area corresponding to the fingertip swipe operation can be determined based on the fingertip start point coordinate and fingertip end point coordinate.
[0029] In some implementations, after step 102, the following may also be included: If the text in the sliding area is not fully displayed, the sliding area is expanded by a preset pixel value in each of the four directions to obtain the expanded sliding area; or, Based on the size and position information of a portion of the text within the sliding area, the sliding area is expanded to obtain an expanded sliding area; or... A prompt message is sent to the terminal, prompting the user to perform a swipe operation again on the target area to determine the new swipe area.
[0030] Here, after obtaining the fingertip swiping area through fingertip recognition, if the text information in the swiping area is not displayed completely, resulting in the swiping area being insufficient to include the entire text information, the swiping area can be adjusted through strategies such as fixed expansion, dynamic expansion, user feedback, or intelligent learning. Specifically, the sliding area can be expanded by a preset pixel value in each of the four directions: up, down, left, and right, so that the expanded sliding area can include complete text information. Alternatively, the sliding area can be dynamically expanded based on the size and position of the text detected in the sliding area, so that the expanded sliding area can include complete text information. For example, if the text detected in the sliding area is located at the bottom of the sliding area and appears to be truncated, the preset pixel value can be expanded downwards in the sliding area until the text in the sliding area is fully displayed. Alternatively, a prompt message can be sent to the terminal to prompt the user to perform a fingertip swipe operation again on the target area to determine a new sliding area that can include complete text information.
[0031] Optionally, the recognition gateway can also have a built-in preset intelligent model. This model learns the user's swiping habits through machine learning and big data analysis, and automatically adjusts the swiping area that does not fully include the text information. For example, when a user swipes their fingertip across a target area, the recognition gateway will call the model to determine in real time whether the text information in the swiping area is fully displayed. If not, the model will automatically adjust the swiping area so that the adjusted swiping area can fully include the text information.
[0032] Step 103: Perform text recognition on the sliding area to obtain the first text in the sliding area.
[0033] In this embodiment, after obtaining the sliding area, the recognition gateway calls the text recognition engine to perform text recognition on the text portion of the sliding area, extracting the editable text content to obtain the text information in the sliding area, i.e., the first text. The first text refers to the erroneous text in the sliding area.
[0034] Here, the text recognition engine is used to convert text in images or scanned documents into editable and searchable digital text for subsequent text correction. Text recognition analyzes pixel patterns in an image and converts them into structured text information. This process may involve steps such as image preprocessing (e.g., noise reduction, grayscale conversion, binarization), character segmentation, feature extraction, and matching.
[0035] In some implementations, step 103 specifically includes: The sliding area is processed to obtain the text area within the sliding area; The text recognition engine is invoked to perform text recognition on the text region, and the first text in the text region is obtained.
[0036] Here, the recognition gateway can first perform preprocessing such as noise reduction and cropping on the sliding area to extract the text area in the sliding area. Then, it calls the text recognition engine to perform text recognition on the text area, extract the text information in the text area, and convert it into editable text information, thereby obtaining the first text in the text area.
[0037] Step 104: Receive confirmation information sent by the terminal.
[0038] The confirmation information is used to confirm the correction of the first text in the sliding area.
[0039] In this embodiment, when a user performs a fingertip swipe operation on a target area of the terminal, a selected swipe area is obtained. If the user's fingertip leaves the target area, a confirmation button will pop up on the side of the swipe area. When the user clicks the confirmation button, the recognition gateway will receive confirmation information sent by the terminal to confirm the error correction processing of the first text in the swipe area.
[0040] Here, when a user swipes their fingertip across a target area on the device, the selected area is highlighted to distinguish it from other unselected areas.
[0041] In practice, receiving confirmation information marks the formal start of the user feedback process. The identification gateway needs to ensure it can accurately capture the user's intent and initiate the corresponding error correction mechanism based on that intent. Furthermore, to prevent accidental operations or malicious interference, the identification gateway can be configured with secondary verification mechanisms, such as pop-up prompts or CAPTCHA verification, thereby ensuring the security and reliability of the error correction process.
[0042] Step 105: Perform rule-based error correction and / or text-based error correction on the first text in the sliding area to obtain the second text.
[0043] In this embodiment, after receiving the confirmation information sent by the terminal, the identification gateway will call the rule correction engine to perform rule correction on the first text in the sliding area, and / or call the text correction engine to perform text correction on the first text in the sliding area, so as to obtain the second text. The second text refers to the correct text corresponding to the first text, that is, the correct text corresponding to the incorrect text.
[0044] Rule-based error correction is a method that uses a knowledge base based on preset grammar rules, spelling rules, and hot word graphs. It corrects common errors in the first text by comparing it to known correct expressions. For example, if the first text contains "login" instead of "login," rule-based error correction will replace it with the correct word according to preset rules.
[0045] Text correction is a process of intelligently analyzing and correcting text based on deep learning models. Compared to rule-based correction, text correction can handle more complex language structures and context-dependent issues, and has higher accuracy and generalization ability. For example, if the first text contains grammatical errors or semantic inconsistencies, the text correction engine can generate more natural and context-appropriate text through contextual reasoning and semantic understanding.
[0046] In some implementations, step 105 specifically includes: The rule-based error correction engine is invoked to perform spelling and grammar correction on the first text in the sliding area, resulting in the second text; or... The rule-based error correction engine is invoked to perform spelling and grammar correction on the first text in the sliding area to obtain the middle text; the text error correction engine is invoked to perform intelligent text correction on the middle text to obtain the second text.
[0047] Here, the recognition gateway first calls the rule-based error correction engine to perform preliminary error correction on the first text, correcting some common spelling and grammar errors. If the rule-based error correction can completely solve the error problems of the first text, the second text is obtained directly; otherwise, a more complex error correction process needs to be performed. In this case, the first text is corrected by rule-based error correction to obtain intermediate text. Then, the recognition gateway will continue to call the text error correction engine, using the intermediate text as input to the text error correction engine. The text error correction model in the text error correction engine will further perform intelligent text error correction on the intermediate text based on deep learning algorithms to correct more complex text errors in the intermediate text, and finally obtain the second text.
[0048] In the rule-based error correction process, the types of spelling and grammar correction include forced correction, full spelling correction, and hot word correction. Furthermore, forced correction has a higher priority than full spelling correction, which in turn has a higher priority than hot word correction.
[0049] In some implementations, a rule-based error correction engine is invoked to perform spelling and grammar correction on the first text in the sliding region to obtain intermediate text, including: Based on preset error correction rules, the first text in the sliding area is forcibly replaced or corrected to obtain the intermediate text; and / or, The correct text is determined based on the full spelling information of the first text in the sliding area, and the intermediate text is obtained by replacing or correcting the first text in the sliding area based on the correct text; and / or, The first text in the sliding area is replaced or modified based on a preset hot word map to obtain the middle text.
[0050] Predefined error correction rules refer to a pre-defined set of language rules used to identify and correct common spelling errors, grammatical errors, and formatting issues. These rules can include lexical substitutions (e.g., replacing "teh" with "the"), grammatical adjustments (e.g., subject-verb agreement, tense consistency), and punctuation corrections. These rules can be obtained by training natural language processing techniques using a large corpus.
[0051] A pre-defined hot word graph is a dynamically updated network structure of popular words that records currently popular or frequently used words as well as their variations. For example, if the metaverse becomes a hot topic during a live broadcast, then the metaverse will be added to the hot word graph and will be considered a higher-priority candidate word.
[0052] Here, in the rule-based error correction process, the rule-based error correction engine first performs forced error correction on the first text. This means that based on preset error correction rules, it directly replaces or corrects errors in the first text without user intervention. For example, when "aple" is detected, the system immediately replaces "aple" with "apple". If forced error correction resolves spelling and grammatical errors in the first text, an intermediate text is obtained. Thus, through the automatic application of preset error correction rules, the initial error correction operation is completed in the shortest possible time, improving overall error correction efficiency. Otherwise, it enters the full-spelling error correction stage. This involves determining the corresponding correct text based on the full-spelling information of the first text, including converting the first text into pinyin form and then using a pinyin matching algorithm to find the correct text that is closest to that pinyin form. The system replaces or corrects errors in the first text based on the correct text. If the full-text error correction can resolve spelling and grammatical errors in the first text, an intermediate text is obtained, which can then use pinyin information for precise error correction, further improving the error correction capability. Otherwise, it enters the hot word error correction stage, which replaces or corrects errors in the first text based on a preset hot word map. This includes checking whether there are related correct expressions in the hot word map when errors are found in the first text, and replacing or correcting the errors in the first text based on the relevant correct expressions to obtain the intermediate text. This allows for optimization of high-frequency words in specific scenarios through the hot word map, improving the timeliness and accuracy of subtitles and helping to reflect the latest trends and hot topics in a timely manner.
[0053] Similarly, in some implementations, a rule-based error correction engine is invoked to perform spelling and grammar correction on the first text in the sliding area to obtain the second text, including: Based on preset error correction rules, the first text in the sliding area is forcibly replaced or corrected to obtain the second text; and / or, The correct text is determined based on the full spelling information of the first text in the sliding area, and the second text is obtained by replacing or correcting the first text in the sliding area based on the correct text; and / or, The first text in the sliding area is replaced or modified based on a preset hot word map to obtain the second text.
[0054] The specific implementation process of the second text can be referred to the specific implementation process of the intermediate text mentioned above, and will not be elaborated further here.
[0055] In some implementations, if there are multiple first texts in the sliding area, the text length of each first text is determined, and the multiple first texts are prioritized in descending order of text length; based on the priority settings, the multiple first texts are sequentially subjected to spelling and grammar correction to obtain multiple second texts or multiple intermediate texts.
[0056] Here, during the rule-based error correction process, the rule-based error correction engine will prioritize long words. That is, when there are multiple first texts in the sliding area, it means that there are multiple possible error correction options. Then, according to the text length of each first text, the priority of each first text can be set from high to low in descending order of text length. Then, according to the priority settings, the first text with the longest text length is selected as the error correction target, and spelling and grammar correction is performed on the first text. This process continues until all first texts have completed the error correction process, resulting in multiple second texts or multiple intermediate texts.
[0057] It should be noted that if rule-based error correction can completely resolve the errors in multiple first texts, then multiple second texts are obtained directly; otherwise, the text correction engine needs to be called to further perform intelligent text correction on the multiple first texts to correct more complex errors. In this case, rule-based error correction on the multiple first texts will result in multiple intermediate texts.
[0058] In some implementations, a text correction engine is invoked to perform intelligent text correction on the intermediate text to obtain the second text, including: The text correction model in the text correction engine is invoked to perform intelligent text correction on the intermediate text, resulting in the second text.
[0059] Here, the text correction engine has a built-in text correction model. After the recognition gateway calls the rule-based error correction engine to perform spelling and grammar correction on the first text in the sliding area to obtain the intermediate text, the rule-based error correction engine will return the intermediate text to the recognition gateway. Then, the recognition gateway will use the intermediate text as input to the text correction engine and call the text correction model in the text correction engine to perform more complex intelligent text correction on the intermediate text based on deep learning algorithms, thereby obtaining the second text.
[0060] The text correction model can be any type of artificial intelligence model with self-learning capabilities, and there are no restrictions on it here.
[0061] In some implementations, before invoking the text correction model in the text correction engine to perform intelligent text correction on the intermediate text, the following may also be included: The corpus annotation engine is invoked to perform word segmentation on the first text in the sliding region, resulting in multiple lexical units. Based on the words in the error lexicon, the correct text corpus containing multiple lexical units is determined. It is then determined whether the proportion of multiple lexical units in the error segmentation lexicon exceeds a preset proportion. If so, the self-learning process of the text correction model is initiated, resulting in the trained text correction model. Here, before calling the text correction model in the text correction engine to perform intelligent text correction on the intermediate text, the text correction engine can first receive the updated text correction model sent by the corpus annotation engine and replace the original model with the updated text correction model to ensure that subsequent text recognition and correction operations can be based on the latest and more accurate model.
[0062] The text correction model update process includes: the recognition gateway sends the first text in the sliding area to the corpus annotation engine, which then performs Natural Language Processing (NLP) segmentation on the first text to obtain multiple lexical units; the recognition gateway then uses Artificial Intelligence Generated Content (AIGC) technology to generate text content that conforms to grammatical rules and semantic logic for each lexical unit, thus obtaining correct text corpus containing multiple lexical units, which serves as an important resource for the text correction model's self-learning optimization; when the number of multiple lexical units exceeds a preset proportion (e.g., 20%) in the incorrectly segmented lexicon, the text correction model's self-learning optimization process is automatically initiated. During this process, the text correction model trains itself based on the correct text corpus containing multiple lexical units to correct biases in recognizing and processing incorrect words, resulting in a final trained text correction model. In this self-learning process, the text correction model continuously optimizes its performance, improving the accuracy and efficiency of error correction.
[0063] The preset ratio can be adjusted according to actual needs to ensure that the text correction model can optimize itself after accumulating enough error data, while avoiding excessively frequent fine-tuning that could affect system performance. Therefore, no restrictions are imposed on it here.
[0064] Step 106: Send the second text to the terminal.
[0065] The terminal is used to replace the first text in the sliding area with the second text.
[0066] In this embodiment of the application, after the identification gateway performs rule-based error correction and / or text-based error correction on the first text in the sliding area to obtain the second text, it returns the second text to the terminal. Thus, the terminal can directly replace the first text in the sliding area with the second text to update the subtitle content and display it to the user.
[0067] In practical implementation, it is necessary to ensure the stability and timeliness of data transmission to avoid a decline in user experience due to network latency or excessive processing time. Furthermore, the terminal side needs to have sufficient rendering capabilities to support the updating and display of dynamic subtitles. For example, during live streaming, the subtitle content should be able to refresh instantly without affecting the smoothness of video playback.
[0068] In the technical solution of this application embodiment, the identification gateway receives fingertip information sent by the terminal; the fingertip information is information generated by the user performing a swipe operation on a target area of the terminal, and the swipe operation is a valid operation; fingertip recognition is performed on the fingertip information to obtain a swipe area; text recognition is performed on the swipe area to obtain a first text in the swipe area; confirmation information is received from the terminal, which is used to confirm the correction of the first text in the swipe area; rule-based error correction and / or text-based error correction are performed on the first text in the swipe area to obtain a second text; the second text is sent to the terminal, which uses the second text to replace the first text in the swipe area. In this way, when a user swipes on a target area of the terminal using their fingertip, the recognition gateway can immediately identify the user's error correction intention based on the received fingertip information. It then performs fingertip recognition and text recognition to obtain the swipe area and the erroneous text within it. If the user determines that the erroneous text in the swipe area needs correction, the recognition gateway will perform secondary correction and send the correct text to the terminal so that the terminal can directly replace the erroneous text with the correct text. This solution not only directly corrects text subtitles based on the user's fingertip operations but also replaces subtitles while improving the accuracy and reliability of the correction. This achieves accurate recognition and efficient processing of user operations, reduces the burden of repetitive operations for users, provides a more convenient and intuitive interaction method, and improves user experience and system performance. This application also proposes a data processing method, which is applied to a terminal. Specifically, the terminal is a terminal device with other functions such as image recognition and fingertip recognition, such as a mobile phone or tablet. Figure 2 This is a flowchart illustrating the data processing method provided in an embodiment of this application, as shown below. Figure 2 As shown, the method includes the following steps: Step 201: In response to the user's swipe operation on the target area, determine whether the swipe operation is a valid operation.
[0069] Step 202: If the swipe operation is determined to be a valid operation, determine the fingertip information corresponding to the swipe operation and send the fingertip information to the recognition gateway.
[0070] In this embodiment, the user first performs a fingertip swipe operation on a target area of the terminal. The terminal responds to the user's swipe operation and analyzes the user's operation in real time using a preset algorithm to determine whether the operation is a valid touch or a mis-touch, thus distinguishing between valid and mis-touch operations. If the operation is determined to be valid, the terminal sends the fingertip information generated during the user's swipe operation to the backend, which then forwards the fingertip information to the recognition gateway. The recognition gateway then performs fingertip recognition on the obtained swipe area and performs text recognition. The key to the above-mentioned mis-touch detection mechanism lies in the accuracy and response speed of the algorithm. Only by accurately identifying whether it is a mis-touch, filtering out invalid operations such as mis-touches, and ensuring that only the user's true error correction intention is recognized can unnecessary errors be avoided, ensuring user experience and system performance.
[0071] Here, the terminal can also send the screen image and / or audio data corresponding to the user's swipe operation to the recognition gateway via the backend, so as to improve the recognition gateway's accuracy in recognizing erroneous text.
[0072] In some implementations, determining whether a sliding operation is a valid operation includes: Determine whether the sliding distance corresponding to the sliding operation reaches a distance threshold, and / or determine whether the sliding speed corresponding to the sliding operation reaches a speed threshold. If the sliding distance corresponding to the sliding operation reaches the distance threshold, and / or the sliding speed corresponding to the sliding operation reaches the speed threshold, then the sliding operation is determined to be a valid operation; or... Collect image data corresponding to the swipe operation, and perform feature recognition on the image data. If fingertip features are identified, the swipe operation is determined to be valid; or... Record the sliding time corresponding to the sliding operation, and determine whether the time interval between the sliding time and the sliding time corresponding to the previous sliding operation is greater than or equal to a time threshold, and / or determine whether the similarity between the sliding trajectory corresponding to the sliding operation and the sliding trajectory corresponding to the previous sliding operation is greater than or equal to a similarity threshold. If the time interval between the sliding time and the sliding time corresponding to the previous sliding operation is greater than or equal to the time threshold, and / or the similarity between the sliding trajectory corresponding to the sliding operation and the sliding trajectory corresponding to the previous sliding operation is greater than or equal to the similarity threshold, then the sliding operation is determined to be a valid operation.
[0073] Here, we're referring to a Software Development Kit (SDK) integrated into the device that provides fingertip swipe functionality. An SDK is a set of tools provided by the manufacturer of a hardware platform, operating system, or programming language, designed to help software developers create, test, and deploy software applications more efficiently. An SDK typically includes programming tools, code examples, technical documentation, debugging and testing tools, and, where necessary, libraries and frameworks for specific programming languages or platforms. An SDK allows developers to access mobile phone hardware (such as cameras and GPS), handle touch input, and utilize other features of the operating system.
[0074] Specifically, the SDK integrated into the terminal has the ability to trigger fingertip recognition by swiping text with fingertips, and also has a built-in anti-mistouch mechanism. When a user swipes text on the terminal screen with their fingertip, the SDK first determines whether the action is a small movement to rule out the possibility of accidental touch. The following methods can be used to eliminate the possibility of accidental touch in the implementation of the anti-mistouch function: (1) Set distance threshold and / or speed threshold. After receiving the user's fingertip swipe action, the SDK determines whether the swipe distance corresponding to the swipe operation reaches the distance threshold and / or whether the swipe speed corresponding to the swipe operation reaches the speed threshold. If the swipe distance corresponding to the swipe operation reaches the distance threshold and / or the swipe speed corresponding to the swipe operation reaches the speed threshold, the swipe operation is determined to be a valid operation. Otherwise, the swipe operation is determined to be a false touch and the operation is ignored.
[0075] (2) After receiving the user's fingertip swipe operation, the SDK collects the image data corresponding to the user's fingertip swipe operation in real time, and uses image processing technology to perform feature recognition on the image data to identify the features of the user's fingertip, such as shape and size. If the fingertip feature is identified, the swipe operation is determined to be a valid operation. If non-fingertip features are not identified or the fingertip features are not obvious, the swipe operation is determined to be a false touch and the operation is ignored.
[0076] (3) Set the time interval between consecutive swipe operations. After receiving the user's fingertip swipe operation, the SDK will record the swipe time corresponding to the swipe operation. If the time interval between the swipe time corresponding to the swipe operation and the swipe time corresponding to the previous swipe operation is greater than or equal to the time threshold, and / or the similarity between the swipe trajectory corresponding to the swipe operation and the swipe trajectory corresponding to the previous swipe operation is greater than or equal to the similarity threshold, then the swipe operation is determined to be a valid operation; otherwise, if the time interval between two consecutive swipe operations (the swipe operation and the previous swipe operation) is less than the time threshold, and / or the similarity between the swipe trajectories corresponding to two consecutive swipe operations is less than the similarity threshold, then the swipe operation is determined to be a false touch and the operation is ignored.
[0077] It's important to note that if the SDK determines the user's swipe gesture is not accidental but still doesn't meet expectations, a self-learning startup phase will be triggered. In this phase, the SDK automatically adjusts its internal parameters and algorithms based on user habits and historical data to adapt to the personalized needs of different users. The core of self-learning startup lies in the system's self-optimization capability; through continuous learning and adjustment, the system can gradually improve the accuracy and efficiency of operations. Finally, if problems persist after self-learning adjustments, the SDK will enter a secondary error correction phase. In this phase, the system uses a higher-level error correction algorithm to re-analyze and correct the user's actions. The key to secondary error correction is its powerful error correction capability, which can effectively handle various complex situations and ensure the accuracy and stability of user operations.
[0078] Step 203: In response to the user's release action on the target area, highlight the sliding area.
[0079] Step 204: In response to the user's confirmation operation on the sliding area, send confirmation information to the identification gateway.
[0080] In this embodiment, when a user slides their fingertip across a target area on the terminal, a selected sliding area is obtained. If the user's fingertip leaves the target area, the terminal will respond to the user's release operation on the target area by highlighting the sliding area to distinguish it from other unselected areas, and a confirmation button will pop up on the side of the sliding area. When the user clicks the confirmation button, the terminal will respond to the user's confirmation operation on the target area by sending confirmation information to the identification gateway to confirm the error correction processing of the first text in the sliding area. Step 205: Receive the second text sent by the identification gateway and replace the first text in the sliding area with the second text.
[0081] In this embodiment of the application, after the identification gateway performs rule-based error correction and / or text-based error correction on the first text in the sliding area to obtain the second text, it will send the second text to the terminal. The terminal will then receive the second text sent by the identification gateway and replace the first text in the sliding area with the second text.
[0082] In the technical solution of this application embodiment, the terminal responds to the user's swipe operation on the target area and determines whether the swipe operation is valid; if the swipe operation is determined to be valid, the terminal determines the fingertip information corresponding to the swipe operation and sends the fingertip information to the recognition gateway, which performs fingertip recognition on the fingertip information and performs text recognition on the obtained swipe area; in response to the user's release operation on the target area, the terminal highlights the swipe area; in response to the user's confirmation operation on the swipe area, the terminal sends confirmation information to the recognition gateway, which confirms the correction of the first text in the swipe area; the terminal receives the second text sent by the recognition gateway and replaces the first text in the swipe area with the second text; the second text is obtained by the recognition gateway after performing rule correction and / or text correction on the first text in the swipe area. Thus, when a user performs a fingertip swipe operation on the target area, the terminal immediately activates the accidental touch detection mechanism to accurately and quickly determine whether the user's fingertip swipe operation is an accidental touch. If it is not an accidental touch, the terminal sends the fingertip information corresponding to the fingertip swipe operation to the recognition gateway. The recognition gateway identifies subtitle errors in the target area, and then uses secondary error correction technology on the erroneous subtitles to effectively deal with complex situations in erroneous subtitles. In the above process, the user only needs to perform an error correction action once, and the correction will be automatically applied to subsequent subtitles without the user having to repeat the operation. This one-time error correction and automatic application function can not only quickly resolve erroneous subtitles in the target area and display more accurate subtitles to the user, but also improve the user experience and system performance.
[0083] This application also proposes a method for rapid feedback of typos in live streaming subtitles. Figure 3 This is a flowchart illustrating the rapid feedback method for live subtitle errors provided in this application embodiment, as shown below. Figure 3 As shown, the method includes the following steps: Step 301: The user watches the live stream on the video client and triggers fingertip recognition by swiping the subtitles on the screen.
[0084] Step 302: The video client performs accidental touch detection on the user's fingertip operation to determine whether it is an accidental touch. If it is, the action is ignored; otherwise, proceed to step 303.
[0085] An SDK integrating fingertip swipe functionality is provided in the video client. This SDK has the ability to trigger fingertip recognition when swiping subtitles and also includes a built-in anti-mistouch mechanism. When a user swipes subtitles on the video client with their fingertip, the SDK first determines whether the action is valid, eliminating the possibility of accidental touches. Specifically, the following strategies are used to prevent accidental touches: (1) Set swipe threshold. After receiving a fingertip swipe action, the SDK first determines whether the swipe distance and speed exceed the preset threshold. Only when the swipe distance and speed reach or exceed the threshold is it considered a valid operation; otherwise, it is considered a mis-touch and the action is ignored.
[0086] (2) Recognize fingertip features. The SDK uses image processing technology to recognize fingertip features, such as shape and size. When non-fingertip features are recognized or the features are not obvious, the system judges them as accidental touches and ignores the corresponding operation.
[0087] (3) Time interval judgment. The SDK records the time of each fingertip swipe. When the time interval between two consecutive swipes is too short and the swipe trajectories are similar, the system considers this to be a series of accidental touches and ignores the corresponding operation.
[0088] Step 303: The video client sends the screen image, audio information, and fingertip information of the user's fingertip operation to the video backend system.
[0089] After determining that the user's fingertip operation is a valid operation, the SDK transmits the screen image, audio information, and fingertip information at the time of the operation to the video backend system. This information includes a screenshot of the user sliding the subtitles, the audio data at that time, and the sliding trajectory and position information of the fingertip.
[0090] Step 304: The video backend system sends the screen image, audio information, and fingertip information of the user's fingertip operation to the recognition gateway.
[0091] The video backend system transmits screen images, audio information, and fingertip data during operation to the recognition gateway. The recognition gateway, acting as the information processing hub, is responsible for receiving and forwarding this information.
[0092] Step 305: The recognition gateway calls the fingertip recognition engine to perform fingertip recognition on the received fingertip information, and obtains the area of fingertip sliding and the coordinate information of fingertip sliding.
[0093] Step 306: The fingertip recognition engine returns the area and coordinates of the fingertip swipe to the recognition gateway.
[0094] The recognition gateway invokes the fingertip recognition engine to process the received fingertip information. The fingertip recognition engine can identify the area and coordinates of the fingertip swipe and return this information to the recognition gateway.
[0095] like Figure 4The diagram illustrates the fingertip recognition process provided in this embodiment. In this diagram, the upper triangle represents the starting point of the fingertip swipe, the lower triangle represents the ending point of the fingertip swipe, and the entire rectangular area represents the boundary of the screen or recognition area. Assume the upper left corner of the screen or recognition area is the origin (0,0), and the lower right corner is (W,H), where W is the width and H is the height.
[0096] The starting point coordinates of a fingertip swipe could be (x1, y1), and the ending point coordinates could be (x2, y2). For example, if the screen size is 800x600 pixels, then the starting point coordinates of the fingertip swipe are: x1=200, y1=300, and the ending point coordinates are: x2=400, y2=450. In this way, a swipe area consisting of a starting point and an ending point is defined.
[0097] Step 307: The recognition gateway performs cropping on the area swiped by the fingertip to extract the text area to be recognized.
[0098] After obtaining the area where the fingertip swipes, OCR technology can be used to recognize the text within that area. OCR technology can analyze pixel patterns in an image and convert them into editable and searchable text.
[0099] During the recognition process, the recognition gateway further processes the area swiped by the fingertip. Through operations such as photo editing and noise reduction, it extracts the text area to be recognized and sends it to the OCR recognition engine for text recognition.
[0100] It should be noted that when the area swiped by the fingertip is insufficient to encompass all subtitle content (smaller than the subtitle area), several strategies can be used to expand the area of judgment: (1) Fixed expansion. When the subtitle content is found to be incomplete, a certain number of pixels can be expanded in each of the four directions (up, down, left, right) of the sliding area, and then OCR recognition can be performed again.
[0101] (2) Dynamic expansion. The direction and size of the expansion are dynamically adjusted based on the size and position of the identified partial text. For example, if the identified text is located at the bottom of the sliding area and appears to be truncated, more pixels can be expanded downwards from the sliding area.
[0102] (3) User feedback. In some cases, the sliding area can be manually adjusted based on user feedback. For example, if a possible caption truncation is detected, the recognition gateway can prompt the user to slide again to include more content.
[0103] (4) Intelligent learning. Through machine learning and big data analysis, the system learns the user's swiping habits and automatically adjusts the determination of the swiping area. This usually requires a large amount of user data and algorithm training to achieve.
[0104] By combining the above strategies, we can more effectively handle situations where the fingertip swipe area is insufficient to cover the complete subtitle content, and improve recognition accuracy and user experience.
[0105] Step 308: The recognition gateway calls the OCR recognition engine to perform OCR recognition on the recognized text area to obtain text information.
[0106] The recognition gateway calls the OCR recognition engine to perform text recognition on the extracted text area, converts the text in the text area into editable text information, and then returns the text information to the recognition gateway, which then passes this information to the required application or system.
[0107] Step 309: The user leaves the screen, a confirmation button pops up, and the user clicks the confirmation button to send confirmation information to the identification gateway.
[0108] When a user's fingertip leaves the screen, the area the user swiped will be highlighted, and a confirmation button will pop up above the swiped area. When the user clicks the confirmation button, a confirmation message will be sent to the recognition gateway to confirm the error correction processing of the text information in the swiped area.
[0109] Step 310: The recognition gateway sends the recognized text information to the corpus annotation engine.
[0110] Step 311: The corpus annotation engine saves the correct text corresponding to the text information and the audio information corresponding to the text information, and returns them to the recognition gateway.
[0111] The recognition gateway returns text information to the corpus annotation engine, which records the text information and corresponding audio information. Researchers then identify erroneous text within the corpus, correct it, and save the correct text and corresponding audio information as training data for the next training task. Finally, the corpus annotation engine feeds the saved information back to the gateway, indicating successful data saving. This training is expected to be used to optimize the performance of the OCR recognition engine and improve recognition accuracy.
[0112] Step 312: The identification gateway calls the rule correction engine to perform the first correction on the text information to obtain the intermediate text.
[0113] The gateway identifies and invokes a rule-based error correction engine to perform the first step of error correction on the text information. The rule-based error correction engine performs preliminary error correction on the identified text information based on preset rules, correcting some common spelling and grammatical errors. The error correction types include forced correction, full spelling correction, and hot word correction.
[0114] During the initial error correction process, the system first attempts forced error correction, directly replacing or correcting erroneous text according to preset correction rules. If forced error correction fails, the system enters the full-spelling error correction stage, attempting to find the correct words based on the full-spelling information of the text. Finally, if the first two steps still fail to resolve the issue, the system will further utilize the hot word correction function, correcting erroneous text by referencing the current hot word graph.
[0115] It should be noted that during the error correction process, the system follows a priority principle: first, it performs forced error correction; then, it performs full-spelling error correction; and finally, it performs hot word error correction. Furthermore, the system prioritizes longer words; when multiple possible correction options exist, the system will prioritize longer words as the correction target, as longer words generally have higher accuracy and reliability.
[0116] Step 313: The identification gateway calls the text correction engine to perform a second correction on the intermediate text to obtain the correct text.
[0117] After the first step of error correction is completed, the recognition gateway uses the obtained intermediate text as input to call the text correction engine to perform a second round of error correction on the intermediate text, finally obtaining the correct text, which is then returned to the recognition gateway. The text correction engine uses deep learning algorithms to perform more refined error correction processing on the intermediate text, enabling it to identify and correct more complex errors.
[0118] Step 314: The identification gateway returns the correct text to the video client.
[0119] The identification gateway returns the correct text to the video client, which then directly replaces the text information in the sliding area with the correct text and displays it to the user.
[0120] It should be noted that when the corpus annotation engine receives text information and identifies errors in the text information, it will use NLP technology to perform word segmentation processing on the text information, dividing it into independent word units, which helps to more accurately identify and understand the errors. The error segmentation lexicon will record all the identified error word units, and based on the proportion of error words, it will activate the model fine-tuning function to realize the self-optimization update of the model.
[0121] Figure 5 This is a schematic diagram illustrating the process of model self-optimization and updating after enabling the model fine-tuning function, as provided in the embodiments of this application. Figure 5As shown, the process includes the following methods: Step 501: Perform NLP segmentation on the text information to obtain multiple lexical units and record them in the error segmentation lexicon.
[0122] Step 502: Use AIGC technology to determine the correct text corpus containing each lexical unit.
[0123] The corpus annotation engine calls AIGC technology to generate correct text corpus containing each word unit. AIGC technology can generate text content that conforms to grammatical rules and semantic logic based on given words or context. These generated correct text corpora will serve as an important resource for the model's self-learning and optimization.
[0124] Step 503: Determine whether the proportion of multiple word units in the erroneous word segmentation lexicon exceeds 20%. If it does, start the model fine-tuning function and enable model self-learning.
[0125] As reasoning progresses over time, the system will continuously accumulate new erroneous words. To determine when to initiate model fine-tuning, it is necessary to define a threshold to determine whether newly added erroneous words exceed 20% of the accumulated vocabulary. This threshold can be adjusted according to actual needs to ensure that the model can self-optimize after accumulating sufficient erroneous data, while avoiding excessively frequent fine-tuning that could impact system performance.
[0126] When the number of newly added erroneous words (multiple word units) exceeds a set threshold (e.g., 20%) in the erroneous word segmentation corpus, the system will automatically activate the model fine-tuning function. During fine-tuning, the system will train the model using previously generated correct text corpora to correct the model's biases in recognizing and processing these erroneous words. Through this self-learning method, the model can gradually improve its accuracy and generalization ability.
[0127] Step 504: Send the updated model to the text correction engine.
[0128] After fine-tuning, the model's performance will be improved, enabling it to better identify and handle similar errors. The corpus annotation engine will send the updated model to the text correction engine, which will replace the original model with the updated one to ensure that subsequent caption recognition and correction operations are based on the latest and more accurate model.
[0129] like Figure 6A The image shown is a schematic diagram of an interface where the selected area is highlighted after the user selects an incorrect subtitle with their fingertip, according to an embodiment of this application. In this diagram, the user selects the incorrect subtitle with their fingertip, and the selected area corresponding to the incorrect subtitle is highlighted. Figure 6BThe diagram shown is an example of an interface where a confirmation button pops up above the selected area after the user releases their fingertip, according to an embodiment of this application. In this diagram, after the user selects an incorrect text and releases their fingertip, a confirmation button pops up above the selected area. Figure 6C The diagram shows an interface where the incorrect subtitles are displayed as correct subtitles after the user clicks the confirmation button in this application embodiment. In the diagram, after the user clicks the confirmation button, the incorrect subtitles in the selected area are automatically corrected to the correct subtitles: "Out of leg".
[0130] In terms of automatic subtitle error correction, the solution of this application can directly correct errors based on the user's fingertip operation, and only one correction operation is needed to automatically apply it to subsequent subtitles. This design greatly improves the efficiency and convenience of error correction, reduces the burden of repetitive operations for users, provides users with a more convenient and intuitive interaction method, meets users' needs for intelligent and personalized services, and helps to improve user experience and satisfaction.
[0131] The solution proposed in this application demonstrates excellent identification and judgment capabilities in identifying and judging erroneous operations. It can accurately distinguish between user fingertip operations and erroneous operations, thereby ensuring the accuracy of error correction. This is crucial for improving user experience and avoiding unnecessary interference.
[0132] The solution proposed in this application possesses strong self-learning capabilities, continuously optimizing the error correction algorithm and model based on user operating habits and historical data. This allows the system to gradually adapt to the personalized needs of different users, improving the targeting and accuracy of error correction.
[0133] The embodiments of this application provide a secondary error correction mechanism. When a primary error correction fails to completely resolve the issue, a secondary AI error correction can be triggered, further improving the accuracy and reliability of the correction. This design provides users with greater error correction assurance and reduces the risk of subtitle errors.
[0134] This application also proposes a data processing device, which is applied to an identification gateway. Specifically, the identification gateway is a gateway device that can call other functional engines such as fingertip recognition engine, OCR recognition engine, rule error correction engine, text error correction engine, and corpus annotation engine. Figure 7 This is a schematic diagram of the structure of the data processing device provided in the embodiments of this application, as shown below. Figure 7 As shown, the device includes: The first receiving unit 701 is used to receive fingertip information sent by the terminal; the fingertip information is information generated by the user performing a swipe operation on the target area of the terminal, and the swipe operation is a valid operation.
[0135] The first processing unit 702 is used to perform fingertip recognition on the fingertip information to obtain a sliding area; and to perform text recognition on the sliding area to obtain the first text in the sliding area.
[0136] The first receiving unit 701 is also used to receive confirmation information sent by the terminal, the confirmation information being used to confirm the error correction of the first text in the sliding area.
[0137] The first processing unit 702 is also used to perform rule-based error correction and / or text-based error correction on the first text in the sliding area to obtain the second text.
[0138] The first sending unit 703 is used to send second text to the terminal, and the terminal is used to replace the first text in the sliding area with the second text.
[0139] In some implementations, the fingertip information includes a fingertip gliding trajectory, which includes fingertip position information; wherein, The first processing unit 702 is specifically used to: call the fingertip recognition engine to recognize the fingertip position information and obtain the fingertip start coordinates and fingertip end coordinates; and determine the sliding area based on the fingertip start coordinates and fingertip end coordinates.
[0140] In some implementations, the first processing unit 702 is further specifically used to: call the rule-based error correction engine to perform spelling and grammar correction on the first text in the sliding area to obtain the second text; or, call the rule-based error correction engine to perform spelling and grammar correction on the first text in the sliding area to obtain intermediate text; and call the text error correction engine to perform intelligent text correction on the intermediate text to obtain the second text.
[0141] In some embodiments, the first processing unit 702 is further specifically used to: forcibly replace or correct the first text in the sliding area based on a preset error correction rule to obtain intermediate text; and / or, determine the correct text based on the full spelling information of the first text in the sliding area, and replace or correct the first text in the sliding area based on the correct text to obtain intermediate text; and / or, replace or correct the first text in the sliding area based on a preset hot word map to obtain intermediate text.
[0142] In some embodiments, the first processing unit 702 is further configured to: if the text in the sliding area is not fully displayed, expand the sliding area by a preset pixel value in each of the four directions to obtain an expanded sliding area; or, expand the sliding area based on the size and position information of some text in the sliding area to obtain an expanded sliding area; or, send a prompt message to the terminal, the prompt message being used to prompt the user to perform a sliding operation again on the target area to determine a new sliding area.
[0143] This application also proposes a data processing device, which is applied to a terminal. Specifically, the terminal is a terminal device with other functions such as image recognition and fingertip recognition, such as a mobile phone or tablet. Figure 8 This is a schematic diagram of the structure of the data processing device provided in the embodiments of this application, as shown below. Figure 8 As shown, the device includes: The second processing unit 801 is used to respond to the user's swipe operation on the target area, determine whether the swipe operation is a valid operation, and determine the fingertip information corresponding to the swipe operation if the swipe operation is determined to be a valid operation.
[0144] The second sending unit 802 is used to send fingertip information to the recognition gateway, which is used to perform fingertip recognition on the fingertip information and to perform text recognition on the obtained sliding area.
[0145] The second processing unit 801 is also used to highlight the sliding area in response to the user's release operation on the target area.
[0146] The second sending unit 802 is also used to send confirmation information to the identification gateway in response to the user's confirmation operation on the sliding area. The confirmation information is used to confirm the error correction of the first text in the sliding area.
[0147] The second receiving unit 803 is used to receive the second text sent by the identification gateway.
[0148] The second processing unit 801 is further configured to replace the first text in the sliding area with the second text; the second text is obtained by the recognition gateway after performing rule correction and / or text correction on the first text in the sliding area.
[0149] In some embodiments, the second processing unit 801 is specifically used for: Determine whether the sliding distance corresponding to the sliding operation reaches a distance threshold, and / or determine whether the sliding speed corresponding to the sliding operation reaches a speed threshold. If the sliding distance corresponding to the sliding operation reaches the distance threshold, and / or the sliding speed corresponding to the sliding operation reaches the speed threshold, then the sliding operation is determined to be a valid operation; or... Collect image data corresponding to the swipe operation, and perform feature recognition on the image data. If fingertip features are identified, the swipe operation is determined to be valid; or... Record the sliding time corresponding to the sliding operation, and determine whether the time interval between the sliding time and the sliding time corresponding to the previous sliding operation is greater than or equal to a time threshold, and / or determine whether the similarity between the sliding trajectory corresponding to the sliding operation and the sliding trajectory corresponding to the previous sliding operation is greater than or equal to a similarity threshold. If the time interval between the sliding time and the sliding time corresponding to the previous sliding operation is greater than or equal to the time threshold, and / or the similarity between the sliding trajectory corresponding to the sliding operation and the sliding trajectory corresponding to the previous sliding operation is greater than or equal to the similarity threshold, then the sliding operation is determined to be a valid operation.
[0150] Those skilled in the art should understand that Figure 7 , Figure 8 The functions of each unit in the data processing device shown can be understood by referring to the relevant description of the aforementioned method. Figure 7 , Figure 8 The functions of each unit in the data processing device shown can be implemented by a program running on the processor or by specific logic circuits.
[0151] Figure 9 This is a schematic diagram of the processing device provided in an embodiment of this application. The processing device may be a terminal device or a network device. Figure 9 The processing device shown includes a processor 901, which can call and run computer programs from memory to implement the methods in the embodiments of this application.
[0152] Optionally, such as Figure 9 As shown, the processing device may further include a memory 902. The processor 901 can retrieve and run computer programs from the memory 902 to implement the methods described in the embodiments of this application.
[0153] The memory 902 can be a separate device independent of the processor 901, or it can be integrated into the processor 901.
[0154] Optionally, such as Figure 9 As shown, the processing device may also include a transceiver 903, which the processor 901 can control to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices.
[0155] The transceiver 903 may include a transmitter and a receiver. The transceiver 903 may further include an antenna, which may be one or more.
[0156] The processing device may specifically be the data processing device of the embodiments of this application, and the processing device can implement the corresponding processes of the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.
[0157] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0158] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0159] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0160] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to the processing device in this application embodiment, and the computer program causes the computer to execute the corresponding processes implemented by the various methods in this application embodiment; for brevity, further details are omitted here.
[0161] This application also provides a computer program product, including computer program instructions. This computer program product can be applied to the processing device in this application embodiment, and the computer program instructions cause the computer to execute the corresponding processes implemented by the various methods in this application embodiment; for brevity, further details are omitted here.
[0162] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0163] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0164] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0165] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0166] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0167] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0168] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data processing method, characterized by, Applied to an identification gateway, the method includes: The receiving terminal sends fingertip information; the fingertip information is generated by the user performing a swipe operation on a target area of the terminal, and the swipe operation is a valid operation. Finger tip recognition is performed on the fingertip information to obtain a sliding area; text recognition is performed on the sliding area to obtain the first text in the sliding area; The terminal sends a confirmation message, which is used to confirm that the first text in the sliding area has been corrected. The first text in the sliding area is subjected to rule-based error correction and / or text error correction to obtain the second text; The second text is sent to the terminal, which is used to replace the first text in the sliding area with the second text.
2. The method of claim 1, wherein, The fingertip information includes a fingertip sliding trajectory, and the fingertip sliding trajectory includes fingertip position information; the step of performing fingertip recognition on the fingertip information to obtain the sliding area includes: The fingertip recognition engine is invoked to identify the fingertip position information, thereby obtaining the coordinates of the fingertip start point and the fingertip end point. The sliding area is determined based on the coordinates of the fingertip starting point and the coordinates of the fingertip ending point.
3. The method of claim 1, wherein, The step of performing rule-based error correction and / or text-based error correction on the first text in the sliding area to obtain the second text includes: The rule-based error correction engine is invoked to perform spelling and grammar correction on the first text in the sliding area to obtain the second text; or... The rule-based error correction engine is invoked to perform spelling and grammar correction on the first text in the sliding area to obtain the intermediate text; the text error correction engine is invoked to perform intelligent text correction on the intermediate text to obtain the second text.
4. The method of claim 3, wherein, The rule-based error correction engine performs spelling and grammar correction on the first text in the sliding area to obtain the intermediate text, including: The first text in the sliding area is forcibly replaced or corrected based on preset error correction rules to obtain the intermediate text; and / or, The correct text is determined based on the full spelling information of the first text in the sliding area, and the first text in the sliding area is replaced or corrected based on the correct text to obtain the intermediate text; and / or, The first text in the sliding area is replaced or modified based on a preset hot word map to obtain the intermediate text.
5. The method of claim 1, wherein, After performing fingertip recognition on the fingertip information to obtain the sliding area, the method further includes: If the text in the sliding area is not fully displayed, then the sliding area is expanded by a preset pixel value in each of the four directions to obtain an expanded sliding area; or, Based on the size and position information of a portion of the text within the sliding area, the sliding area is expanded to obtain an expanded sliding area; or... A prompt message is sent to the terminal, which prompts the user to perform a swipe operation again on the target area to determine a new swipe area.
6. A data processing method, characterized by, Applied to a terminal, the method includes: In response to a user's swipe gesture on a target area, determine whether the swipe gesture is a valid gesture; If the sliding operation is determined to be a valid operation, the fingertip information corresponding to the sliding operation is determined and the fingertip information is sent to the recognition gateway. The recognition gateway is used to perform fingertip recognition on the fingertip information and to perform text recognition on the obtained sliding area. In response to the user's release operation on the target area, the sliding area is highlighted; in response to the user's confirmation operation on the sliding area, confirmation information is sent to the recognition gateway, the confirmation information being used to confirm the correction of the first text in the sliding area; The system receives the second text sent by the identification gateway and replaces the first text in the sliding area with the second text; the second text is obtained by the identification gateway after performing rule correction and / or text correction on the first text in the sliding area.
7. The method of claim 6, wherein, The determination of whether the sliding operation is a valid operation includes: Determine whether the sliding distance corresponding to the sliding operation reaches a distance threshold, and / or determine whether the sliding speed corresponding to the sliding operation reaches a speed threshold. If the sliding distance corresponding to the sliding operation reaches the distance threshold, and / or the sliding speed corresponding to the sliding operation reaches the speed threshold, then determine that the sliding operation is a valid operation; or... The image data corresponding to the sliding operation is acquired, and feature recognition is performed on the image data. If fingertip features are identified, the sliding operation is determined to be a valid operation; or, Record the sliding time corresponding to the sliding operation, and determine whether the time interval between the sliding time and the sliding time corresponding to the previous sliding operation is greater than or equal to a time threshold, and / or determine whether the similarity between the sliding trajectory corresponding to the sliding operation and the sliding trajectory corresponding to the previous sliding operation is greater than or equal to a similarity threshold. If the time interval between the sliding time and the sliding time corresponding to the previous sliding operation is greater than or equal to the time threshold, and / or the similarity between the sliding trajectory corresponding to the sliding operation and the sliding trajectory corresponding to the previous sliding operation is greater than or equal to the similarity threshold, then the sliding operation is determined to be a valid operation.
8. A processing device, characterized in that, include: A processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 7.
10. A computer program product, characterised in that, It includes computer program instructions that cause a computer to perform the method as described in any one of claims 1 to 7.