Data processing method and system, computing device and storage medium
By calculating the text similarity and ratio of the target page image in a preset text object database, the problem of inaccurate retrieval in online reading is solved, and accurate e-book retrieval of paper books is achieved.
Patent Information
- Application Number
- CN202510880994.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-26
AI Technical Summary
When users read online, it is difficult to accurately retrieve e-books that are the same as paper books. Due to factors such as printing differences and damage, the retrieval accuracy is poor.
By acquiring the target page image, extracting the target text, and calculating the text similarity and ratio in the preset text object database, the target text page is determined and accurate retrieval is achieved.
Improved the retrieval accuracy of online reading, ensuring that users can quickly find the corresponding e-book content.
Smart Images

Figure CN120705607A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to data processing methods and systems, computing devices, and storage media. Background Art
[0002] With the development of Internet technology, paper reading has gradually shifted to online reading. Users can obtain e-books and other materials through the Internet for online reading, which is more convenient and easy. For example, users can read novels on a novel reading platform, or users can also browse e-textbooks on an online tutoring platform. However, when users look for books they want to read online, they can usually search on the Internet by taking pictures of paper books. Due to different printing dates or version years of books, the same page of the same book may have differences in some content, or paper books may be damaged or the content may be incorrect, resulting in the retrieved books being different from the content the user wants to read, and the retrieval accuracy is poor, which in turn affects the user's online reading experience. Therefore, there is an urgent need for an effective technical solution to solve the above problems. Summary of the Invention
[0003] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a data processing system, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0004] According to a first aspect of an embodiment of this specification, there is provided a data processing method, including: Acquire a target page image, and extract target text from the target page image; Determining, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; Calculating text similarity between the target text and candidate texts corresponding to each candidate text page, and calculating the text ratio of the target text relative to each candidate text page; According to the text similarities and text ratios corresponding to the candidate text pages, a target text page is determined from the at least one candidate text page, and a visualization page corresponding to the target text page is determined.
[0005] According to a second aspect of the embodiments of this specification, there is provided a data processing device, including: an acquisition module configured to acquire a target page image and extract target text from the target page image; A first determining module is configured to determine, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; a calculation module configured to calculate text similarity between the target text and candidate texts corresponding to each candidate text page, and calculate a text ratio of the target text relative to each candidate text page; The second determination module is configured to determine a target text page from the at least one candidate text page according to the text similarity and text ratio corresponding to the candidate text pages, and determine a visualization page corresponding to the target text page.
[0006] According to a third aspect of the embodiments of this specification, a data processing system is provided, including a client and a server, wherein: The server is configured to obtain a target page image and extract a target text from the target page image; determine, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database; calculate text similarity between the target text and candidate texts corresponding to each candidate text page, and calculate the text ratio of the target text relative to each candidate text page; determine a target text page from the at least one candidate text page based on the text similarity and text ratio corresponding to each candidate text page, determine a visualization page corresponding to the target text page, and send the visualization page corresponding to the target text page to the client, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; The client is configured to display a visual page corresponding to the target text page, wherein the visual page includes a point reading area.
[0007] According to a fourth aspect of the embodiments of this specification, there is provided a computing device, including: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.
[0008] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the computer program / instruction implements the steps of the above-mentioned data processing method when executed by a processor.
[0009] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0010] An embodiment of the present specification provides a data processing method, comprising: acquiring a target page image and extracting target text from the target page image; determining, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; calculating text similarity between the target text and candidate texts corresponding to each candidate text page, and calculating the text ratio of the target text relative to each candidate text page; determining a target text page from the at least one candidate text page based on the text similarity and text ratio corresponding to each candidate text page, and determining a visualization page corresponding to the target text page.
[0011] In the above method, after obtaining the target page image, the target text can be extracted from the target page image, and based on the target text, at least one candidate text page corresponding to the target text can be retrieved in the preset text object database, and the text similarity between the target text and the candidate texts corresponding to each candidate text page and the text ratio of the target text relative to each candidate text page can be calculated. By combining the two considerations of the text similarity between the local target text and the candidate text, and the text ratio of the overall target text relative to each candidate text page, the target text page corresponding to the target page image can be retrieved, the retrieval accuracy can be improved, and the user's online reading experience of the text object corresponding to the target page image can be further enhanced. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 This is a schematic diagram of an application scenario of a data processing method provided by an embodiment of this specification; Figure 2 is a flow chart of a data processing method provided by one embodiment of this specification; Figure 3 is a schematic diagram of an image of a text page in a data processing method provided by an embodiment of this specification; Figure 4 This is a schematic diagram of a visualization page and a point-reading area in a data processing method provided in one embodiment of this specification; Figure 5 This is a flowchart of a data processing method provided by one embodiment of this specification; Figure 6 This is a schematic diagram of the structure of a data processing device provided by one embodiment of this specification; Figure 7 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0012] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0013] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0014] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0015] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0016] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a foundation model. It is pre-trained on a large amount of unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as large language models (LLMs) and multi-modal pre-training models.
[0017] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0018] First, the terms involved in one or more embodiments of this specification are explained.
[0019] OCR: Optical Character Recognition, is a technology that converts text content in an image into editable and searchable text format.
[0020] id: identifier, unique identifier, uniquely identifies an object, entity, user, device, process, etc.
[0021] Textbook: A textbook is an official reading book designated by the school. It is written according to the school curriculum content and is also used to refer to books related to school courses.
[0022] Point reading: By clicking on the patterns, text, numbers and other content on the page, you can identify the specific content in the book and read out the sound file of the corresponding content.
[0023] In this specification, a data processing method is provided. This specification also relates to a data processing device, a data processing system, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0024] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of a data processing method provided according to an embodiment of this specification is shown.
[0025] like Figure 1 As shown, Figure 1 It includes a terminal side device 102 and a cloud side device 104.
[0026] In a specific implementation, a user can log in to the online education platform through the terminal device 102 and take a picture of the target page. The terminal device 102 sends the target page image to the cloud device 104 of the online education platform. The cloud device 104 receives the target page image, extracts the target text in the target page image, and obtains at least one candidate text page corresponding to the target text from a preset text object database. The target text is calculated with respect to the candidate texts corresponding to each candidate text page, and the text ratio of the target text relative to each candidate text page is calculated. Based on the text similarity and text ratio corresponding to each candidate text page, the target text page is determined from the at least one candidate text page, and the visualization page corresponding to the target text page is determined. The cloud device 104 sends the visualization page to the terminal device 102. The terminal device 102 can display the visualization page to the user through the online education platform, so that the user can subsequently perform operations such as clicking and reading on the visualization page. This enables users to accurately search for electronic teaching materials.
[0027] The end-side device 102 may include a browser, an application (APP), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program, a type of lightweight application), or a cloud application. The end-side device may be developed based on a software development kit (SDK) for the corresponding service provided by the server, such as a real-time communication (RTC) SDK. The end-side device may be deployed in an electronic device and may rely on the device or certain apps in the device to operate. The electronic device may have a display and support information browsing, such as a personal mobile terminal such as a mobile phone, tablet computer, or personal computer. Various other types of applications may also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0028] Cloud-side devices 104 can be understood as servers that provide various services, including physical servers and cloud servers. For example, these servers provide communication services to multiple clients, servers that support backend training for models used by clients, and servers that process data sent by clients. It should be noted that cloud-side devices 104 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. Cloud-side devices 104 can also be servers in a distributed system or servers integrated with blockchain. Cloud-side devices 104 can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or intelligent cloud computing servers or intelligent cloud hosts equipped with artificial intelligence technology.
[0029] It is worth noting that the data processing method provided in the embodiments of this specification can be executed by the cloud-side device 104 or by the end-side device 102; in other embodiments, the data processing method provided in the embodiments of this specification can also be executed jointly by the end-side device 102 and the cloud-side device 104.
[0030] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0031] Step 202: Acquire a target page image and extract target text from the target page image.
[0032] Specifically, the data processing method provided in the embodiments of this specification can be applied to the field of book retrieval. Specifically, the electronic version of the book can be retrieved based on the page image of the book photographed by the user, thereby realizing online reading of the book. The data processing method can be applied to the server side, for example, it can be applied to a novel reading platform or an online tutoring platform. In the novel reading platform, the electronic version of the novel can be retrieved based on the page image of the novel photographed by the user based on the data processing method. On the online tutoring platform, the electronic version of the textbook can be retrieved based on the page image of the textbook photographed by the user based on the data processing method. Subsequently, tutoring operations such as point reading can be performed on the electronic version of the textbook to realize online tutoring. For ease of understanding, the embodiments of this specification are explained by taking the application of the data processing method to textbook retrieval as an example, which does not affect the application of the data processing method provided in the embodiments of this specification in other fields.
[0033] The target page image can be understood as an image of a text page in a text object. For example, if the text object is a textbook, the target page image can be an image of any page in the textbook. It can be understood that the target page image can be the content page or cover page of the text object. The target text can be understood as the text content in the target page image, which can include information such as characters, pinyin, numbers, and page numbers.
[0034] In practice, the user can use the client to capture an image of the target page and send it to the server. The server can then receive the target page image and extract the target text from it. In practical applications, the target page image can be subjected to OCR recognition to obtain the target text and the coordinates of the recognition box.
[0035] Step 204: Based on the target text, determine at least one candidate text page corresponding to the target text in a preset text object database, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page.
[0036] Among them, the text object can be a paper book that needs to be retrieved, such as a paper textbook, a paper novel, etc. The text page can be understood as a page in a paper book, which can include a cover page and a content page. The text corresponding to the text page can be understood as the text in the text page, which can include text, numbers, pinyin, page numbers, etc. The preset text object database can be understood as a pre-built database that stores electronic data of text objects. It can be understood that the preset text object database can be determined according to actual needs. For example, for a novel reading platform, the preset text object database can store electronic data of novels. For an online teaching platform, the preset text object database can store electronic data of textbooks. This embodiment of the specification does not limit this. The candidate text page can be understood as a text page containing the target text.
[0037] Specifically, after the target text contained in the target page image is identified, at least one candidate text page containing the target text may be retrieved from a preset text object database.
[0038] For example, the preset text object database may include textbook A, all text pages in textbook A and the text corresponding to each text page, textbook B, all text pages in textbook B and the text corresponding to each text page, textbook C, all text pages in textbook C and the text corresponding to each text page. After obtaining the target text corresponding to the target page image, all text pages of textbook A, all text pages of textbook B and all text pages of textbook C contained in the preset text object database can be searched based on the target text to obtain the text page containing the target text as a candidate text page.
[0039] In actual applications, before obtaining the target page image, the following steps are also included: Performing text recognition on at least one text page included in at least one text object to obtain text corresponding to the at least one text page; The preset text object database including the at least one text object is constructed according to the text corresponding to the at least one text page.
[0040] Specifically, a preset text object database can be constructed in advance based on at least one text object. Taking a textbook as an example, the image of each text page of at least one textbook can be determined, and the text, numbers, pinyin, page numbers, etc. contained in the image of each text page can be identified using a text recognition model to obtain the text of each text page. The text of each text page, the textbook corresponding to the text page, and the page number corresponding to each text page are stored in the database, thereby constructing a preset text object database for an electronic teaching material including at least one textbook. It can be understood that in the preset text object database, text, text pages, and text objects are stored correspondingly, that is, the preset text object database actually also stores the corresponding relationship between text, text pages, and text objects. Among them, the text recognition model can be a trained machine learning model with OCR recognition capabilities, a deep learning model, etc., or it can also be a large model. Moreover, when applied to textbook retrieval, during the training process of the text recognition model, the text recognition model can be trained based on sample images containing text, numbers, pinyin and other information, so that it can have the ability to recognize text, numbers, pinyin and other information, thereby making it suitable for textbook retrieval.
[0041] In practical applications, see Figure 3 , Figure 3 FIG. 1 shows a schematic diagram of an image of a text page in a data processing method provided according to an embodiment of the present specification. Figure 3 As shown, after the image of the text page is recognized using the text recognition model, the following recognition content can be obtained, and the recognition content is stored as the text corresponding to the text page in the preset text object database.
[0042] { "res_id":"12110011021610102001530595019531", "tb_id":"1211001102161", "tb_name":"Chinese Language, Grade 1, Volume 2 (2016 Edition)", "res_path":"*****.mp3", "ori_tree_name":"2 Surname Song", "rect_pos_lst":[ { "rect_pos_x":143.79921, "rect_pos_y":67.29962, "rect_width":220.835648, "rect_height":64.75903, "content":"$LINE_S$$LINE_E$$LINE_S$Surname Song$LINE_E$" } ], "page_num":4, "res_index":18, }, { "res_id":"12110011021610102001530595019532", "tb_id":"1211001102161", "tb_name":"Chinese Language, Grade 1, Volume 2 (2016 Edition)", "res_path":"*****.mp3", "ori_tree_name":"2 Surname Song", "rect_pos_lst":[ { "rect_pos_x":124.540985, "rect_pos_y":166.34314, "rect_width":263.203979, "rect_height":44.4429321, "content":"$LINE_S$$LINE_E$$LINE_S$$LINE_E$$LINE_S$What's your last name? My last name is Li.$LINE_E$" } ], "page_num":4, } Among them, res_id and tb_id are the IDs corresponding to the resource; tb_name is the textbook name; res_path is the audio reading corresponding to the resource; ori_tree_name is the unit name corresponding to the text page; res_pos_lst is the coordinates and text content corresponding to the text content in the text page; page_num is the page number of the text page in the textbook.
[0043] Then, in the subsequent retrieval process, retrieval can be performed based on the target text in the target page image and the recognition box coordinates, the above-mentioned ID contained in the preset text object database, the textbook name, the unit name, the coordinates and text content corresponding to the text content, the page number of the text page in the textbook and other information to obtain at least one candidate text page containing the target text.
[0044] In summary, by pre-building a preset text object database, it is convenient to subsequently retrieve candidate text pages corresponding to the target text from the preset text object database, thereby improving retrieval efficiency.
[0045] Furthermore, the determining, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database includes: According to the target text, determining at least one related text corresponding to the target text in the preset text object database; The text page corresponding to the at least one related text is determined as at least one candidate text page corresponding to the target text.
[0046] The relevant text can be understood as the text whose similarity with the target text reaches the similarity threshold. Alternatively, the relevant text can also be the text that is completely identical to the target text. For example, if the target text is "What is your name", the relevant text corresponding to the target text can be "What is your name" or "What is your name".
[0047] Specifically, the target text can be determined from the text corresponding to each text page included in each text object in the preset text object database. Then, all text pages containing the relevant text can be determined as candidate text pages based on the correspondence between the text and the text page stored in the preset text object database. It is understandable that at least one relevant text can be a text stored in different text pages. For example, for the target text "What is your name", 20 relevant texts are obtained in the preset text object database, all of which are "What is your name". Among these 20 texts, some may be stored in the same text page, and some may be stored in different text pages. For example, there may be one such relevant text in text page 10 of textbook A. It is understandable that "10" may refer to the page number of the text page. There may be two such relevant texts in text page 25 of textbook A, and there may be one such relevant text in text page 20 of textbook B.
[0048] In practical applications, the target text can be searched in a preset text object database to obtain the 20 most similar texts to the target text. These 20 related texts are then aggregated according to their corresponding textbook titles and text page numbers, thereby obtaining the text pages corresponding to these 20 related texts as candidate text pages. The number of related texts obtained can be determined based on actual needs.
[0049] Furthermore, the target text can be divided into multiple sub-target texts. Then, for each sub-target text, a search can be performed in the preset text object database according to the above process, thereby obtaining a text page containing the relevant text corresponding to the sub-target text as a candidate text page. It can be understood that the candidate text page at this time can contain the relevant text corresponding to all sub-target texts, or can contain the relevant text corresponding to any one or more sub-target texts. For example, if the target text is divided into sub-target text 1 and sub-target text 2, then 20 relevant texts corresponding to sub-target text 1 can be obtained from the preset text object database, and these 20 relevant texts can be aggregated according to their corresponding textbook names and text page numbers, thereby obtaining the text pages corresponding to the 20 relevant texts as candidate text pages. In addition, 20 relevant texts corresponding to sub-target text 2 can be obtained from the preset text object database, and these 20 relevant texts can be aggregated according to their corresponding textbook names and text page numbers, thereby obtaining the text pages corresponding to the 20 relevant texts as candidate text pages.
[0050] In summary, by performing an initial screening in the preset text object database based on the target text, multiple candidate text pages are screened out, which facilitates subsequent secondary screening from the multiple candidate text pages to obtain the target text page with the highest matching degree, thereby saving screening time and further improving retrieval efficiency and accuracy.
[0051] Step 206: Calculate the text similarity between the target text and the candidate texts corresponding to each candidate text page, and calculate the text ratio of the target text relative to each candidate text page.
[0052] The candidate text corresponding to a candidate text page can be understood as the candidate text contained in the candidate text page. The text ratio of the target text relative to each candidate text page can be understood as the text ratio of the target text relative to all candidate texts in each candidate text page. It can further be understood as the text ratio of candidate texts similar to the target text in each candidate text page relative to all candidate texts in each candidate text page. Candidate texts similar to the target text can be candidate texts whose text similarity with the target text is greater than a preset similarity threshold.
[0053] Specifically, after determining at least one candidate text page corresponding to the target text, the text similarity between the target text and the candidate text in each candidate text page may be calculated, and the text ratio of the target text relative to each candidate text page may be calculated.
[0054] Then, it can be understood that, for one of the candidate text pages, if the similarity between the target text and the candidate text in the candidate text page is less than the preset similarity threshold, then the subsequent calculation of the text ratio of the target text relative to the candidate text page can be omitted, and it can be directly stated that the candidate text page is not the target text page corresponding to the target page image.
[0055] In practical applications, the calculation of the text similarity between the target text and the candidate texts corresponding to each candidate text page includes: Determining a candidate text to be processed corresponding to a candidate text page to be processed, wherein the candidate text page to be processed is any one of the candidate text pages; Calculate the text similarity between the target text and the candidate text to be processed.
[0056] Specifically, any one of the at least one candidate text page can be used as a candidate text page to be processed, and the candidate text to be processed included in the candidate text page to be processed is determined, and the text similarity between the target text and the candidate text to be processed is calculated to obtain the text similarity between the target text and the candidate text corresponding to each candidate text page.
[0057] For example, the target text corresponds to candidate text page 1, candidate text page 2 and candidate text page 3. Then, candidate text page 1 can be used as the candidate text page to be processed, the candidate text to be processed included in candidate text page 1 can be determined, and the text similarity between the target text and the candidate text to be processed included in candidate text page 1 can be calculated. Similarly, candidate text page 2 can be used as the candidate text page to be processed, the candidate text to be processed included in candidate text page 2 can be determined, and the text similarity between the target text and the candidate text to be processed included in candidate text page 2 can be calculated; candidate text page 3 can be used as the candidate text page to be processed, the candidate text to be processed included in candidate text page 3 can be determined, and the text similarity between the target text and the candidate text to be processed included in candidate text page 3 can be calculated.
[0058] In summary, by calculating the text similarity between the target text and the candidate texts contained in each candidate text page, a data basis is provided for the subsequent determination of the target text page.
[0059] Furthermore, the calculation between the target text and the candidate text to be processed included in the candidate text page to be processed is taken as an example for explanation. It can be understood that the calculation between the target text and the candidate text included in each candidate text page is performed according to the following process, and the specific implementation method is as follows.
[0060] The target text includes at least one sub-target text, and the candidate text to be processed includes at least one sub-candidate text to be processed; The calculating of the text similarity between the target text and the candidate text to be processed includes: The text similarity between each sub-target text in the at least one sub-target text and each sub-candidate text to be processed in the at least one sub-candidate text to be processed is calculated.
[0061] In practical applications, a textbook's text page may contain a large amount of text. For example, a text page in a Chinese textbook may contain a text. Therefore, to facilitate text similarity analysis, the target text in the target page image can be divided into at least one sub-target text. Correspondingly, the candidate text to be processed in the candidate text page can also be divided into at least one sub-candidate text to be processed. It is understandable that the sub-target texts and sub-candidate texts to be processed are also stored in the preset text object database. For example, if the target text is "We walked in the fields: me, my mother, my wife, and my son," then the sub-target texts can be "We walked in the fields," "me, my mother," and "my wife and my son."
[0062] In specific implementations, the longest common subsequence algorithm can be used to calculate the degree of match between each sub-target text and each sub-candidate text to be processed. This degree of match can be understood as the aforementioned text similarity. For a sub-target text, if there exists a sub-candidate text to be processed whose longest common subsequence length is greater than 60% (i.e., the preset similarity threshold), then the text similarity between the sub-target text and the sub-candidate text to be processed exceeds the preset similarity threshold.
[0063] For example, the target text includes sub-target text 1 and sub-target text 2, and the candidate text to be processed includes sub-to-be-processed candidate text 1, sub-to-be-processed candidate text 2, and sub-to-be-processed candidate text 3. Then, the text similarity between sub-target text 1 and sub-to-be-processed candidate text 1 can be calculated, the text similarity between sub-target text 1 and sub-to-be-processed candidate text 2 can be calculated, the text similarity between sub-target text 1 and sub-to-be-processed candidate text 2 can be calculated, the text similarity between sub-target text 1 and sub-to-be-processed candidate text 3 can be calculated, the text similarity between sub-target text 2 and sub-to-be-processed candidate text 1 can be calculated, the text similarity between sub-target text 2 and sub-to-be-processed candidate text 2 can be calculated, and the text similarity between sub-target text 2 and sub-to-be-processed candidate text 3 can be calculated.
[0064] In summary, by dividing the target text and the candidate text to be processed, the efficiency of text similarity calculation is improved, and the accuracy of subsequent selection of target text pages based on text similarity is further improved, avoiding the situation where the retrieval difficulty, similarity calculation efficiency and retrieval accuracy are low when the number of target texts and candidate texts to be processed is large.
[0065] Furthermore, the calculating of the text ratio of the target text to the candidate text pages includes: The text ratio of the target text relative to the candidate text pages is calculated based on the text similarity between the target text and the candidate texts corresponding to the candidate text pages.
[0066] Specifically, the text ratio of candidate texts similar to the target text in each candidate text page relative to all candidate texts in each candidate text page may be calculated based on the text similarity between the target text and the candidate texts included in each candidate text page.
[0067] Furthermore, if the text similarity between the target text and the candidate text corresponding to each candidate text page is greater than a preset similarity threshold, the text ratio of the target text relative to each candidate text page can be calculated. For example, if the text similarity between the target text and the candidate text corresponding to candidate text page 1 is greater than a preset similarity threshold, the text ratio of the target text relative to candidate text page 1 can be calculated. If the text similarity between the target text and the candidate text corresponding to candidate text page 2 is less than or equal to the preset similarity threshold, the text ratio of the target text relative to candidate text page 2 can be omitted.
[0068] In summary, the text similarity between the target text and the candidate texts reflects the text similarity between local paragraphs in the text page, and the text ratio reflects the proportion of candidate texts similar to the target text relative to the entire candidate text page. By combining the local and the overall, the accuracy of determining the target text page is further improved. In a specific implementation, the calculating of the text ratio of the target text relative to each candidate text page based on the text similarity between the target text and the candidate texts corresponding to each candidate text page includes: selecting, from the at least one sub-target candidate text, a sub-target candidate text corresponding to each sub-target text according to text similarities between each sub-target text in the at least one sub-target text and each sub-target candidate text in the at least one sub-target candidate text; The text ratio of the target text relative to the page of candidate text to be processed is calculated according to the number of the sub-target candidate texts to be processed and the number of the at least one sub-candidate text to be processed.
[0069] The sub-target candidate text to be processed corresponding to the sub-target text can be understood as a sub-target candidate text in at least one sub-target candidate text, whose text similarity with the sub-target text is greater than a preset similarity threshold.
[0070] Specifically, based on the text similarity between each sub-target text and each sub-candidate text to be processed, the sub-target candidate text to be processed corresponding to each sub-target text can be selected from at least one sub-candidate text to be processed, and based on the number of sub-target candidate texts to be processed and the number of sub-candidate texts to be processed, the text ratio of the target text relative to the candidate text page to be processed can be calculated. Similarly, the text ratio of the target text relative to each candidate text page to be processed can be calculated.
[0071] In a specific implementation, the selecting, from the at least one sub-target candidate text, the sub-target candidate text corresponding to each sub-target text according to the text similarity between each sub-target text in the at least one sub-target text and each sub-target candidate text in the at least one sub-target candidate text, includes: The sub-to-be-processed candidate texts whose text similarity with the sub-target texts is greater than a preset similarity threshold among the at least one sub-to-be-processed candidate texts are determined as the sub-target to-be-processed candidate texts corresponding to the sub-target texts.
[0072] Specifically, a sub-candidate text to be processed whose text similarity with the sub-target text is greater than a preset similarity threshold may be determined as the sub-target candidate text to be processed corresponding to the sub-target text.
[0073] For example, the target text includes sub-target text 1 and sub-target text 2, and the candidate text to be processed includes sub-to-be-processed candidate text 1, sub-to-be-processed candidate text 2, and sub-to-be-processed candidate text 3. The text similarity between sub-target text 1 and sub-to-be-processed candidate text 1 is calculated, the text similarity between sub-target text 1 and sub-to-be-processed candidate text 2 is calculated, the text similarity between sub-target text 1 and sub-to-be-processed candidate text 3 is calculated, the text similarity between sub-target text 2 and sub-to-be-processed candidate text 1 is calculated, the text similarity between sub-target text 2 and sub-to-be-processed candidate text 2 is calculated, and the text similarity between sub-target text 2 and sub-to-be-processed candidate text 3 is calculated. For sub-target text 1, the sub-to-be-processed candidate text 2 whose text similarity is greater than the preset similarity threshold is used as the sub-target to-be-processed candidate text corresponding to sub-target text 1. For sub-target text 2, the sub-to-be-processed candidate text 1 whose text similarity is greater than the preset similarity threshold is used as the sub-target to-be-processed candidate text corresponding to sub-target text 2. Then, the candidate text page to be processed contains 3 sub-to-be-processed candidate texts and 2 sub-target to-be-processed candidate texts, and the text ratio of the target text relative to the candidate text page to be processed is calculated to be 2 / 3.
[0074] Furthermore, for a particular candidate text page, if at least one of the candidate sub-sub ...
[0075] In summary, by calculating the text ratio, an overall comparison between the candidate text pages and the target text can be achieved, further improving the retrieval accuracy.
[0076] Step 208: Determine a target text page from the at least one candidate text page based on the text similarities and text ratios corresponding to the candidate text pages, and determine a visualization page corresponding to the target text page.
[0077] The target text page can be understood as the text page corresponding to the target page image. The text similarity corresponding to the candidate text page can be understood as the text similarity between the target text and the candidate text corresponding to the candidate text page. The text ratio corresponding to the candidate text page can be understood as the text ratio of the target text to the candidate text page.
[0078] Specifically, after calculating the text similarity and text ratio, a target text page may be selected from at least one candidate text page according to the text similarity and text ratio corresponding to each candidate text page, and a visualization page corresponding to the target text page may be determined.
[0079] In a specific implementation, determining a target text page from the at least one candidate text page according to the text similarity and text ratio corresponding to each candidate text page includes: Among the at least one candidate text page, a candidate text page whose text similarity is greater than a preset similarity threshold and whose text ratio is greater than a preset ratio threshold is determined as the target text page.
[0080] Specifically, a candidate text page of which the text similarity is greater than a preset similarity threshold and the text ratio is greater than a preset ratio threshold among the at least one candidate text page may be determined as the target text page.
[0081] It can be understood that, among at least one candidate text page, if there are multiple candidate text pages whose text similarity is greater than a preset similarity threshold and whose text ratio is greater than a preset ratio threshold, then the candidate text page with the largest text similarity and / or preset ratio threshold among the multiple candidate text pages is determined as the target text page.
[0082] Optionally, after calculating the text similarity between the target text and the candidate texts corresponding to each candidate text page, it can be determined whether the text similarity corresponding to each candidate text page is greater than a preset similarity threshold. If so, the candidate text page is determined to be an intermediate text page; otherwise, the match fails. Based on the text similarity corresponding to the intermediate text page, the text ratio corresponding to the intermediate text page is calculated. This calculation process is similar to the process of calculating the text ratio corresponding to the candidate text page described above and will not be repeated in this embodiment of the present specification. It is determined whether the text ratio corresponding to the intermediate text page is greater than a preset ratio threshold. If so, the intermediate text page is determined to be the target text page; otherwise, the match fails.
[0083] In actual applications, the preset ratio threshold may be 70%. The preset ratio threshold may be dynamically adjusted according to actual needs, and the embodiments of this specification do not limit this.
[0084] Furthermore, determining, among the at least one candidate text page, a candidate text page whose text similarity is greater than a preset similarity threshold and whose text ratio is greater than a preset ratio threshold as the target text page includes: The candidate text page in the at least one candidate text page, in which the text similarity between each sub-target text and the sub-target candidate text to be processed corresponding to each sub-target text is greater than the preset similarity threshold, and the text ratio is greater than the preset ratio threshold, is determined as the target text page.
[0085] Specifically, a candidate text page in which the text similarity between each sub-target text and the sub-target candidate text to be processed corresponding to the sub-target text in at least one candidate text page is greater than a preset similarity threshold and the text ratio is greater than a preset ratio threshold can be determined as a target text page.
[0086] Furthermore, the preset text object database also includes a cover image corresponding to the at least one text object; After obtaining the target page image, the method further includes: performing image matching on the target page image and the cover image corresponding to the at least one text object to obtain an image matching result; In the case where an associated cover image associated with the target page image is determined according to the image matching result, a text object corresponding to the associated cover image is determined, and a visual page corresponding to the text object is determined; or If it is determined according to the image matching result that the target page image is not associated with a cover image, target text in the target page image is extracted.
[0087] The associated cover image associated with the target page image can be understood as the cover image having the highest image similarity with the target page image.
[0088] In actual applications, since the cover of a textbook contains relatively little text content or may not contain text, if the target page image photographed by the user is the cover image of a textbook, it is difficult to retrieve it through the above-mentioned text content extraction and recognition method. Based on this, a cover image corresponding to each text object can be stored in a preset text object database, and the target page image and the cover image stored in the preset text object database can be image matched to obtain an image matching result. If it is determined based on the image matching result that there is an associated cover image associated with the target page image in the preset text object database, the text object corresponding to the associated cover image can be directly determined. If it is determined based on the image matching result that there is no associated cover image associated with the target page image in the preset object database, text content extraction can be performed on the target page image to extract the target text in the target page image.
[0089] Furthermore, when constructing a preset text object database, an image feature extraction model can be used to extract image features from the cover images of all text objects, and the image features of the cover images can be stored in the preset text object database. When searching the preset text object database based on the target page image, the target image features of the target page image can be extracted, and the image feature similarity between the target image features of the target page image and the image features of the cover images stored in the preset text object database can be calculated. The cover image corresponding to the image feature in the preset text object database whose image feature similarity is greater than the preset image feature similarity threshold and whose image feature similarity is the largest is used as the associated cover image of the target page image. In practical applications, the image feature extraction model can be any model with image feature extraction capabilities, such as a trained machine learning model, a deep learning model, a large model, etc., and the embodiments of this specification do not limit this.
[0090] In summary, by combining images and text content, we can achieve comprehensive retrieval of target page images, improve retrieval efficiency and enhance retrieval accuracy.
[0091] In actual applications, after determining the visualization page corresponding to the target text page, the method further includes: The visual page corresponding to the target text page is sent to the client, so that the client displays the visual page corresponding to the target text page, wherein the visual page includes a point reading area.
[0092] For details, see Figure 4 , Figure 4 A schematic diagram of a visualization page and a point-reading area in a data processing method provided according to an embodiment of the present specification is shown. After the server determines the visualization page corresponding to the target text page, the visualization page and the corresponding audio data can be sent to the client, so that the client displays the visualization page and the point-reading area contained in the visualization page. The client can read aloud the text content in the point-reading area in response to the user's click operation on the point-reading area, that is, play the corresponding audio data.
[0093] In addition, the server can also send the visual pages and point-reading audio of all text pages of the text object corresponding to the target text page to the client. In response to the user's page-turning operation, the client can display the visual pages of other text pages of the text object and read them.
[0094] To sum up, in the above method, after obtaining the target page image, the target text can be extracted from the target page image, and based on the target text, at least one candidate text page corresponding to the target text can be retrieved in the preset text object database, and the text similarity between the target text and the candidate text corresponding to each candidate text page and the text ratio of the target text relative to each candidate text page are calculated. By combining the two considerations of the text similarity between the local target text and the candidate text, and the text ratio of the overall target text relative to each candidate text page, the target text page corresponding to the target page image can be retrieved, the retrieval accuracy can be improved, and the user's online reading experience of the text object corresponding to the target page image can be further enhanced. The following combined Figure 5 , taking the application of the data processing method provided in this specification in textbook retrieval and point reading as an example, the data processing method is further explained. Figure 5 A flowchart of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0095] Step 502: Acquire the target page image.
[0096] Step 504: Cover image matching.
[0097] If the match is successful, step 506 is executed; if the match fails, step 508 is executed.
[0098] Specifically, the similarity between the image features of the target page image and the image features of the textbook cover images stored in the database can be calculated. If the similarity is greater than the image similarity threshold, it means that the match is successful, and the cover image with the greatest similarity is used as the associated cover image of the target page image.
[0099] Step 506: Return the visual page and point reading information of the target text page corresponding to the target page image.
[0100] Specifically, the point reading information may be the point reading audio data of the point reading area included in the visual page.
[0101] Step 508: Extract the target text from the target page image.
[0102] Step 510: Retrieve at least one candidate text page corresponding to the target text.
[0103] Specifically, the target text may be searched in the database to retrieve 20 most similar related texts, which are then aggregated according to textbook names and page numbers to obtain at least one candidate text page.
[0104] Step 512: Calculate text similarity.
[0105] Specifically, for each candidate text page, the text similarity between the target text and the candidate text in each candidate text page may be calculated.
[0106] Step 514: Calculate the text ratio.
[0107] Specifically, for each candidate text page, the text ratio of the target text relative to each candidate text page is calculated.
[0108] Step 516: Check whether the text matching is successful based on the text similarity and text ratio.
[0109] If yes, then step 506 is executed; if no, then an empty result is returned.
[0110] Specifically, when the text similarity is greater than a preset similarity threshold and the text ratio is greater than a preset ratio threshold, it is determined that the text matching is successful.
[0111] In the above method, after obtaining the target page image, the target text can be extracted from the target page image, and based on the target text, at least one candidate text page corresponding to the target text can be retrieved in the preset text object database, and the text similarity between the target text and the candidate texts corresponding to each candidate text page and the text ratio of the target text relative to each candidate text page can be calculated. By combining the two considerations of the text similarity between the local target text and the candidate text, and the text ratio of the overall target text relative to each candidate text page, the target text page corresponding to the target page image can be retrieved, the retrieval accuracy can be improved, and the user's online reading experience of the text object corresponding to the target page image can be further enhanced. Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 6 FIG1 shows a schematic diagram of the structure of a data processing device provided by an embodiment of this specification. Figure 6 As shown, the device includes: An acquisition module 602 is configured to acquire a target page image and extract target text from the target page image; A first determining module 604 is configured to determine, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; A calculation module 606 is configured to calculate text similarity between the target text and candidate texts corresponding to each candidate text page, and calculate a text ratio of the target text relative to each candidate text page; The second determining module 608 is configured to determine a target text page from the at least one candidate text page according to the text similarity and text ratio corresponding to the candidate text pages, and determine a visualization page corresponding to the target text page.
[0112] In an optional embodiment, the calculation module 606 is further configured to: Determining a candidate text to be processed corresponding to a candidate text page to be processed, wherein the candidate text page to be processed is any one of the candidate text pages; Calculate the text similarity between the target text and the candidate text to be processed.
[0113] In an optional embodiment, the target text includes at least one sub-target text, and the candidate text to be processed includes at least one sub-candidate text to be processed; The calculation module 606 is further configured to: The text similarity between each sub-target text in the at least one sub-target text and each sub-candidate text to be processed in the at least one sub-candidate text to be processed is calculated.
[0114] In an optional embodiment, the calculation module 606 is further configured to: The text ratio of the target text relative to the candidate text pages is calculated based on the text similarity between the target text and the candidate texts corresponding to the candidate text pages.
[0115] In an optional embodiment, the calculation module 606 is further configured to: selecting, from the at least one sub-target candidate text, a sub-target candidate text corresponding to each sub-target text according to text similarities between each sub-target text in the at least one sub-target text and each sub-target candidate text in the at least one sub-target candidate text; The text ratio of the target text relative to the page of candidate text to be processed is calculated according to the number of the sub-target candidate texts to be processed and the number of the at least one sub-candidate text to be processed.
[0116] In an optional embodiment, the calculation module 606 is further configured to: The sub-to-be-processed candidate texts whose text similarity with the sub-target texts is greater than a preset similarity threshold among the at least one sub-to-be-processed candidate texts are determined as the sub-target to-be-processed candidate texts corresponding to the sub-target texts.
[0117] In an optional embodiment, the second determining module 608 is further configured to: Among the at least one candidate text page, a candidate text page whose text similarity is greater than a preset similarity threshold and whose text ratio is greater than a preset ratio threshold is determined as the target text page.
[0118] In an optional embodiment, the second determining module 608 is further configured to: The candidate text page in the at least one candidate text page, in which the text similarity between each sub-target text and the sub-target candidate text to be processed corresponding to each sub-target text is greater than the preset similarity threshold, and the text ratio is greater than the preset ratio threshold, is determined as the target text page.
[0119] In an optional embodiment, the first determining module 604 is further configured to: According to the target text, determining at least one related text corresponding to the target text in the preset text object database; The text page corresponding to the at least one related text is determined as at least one candidate text page corresponding to the target text.
[0120] In an optional embodiment, the preset text object database further includes a cover image corresponding to the at least one text object; The acquisition module 602 is further configured to: performing image matching on the target page image and the cover image corresponding to the at least one text object to obtain an image matching result; In a case where an associated cover image associated with the target page image is determined according to the image matching result, determining a text object corresponding to the associated cover image, and displaying a visual page corresponding to the text object; or If it is determined according to the image matching result that the target page image is not associated with a cover image, target text in the target page image is extracted.
[0121] In an optional embodiment, the apparatus further includes a sending module configured to: The visual page corresponding to the target text page is sent to the client, so that the client displays the visual page corresponding to the target text page, wherein the visual page includes a point reading area.
[0122] In an optional embodiment, the apparatus further comprises a building module configured to: Performing text recognition on at least one text page included in at least one text object to obtain text corresponding to the at least one text page; The preset text object database including the at least one text object is constructed according to the text corresponding to the at least one text page.
[0123] In the above-mentioned device, after obtaining the target page image, the target text can be extracted from the target page image, and based on the target text, at least one candidate text page corresponding to the target text can be retrieved in the preset text object database, and the text similarity between the target text and the candidate texts corresponding to each candidate text page and the text ratio of the target text relative to each candidate text page can be calculated. By combining the two considerations of the text similarity between the local target text and the candidate text, and the text ratio of the overall target text relative to each candidate text page, the target text page corresponding to the target page image can be retrieved, the retrieval accuracy can be improved, and the user's online reading experience of the text object corresponding to the target page image can be further enhanced.
[0124] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the above-mentioned data processing method.
[0125] Corresponding to the above method embodiment, this specification also provides a data processing system embodiment, which includes a client and a server, wherein: The server is configured to obtain a target page image and extract a target text from the target page image; determine, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database; calculate text similarity between the target text and candidate texts corresponding to each candidate text page, and calculate the text ratio of the target text relative to each candidate text page; determine a target text page from the at least one candidate text page based on the text similarity and text ratio corresponding to each candidate text page, determine a visualization page corresponding to the target text page, and send the visualization page corresponding to the target text page to the client, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; The client is configured to display a visual page corresponding to the target text page, wherein the visual page includes a point reading area.
[0126] In the above system, after obtaining the target page image, the target text can be extracted from the target page image, and based on the target text, at least one candidate text page corresponding to the target text can be retrieved in the preset text object database, and the text similarity between the target text and the candidate texts corresponding to each candidate text page and the text ratio of the target text relative to each candidate text page can be calculated. By combining the two considerations of the text similarity between the local target text and the candidate text, and the text ratio of the overall target text relative to each candidate text page, the target text page corresponding to the target page image can be retrieved, the retrieval accuracy can be improved, and the user's online reading experience of the text object corresponding to the target page image can be further enhanced.
[0127] Figure 7 7 shows a block diagram of a computing device 700 according to one embodiment of the present disclosure. Components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0128] Computing device 700 also includes an access device 740 that enables computing device 700 to communicate via one or more networks 760. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 740 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0129] In one embodiment of the present application, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 7 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0130] Computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 700 can also be a mobile or stationary server.
[0131] The processor 720 is configured to execute the following computer program / instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.
[0132] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the computing device embodiment is generally similar to the data processing method embodiment, so the description is relatively simple. For relevant parts, refer to the description of the data processing method embodiment.
[0133] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0134] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the computer-readable storage medium embodiment is generally similar to the data processing method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the data processing method embodiment.
[0135] An embodiment of the present specification further provides a computer program product, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0136] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-mentioned data processing method.
[0137] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0138] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0139] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0140] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0141] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method, characterized in that: include: Acquire a target page image, and extract target text from the target page image; Determining, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; Calculating text similarity between the target text and candidate texts corresponding to each candidate text page, and calculating the text ratio of the target text relative to each candidate text page; According to the text similarities and text ratios corresponding to the candidate text pages, a target text page is determined from the at least one candidate text page, and a visualization page corresponding to the target text page is determined.
2. The method according to claim 1, characterized in that The calculating of the text similarity between the target text and the candidate texts corresponding to each candidate text page includes: Determining a candidate text to be processed corresponding to a candidate text page to be processed, wherein the candidate text page to be processed is any one of the candidate text pages; Calculate the text similarity between the target text and the candidate text to be processed.
3. The method according to claim 2, characterized in that The target text includes at least one sub-target text, and the candidate text to be processed includes at least one sub-candidate text to be processed; The calculating of the text similarity between the target text and the candidate text to be processed includes: The text similarity between each sub-target text in the at least one sub-target text and each sub-candidate text to be processed in the at least one sub-candidate text to be processed is calculated.
4. The method according to claim 3, characterized in that Calculating the text ratio of the target text relative to the candidate text pages includes: The text ratio of the target text relative to the candidate text pages is calculated based on the text similarity between the target text and the candidate texts corresponding to the candidate text pages.
5. The method according to claim 4, characterized in that Calculating the text ratio of the target text relative to each candidate text page based on the text similarity between the target text and the candidate texts corresponding to each candidate text page includes: selecting, from the at least one sub-target candidate text, a sub-target candidate text corresponding to each sub-target text according to text similarities between each sub-target text in the at least one sub-target text and each sub-target candidate text in the at least one sub-target candidate text; The text ratio of the target text to the page of candidate text to be processed is calculated according to the number of the sub-target candidate texts to be processed and the number of the at least one sub-candidate text to be processed.
6. The method according to claim 5, characterized in that The selecting, from the at least one sub-target candidate text, sub-target candidate texts corresponding to the at least one sub-target texts based on text similarities between the at least one sub-target texts and the at least one sub-target candidate texts, comprises: The sub-to-be-processed candidate texts whose text similarity with the sub-target texts is greater than a preset similarity threshold among the at least one sub-to-be-processed candidate texts are determined as the sub-target to-be-processed candidate texts corresponding to the sub-target texts.
7. The method according to claim 5, characterized in that Determining a target text page from the at least one candidate text page according to the text similarities and text ratios corresponding to the candidate text pages includes: Among the at least one candidate text page, a candidate text page whose text similarity is greater than a preset similarity threshold and whose text ratio is greater than a preset ratio threshold is determined as the target text page.
8. The method according to claim 7, characterized in that The step of determining, among the at least one candidate text page, a candidate text page whose text similarity is greater than a preset similarity threshold and whose text ratio is greater than a preset ratio threshold as the target text page includes: The candidate text page in the at least one candidate text page, in which the text similarity between each sub-target text and the sub-target candidate text to be processed corresponding to each sub-target text is greater than the preset similarity threshold, and the text ratio is greater than the preset ratio threshold, is determined as the target text page.
9. The method according to any one of claims 1 to 8, characterized in that The step of determining, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database includes: According to the target text, determining at least one related text corresponding to the target text in the preset text object database; The text page corresponding to the at least one related text is determined as at least one candidate text page corresponding to the target text.
10. The method according to any one of claims 1 to 8, characterized in that The preset text object database also includes a cover image corresponding to the at least one text object; After obtaining the target page image, the method further includes: performing image matching on the target page image and the cover image corresponding to the at least one text object to obtain an image matching result; In a case where an associated cover image associated with the target page image is determined according to the image matching result, determining a text object corresponding to the associated cover image, and displaying a visual page corresponding to the text object; or If it is determined according to the image matching result that the target page image is not associated with a cover image, target text in the target page image is extracted.
11. The method according to any one of claims 1 to 8, characterized in that After determining the visualization page corresponding to the target text page, the method further includes: The visual page corresponding to the target text page is sent to the client, so that the client displays the visual page corresponding to the target text page, wherein the visual page includes a point reading area.
12. The method according to any one of claims 1 to 8, characterized in that Before obtaining the target page image, the method further includes: Performing text recognition on at least one text page included in at least one text object to obtain text corresponding to the at least one text page; The preset text object database including the at least one text object is constructed according to the text corresponding to the at least one text page.
13. A data processing system, characterized in that: Including client and server, The server is configured to obtain a target page image and extract a target text from the target page image; determine, based on the target text, at least one candidate text page corresponding to the target text in a preset text object database; calculate text similarity between the target text and candidate texts corresponding to each candidate text page, and calculate the text ratio of the target text relative to each candidate text page; determine a target text page from the at least one candidate text page based on the text similarity and text ratio corresponding to each candidate text page, determine a visualization page corresponding to the target text page, and send the visualization page corresponding to the target text page to the client, wherein the preset text object database includes at least one text object, and the text object includes at least one text page and text corresponding to each text page; The client is configured to display a visual page corresponding to the target text page, wherein the visual page includes a point reading area.
14. A computing device, characterized in that include: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium, characterized in that It stores a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 12 when executed by a processor.
16. A computer program product, characterized in that The method comprises a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Intelligent reading recommendation method and device and electronic equipment
CN107679070A
Method and device for automatically identifying pages
CN110209759A
Similarity calculation method based on text and semantics, server and storage medium
CN110222154A
Data processing method and device, equipment and storage medium
CN113377924A
Fraudulent webpage identification method, system and device and storage medium
CN113779956A