Data processing method, apparatus, device, and medium

CN116860705BActive Publication Date: 2026-09-15PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310797848.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2026-09-15
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

[0005]本申请实施例的一个目的旨在提供一种数据处理方法、装置、设备以及介质,旨在改善现有方案中对PDF文本进行标注时的效率较低的技术问题

Benefits of technology

[0020]In the above-described data processing method, apparatus, device, and medium, the solution involves obtaining a first virtual PDF text displayed on a web page from a first PDF text, and displaying the first virtual PDF text on the web page, wherein the first virtual PDF text has the same text content as the first PDF text. The solution also includes receiving a first text to be annotated from the first virtual PDF text displayed by a target user on the web page, obtaining the position information of the first text to be annotated, determining the graphic annotation information corresponding to the first text to be annotated based on the position information and the first text to be annotated, and defining the graphic annotation information as the graphic annotation information of a second text to be annotated in the first PDF text corresponding to the first text to be annotated. Therefore, the graphic annotation information corresponding to the first text to be annotated can be quickly determined based on the first text to be annotated by the user in the first virtual PDF text and its position information, allowing for rapid annotation of the first PDF text and improving the efficiency of PDF annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116860705B_ABST
    Figure CN116860705B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing and medical health, and discloses a data processing method, device, equipment and medium, comprising: obtaining a first virtual PDF text for displaying a first PDF text in a Web page, and displaying the first virtual PDF text through the Web page, wherein the first virtual PDF text is the same as the text content of the first PDF text; receiving a first to-be-labeled text in the first virtual PDF text displayed in the Web page by a target user; obtaining position information of the first to-be-labeled text; determining graphic labeling information corresponding to the first to-be-labeled text according to the position information and the first to-be-labeled text; and determining the graphic labeling information as graphic labeling information of a second to-be-labeled text corresponding to the first to-be-labeled text in the first PDF text. The efficiency of labeling the PDF text is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data processing and medical and health technologies, and in particular to a data processing method, apparatus, equipment and medium. Background Technology

[0002] With the development of science and technology and the progress of the times, people's quality of life has greatly improved in the past few decades, and people have begun to pursue a higher quality of life. To pursue a high quality of life, it is essential to maintain good health. Therefore, people have begun to learn about relevant medical and health knowledge. To learn about medical and health knowledge, one first needs to obtain relevant medical and health information, and then read and understand that information to acquire the corresponding knowledge.

[0003] In today's interconnected world, the primary way people obtain information is by searching for and downloading what they want online. Due to its portability and ease of creation, PDF (Portable Document Format) has become a widely used electronic document format. When reading PDFs, especially medical-related documents such as medical journals or promotional materials, users often highlight key points or areas they don't immediately understand. This allows them to better comprehend the text or search online for information to help them understand the parts they don't yet grasp. Therefore, annotating and commenting on PDFs has become increasingly common.

[0004] In existing solutions, users typically need to download and install the PDF annotation tool and learn how to use it to annotate text. This makes the process cumbersome and results in low efficiency when annotating PDF documents. Summary of the Invention

[0005] One objective of this application is to provide a data processing method, apparatus, device, and medium that aims to improve the low efficiency of annotation of PDF text in existing solutions.

[0006] In a first aspect, embodiments of this application provide a data processing method, the method comprising:

[0007] Obtain a first virtual PDF text that is displayed on a web page, and display the first virtual PDF text through a web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text;

[0008] Receive the first text to be annotated from the first virtual PDF text displayed by the target user on the web page;

[0009] Obtain the position information of the first text to be annotated;

[0010] Based on the location information and the first text to be annotated, determine the graphic annotation information corresponding to the first text to be annotated;

[0011] The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated.

[0012] Secondly, a data processing apparatus is provided, the data processing apparatus comprising:

[0013] The first acquisition unit is used to acquire a first virtual PDF text displayed on a web page, and to display the first virtual PDF text on the web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text.

[0014] The receiving unit is used to receive the first text to be annotated in the first virtual PDF text displayed by the target user on the web page;

[0015] The second acquisition unit is used to acquire the position information of the first text to be labeled;

[0016] The first determining unit is configured to determine the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated.

[0017] The second determining unit is used to determine the graphic annotation information as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated.

[0018] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described data processing method.

[0019] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described data processing method.

[0020] In the above-described data processing method, apparatus, device, and medium, the solution involves obtaining a first virtual PDF text displayed on a web page from a first PDF text, and displaying the first virtual PDF text on the web page, wherein the first virtual PDF text has the same text content as the first PDF text. The solution also includes receiving a first text to be annotated from the first virtual PDF text displayed by a target user on the web page, obtaining the position information of the first text to be annotated, determining the graphic annotation information corresponding to the first text to be annotated based on the position information and the first text to be annotated, and defining the graphic annotation information as the graphic annotation information of a second text to be annotated in the first PDF text corresponding to the first text to be annotated. Therefore, the graphic annotation information corresponding to the first text to be annotated can be quickly determined based on the first text to be annotated by the user in the first virtual PDF text and its position information, allowing for rapid annotation of the first PDF text and improving the efficiency of PDF annotation. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of an application environment for a data processing method according to an embodiment of the present invention;

[0023] Figure 2 A flowchart illustrating a data processing method provided in an embodiment of this application;

[0024] Figure 3 This is a schematic diagram of the coordinate system when performing segmentation processing on PDF text in one embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of a data processing device in one embodiment of the present invention;

[0026] Figure 5 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0029] The data processing method provided in this embodiment of the invention can be applied to, for example, Figure 1 In this application environment, the client communicates with the server via a network. In healthcare scenarios, for example, when a target user needs to process medical data, such as medical diagnostic data, medical assessment data, and medical risk prediction data, the user can perform preliminary viewing and annotation of specific medical data, such as annotating data related to specific diseases (like cancer). In this case, the user might view and annotate the medical data as a PDF text file. Alternatively, the user could annotate PDF texts from medical journals or medical promotional materials. In existing solutions, users typically download an installation package for a PDF text reader from the internet. After downloading, they need to unzip and install the package, and then read the corresponding user manual before they can use the PDF text reader to annotate the PDF text. This makes annotating PDF text cumbersome and inefficient.

[0030] To address the aforementioned problems, this invention provides a data processing method. In this method, a target user can upload PDF text to a web page (global wide area network page) to open and read it. When the target user annotates the PDF text, they can directly perform the corresponding operations on the web page to complete the annotation, improving the efficiency of PDF text annotation. For example, when a target user reads a PDF text of a medical journal (an academic PDF text), they can annotate the corresponding virtual PDF text through the web page. Specifically, when the target user reads a passage of text that interests them or that they don't understand, they can select the passage by dragging the pointer. The web page can then generate graphic annotation information corresponding to that passage, thereby annotating the selected text through the graphic annotation information.

[0031] Therefore, the server can assist the target user in annotating PDF text. In a specific example, when the target user uploads a first PDF text (e.g., a PDF text of a medical journal) to the server, the user selects the local save path for the first PDF text in the selection box. The server then extracts the first PDF text based on the local save path set by the user, performs virtual mapping on the first PDF text to obtain a first virtual PDF text, and displays the first virtual PDF text on the web page. The first virtual PDF text has the same text content as the first PDF text. This can be understood as the first virtual PDF text being a copy of the first PDF text that has been virtualized.

[0032] When the server receives a user's selection of a first text to be annotated from a first virtual PDF text displayed on the web page, the server analyzes this first text to obtain graphic annotation information. For example, if a user selects a text from an academic-type virtual PDF text, the server retrieves the abstract text from that text and the location information (including page numbers) of the selected text within the virtual PDF text. The server calculates the correlation between the selected text and the abstract text, obtaining a second abstract text based on the correlation. Based on the second abstract text and the location information, the server determines the graphic annotation information. Finally, the server displays the annotation information at the corresponding location within the virtual PDF text, completing the annotation of the PDF text and identifying the graphic annotation information as the graphic annotation information for the corresponding text in the academic-type PDF text.

[0033] The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0034] Please see Figure 2 As shown, Figure 2 This is a schematic flowchart of a data processing method provided in an embodiment of the present invention. Figure 2 As shown, the data processing method can be applied to the server side, including the following steps:

[0035] S201: Obtain a first virtual PDF text that is displayed on a web page, and display the first virtual PDF text through the web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text.

[0036] The target user can upload the first PDF file by clicking the "Upload File" button on the web page on the client, allowing the server to retrieve the first PDF file. When the target user clicks the "Upload File" button on the web page, a selection box will pop up, allowing the target user to set the local save path for the first PDF file to complete the upload.

[0037] Once the target user completes the local storage path setting for the first PDF text in the selection box, the web page will extract the first PDF text according to the local storage path set by the user in the selection box, then perform virtual mapping on the first PDF text to obtain the first virtual PDF text, and display the first virtual PDF text on the web page.

[0038] The fact that the first virtual PDF text and the first PDF text have the same text content can be understood as the same text information in the first virtual PDF text and the same text format in the first virtual PDF text. Specifically, the first virtual PDF text is the text obtained by copying the first PDF text and then virtualizing it.

[0039] If the target user wants to annotate a PDF text that has been deleted or a file for which the target user cannot find the PDF text, the first PDF text can be obtained from the server. Specifically, for example, the src tag information (source tag information) of the first PDF text in the server's HTML file (Hypertext Markup Language file) can be obtained. Based on the src tag information, the URL address (standard resource address) of the first PDF text can be obtained. The first PDF text can be obtained from the server using the obtained URL address of the first PDF text, and then the first virtual PDF text corresponding to the first PDF text can be obtained. The specific method of obtaining the first virtual PDF text can be referred to in the previous embodiment, which will not be repeated here.

[0040] The first virtual PDF text is displayed on a web page according to a pre-defined display method. Specifically, it can be displayed using a pre-defined web page background, a pre-defined web page layout, and a pre-defined display position for the PDF page numbers.

[0041] S202: Receive the first text to be annotated from the first virtual PDF text displayed by the target user on the web page.

[0042] It can receive the first position information of the cursor when the target user starts to select the first text to be annotated and the second position information of the cursor when the target user finishes to select the first text to be annotated, and obtain the text between the position corresponding to the first position information and the position corresponding to the second position information to obtain the first text to be annotated in the first virtual PDF text.

[0043] S203: Obtain the position information of the first text to be annotated.

[0044] like Figure 3As shown, this can be achieved by dividing the first virtual PDF text into a grid, using the bottom left corner of the last grid in the first virtual PDF text as the origin, the long side of the first virtual PDF text as the Y-axis, and the short side of the first virtual PDF text as the X-axis, with the long side perpendicular to the short side. The coordinate position information of the four vertices of the smallest unit grid corresponding to each character in the first text to be annotated is obtained. The coordinate position information of the four vertices of the smallest unit grid corresponding to each character in the first text to be annotated is determined as the elements in the coordinate position information set, and the above coordinate position information set is determined as the position information of the first text to be annotated.

[0045] S204: Based on the location information and the first text to be annotated, determine the graphic annotation information corresponding to the first text to be annotated.

[0046] Specifically, this can be achieved by obtaining the type information of the first PDF text. If the type information of the first PDF text is academic, then the first abstract text in the first virtual PDF text is obtained. If the correlation between the first abstract text and the first text to be annotated is higher than a preset correlation, then the second abstract text of the chapter where the first text to be annotated is located in the first virtual PDF text is obtained. Based on the position information corresponding to the second abstract text and the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated is determined. The graphic annotation information may include graphic annotation text information and graphic annotation attribute information.

[0047] If the type information of the first PDF text is promotional, then obtain the first keyword set of the first text to be annotated, obtain the first promotional text set based on the first keyword set, determine the reference keyword with the highest number of occurrences in the first virtual PDF text based on the number of times each keyword extracted from the first virtual PDF text appears, obtain the semantic information of the target keyword, determine the semantic information as the main information of the first virtual PDF text, obtain the first promotional text corresponding to the main information from the first promotional text set, and obtain the graphic annotation information corresponding to the first text to be annotated based on the position information of the first promotional text and the first text to be annotated.

[0048] Alternatively, K first associated texts corresponding to the first text to be labeled can be obtained by acquiring the first semantic information of the first text to be labeled. The first text theme information can be determined based on the K first associated texts and the first text to be labeled. The graphic annotation information corresponding to the first text to be labeled can be determined based on the first text theme information and the location information.

[0049] S205: The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated.

[0050] The method can be to obtain the second text to be annotated in the first PDF text that corresponds to the first text to be annotated, based on the position information and page number information of the first text to be annotated. The graphic annotation information corresponding to the first text to be annotated can then be used as the graphic annotation information for the second text to be annotated. Alternatively, after determining the graphic annotation information for the second text to be annotated in the first PDF text that corresponds to the first text to be annotated, the graphic annotation information can be displayed in the corresponding position in the first PDF text according to a pre-set style.

[0051] In one possible implementation, when obtaining the position information of the first text to be annotated, the position information of the first text to be annotated can be quickly obtained by acquiring the first coordinate position information of the first character and the second coordinate position information of the last character when the target user selects the first text to be annotated. This improves the efficiency of obtaining the position information of the first text to be annotated, thereby improving the efficiency of data processing. The method specifically includes:

[0052] A1. Obtain the first coordinate position information of the first character in the first text to be annotated selected by the target user;

[0053] This can be achieved by dividing the first virtual PDF text into a grid, establishing a Cartesian coordinate system with the bottom left corner of the last grid cell in the first virtual PDF text as the origin, the longer side of the first virtual PDF text as the Y-axis, and the shorter side of the first virtual PDF text as the X-axis. The length and width of the smallest unit in the grid are equal to the length and width of each character in the first virtual PDF text. The coordinate position information W of the four vertices of the smallest unit grid containing the first character in the first text to be annotated is then obtained. LS1 W RS1 W LX1 W RX1 Among them, W LS1 This corresponds to the coordinates of the top-left vertex of the smallest unit grid corresponding to the first character of the first text to be annotated, W. RS1 This corresponds to the coordinates of the top-right vertex of the smallest unit grid corresponding to the first character of the first text to be annotated, W. LX1 This corresponds to the coordinates of the bottom left vertex of the smallest unit grid corresponding to the first character of the first text to be annotated, W. RX1 This corresponds to the coordinates of the bottom right vertex of the smallest unit grid corresponding to the first character of the first text to be annotated, W. LS1 =(X L1 ,Y S1 W RS1 =(X R1 Y S1 WLX1 =(X L1 Y X1 W RX1 =(X R1 Y X1 ), W LS1 W RS1 W LX1 W RX1 The first coordinate position information of the first character in the first text to be annotated is determined. The X... L1 X R1 Y S1 Y X1 These represent the x and y coordinates of the four vertices of the smallest unit grid corresponding to the first character of the first text to be annotated.

[0054] A2 obtains the second coordinate position information of the last character in the first text to be annotated selected by the target user;

[0055] The method for obtaining the first coordinate position information of the first character in the first text to be annotated is used to obtain the coordinate position information of the four vertices of the smallest unit grid where the last character in the first text to be annotated is located. LSn W RSn W LXn W RXn Among them, W LSn =(X Ln ,Y Sn W RSn =(X Rn Y Sn W LXn =(X Ln Y Xn W RXn =(X Rn Y Xn ), W LSn W RSn W LXn W RXn The second coordinate position information of the last character in the first text to be annotated was determined.

[0056] A3. Obtain the coordinate position information of each character between the first position indicated by the first coordinate position information and the second position indicated by the second coordinate position information to obtain a set of coordinate position information;

[0057] For example: If the width of the first virtual PDF text is N, and the length of the smallest unit grid after dividing the first virtual PDF text into a grid is i and the height is h; then the coordinate information of the second character in the first text to be annotated selected by the target user is W. LS2 W RS2 WLX2 W RX2 Among them, W LS2 =(X L1 Y S1 W LX2 =(X L1 Y X1 W RS2 =(X R2 Y S2 W RX2 =(X R2 Y X2 ), determine X R2 =X R1 If +i>N is not true, then X R2 =X R1 +i、X L2 =X L1 +i、Y S2 =Y S1 Y X2 =Y X1 If true, then X L2 =0, X R2 =i, Y S2 =Y S1 -h、Y X2 =Y X1 -h; The method of obtaining the coordinate position information of the second character in the first text to be annotated selected by the target user is used to obtain the coordinate position information of all characters between the first and last characters in the first text to be annotated selected by the target user, thereby obtaining a set of coordinate position information.

[0058] A4. The coordinate position information, the first coordinate position information, and the second coordinate position information in the coordinate position information set are determined as the position information of the first text to be labeled.

[0059] The coordinate position information in the coordinate position information set, the first coordinate position information, and the second coordinate position information can be determined as the position information of the first text to be labeled.

[0060] By obtaining the coordinates of the first and last characters in the first text to be annotated by the target user, the positional information of the first text to be annotated can be obtained quickly and accurately, which improves the efficiency of obtaining the positional information of the first text to be annotated and improves the efficiency of data processing.

[0061] In one possible implementation, if the first type information of the first PDF text is academic, then the correlation between the first text to be annotated and the first abstract text in the first virtual PDF text is obtained to determine whether the first virtual PDF text contains a second abstract text. If the first virtual PDF text contains a second abstract text, then the graphic annotation information corresponding to the second abstract text and the first text to be annotated is used, specifically:

[0062] B1. Obtain the first type information of the first PDF text;

[0063] It can be that the first type information of the first text to be labeled is obtained through a classifier, which is a model that is pre-trained with a large amount of text semantic information and is used to classify the first text to be labeled.

[0064] B2. If the type indicated by the first type information is academic, then obtain the first summary text of the first virtual PDF text;

[0065] If the type indicated by the first type of information is academic, then use "Abstract" as the keyword to search in the first virtual PDF text, find the chapters with "Abstract" as the chapter title, determine all the text under the chapters with "Abstract" as the chapter title as the first abstract text, and obtain all the text under the chapters with "Abstract" as the chapter title to obtain the first abstract text.

[0066] B3. Obtain the target correlation degree between the first text to be annotated and the first summary text;

[0067] Obtain the delimiter between the first text to be annotated and the first summary text. Separate the first text to be annotated and the first summary text according to the delimiter, so as to obtain M short sentences and a second summary text set divided by the first text to be annotated, wherein the second summary text set includes Q second sub-summary texts.

[0068] M short sentences and Q second sub-summary texts are sequentially encoded using word vectors to obtain word vectors corresponding to the M short sentences and the Q second sub-summary texts, respectively. A first Transformer model is then used to sequentially calculate the knowledge representations corresponding to the M short sentences and the Q second sub-summary texts, respectively. The first Transformer model is used to obtain the knowledge representations of the Q second sub-summary texts. Based on the knowledge representations corresponding to the M short sentences and the Q second sub-summary texts, semantic information corresponding to the first text to be annotated and the first summary text are obtained. This can be achieved by weighted averaging of the knowledge representations corresponding to the M short sentences and weighted averaging of the knowledge representations corresponding to the Q second summary texts. Finally, the correlation between the first text to be annotated and the first summary text is calculated based on the semantic information of the first text to be annotated and the first summary text.

[0069] B4. If the target relevance is higher than the preset relevance, then obtain the second summary text of the chapter where the first text to be annotated is located in the first virtual PDF text;

[0070] The system determines whether the correlation between the first text to be annotated and the first summary text is higher than a preset correlation. If the correlation is higher, the system retrieves the second summary text of the chapter containing the first text to be annotated in the first virtual PDF text based on the page number information corresponding to the first text to be annotated. The preset correlation can be determined by empirical values ​​or historical data.

[0071] B5. Determine the graphic annotation information based on the second summary text and the location information.

[0072] The graphic annotation information includes graphic annotation attribute information and graphic annotation text information. The text information corresponding to the second summary text is determined as the graphic annotation text information in the graphic annotation information corresponding to the first text to be annotated. The graphic annotation attribute information in the graphic annotation information corresponding to the first text to be annotated is obtained based on the position information of the first text to be annotated. The graphic annotation information corresponding to the first text to be annotated is determined based on the graphic annotation text information and the graphic annotation attribute information in the graphic annotation information corresponding to the first text to be annotated. Here, graphic annotation text information refers to the text information in the graphic annotation information, specifically including the text information of the second summary text, etc.; graphic annotation attribute information refers to the appearance parameters of the graphic annotation box, specifically including the shape, size, and position information of the graphic annotation information in the first PDF text, etc. The shape and size of the graphic annotation information can be determined based on the coordinate position information of all characters in the first text to be annotated, specifically by connecting the coordinate positions of all characters in the first text to be annotated with straight lines; the shape and size of the closed shape formed by the straight lines are the shape and size of the graphic annotation information. Since graphic annotations are displayed in a specific area of ​​the PDF text, the position information of the graphic annotation box in the graphic annotation attribute information is a fixed coordinate position information in the first PDF text.

[0073] In this example, the similarity between the first text to be annotated and the first summary text is obtained. The similarity between the first text to be annotated and the first summary text is compared with a preset similarity to determine whether there is a second summary text in the first virtual PDF text that is in the chapter where the first text to be annotated is located. If there is a second summary text, the graphic annotation information is quickly obtained based on the position data corresponding to the second summary text and the first text to be annotated to quickly annotate the first PDF text.

[0074] In one possible implementation, if the second type of information of the first PDF text is promotional, then the keyword set in the first text to be annotated is obtained, and each first keyword in the first keyword set corresponds to a first promotional text to obtain a first promotional text set. The first promotional text is determined based on the promotional text set, and the graphic annotation information corresponding to the first promotional text and the first text to be annotated is obtained. The specific steps are as follows:

[0075] C1. Obtain the second type of information of the first text to be annotated;

[0076] The second type of information of the first text to be labeled can be obtained by a classifier. The classifier is trained with a large amount of text semantic information. The classifier can obtain the second type of information of the first text to be labeled by obtaining the semantic information corresponding to the first text to be labeled.

[0077] C2. If the type indicated by the second type of information is publicity, then extract keywords from the first text to be annotated to obtain a first set of keywords;

[0078] A general keyword extraction method can be used to extract keywords from the first text to be annotated, so as to obtain the first keyword set.

[0079] C3. Obtain the first promotional text corresponding to each first keyword in the first keyword set to obtain the first promotional text set;

[0080] Each first keyword in the first keyword set is used to match the first virtual PDF text to obtain the first promotional text in the first virtual PDF text that contains the first keywords in the first keyword set. The first promotional text corresponding to each first keyword in the first keyword set is determined as the first promotional text set.

[0081] C4. Obtain the topic information of the first virtual PDF text;

[0082] Multiple reference keywords can be obtained from a first virtual PDF text. The frequency of each reference keyword within the first virtual PDF text can be calculated, and the reference keyword with the highest frequency is identified as the target keyword of the first virtual PDF text. The semantic information of this target keyword is then extracted and used as the main body information of the first virtual PDF text. For example, obtaining multiple reference keywords from a first virtual PDF text could involve extracting keywords from a summary paragraph of the first virtual PDF text. Of course, other methods can also be used to obtain multiple keywords; this is merely an example.

[0083] C5. Determine the first promotional text corresponding to the theme information from the first set of promotional texts;

[0084] The matching degree between the first promotional text in the first promotional text set and the topic information is calculated. The first promotional text with the highest matching degree in the first promotional text set is obtained. The matching degree between the first promotional text in the first promotional text set and the topic information can be calculated by calculating the loss value between the semantic vector representation of the first promotional text in the first promotional text set and the semantic vector representation of the topic information. The smaller the loss value between the semantic vector representation of the first promotional text in the first promotional text set and the semantic vector representation of the topic information, the higher the matching degree between the first promotional text in the first promotional text set and the topic information, and vice versa.

[0085] C6. Based on the first promotional text and the location information, determine the graphic annotation information corresponding to the first text to be annotated.

[0086] For a detailed method for determining the graphic annotation information corresponding to the first text to be annotated based on the first promotional text and the location information, please refer to the method shown in B5, which will not be repeated here.

[0087] In this example, the graphic annotation information corresponding to the first text to be annotated is determined by obtaining the first keyword set corresponding to the first text to be annotated and the first promotional text corresponding to the first keyword set in the first virtual PDF text. This improves the speed of obtaining graphic annotation information and thus improves the efficiency of data processing.

[0088] In one possible implementation, the first text topic information can be determined by obtaining the first semantic information of the first text to be labeled, and the graphic annotation information of the first text to be labeled can be determined based on the first text topic information and the position information of the first text to be labeled. The specific steps are as follows:

[0089] D1. Obtain the first semantic information of the first text to be annotated;

[0090] Obtain the delimiters in the first text to be annotated, wherein the delimiters may be periods, semicolons, etc. in the first text to be annotated. Divide the first text to be annotated into M short sentences according to the delimiters in the first text to be annotated. Encode the M short sentences sequentially using the word vector method to obtain word vectors corresponding to the M short sentences. Use a second Transformer model to calculate the knowledge representations of the M short sentences sequentially to obtain M knowledge representations. The second Transformer model is used to obtain the knowledge representations of the M short sentences divided from the first text to be annotated. Average the M knowledge representations to obtain the semantic information of the first text to be annotated.

[0091] Based on the first semantic information, D2 determines K first associated texts corresponding to the first text to be annotated from the first virtual PDF text;

[0092] Obtain the delimiter of the first virtual PDF text. Divide the first virtual PDF text into N second texts based on the delimiter. Obtain the semantic information corresponding to each of the N second texts. Calculate the correlation between the semantic information of the first text to be annotated and the semantic information of each of the N second texts. Determine whether the correlation between the semantic information of the first text to be annotated and the semantic information of each of the N second texts exceeds a preset correlation threshold. If the correlation between the semantic information of the first text to be annotated and the semantic information of each of the N texts is greater than or equal to the preset correlation threshold, then the second text is determined to be associated with the first text to be annotated, and is identified as the first associated text. If the correlation between the semantic information of the first text to be annotated and the semantic information of each of the N texts is less than the preset correlation threshold, then the second text is considered not associated with the first text to be annotated, thus obtaining K first associated texts corresponding to the first text to be annotated. The preset correlation threshold can be determined by empirical values ​​or historical data.

[0093] D3. Determine the topic information of the first text based on the first text to be annotated and K first associated texts;

[0094] This can be achieved by extracting keywords from the first text to be annotated to obtain a second keyword set, extracting keywords from K first associated texts to obtain a third keyword set corresponding to each of the K first associated texts, determining a first keyword set based on the second keyword set and the third keyword sets corresponding to the K first associated texts, and determining the topic information of the first text based on the first keyword set.

[0095] D4. Based on the first text topic information and the location information, determine the graphic annotation information corresponding to the first text to be annotated.

[0096] For a detailed method of determining the graphic annotation information corresponding to the first text to be annotated based on the first text topic information and the location information, please refer to the method shown in step B5, which will not be repeated here.

[0097] In this example, by obtaining K first associated texts from among the N semantic representations corresponding to the first virtual PDF text, whose correlation with the semantic representation corresponding to the first text to be annotated is greater than or equal to a preset correlation, the topic information of the first text can be quickly determined based on the K first associated texts. The graphic annotation information can be determined based on the position information corresponding to the first text information and the first text to be annotated information, thereby quickly completing the annotation of the first PDF text and improving the efficiency of data processing.

[0098] In one possible implementation, the first text topic of the first text to be annotated can be determined by obtaining the first set of target keywords of the first text to be annotated. The specific steps are as follows:

[0099] E1. Extract keywords from the first text to be annotated to obtain a second keyword set, and extract keywords from K first associated texts to obtain a third keyword set corresponding to each of the K first associated texts;

[0100] This can be achieved by using a general text extraction method to extract keywords from the first text to be annotated, thereby obtaining a second keyword set; and by extracting keywords from each of the K first associated texts, thereby obtaining a third keyword set corresponding to each of the K first associated texts.

[0101] E2. Obtain a first target keyword set from the second keyword set and the K third keyword sets, wherein the keywords in the first target keyword set are keywords that exist in both the second keyword set and the K third keyword sets;

[0102] Take the intersection of the K sets of third keywords and merge them into a single keyword set to obtain a fourth keyword set. Take the intersection of the fourth keyword set and the second keyword set to obtain a first target keyword set. The first target keyword in the first keyword set is a keyword shared by the second keyword set and the K sets of third keywords.

[0103] E3. Determine the first text topic information based on the first target keyword in the first target keyword set;

[0104] This can be achieved by performing semantic calculations on each first target keyword in the first target keyword set to obtain the semantic information of each first target keyword in the first target keyword set, determining the semantic representation of each first target keyword in the first target keyword set as an element in the first semantic information set to obtain the first semantic information set, obtaining the first reference text topic information corresponding to each first semantic information in the first semantic information set based on the first semantic information set, and performing information fusion processing based on the first reference text topic information in the first reference text topic information set to obtain the first text topic information.

[0105] In this example, a first target keyword set can be obtained by acquiring the set of third keywords corresponding to the first text to be annotated and the K keywords shared by the K sets of third keywords corresponding to the first associated text. The semantic representation of the first target keyword in the first target keyword set is then compared with the semantic representation of the first text to be annotated to calculate the correlation, thus obtaining the first text's topic information. By extracting keywords and calculating their semantic representations, the need for semantic calculations on the first PDF text when obtaining the first text's topic information is avoided, reducing computational load and improving data processing efficiency.

[0106] In one possible implementation, each first target keyword in the first target keyword set is encoded separately. Semantic analysis is then performed on the first encoded vector obtained from the encoding of each first target keyword in the first target keyword set to obtain a first semantic information set. The first reference text topic information corresponding to each first semantic information in the first semantic information set is then obtained to obtain a first reference text topic information set. By quickly obtaining the first text topic information from the first reference text topic information set, the graphic annotation information corresponding to the first text to be annotated can be quickly obtained based on the first text topic information, thereby improving the efficiency of data processing.

[0107] F1. Encode each first target keyword in the first target keyword set to obtain a first encoding vector corresponding to each first target keyword;

[0108] Each first target keyword in the first target keyword set is encoded using the word vector method to obtain the word vector corresponding to each first target keyword in the first target keyword set. The word vector corresponding to each first target keyword in the first target keyword set is determined as the first encoding vector corresponding to each first target keyword.

[0109] F2. Perform semantic analysis on the first encoding vector corresponding to each first target keyword to obtain a second semantic information set. The second semantic information in the second semantic information set corresponds one-to-one with the first target keywords in the first target keyword set.

[0110] The third Transformer model is used to perform semantic analysis on the first encoding vector corresponding to each first target keyword to obtain the semantic information corresponding to each first target keyword. The third Transformer model is used to obtain the semantic information corresponding to each first target keyword and determine the semantic information corresponding to each first target keyword as an element in the second semantic information set to obtain the second semantic information set.

[0111] F3. Determine the first reference text topic information corresponding to each second semantic information in the second semantic information set, so as to obtain the first reference text topic information set;

[0112] Obtain the delimiter of the first virtual PDF text, divide the first virtual PDF text into N second texts according to the delimiter, obtain the semantic information corresponding to each of the N second texts to obtain the second text semantic information set, obtain the second text corresponding to the second text semantic information set with the highest matching degree of each second semantic information in the second text semantic information set, and determine the second text corresponding to the second text semantic information set with the highest matching degree of the second text semantic information set as the first reference text topic information corresponding to each second semantic information in the second semantic information set to obtain the first reference text topic information set.

[0113] F4. Perform information fusion processing on the first reference text topic information in the first reference text topic information set to obtain the first text topic information.

[0114] The first text topic information can be obtained by fusing the first reference text topic information in the first reference text topic information set using a general weighted average algorithm.

[0115] In this example, by obtaining the first encoding vector corresponding to each first target keyword in the first target keyword set, semantic analysis is performed on each first keyword based on the first encoding vector, to obtain a first semantic information set. Based on the first semantic information set, first reference text topic information corresponding one-to-one with each first semantic information in the first semantic information set is obtained, to obtain a first reference text topic information set. Based on the first reference text topic information set, the first text topic information is determined, thereby quickly obtaining the first text topic information. Based on the first text topic information and combined with the position information corresponding to the first text to be labeled, the graphic annotation information corresponding to the first text to be labeled can be quickly determined, improving the efficiency of data processing.

[0116] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0117] In one embodiment, a data processing apparatus is provided, the apparatus comprising: a data processing device corresponding one-to-one with the data processing method of the above embodiments. For example... Figure 3 As shown, the data processing device includes a first acquisition unit 301, a receiving unit 302, a second acquisition unit 303, a determining unit 304, and a second determining unit 305. Detailed descriptions of each functional module are as follows:

[0118] The first acquisition unit 301 is used to acquire a first virtual PDF text displayed on a web page, and to display the first virtual PDF text through a web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text.

[0119] The receiving unit 302 is used to receive the first text to be annotated in the first virtual PDF text displayed by the target user on the Web page;

[0120] The second acquisition unit 303 is used to acquire the position information of the first text to be labeled;

[0121] The determining unit 304 is used to determine the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated.

[0122] The second determining unit 305 is used to determine the graphic annotation information as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated.

[0123] In one possible implementation, the second acquisition unit 303 is specifically used for:

[0124] Obtain the first coordinate position information of the first character in the first text to be annotated selected by the target user;

[0125] Obtain the second coordinate position information of the last character in the first text to be annotated selected by the target user;

[0126] Obtain the coordinate position information of each character between the first position indicated by the first coordinate position information and the second position indicated by the second coordinate position information to obtain a set of coordinate position information;

[0127] The coordinate position information in the set of coordinate position information, the first coordinate position information, and the second coordinate position information are determined as the position information of the first text to be labeled.

[0128] In one possible implementation, the determining unit 304 is used for:

[0129] Obtain the first type information of the first PDF text;

[0130] If the type indicated by the first type information is academic, then obtain the first summary text of the first virtual PDF text;

[0131] Obtain the target correlation degree between the first text to be annotated and the first summary text;

[0132] If the target relevance is higher than the preset relevance, then obtain the second summary text of the chapter where the first text to be annotated is located in the first virtual PDF text;

[0133] The graphic annotation information is determined based on the second summary text and the location information.

[0134] In one possible implementation, the determining unit 304 is further used for:

[0135] Obtain the second type of information from the first text to be annotated;

[0136] If the type indicated by the second type of information is promotional, then keywords are extracted from the first text to be annotated to obtain a first set of keywords;

[0137] Obtain the first promotional text corresponding to each first keyword in the first keyword set to obtain the first promotional text set;

[0138] Obtain the topic information of the first virtual PDF text;

[0139] Determine the first promotional text corresponding to the theme information from the first set of promotional texts;

[0140] Based on the first promotional text and the location information, determine the graphic annotation information corresponding to the first text to be annotated.

[0141] In one possible implementation, the determining unit 304 is further used for:

[0142] Obtain the first semantic information of the first text to be annotated;

[0143] Based on the first semantic information, K first associated texts corresponding to the first text to be annotated are determined from the first virtual PDF text;

[0144] Based on the first text to be annotated and K first associated texts, determine the topic information of the first text;

[0145] Based on the first text topic information and the location information, determine the graphic annotation information corresponding to the first text to be annotated.

[0146] In one possible implementation, regarding the determination of the first text topic information based on the first text to be annotated and K first associated texts, the determining unit 304 is configured to:

[0147] Keyword extraction is performed on the first text to be annotated to obtain a second keyword set, and keyword extraction is performed on K first associated texts to obtain a third keyword set corresponding to each of the K first associated texts;

[0148] A first target keyword set is obtained from the second keyword set and the K third keyword sets, wherein the keywords in the first target keyword set are keywords that exist in both the second keyword set and the K third keyword sets;

[0149] The first text topic information is determined based on the first target keyword in the first target keyword set.

[0150] In one possible implementation, in determining the first text topic information based on the first target keyword in the first target keyword set, the determining unit 304 is configured to:

[0151] Each first target keyword in the first target keyword set is encoded to obtain a first encoding vector corresponding to each first target keyword.

[0152] Semantic analysis is performed on the first encoding vector corresponding to each first target keyword to obtain a second semantic information set, and the second semantic information in the second semantic information set corresponds one-to-one with the first target keywords in the first target keyword set;

[0153] Determine the first reference text topic information corresponding to each second semantic information in the second semantic information set, so as to obtain the first reference text topic information set;

[0154] Information fusion processing is performed on the first reference text topic information in the first reference text topic information set to obtain the first text topic information.

[0155] In this example, by obtaining the first virtual PDF text displayed on a web page, receiving the first text to be annotated in the first virtual PDF text displayed by the target user on the web page, the first text to be annotated in the first virtual PDF text is obtained. The location information of the first text to be annotated is obtained. Based on the location information and the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated is determined. The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated, so as to obtain the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated. Thus, by obtaining the location information of the first text to be annotated and combining it with the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated can be determined. The graphic annotation information can be determined as the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated, thereby realizing the rapid annotation of the first PDF text and improving the efficiency of data processing.

[0156] Specific limitations regarding the data processing device can be found in the limitations regarding the data processing method described above, and will not be repeated here. Each module in the aforementioned data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0157] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a data processing method on the server side.

[0158] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0159] Obtain a first virtual PDF text that is displayed on a web page, and display the first virtual PDF text through a web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text;

[0160] Receive the first text to be annotated from the first virtual PDF text displayed by the target user on the web page;

[0161] Obtain the position information of the first text to be annotated;

[0162] Based on the location information and the first text to be annotated, determine the graphic annotation information corresponding to the first text to be annotated;

[0163] The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated.

[0164] In this example, by obtaining the first virtual PDF text displayed on a web page, receiving the first text to be annotated in the first virtual PDF text displayed by the target user on the web page, the first text to be annotated in the first virtual PDF text is obtained. The location information of the first text to be annotated is obtained. Based on the location information and the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated is determined. The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated, so as to obtain the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated. Thus, by obtaining the location information of the first text to be annotated and combining it with the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated can be determined. The graphic annotation information can be determined as the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated, thereby realizing the rapid annotation of the first PDF text and improving the efficiency of data processing.

[0165] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0166] Obtain a first virtual PDF text that is displayed on a web page, and display the first virtual PDF text through a web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text;

[0167] Receive the first text to be annotated from the first virtual PDF text displayed by the target user on the web page;

[0168] Obtain the position information of the first text to be annotated;

[0169] Based on the location information and the first text to be annotated, determine the graphic annotation information corresponding to the first text to be annotated;

[0170] The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated.

[0171] In this example, by obtaining the first virtual PDF text displayed on a web page, receiving the first text to be annotated in the first virtual PDF text displayed by the target user on the web page, the first text to be annotated in the first virtual PDF text is obtained. The location information of the first text to be annotated is obtained. Based on the location information and the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated is determined. The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated, so as to obtain the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated. Thus, by obtaining the location information of the first text to be annotated and combining it with the first text to be annotated, the graphic annotation information corresponding to the first text to be annotated can be determined. The graphic annotation information can be determined as the graphic annotation information of the second text to be annotated in the first PDF text corresponding to the first text to be annotated, thereby realizing the rapid annotation of the first PDF text and improving the efficiency of data processing.

[0172] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0173] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0174] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0175] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A data processing method, characterized in that, The method includes: Obtain a first virtual PDF text that is displayed on a web page, and display the first virtual PDF text through a web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text; Receive the first text to be annotated from the first virtual PDF text displayed by the target user on the web page; Obtain the position information of the first text to be annotated; Based on the location information and the first text to be annotated, determine the graphic annotation information corresponding to the first text to be annotated; The graphic annotation information is determined as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated; The step of determining the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated includes: Obtain the first type information of the first PDF text; If the type indicated by the first type information is academic, then obtain the first summary text of the first virtual PDF text; Obtain the target correlation degree between the first text to be annotated and the first summary text; If the target relevance is higher than the preset relevance, then obtain the second summary text of the chapter where the first text to be annotated is located in the first virtual PDF text; The graphic annotation information is determined based on the second summary text and the location information; Alternatively, determining the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated includes: Obtain the second type of information from the first text to be annotated; If the type indicated by the second type of information is promotional, then keywords are extracted from the first text to be annotated to obtain a first set of keywords; Obtain the first promotional text corresponding to each first keyword in the first keyword set to obtain the first promotional text set; Obtain the topic information of the first virtual PDF text; Determine the first promotional text corresponding to the theme information from the first set of promotional texts; Based on the first promotional text and the location information, determine the graphic annotation information corresponding to the first text to be annotated; Alternatively, determining the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated includes: Obtain the first semantic information of the first text to be annotated; Based on the first semantic information, K first associated texts corresponding to the first text to be annotated are determined from the first virtual PDF text; Based on the first text to be annotated and K first associated texts, determine the topic information of the first text; Based on the first text topic information and the location information, determine the graphic annotation information corresponding to the first text to be annotated.

2. The data processing method as described in claim 1, characterized in that, Obtain the position information of the first text to be annotated, including: Obtain the first coordinate position information of the first character in the first text to be annotated selected by the target user; Obtain the second coordinate position information of the last character in the first text to be annotated selected by the target user; Obtain the coordinate position information of each character between the first position indicated by the first coordinate position information and the second position indicated by the second coordinate position information to obtain a set of coordinate position information; The coordinate position information in the set of coordinate position information, the first coordinate position information, and the second coordinate position information are determined as the position information of the first text to be labeled.

3. The data processing method as described in claim 1, characterized in that, The step of determining the first text topic information based on the first text to be labeled and K first associated texts includes: Keyword extraction is performed on the first text to be annotated to obtain a second keyword set, and keyword extraction is performed on K first associated texts to obtain a third keyword set corresponding to each of the K first associated texts; A first target keyword set is obtained from the second keyword set and the K third keyword sets, wherein the keywords in the first target keyword set are keywords that exist in both the second keyword set and the K third keyword sets; The first text topic information is determined based on the first target keyword in the first target keyword set.

4. The data processing method as described in claim 3, characterized in that, Determining the first text topic information based on the first target keyword in the first target keyword set includes: Each first target keyword in the first target keyword set is encoded to obtain a first encoding vector corresponding to each first target keyword. Semantic analysis is performed on the first encoding vector corresponding to each first target keyword to obtain a second semantic information set, and the second semantic information in the second semantic information set corresponds one-to-one with the first target keywords in the first target keyword set; Determine the first reference text topic information corresponding to each second semantic information in the second semantic information set, so as to obtain the first reference text topic information set; Information fusion processing is performed on the first reference text topic information in the first reference text topic information set to obtain the first text topic information.

5. A data processing apparatus, characterized in that, The data processing device includes: The first acquisition unit is used to acquire a first virtual PDF text displayed on a web page, and to display the first virtual PDF text on the web page, wherein the text content of the first virtual PDF text is the same as that of the first PDF text. The receiving unit is used to receive the first text to be annotated in the first virtual PDF text displayed by the target user on the web page; The second acquisition unit is used to acquire the position information of the first text to be labeled; The first determining unit is configured to determine the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated. The second determining unit is used to determine the graphic annotation information as the graphic annotation information of the second text to be annotated in the first PDF text that corresponds to the first text to be annotated; The step of determining the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated includes: Obtain the first type information of the first PDF text; If the type indicated by the first type information is academic, then obtain the first summary text of the first virtual PDF text; Obtain the target correlation degree between the first text to be annotated and the first summary text; If the target relevance is higher than the preset relevance, then obtain the second summary text of the chapter where the first text to be annotated is located in the first virtual PDF text; The graphic annotation information is determined based on the second summary text and the location information; Alternatively, determining the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated includes: Obtain the second type of information from the first text to be annotated; If the type indicated by the second type of information is promotional, then keywords are extracted from the first text to be annotated to obtain a first set of keywords; Obtain the first promotional text corresponding to each first keyword in the first keyword set to obtain the first promotional text set; Obtain the topic information of the first virtual PDF text; Determine the first promotional text corresponding to the theme information from the first set of promotional texts; Based on the first promotional text and the location information, determine the graphic annotation information corresponding to the first text to be annotated; Alternatively, determining the graphic annotation information corresponding to the first text to be annotated based on the location information and the first text to be annotated includes: Obtain the first semantic information of the first text to be annotated; Based on the first semantic information, K first associated texts corresponding to the first text to be annotated are determined from the first virtual PDF text; Based on the first text to be annotated and K first associated texts, determine the topic information of the first text; Based on the first text topic information and the location information, determine the graphic annotation information corresponding to the first text to be annotated.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the data processing method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the data processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method, device and equipment for labeling PDF document

    CN113221563A

  • Text information reading method and device and terminal

    CN116341489A