Code extraction method and device, electronic equipment and medium

By generating and formatting PDF training samples, training a large language model based on Transformer, solving the problem of indistinguishable code and text in PDF files, and achieving the effect of quickly and accurately extracting code from PDF files.

CN119961674APending Publication Date: 2025-05-09BEIJING WUWEN CORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510043287.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately extract code from PDF files, especially because the code and text in PDF files do not have special logos, making it difficult for computers to distinguish and extract code information.

Method used

By generating multi-page PDF training samples, using code snippets and text snippets for formatting and random mixing, converting them into the second format training samples, training a large language model based on the Transformer structure, thereby identifying and extracting the code in the PDF file.

Benefits of technology

It realizes the extraction of code from PDF files accurately, quickly and conveniently without OCR, solving the problems of difficulty and inefficiency in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961674A_ABST
    Figure CN119961674A_ABST
Patent Text Reader

Abstract

The invention provides a code extraction method and device, electronic equipment and a medium. A method for extracting codes from a portable document format (PDF) file is characterized by comprising the following steps of: generating a plurality of pages of PDF training samples by utilizing a plurality of code fragments and a plurality of character fragments; converting each page of PDF training samples in the multiple pages of PDF training samples into training samples in a second format; training a code extraction model by using the second format training sample; and extracting codes from the PDF file by utilizing the trained code extraction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text processing technology, and in particular to a method, device, electronic device and medium for extracting codes. Background Art

[0002] With the rapid development of information technology, Portable Document Format (PDF), as a cross-platform and cross-application document format, has been widely used in various text processing scenarios. PDF files are not only highly readable and printable, but also able to maintain the original format and layout of the file, thereby ensuring the consistency and integrity of the file. However, this feature also increases the difficulty of parsing and extracting information from PDF files. Tools or applications based on Optical Character Recognition (OCR) can extract text information from PDF files to a certain extent, but their processing speed is slow and accuracy cannot be guaranteed.

[0003] On the other hand, when technicians engaged in software development process PDF files, they sometimes need to extract the codes for use. "Extracting" codes may include, for example, identifying the codes, copying the codes as editable text, separating the codes from other texts such as text, and saving the codes to other files. Since all texts in PDF files are stored in lines, and code snippets do not have special marks or identifiers, even if the text in the PDF file is extracted by OCR, it is impossible to directly distinguish which lines are codes and which lines are ordinary text. Most of the existing PDF file processing tools and applications on the market also focus on basic functions such as viewing, editing and printing PDF files, and there is no perfect solution for the need to accurately extract codes from PDF files. This not only limits the application of PDF files in code management and sharing, but also increases the difficulty and cost of developers to obtain code information from PDF files. Summary of the invention

[0004] The present application provides a method, device, electronic device and medium for extracting codes from a portable document format PDF file, so that codes can be accurately, quickly and conveniently extracted from a PDF file without OCR.

[0005] In order to achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, a method for extracting code from a portable document format PDF file is provided, the method comprising: generating a multi-page PDF training sample using a plurality of code snippets and a plurality of text snippets; converting each page of the multi-page PDF training sample into a second format training sample; training a code extraction model using the second format training sample; and extracting code from the PDF file using the trained code extraction model.

[0007] Optionally, each page of the multiple pages of PDF training samples includes multiple lines, and each line is a code line or a text line.

[0008] Optionally, the generating of multiple pages of PDF training samples using multiple code snippets and multiple text snippets includes: generating multiple code snippets and multiple text snippets; formatting the generated multiple code snippets and multiple text snippets into multiple code blocks and multiple text blocks, wherein each of the multiple code blocks includes one or more code lines, and each of the multiple text blocks includes one or more text lines; and randomly mixing the multiple code blocks and the multiple text blocks so that code blocks and text blocks exist alternately in each page of PDF training samples.

[0009] Optionally, converting each page of the multiple pages of PDF training samples into a second format training sample includes: assigning a label to each of the multiple lines of each page of the PDF training sample, the label indicating whether the line is a text line or a code line; obtaining the coordinates of each of the multiple lines; converting each page of the PDF training sample into a second format training sample including the text or code of each of the multiple lines, the label of each line and the coordinates of each line.

[0010] Optionally, extracting code from a PDF file using the trained code extraction model includes: enabling the trained code extraction model to infer the PDF file to obtain a label for each line in the PDF file; and identifying code blocks and text blocks in the PDF file based on the label of each line.

[0011] Optionally, after identifying the code blocks and text blocks in the PDF file, the method further includes: enabling the trained code extraction model to infer the code blocks and text blocks in the PDF file to obtain semantic relationships between the code blocks and text blocks; and rearranging the code blocks and text blocks according to the semantic relationships to obtain semantically continuous code fragments and text fragments.

[0012] Optionally, the code extraction model is a large language model based on a Transformer structure.

[0013] In a second aspect, a device for extracting code from a portable document format PDF file is provided, the device comprising: a sample generation module, used to generate a multi-page PDF training sample using a plurality of code snippets and a plurality of text snippets; a format conversion module, used to convert each page of the multi-page PDF training sample into a second format training sample; a model training module, used to train a code extraction model using the second format training sample; and a code extraction module, used to extract code from a PDF file using the trained code extraction model.

[0014] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for extracting code from a portable document format PDF file as described in an embodiment of the present application.

[0015] In a fourth aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute the method for extracting code from a portable document format PDF file as described in an embodiment of the present application.

[0016] The method, device, electronic device and medium for extracting codes from a portable document format PDF file provided in the present application can accurately, quickly and conveniently extract codes from a PDF file without using OCR. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A schematic diagram of an example PDF file provided in an embodiment of the present application;

[0018] Figure 2 A flowchart of a method for extracting codes from a PDF file provided in an embodiment of the present application;

[0019] Figure 3 A schematic diagram of a page of PDF training samples provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of the structure of the code extraction model provided in the embodiment of the present application;

[0021] Figure 5 A schematic diagram of a PDF file rearranged by a code extraction model provided in an embodiment of the present application;

[0022] Figure 6 A schematic diagram of the structure of a device for extracting codes from a PDF file provided in an embodiment of the present application;

[0023] Figure 7 A schematic diagram of the structure of a method or device that can implement an embodiment of the present application or realize an embodiment of the present application is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0024] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, a detailed description of well-known model training methods is omitted to prevent unnecessary details from obstructing the description of the present application.

[0025] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and appended claims of the present application, the singular expressions "a", "said", "above", and "the" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise.

[0026] With the rapid advancement of information technology, PDF files, as a cross-platform and cross-application document format, are widely used in text processing. Its high readability and printability as well as its ability to maintain the original format and layout of the file ensure the consistency and integrity of the file. However, this feature also increases the difficulty of parsing and extracting information from PDF files.

[0027] In PDF files, text, images, codes and other information are stored in a specific way. For example, text and code are usually stored in the form of coordinates in lines, and code snippets do not have special marks that can distinguish them from ordinary text. Figure 1 For example. Figure 1 A schematic diagram of an example PDF file provided in an embodiment of the present application. Figure 1 In the example PDF file, a code snippet is inserted into a text snippet, and both the text and the code are presented in the form of lines. Although humans can distinguish which lines are codes and which lines are text when reading a PDF file, it is difficult for a computer to parse and extract the code from a PDF file because both code lines and text lines are stored in line coordinates without distinguishing whether the content is code or text. In the related art, for example, all text in a PDF file can be extracted through OCR, but due to the above-mentioned characteristics of a PDF file, it is still impossible to quickly and accurately extract the code from it.

[0028] In view of this, the embodiments of the present application propose a method, device, electronic device and medium for extracting codes from a portable document format PDF file, so that codes can be accurately, quickly and conveniently extracted from a PDF file without using OCR.

[0029] The following first introduces the method for extracting code from a PDF file provided in an embodiment of the present application, see Figure 2 . Figure 2 The following is a flow chart of a method for extracting code from a PDF file provided in an embodiment of the present application. Figure 2 The illustrated method 20 for extracting codes from a PDF file comprises the following steps.

[0030] S201: Generate a multi-page PDF training sample using a plurality of code snippets and a plurality of text snippets.

[0031] A code snippet can be understood as one or more lines of code that are semantically continuous. Similarly, a text snippet can be understood as one or more lines of text that are semantically continuous. In this case, each page of the generated multi-page PDF training sample can include multiple lines, each of which is a code line or a text line. In addition to the semantic continuity of code lines in a code snippet and the semantic continuity of text lines in a text snippet, code snippets and text snippets can also be semantically continuous. For example, it can be text that introduces a certain algorithm or function and the code that implements the algorithm or function, or it can be a code example and text that explains the code example.

[0032] In some embodiments, generating multiple pages of PDF training samples using multiple code snippets and multiple text snippets may, for example, include: generating multiple code snippets and multiple text snippets; formatting the generated multiple code snippets and multiple text snippets into multiple code blocks and multiple text blocks, wherein each of the multiple code blocks includes one or more code lines, and each of the multiple text blocks includes one or more text lines; and randomly mixing the multiple code blocks and the multiple text blocks so that the code blocks and the text blocks exist alternately in each page of the PDF training sample.

[0033] The generation of multiple code snippets and multiple text snippets can be performed, for example, by crawling online articles, automatically generating using artificial intelligence tools, extracting from existing databases, etc. Formatting the generated multiple code snippets and multiple text snippets can include, for example, splitting the multiple code snippets and multiple text snippets into code blocks including a specific number of code lines and text blocks including a specific number of text lines. The specific number is not specifically limited, and preferably can be, for example, 3 to 50 lines. The number of code lines included in different code blocks can be different, and the number of text lines included in different text blocks can also be different. Formatting the generated multiple code snippets and multiple text snippets in this way can format the multiple code snippets and multiple text snippets into multiple code blocks and multiple text blocks, wherein each code block includes one or more code lines, and each text block includes one or more text lines. The obtained multiple code blocks and multiple text blocks are randomly mixed, so that each page of PDF training samples in which code blocks and text blocks alternately exist can be obtained. "Random mixing" can be understood, for example, as randomly rearranging code blocks and text blocks without changing the semantic order between code blocks and without changing the semantic order between text blocks, but only changing the semantic order between code blocks and text blocks.

[0034] For example, capital letters A, B, C, etc. are used to represent text snippets, and lowercase letters a, b, c, etc. are used to represent code snippets. The multiple text snippets and multiple code snippets before formatting are, for example, AaBCbc, indicating that these text snippets and code snippets are text snippet A, code snippet a, text snippet B, text snippet C, code snippet b, and code snippet c in semantic order. Assume that text snippet A is formatted into text blocks A1 and A2, code snippet a remains a after formatting, text snippet B is formatted into text blocks B1, B2, and B3, text snippet C is formatted into text blocks C1 and C2, code snippet b is formatted into code blocks b1 and b2, and code snippet c is formatted into code blocks c1 and c2. Therefore, the multiple text blocks and multiple code blocks after formatting are A1A2aB1B2B3C1C2b1b2c1c2. By randomly mixing these text blocks and code blocks, for example, we can obtain a page of PDF training samples with the following structure: A1aA2b1B1B2b2B3C1c1c2C2.

[0035] For an example of a one-page PDF training sample generated in the above way, see Figure 3 . Figure 3 This is a schematic diagram of a page of PDF training samples provided in an embodiment of the present application. Figure 3In the PDF training samples in , text blocks including multiple text lines and code blocks including multiple code lines exist alternately. It can be seen that the text blocks formatted by semantically continuous text fragments are interspersed with code blocks formatted by semantically continuous code fragments. It can also be understood that a code fragment is split into multiple code blocks and the multiple code blocks are respectively inserted into multiple positions of a text fragment, thereby splitting the text fragment into multiple text blocks, and obtaining PDF training samples in which code blocks and text blocks exist alternately.

[0036] S202: Convert each page of the generated multiple pages of PDF training samples into a second format training sample.

[0037] In some embodiments, converting each page of the generated multiple pages of PDF training samples into a second format training sample may, for example, include: assigning a label to each of the multiple lines of each page of the PDF training sample, the label indicating whether the line is a text line or a code line; obtaining the coordinates of each of the multiple lines; and converting each page of the PDF training sample into a second format training sample including text or code for each of the multiple lines, a label for each line, and the coordinates of each line.

[0038] The labels assigned to each of the multiple rows of each PDF training sample page can be found, for example, Figure 3 The label of each line on the right side of the text in the PDF file is "text", and the label "code" indicates that the line is a text line, and the label "code" indicates that the line is a code line. It should be understood that the form of the label is not limited to this, as long as it can indicate whether each line is a text line or a code line. The method for obtaining the coordinates of each line in the multiple lines is not particularly limited, and any known method for obtaining the text line coordinates of PDF files can be used. The acquired coordinates can be, for example, a four-dimensional array [x0, y0, x1, y1], where x0 and y0 represent the horizontal and vertical coordinates of the starting position of the text box of the text line, and x1 and y1 represent the length and width of the text box of the text line. The second format training sample can be, for example, a training sample in a lightweight data exchange format, such as a training sample in JSON (JavaScript Object Notation) format. The second format is also not particularly limited, as long as it can include the text or code of each line in the multiple lines of the PDF training sample, the label of each line, and the coordinates of each line.

[0039] A specific example of a second format training sample when the second format is JSON is given below.

[0040] Input: Each line is a row of training samples, including row content and row coordinates

[0041]

[0042] S203: Train the code extraction model using the second format training samples.

[0043] In some embodiments, the code extraction model may be, for example, a large language model based on a Transformer structure. In such a code extraction model, for example, the encoder and decoder of the Transformer structure may be set to N layers respectively, and, for example, multiple graphics processing units (GPUs) may be used simultaneously to train the code extraction model in parallel. Figure 4 A schematic diagram of the structure of the code extraction model provided in the embodiment of the present application. Figure 4 In the model, the encoder part contains Nx layers (for example, it can be 6 layers or more), each layer performs operations such as multi-head attention mechanism (Multi-Head Attention) and feed-forward neural network (Feed-Forward Neural Network), and stabilizes the training process through residual connection (Add) and layer normalization (Norm). The decoder part also contains Nx layers, which has a similar structure to the encoder, but also contains an additional masked multi-head attention mechanism (Masked Multi-Head Attention) to utilize the output information of the encoder. Increasing the number of layers of the encoder and decoder is usually aimed at increasing the depth and complexity of the model, so as to better capture the characteristics of the input data and learn complex function mappings, thereby improving the performance of the model. The training samples of the code extraction model are second format training samples, so that the code extraction model can learn the characteristics of text lines and code lines, so that it can infer which lines in a given PDF file are codes, and then perform code extraction.

[0044] S204: Extracting codes from the PDF file using the trained code extraction model.

[0045] In some embodiments, extracting code from a PDF file using a trained code extraction model may include, for example: enabling the trained code extraction model to infer the PDF file to obtain a label for each line in the PDF file; and identifying code blocks and text blocks in the PDF file based on the label for each line. As described above, since the code extraction model is trained by a second format training sample including text or code for each line of the PDF training sample, a label for each line, and coordinates for each line, it is possible to learn the features of text lines and code lines, thereby being able to infer which lines in the PDF file are code and which lines are text, that is, to obtain a label for each line in the PDF file. Subsequently, based on the label for each line, one or more continuous lines labeled as code are identified as code blocks, and one or more continuous lines labeled as text are identified as text blocks. Thus, the trained code extraction model is able to extract code from a PDF file.

[0046] In some cases, code blocks or text blocks in a PDF file may not be semantically continuous. For example, Figure 1 In the PDF file shown in , the code in the box is inserted into a paragraph of text, so that the paragraph of text is separated by the code and becomes discontinuous. In these cases, after identifying the code blocks and text blocks in the PDF file, for example, the trained code extraction model can be used to infer the code blocks and text blocks in the PDF file to obtain the semantic relationship between the code blocks and the text blocks, and the code blocks and the text blocks are rearranged according to the semantic relationship to obtain semantically continuous code fragments and text fragments. The rearrangement can be understood as exchanging the order of code blocks and text blocks, so that the code blocks that may have been separated before form semantically continuous code fragments, and the text blocks that may have been separated before form semantically continuous text fragments, and the code fragments and text fragments can also be semantically continuous. For example, for Figure 1 For the PDF file shown in the figure, the code extraction model can identify 5 text blocks (5 text blocks separated by blank lines or natural paragraphs) and 5 code blocks (5 code blocks separated by blank lines), but from a semantic point of view, the 5 code blocks should be arranged between the first text block and the second text block. In this case, the trained code extraction model can be used to infer the 5 code blocks and 5 text blocks in the PDF file to obtain the semantic relationship between these code blocks and text blocks, and rearrange these code blocks and text blocks into semantically continuous code segments and text segments according to the semantic relationship. The rearranged PDF file is, for example, Figure 5 shown. Figure 5 A schematic diagram of a PDF file rearranged by a code extraction model provided in an embodiment of the present application includes two text segments and one code segment, and the text segments and the code segments are semantically continuous.

[0047] After rearrangement, for example, the code snippets and text snippets can be rearranged so that both the text and the code have a good visual effect. The text snippets can be rearranged according to the following text snippet layout rules, and the code snippets can be rearranged according to the following code snippet layout rules.

[0048] Examples of text snippet layout rules include:

[0049] 1) Perform hard line breaks: Hard line breaks are a crucial step in typesetting rules. It involves inserting line breaks at appropriate locations to ensure that the text is formatted as expected when displayed or printed. Hard line breaks usually occur at the end of a sentence, the end of a paragraph, or when a line of text reaches a predetermined width limit. Through hard line breaks, the layout of the text can be effectively controlled, making it easier to read and understand.

[0050] 2) Perform line splicing: Line splicing is another important step in the typesetting process. In some cases, the text may be split into multiple incomplete lines for various reasons, which will affect the overall reading experience of the text. The purpose of line splicing is to reassemble these incomplete lines into complete sentences or paragraphs to make the text more visually coherent and complete. This usually involves logical analysis of the text to determine which lines should be spliced ​​together.

[0051] 3) Filtering space symbols: Filtering space symbols is an important step in typesetting rules. In the text, extra spaces, tabs or other unnecessary symbols may cause typesetting confusion and affect the readability of the text. Therefore, it is necessary to filter these space symbols, delete extra spaces, and unify the size and number of spaces to ensure the neatness and consistency of the text. By filtering space symbols, the text can be made visually clearer and more standardized.

[0052] For example, code snippet typography rules may include:

[0053] 1) Keep spaces: In code paragraphs, spaces usually have specific meanings and uses. They may be separators between operators and operands, or they may be tools used for indentation to indicate code blocks or levels. Therefore, when typesetting the code, you must ensure that spaces are properly preserved. This means that spaces cannot be deleted or added at will, otherwise the meaning of the code may change or cause syntax errors. Keeping appropriate spaces helps the readability and maintainability of the code, making it easier for other developers to understand and modify the code.

[0054] 2) Convert the format of lightweight markup language to facilitate downstream applications: Lightweight markup language allows people to write documents in a plain text format that is easy to read and write. Converting code paragraphs to lightweight markup language format can make them easier to use in downstream applications. This conversion usually involves highlighting the code, adding comments, inserting code blocks, etc., so that the code appears clear and beautiful in the lightweight markup language document. Through the conversion of the lightweight markup language format, downstream applications can more conveniently display, edit or parse the code, which improves the accessibility and reusability of the code. In addition, the wide acceptance and compatibility of the lightweight markup language format also makes code paragraphs easier to share and transfer between multiple platforms and tools.

[0055] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0056] The above describes the method for extracting code from a PDF file provided in an embodiment of the present application. The following describes the device for extracting code from a PDF file provided in an embodiment of the present application. Figure 6 . Figure 6 The schematic diagram of the structure of the device for extracting code from PDF file provided by the embodiment of the present application. For the convenience of explanation, only the part related to the embodiment of the present application is shown. Figure 6 As shown, the device 60 for extracting code from a PDF file includes: a sample generation module 601, a format conversion module 602, a model training module 603 and a code extraction module 604.

[0057] The sample generation module 601 is used to generate multiple pages of PDF training samples using multiple code snippets and multiple text snippets.

[0058] The format conversion module 602 is used to convert each page of the generated multiple pages of PDF training samples into a second format training sample.

[0059] The model training module 603 is used to train the code extraction model using the second format training samples.

[0060] The code extraction module 604 is used to extract codes from the PDF file using the trained code extraction model.

[0061] The specific functions of each module can be found in the description of the method for extracting codes from a portable document format PDF file, and will not be described in detail here.

[0062] The method and device for extracting codes from a PDF file provided in the present application can accurately, quickly and conveniently extract codes from a PDF file without using OCR.

[0063] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0064] An embodiment of the present application also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for extracting code from a portable document format PDF file as described in an embodiment of the present application.

[0065] On the other hand, an embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute a method for extracting code from a portable document format PDF file as described in an embodiment of the present application.

[0066] An embodiment of the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the method for extracting code from a portable document format PDF file as described in an embodiment of the present application.

[0067] The embodiment of the present application also provides a computer program, which includes program instructions. When the program instructions are executed by a computer, the computer executes the method for extracting code from a portable document format PDF file as described in the embodiment of the present application.

[0068] Figure 7A schematic diagram of a method or device for implementing an embodiment of the present application provided for an embodiment of the present application may include more or fewer devices than shown in the figure in some embodiments. In some embodiments, it can be implemented using a single or multiple devices. In some embodiments, it can be implemented using cloud or distributed devices.

[0069] like Figure 7 As shown, the device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to the program and / or data stored in the read-only memory (ROM) 1002 or the program and / or data loaded from the storage part 1008 to the random access memory (RAM) 1003. Processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, processor 1001 can include a general main processor and one or more special coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. In RAM 1003, various programs and data required for the operation of device 1000 are also stored. Processor 1001, ROM 1002 and RAM 1003 are connected to each other via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0070] The processor and the memory are used together to execute the program stored in the memory. When the program is executed by the computer, the methods, steps or functions described in the above embodiments can be implemented.

[0071] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed, so that a computer program read therefrom is installed into the storage section 1008 as needed. Figure 7 Only some components are schematically shown in the figure, which does not mean that the device 1000 only includes Figure 7 Components shown.

[0072] The systems, devices, modules or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server or a combination thereof.

[0073] The storage media of the embodiments of the present application include permanent and non-permanent, removable and non-removable items that can be used to store information by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0074] The methods, programs, systems, devices, etc. of the embodiments of the present application can be executed or implemented in a single or multiple networked computers, or can be practiced in a distributed computing environment. In the embodiments of the present application, in these distributed computing environments, tasks can be performed by remote processing devices connected through a communication network.

[0075] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, those skilled in the art can imagine that the implementation of the functional modules / units or controllers and related method steps described in the above embodiments can be implemented in software, hardware or a combination of software / hardware.

[0076] Unless explicitly stated, the actions or steps of the methods, programs, and embodiments of the present application do not have to be performed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0077] In this article, multiple embodiments of the present application are described, but for the sake of simplicity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the various embodiments may be omitted. In this article, "one embodiment", "some embodiments", "example", "specific example", or "some examples" are intended to be applicable to at least one embodiment or example according to the present application, rather than all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. In the absence of mutual contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0078] The exemplary systems and methods of the present application have been specifically shown and described with reference to the above embodiments, which are merely examples of the best modes for implementing the present systems and methods. It will be appreciated by those skilled in the art that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present application as defined in the appended claims.

Claims

1. A method for extracting code from a portable document format PDF file, characterized in that The method comprises: Generate multi-page PDF training samples using multiple code snippets and multiple text snippets; Convert each page of the multiple pages of PDF training samples into a second format training sample; Using the second format training sample to train the code extraction model; The trained code extraction model is used to extract codes from the PDF file.

2. The method according to claim 1, characterized in that Each page of the PDF training samples in the multiple pages of PDF training samples includes multiple lines, and each line is a code line or a text line.

3. The method according to claim 2, characterized in that The method of generating a multi-page PDF training sample by using a plurality of code snippets and a plurality of text snippets comprises: Generate multiple code snippets and multiple text snippets; Formatting the generated multiple code snippets and the multiple text snippets into multiple code blocks and multiple text blocks, wherein each of the multiple code blocks includes one or more code lines, and each of the multiple text blocks includes one or more text lines; The multiple code blocks and the multiple text blocks are randomly mixed, so that the code blocks and the text blocks exist alternately in each page of the PDF training sample.

4. The method according to claim 3, characterized in that The converting each page of the multiple pages of PDF training samples into a second format training sample comprises: Assigning a label to each of the plurality of lines of each page of the PDF training sample, the label indicating whether the line is a text line or a code line; Obtaining coordinates of each row in the plurality of rows; Each page of the PDF training sample is converted into a second format training sample including text or code of each line of the plurality of lines, a label of each line, and coordinates of each line.

5. The method according to claim 4, characterized in that The extracting code from the PDF file using the trained code extraction model includes: Enable the trained code extraction model to infer the PDF file to obtain a label for each line in the PDF file; According to the label of each line, the code block and the text block in the PDF file are identified.

6. The method according to claim 5, characterized in that After identifying the code blocks and text blocks in the PDF file, the method further includes: The trained code extraction model is used to infer the code blocks and text blocks in the PDF file to obtain the semantic relationship between the code blocks and the text blocks; The code blocks and text blocks are rearranged according to semantic relationships to obtain semantically continuous code fragments and text fragments.

7. The method according to any one of claims 1 to 6, characterized in that The code extraction model is a large language model based on the Transformer structure.

8. A device for extracting code from a portable document format PDF file, characterized in that: The device comprises: A sample generation module, used to generate a multi-page PDF training sample using multiple code snippets and multiple text snippets; A format conversion module, used for converting each page of the multiple pages of PDF training samples into a second format training sample; A model training module, used for training a code extraction model using the second format training samples; The code extraction module is used to extract codes from the PDF file using the trained code extraction model.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for extracting codes from a portable document format PDF file according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method for extracting codes from a portable document format PDF file according to any one of claims 1 to 7.