Text processing method and apparatus using same
By grouping and compressing long texts, the problem of input length limitation of large language models in long text scenarios is solved, achieving more efficient text processing and better semantic retention.
Patent Information
- Application Number
- PCT/CN2024/137366
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-19
AI Technical Summary
The existing large language model is not effective in long text scenarios, mainly due to the limitation of input text length and position encoding that does not adapt to long text, resulting in reduced processing effects and increased inference costs.
By grouping text sequences into text groups with high semantic similarity, the subtext sequences in each group are reduced relative to the original text sequence length, and each text group is compressed after grouping to increase the effective input length of the language model.
While increasing the effective input length of the language model, the processing accuracy is maintained, the loss of semantic information during compression is reduced, and the compression effect is improved.
Smart Images

Figure CN2024137366_19062025_PF_FP_ABST
Abstract
Description
Text processing method and device thereof
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 15, 2023, with application number 202311734367.1 and application name “A text processing method and device thereof”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence, and in particular to a text processing method and device thereof. Background Art
[0003] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0004] Large language models are neural network models with numerous parameters, trained on large amounts of text data. They can accurately understand the meaning of text or generate specific natural language text, effectively handling a variety of natural language processing tasks such as translation, conversation, and code generation. Therefore, they have great commercial value. However, current large language models often have limitations on the length of input text, making them ineffective for long text scenarios such as books, academic papers, and legal contracts.
[0005] There are two main reasons for the limited length of input text for language models. First, language models are often trained on short texts. The positional encoding of large language models is not directly suitable for long texts, resulting in a significant decrease in the effectiveness of large language models on long texts. Second, the inference cost of language models increases quadratically with the length of the input text. Longer texts place a greater load on the inference engine and platform.
[0006] Therefore, there is an urgent need for a method that can increase the effective input length of the language model. Summary of the Invention
[0007] The present application provides a text processing method that groups multiple sub-text sequences in a text sequence and performs subsequent language model processing at the text group granularity. Due to the semantic proximity between the sub-text sequences within the text group, the processing accuracy of the language model can be guaranteed while increasing the effective input length of the language model.
[0008] In a first aspect, the present application provides a text processing method, comprising: obtaining a text sequence; the text sequence includes multiple sub-text sequences; grouping the multiple sub-text sequences to obtain multiple text groups; the semantic similarity between the sub-text sequences included in the same text group is greater than a first threshold; and the multiple text groups are used to obtain model output through a language model.
[0009] The present application groups multiple sub-text sequences in a text sequence, and the sub-text sequences included in each group are reduced relative to the length of the text sequence. Moreover, due to the semantic proximity between the sub-text sequences within the text group, that is, the sub-text sequences within the text group are continuous (a text group composed of semantically similar sub-text sequences can express a common or similar theme. In this case, the continuity and coherence of the sub-text sequences within the text group are higher). The continuous sub-text sequences can prevent the grouped text groups from losing too much semantic information. When the subsequent language model processing is performed with the text group as the granularity, the accuracy of the model output can be guaranteed. Furthermore, while the effective input length of the language model is increased, the processing accuracy of the language model is guaranteed.
[0010] Among them, text sequences are one of the most commonly used serialized data types. A text sequence can also be referred to as text data. Text data can be considered a sequence of characters or words. In other words, a text sequence can be data obtained by arranging sub-texts or words in a certain order. Each sub-text sequence can be a continuous partial sequence of a text sequence.
[0011] After obtaining multiple text groups, each text group may be subjected to other processing such as compression and sequence expansion.
[0012] In one possible implementation, the method further includes: compressing each of the multiple text groups to obtain multiple processing results; the multiple text groups are used to obtain model outputs through a language model, including: the multiple processing results are used to obtain model outputs through a language model.
[0013] This application compresses a long text sequence (when a text sequence exceeds the maximum input length supported by the language model, the text sequence can be called a long text) input to a language model into a shorter natural language text, which is equivalent to extending the input length of the large language model. In addition, in the embodiment of the present application, the sub-text sequences within each text group obtained after grouping describe content with similar semantics, which can maintain the coherence of the content of each sub-text sequence (the semantics of text content with higher continuity and coherence are similar), thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0014] In a possible implementation, the semantic similarity between subtext sequences included in different text groups is less than the first threshold.
[0015] In order to ensure a high degree of continuity between subtext sequences within the same text group (that is, their positions in the original text sequence are continuous or very close, so that the compression process can recognize richer continuous semantic information, thereby improving compression quality), in addition to considering semantic similarity, the positional similarity between subtext sequences in the text sequence can also be considered when grouping subtext sequences. That is, through grouping, the positional similarity between subtext sequences included in the same text group in the text sequence is greater than a second threshold. In addition, through grouping, the positional similarity between subtext sequences included in the same text group in the text sequence can also be less than the second threshold.
[0016] In a possible implementation, at least two subtext sequences among the multiple subtext sequences overlap. The overlap can improve the continuity between the subtext sequences.
[0017] In a possible implementation, the method further includes: splicing the multiple processing results according to the positional relationship between the multiple text groups in the text sequence to obtain a spliced processing result.
[0018] In a possible implementation, the method further includes: obtaining a processing result of the text sequence through a language model according to the multiple text groups.
[0019] In one possible implementation, compressing each of the multiple text groups to obtain multiple processing results includes: compressing each of the multiple text groups through the language model or other machine learning models other than the language model to obtain multiple processing results.
[0020] Machine learning used for compression also has limitations on the length of input text. However, through the method of this application, text sequences are grouped, and the long text input to the compression model is segmented and grouped to obtain shorter text groups during compression. This is equivalent to extending the input length of the compression model, and the sub-text sequences included in each text group describe content with similar semantics, maintaining the coherence of the content of each sub-text sequence, thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0021] In one possible implementation, the model output is a result of performing one of the following tasks on the text sequence: a summary extraction task, a text generation task, a dialogue task, a question-answering task, a text translation task, or a knowledge retrieval task.
[0022] In a second aspect, the present application provides a text processing device, comprising:
[0023] An acquisition module, configured to acquire a text sequence; the text sequence includes a plurality of sub-text sequences;
[0024] A grouping module is used to group the multiple sub-text sequences to obtain multiple text groups; the semantic similarity between the sub-text sequences included in the same text group is greater than a first threshold; the multiple text groups are used to obtain model outputs through a language model.
[0025] In a possible implementation, the apparatus further includes:
[0026] a compression module, configured to compress each of the plurality of text groups to obtain a plurality of processing results;
[0027] The multiple text groups are used to obtain a model output through a language model, including:
[0028] The multiple processing results are used to obtain a model output through a language model.
[0029] In a possible implementation, the semantic similarity between subtext sequences included in different text groups is less than the first threshold.
[0030] In a possible implementation, the position similarity between subtext sequences included in the same text group in the text sequence is greater than a second threshold.
[0031] In a possible implementation, the position similarity between subtext sequences included in the same text group in the text sequence is less than the second threshold.
[0032] In a possible implementation, at least two subtext sequences among the multiple subtext sequences overlap.
[0033] In a possible implementation, the compression module is further configured to:
[0034] The multiple processing results are spliced according to the positional relationship between the multiple text groups in the text sequence to obtain a spliced processing result.
[0035] In a possible implementation, the apparatus further includes:
[0036] The task processing module is used to obtain a processing result of the text sequence based on the multiple text groups through a language model.
[0037] In a possible implementation, the compression module is specifically configured to:
[0038] Each of the multiple text groups is compressed using the language model or other machine learning models other than the language model to obtain multiple processing results.
[0039] In one possible implementation, the model output is a result of performing one of the following tasks on the text sequence:
[0040] Summary extraction tasks, text generation tasks, dialogue tasks, question-answering tasks, text translation tasks, and knowledge retrieval tasks.
[0041] In a third aspect, an embodiment of the present application provides a text processing device, which may include a memory, a processor, and a bus system, wherein the memory is used to store programs, and the processor is used to execute the programs in the memory to perform the first aspect and any optional method thereof.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned first aspect and any optional method thereof.
[0043] In a fifth aspect, an embodiment of the present application provides a computer program, which, when executed on a computer, enables the computer to execute the above-mentioned first aspect and any optional method thereof.
[0044] In a sixth aspect, the present application provides a chip system comprising a processor configured to support execution of a text processing device to implement the functions described in the aforementioned aspects, such as transmitting or processing data or information described in the aforementioned methods. In one possible design, the chip system further comprises a memory configured to store program instructions and data necessary for executing the device or training the device. The chip system may consist solely of a chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] FIG1A is a schematic diagram of a structure of an artificial intelligence main framework;
[0046] Figures 1B to 1C are schematic diagrams of the application system framework of the present application;
[0047] FIG1D is a schematic diagram of an optional hardware structure of a terminal;
[0048] FIG2 is a schematic diagram of the structure of a server;
[0049] FIG3 is a schematic diagram of a system architecture of the present application;
[0050] Figure 4 shows a process of cloud services;
[0051] FIG5 is a flowchart of a text processing method provided in an embodiment of the present application;
[0052] FIG6 is a flowchart illustrating text sequence compression;
[0053] FIG7 is a flowchart of a text processing method provided in an embodiment of the present application;
[0054] FIG8 is a schematic diagram of the beneficial effects of the present application;
[0055] FIG9 is a schematic structural diagram of a text processing device provided in an embodiment of the present application;
[0056] FIG10 is a schematic diagram of a structure of an execution device provided in an embodiment of the present application;
[0057] FIG11 is a schematic diagram of a structure of a training device provided in an embodiment of the present application;
[0058] FIG12 is a schematic structural diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0060] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0061] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0062] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent deviations in measurements or calculations that one of ordinary skill in the art would recognize. Furthermore, the use of "may" when describing embodiments of the present application refers to "one or more possible embodiments." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.
[0063] First, the overall workflow of an AI system will be described. See Figure 1A, which shows a schematic diagram of the main AI framework. This AI framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.
[0064] (1) Infrastructure
[0065] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.
[0066] (2) Data
[0067] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0068] (3) Data processing
[0069] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0070] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0071] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0072] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0073] (4) General ability
[0074] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0075] (5) Smart products and industry applications
[0076] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.
[0077] This application can be applied to the field of text processing in the field of artificial intelligence. Taking text processing as an example, the following will introduce multiple application scenarios that have been implemented in products.
[0078] First, the application scenario of this application is introduced.
[0079] This application can be applied to, but is not limited to, applications with text generation functions (hereinafter referred to as text generation applications) or cloud services provided by cloud-side servers, etc., which are described below:
[0080] 1. Text Generation Applications
[0081] The product form of the embodiment of the present application can be a text generation application. The application with text generation function can be run on a terminal device or a cloud-side server.
[0082] In a possible implementation, a text generation application can implement a task of generating text based on an input text generation request (including a text sequence to be processed) to obtain a processing result.
[0083] In particular, the text sequence to be processed is a very long text sequence.
[0084] Among them, text generation requests can indicate generative tasks such as summarization, dialogue, question answering, and translation.
[0085] For example, a text generation request may be: Please summarize the following text into an abstract of no more than 100 words.
[0086] For example, a text generation request may be: Use the four characters "Pangu Zhizi" to start each sentence, and write a seven-character quatrain with hidden acrostic poetry, rhyming with "a".
[0087] In one possible implementation, a user can open a text generation application installed on a terminal device and input a text processing instruction containing a text sequence to be processed. The text generation application can process the text processing instruction containing the text sequence to be processed through the method provided in an embodiment of the present application and present the processing result to the user (the presentation method may be but is not limited to display, saving, uploading to the cloud side, etc.).
[0088] In one possible implementation, a user can open a text generation application installed on a terminal device and input a text processing instruction containing a text sequence to be processed. The text generation application can send the text processing instruction containing the text sequence to be processed to a server on the cloud side. The server on the cloud side processes the text processing instruction containing the text sequence to be processed through the method provided in an embodiment of the present application, and transmits the processing result back to the terminal device. The terminal device can present the processing result to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).
[0089] Optionally, the input and output may occur on a payment dialogue interface, a terminal application service that calls a related interface, or a built-in function of the terminal device.
[0090] Next, the text generation application in the embodiment of this application is introduced from the perspective of functional architecture and product architecture that implements the functions.
[0091] Referring to FIG. 1B , FIG. 1B is a schematic diagram of the functional architecture of a text generation application in an embodiment of the present application:
[0092] In one possible implementation, as shown in FIG1B , a text generation application 102 may receive input parameters 101 (e.g., including text processing instructions including a text sequence to be processed) and generate processing results 103. The text generation application 102 may be executed on, for example, at least one computer system and include computer code that, when executed by one or more computers, causes the computers to execute the method provided in the embodiments of the present application.
[0093] Referring to FIG. 1C , FIG. 1C is a schematic diagram of the entity architecture for running a text generation application in an embodiment of the present application:
[0094] Referring to FIG1C , FIG1C shows a schematic diagram of a system architecture. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (FIG1C illustrates one server as an example), and the server 200 may provide the method provided in the embodiments of the present application to one or more terminals.
[0095] Among them, a text generation application can be installed on the terminal 100. The above application and web page can provide an interface. The terminal 100 can receive relevant parameters entered by the user on the text generation interface and send the above parameters to the server 200. The server 200 can obtain processing results based on the received parameters and return the processing results to the terminal 100.
[0096] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the need for the cooperation of the server, and the embodiments of the present application are not limited to this.
[0097] Next, the product form of the terminal 100 in FIG1C is described;
[0098] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0099] FIG1D shows a schematic diagram of an optional hardware structure of the terminal 100 .
[0100] 1D , the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, and a power supply 190. Those skilled in the art will appreciate that FIG1D is merely an example of a terminal or multi-function device and does not limit the terminal or multi-function device. The terminal or multi-function device may include more or fewer components than shown, or may combine certain components or have different components.
[0101] The input unit 130 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the portable multifunction device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can detect user touch operations on or near it (for example, operations performed on or near the touch screen using a finger, joint, stylus, or any other suitable object) and drive corresponding connected devices according to pre-set programs. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and transmit them to the processor 170. It can also receive and execute commands sent by the processor 170; the touch signals include at least touch point coordinate information. The touch screen 131 provides an input interface and an output interface between the terminal 100 and the user. Touch screens can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.
[0102] The input device 132 may receive input text processing instructions including a text sequence to be processed, and the like.
[0103] The display unit 140 may be used to display information input by the user or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In an embodiment of the present application, the display unit 140 may be used to display the interface of a text generation application, processing results, etc.
[0104] Memory 120 can be used to store instructions and data. It primarily includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files and text. The instruction storage area can store software units such as the operating system, applications, and instructions required for at least one function, or subsets or extensions thereof. It may also include non-volatile random access memory (RAM). It provides processor 170 with management functions for the hardware, software, and data resources within the computing and processing device, supporting control software and applications. It is also used to store multimedia files and running programs and applications.
[0105] The processor 170 is the control center of the terminal 100. It connects all components of the terminal 100 using various interfaces and circuits. By executing instructions stored in the memory 120 and accessing data stored therein, it executes various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units. Preferably, the processor 170 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory may be implemented on a single chip; in other embodiments, they may be implemented on separate chips. The processor 170 may also generate corresponding operational control signals and send them to the corresponding components of the computing and processing device. It may also read and process data in the software, particularly the data and programs in the memory 120, to enable the various functional modules therein to perform their corresponding functions, thereby controlling the corresponding components to operate as instructed.
[0106] Among them, the memory 120 can be used to store software codes related to the text processing method, the processor 170 can execute the steps of the text processing method of the chip, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.
[0107] The RF unit 110 (optional) can be used to send and receive information or receive and send signals during a call. For example, after receiving downlink information from the base station, it is passed to the processor 170 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices via wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0108] In this embodiment of the present application, the radio frequency unit 110 may send a text processing instruction including a text sequence to be processed to the server 200 , and receive a processing result sent by the server 200 .
[0109] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.
[0110] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.
[0111] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100 .
[0112] Although not shown, the terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied to the terminal 100 shown in FIG1D .
[0113] Next, the product form of the server 200 in FIG1C is described;
[0114] FIG2 provides a schematic diagram of the structure of a server 200. As shown in FIG2, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.
[0115] Bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, FIG2 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.
[0116] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0117] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard drive (HDD), or solid state drive (SSD).
[0118] The memory 204 may be used to store software codes related to the text processing method, and the processor 202 may execute the steps of the text processing method of the chip, and may also schedule other units to implement corresponding functions.
[0119] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0120] It should be understood that the steps related to the model reasoning process in the embodiments of this application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and server is not limited to the processor-memory architecture described above. The system architecture provided in the embodiments of this application is described in detail below with reference to Figure 3.
[0121] FIG3 is a schematic diagram of the system architecture provided by an embodiment of the present application. As shown in FIG3 , the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data acquisition system 560 .
[0122] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.
[0123] The execution device 510 may be a terminal device or a server that runs the above-mentioned text generation application.
[0124] The data collection device 560 is used to collect training samples. After collecting the training samples, the data collection device 560 stores these training samples in the database 530.
[0125] The training device 520 can train the neural network to be trained (such as the language model in the embodiment of the present application) based on the training samples maintained in the database 530 to obtain the target model / rule 501.
[0126] It should be understood that the training device 520 can perform a pre-training process on the neural network to be trained based on the training samples maintained in the database 530, or fine-tune the model based on the pre-training.
[0127] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.
[0128] The target model / rule 501 obtained through training with the training device 520 can be applied to different systems or devices, such as the execution device 510 shown in FIG3 . The execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, an in-vehicle terminal, etc., or a server, etc.
[0129] Specifically, the training device 520 may transfer the trained model to the execution device 510 .
[0130] In Figure 3, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices. The user can input data (such as text processing instructions containing a text sequence to be processed, etc.) into the I / O interface 512 through the client device 540.
[0131] Preprocessing module 513 and preprocessing module 514 are used to preprocess the input data received by I / O interface 512. It should be understood that preprocessing module 513 and preprocessing module 514 may be absent or only one preprocessing module may be present. If preprocessing module 513 and preprocessing module 514 are absent, computing module 511 may be used directly to process the input data.
[0132] When the execution device 510 preprocesses the input data, or when the computing module 511 of the execution device 510 performs calculations and other related processing, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.
[0133] Finally, the I / O interface 512 provides the processed results to the client device 540 and thus to the user.
[0134] In the scenario shown in FIG3 , the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another scenario, client device 540 can automatically send input data to I / O interface 512. If user authorization is required for client device 540 to automatically send input data, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, and the specific presentation form can be a display, sound, action, or other specific method. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results output from I / O interface 512 as new sample data, and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 directly stores the input data input into I / O interface 512 and output results output from I / O interface 512 as new sample data in database 530.
[0135] It is worth noting that FIG3 is merely a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in FIG3 , the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.
[0136] From the inference side of the model:
[0137] In the embodiment of the present application, the computing module 511 of the above-mentioned execution device 510 can obtain the code stored in the data storage system 550 to implement the steps related to the model reasoning process in the embodiment of the present application.
[0138] In an embodiment of the present application, the computing module 511 of the execution device 510 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0139] Specifically, the computing module 511 of the execution device 510 can be a hardware system with an execution instruction function, and the steps related to the model reasoning process provided in the embodiment of the present application can be software codes stored in the memory. The computing module 511 of the execution device 510 can obtain the software code from the memory and execute the obtained software code to implement the steps related to the model reasoning process provided in the embodiment of the present application.
[0140] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to the model reasoning process provided in the embodiment of the present application can also be implemented by the hardware system that does not have the function of executing instructions in the computing module 511 of the execution device 510, which is not limited here.
[0141] From the training side of the model:
[0142] In an embodiment of the present application, the above-mentioned training device 520 can obtain the code stored in the memory (not shown in Figure 3, which can be integrated into the training device 520 or deployed separately from the training device 520) to implement the steps related to model training in the embodiment of the present application.
[0143] In an embodiment of the present application, the training device 520 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0144] It should be understood that the training device 520 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to model training provided in the embodiments of the present application can also be implemented by the hardware system in the training device 520 that does not have the function of executing instructions, which is not limited here.
[0145] 2. Text generation cloud services provided by the server:
[0146] In a possible implementation, the server may provide a text generation service to the terminal side through an application programming interface (API).
[0147] Among them, the terminal device can send relevant parameters (such as text processing instructions containing the text sequence to be processed) to the server through the API provided by the cloud. The server can obtain processing results based on the received parameters, etc., and return the processing results to the terminal.
[0148] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.
[0149] FIG4 shows a process of using a text generation cloud service provided by a cloud platform.
[0150] 1. Activate and purchase the text generation service.
[0151] 2. Users can download the software development kit (SDK) for the text generation service. Cloud platforms usually provide multiple development versions of the SDK for users to choose based on their development environment requirements, such as JAVA version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.
[0152] 3. After the user downloads the corresponding version of the SDK to the local computer as needed, import the SDK project into the local development environment, configure and debug it in the local development environment. The local development environment can also be used to develop other functions, forming an application that integrates text generation capabilities.
[0153] 4. When a text generation application needs to generate text during use, it can trigger a text generation API call. When the application triggers the text generation function, it initiates an API request to the running instance of the text generation service in the cloud environment. The API request carries a text processing instruction containing the text sequence to be processed. The running instance in the cloud environment processes the text processing instruction containing the text sequence to be processed and obtains the processing result.
[0154] 5. The cloud environment returns the processing results to the application, thereby completing a method call provided in an embodiment of the present application.
[0155] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.
[0156] (1) Neural Network
[0157] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:
[0158] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0159] (2) Deep Neural Networks
[0160] Deep Neural Network (DNN), also known as multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. Based on the position of different layers in DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscript is the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).
[0161] (3) Backpropagation algorithm
[0162] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.
[0163] (4) Loss function
[0164] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.
[0165] (6) Embedding: A low-dimensional embedding representation of high-dimensional words or sentences, so that words or sentences with similar embedding vectors have similar meanings.
[0166] (7) Token: The element generated by word segmentation of text.
[0167] (8) Tokenizer: A tokenizer that is used to tokenize text based on a vocabulary and generate metrics.
[0168] (9) Foundation model: A foundation model usually refers to a large language model that has been pre-trained on a large scale using data from different fields.
[0169] (10) Pre-training: Pre-training is to use self-supervised methods to learn from a large amount of (text) input information without labeling. The models used are usually divided into autoregressive models and autoencoding models.
[0170] Large language models are neural network models with numerous parameters, trained on large amounts of text data. They can accurately understand the meaning of text or generate specific natural language text, effectively handling a variety of natural language processing tasks such as translation, conversation, and code generation. Therefore, they have great commercial value. However, current large language models often have limitations on the length of input text, making them ineffective for long text scenarios such as books, academic papers, and legal contracts.
[0171] There are two main reasons for the limited length of input text for language models. First, language models are often trained on short texts. The positional encoding of large language models is not directly suitable for long texts, resulting in a significant decrease in the effectiveness of large language models on long texts. Second, the inference cost of language models increases quadratically with the length of the input text. Longer texts place a greater load on the inference engine and platform.
[0172] Therefore, there is an urgent need for a method that can increase the effective input length of the language model.
[0173] In order to solve the above problems, the present invention provides a text processing method. The text processing method of the present invention is described in detail below with reference to the accompanying drawings.
[0174] Referring to Figure 5, Figure 5 is a flow chart of a text processing method provided in an embodiment of the present application. As shown in Figure 5, an embodiment of the present application provides a text processing method, which can be but is not limited to a feedforward process for model pre-training, a feedforward process during model fine-tuning, a model inference process, etc. Specifically, the method can include steps 501 to 502, and these steps are described in detail below.
[0175] 501. Acquire a text sequence; the text sequence includes multiple sub-text sequences.
[0176] In the feedforward process of model training (such as pre-training or fine-tuning), text sequences can be training samples that the language model needs to process.
[0177] In the feedforward process of model training, text sequences can be training samples that the language model needs to process.
[0178] During the model inference process, the text sequence may be a text sequence to be processed by the language model. For example, the text sequence may be carried in a text processing request.
[0179] In a possible implementation, the length of the text sequence is long and exceeds the processing length supported by the input of the language model. Therefore, the text sequence needs to be compressed.
[0180] In an embodiment of the present application, a text sequence can be divided (or, referred to as segmented) to obtain multiple sub-text sequences. For example, the text sequence can be a text paragraph including multiple sentences. The text sequence can be segmented at the sentence level, that is, each sub-text sequence can include one or more sentences. For example, the number of characters in each sub-text sequence obtained after segmentation can be controlled to be less than or equal to a preset value.
[0181] The different sub-text sequences after division may overlap or be completely staggered, and the overlap can improve the continuity of the segmented content. For example, at least two of the multiple sub-text sequences obtained after dividing the text sequence may overlap.
[0182] In a possible implementation, the text sequence includes m sentences, and the subtext sequence includes n sentences, where n<m.
[0183] For example, the text sequence may include 10 sentences:
[0184] Sentence 1, sentence 2, sentence 3, sentence 4, sentence 5, sentence 6, sentence 7, sentence 8, sentence 9, sentence 10.
[0185] When the subtext sequences are completely staggered, the multiple subtext sequences obtained after division may include the following subtext sequences:
[0186] Subtext sequence 1: sentence 1, sentence 2, sentence 3;
[0187] Subtext sequence 2: sentence 4, sentence 5, sentence 6, sentence 7;
[0188] Subtext sequence 3: sentence 8, sentence 9.
[0189] When there is overlap between subtext sequences, the multiple subtext sequences obtained after division may include the following subtext sequences:
[0190] Subtext sequence 1: sentence 1, sentence 2, sentence 3, sentence 4;
[0191] Subtext sequence 2: sentence 4, sentence 5, sentence 6, sentence 7;
[0192] Subtext sequence 3: sentence 8, sentence 9.
[0193] Among them, there is an overlapping part (sentence 4) between subtext sequence 1 and subtext sequence 2.
[0194] 502. Group the multiple subtext sequences to obtain multiple text groups; the semantic similarity between the subtext sequences included in the same text group is greater than a first threshold; and the multiple text groups are used to obtain model outputs through a language model.
[0195] The embodiments of this application hope to expand the length limit of the input text that the language model can process and increase the effective input length of the language model. There are often two main reasons for the length limit. First, the language model is often trained on short texts, and the position encoding of the large language model is not directly adapted to the situation of long texts, resulting in a significant decrease in the effectiveness of the large language model on long texts. Second, the inference cost of the language model increases quadratically with the length of the input text. The longer the text, the higher the load on the inference engine and platform.
[0196] Therefore, it is necessary to perform semantic compression on the input text. However, on the one hand, model compression also has the problem of input text length limitation. Therefore, before performing semantic compression, the text sequence can be divided, and each sub-text sequence obtained after the division can be compressed separately, and then the processing results can be spliced, which is equivalent to a divide-and-conquer and then merge compression method.
[0197] However, the division of text sequences will destroy the coherence in the text, resulting in the loss of a lot of text semantic information when compressing each sub-text sequence. Therefore, the idea of this application is to design a grouping method for sub-text sequences, so that the sub-text sequences included in each text group can describe content with similar semantics, maintain the coherence of the content of each sub-text sequence, thereby reducing the loss of text semantic information during compression, improving the compression effect, and then improving the processing effect of the language model.
[0198] Natural language text sequences often have a hierarchical structure. For example, books often have chapters, each covering a complete content. Together, the chapters form a complete book. While not all texts have a chapter structure, it can be implicitly modeled using topic paragraphs. Each topic paragraph describes similar content, but differs from the preceding and following content. By dividing long texts into topic paragraphs, the content of each topic can be independently compressed while maintaining coherence. By then summarizing the compressed content in sequence, more effective semantic processing results for long texts can be achieved.
[0199] Each topic paragraph obtained by the above division can be a text group in the embodiment of the present application. In other words, the sub-text sequences in a text group can express the same topic.
[0200] In one possible implementation, the semantic similarity between subtext sequences included in the same text group is greater than a first threshold, while the semantic similarity between subtext sequences included in different text groups is less than the first threshold. In other words, each text group obtained after the division may include semantically similar subtext sequences, while the semantic differences between subtext sequences included in different text groups are relatively large.
[0201] Next, we introduce an example of how to group multiple subtext sequences:
[0202] In one possible implementation, a pre-trained embedding model can be used to embed each subtext sequence to obtain an embedding vector. A weighted graph can then be constructed based on the embedding vectors of multiple subtext sequences. The similarity between embedding vectors is the strength of the connection between segments in the graph (this connection strength is represented by a weight value, so the graph structure can be called a weighted graph). Based on this, a weighted graph modeling of subtext sequences can be obtained. After obtaining the weighted graph of a long text, continuous community structures can be found in the graph, which is equivalent to finding the divisions of the main paragraphs. Each community structure corresponds to a text group. The community structure has strong internal connections but weak external connections, which is consistent with the concept of implicitly modeling the main paragraphs. That is, each community structure can include multiple graph nodes, each corresponding to a subtext sequence. The similarity between subtext sequences within each community structure is high, while the similarity between subtext sequences in different community structures is low.
[0203] In addition, in order to ensure that the subtext sequences within the same text group have a high degree of continuity (that is, their positions in the original text sequence are continuous or very close, so that the compression process can recognize richer continuous semantic information, thereby improving compression quality), in addition to considering semantic similarity, the positional similarity between the subtext sequences in the text sequence can also be considered when grouping the subtext sequences. That is, through grouping, the positional similarity between the subtext sequences included in the same text group in the text sequence is greater than a second threshold. In addition, through grouping, the positional similarity between the subtext sequences included in the same text group in the text sequence can also be less than the second threshold.
[0204] Next, we introduce an example of how to group multiple subtext sequences:
[0205] Taking the above weighted graph implementation as an example, in order to obtain a continuous community structure, the connections between nodes that are far apart can be masked. For example, the weights of the connections between nodes that are far apart in the graph can be set to 0. This ensures that the nodes contained in the main paragraph are also segmented in a sequential order.
[0206] The multiple text groups are used to obtain model output through a language model.
[0207] After obtaining multiple text groups, each text group may be subjected to other processing such as compression and sequence expansion.
[0208] Taking compression as an example, an embodiment of the present application can compress each text group in the multiple text groups to obtain multiple processing results, and then the multiple text groups can be used to obtain model outputs through a language model, including: the multiple processing results are used to obtain model outputs through a language model, and the multiple processing results are used to obtain processing results of the text sequence through a language model.
[0209] In one possible implementation, the compression is semantic compression for reducing text length. Using semantic compression, long text input to the language model is compressed into shorter natural language text while retaining important semantic information, which is equivalent to extending the input length of the large language model. Furthermore, in the embodiments of the present application, the subtext sequences within each text group obtained after grouping describe content with similar semantics, which can maintain the coherence of the content of each subtext sequence, thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0210] For example, multiple processing results can be obtained based on the multiple text groups through the language model or other machine learning models other than the language model. Machine learning used for compression also has a limit on the length of the input text. However, through the method of this application, the text sequences are grouped, and the long texts input to the compression model are divided into shorter text groups during compression. This is equivalent to extending the input length of the compression model, and the sub-text sequences included in each text group describe content with similar semantics, maintaining the coherence of the content of each sub-text sequence, thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0211] The plurality of processing results are spliced according to the positional relationship between the plurality of text groups in the text sequence to obtain a spliced processing result. After the compressed contents are aggregated in order, a semantic processing result of the long text with better effect can be obtained.
[0212] In a possible implementation, the processing result of the text sequence can be obtained based on the multiple processing results through a language model.
[0213] Optionally, the plurality of processing results may be expanded by a sequence expansion method, and then the processing result of the text sequence may be obtained by using a language model based on the plurality of expanded processing results.
[0214] Among them, the language model can be a large language model (LLM), and the language model can perform text processing tasks to obtain processing results. For example, the text processing tasks can be but are not limited to summary generation, question and answer response, text translation, etc.
[0215] Next, a process of an embodiment of the present application is introduced with reference to a specific example.
[0216] Referring to Figure 6, the long text is first divided into sentence-level segments. The number of characters in each segment can be controlled, and characters between segments can be overlapped to ensure continuity of content. Once the sequential segments are obtained, a pre-trained sentence embedding model can be used to embed the segments into vectors. The similarity between vectors represents the strength of the connections between segments in the graph, resulting in weighted graph modeling of the text. After obtaining the graph structure of the long text, finding the continuous community structure within the graph is equivalent to finding the topic paragraphs. The community structure has strong internal connections but weak external connections, which is consistent with the concept of implicitly modeling topic paragraphs. To obtain a continuous community structure, the connections between distant segments are masked, i.e., their graph values are set to 0. Using the graph segmentation algorithm, the resulting segments are the desired topic paragraphs, and the topic paragraphs contain adjacent, sequentially arranged segments. By compressing the model and then sequentially aggregating them, the semantic processing results of the long text are obtained. Although the length is reduced, the key semantic information is preserved.
[0217] The embodiment of the present application can be applied to scenarios where a large language model is used to process text that exceeds its maximum text length, and can also help reduce the memory limitations encountered by large model reasoning. Figure 7 is a schematic diagram of an application architecture of an embodiment of the present application.
[0218] Among them, the word segmenter can be used to specify the action of dividing the text sequence into multiple sub-text sequences, the body division can perform the action of grouping multiple sub-text sequences, the block summary can perform the action of compressing each text group, and the text reorganization can perform the action of splicing multiple processing results.
[0219] The embodiments of this application can effectively expand the input length of large language models for natural language tasks. The following are the gain results compared with other expansion methods from the dataset.
[0220] Table 1
[0221] Long text summarization tasks can be accomplished by directly inputting length-based chunk summaries into a large language model and then summarizing them. This technical solution is a natural extension of this solution from a semantic perspective. As shown in Table 1, the solution provided by the embodiment of this application achieves significant gains compared to existing solutions.
[0222] Question-answering tasks are more complex than summary tasks and cannot be handled by simply splitting and then merging. All information must be processed at once. The key retrieval task is a question-answering task that can be generated to any length and is a commonly used dataset for testing input length expansion schemes for large models. Referring to Table 2, the results of llama2 in Table 2 show that llama2 has an input limit of 4,000 words, and the model's performance drops sharply after exceeding the limit. However, when this scheme is applied to llama2, the input length can be expanded to 30,000 words with an accuracy rate of over 90%. This scheme can also be easily stacked with other expansion schemes to further increase the input text length limit.
[0223] Table 2
[0224] LongBench is a comprehensive dataset for long-text tasks, encompassing four types of tasks: single-text question answering, multi-document question answering, summarization, and few-shot learning. We evaluated each type on three English datasets. The results in Figure 8 show that the method presented in this embodiment outperforms the best-performing extension method (YARN) in academia on most tasks, and also outperforms the original model on most tasks.
[0225] 9 , which is a schematic diagram of the structure of a text processing device provided in an embodiment of the present application. As shown in FIG9 , a text processing device 900 provided in an embodiment of the present application includes:
[0226] An acquisition module 901 is configured to acquire a text sequence, wherein the text sequence includes a plurality of subtext sequences;
[0227] The specific description of the acquisition module 901 can refer to the description of step 501 in the above embodiment, which will not be repeated here.
[0228] Grouping module 902 is configured to group the multiple subtext sequences into multiple text groups; the semantic similarity between the subtext sequences within a text group is greater than a first threshold; and the multiple text groups are used to generate model outputs using a language model. A detailed description of grouping module 902 can be found in the description of step 502 in the above embodiment and will not be repeated here.
[0229] In a possible implementation, the apparatus further includes:
[0230] A compression module 903, configured to compress each of the plurality of text groups to obtain a plurality of processing results;
[0231] The multiple text groups are used to obtain a model output through a language model, including:
[0232] The multiple processing results are used to obtain a model output through a language model.
[0233] In a possible implementation, the semantic similarity between subtext sequences included in different text groups is less than the first threshold.
[0234] In a possible implementation, the position similarity between subtext sequences included in the same text group in the text sequence is greater than a second threshold.
[0235] In a possible implementation, the position similarity between subtext sequences included in the same text group in the text sequence is less than the second threshold.
[0236] In a possible implementation, at least two subtext sequences among the multiple subtext sequences overlap.
[0237] In a possible implementation, the compression module 903 is further configured to:
[0238] The multiple processing results are spliced according to the positional relationship between the multiple text groups in the text sequence to obtain a spliced processing result.
[0239] In a possible implementation, the apparatus further includes:
[0240] The task processing module is used to obtain a processing result of the text sequence based on the multiple text groups through a language model.
[0241] In a possible implementation, the compression module 903 is specifically configured to:
[0242] Each of the multiple text groups is compressed using the language model or other machine learning models other than the language model to obtain multiple processing results.
[0243] In one possible implementation, the model output is a result of performing one of the following tasks on the text sequence:
[0244] Summary extraction tasks, text generation tasks, dialogue tasks, question-answering tasks, text translation tasks, and knowledge retrieval tasks.
[0245] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 10. Figure 10 is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1000 can be specifically manifested as a virtual reality VR device, a mobile phone, a tablet, a laptop computer, a smart wearable device, a monitoring data processing device or a server, etc., which is not limited here. Specifically, the execution device 1000 includes: a receiver 1001, a transmitter 1002, a processor 1003 and a memory 1004 (wherein the number of processors 1003 in the execution device 1000 can be one or more, and Figure 10 takes one processor as an example), wherein the processor 1003 may include an application processor 10031 and a communication processor 10032. In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003 and the memory 1004 may be connected via a bus or other means.
[0246] The memory 1004 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1003. A portion of the memory 1004 may also include non-volatile random access memory (NVRAM). The memory 1004 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0247] Processor 1003 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0248] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1003. Processor 1003 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1003. The above processor 1003 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1003 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1004. Processor 1003 reads information from memory 1004 and, in conjunction with its hardware, completes the steps involved in the model inference process in the above method.
[0249] Receiver 1001 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1002 can be used to output digital or character information through the first interface. Transmitter 1002 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1002 can also include a display device such as a display screen.
[0250] The present application also provides a training device. Please refer to FIG. 11 , which is a schematic diagram of the structure of a training device provided by an embodiment of the present application. Specifically, the training device 1100 is implemented by one or more servers. The training device 1100 may vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1111 (e.g., one or more processors), a memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) storing application programs 1142 or data 1144. The memory 1132 and storage medium 1130 may be either ephemeral or persistent storage. The program stored in the storage medium 1130 may include one or more modules (not shown), each of which may include a series of instruction operations on the training device. Furthermore, the CPU 1111 may be configured to communicate with the storage medium 1130 to execute the series of instruction operations in the storage medium 1130 on the training device 1100.
[0251] The training device 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input and output interfaces 1158; or, one or more operating systems 1141, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0252] In the embodiment of the present application, the central processing unit 1111 is used to execute actions related to model training in the above embodiments.
[0253] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0254] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0255] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the text processing method described in the above embodiment, or so that the chip in the training device executes the text processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0256] Specifically, see Figure 12, which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1200. NPU 1200 is mounted on the host CPU (host CPU) as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1203, which is controlled by controller 1204 to extract matrix data from memory and perform multiplication operations.
[0257] In some implementations, the arithmetic circuit 1203 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1203 is a two-dimensional systolic array. The arithmetic circuit 1203 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1203 is a general-purpose matrix processor.
[0258] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1201 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1208.
[0259] Unified memory 1206 is used to store input and output data. Weight data is directly transferred to weight memory 1202 through the Direct Memory Access Controller (DMAC) 1205. Input data is also transferred to unified memory 1206 through the DMAC.
[0260] BIU stands for Bus Interface Unit, i.e., bus interface unit 1210 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1209 .
[0261] The bus interface unit 1210 (BIU) is used for the instruction fetch memory 1209 to obtain instructions from the external memory, and is also used for the storage unit access controller 1205 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0262] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1206 or transfer weight data to the weight memory 1202 or transfer input data to the input memory 1201.
[0263] The vector calculation unit 1207 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1203, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0264] In some implementations, the vector calculation unit 1207 can store the processed output vector in the unified memory 1206. For example, the vector calculation unit 1207 can apply a linear function or a nonlinear function to the output of the operation circuit 1203, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, the vector calculation unit 1207 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1203, for example, for use in subsequent layers in a neural network.
[0265] An instruction fetch buffer 1209 connected to the controller 1204 is used to store instructions used by the controller 1204;
[0266] Unified memory 1206, input memory 1201, weight memory 1202, and instruction fetch memory 1209 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0267] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0268] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0269] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0270] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0271] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A text processing method, characterized in that: The method comprises: Acquire a text sequence; the text sequence includes a plurality of sub-text sequences; The multiple sub-text sequences are grouped to obtain multiple text groups; the semantic similarity between the sub-text sequences included in the same text group is greater than a first threshold; and the multiple text groups are used to obtain model outputs through a language model.
2. The method according to claim 1, characterized in that The method further comprises: compressing each text group in the plurality of text groups to obtain a plurality of processing results; The multiple text groups are used to obtain model outputs through a language model, including: The multiple processing results are used to obtain a model output through a language model.
3. The method according to claim 1 or 2, characterized in that: The semantic similarity between sub-text sequences included in different text groups is less than the first threshold.
4. The method according to any one of claims 1 to 3, characterized in that: The position similarity between the subtext sequences included in the same text group in the text sequence is greater than a second threshold.
5. The method according to any one of claims 1 to 4, characterized in that: The position similarity between subtext sequences included in the same text group in the text sequence is less than the second threshold.
6. The method according to any one of claims 1 to 5, characterized in that: There is overlap between at least two subtext sequences among the plurality of subtext sequences.
7. The method according to any one of claims 2 to 6, characterized in that: The method further comprises: The multiple processing results are spliced according to the positional relationship between the multiple text groups in the text sequence to obtain a spliced processing result.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: According to the multiple text groups, a processing result of the text sequence is obtained through a language model.
9. The method according to any one of claims 2 to 8, characterized in that: The compressing each text group in the multiple text groups to obtain multiple processing results includes: Each of the multiple text groups is compressed using the language model or other machine learning models except the language model to obtain multiple processing results.
10. The method according to any one of claims 1 to 9, characterized in that: The model output is the result of performing one of the following tasks on the text sequence: Summary extraction task, text generation task, dialogue task, question and answer task, text translation task, knowledge retrieval task.
11. A text processing device, characterized in that: The device comprises: An acquisition module, used for acquiring a text sequence; the text sequence includes a plurality of sub-text sequences; A grouping module is used to group the multiple sub-text sequences to obtain multiple text groups; the semantic similarity between the sub-text sequences included in the same text group is greater than a first threshold; and the multiple text groups are used to obtain model outputs through a language model.
12. The device according to claim 11, characterized in that The device also includes: A compression module, used for compressing each text group in the plurality of text groups to obtain a plurality of processing results; The multiple text groups are used to obtain model outputs through a language model, including: The multiple processing results are used to obtain a model output through a language model.
13. The device according to claim 11 or 12, characterized in that The semantic similarity between sub-text sequences included in different text groups is less than the first threshold.
14. The device according to any one of claims 11 to 13, characterized in that: The position similarity between the subtext sequences included in the same text group in the text sequence is greater than a second threshold.
15. The device according to any one of claims 11 to 14, characterized in that: The position similarity between subtext sequences included in the same text group in the text sequence is less than the second threshold.
16. The device according to any one of claims 11 to 15, characterized in that There is overlap between at least two subtext sequences among the plurality of subtext sequences.
17. The device according to any one of claims 12 to 16, characterized in that The compression module is also used for: The multiple processing results are spliced according to the positional relationship between the multiple text groups in the text sequence to obtain a spliced processing result.
18. The device according to any one of claims 11 to 17, characterized in that The device also includes: The task processing module is used to obtain the processing result of the text sequence according to the multiple text groups through the language model.
19. The device according to any one of claims 11 to 18, characterized in that The model output is the result of performing one of the following tasks on the text sequence: Summary extraction task, text generation task, dialogue task, question and answer task, text translation task, knowledge retrieval task.
20. The device according to any one of claims 12 to 19, characterized in that The compression module is specifically used for: Each of the multiple text groups is compressed using the language model or other machine learning models except the language model to obtain multiple processing results.
21. A computer storage medium, characterized in that The computer storage medium stores one or more instructions which, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 10.
22. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 10.
23. A system comprising at least one processor and at least one memory; the processor and the memory are connected via a communication bus and communicate with each other; The at least one memory is used to store code; The at least one processor is configured to execute the code to perform the method according to any one of claims 1 to 10.
24. A chip, comprising a processor, characterized in that: The processor is used to support the text processing device to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Text processing method and device
CN120163131A
Semantic segmentation model training method, intention understanding method and device
CN115018516A
Abstract generation method and device, intelligent terminal and storage medium
CN116483988A
Intelligent question-answering system and method suitable for vertical field and application of intelligent question-answering system and method
CN116805001A
Abstract generation method and device, equipment, storage medium and product
CN116894089A
Cited By
Method for realizing semantic chain retrieval enhanced generation processing aiming at knowledge questions and answers of long document
CN121029966A